Hybrid memory compression
The hybrid memory arrangement dynamically compresses data in near memory to create a dynamic cache, addressing inefficiencies in existing systems by maximizing memory capacity utilization and reducing metadata overhead, thus improving system performance.
Patent Information
- Application Number
- PCT/SE2025/050037
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-05
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-24
AI Technical Summary
Existing hybrid memory systems with a two-level hierarchy face inefficiencies due to wasted NM capacity and significant metadata overhead, leading to suboptimal utilization of bandwidth and storage resources.
A hybrid memory arrangement that dynamically compresses data in near memory (NM) to create a dynamic cache, exposing the entire NM and far memory (FM) capacity to the system, while minimizing metadata overhead by using ECC bits and storing metadata in FM, allowing fine-grain management of bandwidth-demanding data.
The solution effectively utilizes the entire memory capacity, reducing metadata overhead and avoiding costly swap operations, thereby enhancing system performance and efficiency.
Smart Images

Figure SE2025050037_24072025_PF_FP_ABST
Abstract
Description
HYBRID MEMORY COMPRESSIONTECHNICAL FIELD
[0001] The present disclosure generally relates to the field of data compression in memories in electronic computers. More specifically, the present disclosure relates to a hybrid memory arrangement providing a two-level main memory hierarchy in a computer system, to a computer system that comprises such a hybrid memory arrangement as well as one or more processing units and one or more levels of cache memory, to a method of providing a two- level main memory hierarchy in a computer system, and to a hybrid memory controller for use in a hybrid memory arrangement.BACKGROUND
[0002] Data compression is a general technique to store and transfer data more efficiently by coding frequent collections of data more efficiently than less frequent collections of data. It is of interest to generally store and transfer data more efficiently for several reasons. In computer memories, for example memories that keep data and computer instructions that processing devices operate on, for example main memory or cache memories, it is of interest to store said data more efficiently, say K times, as it then can reduce the size of said memories potentially by K times, using potentially K times less communication capacity to transfer data between one memory to another memory and with potentially K times less energy expenditure to store and transfer said data inside or between computer systems and / or between memories. Alternatively, one can potentially store K times more data in available computer memory than without data compression. This can be of interest to achieve potentially K times higher performance of a computer without having to add more memory, which can be costly or can simply be less desirable due to resource constraints. As another example, the size and weight of a smartphone, a tablet, a lap / desktop or a set-top box are limited as a larger or heavier smartphone, tablet, a lap / desktop or a set-top box could be of less value for an end user; hence potentially lowering the market value of such products. Yet, more memory capacity can potentially increase the market value of the product as more memory capacity can result in higher performance and hence better utility of the product.
[0003] To summarize, in the general landscape of computerized products, including isolated devices or interconnected ones, data compression can potentially increase the performance, lower the energy expenditure, or lower the cost and space consumed by memory. Therefore, data compression has a broad utility in a wide range of computerized products beyond those mentioned here.
[0004] Dynamic random-access memory (DRAM) is plagued by limited bandwidth. To mitigate it, heterogeneous memory systems comprising a two-level main-memory hierarchy, a.k.a. hybrid memory (HM), is an attractive way of addressing this deficiency. The first level of hybrid memory is sometimes referred to as near memory (NM) whereas the second level is sometimes referred to as far memory (FM). In some implementations of hybrid memory, High-Bandwidth Memory (HBM) comprises the first level whereas DRAM comprises the second level. HBM typically offers substantially higher bandwidth than DRAM but the cost of HBM can be substantially higher than DRAM. As a result, a hybrid memory implementation with HBM and DRAM can offer a main memory that has as high a bandwidth as HBM and is as large as well as having nearly as low a cost as DRAM.
[0005] Prior art has investigated two broad approaches to manage hybrid memories: cached and flat hybrid memories. In a cached hybrid memory, NM is used as a cache for FM, managed transparently to the operating system. However, for a hybrid memory using HBM and DRAM as NM and FM, respectively, as we consider in this patent disclosure, the amount of DRAM is often a small factor, say eight, more than the amount of HBM. Hence, by not exposing NM to the (operating) system, a significant amount of memory capacity is wasted.
[0006] On the other hand, in a flat hybrid memory, NM as well as FM contribute to the flat physical address space making the entire hybrid-memory capacity available to the (operating) system. Here, bandwidth-demanding pages mapped in FM (e.g., DRAM) are typically swapped by less bandwidth-demanding pages residing in NM (e.g., HBM). Unfortunately, changing the page mapping entails significant operating system-induced overhead along with the traffic overhead of swapping pages. To reduce the former type of overhead, prior art has proposed remapping mechanisms at the hardware level and considered smaller grain sizes. However, the metadata needed to track finer-grain access units for remapping can consume a significant portion of the NM capacity and can require significant on-chip memory resources for keeping remapping metadata.
[0007] Prior art has also considered a middle ground between cache and flat hybrid memories by statically setting aside a portion of NM to cache data from FM to avoid costly swap operations. The rest of the NM capacity is available to the system in flat mode. In one approach, fine-grain swapping and caching is allowed. However, metadata still consumes precious space in NM. In another approach, data is compressed in NM to agnostically expand its capacity for cache or flat space. However, cache space is still statically set aside from the flat space and compression necessitates a staging area to stabilize compressed data. Both contribute to less NM capacity being available to the system.SUMMARY
[0008] As used in this document, near memory, NM, and far memory, FM, are first and second tiers, respectively, of a two-level computer main memory hierarchy. Typically, the NM has higher bandwidth (or faster or more performant in other ways) but can be more costly than the FM. We disclose an invention that does hybrid memory compression. Unlike prior art, the disclosed method and system (i) expose the entire NM plus FM capacity in flat mode to the system, and (ii) dynamically expose a cache in NM from capacity made available from compressing data in NM. The freed-up cache is used to bring more bandwidthdemanding FM data into NM to avoid costly swap operations. Through its novel management along with a carefully crafted metadata layout, metadata can be kept in FM. The present disclosure imposes virtually no area overhead in NM for metadata or for staging.
[0009] The present disclosure unlocks space for caching by selectively compressing data in NM. This is done by dynamically gauging compressibility and bandwidth demands of fine- grain access units in FM. By allowing fine-grain management of hybrid memory, it carefully manages compressed blocks in NM (e.g. HBM) with a minimum of metadata needed. For example, it compresses NM blocks where they originally are mapped and may use surplus ECC bits (in NM, e.g. HBM) in a clever way to locate a compressed block and eliminate the use of remap tables in NM altogether. As used in this document, the term ECC bits refers to error correction code bits.
[0010] Accordingly, a first inventive aspect is a hybrid memory arrangement providing a two-level main memory hierarchy in a computer system. The hybrid memory arrangement comprises a near memory, comprising one or more near-memory modules controllable by one or more near-memory controllers. The hybrid memory arrangement further comprises a far memory, comprising one or more far-memory modules controllable by one or more far- memory controllers. A hybrid memory controller of the hybrid memory arrangement is configured for monitoring access requests to the far memory to identify bandwidthdemanding memory contents of the far memory, compressing memory contents in the near memory to form a dynamic cache from memory space freed up by the compression, and accommodating at least some of the detected bandwidth-demanding memory contents of the far memory in the dynamic cache in the near memory.
[0011] The hybrid memory controller may be configured to identify bandwidth-demanding memory contents of the far memory as memory contents for which access is requested more frequently than other memory contents of the far memory. In some embodiments, the hybrid memory controller is configured for monitoring access requests to the far memory by incrementing a reference counter for an access unit of the far memory each time access to it is requested, wherein the access unit is identified as bandwidth-demanding when the reference counter meets or exceeds a threshold value. Within the general context of the present disclosure, the term access unit refers to a memory page, a memory subpage, a memory superblock or a memory block, wherein a memory page encompasses a plurality of memory subpages, a memory subpage encompasses a plurality of memory superblocks, and a memory superblock encompasses a plurality of memory blocks. Hence, depending on implementation, the hybrid memory controller may be configured for monitoring access requests to access units being memory pages, memory subpages, memory superblocks or memory blocks.
[0012] The hybrid memory controller may be configured for accommodating an access unit of detected bandwidth-demanding memory contents of the far memory in a congruent access unit of compressed memory contents in the near memory, the congruent access unit having a same relative start address and a same size within a memory page addressable by an operating system running on the computer system, as the access unit of detected bandwidthdemanding memory contents of the far memory.
[0013] Within the general context of the present disclosure, congruence has the following meaning. A congruence group is a group of memory pages, among which the first N pages are mapped to near memory whereas the subsequent KxN pages are mapped to far memory. Then page I <=N in NM and pages I+N, I+2N, . . .1 + KN form a congruence group. The size of the congruence group is K+l. An access unit in the form of a memory page in near memory is congruent with an access unit in the form of another memory page in far memory, or vice versa, if the two pages belong to the same congruence group. An access unit of a smaller granularity than a memory page in near memory is congruent with an access unit of the same size in far memory if the two access units belong to congruent pages and have the same relative start address within the memory page.
[0014] In some embodiments, the hybrid memory controller is configured to manage a total memory space addressable by an operating system running on the computer system as follows. Subpages of memory pages are handled in three different modes: non-swap mode, swap mode and cache / compress mode.• For a subpage being in non-swap mode, the hybrid memory controller is configured for directing an access request to the near memory or the far memory depending on apage address mapping. If the subpage is in the far memory, a reference counter of the subpage is incremented. If the reference counter meets or exceeds a threshold value, a change to swap mode is made for the subpage.• For a subpage being in swap mode, the hybrid memory controller is configured for, upon entry into swap mode, swapping all access units of the subpage of the far memory with corresponding congruent access units of a congruent subpage in the near memory. For all access units of the subpage, compressibility is investigated for each pair formed by an access unit, or a sub-access unit thereof, of the subpage of the far memory and a corresponding congruent access unit, or a sub-access unit thereof, of a congruent subpage in the near memory, wherein sufficient compressibility is when the pair of compressed access units, or sub-access units thereof, will fit together in the congruent access unit, or sub-access unit thereof, of the congruent subpage in the near memory. When a fraction of the pair of access units, or the sub-access units thereof, of the subpage exhibiting sufficient compressibility meets or exceeds a compression threshold, a change cache / compress mode is made for the subpage.• For a subpage being in cache / compress mode, the hybrid memory controller is configured, for each pair of access units, or sub-access units thereof, of the subpage that exhibits sufficient compressibility, for accommodating the compressed access unit, or sub-access unit thereof, of the subpage of the far memory together with the corresponding compressed congruent access unit, or sub-access unit thereof, of the congruent subpage in the near memory in the congruent access unit, or sub-access unit thereof, of the congruent subpage in the near memory.
[0015] In some embodiments, the hybrid memory controller is configured for storing metadata for the handling of the three different modes of subpages of memory pages in the far memory.
[0016] In some embodiments, metadata indicative of a size of the compressed access unit, or sub-access unit thereof, of the far memory is encoded in unused control bits or unused parts of the congruent access unit, or sub-access unit thereof, of the congruent subpage in the near memory. Such unused control bits may, for instance, be ECC bits.
[0017] In some embodiments, the hybrid memory controller is implemented on a microprocessor chip which furthermore comprises one or more processing units, one or more levels of cache memory, said one or more near-memory controllers and said one or more far- memory modules. The hybrid memory controller is operatively connected between a last level cache of said one or more levels of cache memory and said one or more near-memory controllers and far-memory modules.
[0018] In other embodiments, the hybrid memory controller is not a separate unit, but the functionality of the hybrid memory controller is integrated with said one or more near- memory controllers.
[0019] A general benefit of the present disclosure is that it exposes the entire NM and FM capacity to the system. Beneficially, metadata for said accommodating (of at least some of the detected bandwidth-demanding memory contents of the far memory in said dynamic cache in the near memory) is stored in the far memory FM, and the total memory capacity (i.e., the total addressable memory space) exposed by the hybrid memory arrangement to an operating system running on the computer system is the total of the capacities of the near memory and the far memory. Data is compressed to free up NM capacity (e.g. HBM) tocache bandwidth-demanding blocks from FM (e.g. DRAM). This includes techniques for dynamically assessing compressibility and bandwidth demand of fine-grain access units in FM.
[0020] Another general benefit of the present disclosure is its metadata layout that keeps the overhead low by, among other techniques, placing compressed blocks at the same place as uncompressed blocks and using, for instance, surplus ECC bits to store metadata in NM compactly. This eliminates the need for any metadata in NM and leads to modestly-sized on- chip memory structures to cache metadata from FM.
[0021] A second inventive aspect is a computer system that comprises one or more processing units, one or more levels of cache memory and a hybrid memory arrangement. The hybrid memory arrangement provides a two-level main memory hierarchy in the computer system and comprises a near memory, comprising one or more near-memory modules controllable by one or more near-memory controllers. The hybrid memory arrangement further comprises a far memory, comprising one or more far-memory modules controllable by one or more far-memory controllers. A hybrid memory controller of the hybrid memory arrangement is configured for monitoring access requests to the far memory to identify bandwidth-demanding memory contents of the far memory, compressing memory contents in the near memory to form a dynamic cache from memory space freed up by the compression, and accommodating at least some of the detected bandwidth-demanding memory contents of the far memory in the dynamic cache in the near memory.
[0022] The hybrid memory arrangement in the computer system according to the second inventive aspect may be configured as defined for the hybrid memory arrangement according to the first inventive aspect, including any of its embodiments as presented herein. Again, beneficially, metadata for said accommodating (of at least some of the detected bandwidthdemanding memory contents of the far memory in said dynamic cache in the near memory) is stored in the far memory FM, and the total memory capacity (i.e., the total addressable memory space) exposed by the hybrid memory arrangement to an operating system running on the computer system is the total of the capacities of the near memory and the far memory.
[0023] A third inventive aspect is a method of providing a two-level main memory hierarchy in a computer system. The two-level main memory hierarchy comprises a near memory and a far memory, together forming a hybrid memory in the computer system. The method involves monitoring access requests to the far memory to identify bandwidth-demanding memory contents of the far memory. The method further involves compressing memory contents in the near memory to form a dynamic cache from memory space freed up by the compression, and accommodating at least some of the detected bandwidth-demanding memory contents of the far memory in the dynamic cache in the near memory. Again, beneficially, metadata for said accommodating (of at least some of the detected bandwidth-demanding memory contents of the far memory in said dynamic cache in the near memory) is stored in the far memory FM, and the total memory capacity (i.e., the total addressable memory space) exposed by the hybrid memory arrangement to an operating system running on the computer system is the total of the capacities of the near memory and the far memory.
[0024] The method according to the third inventive aspect may comprise the functionality defined for the hybrid memory arrangement according to the first inventive aspect, including any of its embodiments as presented herein.
[0025] A fourth inventive aspect is a hybrid memory controller for use in a hybrid memory arrangement that provides a two-level main memory hierarchy in a computer system, the two- level main memory hierarchy comprising a near memory and a far memory, together forming a hybrid memory in the computer system. The hybrid memory controller of the fourth inventive aspect is configured for monitoring access requests to the far memory to identify bandwidth-demanding memory contents of the far memory, compressing memory contents in the near memory to form a dynamic cache from memory space freed up by the compression, and accommodating at least some of the detected bandwidth-demanding memory contents of the far memory in the dynamic cache in the near memory. Again, beneficially, metadata for said accommodating (of at least some of the detected bandwidth-demanding memory contents of the far memory in said dynamic cache in the near memory) is stored in the far memory FM, and the total memory capacity (i.e., the total addressable memory space) exposed by the hybrid memory arrangement to an operating system running on the computer system is the total of the capacities of the near memory and the far memory.
[0026] The hybrid memory controller of the fourth inventive aspect may further be configured as the hybrid memory controller in the hybrid memory arrangement according to the first inventive aspect, including any of its embodiments as presented herein.
[0027] The present disclosure is agnostic to the type of memory technologies considered for near memory and far memory. In one embodiment, High-Bandwidth Memory (HBM) could be used as near memory and Dynamic Random Access Memory (DRAM) could be used as far memory. In another embodiment, DRAM could be used as near memory and Non- Volatile Memory (NVM) could be used as far memory. In yet another embodiment, a DRAM stacked on top of a microprocessor die could be used as near memory and DRAM off of the microprocessor die could be used as far memory. Someone skilled in the art can consider other memory technologies for near memory and far memory and all such embodiments are contemplated.
[0028] The present disclosure is also agnostic to the choice of compression / decompression algorithms / accelerators, lossless or lossy. In one embodiment, base-delta immediate compression (BDI) could be used to compress far memory-mapped pages, subpages or superblocks together with near memory-mapped pages, subpages or superblocks. In another embodiment, any entropy -based compression algorithm could be used for the same purpose. Someone skilled in the art can consider other compression algorithms / accelerators for near memory and far memory and all such embodiments are contemplated.
[0029] A general benefit with the present disclosure is that the total memory capacity exposed by the hybrid memory arrangement to an operating system running on the computer system is the total of the capacities of the near memory and the far memory.
[0030] Other aspects, as well as objectives, features and advantages of the disclosed embodiments will appear from the following detailed patent disclosure, from the attached dependent claims as well as from the drawings.
[0031] Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise herein. All references to "a / an / the [element, device, component, means, step, etc.]" are to be interpreted openly as referring to at least one instance of the element, device, component, means, step, etc., unlessexplicitly stated otherwise. The steps of any method disclosed herein do not have to be performed in the exact order disclosed, unless explicitly stated.BRIEF DESCRIPTION OF THE DRAWINGS
[0032] FIG. 1 depicts a computer system comprising a microprocessor chip with one or a plurality of processing units, an exemplary cache hierarchy of three levels, one or a plurality of memory controllers for near and far memory connected to one or a plurality of off-chip near and far memories.
[0033] FIG. 2A depicts the same computer system as in FIG. 1 extended with a hybrid memory controller configured to manage the placement of data in far and near memory to allow for data in near memory to be compressed to free up capacity for caching data in far memory.
[0034] FIG. 2B depicts an embodiment of the hybrid memory controller in FIG. 2A, comprising a metadata cache, a control unit, a compression and decompression accelerator.
[0035] FIG. 3 depicts how pages mapped to near and far memory form congruence groups.
[0036] FIG. 4 depicts how blocks in pages in far and near memory forming congruence groups can be compressed together in near memory if they are sufficiently compressible allowing data in far memory to be cached in near memory.
[0037] FIG. 5 depicts the metadata layout used to handle subpage and superblock portions of a page with respect to each operating mode.
[0038] FIG. 6A depicts how blocks in pages in far and near memory forming congruence groups can be in three different operating modes: baseline, cache / compress and swap modes.
[0039] FIG. 6B depicts how congruent blocks can be stored in near memory when they are compressed sufficiently to free up cache space for additional far memory blocks to be cached in near memory.
[0040] FIG. 7 depicts the flow graph of a method for handling read / write requests to pages mapped to far memory when the requested subpage is in swap mode.
[0041] FIG. 8 depicts the flow graph of a method for handling read / write requests to pages mapped to far memory when the requested subpage is in cache / compress mode.
[0042] FIG. 9 depicts the flow graph of a method for deciding whether to cache or swap a superblock.
[0043] FIG. 10 depicts the flow graph of a method for handling far memory write requests.
[0044] FIG. 11 is an alternative embodiment wherein FIG. 11 depicts the same computer system as in FIG. 2A but where the hybrid memory controller is now integrated in each nearmemory device.
[0045] FIG. 12 shows and alternative embodiment wherein the hybrid memory controller is integrated in a high-bandwidth memory device.DETAILED DESCRIPTION
[0046] An exemplary embodiment of a computer system 100 is depicted in FIG. 1. This system comprises a microprocessor chip 110 and two types of memory comprising an exemplary two-level hybrid memory hierarchy where a plurality of one type of M memory modules denoted near memory (NM) are shown as NMi 151 to NMM 152 and a plurality of another type of L memory modules denoted far memory (FM) are shown as FMi 153 to FML 154. Together, the near memory NM and far memory FM constitute a hybrid memory HM. The microprocessor chip 110 could be a discrete system or integrated on a system-on-a-chip (SoC) in any available technology. The microprocessor 110 comprises one or several processing units, denoted Pi 131, P2 132 through PN 133 sometimes called CPU or core and a memory hierarchy. In one exemplary embodiment the near memory is made up of High- Bandwidth Memory (HBM) and the far memory is made up of Dynamic Random Access Memory (DRAM). However, near and far memory could be realized with other combinations. For example, near memory in one embodiment could be realized with DRAM whereas the far memory could be realized as non-volatile memory such as phase-change memory (PCM) or Magnetic RAM (MRAM).In the embodiment of FIG. 1, LLC requests are routed to an HBM or DRAM controller, depending on whether a page is mapped to NM or FM (HBM and DRAM, respectively, in a preferred embodiment) as dictated by the virtual / physical page address mapping.
[0047] The memory hierarchy, on the other hand, comprises several cache levels, e.g., three levels as is shown exemplary in FIG. 1 and denoted Cl, C2, and C3. These levels can be implemented in the same or different memory technologies, e.g., SRAM, DRAM, or any type of non-volatile memory technology including, for example, Spin-Transfer Torque Random Access Memories (STT-RAM). The number of cache levels may vary in different embodiments and the exemplary embodiment 100 depicts three levels where the last cache level is C3 120. These levels are connected using some kind of interconnection means (e.g., bus or any other interconnection network). In the exemplary embodiment, levels Cl and C2 are private to, and only accessible by, a respective processing unit z denoted Pi (e.g., Pi in FIG. 1). It is well known to someone skilled in the art that alternative embodiments can have any number of private cache levels or, as an alternative, all cache levels are shared as illustrated by the third level C3 120 in FIG. 1. Regarding the inclusion of the data in the cache hierarchy, any embodiment is possible and can be appreciated by someone skilled in the art. For example, Cl can be included in C2 whereas C2 can be non-inclusive with respect to level C3 120. Someone skilled in the art can appreciate alternative embodiments.
[0048] The computer system 100 of FIG. 1 comprises one or a plurality of memory controllers where memory controllers are also of two types with M near memory controllers denoted NMCTRLi 141 to NMCTRLM 142 and L far memory controllers denoted FMCTRLi 143 to FMCTRLL 144. The last cache level (C3 120 in FIG. 1) is in this embodiment connected to the memory controllers, which in turn are connected to one or a plurality of memory modules of respective types. In other embodiments the last cache level can be realized as a banked memory with, for example, one bank per processing unit. Memory controllers can be integrated on the microprocessor chip 110 or can be implemented outside of the microprocessor chip for example tightly with the near memories NMi 151 andNMM 152. An operating system (OS) 170 is running on the computer system 100, as is of course commonplace per se. Finally, the computer system 100 performs one or more tasks. A task can be any software application or part of it that can be run by the operating system 170 of the computer system 100.
[0049] In a first embodiment according to FIG. 1 denoted Baseline System 1 (BL1), the operating system 170 manages the two levels of the hybrid memory in flat mode, where flat mode means that the capacity of each of the two levels are exposed to the operating system 170, so the entire capacity is the sum of the NM and FM capacities. Pages are in this embodiment interleaved using the concept of congruence groups known from prior art as shown in FIG. 3. Here, the page size is N and page A 312 occupies addresses 0 314 to N- 1 316 and is mapped to NM whereas pages B 322, C 332, D 342 and E 352 are mapped to FM, assuming a congruence group with one NM page and four FM pages. For the meaning of congruence, reference is made to the Summary section. The following is recalled. Suppose that, for a given set of memory pages, the first N pages are mapped to NM whereas the subsequent KxN pages are mapped to FM. Then page I <=N in NM and pages I+N, I+2N, . . .1 + KN form a congruence group. Only one of the pages in FM belonging to a congruence group can be swapped with a page in NM belonging to the same congruence group.(swapping being explained in more detail later). We say that the size of the congruence group is K+l. As the skilled person will readily understand, this is analogous to a direct-mapped cache whose size is N whereas the size of memory is KN. Then, K blocks in memory are mapped to the same block frame in the cache and only one of said blocks can be in the direct- mapped cache at any point in time.
[0050] The second embodiment, denoted Baseline System 2 (BL2) is built on top of BL1. When a page (or portion of it) mapped to FM is deemed bandwidth demanding, it is remapped to NM transparently to the operating system. This involves a swap operation with the page (or portion of it) congruent to it and located in NM. For example, block K in page C 332 in FM is congruent with block K in page A 312 in FM and the two blocks will be swapped with each other (see FIG. 3).
[0051] The granularity chosen as the portion of a page to monitor bandwidth demand will determine the amount of metadata needed to keep track of it. The finer the grain size the more metadata is needed. This will push towards larger grain sizes. On the other hand, a too large grain-size can lead to over-fetching of data due to limited spatial locality and result in too high traffic overhead for swapping. Therefore, the tradeoff when selecting a grain size is between metadata overhead and spatial locality and grain sizes other than pages and blocks will be considered in this patent disclosure. In the remainder of the document, and to clarify terminology, we will deal with the following grain sizes of access units: pages, subpages, superblocks and blocks, where the size of pages =< subpages =< superblocks =< blocks. As exemplary sizes of these access units, we will consider 4 KB, 2 KB, 512 B and 64 B for pages, subpages, superblocks and blocks, respectively, if not stated otherwise, where B refers to bytes.
[0052] The objective of the disclosed system and method is to free up capacity in NM using sufficiently fast compression techniques and use the freed-up capacity to cache bandwidthdemanding FM blocks. Without loss of generality, NM could exemplary be HBM devices whereas FM could exemplary be DRAM devices.
[0053] Just like in the BL1 and BL2 embodiments, pages are mapped by the operating system to NM and FM using congruence groups. Turning to FIG. 2A, it shows the same embodiment as FIG. 1 but extended with a functional block, denoted HMCTRL 260. The rest of the elements in FIG. 2A may be same as or equivalent to the corresponding elements in FIG. 1 (with like references numerals representing like elements, such that microprocessor chip 210, processing unit 231 and operating system 270 in FIG. 2A may be the same as or equivalent to microprocessor chip 110. processing unit 131 and operating system 170 in FIG. 1, etc.). HMCTRL 260 intercepts all C3 220 requests. Initially, HMCTRL 260 adopts the policy of BL1 to all FM pages meaning that none of them are subject to swapping from the very start. This is referred to as non-swap mode. However, when a FM-mapped page is accessed, it will be tracked by a reference counter at the granularity of a subpage. The reference counter will be incremented for each access to said subpage. For as long as the reference counter is below a preset threshold, all superblocks of the tracked subpage will be accessed from FM and will not be swapped. However, when the reference counter exceeds a preset threshold, we say that the subpage is bandwidth demanding and the subpage will turn into swap mode. From this point, all accessed FM superblocks associated with a subpage in swap mode will be swapped with their corresponding NM superblocks belonging to the same congruence group. We will provide details on how devices are configured and methods set up to make this happen further down below.
[0054] Once a subpage switches to swap mode, attempts will be made to gain cache space in NM through compression. This is done by attempting to compress a FM superblock in swap mode together with its corresponding congruent NM superblock. We note that the present disclosure is agnostic to the choice of compression algorithm and that any reasonably fast compression algorithm in prior art can be used. The present disclosure is agnostic to the type of memory technologies considered for NM and FM. In one embodiment. In one embodiment, base-delta immediate compression (BDI) could be used to compress FM- mapped pages, subpages or superblocks together with NM-mapped pages, subpages or superblocks. In another embodiment, any entropy-based compression algorithm could be used for the same purpose. Someone skilled in the art can consider other compression algorithms / accelerators for NM and FM and all such embodiments are contemplated.
[0055] When a FM superblock is requested, HMCTRL 260 will also request the corresponding congruent NM superblock. Next, for each pair of blocks in the two superblocks, it will be decided whether these two blocks compress sufficiently well, meaning that they can be accommodated within the same block frame of, say, 64 B. If a certain fraction of the blocks within a requested superblock is sufficiently well compressed, using above definition, the corresponding subpage is deemed to be in cache / compress mode. All sufficiently well compressible blocks inside said superblock will be compressed. An uncompressible FM-mapped block will stay uncompressed in FM.
[0056] From now on, an attempt is made to compress all superblocks inside the subpage to be placed in NM. HMCTRL 260 will then update the metadata table, to be explained in detail later, so that subsequent requests for the remapped superblocks are destined to NM with no involvement of FM. Otherwise, the subpage will remain in swap mode and the requested superblock will be swapped with the corresponding congruent superblock in NM, to be explained in detail later.
[0057] FIG. 2B shows, at 2000, an embodiment of HMCTRL 260 in FIG. 2A. It comprises four functional blocks: a metadata cache 2100, a control unit 2200, a compression accelerator2300 and a decompression accelerator 2400. As will be explained later in detail, the metadata table is in some embodiments stored in its entirety in FM and portions of it are cached in the HMCTRL 260. The cache for metadata table entries is shown in FIG. 2B at 2100. Portions of the memory address is used to index into the metadata cache and the rest is used as a tag, just like in any cache structure known in prior art. This cache structure can on one extreme be a direct-mapped cache and on another extreme a fully associative cache or be set associative. To allow superblocks within a subpage to be compressed, a compression accelerator 2300 is available. Similarly, to allow a compressed superblock to be decompressed, a decompression accelerator 2400 is available. These accelerators could be designed to accelerate any known compression method, such that entropy encoding, base delta immediate compression, or other methods known in prior art. The detailed operation of the HMCTRL will be explained subsequently and the control unit implementing the operation is another block 2200 in HMCTRL 260 (i.e., 2000).
[0058] To see how compressible blocks are compressed, FIG. 4 shows three contiguous compressed blocks (blocks N-l, N, and N+l) from two congruent pages A 410 and B 420. Here, blocks N-l 422, N 424, and N+l 426 from page B in FM are compressed and stored together with their corresponding congruent blocks N-l 412, N 414, and N+l 416 from page A in NM 410. If the LLC subsequently requests block N in page B 424, HMCTRL 260 will reroute the request to NM based on the metadata to the cached th block in page A. In the case that the FM block and the corresponding congruent NM block cannot fit into the 64-B block frame, the FM block will remain in FM and HMCTRL 260 will verify the response from NM and then forward the request to FM, to be explained in detail further down.Metadata Layout
[0059] For HMCTRL 260 to decide which action to take for each LLC request, it uses a metadata cache. Going back to FIG. 2A, C3 220 requests will either be routed to NM or FM.
[0060] Recall that HMCTRL 260 initially operates the hybrid memory in flat mode, where a page is mapped to NM or FM by the operating system 270. However, when the reference count of a referenced FM-mapped subpage exceeds a preset threshold, it will be remapped from FM to NM at the granularity of superblocks. From this point, requested FM-mapped superblocks belonging to that subpage will be swapped with the corresponding congruent NM-mapped superblock.
[0061] The layout of the metadata table is shown in FIG. 5. Starting with box 510 the metadata table associates an entry with each congruence group: Metadata congruence group 0 512 and Metadata congruence group 1 514. Hence, it has as many entries as the number of pages in NM. Metadata table is preferably stored in FM (this will save precious resources from the NM), or alternatively in NM, but metadata of the metadata table is cached in the metadata cache in HMCTRL 260 / 2000 of FIG. 2A and Fig. 2B (the metadata cache being indicated at 2100 in FIG. 2B).. The metadata entry for a congruence group is constructed to track one out of all FM pages belonging to a congruence group. Hence, box 520 shows the metadata for exemplary Metadata congruence group 1 and assuming two subpages per page, a metadata entry for a congruence group needs a Tag 522 of 2 bits to designate one out of exemplary four tracked FM pages in the congruence group and the two subpages - Subpage 0 metadata 524 and subpage 1 metadata 526 (e.g., 2 KB each) belonging to the tracked page (e.g., 4 KB), with 15 bits of meta data for each subpage.
[0062] The 15-bit metadata field for each subpage is shown in box 530. To the left, a single bit (Swap / Comp) 531 together with the content of the reference counter (Reference Counter) 540 designates whether the subpage is in non-swap, swap or cache / compress mode. If the reference count is below a preset threshold, the subpage is in non-swap mode and requests will be destined to NM or FM based on the virtual-to-physical address mapping. If the reference count is above a preset threshold, the Swap / Comp flag 531 designates whether the subpage is in swap mode (flag is set) or in cache / compress mode (the flag is reset). Next, there is one valid bit for each of the four superblocks, 532, 533,. . ., 535 with an exemplary size of 512 B that belong to a subpage with an exemplary size of 2 KB. A valid superblock bit designates that the FM-mapped superblock is swapped (Swap / Comp bit 531 set) or compressed in NM together with its corresponding NM-mapped superblock (Swap / Comp bit 531 cleared). There are also 4 dirty superblock bits, 536, 537, . . . , 539. Whenever a request is written back in cache / compress mode, the superblock dirty bit will be set. Finally, the Reference Counter for a superblock 540 uses 6 bits.Metadata Cache
[0063] Recall that we assume that the metadata table is stored in FM. HMCTRL 260 of FIG. 2A is configured to cache contents of the metadata table using a metadata cache, illustrated in FIG. 2B as metadata cache 2100. In one embodiment, each metadata cache entry (say 64 B) in the metadata cache contains 16 consecutive metadata entries of 32 bits each. Thus, the exemplary metadata cache design is indexed by an address corresponding to the requested congruence group, stripping out the least significant 4 bits. Given that the NM size is C = 2Cand the page size is P = 2Pthere are N = 2C'Pcongruence groups. The congruence group can be stripped out from the physical page number taking the most significant c bits.
[0064] On a metadata cache hit, HMCTRL 260 in FIG. 2A will update the reference counter and retrieve the request's corresponding metadata entry from the metadata cache. Conversely, on a metadata cache miss, HMCTRL 260 of FIG. 2 A will first evict an entry to make room for the requested metadata entry and, if needed, write back the evicted metadata entry to where the metadata table is stored, for instance in FM. Then, the missing metadata entry from FM is fetched into the metadata cache.Near-Memory Metadata Support
[0065] FIG. 6 A shows the mapping of blocks N-l, N, and N+l in the baseline inside two pages (A 612 and B 614) in the cache & compress mode (page A 622 and page B 624) and in swap mode (page A 632 and page B 634). Here, pages A and B belong to the same congruence group and page A is mapped to NM whereas page B is mapped to FM. Recall that C3 220 block requests, as intercepted by the HMCTRL 260 in FIG. 2A, will be destined to NM in three cases with reference to FIG. 6A.
[0066] The first case is in non-swap mode, when the block is mapped to NM. This corresponds to the baseline 610 in FIG. 6A. The second case is in swap mode, when the Swap / Comp bit 530 in FIG. 5 is set and the superblock valid bit is set, the one out of 532, 533, . . ., 535 that corresponds to the superblock. This corresponds to Swap 630 in FIG. 6A. Here, all blocks in page A 612 have been swapped with the blocks in page 614. Finally, the third case is in cache / compress mode when the Swap / Comp bit 530 in FIG. 5 is reset and the superblock valid bit is set (the one out of 532, 533, . . ., 535 in FIG. 5 that corresponds to the superblock). This corresponds to Cache / compress 620 in FIG. 6A.
[0067] In cache / compress mode it is not certain that the requested block is in NM as a block may not compress sufficiently. Therefore, block-level information must be maintained in NM whether the FM-mapped block is compressed together with the corresponding congruent NM-mapped block. If not, the FM-block is placed in FM and the request must be rerouted to FM.
[0068] HBM devices sometimes can associate 16 ECC bits with each 32-B access unit. When the NM-mapped and FM-mapped blocks in the same congruence group are compressed to fit into a 64-B block frame (two 32-B access units), in one embodiment one can use 6 unused ECC bits (out of 16) to encode the validity and size of the compressed NM-mapped and FM- mapped blocks. If the FM-mapped block is not compressed, it is stored in FM. This case is recorded by setting all six ECC bits to zero. If the FM and NM-mapped blocks are compressed, their compressed sizes can be recorded in the unused respective six ECC bits. For example, if the compressed size is 63 bytes, the 6 ECC bits can be encoded T 11111' and if the compressed size is 2 bytes, the 6 ECC bits can be encoded '000010'.
[0069] As shown in FIG. 6B, we consider three consecutive blocks denoted N-l, N and N+l in two congruent pages 662 and 664. Here, we assume that page 662 is mapped to NM and page 664 is mapped to FM. Now assume that blocks N in the subpage of congruent pages 662 and 664 are in cache / compress mode. This means that the two blocks denoted N, 666 and 668 in FIG. 6B, in the two congruent pages 662 and 664, respectively, will be compressed and placed in NM if they fit into the 64-block frame. This is illustrated by arrows 666 and 668 in FIG. 6B. For the two blocks, the NM block is placed starting at the original address of the block frame whereas the corresponding congruent FM-mapped block is mapped to the end of the block frame. Pointer 676 and Pointer 678 refer to the 6 ECC bits and are in effect interpreted as the location of the last byte of each compressed block.
[0070] The use of such unused ECC bits to encode the validity and size of the compressed NM-mapped and FM-mapped blocks may reduce the memory waste and lower the overhead of metadata. In case ECC bits cannot be used, an alternative embodiment is to store metadata needed for compression, i.e., the size of the compressed block as part of the unused portions of the block. The only metadata needed outside of NM would be to designate whether the block is compressed, a single bit per block. A further alternative is a scheme that combines unused ECC bits and metadata bits stored in memory, as can be realized by someone skilled in the art. In yet another alternative, memory bits reserved for a certain purpose that are not used for said purpose, can be repurposed to be used as metadata bits by those skilled in the art. All such alternatives are contemplated.Detailed Operation
[0071] We now review in detail the operation of HMCTRL 260 in FIG. 2A. Recall that HMCTRL 260 initially, when in non-swap mode with respect to an LLC request, will send the LLC request to NM or FM depending on its virtual-to-physical mapping.
[0072] A subpage will make a mode change from the non-swap mode to swap mode when the reference count exceeds a preset threshold. FIG. 9 shows the flow graph of going from swap mode to cache / compress mode for a subpage. For each subpage in swap mode and at each FM superblock request for that subpage, the first action is to request the FM superblock along with the corresponding congruent NM superblock 910. All blocks in the twosuperblocks will be pair-wise compressed 920. A pair of blocks that are compressed and can fit into a 64B block frame will be successfully compressed. If the fraction of NM and FM block pairs that are successfully compressed to fit within a block frame are above a compression threshold (TH) 930, the valid bit for said superblock will be set and the subpage will be in cache / compress mode 940. Otherwise, the subpage is set to swap mode 950. In both cases, the process will thereafter terminate 960.Transaction Flow for Last-level Cache Requests
[0073] We now consider the transaction flow associated with LLC read or write requests to FM-mapped pages as shown in FIG. 7 and FIG. 8 in the case the subpage is in swap mode (FIG. 7) and cache / compress mode (FIG. 8). As shown in FIG. 7, in swap mode, upon receiving a FM read or write request 710, HMCTRL 260 will first increment the reference counter of the referenced subpage 720. It will then check whether the superblock is already in NM as a result of an earlier swap operation. If the superblock Valid bit is set 730, HMCTRL 260 will forward the request to NM 750. If the valid bit is cleared and the reference counter for the subpage is below a preset threshold 740, HMCTRL 260 will swap the superblock in FM with the superblock in NM 760. This applies to both read and write FM requests in swap mode. Going back to the decision box 740, if the reference counter is above a preset threshold, the request is forwarded to FM 770. In all cases, the action is terminated in box 780.
[0074] When considering the transactions of read requests to subpages in cache / compress mode 805 in FIG. 8, HMCTRL 260 first increments the reference counter of the subpage 810 and then verifies the validity of the superblock 815. In case the superblock is valid, the request will be forwarded to NM 820. Whether or not the block is valid will be checked in NM 835. This may in some embodiments involve checking the ECC bits 830. In case the block is valid, a response will be returned 845 to HMCTRL 260. If the block is not valid, a request will be forwarded to FM 840.
[0075] It is possible that some of the FM blocks within the superblock are cached in the compressed NM, while other FM blocks remain in FM. The latter can apply to FM blocks that cannot be compressed and stored within a 64-B block frame along with the corresponding congruent and compressed NM block. For this reason, HMCTRL 260 will examine the response from NM, including the data and relevant ECC bits. If the ECC bits are non-zero, the FM block is compressed and cached and HMCTRL will respond to LLC. If the ECC bits are zero, however, HMCTRL will forward the request to FM 840.
[0076] Going back to the decision box that checks the validity of the superblock 815, in case the superblock is not valid, the next action is to check the reference counter 825 for the corresponding subpage. In case it is above a preset threshold, the superblock will be in cache / compress mode and all blocks in the superblock will be attempted to be compressed with their congruent counterparts in FM 850. On the other hand, if the reference counter for the corresponding subpage is not above a preset threshold, the request will be forwarded to FM 855. Regardless of the outcome of the tests in. the decision boxes, the process will eventually end up in box 860 where the process terminates.
[0077] The process for handling FM write request hits to subpages in cache / compress mode is illustrated in Figure 10. The issue here is that a block written back from the last-level cache, e.g., C3 220 in FIG. 2A may, after compression, change in size and may not fitanymore. First, a test is carried out whether the superblock is valid 1010. If not, the block does not exist in NM and the request is forwarded to FM 1050. If the superblock is valid, an attempt is made to compress the block with the corresponding congruent block 1020. Second, a test is carried out whether the size of the compressed written-back block and the size of the congruent block can fit in a 64B block frame 1030. If it can, the block is written back and the ECC bits are updated to reflect its new size. Meanwhile, the superblock dirty bit is set. However, if it exceeds the size but can still fit into the 64-B block frame together with the NM-mapped block, the block is written into NM 1040 and the process terminates 1060. Finally, if it does not fit, the block will be forwarded to FM 1050 and in embodiments using the unused ECC bits of the congruent block in NM these bits will be reset to reflect that it is not valid.Other Relevant Operations
[0078] When a page in a congruence group is being tracked, it can happen that another page in that same congruence group will be accessed. For as long as the first page is not in swap mode, accesses to other pages inside the same congruence group will be disregarded. However, when the preset threshold is exceeded and the page will turn into swap mode, accesses by other pages inside the congruence group will decrement the reference counter. If it hits zero, the page will not be considered bandwidth demanding anymore and will make a transition from either cache / compress or swap mode to non-swap mode. This transition necessitates that all superblocks that have been potentially migrated to NM in noncompressed or compressed fashion must move back non-compressed to FM. We will call this operation page consolidation.
[0079] Page consolidation is also needed when the mapping is changed for a page by the operating system 270. Then, typically, TLB entries to the page must be invalidated (TLB shoot-down) and all blocks from said page must be evicted too. Page consolidation does exactly the latter as follows.
[0080] The page metadata is consulted. For each subpage and each superblock of that subpage, if the superblock is in NM in swap mode, it will be swapped with the superblock in FM. If it is cached and compressed in NM, it will be decompressed and written back if the superblock is dirty or silently evicted if it is not dirty.
[0081] The embodiment of FIG 2A, wherein a hybrid memory controller HMCTRL 260 is integrated on a microprocessor chip 210, has been disclosed so far. Alternatively, it can be valuable to avoid integrating any functionality of the present disclosure on a microprocessor chip 210. Instead, it can be of value to integrate the functionality of the hybrid memory controller HMCTRL within the near-memory devices. In the following, we contemplate such an alternative embodiment with reference to FIG. 11.
[0082] The difference between the embodiments of FIG. 2A and FIG. 11 is that the hybrid memory controller 260 in FIG. 2 A has been replaced with an interconnect 1160 that connects the last-level cache C3 1120 with the near and far memory controllers 1141-1142 (NMCTRLI-NMCTRLM) versus 1143-1144 (FMCTRLi and FMCTRLL), respectively. The functionality of the hybrid memory controller HMCTRL is no longer a separate unit like 260 in FIG. 2A but integrated with the near memory controllers 1141-1142. This can be seen in FIG. 11. In the embodiment of FIG. 11, the operating system 1170 will decide whether a pageis mapped to a far memory or a near memory device, and the virtual-to-physical address translation will decide whether to destine the request to a near or a far memory device.
[0083] In FIG. 12 we show an embodiment using High-Bandwidth Memory (HBM) as nearmemory devices. An HBM device comprises a plurality of memory dies 1220, 1230 and 1240, say K, as in FIG. 12. These can be stacked on top of each other and can be connected to a logical layer 1210 using a plurality of through-silicon vias (TSVs), say N as seen at 1260, 1270, 1280 in FIG. 12. In an alternative embodiment, the hybrid memory controller 260 of FIG. 2A can be integrated in the logical layer of all near-memory devices which exemplary can be HBM devices.
[0084] Someone skilled in the art can realize that the physical placement of the hybrid memory controller will not affect neither its functionality nor its operation and what has been said about how it can be configured and operated applies to this alternative embodiment.Concluding Remarks
[0085] What has been disclosed above is a hybrid memory arrangement providing a two- level main memory hierarchy in a computer system 200; 1100. The hybrid memory arrangement comprises a near memory NM, comprising one or more near-memory modules 251-252; 1151-1152 controllable by one or more near-memory controllers 241-242; 1141- 1142. The hybrid memory arrangement further comprises a far memory FM, comprising one or more far-memory modules 253-254; 1153-1154 controllable by one or more far-memory controllers 251-252; 1151-1152. The hybrid memory arrangement moreover comprises a hybrid memory controller 260 configured for monitoring access requests to the far memory FM to identify bandwidth-demanding memory contents of the far memory FM, compressing memory contents in the near memory NM to form a dynamic cache from memory space freed up by the compression, and accommodating at least some of the detected bandwidthdemanding memory contents of the far memory FM in the dynamic cache in the near memory NM. The entire capacity of NM and FM is available for use by the operating system 270; 1170. Arrangements using one or many NMs and one or many FMs are contemplated. This means that the total memory capacity exposed to the operating system 270; 1170 is the total of the capacity of NMs and FMs.
[0086] The hybrid memory arrangement will dynamically monitor the access frequency and compressibility of pages, subpages (portions or all of pages) and superblocks (portions or all of subpages). Based on the access frequency and the compressibility of pages, subpages and superblocks, the hybrid memory arrangement will decide which of the FM allocated pages, subpages or superblocks are to be migrated from FM to NM and which of the FM allocated pages, subpages or superblocks are to be compressed together with NM allocated pages, subpages or superblocks to allow pages, subpages and superblocks to be located in NM without using any physical space therein. In addition, the hybrid memory arrangement is configured and operated in a way that is transparent to the software, e.g., the operating system. The hybrid memory arrangement compresses data stored in NM, and in the space created it stores (also in compressed form) data from the FM. Essentially, it creates a dynamic cache with this space created by memory compression.
[0087] As used in this document, thresholds used in for the dynamic monitoring and dynamic decisions are exemplary and other thresholds can be used by those skilled in the art, depending on the tuning of the target implementation.
[0088] What has additionally been disclosed in the present disclosure is a computer system 200; 1100 comprising one or more processing units 231-233; 1131-1133, one or more levels of cache memory C1-C3 and a hybrid memory arrangement. The hybrid memory arrangement provides a two-level main memory hierarchy in the computer system 200; 1100 and comprises a near memory NM, comprising one or more near-memory modules 251-252; 1151-1152 controllable by one or more near-memory controllers 241-242; 1141-1142. The hybrid memory arrangement further comprises a far memory FM, comprising one or more far-memory modules 253-254; 1153-1154 controllable by one or more far-memory controllers 251-252; 1151-1152. The hybrid memory arrangement moreover comprises a hybrid memory controller 260 configured for monitoring access requests to the far memory FM to identify bandwidth-demanding memory contents of the far memory FM, compressing memory contents in the near memory NM to form a dynamic cache from memory space freed up by the compression, and accommodating at least some of the detected bandwidthdemanding memory contents of the far memory FM in the dynamic cache in the near memory NM.
[0089] What has further been disclosed is a method of providing a two-level main memory hierarchy in a computer system 200; 1100, wherein the two-level main memory hierarchy comprises a near memory NM and a far memory FM, together forming a hybrid memory HM in the computer system 200; 1100. The method involves monitoring access requests to the far memory FM to identify bandwidth-demanding memory contents of the far memory FM. The method further involves compressing memory contents in the near memory NM to form a dynamic cache from memory space freed up by the compression, and accommodating at least some of the detected bandwidth-demanding memory contents of the far memory FM in the dynamic cache in the near memory NM.
[0090] Finally, what has also been disclosed is a hybrid memory controller 260 for use in a hybrid memory arrangement that provides a two-level main memory hierarchy in a computer system 200; 1100, the two-level main memory hierarchy comprising a near memory NM and a far memory FM, together forming a hybrid memory HM in the computer system 200; 1100. The hybrid memory controller 260 is configured for:• monitoring access requests to the far memory FM to identify bandwidth-demanding memory contents of the far memory FM,• compressing memory contents in the near memory NM to form a dynamic cache from memory space freed up by the compression, and• accommodating at least some of the detected bandwidth-demanding memory contents of the far memory FM in the dynamic cache in the near memory NM.
Claims
CLAIMS1. A hybrid memory arrangement providing a two-level main memory hierarchy in a computer system (200; 1100), the hybrid memory arrangement comprising: a near memory (NM), comprising one or more near-memory modules (251-252; 1151-1152) controllable by one or more near-memory controllers (241-242; 1141-1142); a far memory (FM), comprising one or more far-memory modules (253-254; 1153- 1154) controllable by one or more far-memory controllers (251-252; 1151-1152); and a hybrid memory controller (260) configured for monitoring access requests to the far memory (FM) to identify bandwidth-demanding memory contents of the far memory (FM), compressing memory contents in the near memory (NM) to form a dynamic cache from memory space freed up by the compression, and accommodating at least some of the detected bandwidth-demanding memory contents of the far memory (FM) in the dynamic cache in the near memory (NM).
2. The hybrid memory arrangement as defined in claim 1, wherein metadata for said accommodating is stored in the far memory (FM) and wherein the total memory capacity exposed by the hybrid memory arrangement to an operating system (270; 1170) running on the computer system (200; 1100) is the total of the capacities of the near memory (NM) and the far memory (FM).
3. The hybrid memory arrangement as defined in claim 1 or 2, wherein the hybrid memory controller (260) is configured to identify bandwidth-demanding memory contents of the far memory (FM) as memory contents for which access is requested more frequently than other memory contents of the far memory (FM).
4. The hybrid memory arrangement as defined in any preceding claim, wherein the hybrid memory controller (260) is configured for monitoring access requests to the far memory (FM) by incrementing a reference counter for an access unit of the far memory (FM) each time access to it is requested, an access unit being one of a memory page, a memory subpage, a memory superblock and a memory block, and wherein the access unit is identified as bandwidth-demanding when the reference counter meets or exceeds a threshold value.
5. The hybrid memory arrangement as defined in any preceding claim, wherein the hybrid memory controller (260) is configured for accommodating an access unit of detected bandwidth-demanding memory contents of the far memory (FM) in a congruent access unit of compressed memory contents in the near memory (NM), the congruent access unit having a same relative start address and a same size within a memory page addressable by an operating system (270; 1170) running on the computer system (200; 1100), as the access unit of detected bandwidth-demanding memory contents of the far memory (FM).
6. The hybrid memory arrangement as defined in any preceding claim, wherein the hybrid memory controller (260) is configured to manage a total memory space addressable by an operating system (270; 1170) running on the computer system (200; 1100), as follows: handling subpages of memory pages in three different modes: non-swap mode, swap mode and cache / compress mode; for a subpage being in non-swap mode: directing an access request to the near memory (NM) or the far memory (FM) depending on a page address mapping, if the subpage is in the far memory (FM), incrementing a reference counter of the subpage, and if the reference counter meets or exceeds a threshold value, changing to swap mode for the subpage; for a subpage being in swap mode: upon entry into swap mode, swapping all access units of the subpage of the far memory (FM) with corresponding congruent access units of a congruent subpage in the near memory (NM), for all access units of the subpage, investigating compressibility of each pair formed by an access unit, or a sub-access unit thereof, of the subpage of the far memory (FM) and a corresponding congruent access unit, or a sub-access unit thereof, of a congruent subpage in the near memory (NM), wherein sufficient compressibility is when the pair of compressed access units, or sub-access units thereof, will fit in the congruent access unit, or sub-access unit thereof, of the congruent subpage in the near memory (NM); when a fraction of the pair of access units, or the sub-access units thereof, of the subpage exhibiting sufficient compressibility meets or exceeds a compression threshold, changing to cache / compress mode for the subpage; and for a subpage being in cache / compress mode: for each pair of access units, or sub-access units thereof, of the subpage that exhibits sufficient compressibility, accommodating the compressed access unit, or sub-access unit thereof, of the subpage of the far memory (FM) together with the corresponding compressed congruent access unit, or sub-access unit thereof, of the congruent subpage in the near memory (NM) in the congruent access unit, or sub-access unit thereof, of the congruent subpage in the near memory (NM).
7. The hybrid memory arrangement as defined in claim 6, wherein the hybrid memory controller (260) is configured for storing metadata for the handling of the three different modes of subpages of memory pages in the far memory (FM).
8. The hybrid memory arrangement as defined in claim 6 or 7, wherein metadata indicative of a size of the compressed access unit, or sub-access unit thereof, of the farmemory (FM) is encoded in unused control bits or unused parts of the congruent access unit, or sub-access unit thereof, of the congruent subpage in the near memory (NM).
9. The hybrid memory arrangement as defined in any preceding claim, wherein the near memory (NM) is High-Bandwidth Memory, HBM, and the far memory (FM) is Dynamic Random Access Memory, DRAM.
10. The hybrid memory arrangement as defined in any of claims 1-8, wherein the near memory (NM) is DRAM and the far memory (FM) is Non-Volatile Memory, NVM.
11. The hybrid memory arrangement as defined in any of claims 1-8, wherein the near memory (NM) is DRAM stacked on top of a microprocessor die and the far memory (FM) is DRAM off of the microprocessor die.
12. The hybrid memory arrangement as defined in any preceding claim, wherein the hybrid memory controller (260) is implemented on a microprocessor chip (210) which furthermore comprises one or more processing units (231-233), one or more levels of cache memory (C1-C3), said one or more near-memory controllers (241-242) and said one or more far-memory modules (253-254), the hybrid memory controller (260; 1160-1170) being operatively connected between a last level cache (220) of said one or more levels of cache memory (C1-C3) and said one or more near-memory controllers (241-242) and far-memory modules (253-254).
13. The hybrid memory arrangement as defined in any of claims 1-11, wherein the functionality of the hybrid memory controller is integrated with said one or more near- memory controllers (1141-1142).
14. A computer system (200; 1100) comprising one or more processing units (231- 233; 1131-1133), one or more levels of cache memory (C1-C3) and a hybrid memory arrangement, the hybrid memory arrangement providing a two-level main memory hierarchy in the computer system and comprising: a near memory (NM), comprising one or more near-memory modules (251-252; 1151-1152) controllable by one or more near-memory controllers (241-242; 1141-1142); a far memory (FM), comprising one or more far-memory modules (253-254; 1153- 1154) controllable by one or more far-memory controllers (251-252; 1151-1152); and a hybrid memory controller (260) configured for monitoring access requests to the far memory (FM) to identify bandwidth-demanding memory contents of the far memory (FM), compressing memory contents in the near memory (NM) to form a dynamic cache from memory space freed up by the compression, and accommodating at least some of the detected bandwidth-demanding memory contents of the far memory (FM) in the dynamic cache in the near memory (NM).
15. The computer system (200; 1100) as defined in claim 14, wherein metadata for said accommodating is stored in the far memory (FM) and wherein the total memory capacity exposed by the hybrid memory arrangement to an operating system (270; 1170) running on the computer system (200; 1100) is the total of the capacities of the near memory (NM) and the far memory (FM).
16. The computer system (200; 1100) as defined in claim 14 or 15, with the hybrid memory arrangement as defined in any of claims 2-13.
17. A method of providing a two-level main memory hierarchy in a computer system (200; 1100), the two-level main memory hierarchy comprising a near memory (NM) and a far memory (FM), together forming a hybrid memory (HM) in the computer system (200; 1100), the method involving: monitoring access requests to the far memory (FM) to identify bandwidth-demanding memory contents of the far memory (FM); compressing memory contents in the near memory (NM) to form a dynamic cache from memory space freed up by the compression; and accommodating at least some of the detected bandwidth-demanding memory contents of the far memory (FM) in the dynamic cache in the near memory (NM).
18. The method as defined in claim 17, wherein metadata for said accommodating is stored in the far memory (FM) and wherein the total memory capacity exposed by the hybrid memory arrangement to an operating system (270; 1170) running on the computer system (200; 1100) is the total of the capacities of the near memory (NM) and the far memory (FM).
19. The method as defined in claims 17 or 18, further comprising the functionality defined for the hybrid memory arrangement in any of claims 2-13.
20. A hybrid memory controller (260) for use in a hybrid memory arrangement that provides a two-level main memory hierarchy in a computer system (200; 1100), the two-level main memory hierarchy comprising a near memory (NM) and a far memory (FM), together forming a hybrid memory (HM) in the computer system (200; 1100), the hybrid memory controller (260) being configured for: monitoring access requests to the far memory (FM) to identify bandwidth-demanding memory contents of the far memory (FM), compressing memory contents in the near memory (NM) to form a dynamic cache from memory space freed up by the compression, and accommodating at least some of the detected bandwidth-demanding memory contents of the far memory (FM) in the dynamic cache in the near memory (NM).
21. The hybrid memory controller (260), wherein metadata for said accommodating is stored in the far memory (FM) and wherein the total memory capacity exposed by the hybridmemory arrangement to an operating system (270; 1170) running on the computer system (200; 1100) is the total of the capacities of the near memory (NM) and the far memory (FM).
22. The hybrid memory controller (260) as defined in claim 20 or 21, further configured as the hybrid memory controller in the hybrid memory arrangement according to any of claims 2-13.
Citation Information
Patent Citations
Bus attached compressed random access memory
US20090254705A1
Intelligent far memory bandwith scaling
US20140092678A1
Hybrid multi-level memory architecture
US20150006805A1
Dynamic memory expansion by data compression
US20170004069A1
Two-level system main memory
US20180004432A1