Memory management for compressed memories

The compressed memory manager decouples fragment allocation and deallocation, using a fragmented memory layout with pre-allocated fragments of varying sizes, to address the challenges of managing variable-sized memory pages and reducing fragmentation in compressed memory systems, resulting in improved memory utilization and stability.

WO2025136191A1PCT designated stage expired Publication Date: 2025-06-26ZEROPOINT TECH AB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2024/051075
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-20
Filing Date
2024-12-13
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing compressed memory systems face challenges in efficiently managing variable-sized memory pages and locating compressed data due to internal and external fragmentation, which leads to low memory utilization and stability issues.

Method used

The proposed solution involves a compressed memory manager that decouples the allocation and deallocation of free memory fragments from the actual utilization, using a fragmented memory layout with pre-allocated fragments of varying sizes, and employing a Two-Level Segregated Fit memory allocator to minimize fragmentation.

Benefits of technology

This approach significantly accelerates compression and decompression operations while maintaining high memory optimization, reducing internal and external fragmentation, and improving memory utilization without stability issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2024051075_26062025_PF_FP_ABST
    Figure SE2024051075_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A compressed memory management method for a computer system having a computer subsystem, a memory subsystem and a memory compression arrangement. The compressed memory management method involves the following. Managing a compressed memory layout which divides the physical address space of the memory subsystem into a plurality of fragments of memory. Pre-allocating free fragments in the compressed memory layout. Receiving and handling fragment allocation requests from the memory compression arrangement by providing one or more of the pre-allocated free fragments for use by the memory compression arrangement to accommodate pieces of compressed memory. Receiving and handling fragment release requests pertaining to one or more used fragments by providing at least some of said one or more used fragments for deallocation in the compressed memory layout. The managing of the compressed memory and pre allocating of fragments are decoupled from the memory compression arrangement.
Need to check novelty before this filing date? Find Prior Art

Description

MEMORY MANAGEMENT FOR COMPRESSED MEMORIESTECHNICAL FIELD

[0001] The present disclosure generally relates to the field of data compression in memories in electronic computers. More particularly, the present disclosure relates to a compressed memory manager for use in a computer system, to a computer system comprising such a compressed memory manager, and to an associated compressed memory management method.BACKGROUND

[0002] Data compression is a general technique to store and transfer data more efficiently by coding frequent collections of data more densely than less frequent collections of data. It is of interest to generally store and transfer data more efficiently for a number of reasons. In computer memories, for example memories that keep data and computer instructions that processing devices operate on, for example main memory or cache memories, it is of interest to store said data more efficiently, say K times, as it then can reduce the size of said memories potentially by K times, using potentially K times less communication capacity to transfer data between one memory to another memory and with potentially K times less energy expenditure to store and transfer said data inside or between computer systems and / or between memories. Alternatively, one can potentially store K times more data in available computer memory than without data compression. This can be of interest to achieve potentially K times higher performance of a computer without having to add more memory, which can be costly or can simply be less desirable due to resource constraints. As another example, the size and weight of a smartphone, a tablet, a lap / desktop or a set-top box are limited as a larger or heavier smartphone, tablet, lap / desktop or set-top box could be of less value for an end user; hence potentially lowering the market value of such products. Yet, more memory capacity can potentially increase the market value of the product as more memory capacity can result in higher performance and hence better utility of the product.

[0003] To summarize, in the general landscape of computerized products, including isolated devices or interconnected ones, data compression can potentially increase the performance, reduce the energy expenditure or reduce the cost and area consumed by memory. Therefore, data compression has a broad utility in a wide range of computerized products beyond those mentioned here.

[0004] Compressed memory systems in prior art typically compress a memory page when it is created, either by reading it from disk or through memory allocation. Compression can be done using a variety of well-known methods by software routines or by hardware accelerators. When the processors request data from memory, data must be first decompressed before serving the requesting processor. As such requests end up on the criticalmemory access path, decompression is typically hardware accelerated to impose minimal impact on the memory access time.

[0005] One problem with compressed memory systems is that variable-sized memory pages must be managed. Unlike in a conventional memory system, without compression, where all pages have the same size, the size of each page in a compressed memory system depends on the compressibility of the page which is highly data dependent. Moreover, the size of each page can vary during its lifetime in main memory owing to modifications of the data in the page which may affect its compressibility. For this reason, a fundamental problem is how to dynamically manage the amount of memory freed up through compression.

[0006] Prior art of managing the freed-up memory space in compressed memory systems takes a variety of approaches. Typically, memory management is an integral part of the operating system and more specifically the virtual memory management (VMM) routines. The VMM is tasked with managing the available physical main memory space efficiently. At boot time, the VMM typically establishes how much physical main memory is available. Based on this, it can decide which pages should be resident in the available physical memory space on a demand basis; when required, the VMM typically selects as a victim a page that is less likely to be accessed soon to be paged out to disk. In a compressed memory, the amount of available memory capacity will however change dynamically. For this reason, conventional VMM routines do not work as they assume a fixed amount of physical memory at boot time.

[0007] An alternative is to modify the VMM to be aware of the amount of memory available at any point in time. This is however highly undesirable as changes to operating systems must apply to computer systems whether or not they use compression. Instead, several approaches in prior art solve the management problem in a transparent way to the operating system. An alternative approach is to let the operating system boot while assuming a larger amount of physical memory than what is available. If the amount of memory freed up by compression is less than that amount, pages are swapped out to disk to lower the memory pressure. This solution may however suffer from stability problems due to thrashing and may not be desirable.

[0008] Clearly, systems and methods are needed to manage free space in a compressed memory that are transparent to operating systems and that allow all memory to be compressed without suffering from stability problems. Another problem related to compressed memory systems is how to efficiently locate an item in the compressed memory. In a compressed memory, each individual cache line or memory block (henceforth, blocks and lines are used interchangeably), like memory pages, has variable sizes. In addition, variable-sized compressed data must be compacted, which may mean that compressed data items are packed together. This may change the memory layout. Hence, memory pages and blocks are not necessarily placed in the original locations where they would be placed in a conventional uncompressed memory system. To locate a compressed block, an additionallevel of address translation is typically needed between conventional physical memory addresses and compressed memory addresses.

[0009] One approach in prior art to orchestrate address translation between physical and compressed addresses partitions the physical memory into fixed-size units, referred to as segments. For example, a 4-KB memory page can be partitioned into four segments making the segment size 1 KB. Now, if a page is compressed by a factor of two, it would need two segments instead of four as in a conventional uncompressed memory system. To locate a memory block using the concept of segments, metadata is needed as part of the address translation to locate the segment in which the requested memory block resides. A large segment has the advantage of reducing the amount of metadata needed to locate the segment but has the disadvantage of leading to more internal fragmentation than a smaller segment and hence lower utilization of the memory. This is because if a 4-KB page is compressed to 1280 bytes, in the previous segmentation approach, it would take 2KB in space, underutilizing 768B (=2048B-1280B). In addition, this scheme could suffer from external fragmentation; a previously compressed block can change in size due to subsequent writes causing said block to shrink in size needing fewer segments than originally allocated. This can lead to poor utilization of the created segments due to fragmentation that cannot be addressed unless the memory is recompacted. Hence, this approach either suffers from a large amount of metadata or poor utilization of the memory.

[0010] Another approach in prior art for the aforementioned embodiment of compressing 4- KB pages is to use pools of different page sizes, segmenting in this way different parts of the memory with different fixed-size segment sizes. For example, in one example the memory is divided into two regions; one region is segmented into 1-KB segments and another region into 2-KB segments. This addresses the issue of external fragmentation, as groups of pages compressed by the same size are allocated utilizing segments of a similar size. If there is a significant change in size so that the compressed page cannot fit into that allocated segment or is much larger, it is moved to another part of the memory with a more appropriate segment size that matches the new compressed size. This scheme alleviates external fragmentation at the cost of data migration. However, because of the grouping, the metadata size can be potentially reduced. Importantly, this approach still suffers from internal fragmentation; this is because if a page after compression has a size of 1050 bytes, it cannot allocate a segment of 1KB (1024 bytes), so it is enforced to utilize a segment of 2KB (2048 bytes), leaving 998 B (bytes) unutilized for the lifetime of the page until it is modified again. Another disadvantage of this approach is the inflexibility of having a fixed number of segments of a certain size. This can make internal fragmentation more severe if the compressed memory runs out of a certain segment size and thus future compressed pages need to rely on larger segment sizes. As an alternative, the larger segment can be split into smaller segments to tackle the internal fragmentation issue by further complicating the management and metadata layout. The coarse segment can be repurposed by dividing it into smaller segments while being part of the coarse segment memory area.

[0011] Memory systems typically manage data at the granularity of memory blocks. Since memory blocks have variable sizes in a compressed memory system, there is an additional problem of locating them within a compressed memory page. One approach in prior art is to use pointers to locate the blocks inside a memory page or memory pages within the compressed memory. For example, if the page size is 4 KB, a 12-bit pointer can locate any block inside a page at the byte granularity. Assuming a block size of 64 B, there are 64 blocks in a 4-KB memory page. Hence, as many as 64 x 12 bits = 768 bits of metadata per memory page is needed per memory page. Similarly, locating compressed pages within the segmented memory, pointers to the allocated segment(s) must be associated with the address of the compressed page. Said pointers to segments are addresses address space of the compressed memory and their width is memory size (in bytes) / segment size. For example, for a 1-TiB memory segmented in IKiB segments, the pointer must be wide enough to address a memory of 230segments requiring a 30-b segment pointer.

[0012] Another approach in prior art associates with each block the size of that block. For example, if 64 different sizes are accommodated, the metadata per memory block to encode the size would need 6 bits. In this case, the amount of metadata associated with a memory page to locate each block is 64 x 6 = 384 which is substantially less than using pointers. However, to locate a block, for example the last block, would require to sum all the sizes of all the blocks prior to that block in the memory page which can cause a substantial latency in the address translation process. To shorten that latency, one can restrict the size of compressed blocks. However, this can potentially reduce the utilization of the compressed memory. For example, let us suppose that four blocks are compressed by a factor two, four, eight and sixteen. With the pointer-based approach one would enjoy the full potential of compression leading to a compression factor of 4 / (l / 2 + 1 / 4 +1 / 8 + 1 / 16) = 4.3, whereas restricting the size to four would yield a compression factor of 4 / (l / 2 + 1 / 4 + 1 / 4 + 1 / 4) = 3.2.

[0013] The metadata needed to locate a memory page and / or a memory block in a compressed memory system is typically cached on the processor chip similarly with what a translation-lookaside-buffer does for virtual-to-physical address translation in conventional computer systems to avoid having to access memory to locate a memory block in the physical memory. Since such address translation mechanisms must be fast and hence are of limited size, they can typically not keep all the metadata needed. Hence, all metadata needs to be placed in memory (or storage). There is a need to limit the amount of metadata to save memory space and to cut down on the extra memory accesses needed to bring metadata from memory to the address translation mechanism on the processor chip.

[0014] Hence, compressed memory management approaches can suffer from internal and external fragmentation when they attempt to utilize the compressed memory space. This results in low utilization of the memory. Systems, devices and methods are clearly needed to manage the allocated and free space in a compressed memory substantially more efficiently than approaches known from prior art.

[0015] Furthermore, systems, devices and methods are needed to locate data in the modified physical address space or in an extra compressed address space more efficiently than approaches known in prior art.SUMMARY

[0016] It is an object of the present invention to offer improvements in the field of data compression in memories in electronic computers, and to solve, eliminate or mitigate one or more of the problems referred to above.

[0017] As will be seen, the disclosed invention provides compressed memory management that in effect decouples the allocation of free memory space from the actual utilization of the space. The decoupling allows to significantly accelerate the compression and decompression operations whilst the actual compressed memory management remains highly optimized. The disclosed invention is applicable to disaggregated memory systems in which a plurality of compute nodes with processing elements and local memories share a common pool of memory. This is referred to as disaggregated memory systems. The disclosed invention is also applicable to directly attached memory systems in which a compute node with processing elements (e.g. a multi- or many-core microprocessor with additional accelerators) shared on- or off-chip memory.

[0018] A first aspect of the present invention is a compressed memory manager for use in a computer system having a computer subsystem, a memory subsystem and a memory compression arrangement. As such, the memory compression arrangement comprises a compressor and a decompressor for compressing and decompressing, respectively, memory contents in a physical address space of the memory subsystem into a compressed address space used by the computer subsystem. The compressed memory manager itself is configured for managing a compressed memory layout which divides the physical address space of the memory subsystem into a plurality of fragments of memory, the fragments being of two or more different sizes.

[0019] Moreover, the compressed memory manager is configured for pre-allocating free fragments in the compressed memory layout.

[0020] In addition, the compressed memory manager is configured for receiving and handling fragment allocation requests from the memory compression arrangement by providing one or more of the pre-allocated free fragments for use by the memory compression arrangement to accommodate pieces of compressed memory.

[0021] The compressed memory manager is further configured for receiving and handling fragment release requests pertaining to one or more used fragments by providing at least some of said one or more used fragments for deallocation in the compressed memory layout.

[0022] Pursuant to the invention, the managing of the compressed memory layout, the preallocating of free fragments and the deallocation of used fragments in the compressed memory layout are performed decoupled from and asynchronously with the receiving and handling of the fragment allocation requests and fragment release requests from the memory compression arrangement.

[0023] A second aspect of the present invention is a computer system comprising a computer subsystem, a memory subsystem and a memory compression arrangement that comprises a compressor and a decompressor for compressing and decompressing, respectively, memory contents in a physical address space of the memory subsystem into a compressed address space used by the computer subsystem. The computer system according to the second aspect of the present invention furthermore comprises a compressed memory manager according to the aforementioned first aspect of the present invention.

[0024] A third aspect of the present invention is a compressed memory management method for a computer system having a computer subsystem, a memory subsystem and a memory compression arrangement. The memory compression arrangement comprises a compressor and a decompressor for compressing and decompressing, respectively, memory contents in a physical address space of the memory subsystem into a compressed address space used by the computer subsystem.

[0025] The compressed memory management method according to the third aspect of the present invention comprises the following:• managing a compressed memory layout which divides the physical address space of the memory subsystem into a plurality of fragments of memory, the fragments being of two or more different sizes;• pre-allocating free fragments in the compressed memory layout;• receiving and handling fragment allocation requests from the memory compression arrangement by providing one or more of the pre-allocated free fragments for use by the memory compression arrangement to accommodate pieces of compressed memory; and• receiving and handling fragment release requests pertaining to one or more used fragments by providing at least some of said one or more used fragments for deallocation in the compressed memory layout.

[0026] According to the method, the managing of the compressed memory layout, the preallocating of free fragments and the deallocating of used fragments in the compressed memory layout are performed decoupled from and asynchronously with the receiving and handling of the fragment allocation requests and fragment release requests from the memory compression arrangement.

[0027] Further aspects, features and advantages of the present invention and embodiments thereof will appear from the following technical description, from the drawings as well from the patent claims.BRIEF DESCRIPTION OF DRAWINGS

[0028] FIG. 1 A depicts an example of a conventional system with a host CPU subsystem connected to a memory subsystem which is further connected to a MEDIA subsystem.

[0029] FIG. IB depicts an example of a conventional system with disaggregated compute and memory resources, comprising an exemplary CPU subsystem (Host) connected to a memory subsystem over a CXL-based communication interface. The memory subsystem is further connected to a MEDIA subsystem.

[0030] FIG. 1C depicts an example of a conventional system of FIG. 1 A with compression and decompression.

[0031] FIG. ID depicts an example of a conventional system of FIG. 1 A with compression and decompression inside the memory subsystem.

[0032] FIG. IE depicts an example of the system of FIG. IB with compression and decompression.

[0033] FIG. IF depicts an example of the system of FIG. IB with compression and decompression inside the disaggregated memory subsystem.FIG. 2 depicts an example of a state of art memory compression device.

[0034] FIG. 3 depicts an example of a buddy allocator tree.

[0035] FIG. 4 depicts an embodiment of the present invention, based on the memory compression device of FIG. 2 and with a compressed memory management (CMM) device as disclosed in this patent application.

[0036] FIG. 5 depicts an embodiment of the memory compression device of FIG. 4 wherein the CMM mechanism is partitioned to a CMM-FGR (CMM Fragment Generator and Replenisher unit) and a CMM-FQA (CMM Fragment Quick Access unit).

[0037] FIG. 6 illustrates a fragment ID, consisting of one address pointer field and one order field.

[0038] FIG. 7 depicts an exemplary communication sequence between the two entities of the CMM: CMM-FGR and CMM-FQA.

[0039] FIG. 8 depicts an embodiment of the CMM-FQA, consisting of one DST ID assembler, one DST ID disassembler, numerous fragment queues, and one recycle queue.

[0040] FIG. 9 illustrates a DST ID assembler table.

[0041] FIG. 10 illustrates a DST ID assembler control flow.

[0042] FIG. 11 illustrates CMM-FQA queues.

[0043] FIG. 12 illustrates watermarks of a fragment queue.

[0044] FIG. 13 illustrates a method to replenish a fragment queue.

[0045] FIG. 14 illustrates watermarks of recycle queues.

[0046] FIG. 15 illustrates a method to drain the recycle queue.

[0047] FIG. 16 depicts a memory compression device that shows the flow of metadata generation and bookkeeping in an embodiment of a memory compression system.

[0048] FIG. 17 depicts a memory compression device and shows the flow of how a memory compression device responds to a read request by decompressing a previously compressed page, including fetching the metadata to locate the compressed fragment in memory.

[0049] FIG. 18 depicts an embodiment of a baseline system with compression and memory management, comprising an exemplary CPU subsystem (Host) connected to a memory subsystem, which is further connected to a MEDIA subsystem which contains compressed data managed by the compressed memory manager disclosed in this patent application.

[0050] FIG. 19 depicts an embodiment of a baseline system with disaggregated compute and memory resources employing compression and memory management, comprising an exemplary CPU subsystem (Host) connected to a memory subsystem over a CXL-based communication interface. The memory subsystem is further connected to a MEDIA subsystem which contains compressed data managed by the compressed memory manager disclosed in this patent application.

[0051] FIG. 20 is a flowchart to illustrate a compressed memory management method generally according to the present invention.TECHNICAL DESCRIPTION

[0052] An example of a conventional computer system is depicted in FIG. 1 A. It comprises a computer subsystem (referred to as CPU subsystem but is not limited to CPUs only) 110, a memory subsystem 140, MEDIA 150 wherein the computer and memory subsystems are considered as part of the same system, for example the same system-on-chip (SoC) or the same application specific integrated circuit (ASIC). The computer subsystem runs applications 114 and operating systems and / or hypervisors (HV) 118, or other types of software or firmware. All or part of said types of software programs may need to communicate to the memory subsystem to potentially fetch data and instructions from the MEDIA 150 and to write data to the MEDIA 150. Despite not being shown in the figure, the CPU subsystem 110 usually comprises one or a plurality of cores, a cache hierarchy, TLBs and memory management units (MMUs) to be able to locate data and instructions in thetarget MEDIA 150. MEDIA 150 corresponds to one or a plurality of MEDIA modules, for example, DRAM modules.

[0053] A second example of a conventional computer system is depicted in FIG. IB. It comprises a computer subsystem 110, a memory subsystem 140, MEDIA 150 wherein the computer and memory subsystems are disaggregated. Disaggregation means that the computer subsystem 110 is implemented in one SoC and the memory subsystem 140 is implemented in a different SoC. In this exemplary embodiment the disaggregated subsystems can be connected via Compute Express Link (CXL)-based channels: CXL.io 120 and CXL.mem 130, wherein the former is used for control and the latter is used for data transfers. CXL is an open standard for high-speed device to device (e.g., CPU to memory) interconnection. In an alternative embodiment, the interconnection of the disaggregated systems could be done with Ethernet-based channels. In yet another embodiment, the interconnection could be done with a proprietary interface. The present invention is not limited to systems with specific interconnects or interfaces when it comes to integrating aggregated or disaggregated computer systems.

[0054] By MEDIA the present invention relates to any memory technology incorporated in a computer system, such as DRAM, Flash, SRAM, MRAM, HBM, PCM, STT-RAM and in general volatile or non-volatile memory. The terms of memory as well as the more general term Media (or MEDIA) are used interchangeably.

[0055] By computer system, the present disclosure refers to different types of computer systems such as aggregated or segregated compute and memory (sub-)systems, where the CPU, SRAM, DRAM and other media may reside in the chip; disaggregated compute and memory (sub-)systems, where the CPU and other Media may reside in different chips, or different racks in a server system or even different cabinets in a datacenter system. By compressed address space (CAS) or compressed memory space (CMS), the present disclosure refers to the address space of the memory that stores compressed data in the memory compression system. By computer subsystem, the present disclosure refers to any subsystem performing some form of computation, for example, CPU (central processing unit), GPU (graphics processing unit), NPU (neural processing unit), DPU (data processing unit), or any type of application or domain specific accelerator, etc.

[0056] An example of disaggregated systems that the described invention can be applied to is for Compute-Express Link (CXL) disaggregated memory systems. However, the scope is not limited to the CXL based protocol and can be expanded to other memory systems that enable disaggregated memory systems.

[0057] Computer systems, as exemplified by FIG. 1 A and IB, can suffer from a limited capacity of the memory 150. A limited memory capacity can manifest itself in more memory requests that will have to be serviced at the next level of the memory hierarchy typically realized as the storage level of the memory hierarchy. Such storage-level accesses are, in general, considerably slower and may result in a considerable loss in performance and energyefficiency. Increasing the memory capacity can mitigate these drawbacks. However, more memory capacity can increase the cost of the computer system both at the component level (CAPEX) or in terms of energy expenditure (OPEX). In addition, more memory may consume more space, which may limit the utility of the computer system, in particular in form-factor constrained products including for example mobile computers such as tablets, smart phones, wearables and small computerized devices connected to the Internet (also known as loT (Internet-of-Things) devices. Finally, the bandwidth provided by the memory modules may not be sufficient to respond to requests from the processing units. This may manifest itself in lower performance.

[0058] To address particularly the problem of a limited MEDIA capacity and limited MEDIA bandwidth, the exemplary systems of FIG. 1 A and IB can be modified to allow data and instructions to be compressed in the MEDIA. In a first example, the compression can happen in the memory subsystem by using a compression / decompression device 142 (see “Comp / Dcmp” at 142 in FIG. ID and FIG. IF) integrated in the memory subsystem. This way data received by the memory subsystem 140 is compressed before being transmitted and stored in the MEDIA 150. When MEDIA data is requested by the CPU subsystem, the memory subsystem requests data from the MEDIA, decompresses the compressed data (by the compression / decompression device 142) and returns the data in uncompressed form to the CPU subsystem. In an alternative embodiment, the compression can occur by a standalone compression / decompression mechanism 120 (see “Comp / Dcmp” in FIG. 1C and FIG. IE), wherein data arrives at the memory subsystem in compressed form. Similarly, decompression is executed on the compression / decompression mechanism 120 after it is provided by the memory subsystem with compressed data read from the MEDIA. In yet another example, the compression / decompression mechanism can be integrated and be part of the CPU subsystem, for example by being integrated alongside the memory controllers, e.g., MEM CTRL 146 of FIG. 1A.

[0059] No matter where the compression / decompression operations occur in a system employing memory compression, expanding the memory capacity with compressed data requires a mechanism that manages the compressed memory. The compressed memory manager is responsible for keeping track of the memory areas currently being used and areas being free, and for providing to the compressor / decompressor mechanism with the information needed to place the compressed data in memory, and to retrieve the compressed data from memory. The compressed memory management is different from the conventional memory management done for example by the operating system (OS), because the compressed memory is of variable size due to the variable sizes of compressed data.tae state of the art — roblem statement

[0060] FIG. 2 depicts an example of a state-of-art memory compression device 200. It comprises a Compressed Memory Manager (CMM) 210, a memory compression arrangement 201 that includes a compressor 220, a remapper 230, a recompactor 240 and a decompressor 260, and MEDIA 270. The MEDIA unit 270 represents the actual memorystorage wherein data is stored in compressed form. The compressor 220 compresses data by implementing any data compression algorithm. The compressor can be specialized to compress blocks (for example, DRAM blocks or cache lines), pages (aligned to operating system pages) or blocks of other granularities (for example, 1-KB or 16-KB blocks or blocks on the order of MBs or GBs of data). The present invention is not limited to specific compression algorithms or specific compression granularities. The decompressor 260 decompresses data that was previously compressed by compressor 220. The remapper 230 stores metadata related to where a compressed block or page is located in the compressed MEDIA and uses it to locate said compressed block / page upon a read request.

[0061] FIG. 2 also depicts the CPU subsystem 250 connected to the compressor 220. The CPU subsystem 250 corresponds to the CPU subsystem 110 as depicted in FIGS 1 A-1F. The CMM 210 is further connected to the compressor 220, the remapper 230, the recompactor 240, and potentially the decompressor 260 and the CPU subsystem 250. The compressor 220 receives data from the CPU subsystem 250 and performs compression to form compressed data blocks or compressed data pages. In order to determine where in the memory it can send the compressed block / page, it requires the location metadata of a free memory area that can fit this compressed block / page. This information is provided by the CMM 210 in the form of fragments. We will provide a definition of the concept fragment later. This is usually done by providing the size of the compressed block / page to the CMM, so that the CMM can allocate a fragment and provide its metadata back to the compressor. The transmission of size is reflected by the arrow with label “size” in FIG. 2, whilst the fragment allocation is reflected by the arrow “Alloc fragments” from 210 to 220 in FIG. 2. The compressed memory is allocated and managed in a number of memory chunks called fragments. The fragment definition is further expanded below.

[0062] Similarly, specific fragments can become unused. In one example wherein read data is considered obsolete, memory fragments can become unused after the decompressor 260 decompresses a specific block or page to respond to a read request. In a second example, the CPU subsystem 250 can decide at any point in time to free up memory fragments that are not going to be used again. In a third example, the recompactor 240 can recompact part of the compressed memory releasing freed-up memory fragments. Recompaction is defined as a mechanism to reorganize compressed data in such a way so that contiguous free space is accumulated out of smaller pieces of free space that exists between various compressed data chunks. In yet another embodiment, the remapper 230 which keeps metadata information about the mapping to the compressed memory space may have its own heuristics and can decide that a specific group of fragments is not used anymore. When specific fragments become unused, they need to be deallocated. A deallocation fragment request is reflected with the “dealloc fragment” arrows in FIG. 2.

[0063] Fragments can be fixed-size segments'. Prior approaches that are described in paragraphs

[0009] -

[0011] take the design-time decision to segment the memory in fixed- size segments (e.g., 1KB) and allocating a variable number of them to store a compressedpage based on its compressed size. This causes high internal and external fragmentation. This is undesirable.

[0064] Fragments can be collected out of pools of fixed-size segments'. Prior approaches that are described in paragraph

[0010] take this approach and try to decide the mapping and locations of the compressed memory data by allocating a segment belonging to a specific pool (e.g., a region with 1KB segments, another region with 2KB segments, etc.) based on the size of the compressed block or page. This causes high internal fragmentation. This is undesirable.

[0065] As an alternative, ^ragmen A can be of variable size. This can happen if the compressed memory management implements a data structure such as a buddy allocator tree, which minimizes external fragmentation at static metadata cost. A buddy allocator divides the memory into fragments by always splitting one memory fragment into two smaller, equally- sized, fragments called buddies. If two buddies are free, the buddy allocator immediately coalesces them into one contiguous fragment. The allocation scheme is implemented by a perfect binary tree (a buddy allocator tree) where allocation and deallocation is made inOfieightOfTree') which is O(log2(totalMemorySizey). External fragmentation is minimized because the buddy allocator prioritizes the most fragmented subtree. Moreover, the tree is automatically defragmented since a freed up fragment is automatically merged with free fragments (creating larger fragments) during a deallocation. The metadata cost is static since the size of the buddy allocator tree is static.

[0066] An exemplary illustration of a buddy allocator tree is depicted in Figure 3. 300, 330, 370, and 380 form a perfect binary tree. Note that the tree nodes between 512GB and 8192B are not shown in the figure; they are abstracted by 330. In the example, the entire memory space spans 1TB (node 303). Each parent node is split into two child nodes which represent a smaller memory fragment. For example, 371 represents 8192 bytes of data, and its children (372 and 373) represent 4096 bytes of data. The leaf nodes of the binary tree either represent free data fragments (373, 375, and 389) or allocated fragments (385 and 378). The bottompart of the figure (390) illustrates how the leaf nodes of the binary tree map to space in the memory creating memory fragments to be used by the memory compression system. For example, the allocated node 385 is of size 512 bytes and maps to memory fragment 392 in the memory. In a similar manner, the free node 389 maps to a memory fragment 394 in the memory.

[0067] Although the buddy allocator tree algorithm seems to minimize the fragmentation at a static metadata cost, traversing the tree to find the most suitable fragment everytime the compressor needs a fragment to compress a block or page, is unfortunately a slow operation. This can significantly slowdown memory compression which usually needs to operate on the order of a few nanoseconds (ns). This is undesirable.

[0068] Instead of a buddy allocator tree, the chosen data structure to implement the compressed memory management of FIG. 2 can be a Two-Level Segregated Fit memoryallocator (TLSF) which minimizes internal fragmentation and makes the metadata cost dynamic, i.e. depending on how much data is allocated. The TLSF data structure comprises two levels of arrays which keep lists of free memory fragments. The first-level array divides free memory fragments in classes that are powers of two apart, and the second level divides each first-level class linearly. Each array of lists marks whether it contains free memory fragments or not. Fixed-size pointers to the free memory fragments are stored within the memory fragments themselves. Since the amount of required metadata varies with the number of allocated fragments and the size of the allocated fragments, the metadata cost is dynamic. Similar to the buddy allocator, traversing the TLSF lists to locate a suitable fragment for compressing a page can lead to high latencies making memory compression impractical. This is undesirable.

[0069] The methods, devices and systems of the current patent disclosure teach how to decouple the allocation of free space from the actual utilization of the space and vice versa. This way, compressed memory management does not need to be restricted to be implemented with rather simple mechanisms that suffer from high fragmentation. The current patent disclosure demonstrates how to use more complex allocation strategies (such as the buddy allocator tree) without being on the critical path. Henceforth, although allocation or deallocation can take significantly more time to traverse said complex data structures, the compressed memory management will not hold the compression and decompression operations whilst the actual compressed memory management remains highly optimized.Introduction to CMM device of the present disclosure

[0070] The compressed memory management (CMM) methods and devices disclosed in the present invention comprise features related to the allocation method of free space to be used by compressed blocks, deallocation of free space that is released out of any compressed blocks, variable size compressed memory layout management including tracking and bookkeeping the reserved and free memory chunks, as well as various maintenance operations such as defragmentation.

[0071] FIG. 4 depicts an embodiment of a memory compression device 400, being based on the memory compression device 200 of FIG. 2 but modified and improved to implement a compressed memory management method and device as disclosed in the present patent application. Similar to FIG. 2, a compressor 420 and a decompressor 460 are part of a memory compression arrangement 401. The memory compression device 400 comprises a CMM device 410 that further comprises an allocator unit 412 (referred to as “allocator” in the following), a deallocator unit 414 (“deallocator”), a memory layout management unit 416 (“layout manager”), and a maintenance controller unit 418 (“maintenance controller”).

[0072] The layout manager 416 builds and continuously maintains the compressed memory layout of the compressed memory. The layout manager 416 keeps the memory layout in a data structure and bookkeeps it in terms of fragments. A fragment follows the same definition as defined earlier in paragraphs

[0063] -

[0065] , In one embodiment, said data structure canbe a buddy tree data structure as described in paragraphs

[0065] -

[0066] , In an alternative embodiment, this data structure can be a TLSF data structure as described in paragraph

[0068] , Alternative data structures can be implemented by those skilled in the art.

[0073] The allocator 412 and deallocator 414 can be organized in a producer-consumer manner (412a and 412b as well as 414a and 414b). This can facilitate the utilization and the efficiency of the memory compression system by decoupling the creation of fragments from the use of those and similarly decoupling the release of fragments from the actual recycling (or deallocation), i.e., addition back to the layout manager 416 data structure. In the case of an allocation operation, the producer 412a can traverse the data structures of the layout manager 416, retrieve possible free fragments and create a list of new free fragments. The list of free fragments is consumed (412b) when the compressor 420 receives a task to compress a block or page. This is reflected with the arrow labeled as New fragments in FIG. 4. In a similar way, in the case of a deallocation operation, the producer 414b collects the released fragments by one or a plurality of requestors (420, 430, 440, 450, 460) and creates a list of old free fragments. This is consumed (414a) when the deallocator updates the compressed memory layout. Designing the allocator and / or the deallocator into a producer-consumer model can result in significant benefits, as disclosed in paragraph

[0077] and onwards.

[0074] As mentioned above, in the disclosed embodiment the CMM unit 410 contains a maintenance controller 418. One example of a maintenance operation by the maintenance controller 418 is to defragment the compressed memory, which is detailed in paragraphs

[0135] -

[0139] below. Defragmentation usually requires recompacting (part of) the compressed memory to create more contiguous free space. This can be assisted by the recompactor 440, which also provides size to the allocator 412 and receives back fragments from it (arrow with “size” and “fragments” between 412 and 440 in FIG. 4).

[0075] When it comes to storing and maintaining the fragments utilized by the compressor, in the exemplary embodiment of FIG. 4, the actual population and insertion of metadata to the remapper is done by the compressor 420 (arrow with “insert <fragment_id> metadata”). Metadata is described in more detail in paragraphs

[0084] -

[0087] ,

[0076] The following section will describe how fragments are organized, maintained, used and combined by the compression device. As disclosed in paragraph

[0094] and onwards, the allocator can provide one or a plurality of fragments to the compressor for handling the storage of specific compressed memory block / page.CMM device as a combination of FGR and FQA

[0077] Taking the CMM device 410 of FIG. 4 into a step further, the CMM allocator’s producer 412a, the CMM deallocator’s consumer 414a, the CMM layout manager 416 and the CMM maintenance controller 418 can be organized together forming a Fragment Generator and Replenisher (FGR) mechanism, while the CMM allocator consumer 412b and the CMM deallocator producer 414b can be organized together forming a Fragment Quick Access (FQA) mechanism. FIG. 5 depicts this exemplary embodiment wherein the CMMallocator 412 / 512 is divided between the CMM Fragment Generator and Replenisher 580 and the CMM Fragment Quick Access unit 590 into two parts 582 and 592. Similarly the CMM deallocator 414 / 514 is also divided accordingly, into two parts 584 and 594. Hence, the CMM FGR mechanism can be implemented in a way that maintains flexibility and minimizes complexity, whereas the CMM FQA mechanism can be implemented in a way that operates on the same speed as the rest of the memory compression device, the CPU subsystem or the memory subsystem depending on where in the system, it is integrated into. Again, compressor 520 and decompressor 560 are part of a memory compression arrangement 501.

[0078] In an alternative embodiment of FIG. 5, the CMM FQA mechanism can be instantiated as a device implemented in hardware in the same form and speed as the rest of the memory compression device, whilst the CMM FGR mechanism can be implemented as a software (or firmware) method running on a processing unit in the system. In yet an alternative embodiment of FIG. 5, the CMM FQA mechanism can be instantiated as a device implemented in hardware in the same form and speed as the rest of the memory compression device, whilst the CMM FGR mechanism can be implemented in programmable hardware (e.g., FPGA). In either of these embodiments, this design choice can result in significant benefits because the software-implemented method or the programmable device can allow to retain a lot of flexibility when it comes to the implementation of the memory compression layout construction, management, allocation, deallocation and maintenance algorithms, whilst the list of new and old free fragments are implemented in fast hardware significantly accelerating the actual memory compression. The present invention is not limited to the disclosed embodiments, other similar embodiments can be realized depending on the various performance / area / power / design-complexity / design-cost / other trade-offs, by those skilled in the art. For example, both the CMM FQA and CMM FGR can be implemented in hardware by those skilled in the art, in the same or different clock domains depending on speed requirements.

[0079] The CMM FGR mechanism guides the memory compression system into which areas of the memory it should use. This happens asynchronously and periodically by the CMM FGR communicating to the CMM FQA and providing it with new free fragments or receiving the old free fragments, by utilizing an interface in the compressed memory system. In the exemplary embodiment wherein the CMM FGR method is implemented in software and the CMM FQA mechanism is implemented in hardware, the interface is a SW7HW interface implemented with memory mapped registers. In an alternative embodiment, the SW / HW interface can be implemented with specific commands. In yet another embodiment, the communication between the CMM FGR and CMM FQA can happen with shared memory or other state of art means existing in a computing system. In all such embodiments, the communication is done in a proactive way. In an alternative embodiment, the communication could be done reactively. In such a case, the CMM FGR provides new free fragments to the CMM FQA. Then CMM FQA maintains watermarks that indicate the amount of free fragments used by the memory compression device. Thus, by relying on interrupts that are triggered by the CMM FQA when certain watermarks are crossed the CMM FGR can providemore new free fragments. In other embodiments a combination of proactive and reactive communication mechanisms can be also used.

[0080] The CMM memory allocation algorithm operates as follows: in the CMM FGR mechanism, the allocator traverses the compressed memory layout to allocate free fragments to be able to store compressed pages. Information about these fragments is provided to the CMM FQA device so that the memory compression operations can be handled without involving the CMM FGR. Prior art allocation schemes (e.g., buddy allocator) suffer from internal fragmentation. An allocator that suffers from internal fragmentation can lead to a significant memory waste if the requested memory page size is slightly larger than a power of two. For example, if the requested memory page is 1025B, it has to be allocated to a fragment of 2048B, leaving a waste of 1023B. Alternative allocators (e.g., TLSF) can be used improving on internal fragmentation.

[0081] The current disclosure alleviates this problem by serving the allocation of a compressed page with multiple fragments. The fragments are allocated by the CMM FGR but the actual decision about how to combine fragments of different sizes to serve a compressed page, is handled by the CMM FQA. The bookkeeping is done with a low-metadata mechanism by introducing pointers to the allocated fragments: essentially a requested memory page can be placed into one or a plurality of several fragments. In an exemplary embodiment with two fragments per compressed page, if the smallest fragment is 64B, a 1025B memory fragment can be split into two pieces and placed into one 1024B fragment and one 64B fragment, wasting only 63B, and achieving a significant reduction in internal fragmentation.

[0082] In one embodiment, the number of fragments used to store a compressed page is two fragments. In a second embodiment, the number of fragments used to store a compressed page is three. In other embodiments, the number of fragments can be variable using extra metadata to denote how many fragments to be expected.

[0083] This way, the CMM mechanism not only reduces the external fragmentation but also significantly improves the internal fragmentation.Afe wfo

[0084] The CMM FGR allocates buffers of different sizes to be used by the CMM FQA, denoted as DST ID and comprise one or a plurality of fragments IDs. Each fragment ID contains two fields: one pointer, which is an address pointing into the compressed address space, and one order which tells the size of the allocated space. In the present embodiment, the order is defined as: the log2 of the size minus the log2 of the smallest allocatable size + 1. In alternative embodiments, the order field could be the size in bytes or other encoding selected by the skilled person.

[0085] In one embodiment, valid fragment sizes can be 64B, 128B, 256B, 512B, 1024B, 2048B, and 4096B. In a second embodiment, valid fragment sizes can be 256B, 512B,1024B, 2048B and 4096B. In essence, the number of valid fragment sizes is a fine tradeoff between the metadata cost that can be invested in area and flexibility for the utilization of the compressed space. In yet another embodiment, the amount and type of valid fragment sizes can be decided when the system is initialized using configuration options (e.g., smallest allocatable size and largest allocatable size). An invalid fragment is determined when the order field is equal to zero.

[0086] In one embodiment, for space efficiency the pointer and order of each fragment are packed together. The LSBs of the pointer are just zeros, since every pointer is aligned with a minimum alignment (e.g., 64 bytes for a DRAM based memory). The MSBs of the pointer are typically zeros since the entire address is not needed, if it is for example 64 bits. An exemplary fragment ID is depicted in FIG 6. Alternative embodiments where the packing of pointer and order of fragments is differentiated to optimize for other metrics can be appreciated from an expert in the art.

[0087] The fragments inside a DST ID can be spread over the entire compressed memory, but the memory space within each fragment is contiguous.Detailed CMM-FGR and CMM-FQA disclosuresCMM-FGR and CMM-FQA co-design

[0088] The CMM-FGR method pre-allocates memory locations of varying size using Buddy (or other allocation algorithm) off the critical path. The CMM-FQA device comprises hardware-implemented queues to store the pre-allocated memory fragments to service the allocation requests that come from usually the compressor or the recompactor, and in general the memory compression device. The CMM-FQA pops fragments from the queues and assembles them into DST IDs, each of which comprises pointers to pre-allocated memory fragments in the compressed address space.

[0089] FIG. 7 depicts an example of a typical communication sequence between 710 the CMM-FQA and 760 the CMM-FGR. The CMM-FQA receives a request 720 for a DST ID. The CMM-FQA performs a queue lookup 730, finds fragments that can be formed to a suitable-sized DST ID, and 740 sends a response containing the DST ID. In this exemplary embodiment, the popping of the queues lowers the fill level of one queue below a defined watermark which causes a 750 replenish interrupt, signaling the CMM-FGR to refill the queue. The interrupt handler forwards the request to the allocator 764 in the CMM-FGR, and sends an 770 allocation request to the Buddy instance 769. The 769 unit performs the allocations in the buddy tree and returns the allocated slots 780. The allocator 764 receives them and 790 enqueues the fragments to the hardware queues in the 710 CMM-FQA through a register interface.

[0090] In the previous embodiment, the allocator 764 and the Buddy instance 769 are separate units, that is because the CMM-FGR is simplified by dividing specific functionalities across several modules. In an alternative embodiment, the CMM-FGR mayhave multiple implementations of specific allocators to choose between and thus a decision allocator needs to coordinate the specific ones. In alternative embodiments, the units 764 and 769 can be one unit.

[0091] Watermarks can be defined statically or dynamically by the CMM, the memory compression system or other agents through a configuration interface or alternative means realized by those skilled in the art.

[0092] The signaling to the CMM-FGR to refill the queue can be realized in different ways by experts depending on the target system specifications. In one embodiment, the signaling can be done through interrupts should there be a support of interrupt handlers. In an alternative embodiment it can be done through registers where the CMM-FGR is polling them to determine if it needs to take an action (e.g., to refill the queues).

[0093] The interface between the CMM-FGR and CMM-FQA can be done in different ways as disclosed in paragraph

[0079] ,Embodiment CMM-FQA

[0094] FIG. 8 depicts an embodiment 800 of the CMM-FQA mechanism. It shows a deviceside interface that can be implemented in hardware: a DST ID request interface 801, a response interface 802, and a recycled DST ID interface 803. A request 801 arrives at a DST ID assembler (fragments assembler) 810, which checks the fill level of one or a plurality of M fragment queues 820. The DST ID assembler 810 builds a DST ID for the requested size (part of request 801) based on the queue fill levels and returns it at 802.

[0095] The recycle interface 803 receives a recycled DST ID, which arrives at the DST ID disassembler 830. It splits the DST ID into fragments and puts them into the recycle queue 840.

[0096] There is also an interface 808 from the CMM-FGR, wherein the CMM-FGR pushes fragments into the fragment queues 820, and an interface 809 towards the CMM-FGR, wherein the CMM-FGR fetches fragments from the 840 recycle queue.DST ID Assembler

[0097] The DST ID assembler 810 of the CMM-FQA comprises an assembler controller 812 method and one or a plurality of assembler tables 814. Each table 814 maps from a compressed page size to combinations of fragment queues. Each entry in the table 814 holds a number of different fragment combinations to build a specific DST ID size. The DST ID assembler is aware of which queues are non-empty, and selects a fragment combination based on the available fragment queues. It returns an assembled DST ID. The assembler supports two or a plurality of tables in order to enable run-time reprogramming of the assembler but one table can be also preferred depending on design conditions and trade-offs. In the case where there are two or a plurality of tables, the CMM-FGR selects which one to use by setting an example configuration bit. This way, one of the tables can be used while theother(s) is programmed. The CMM-FGR can switch to (any of) the other table(s) when it is fully updated.

[0098] The DST ID assembler 810 can build DST IDs of all sizes possible to create from the smallest fragment of size X up to a full page. If the page is equivalent to a Linux operating system page, then an exemplary full page size is 4096B. The present invention is not limited to specific page sizes or operating systems. Any full page size can be realized as an alternative. Moreover, the full page size can be selected beyond the conventional page sizes and instead selected to align to the smallest allocatable fragment, i.e., X. That means that the number of entries in the table is equal to 4096 / X ( i.e. 64 for X = 64B and 32 for X = 128B).

[0099] FIG. 9 illustrates an example of a DST ID assembler table 900, i.e. one of the aforementioned assembler tables 814 of the DST ID assembler 810. In this embodiment, the smallest fragment size is 64B and each entry (910, 920, 930, 940, 950, 970, 980) holds three fragment combinations. Note that in this exemplary embodiment the assembler table 900 contains 7 entries, while the full table would comprise 64 entries because the number of entries in the said assembler table depends on the smallest and largest fragment size as described previously. FIG. 9 depicts that the entry marked 940, which is used for creating a DST ID of size 2048B, contains three ways of combining fragments to create a DST ID of size 2048B. It can be created by using one fragment of size 2048B 990a, or two fragments of size 1024B 990b, or three fragments of sizes 1024B, 512B, and 512B 990c. In general, each entry can hold up to N fragment combinations, where N is the number of programmable fragment combinations per size.

[0100] Other fragment sizes or various number of combinations (N) can be realized in alternative embodiments.

[0101] In an alternative embodiment, the contents of the DST ID assembler table are memory-mapped and programmed by the CMM-FGR mechanism. This way, the combinations of fragments that are valid to use in a DST ID is decided and can be modified by the CMM-FGR at different times, depending on the design and configuration capabilities. In one embodiment, the combination can be modified at boottime; in an alternative embodiment it can happen at runtime.

[0102] In an exemplary embodiment of the DST ID assembler table, the DST ID assembler table is implemented in hardware as continuous memory and a DST ID table entry has a width determined by the following parameters: a. N: The number of combinations each entry can hold (cf. paragraphs

[0099] -

[0100] ) b. DST ID size: Number of fragments in a DST ID c. qPtr_size: Number of bits in a queue ptr: log2(M) + 1, where M is the number of fragment queues (e.g., 910-980 in FIG. 9) and + 1 is accounting for a valid bit.

[0103] The qPtr is defined as a pointer to one of the fragment queues and contains two fields: one ptr field and one valid bit field.

[0104] The controller of the assembler is outlined in FIG. 10. First, a 1010 request of size req size is received. In the next step 1020 the size req size will first be rounded up to req size’, the nearest multiple of the smallest fragment, and then it will access the assembler table entry of that req size’ .

[0105] In the next state 1030, the controller loops through the combinations in the entry, one by one, and in state 1040 it sees if the fill levels of the required fragment queues are high enough to assemble a DST ID. The assembler creates a DST ID out of the first combination it finds that can be created given the current fragment queue fill levels. If it finds a valid combination, it breaks the loop, and in state 1050 it assembles and returns the DST ID.

[0106] If all combinations have been evaluated but it finds no valid combination, which can happen if any of the required fragment queues are empty, it goes to state 1060 where it will make a final try by finding a fragment queue which is large enough to fit the entire compressed page. If it succeeds, it transitions to state 1050 to assemble and return a DST ID.

[0107] If the final try in 1060 fails, the DST ID assembler will retry the allocation from the beginning (1020) and keep trying until it can assemble a DST ID. In an alternative embodiment, it can cancel the current allocation. This would require the system to have an allocation failure provisioning. One such provisioning could be to introduce an out-of- capacity state. If the CMM-FGR is out of capacity and cannot allocate more fragments, it will signal this to the CMM-FQA. If this bit is set, the assembler returns an invalid DST ID (1080), which signals to the requester that the CMM is out of capacity.

[0108] The present disclosure is not limited to those implementation alternatives, as other alternatives can be derived by those skilled in the art. For example, the priority to select the first combination as described in paragraph

[0105] can be replaced by monitoring the frequency of combinations selected and adapt to use a least frequent combination.DST ID Disassembler

[0109] The DST ID disassembler splits the DST ID into its fragments which are either pushed into their corresponding fragment queues or recycled. The CMM-FQA identifies which fragment queue to put the disassembled fragments in by inspecting the order field in the fragment. If the corresponding fragment queue is full, the fragment shall instead be pushed into the recycle queue. If the recycle queue is also full, the process needs to wait until the recycle queue is empty.Fragment Queue

[0110] The CMM-FQA contains one or plurality of fragment queues and one or a plurality of recycle queues. The exemplary embodiment of FIG. 11 gives an overview of the CMM-FQA, which comprises M fragment queues, 1 recycle queue, and 1 DST ID dissassembler.

[0111] A fragment queue holds the fragments that are allocated by the CMM-FGR. Each fragment queue needs to be initiated by the CMM by configuring its order through the FQ ORDER register. The CMM-FGR must only push fragments of the size dictated by the FQ ORDER to the fragment queue.

[0112] In the exemplary embodiment of FIG. 11, each fragment queue, for example 1110, is built from three parts: one FIFO queue 1111, one control logic block 1112, and one SRAM buffer 1113. The FIFO queue 1111 is dequeued by the memory compression device, e.g., the compressor or the recompactor . The memory compression device also enqueues fragments into the FIFO queue, either from the SRAM buffer 1113 or from the DST ID disassembler 1160. A small control logic block 1112 decides whether the FIFO should be enqueued from the SRAM buffer 1113 or from the disassembler 1160. The CMM-FQA can be designed to re-use as many (old) free fragments as possible to keep the pressure on the CMM-FGR to a minimum level; in this case, the control logic 1112 prioritizes the flow from the disassembler. Furthermore, it only pushes fragments from the SRAM (1113 or 1133) into the FIFO (1111 or 1131) up until a certain point, which is when its fill level passes a reclaim threshold. In this exemplary embodiment, the reclaim threshold is configured by the CMM-FGR through a register, e.g., FQ RECLAIM TH. Hence, if the inequality below is true, the controller does not push any more fragments from the SRAM into the FIFO.FIFO fill level > FQ RECLAIM TH

[0113] In an exemplary embodiment, the memory compression device prioritizes pushing fragments from the DST ID disassembler into the FIFO queue instead of pushing them into the recycle queue. The memory compression device will only push into the recycle queue 1190 if the FIFO queue of the corresponding size is full. Alternative embodiments with different prioritizations can be realized by those skilled in the art.

[0114] The SRAM buffers are written by the CMM-FGR. The CMM-FGR writes one fragment at a time and keeps track of the fill level of the buffer to make sure that no illegal writes are carried out.Recycle Queue

[0115] The recycle queue 1190 (FIG. 11) is used to return fragments to the allocator in the CMM-FGR. Fragments in the recycle queue 1190 are eventually popped by the CMM-FGR, which deallocates them or re-uses them by pushing them back into their respective fragment queue (e.g., from 1190 to 1133 through the CMM-FGR).

[0116] In one embodiment, the recycle queue 1190 can be implemented as a FIFO. The DST ID disassembler 1160 pushes fragments into the FIFO recycle queue. In this exemplary embodiment, the CMM-FGR pops fragments through a single register which exposes the head of the FIFO. Reading this register returns the head of the FIFO and causes the hardware to pop the FIFO and put the new head position in the register. It is the responsibility of the CMM-FGR to keep track of the fill level of the recycle queue and make sure that no illegal reads are carried out but other safe-guard mechanisms can be realized in alternative embodiments. The advantage of this mechanism is that it can allow flexibility: For example, the CMM-FGR can read many fragments in a burst by carrying out DMA reads with stride 0 to the register which exposes the head.Allocation request

[0117] When an allocation request of a certain size arrives at the CMM-FQA, it is propagated to the DST ID assembler which builds a DST ID of a suitable size. The pre-allocated fragments are structured into M number of queues, organized per size, where X is the smallest size available (e.g. 64B or 128B). The DST ID assembler finds a combination of fragments, pops the queues, builds a DST ID, and returns it to the memory compression device. Note that the popped fragments might come from the same queue, e.g. a 3072B DST ID may comprise 3x 1024B fragments. If the allocation step fails, the assembler will either immediately retry or, if the CMM-FGR is out of capacity, an invalid DST ID is returned. The exact steps are described in paragraphs

[0097] -

[0108] , The DST ID assembler retries the allocation immediately if it fails because it expects the CMM-FGR to top up the queues that are low on fragments as soon as possible. The queues are monitored and replenished by the CMM-FGR based on configurable watermarks, which indicates to the CMM-FGR that the queue level is low. If the fill level is critically low, an interrupt is sent to the CMM-FGR to immediately fill the queue. This is described in the following.Queue replenishment

[0118] The fragment queues are filled pro-actively by the CMM-FGR. A replenishment can be triggered by two different events: either when the queue passes a low watermark, or when it passes a critical watermark, as depicted in the exemplary embodiment of FIG. 12. FIG. 12 depicts a fragment queue 1250 with three watermarks: 1251 high, 1253 low, and 1257 critical. The watermarks encompass both the SRAM buffer 1113 (FIG. 11) and the FIFO 1111 (FIG. 11), hence 1250 in the picture is an abstraction of the entire storage area in the fragment queue.

[0119] When the low watermark 1253 is passed, the CMM-FQA will set a replenish flag 1255 in the FQ LOW WATERMARK STATUS register. This register contains one bit per queue. The CMM-FGR will periodically read the FQ LOW WATERMARK STATUS register to detect queues which need to be replenished. If such a queue is detected, the CMM- FQA will trigger a replenishment operation to refill the queue.

[0120] Passing the critical watermark 1257 signals that the queue is almost empty and an interrupt 1259 needs to be raised to the CMM-FGR to immediately start the process to refill the queue. In addition, a flag is set in the FQ CRITICAL WATERMARK STATUS register. Alternative signaling mechanisms shall be introduced depending on the exact realization of the CMM-FGR, CMM-FQA and the in-between interface depending on the memory compression system capabilities, as disclosed in paragraphs

[0078] -

[0079] ,

[0121] These inequalities are applied as part of the exemplary embodiment of FIG. 12:FQ HIGH WATERMARK STATUS = FQ FILL LEVEL > FQ HIGH WATERMARKFQ LOW WATERMARK STATUS = FQ FILL LEVEL < FQ LOW WATERMARKFQ CRITICAL WATERMARK STATUS = FILL LEVEL < FQ CRITICAL WATERMARK

[0122] During a replenishment, the CMM-FGR will allocate n fragments of the requested size and push them into the SRAM buffer (e.g., 1113 or 1133 in FIG. 11). In the present embodiment, the size of n (how many fragments to allocate) is as many fragments as fits between the high and low watermark. Other more sophisticated embodiments filling the fragment queue can be realized.

[0123] FIG. 13 depicts the replenishment process. It is triggered from two different sources 1310: either if a triggered low watermark is detected by the CMM-FGR 1314, or if a critical watermark is triggered 1316 signalling accordingly the CMM-FGR 1319.

[0124] After that, the CMM-FGR allocation is triggered 1320. If it succeeds, the queue which triggered the process is populated with the allocated fragments 1340. However, if the allocation fails, there is a check 1350 to see if the memory is full or fragmented to a degree where the allocation algorithm can not satisfy the request with large enough fragments. If the memory is full, the process sets the out-of-memory flag 1370 and ends 1390.

[0125] If the memory is not full, but the allocation fails, the memory has to be defragmented 1360. For example, in the embodiment of an FGR buddy allocator, if every other leaf node in the buddy tree is allocated with a small size, no buddies can be merged into larger fragments. In such a case, the memory has to be defragmented before the buddy allocation algorithm can service the request. After the defragmentation, the allocation can be re-triggered 1320. It is worth noting that the current patent disclosure is not limited to either the number of watermarks nor their type (e.g., high, low, critical); different number of watermarks, types or conditions can be introduced in alternative embodiments of the fragment queues.Deallocation request

[0126] Memory can be directly deallocated in different ways as mentioned in paragraph

[0062] by triggering a fragment release operation. In such a case, any fragments allocated for one or a plurality of selected pages will be returned to the CMM. The CMM FGR deallocation can be triggered by various sources as defined in paras

[0062] and

[0073] , This is done first by the various resources pushing release fragment requests to the CMM FQAdeallocator producer 594a. The 594a unit gathers the released fragments and passes them to the CMM FGR, which further does the garbage collection by updating the data structure in the layout manager 586 so that the respective fragments become “new free”. This operation has the same effect as if only zeros are written to a page which previously had allocated fragments: the fragments will be returned to the CMM (reflected with an arrow from the compressor 520 to the CMM FQA deallocator 594 and further to the CMM deallocator 584 and the page will be marked as zero in the remapper. The use of the remapper is disclosed later in the patent application.

[0127] When a deallocation request (or release fragment in FIG. 2 and FIG. 5) is received, the DST ID disassembler splits the DST ID into fragments, which are pushed either into a fragment queue of the corresponding size, or if the fragment queue is full, into the recycle queue. From the perspective of the memory compression device, that is the only step carried out by the CMM during a deallocation. The draining of the recycle queue is handled off the critical path by the CMM-FGR.

[0128] The recycle queue, 1450 in FIG. 14, has watermarks similar to the fragment queues. However, the watermarks are used differently: the critical watermark 1454 indicates that the recycle queue is almost full. If the queue length passes the critical watermark, the CMM-FGR must be signalled (an interrupt 1451 in the embodiment of FIG. 14) to immediately start the draining of the recycle queue. If the queue length passes the high watermark 1459, a flag will be set 1459 in the RECYCLE HIGH WATERMARK STATUS register. The CMM-FGR will periodically scan the RECYCLE HIGH WATERMARK STATUS register and drain the recycle queue to avoid reaching the critical watermark. For clarity, there inequalities apply:RECYCLE HIGH WATERMARK STATUS = RQ FILL LEVEL > HIGH WATERMARKRECYCLE CRITICAL WATERMARK STATUS = RQ FILL LEVEL > CRITICAL WATERMARK

[0129] The method of FIG. 15 depicts the two sources 1510 which may trigger the recycle draining: a high watermark is passed 1513, or a critical watermark is passed 1516. As for the high watermark, the CMM-FGR notices the triggering of the high watermark by scanning the status registers. However, if the critical watermark is passed, the CMM-FQA raises an interrupt 1519. Alternative signaling mechanisms can be introduced by those skilled depending on interface capabilities in the memory compression system.

[0130] In 1530, the CMM-FGR will make a choice in how to drain the recycle queue. It can decide to push the fragments back into the fragment queue of the corresponding size to be used again by another allocation request (1550). Alternatively, the fragment is completely removed from the loop and returned to the allocator (buddy algorithm 1570 in the exemplary embodiment of FIG. 15). This makes it possible for the buddy algorithm to free up the allocated fragment and possibly merge smaller fragments into larger ones, depending on what fragments are unallocated. It is worth noting that the current patent disclosure is not limited to either the number of watermarks nor their type (e.g., high, low, critical); different number ofwatermarks, types or conditions can be introduced in alternative embodiments of the fragment queues.Emptying the fragment queues

[0131] In some scenarios, the CMM-FGR might want to re-balance the fragments in the queues, e.g. to deallocate the currently allocated fragments and allocate other sizes. This can be achieved by a mechanism which empties targeted queues and pushes them to the recycle queue. In one embodiment, this mechanism is implemented in the CMM-FQA device in hardware to accelerate this process.

[0132] Emptying the queues are carried out by lowering the high watermark for each queue (FIG. 12). The CMM-FGR reconfigures the high watermark below the current queue size. This signals to the CMM-FQA to pop fragments from the queues and place them into the recycle queue. This way, the CMM-FGR can access the fragments in the recycle queue and carry out deallocation operations.Mappings and metadata management

[0133] FIG. 16 depicts an embodiment of an exemplary memory compression system wherein the memory compression device is integrated with an embodiment of the CMM mechanism as disclosed in this patent application, for example, as a combination of CMM- FGR and CMM-FQA. For a page compression operation (and thus CMM allocation but not CMM deallocation), the operation is as follows: A memory write request 1610 and data 1620 is received and is compressed by the memory compression device 1630. The memory compression device 1630 picks up the fragment information DST ID 1650, which is provided by the CMM 1640. This information, which is the destination address(es) in the compressed address space, can be bookkept in internal metadata 1660 managed by memory compression. This metadata is stored in memory 1680 in a specific region 1684 managed solely by the memory compression device.

[0134] FIG. 17 depicts an embodiment of an exemplary memory compression system wherein the memory compression device 1720 decompresses a page that was compressed with a system such as the exemplary

[0139] embodiment of FIG. 16 that integrates a memory compression device and the CMM mechanism disclosed in this patent application. For a page decompression operation, the operation is as follows. A read request with an address 1710 is by the memory compression device 1720. The Remapper mechanism 1724 uses the request’s address 1710 to map into the address(es) of the compressed page in the compressed address space in the memory. For this operation, it may need to read metadata in the metadata region 1734 in the memory 1730. The memory compression device uses the remapped address (src 1739) to read 1728 the compressed page from memory 1730 (memory block by memory block 1780) and does the decompression 1750. Data is transferred to the requestor, responding 1770 to specific memory read requests 1710 being received earlier and bookkept by the memory compression device 1720.CMM maintenance

[0135] The CMM method and device in a memory compression system is designed to minimize internal and external fragmentation. In some exemplary embodiments of the CMM, disclosed further in the patent application, the compressed memory layout manager is implemented with data structures that are optimal to handle internal fragmentation or external fragmentation or potentially both types of fragmentation. In other embodiments, other data structures may be selected by those skilled in the art depending on the performance / memory / - latency trade-offs with respect to data structure traversal, data structure management in memory or other metrics.

[0136] In certain conditions, and perhaps independently of the data structure that implements the memory management layout, especially when the memory capacity is low or critical, the CMM may consider relocating stored pages so that free memory is coalesced. This is referred to as defragmentation and is an example maintenance operation of the CMM handled by the maintenance controller 218 (FIG. 2A and 2B) or 588 (FIG. 5). The memory layout management 586 based on the monitoring of the free and used fragments in the memory space, can identify the memory fragmentation degree. Based on this and potentially other factors in the system such as the current memory pressure, the criticality of the relocation operation on the system’s performance and the benefit of creating more contiguous memory, the maintenance controller 586 decides whether and what pages to relocate to achieve better defragmentation.

[0137] The maintenance controller 518 (FIG. 5) selects the pages that will be relocated by the memory compression system and programs the recompactor unit 540 to execute the recompaction operation. The programming essentially includes metadata information related to the locations of the pages that should be read from the memory. The new locations are selected by CMM FGR allocator unit 582 similarly to how the compression operation is executed. The recompactor 540 in this exemplary embodiment will be executed as follows: a. It reads the compressed page from the memory 570. b. It provides the size of the compressed page to the allocator unit 592, and the allocator unit returns with one or a plurality of fragments. c. The recompactor 540 writes the compressed page to the memory 570 in the designated memory area indicated by the fragment information. d. The recompactor 540 releases the fragments used by the page by communicating this information to the deallocator unit 594.

[0138] The exemplary recompaction operation is executed without needing to decompress and recompress the page, however in other embodiments, it might be deemed desirable to recompress the page as well. For example, this can be preferred if there are different algorithms so that compression with all algorithms is retried or because the compressionalgorithm used can adapt at runtime thus it could be deemed to provide better compression performance.

[0139] The maintenance controller 588 is responsible for handling other maintenance operations. In a second embodiment, the maintenance controller 588 would be responsible for rebalancing the queues. The CMM FGR can monitor the queue usage and notice that a specific queue, programmed to hold a specific fragment size, is barely used, while some other queue with a different fragment size is of high demand. There are two possible actions that the CMM FGR can take: 1) Reprogram the queues to hold different sizes (i.e. low-demand fragment queue becomes one more high-demand fragment queue). This also requires an update of the assembler table. 2) Only re-program the assembler table so that the low-demand queue is used more frequently. For example, assuming that 64B fragments are never used at all, while 128B fragments are always used. Thus, the assembler table could be updated so that the 128B entry contains [64B, 64B] instead of [128B], to-reroute all 128B accesses to the 64B queue instead.System with memory compression device and CMM method / device

[0140] The compressed memory management (CMM) device disclosed in the present invention comprises features that in an applied system have benefits to be applied in different parts of the system, for example software (SW) and hardware (HW). In the present embodiment, the features added in SW give the benefit for more flexibility, whilst the features applied to the HW can provide higher performance and more energy efficiency. Importantly, selecting the SW features and HW features are crucial for the performance of the compressed system, notably performance of critical operations such as memory writes require to know the target location in a compressed memory on the order of nanoseconds, thus invoking software routines would lead to severe latency overheads and degraded performance. This way, a HW-SW codesign can be preferred so that the time critical functionalities are implemented in hardware while flexibility is maintained in SW.

[0141] FIG. 18 depicts the embodiment of a system of FIG. 1C or FIG. ID employing memory compression wherein the compressed memory is managed by a CMM mechanism as disclosed in the present invention. The embodiment of FIG. 18 comprises an exemplary CPU subsystem (Host) 1810 connected to a memory subsystem 1840, which is further connected to the memory 1850 which contains compressed data, compressed by the memory compression device 1844. The critical [memory access] path, as referred to previously in this document, can be seen at 1861-1864. In alternative, the compression and decompression can occur along / inside the CPU subsystem 1810 as depicted in FIG. 1C. A system that employs memory compression requires a mechanism to manage the compressed memory. In the embodiment of FIG. 18, the CMM mechanism is composed of CMM-FQA functionality 1844a and CMM-FGR functionality 1812 (being an embodiment of CMM-FGR 580). In the disclosed embodiment, the CMM is a computer-executable method (e.g., software subroutine) running on the host CPU. The CMM-FGR functionality 1812 comprises an allocation method of free space (1812d), deallocation of free space that is released out of anycompressed blocks (1812c), compressed memory layout management (1812a) and other layout maintenance operations such as defragmentation (1812b). The memory compression device 1844 further comprises a compression controller 1844a, which instantiates the CMM- FQA disclosed in this patent application, such as 590 or 800, and a remapper 1844b similar to 1724. Experts in the field can realize other memory compression systems similar to the embodiment of FIG. 18 but wherein the remapper is built from state of art as long as it maintains the fragment information as created by the CMM disclosed in this patent application.

[0142] FIG. 19 depicts the embodiment of the system of FIG. IE or FIG. IF employing memory compression wherein the compressed memory is managed by a CMM mechanism as disclosed in the present invention. The embodiment of FIG. 19 comprises an exemplary CPU subsystem (Host) 1910 connected to a memory subsystem 1940, which is further connected to the memory 1950 which contains compressed data compressed by the memory compression device 1844. The critical [memory access] path, as referred to previously in this document, can be seen at 1961-1964. In alternative, the compression and decompression can occur along / inside the CPU subsystem 1910 as depicted in FIG. IE). A system that employs memory compression requires a mechanism to manage the compressed memory. In the embodiment of FIG. 19, the CMM mechanism is composed of the CMM device 1943 and the CMM functionality 1941 (being an embodiment of CMM-FGR 580). This is similar to the embodiment of FIG. 18 but with the difference that both the CMM sub-mechanisms (CMM- FGR and CMM-FQA) are integrated in the memory subsystem 1940. In this embodiment, the system represents a disaggregated compute-memory system wherein the memory subsystem is built as a separate chip and moreover it comprises a CPU 1945 that allows the execution of software-implemented methods, such as the CMM method (CMM-FGR) 1941, directly on the memory subsystem. In both embodiments of FIG. 18 and FIG. 19, the CMM-FGR method 1941 (1812 in FIG. 18) and the CMM-FQA device 1943a (1844 in FIG. 18) must communicate to each other to optimize the compressed memory management combining the flexibility of the method and speed of the device. The communication between the CMM- FGR method and CMM-FQA device can happen through directly connected wires, memory mapped register or other similar means available in said systems to allow the connection of CPU-executable methods (software subroutine) and devices, as has been discussed above.Summarizing the disclosed invention

[0143] As will have been understood from the disclosure above, and in particular FIG. 4, 5, 18 and 19 with their associated description, one aspect of the present invention is a compressed memory manager (CMM) for use in a computer system, e.g. 1800; 1900, that has a computer subsystem 450; 550, a memory subsystem 470; 570 and a memory compression arrangement 401; 501. The memory compression arrangement 401; 501 comprises a compressor 420; 520 and a decompressor 460; 580 for compressing and decompressing, respectively, memory contents in a physical address space of the memory subsystem 470; 570 into a compressed address space used by the computer subsystem 450; 550.

[0144] The CMM is configured for managing a compressed memory layout which divides the physical address space of the memory subsystem 470; 570 into a plurality of fragments of memory, the fragments being of two or more different sizes. The CMM is further configured for pre-allocating free fragments in the compressed memory layout, for receiving and handling fragment allocation requests from the memory compression arrangement 401; 501 by providing one or more of the pre-allocated free fragments for use by the memory compression arrangement 401; 501 to accommodate pieces of compressed memory, as well as for receiving and handling fragment release requests that pertain to one or more used fragments by providing at least some of said one or more used fragments for deallocation in the compressed memory layout.

[0145] Notably, the managing of the compressed memory layout, the pre-allocating of free fragments and the deallocation of used fragments in the compressed memory layout are performed decoupled from and asynchronously with the receiving and handling of the fragment allocation requests and fragment release requests from the memory compression arrangement 401; 501.

[0146] Advantageously, the CMM is configured for performing the managing of the compressed memory layout, the pre-allocating of free fragments and the deallocation of used fragments in the compressed memory layout off a memory access path (cf. 1861-1864 in FIG. 18; 1961-1964 in FIG. 19) in which the computer subsystem 450; 550 accesses the memory subsystem 470; 570 through the memory compression arrangement 401; 501. At the same time, the CMM is configured for performing the receiving and handling of the fragment allocation requests and fragment release requests from the memory compression arrangement 401; 501 in line with the aforementioned memory access path 1861-1864; 1961-1964.

[0147] The CMM advantageously has fragment recycling functionality. To this end, the CMM is configured for handling some of the aforementioned fragment release requests by providing said one or more used fragments like said pre-allocated free fragments, thereby recycling said one or more used fragments for new use by the memory compression arrangement 401; 501.

[0148] As has been presented above with reference to FIG. 4 and FIG. 5, the CMM conveniently has functional components that include an allocator 412; 512, a deallocator 414; 514 and a layout manager 416; 516. The layout manager 416; 516 is configured for managing the compressed memory layout. The allocator 412; 512 is configured for controlling the preallocating of the free fragments in the compressed memory layout by the layout manager 416; 516, and for the receiving and handling of the fragment allocation requests. The deallocator 414; 514 is configured for the receiving and handling of the fragment release requests, and for controlling the deallocation of said at least some of said one or more used fragments in the compressed memory layout by the layout manager 416; 516.

[0149] Advantageously, the CMM further comprises a maintenance controller 418; 518 configured for causing a recompactor 440; 540 of the memory compression arrangement 401;501 to perform defragmentation of compressed memory contents in the physical address space of the memory subsystem 470; 570 by reorganizing pieces of compressed memory into contiguous accumulated free memory space out of smaller pieces of free memory space between pieces of compressed data. More specifically, as was described particularly in paragraph

[0137] , the maintenance controller 418; 518 may cause the recompactor 440; 540 of the memory compression arrangement 401; 501 to perform defragmentation of the compressed memory contents by instructing the recompactor 440; 540 of the memory compression arrangement 401; 501 of a piece of compressed memory to be reorganized. As the recompactor 440; 540 of the memory compression arrangement 401; 501 has read the piece of compressed memory to be reorganized as currently stored in one or more fragments of the physical address space of the memory subsystem 470; 570, the allocator 412; 512 will receive a fragment allocation request indicating a size of the piece of compressed memory to be reorganized. In response, the allocator 412; 512 will provide one or more of the preallocated free fragments, having a total size satisfying the indicated size, thereby enabling the memory compression arrangement 401; 501 to store the read piece of compressed memory in the provided one or more of pre-allocated free fragments. The deallocator 412; 512 will subsequently receive, from the recompactor 440; 540 of the memory compression arrangement 401; 501, a fragment release request that pertains to the one or more fragments from which the piece of compressed memory to be reorganized was previously stored.

[0150] As will have been understood particularly from FIG. 4, 5, 7, 18 and 19 with their associated description, the CMM beneficially has a co-design by being divided into a fragment generator and replenisher unit (CMM-FGR) 580; 760; 1812; 1941 and a fragment quick access unit (CMM-FQA) 590; 710; 1844a; 1943a. Here, the CMM-FGR 580; 760; 1812; 1941 comprises the layout manager 516 and a producer part 512a of the allocator 512, the producer part 512a being configured for the controlling of the pre-allocating of the free fragments in the compressed memory layout by the layout manager 516. The CMM-FGR 580; 760; 1812; 1941 further comprises a consumer part 514a of the deallocator 514, the consumer part 514a being configured for the controlling of the deallocation of said at least some of said one or more used fragments in the compressed memory layout by the layout manager (516). The CMM-FQA 590; 710; 1844a; 1943a, on the other hand, comprises a consumer part 512b of the allocator 512, the consumer part 512b being configured for the receiving and handling of the fragment allocation requests, as well as a producer part 514bof the deallocator 514, the producer part 514b being configured for the receiving and handling of the fragment release requests. The CMM-FGR 580; 760; 1812; 1941 may further comprise the aforementioned maintenance controller 518.

[0151] As will have been understood particularly from FIG. 8, 9 and 10 with their associated description, the producer part 512a of the CMM-FGR 580; 760; 1812; 1941 is configured to repeatedly request pre-allocation of fragments in a plurality of sizes and to add the preallocated fragments to different fragment queues 820, there being at least one fragment queue for each fragment size. The consumer part 512b of the CMM-FQA 590; 710; 1844a; 1943a is configured for assembling a set of fragments selected from the different fragment queues 820to match a size of a piece of compressed memory requested by the memory compression arrangement 401; 501 in a fragment allocation request, and for responding to the fragment allocation request with metadata DST ID representing, for each selected fragment, a start address in the compressed memory space together with information indicative of the memory size allocated by the assembled set of fragments.

[0152] More specifically, in an advantageous embodiment, the CMM-FQA 590; 710; 1844a; 1943a comprises a fragments assembler 810 configured for assembling said set of fragments. The fragments assembler 810 comprises an assembler controller 812 and one or more assembler tables 814; 900. Each assembler table 900 represents a plurality of sizes of compressed memory pages, wherein each entry 940 in the assembler table 900 defines a number of combinations 990a, 990b, 990c of fragment sizes that together will make up the respective compressed memory page size. The assembler controller 812 is configured for handling a fragment allocation request for a piece of compressed memory of a certain requested size by accessing the assembler table 900 to retrieve an entry 940 therein that represents a compressed memory page size suitable for the requested size. The assembler controller 812 investigates the combinations 990a, 990b, 990c of fragment sizes defined by the retrieved entry 940 with respect to capacities of the fragment queues 820 in terms of availability of pre-allocated fragments of respective sizes therein (cf. the queue fill levels discussed above with reference to FIG. 8-10). The assembler controller 812 assembles said set of fragments by selecting from the fragment queues 820 in accordance with a successfully investigated combination of fragment sizes.

[0153] The CMM co-design provides substantial benefits in terms of efficiency and flexibility of implementation. For instance, the CMM-FQA 590; 710; 1844a; 1943a may beneficially be implemented in hardware and separately from the CMM-FGR 580; 760; 1812; 1941. Advantageously, the CMM-FQA 590; 710; 1844a; 1943a may be implemented in the same hardware implementation as the memory compression arrangement 401; 501. Conveniently, on the other hand, in some embodiments the CMM-FGR 580; 760; 1812; 1941 is implemented as software or firmware executed by a processing unit of the computer subsystem 450; 550, or as a field-programmable gate array.

[0154] In some embodiments of the computer system 1800; 1900, the computer subsystem 450; 550 and memory subsystem 470; 570 are disaggregated and implemented in different systems-on-a-chip (SoC:s). In other embodiments of the computer system 1800; 1900, the computer subsystem 450; 550 and memory subsystem 470; 570 are implemented in a common system-on-a-chip (SoC).

[0155] Another aspect of the present invention is a computer system 1800; 1900 comprising a computer subsystem 450; 550, a memory subsystem 470; 570, a memory compression arrangement 401; 501 and a compressed memory manager CMM. The memory compression arrangement 401; 501 comprises a compressor 420; 520 and a decompressor 460; 560 for compressing and decompressing, respectively, memory contents in a physical address space of the memory subsystem 470; 570 into a compressed address space used by the computersubsystem 450; 550. As discussed above, the compressed memory manager CMM is configured for managing a compressed memory layout which divides the physical address space of the memory subsystem 470; 570 into a plurality of fragments of memory, the fragments being of two or more different sizes. The CMM is further configured for preallocating free fragments in the compressed memory layout, for receiving and handling fragment allocation requests from the memory compression arrangement 401; 501 by providing one or more of the pre-allocated free fragments for use by the memory compression arrangement 401; 501 to accommodate pieces of compressed memory, as well as for receiving and handling fragment release requests that pertain to one or more used fragments by providing at least some of said one or more used fragments for deallocation in the compressed memory layout. The managing of the compressed memory layout, the preallocating of free fragments and the deallocation of used fragments in the compressed memory layout are performed decoupled from and asynchronously with the receiving and handling of the fragment allocation requests and fragment release requests from the memory compression arrangement 401; 501.

[0156] In the computer system 1800; 1900 according to this aspect of the present invention, the compressed memory manager CMM may be further configured like the compressed memory manager CMM in the previously described aspect of the present invention (cf. paragraph

[0143] onwards).

[0157] As is shown in the flowchart diagram of FIG. 20, yet another aspect of the present invention is a compressed memory management method 2000 for a computer system 1800; 1900. As with the previously described aspects, the computer system 1800; 1900 has a computer subsystem 450; 550, a memory subsystem 470; 570 and a memory compression arrangement 401; 501. The memory compression arrangement 401; 501 comprises a compressor 420; 520 and a decompressor 460; 560 for compressing and decompressing, respectively, memory contents in a physical address space of the memory subsystem 470; 570 into a compressed address space used by the computer subsystem (450; 550).

[0158] The compressed memory management method 2000 comprises managing 2010 a compressed memory layout which divides the physical address space of the memory subsystem 470; 570 into a plurality of fragments of memory, the fragments being of two or more different sizes.

[0159] The compressed memory management method 2000 moreover comprises preallocating 2020 free fragments in the compressed memory layout.

[0160] The compressed memory management method 2000 further comprises receiving and handling 2030 fragment allocation requests from the memory compression arrangement 401; 501 by providing one or more of the pre-allocated free fragments for use by the memory compression arrangement 401; 501 to accommodate pieces of compressed memory.

[0161] In addition, the compressed memory management method 2000 comprises receiving and handling 2040 fragment release requests pertaining to one or more used fragments byproviding at least some of said one or more used fragments for deallocation 2050 in the compressed memory layout.

[0162] According to the method 2000, and as can be seen at 2060 in FIG. 20, the managing 2010 of the compressed memory layout, the pre-allocating 2020 of free fragments and the deallocating 2050 of used fragments in the compressed memory layout are performed decoupled from and asynchronously with the receiving and handling 2030, 2040 of the fragment allocation requests and fragment release requests from the memory compression arrangement 401; 501.

[0163] The compressed memory management method 2000 may further comprise any or all of the functionality recited for the compressed memory manager CMM in the previously described aspects of the present invention (cf. paragraph

[0143] onwards and paragraph

[0155] onwards, respectively).

Claims

CLAIMS1. A compressed memory manager (CMM) for use in a computer system (1800; 1900) having a computer subsystem (450; 550), a memory subsystem (470; 570) and a memory compression arrangement (401; 501), the memory compression arrangement (401; 501) comprising a compressor (420; 520) and a decompressor (460; 560) for compressing and decompressing, respectively, memory contents in a physical address space of the memory subsystem (470; 570) into a compressed address space used by the computer subsystem (450; 550), the compressed memory manager (CMM) being configured for: managing a compressed memory layout which divides the physical address space of the memory subsystem (470; 570) into a plurality of fragments of memory, the fragments being of two or more different sizes; pre-allocating free fragments in the compressed memory layout; receiving and handling fragment allocation requests from the memory compression arrangement (401; 501) by providing one or more of the pre-allocated free fragments for use by the memory compression arrangement (401; 501) to accommodate pieces of compressed memory; and receiving and handling fragment release requests pertaining to one or more used fragments by providing at least some of said one or more used fragments for deallocation in the compressed memory layout, wherein the managing of the compressed memory layout, the pre-allocating of free fragments and the deallocation of used fragments in the compressed memory layout are performed decoupled from and asynchronously with the receiving and handling of the fragment allocation requests and fragment release requests from the memory compression arrangement (401; 501).

2. The compressed memory manager (CMM) as defined in claim 1, wherein the compressed memory manager (CMM) is configured for performing the managing of the compressed memory layout, the pre-allocating of free fragments and the deallocation of used fragments in the compressed memory layout off a memory access path (1861-1864; 1961- 1964) in which the computer subsystem (450; 550) accesses the memory subsystem (470; 570) through the memory compression arrangement (401; 501), while performing the receiving and handling of the fragment allocation requests and fragment release requests fromthe memory compression arrangement (401; 501) in line with said memory access path (1861-1864; 1961-1964).

3. The compressed memory manager (CMM) as defined in claim 1 or 2, wherein the compressed memory manager (CMM) is configured for handling some of said fragment release requests by providing said one or more used fragments like said pre-allocated free fragments, thereby recycling said one or more used fragments for new use by the memory compression arrangement (401; 501).

4. The compressed memory manager (CMM) as defined in any preceding claim, comprising an allocator (412; 512), a deallocator (414; 514) and a layout manager (416; 516), wherein the layout manager (416; 516) is configured for managing the compressed memory layout, wherein the allocator (412; 512) is configured for controlling the pre-allocating of the free fragments in the compressed memory layout by the layout manager (416; 516), and for the receiving and handling of the fragment allocation requests, and wherein the deallocator (414; 514) is configured for the receiving and handling of the fragment release requests, and for controlling the deallocation of said at least some of said one or more used fragments in the compressed memory layout by the layout manager (416; 516).

5. The compressed memory manager (CMM) as defined in claim 4, further comprising a maintenance controller (418; 518) configured for causing a recompactor (440; 540) of the memory compression arrangement (401; 501) to perform defragmentation of compressed memory contents in the physical address space of the memory subsystem (470; 570) by reorganizing pieces of compressed memory into contiguous accumulated free memory space out of smaller pieces of free memory space between pieces of compressed data.

6. The compressed memory manager (CMM) as defined in claim 5, further configured as follows: the maintenance controller (418; 518) causing the recompactor (440; 540) of the memory compression arrangement (401; 501) to perform defragmentation of the compressed memory contents by instructing the recompactor (440; 540) of the memory compressionarrangement (401; 501) of a piece of compressed memory to be reorganized; the allocator (412; 512) receiving, from the recompactor (440; 540) of the memory compression arrangement (401; 501) after the recompactor (440; 540) having read the piece of compressed memory to be reorganized as currently stored in one or more fragments of the physical address space of the memory subsystem (470; 570), a fragment allocation request indicating a size of the piece of compressed memory to be reorganized; the allocator (412; 512) in response providing one or more of the pre-allocated free fragments, having a total size satisfying the indicated size, thereby enabling the memory compression arrangement (401; 501) to store the read piece of compressed memory in the provided one or more of pre-allocated free fragments; and the deallocator (412; 512) subsequently receiving, from the recompactor (440; 540) of the memory compression arrangement (401; 501), a fragment release request pertaining to the one or more fragments from which the piece of compressed memory to be reorganized was previously stored.

7. The compressed memory manager (CMM) as defined in any of claims 4-6, divided into a fragment generator and replenisher unit (CMM-FGR; 580; 760; 1812; 1941) and a fragment quick access unit (CMM-FQA; 590; 710; 1844a; 1943a), wherein the fragment generator and replenisher unit (CMM-FGR; 580; 760; 1812; 1941) comprises: the layout manager (516), a producer part (512a) of the allocator (512), the producer part (512a) being configured for the controlling of the pre-allocating of the free fragments in the compressed memory layout by the layout manager (516), and a consumer part (514a) of the deallocator (514), the consumer part (514a) being configured for the controlling of the deallocation of said at least some of said one or more used fragments in the compressed memory layout by the layout manager (516), and wherein the fragment quick access unit (CMM-FQA; 590; 710; 1844a; 1943a) comprises: a consumer part (512b) of the allocator (512), the consumer part (512b) being configured for the receiving and handling of the fragment allocation requests, and a producer part (514b) of the deallocator (514), the producer part (514b) being configured for the receiving and handling of the fragment release requests.

8. The compressed memory manager (CMM) as defined in claim 5 and 7, wherein the fragment generator and replenisher unit (CMM-FGR; 580; 760; 1812; 1941) further comprises the maintenance controller (518).

9. The compressed memory manager (CMM) as defined in any of claims 7-8, wherein the producer part (512a) of the fragment generator and replenisher unit(CMM-FGR; 580; 760; 1812; 1941) is configured to repeatedly request pre-allocation of fragments in a plurality of sizes and to add the pre-allocated fragments to different fragment queues (820), there being at least one fragment queue for each fragment size, and wherein the consumer part (512b) of the fragment quick access unit (CMM-FQA; 590; 710; 1844a; 1943a) is configured for assembling a set of fragments selected from the different fragment queues (820) to match a size of a piece of compressed memory requested by the memory compression arrangement (401; 501) in a fragment allocation request, and for responding to the fragment allocation request with metadata (DST ID) representing, for each selected fragment, a start address in the compressed memory space together with information indicative of the memory size allocated by the assembled set of fragments.

10. The compressed memory manager (CMM) as defined in claim 9, wherein the fragment quick access unit (CMM-FQA; 590; 710; 1844a; 1943a) comprises a fragments assembler (810) configured for assembling said set of fragments, the fragments assembler (810) comprising an assembler controller (812) and one or more assembler tables (814; 900), wherein each assembler table (900) represents a plurality of sizes of compressed memory pages, each entry (940) in the assembler table (900) defining a number of combinations (990a, 990b, 990c) of fragment sizes that together will make up the respective compressed memory page size, and wherein the assembler controller (812) is configured for handling a fragment allocation request for a piece of compressed memory of a certain requested size by: accessing the assembler table (900) to retrieve an entry (940) therein that represents a compressed memory page size suitable for the requested size; investigating the combinations (990a, 990b, 990c) of fragment sizes defined by the retrieved entry (940) with respect to capacities of the fragment queues (820) in terms of availability of pre-allocated fragments of respective sizes therein; and assembling said set of fragments by selecting from the fragment queues (820) in accordance with a successfully investigated combination of fragment sizes.

11. The compressed memory manager (CMM) as defined in any of claims 7-10, wherein the fragment quick access unit (CMM-FQA; 590; 710; 1844a; 1943a) is implemented in hardware and separately from the fragment generator and replenisher unit (CMM-FGR; 580; 760; 1812; 1941).

12. The compressed memory manager (CMM) as defined in claim 11, wherein the fragment quick access unit (CMM-FQA; 590; 710; 1844a; 1943a) is implemented in the same hardware implementation as the memory compression arrangement (401; 501).

13. The compressed memory manager (CMM) as defined in claim 11, wherein the fragment generator and replenisher unit (CMM-FGR; 580; 760; 1812; 1941) is implemented as software or firmware executed by a processing unit of the computer subsystem (450; 550), or as a field-programmable gate array.

14. A computer system (1800; 1900) comprising: a computer subsystem (450; 550); a memory subsystem (470; 570); a memory compression arrangement (401; 501), the memory compression arrangement (401; 501) comprising a compressor (420; 520) and a decompressor (460; 560) for compressing and decompressing, respectively, memory contents in a physical address space of the memory subsystem (470; 570) into a compressed address space used by the computer subsystem (450; 550); and a compressed memory manager (CMM) being configured for: managing a compressed memory layout which divides the physical address space of the memory subsystem (470; 570) into a plurality of fragments of memory, the fragments being of two or more different sizes; pre-allocating free fragments in the compressed memory layout; receiving and handling fragment allocation requests from the memory compression arrangement (401; 501) by providing one or more of the pre-allocated free fragments for use by the memory compression arrangement (401; 501) to accommodate pieces of compressed memory; and receiving and handling fragment release requests pertaining to one or more used fragments by providing at least some of said one or more used fragments for deallocation inthe compressed memory layout, wherein the managing of the compressed memory layout, the pre-allocating of free fragments and the deallocation of used fragments in the compressed memory layout are performed decoupled from and asynchronously with the receiving and handling of the fragment allocation requests and fragment release requests from the memory compression arrangement (401; 501).

15. The computer system (1800; 1900) as defined in claim 14, with the compressed memory manager (CMM) further configured as defined for the compressed memory manager (CMM) in any of claims 2-13.

16. The computer system (1800; 1900) as defined in claim 14 or 15, wherein the computer subsystem (450; 550) and memory subsystem (470; 570) are disaggregated and implemented in different systems-on-a-chip (SoC:s).

17. The computer system (1800; 1900) as defined in claim 14 or 15, wherein the computer subsystem (450; 550) and memory subsystem (470; 570) are implemented in a common system-on-a-chip (SoC).

18. A compressed memory management method (2000) for a computer system (1800; 1900) having a computer subsystem (450; 550), a memory subsystem (470; 570) and a memory compression arrangement (401; 501), the memory compression arrangement (401; 501) comprising a compressor (420; 520) and a decompressor (460; 560) for compressing and decompressing, respectively, memory contents in a physical address space of the memory subsystem (470; 570) into a compressed address space used by the computer subsystem (450; 550), the compressed memory management method (2000) comprising: managing (2010) a compressed memory layout which divides the physical address space of the memory subsystem (470; 570) into a plurality of fragments of memory, the fragments being of two or more different sizes; pre-allocating (2020) free fragments in the compressed memory layout; receiving and handling (2030) fragment allocation requests from the memory compression arrangement (401; 501) by providing one or more of the pre-allocated free fragments for use by the memory compression arrangement (401; 501) to accommodate pieces of compressed memory; andreceiving and handling (2040) fragment release requests pertaining to one or more used fragments by providing at least some of said one or more used fragments for deallocation in the compressed memory layout, wherein the managing (2010) of the compressed memory layout, the pre-allocating (2020) of free fragments and the deallocating (2050) of used fragments in the compressed memory layout are performed decoupled from and asynchronously with (2060) the receiving and handling (2030, 2040) of the fragment allocation requests and fragment release requests from the memory compression arrangement (401; 501).

19. The compressed memory management method (2000) as defined in claim 18, further comprising the functionality recited for the compressed memory manager (CMM) in any of claims 2-13.

Citation Information

Patent Citations

  • Space management method of flash memory device, storage control chip and flash memory device

    CN116414727A

  • Memory controller and operating method thereof

    US20190278716A1

  • Managing free space in a compressed memory system

    US20220011941A1

  • Method and apparatus for performing pre-allocation of memory to avoid triggering garbage collection operations

    US6349312B1

  • Method and system for efficiently managing a dynamic memory in embedded system

    WO2007097581A1