Dynamic granularity compression method and device based on memory page association rule
By adopting a dynamic granular compression method based on memory page association rules on mobile devices, page association relationships are mined and compressed, the problem of tight memory resources of mobile devices and the waste of CPU bandwidth is solved, and the system performance and user experience are significantly improved.
Patent Information
- Application Number
- CN202510419192.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-24
AI Technical Summary
Memory resources on mobile devices are tight, the existing page-based memory compression method wastes CPU bandwidth, introduces significant context switching overhead, and cannot effectively mine the association rules of memory pages.
A dynamic granular compression method based on memory page association rules is adopted. By collecting and analyzing the memory access modes of the application, a frequent pattern tree is generated using the FP-Growth association rule mining algorithm, page association relationships are mined, and memory compression of the associated pages is performed.
It significantly improves system performance, reduces response delay, improves user experience, increases the average application startup speed by 1.55 times, and increases the photography processing speed and frame rate by 1.42 times and 1.31 times respectively.
Smart Images

Figure CN120196446A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of memory management and memory compression, and particularly to a dynamic granularity compression method and device based on memory page association rules. Background Art
[0002] With the continuous growth of memory requirements for mobile applications and the steady increase in the number of cached applications, memory resources on mobile devices are becoming increasingly strained. This trend will become even more pronounced with the implementation of more and more memory-intensive tasks (such as AI, AR / VR, and Transformer). To address this issue, mobile systems (such as Android and iOS) employ memory compression technology. By compressing less important pages, space can be saved for new memory requirements. In mobile systems, when an application is running in the foreground, there are usually active applications in the background. Most users do not clear these background applications, which consume a large amount of memory resources. When memory runs out, the mobile system reclaims memory pages according to the design principles of Linux, such as the LRU (Least Recently Used) algorithm. However, this memory activity by background processes can have a negative impact on foreground processes.
[0003] In the past few decades, research on optimizing memory compression has mainly focused on improving compression algorithms, reducing their impact, or leveraging application characteristics for optimization. Existing solutions mainly perform memory compression at the page granularity, which is reasonable because memory management and access are both based on pages. However, page-based compression methods waste CPU bandwidth and introduce significant context switching overhead. Research shows that when the system compresses frequently, the response latency may increase by up to 2.31 times.
[0004] By analyzing the page access characteristics of popular applications, it is found that there is a high degree of correlation between a large number of pages. These pages are either frequently accessed together or rarely accessed. This internal correlation is implicit and difficult to directly detect. If the association rules of memory pages can be mined and highly correlated pages can be compressed together, system performance is expected to be further improved while minimizing the impact of read amplification. Summary of the Invention
[0005] To overcome the problems existing in the above technologies, the object of the present invention is to provide a dynamic granularity compression method and device based on memory page association rules. The basic idea is to collect and analyze the memory access patterns of applications, obtain page association relationships based on data mining methods, and perform memory compression on associated pages. It consists of three key parts: a footprint stream generator, a frequent pattern tree linked list, and an adaptive compression region.
[0006] The specific technical solution for achieving the object of the present invention is as follows:
[0007] A dynamic granularity compression method based on memory page association rules, comprising the following steps:
[0008] Step 1: Maintain a first-in-first-out circular queue to record the access history of anonymous pages; when an anonymous page is accessed, obtain its physical address and insert it as an "item" into the circular queue; all items in the list form a data stream;
[0009] Step 2: Add a new daemon process in the kernel as an independent control unit to convert the order record of items in the data stream into transactions; use a fixed-size sliding window to generate transactions, and pack the generated transaction order stream to form a "footprint stream";
[0010] Step 3: Adopt the FP-Growth association rule mining algorithm to maintain a frequent pattern tree to represent the association relationship between application anonymous pages; each frequent pattern tree represents an application; a branch path of the tree represents a transaction, and except for the root node, each node represents an anonymous page; the root node is used to distinguish different applications, record the UID of the application, and point to the inactive LRU linked list; the root nodes are sorted according to the application activity, and the less active the application, the closer it is to the head node of the inactive LRU linked list; the frequent pattern tree is constructed in descending order of the occurrence frequency of transactions, and the overlapping parts of the same branch paths between different transactions are used to keep the tree compact; when an anonymous page appears in a transaction, update its count and adjust the sorting of tree nodes according to its occurrence frequency;
[0011] Step 4: When the memory requirement is small, follow the system's page reclaiming process to separately remove and compress the inactive anonymous pages from the inactive LRU linked list; when there is a large memory requirement, scan the associated pages under the corresponding UID of each frequent pattern tree root node one by one from beginning to end;
[0012] Step 5: Batch move the associated pages to the swap cache, merge them into a large block, call the compression algorithm to compress them into smaller blocks, and store them in the compressed area; when accessing the compressed area, perform different decompressions according to the difference in granularity during compression.
[0013] Further, the use of a fixed-size sliding window to generate transactions in step two specifically includes:
[0014] When the window size is w, there are three cases for generating transactions:
[0015] (1) When there are X ≥ w records in the queue, pack these w records into a transaction and remove them from the queue, then move the queue head pointer forward by w steps, and the window slides to the new head position;
[0016] (2) If records are found in the queue, the window will pause for T time; when the pause time ends, these X records are packaged into a transaction and removed, and the sliding window also moves forward X steps;
[0017] (3) If , it will wait until one of the above two conditions is met before generating a transaction.
[0018] Further, when there is a large-scale memory requirement as described in Step 4, the associated pages under the corresponding UID of the root node of each frequent pattern tree are scanned one by one from beginning to end, specifically including:
[0019] After the system enters the reclaim and compression process, the function _alloc_pages_nodemask() will be called; the parameter order is used to determine that the number of pages to be reclaimed and compressed in this round is 2 order pages; when there are N ≥ 2 order mutually associated pages in the first frequent pattern tree, these N pages will be selected at one time; if the number of associated pages in the tree is less than 2 order , then continue to check the next tree; not all pages in the frequent pattern tree are pairwise associated, and only the associated pages are selected through the grouping mechanism.
[0020] Further, as described in Step 5, the associated pages are batch-moved to the swap cache, merged into a large block, and then compressed into a smaller block by calling a compression algorithm and stored in the compression area, specifically including:
[0021] The associated pages are merged into a large block in the swap cache, and then the Block I / O of the large block is handed over to the adaptive compression engine; a compression algorithm is called to compress the Block I / O of the large block, and the adaptive compression engine will parse the structure of the large block and read the pages therein; first, each page is copied to a compression buffer, and then compressed; after compression, the binary format data in the compression buffer is converted into a smaller block; the compressed block is stored in the compression area, and its address is dynamically managed. In the kernel, Block I / O represents a block I / O request being processed.
[0022] Further, the specific method of dynamically managing its address includes:
[0023] For the blocks generated by compression, the same "object + handle" structure as ZRAM is used for management; since the blocks generated by compression usually exceed the capacity of a single 4KB slot, the blocks generated by compression are split into multiple objects and stored separately; the objects are indexed by two layers of handles, and the second layer of handle is called vhandle; for the pages compressed at the block granularity, they will carry flag bits, and the page table entries of these pages all point to the same handle; the handle points to a metadata structure, which contains two parts: an array used to record the page table information of all pages in this block; a pointer used to index the vhandles corresponding to the respective objects into which this block is split.
[0024] Further, when accessing the compressed area as described in step five, different decompressions are performed according to the difference in the granularity during compression, specifically including:
[0025] When accessing a page located in the compressed area, a page fault will be triggered; the page fault handling function is modified to additionally check the flag bit before the normal operation; if the target page is compressed at the independent 4KB granularity, the decompression method is the same as that of ZRAM; if the target page is compressed at the block granularity, the entire block containing the target page will be decompressed at one time; according to the vhandle list in the metadata, find and decompress all the objects corresponding to this block; according to the page table information in the metadata, obtain the pages corresponding to each part in the block; after decompression, return the processing result of the page fault to the upper layer, and the other non-access pages decompressed together will remain in the swap cache.
[0026] A dynamic granularity compression device based on the memory page association rule is used to implement the dynamic granularity compression method based on the memory page association rule, specifically including:
[0027] Footprint stream generator: Obtain the physical address of the accessed anonymous page and regard it as an item; convert the sequential record of the items into a transaction; the transaction sequence stream forms a footprint stream;
[0028] Frequent pattern tree linked list: Maintain multiple frequent pattern trees, representing each application; as the footprint stream is generated, each tree is updated accordingly; when there is a large-scale memory requirement, scan the associated pages of each tree one by one from head to tail, and pack the associated pages waiting for compression;
[0029] Adaptive compression area: Compress the blocks formed by packing the associated pages, put the compressed blocks into the compression area, and dynamically manage their addresses; realize decompressing the pages in the compressed blocks without destroying the system ZRAM mechanism, and be compatible with decompressing the pages compressed at the 4KB granularity.
[0030] Compared with the prior art, the method proposed by the present invention effectively alleviates the response problem caused by insufficient memory and significantly improves the user experience. Experiments have shown that the average startup speed of the application has increased by 1.55 times, and the photography processing speed and frame rate have increased by 1.42 times and 1.31 times respectively. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a flowchart for generating the transaction footprint stream of the footprint stream generator module;
[0032] Figure 2 It is a flowchart for address management of the adaptive compression area module;
[0033] Figure 3 It is a flowchart for the operation of the device of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0034] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0035] Embodiment
[0036] As Figure 3 shown, it is a flowchart for the operation of the dynamic granularity compression device based on the memory page association rule. Three components are designed: the footprint stream generator, the frequent pattern tree linked list, and the adaptive compression area. The footprint stream generator collects the page access traces and packs them into a transaction footprint stream. The frequent pattern tree linked list generates and updates the frequent pattern tree according to the footprint stream. The associated pages mined are compressed and stored and managed by the adaptive compression area.
[0037] Refer to Figure 1 , which is a flowchart for generating the footprint stream of the footprint stream generator module of the present invention. First, when an anonymous page is accessed, its physical address is obtained and inserted into a circular queue as an item. Next, a fixed-size sliding window is used to generate transactions. Finally, the generated transaction sequence stream is packed to form a footprint stream. Specifically, assume that the following physical page addresses are accessed in sequence: 0x3E, 0x29, 0x313, 0x32, 0x10, 0x2F, 0x45, 0x3E, 0x29, 0x32, 0x10, 0x2F, 0x11, 0x3E, 0x29, 0x13, 0x28, 0x29, 0x13, 0x37, 0x2F, 0x28, 0x3E, 0x51, 0x29, 0x10, 0x10. The sliding window size w and the queue length Lq are 6 and 10 respectively, X is the number of items in the current queue, t0 is the time point of the earliest access, and t26 is the time point of the latest access. tᵢ represents the moment when the queue head pointer is at the i-th record. At t5, since 6 pages are accessed within T time (i.e., t5 t0 < T, X = 6), the window packs these 6 pages into a transaction and clears them from the queue, then slides the window forward 6 steps, and then generates a new transaction at t11. After that, another 4 new records (X = 4) arrive before timeout, so at t11 + T, these 4 records (0x11, 0x3E, 0x29, 0x13) are packed into a transaction. After that, two more transactions are generated in the same way at t15 + T and t26 respectively.
[0038] The frequent pattern tree linked list module has a backbone composed of several nodes. The first node points to the inactive LRU linked list of the anonymous page, and the remaining nodes serve as the root nodes of each frequent pattern tree. Different from the traditional frequent pattern tree with an empty node as the root, the root nodes in the frequent pattern tree linked list module are used to distinguish different applications and are represented by recording the UID of the application. The associated pages of the same application will be mounted into the frequent pattern tree under its corresponding UID. As the system runs, the footprint stream will be continuously generated, and each tree in the frequent pattern tree linked list module will be updated accordingly. These UID nodes on the backbone of the frequent pattern tree linked list module will be sorted according to the application activity, and the less active the application, the closer it is to the head of the linked list of the frequent pattern tree linked list module. During normal use, the original page recycling process is still followed: the system directly removes and compresses the inactive pages from the inactive LRU linked list, which is the same as the original scheme. When there is a large-scale memory requirement, the associated pages under the corresponding UID will be scanned one by one from the head to the tail of the frequent pattern tree linked list module and packed for compression.
[0039] In the adaptive compression area module, all candidate associated pages are batch-moved into the swap cache. They will be merged into a block in the swap cache, and then the Block I / O (BIO) of this block is given to the compression engine. In the kernel, BIO represents a block I / O request being processed. The adaptive compression area module calls a compression algorithm to compress the BIO of this block. The compression engine will parse the block structure and read the pages in it. Since the compression algorithm needs to operate on continuous memory during implementation, the adaptive compression area module will first copy each physical page to a buffer and then perform compression. After compression, the binary format data in the buffer is converted into a smaller block. Since the number of pages in the buffer is uncertain, the size of the compressed block is also not fixed. The adaptive compression area module stores the compressed block in the compression area and dynamically manages its address. Refer to Figure 2, which is the address management flowchart of the adaptive compression region module of the present invention. Page1 and page2 are compressed separately at the page granularity (narrow track), while page3 to page6 are combined into a block (wide track) for common compression. The compression result of the latter is divided into object3 and object4, which are indexed by two handles. For the sake of distinction, the second-level handle is called vhandle. When accessing a page in the adaptive compression region module, a page fault will be triggered. By modifying the page fault handling function, it is made to additionally check the page flag bit before the regular operation. If the target page is compressed at the independent granularity, the decompression method is the same as that of ZRAM; otherwise, the adaptive compression region module will decompress the entire data block containing the target page at one time. According to the vhandle list, all the objects corresponding to this large block can be found and decompressed. Since the metadata records the correspondence between these compressed pages and their page table entries, the adaptive compression region module knows which page each part of the block corresponds to. After decompression, the adaptive compression region module returns the result of handling the page fault to the upper layer, and the other decompressed pages will remain in the swap cache.
[0040] This implementation incurs three types of overhead: read amplification, energy consumption, and memory overhead. Read amplification overhead refers to the phenomenon where the actual amount of data that needs to be read by the system to satisfy a single user request is more than the amount of data requested. If the read amplification is too high, it will lead to an increase in the I / O burden and affect the system performance. The evaluation results show that 92.6% of the batch decompressed pages can be accessed quickly, so the read amplification ratio does not exceed 1.08. In addition, batch decompression will prolong the decompression latency of the target page. The experimental results show that the time to decompress 64 pages is 17.6% longer than that of a single page, but the whole process can still be completed within the sub-microsecond level. More importantly, the access latency of the pre-fetched pages (i.e., the associated pages decompressed in advance) is effectively reduced. Considering the overall performance gain and the high accuracy of association mining, the impact of batch decompression on the system performance can be ignored. The energy consumption overhead is found through experiments that the average energy consumption only increases by 0.69%. Compared with the continuous power consumption of the touch screen, this energy consumption impact can be ignored. The memory overhead is divided into three parts. First, only an extra memory space is occupied when a variable in the metadata points to the vhandle, while the vhandle itself structure is retained in the original kernel, so no extra memory consumption is introduced. Assuming that there are 1,048,576 anonymous pages (about 4GB of memory) in the system, and 30% of the pages are highly correlated. Based on this calculation, only an extra 384KB is needed to store the metadata. In addition, when the associated page is decompressed, the corresponding variable will be deleted, so the metadata will not accumulate infinitely. Second, a part of the memory is used for the item circular queue and the transaction ring buffer. Each item in the circular queue uses an unsigned long integer, that is, 64 bits. The sliding window width is defaulted to 32, so the size of each transaction is 64×32 bits. The evaluation shows that caching 128 items and 512 transactions can meet most cases. Therefore, the item circular queue consumes 1KB (128 × 64 bits), and the transaction ring buffer consumes 128KB (512 × 32 × 64 bits). Finally, the statistical results show that the number of pages in the same association group generally does not exceed 32, so the upper limit of the compression buffer is set to 128KB (32 × 4KB). Generally speaking, these memory overheads in the order of hundreds of KB are acceptable compared with the GB-level memory available on the device.
Claims
1. A dynamic granularity compression method based on memory page association rules, characterized in that: The following steps are involved: Step 1: Maintain a first-in-first-out circular queue to record the access history of anonymous pages; when an anonymous page is accessed, obtain its physical address and insert it into the circular queue as an "item"; all items in the list constitute a data stream; Step 2: Add a daemon process to the kernel as an independent control unit to convert the sequential records of items in the data stream into transactions; use a fixed-size sliding window to generate transactions, and package the generated transaction sequence stream into a footprint stream; Step 3: Use the FP-Growth association rule mining algorithm to maintain a frequent pattern tree to represent the association relationship between anonymous application pages; each frequent pattern tree represents an application; A branch path of the tree represents a transaction. Except for the root node, each node represents an anonymous page. The root node is used to distinguish different applications, record the application UID, and point to the inactive LRU list. The root nodes are sorted according to the activity of the application. The less active the application, the closer it is to the head node of the inactive LRU list. The frequent pattern tree is constructed in descending order of transaction frequency. The same branch paths between different transactions are partially overlapped to maintain the compactness of the tree. When an anonymous page appears in a transaction, its count is updated, and the tree node sorting is adjusted according to its frequency of occurrence. Step 4: When the memory demand is small, the system's page recycling process is followed to remove and compress the inactive anonymous pages from the inactive LRU list. When the memory demand is large, the associated pages under the corresponding UID of each frequent pattern tree root node are scanned one by one from beginning to end. Step 5: Move the related pages to the swap cache in batches, merge them into a large block, call the compression algorithm to compress them into smaller blocks, and store them in the compressed area; when accessing the compressed area, perform different decompressions based on the difference in granularity during compression.
2. The dynamic granularity compression method according to claim 1, characterized in that: Step 2 uses a fixed-size sliding window to generate transactions, including: When the window size is w, there are three situations for generating transactions: (1) When there are X ≥ w records in the queue, these w records are packaged into a transaction and removed from the queue. Then the queue head pointer moves forward w steps, and the window slides to the new head position. (2) If found in the queue records, the window will pause T time; when the pause time is over, these X records are packaged into a transaction and removed, and the sliding window also moves forward X steps; (3) If , it will wait until one of the above two conditions is met before generating a transaction.
3. The dynamic granularity compression method according to claim 1, characterized in that: When a large-scale memory requirement occurs as described in step 4, the associated pages under the corresponding UID of each frequent pattern tree root node are scanned one by one from beginning to end, including: After the system enters the recycling and compression process, the function _alloc_pages_nodemask() will be called; the parameter order is used to determine the number of pages that need to be recycled and compressed in this round, which is 2. order Page; when N ≥ 2 in the first frequent pattern tree order If there are less than 2 related pages in the tree, these N pages will be selected at once. order , then continue to check the next tree; not all pages in the frequent pattern tree are associated with each other, and only associated pages are selected through the grouping mechanism.
4. The dynamic granularity compression method according to claim 1, characterized in that: Step 5 moves the associated pages to the swap cache in batches, merges them into a large block, compresses them into smaller blocks using a compression algorithm, and stores them in the compression area. Specifically, the following steps are performed: Merge the associated pages into a large block in the swap cache, and then hand over the large block of Block I / O to the adaptive compression engine; call the compression algorithm to compress the large block of Block I / O, and the adaptive compression engine will parse the structure of the large block and read the pages in it; first copy each page to a compression buffer, and then compress it; after compression, the binary format data in the compression buffer is converted into a smaller block; store the compressed block in the compression area, and dynamically manage its address; where Block I / O represents a block I / O request being processed.
5. The dynamic granularity compression method according to claim 4, characterized in that: The dynamic management of the address specifically includes: For the blocks generated by compression, the same "object + handle" structure as ZRAM is used for management; because the blocks generated by compression usually exceed the capacity of a single 4KB slot, the compressed blocks are divided into multiple objects and stored separately; objects are indexed by two levels of handles, the second level of handle is called vhandle; for pages compressed at block granularity, they will have a mark bit, and the page table entries of these pages all point to the same handle; the handle points to a metadata structure, which contains two parts: an array used to record the page table information of all pages in this block; a pointer used to index the vhandle corresponding to each object divided by this block.
6. The dynamic granularity compression method according to claim 1, characterized in that: In step 5, when accessing the compressed area, different decompressions are performed according to the difference in granularity during compression, including: When a page located in the compressed area is accessed, a page fault interrupt will be triggered; the page fault interrupt processing function has been modified to make it check the mark bit before normal operation; if the target page is compressed with an independent 4KB granularity, the decompression method is the same as ZRAM; if the target page is compressed with a block granularity, the entire block containing the target page will be decompressed at one time; according to the vhandle list in the metadata, all objects corresponding to the block are found and decompressed; according to the page table information in the metadata, the pages corresponding to each part of the block are obtained; after decompression, the processing result of the page fault interrupt is returned to the upper layer, and other non-accessed pages decompressed together will remain in the swap cache.
7. A dynamic granularity compression device based on memory page association rules, used to implement any one of the dynamic granularity compression methods based on memory page association rules of claims 1-6, characterized in that: Specifically include: Footprint stream generator: obtains the physical address of the anonymous page being accessed and regards it as an item; converts the sequential record of the item into a transaction; the transaction sequence stream forms a footprint stream; Frequent pattern tree linked list: maintains multiple frequent pattern trees to represent each application; as the footprint stream is generated, each tree is updated accordingly; when large-scale memory requirements arise, the associated pages of each tree are scanned from beginning to end, and the associated pages are packaged for compression; Adaptive compression area: compresses blocks packed with related pages, puts the compressed blocks into compression areas, and dynamically manages their addresses; decompresses pages in compressed blocks without destroying the system ZRAM mechanism, and is compatible with decompressing pages compressed with 4KB granularity.