Page processing method and page processing system
Through fine-grained large-page perception and coordinated distribution methods, pages suitable for deletion or improvement are identified, which solves the problem of the reduction in the number of large pages in the virtualized system, and achieves efficient memory management and performance improvement.
Patent Information
- Application Number
- CN202510417752.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-11
AI Technical Summary
In the virtualized system, memory deduplication technology is greatly affected by page size, resulting in a reduction in the number of large pages, affecting the system's memory access performance, and it is difficult to maximize memory saving and maintain high performance at the same time.
By calculating the page's access frequency, write ratio, repeat ratio and zero page ratio, a fine-grained large-page perception method is used to identify pages suitable for deletion or promotion, avoid unnecessary large-page splitting, and use sliding hash comparison to obtain page features, and coordinate distribution and conversion processing.
It improves the number of large pages in the system, optimizes memory utilization, reduces memory access latency, improves overall memory access performance, and reduces waste of system resources.
Smart Images

Figure CN120295716A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer memory data storage, and particularly to a page processing method and a page processing system. Background Art
[0002] Cloud computing provides a flexible way to access system resources, enabling rapid resource configuration and release with little user intervention and almost no involvement of service providers. In recent years, cloud computing has received extensive attention in both the industrial and academic communities. Many companies choose to deploy applications and services to the cloud rather than purchasing and maintaining expensive hardware infrastructure themselves. To meet the needs of different customers, cloud computing service providers typically adopt a virtualization architecture, running multiple virtual machines (VMs) on a single physical host to achieve isolation of different customer applications and resource sharing.
[0003] With the continuous increase in cloud computing demand, the memory capacity of physical machines has become an important factor restricting the number of virtual machines, thereby affecting the number of customers that service providers can support without sacrificing performance. Therefore, developing efficient memory management technologies, especially memory deduplication, has become the key to reducing memory occupancy. Memory deduplication technology saves memory space by merging duplicate memory pages. Memory deduplication technology enables cloud service providers to provide more virtual machines to tenants, so it plays a crucial role in realizing over-allocation of memory resources.
[0004] Kernel Same-page Merging[1](KSM) is a memory deduplication technology in the Linux kernel, aiming to reduce memory occupancy. It scans the memory to find duplicate pages in anonymous pages. Once a set of duplicate pages is found, these pages will be merged, that is, they will be mapped to the same physical page. The merged page is read-only. Therefore, if a process needs to modify the page after merging, the kernel will mark the page as "copy-on-write" (COW). In this way, when the process attempts to modify the page, the kernel will automatically copy the content of the page to a newly allocated page, thus allowing modification without affecting other processes sharing the page.
[0005] Meanwhile, the application accesses data through virtual addresses, and the operating system accelerates the translation from virtual addresses to physical addresses by using the Translation Lookaside Buffer (TLB) and the Page Walk Cache (PWC). However, frequent access to the TLB makes it a hot spot structure in the processor. Increasing the TLB size will significantly increase power consumption, and every memory access requires querying the TLB, so the lookup speed must match the CPU main frequency. Therefore, to expand the coverage of the TLB, the operating system introduces Huge Pages, which increase the TLB hit rate by increasing the size of the address translation unit and reduce the length of page table traversal. In a virtualized system, due to the additional overhead of the Memory Management Unit (MMU), two-dimensional page tables make Huge Pages particularly important. Although Huge Pages can significantly improve memory access performance, they may also cause problems such as internal fragmentation, page allocation latency, and memory ballooning. Therefore, the operating system uses the asynchronous Huge Page promotion program Khugepage[2] to merge multiple base pages into a single Huge Page to improve memory efficiency, but this process also requires unified access attributes for the pages.
[0006] In a virtualized system using Huge Page memory management, although there is a large amount of duplicate data in memory, it is difficult to find duplicate Huge Pages. In other words, page-based memory deduplication technology is greatly affected by page size and usually can only delete smaller memory pages. To achieve more efficient memory deduplication, KSM splits Huge Pages containing duplicate data (such as 2MB pages) into small pages (such as 4KB pages) and performs deduplication in units of small pages. Although this strategy can save more memory space, splitting a large number of Huge Pages will significantly increase the number of page table entries in the system, thereby reducing the TLB hit rate and memory access performance. At the same time, since Khugepage has a lower priority than KSM deduplication and cannot promote Huge Pages when encountering shared pages, it cannot effectively merge frequently accessed small pages back into Huge Pages, resulting in a gradual decrease in the number of Huge Pages in a long-running system, thus affecting the overall performance.
[0007] Currently, the main challenge is how to maximize memory savings while maintaining high memory access performance for virtual machines. KSM splits Huge Pages into small pages for memory deduplication, resulting in a rapid decrease in the number of Huge Pages, while the number of shared read-only page table entries and shared pages in the system increases. These shared resources then affect the promotion of Huge Pages, leading to a continuous decline in the system's memory access performance. Summary of the Invention
[0008] The object of the present invention is to overcome the above-mentioned defects or problems in the background technology or to provide a material basis for overcoming the above-mentioned defects or problems in the background technology, and to provide a page processing method and a page processing system.
[0009] To achieve the above object, the present invention and its preferred embodiments adopt the following technical solutions, but the embodiments are not limited to the following solutions:
[0010] Solution 1, a page processing method, including the following steps:
[0011] When the redundancy is greater than the redundancy threshold and greater than the popularity, and the distinctness is greater than the distinctness threshold, and it is a large page, split the large page; when the redundancy is greater than the redundancy threshold and greater than the popularity, and the distinctness is greater than the distinctness threshold, and it is a small page, perform deduplication; when the popularity is greater than the popularity threshold and greater than the redundancy, and the distinctness is greater than the distinctness threshold, and it is a small page, enter the promotion step.
[0012] Among them, calculate the popularity based on the access frequency and write ratio of the corresponding page, calculate the redundancy based on the repetition ratio and zero-page ratio, and calculate the distinctness based on the redundancy and popularity.
[0013] Solution 2, based on Solution 1, for large pages:
[0014] Popularity = (write ratio + access frequency / 2) * number of sampled sub-pages, redundancy = (repetition ratio + zero-page ratio) * number of sampled sub-pages, distinctness = |popularity - redundancy|;
[0015] For small pages:
[0016] Popularity = ((write ratio + access frequency) / 2) * number of small pages within the large page range, redundancy = (repetition ratio + zero-page ratio) * number of small pages within the large page range, distinctness = |popularity - redundancy|.
[0017] Solution 3, based on Solution 2, for large pages:
[0018] Calculate the access frequency of the large page = access count / total scanning period according to the total scanning period and the access count of the access counter;
[0019] Determine the total number of sampled pages according to the interval of initialized sliding sampling and the number of fragment pages; move each sampled fragment according to the step size, and determine the sampled sub-pages that need to be hashed according to the moved sampled fragments; determine the sampled sub-pages, and perform hashing calculation on the sampled sub-pages inside the sampled large page; compare the second half of each sampled fragment with the first half of the previous sampled fragment. If it is the first sampling, no comparison is made. Determine whether the content of the sub-page has changed according to whether the hash value has changed; record the number of sub-pages whose content has changed, and combine this value with the number of sampled sub-pages to calculate the write ratio of the content that has changed = number of sub-pages whose content has changed / number of sampled sub-pages;
[0020] Hash-record the first half of each sampled segment; compare the sampled sub-page hash value with the zero page. If they are equal, increment the zero page count; record the zero page count, and calculate in combination with the number of sampled sub-pages to obtain the zero page ratio = zero page count / number of sampled sub-pages;
[0021] Perform bitwise operations of simple hashing on the hash value of the sampled sub-page to obtain multiple secondary hash values generated by this hash value; put these simple hashes into a Bloom filter; record the duplicate sub-pages calculated by the Bloom filter, and calculate in combination with the number of sampled sub-pages to obtain the duplicate ratio = number of duplicate sub-pages / number of sampled sub-pages.
[0022] Solution Four, based on Solution Two, for small pages:
[0023] Obtain how many small pages in the currently split large page have been shared by deduplication, record the number of these shared pages, and calculate the duplicate ratio = number of shared pages / number of small pages within a large page range; by calculating the small page hash and comparing it with the stored old hash value, obtain the write ratio where the content has changed during two scans = number of pages with content change / number of small pages within a large page range; by statistically counting the access frequencies of each small page, obtain the access frequency of the overall split large page = number of small pages with access bit 1 within a large page range / number of small pages within a large page range; by comparing the calculated small page with the zero page hash, obtain the zero page ratio = number of zero pages within a large page range / number of small pages within a large page range.
[0024] Solution Five, based on Solution One, calculate the popularity and redundancy of the page; and compare the redundancy and popularity of the page;
[0025] If the redundancy is greater than or equal to the popularity, then compare the redundancy and the redundancy threshold; if the redundancy is greater than the redundancy threshold, calculate the distinctness of the page access characteristics and compare the distinctness and the distinctness threshold. If the distinctness is greater than the distinctness threshold, then check whether the page is a large page. If so, perform splitting to allocate the eligible split pages to the deduplication area, screen out the cold pages for deduplication, and the distributed pages are processed; if it is a small page, directly perform deduplication, and the distributed pages are processed;
[0026] If the popularity is greater than the redundancy, then compare the popularity and the popularity threshold; if the popularity is greater than the popularity threshold, calculate the distinctness of the page access characteristics and compare the distinctness and the distinctness threshold. If the distinctness is greater than the distinctness threshold, if it is a small page, judge whether it is suitable for further promotion and perform corresponding processing, and ensure that the pages entering the promotion area have the least sharing, and the distributed pages are processed.
[0027] Solution Six, based on Solution Five, if the visibility is less than the visibility threshold, or the redundancy is greater than the popularity and less than the redundancy threshold, or the popularity is greater than the redundancy and less than the popularity threshold, or the popularity is greater than the redundancy, the popularity is greater than the popularity threshold, the visibility is greater than the visibility threshold, and it is a large page, then the page is allocated to the waiting area.
[0028] For the pages in the waiting area, rescan them after waiting for several rounds and reduce their waiting count. When the waiting count reaches 0, they leave the waiting area for processing.
[0029] For the fuzzy feature pages that are frequently accessed but rarely written, wait in the waiting area for further processing, that is, wait for the next round of scanning.
[0030] Solution Seven, based on Solution Five, further includes the step of monitoring the pages after distribution processing and performing quick conversion for the pages with unreasonable processing:
[0031] Detect whether the page needs quick conversion. For the pages that need quick conversion, select to bypass the sequential scan.
[0032] Collect and analyze the pages that do not achieve the expected effect after deduplication and promotion processing, and collect the pages.
[0033] Add the pages for quick conversion to the front of the quick queue to ensure priority processing, and limit each page to only get one priority conversion opportunity per round.
[0034] Capture the free pages released by deduplication and perform a synchronous zeroing operation when they are added to the free list. Use non-compiler optimization instructions to ensure that the released pages are correctly zeroed.
[0035] Regularly scan the cold shared pages and merge them into continuous memory areas.
[0036] Solution Eight, based on Solution Three or Solution Four, further includes an initialization step:
[0037] Initialize the scanning threads for in-memory deduplication and large page promotion, set the initial number of pages scanned each time and the sleep interval after reaching that number of pages; initialize the Bloom filter and the intervals for sliding sampling hashing, the number of sampled sub-pages, the step size, and the number of segments of the sampled pages, and calculate the hash value of the zero page; initialize the access counter for storing the page access frequency and the data structure for the sampled hash values; initialize the communication linked list for deduplication and promotion.
[0038] Solution Nine, a page processing system, which is suitable for executing a page processing method described in any one of Solutions One to Eight.
[0039] As described above for the present invention and its preferred embodiments, compared with the prior art, the technical solutions of the present invention and its preferred embodiments have the following beneficial effects due to the following technical means:
[0040] 1. The page processing method proposed by the present invention performs deduplication on highly redundant pages suitable for deduplication and promotion on hot pages suitable for promotion. While maximizing memory savings, it retains high performance by avoiding unnecessary large page splits, thereby increasing the number of large pages in the system and improving the overall memory access performance of the system. By precisely detecting and identifying different types of pages, it ensures that pages are reasonably deduplicated or promoted according to their characteristics (i.e., access frequency, write ratio, repetition ratio, and zero page ratio, as well as the calculated heat and redundancy). This method can reduce system resource waste caused by unreasonable page distribution.
[0041] 2. The page detection method based on fine-grained large page awareness proposed by the present invention can, without splitting large pages, obtain accurate page characteristics (i.e., access frequency, write ratio, repetition ratio, and zero page ratio) through sliding hash comparison of pages. It performs deduplication on highly redundant pages suitable for deduplication and promotion on hot pages suitable for promotion. While maximizing memory savings, it retains high performance by avoiding unnecessary large page splits, thereby increasing the number of large pages in the system and improving the overall memory access performance of the system.
[0042] 3. The coordinated distribution method based on page detection proposed by the present invention ensures reasonable deduplication or promotion operations on pages according to their characteristics through precise detection and identification of different types of pages. This method can reduce system resource waste caused by unreasonable page distribution. At the same time, through the setting of the waiting area, the coordinated distribution module improves the effective utilization of large pages, avoids incorrect processing of pages with fuzzy characteristics, optimizes the system's memory allocation strategy, improves memory utilization, and reduces memory access latency.
[0043] 4. The adaptation conversion method based on coordinated distribution proposed by the present invention can preferentially process specific pages by combining fine-grained page awareness and coordinated distribution mechanisms, avoiding the processing delay caused by sequential scanning in traditional methods. This method ensures that pages that need to be quickly converted can bypass long waiting times and complete the conversion quickly, improving the system's adaptability to large pages. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0045] Figure 1 The system structure module diagram of the present invention: The present invention includes a fine-grained large page detection module, a page coordination and distribution module, and a page conversion and adaptation module. When starting the scan, each structure is initialized to place the newly added content in the deduplication and promoted system. Through fine-grained detection, the sampled sub-pages in the large page are statistically analyzed with low overhead. After obtaining sufficiently distinct page features, the page is subjected to coordinated distribution processing and a waiting area is set. At the same time, the overhead during page conversion is optimized through the adaptation and conversion module, especially creating conditions for friendly allocation of large pages.
[0046] Figure 2 The page detection module based on fine-grained large page perception of the present invention: The present invention realizes the comprehensive detection and estimation of the internal sub-pages of the large page through the fine-grained detection module. Among them, the sliding window method is used to sample the sub-pages, perform low-overhead statistics, generate page features, and improve the detection efficiency of zero pages by comparing with the zero hash value. The design of this module can effectively detect redundant data in the large page and avoid the deficiency that the fixed sampling method cannot cover all sub-pages.
[0047] Figure 3 The coordination and distribution module based on page detection of the present invention: The present invention classifies the page features generated by the detection module through the coordination and distribution module, and uses the waiting area to temporarily store the pages with fuzzy features to avoid their being blindly processed and causing performance loss. For the pages with distinct features, they are classified according to popularity or repetition rate, and deduplication and promotion operations are respectively performed to improve the overall memory utilization rate and reduce unnecessary overhead.
[0048] Figure 4 The adaptation and conversion module based on coordination and distribution of the present invention: The present invention optimizes the overhead during the conversion of memory pages through the adaptation and conversion module, especially creating conditions for friendly allocation of large pages. This module first performs pre-conversion layout optimization on the pages to be converted according to the page classification results provided by the coordination and distribution module to reduce fragmentation.
[0049] Figure 5 The flowchart of the page detection part of the present invention: First, relevant structures are initialized, and the sub-pages in the large page are sampled through the sliding window method. After calculating the checksum or hash value of the sampled sub-pages, redundancy detection is performed. Record the status and continue the detection to ensure that the data features in the large page are comprehensively and accurately captured.
[0050] Figure 6Flowchart of processing a page according to the present invention: According to the feature data provided by the page detection module, the present invention classifies and processes the page. Pages with prominent features are directly distributed or subjected to duplicate deletion operations, while pages with ambiguous features are stored in the waiting area to avoid performance losses. After executing the page processing strategy on the classification result, the page status is updated and the operation result is recorded to optimize subsequent memory scheduling and allocation strategies. Detailed implementation manners
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are the preferred embodiments of the present invention and should not be regarded as excluding other embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0052] In the claims, the description and the above-mentioned drawings of the present invention, unless otherwise clearly defined, when using terms such as "first", "second", or "third", etc., are for distinguishing different objects rather than for describing a specific order.
[0053] In the claims, the description and the above-mentioned drawings of the present invention, unless otherwise clearly defined, for orientation terms, when using terms such as "center", "horizontal", "vertical", "level", "vertical", "top", "bottom", "inner", "outer", "up", "down", "front", "rear", "left", "right", "clockwise", "counterclockwise", etc. to indicate the orientation or position relationship, it is based on the orientation and position relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation, so it should not be construed as limiting the specific protection scope of the present invention.
[0054] In the claims, the description and the above-mentioned drawings of the present invention, unless otherwise clearly defined, when using terms such as "fixed connection" or "fixedly connected", should be understood in a broad sense, that is, any connection method without displacement relationship and relative rotation relationship between the two, that is, including non-removable fixed connection, removable fixed connection, being integrated as one, and being fixed connected through other devices or elements.
[0055] In the claims, the description and the above-mentioned drawings of the present invention, when using terms such as "comprising", "having" and their variants, are intended to mean "including but not limited to".
[0056] Refer to Figures 1-6 ;
[0057] A page processing system, which is suitable for a page processing method.
[0058] The described page processing method includes the following steps:
[0059] (1) Initialize the data structures required for scanning and subsequent operations.
[0060] Specifically, it includes the following steps:
[0061] (1.1) Initialize the scanning threads ksmd and khugepaged for memory deduplication and large page promotion, set the initial number of pages scanned each time and the sleep interval after reaching that number of pages. For example, set the initial number of pages scanned each time to 10000 and the sleep interval after reaching that number of pages to 20 ms;
[0062] (1.2) Initialize the Bloom filter, as well as the interval, number of sampled sub - pages, step size, and number of segments of the sampled page for sliding sampling hashing, and calculate the hash value zerohash of the zero page for subsequent comparison. For example, the interval of sliding sampling hashing can be set to 28, the number of sampled sub - pages to 64, the step size to 2, and the number of segments of the sampled page to 4.
[0063] Among them, a zero page is a page that is all 0. For example, when the system allocates a large page for an application, it clears all the pages inside this large page. If the application only uses a part of the pages inside the large page, then the other part is still a zero page. The existence of more zero pages indicates some waste in the large page.
[0064] (1.3) Initialize the access counter for storing the access frequency of pages and the data structure pnode of the sampled hash value to store information such as the page access frequency Fac, page redundancy R, and heat H obtained during the scanning process;
[0065] (1.4) Initialize the communication linked list plist for deduplication and promotion to index the pages that need to be quickly converted;
[0066] (2) Scan the memory pages of each process to detect the characteristics of large and small pages
[0067] Specifically, it includes the following steps:
[0068] (2.1) Scan the memory pages in all memory areas registered for deduplication and large page promotion, classify the scanned pages, and conduct fine - grained detection when a large page is captured. When a small page is captured, jump to step (2.15) to detect its heat and redundancy in a system - supported manner;
[0069] (2.2) Obtain the large page features as follows. According to the access counter, obtain the access frequency Fac of the large page during the scan. When the large page is accessed during this scan, increment the counter and clear the page access flag for the next record. Calculate the access frequency of the large page = access count / total scan period based on the total scan period and the access count of the access counter;
[0070] (2.3) Determine the total number of sampled pages according to the initialized sliding sampling interval and the number of fragmented pages;
[0071] (2.4) Move each sampled segment seg according to the step size, and determine the sampled sub - pages that need to be hashed according to the moved sampled segment (each large page contains several sampled sub - pages);
[0072] (2.5) Determine the sampled sub - pages and perform hash calculation on the sampled sub - pages inside the specified large page;
[0073] (2.6) As Figure 2 , compare the second half segβ of each sampled segment with the first half sub - page segα of the previous sampled segment by hashing. If it is the first sampling, no comparison is made. Determine whether the content of the sub - page has changed according to whether the hash value has changed;
[0074] (2.7) Record the number of sub - pages whose content has changed, and combine this value with the number of sampled sub - pages to calculate and obtain the write ratio Rw = number of sub - pages whose content has changed / number of sampled sub - pages;
[0075] (2.8) Hash - record the first half sega of each sampled segment for sliding comparison during subsequent scan hash calculations;
[0076] (2.9) Compare the hash value of the sampled sub - page with the zero page. If they are equal, increment the number of zero pages;
[0077] (2.10) Record the number of zero pages, and combine it with the number of sampled sub - pages to calculate and obtain the zero page ratio Rz = number of zero pages / number of sampled sub - pages;
[0078] (2.11) Perform bitwise operations of simple hashing on the hash value of the sampled sub - page to obtain multiple secondary hash values generated by this hash value;
[0079] (2.12) Put these simple hashes (i.e., the secondary hash values described in 2.11) into the Bloom filter;
[0080] (2.13) When all the bits corresponding to these simple hashes in the Bloom filter are greater than or equal to 1, it proves that this page is very likely to have a duplicate page, otherwise there must be no duplicate page;
[0081] (2.14) Record the duplicate sub - pages calculated by the Bloom filter, combine with the number of sampled sub - pages to calculate, and obtain the duplicate ratio Rd = the number of duplicate sub - pages / the number of sampled sub - pages; thus, the large - page detection process ends;
[0082] (2.15) The following is to obtain the characteristics of small pages. First, obtain how many small pages in the currently split large page have been shared by deduplication, record the number of these shared pages, and calculate the duplicate ratio Rd = the number of shared pages / the number of small pages within the range of one large page (the split large page);
[0083] (2.16) By calculating the small - page hash and comparing it with the stored old hash value, obtain the write ratio Rw = the number of pages with content changes / the number of small pages within the range of one large page during two scans;
[0084] (2.17) By counting the access frequency of each small page, obtain the access frequency Fac of the overall split large page = the number of small pages with the access bit set to 1 within the range of one large page / the number of small pages within the range of one large page;
[0085] (2.18) By comparing the calculated small - page hash with the zero - page hash, obtain the zero - page ratio Rz = the number of zero pages within the range of one large page / the number of small pages within the range of one large page. Thus, the detection of small pages ends.
[0086] (3) Perform page distribution processing according to page characteristics
[0087] When the redundancy is greater than the redundancy threshold and greater than the popularity, and the distinctiveness is greater than the distinctiveness threshold and it is a large page, split the large page; when the redundancy is greater than the redundancy threshold and greater than the popularity, and the distinctiveness is greater than the distinctiveness threshold and it is a small page, perform deduplication; when the popularity is greater than the popularity threshold and greater than the redundancy, and the distinctiveness is greater than the distinctiveness threshold and it is a small page, enter the promotion step.
[0088] Among them, calculate the popularity based on the access frequency and write ratio of the corresponding page, calculate the redundancy based on the duplicate ratio and zero - page ratio, and calculate the distinctiveness based on the redundancy and popularity.
[0089] Specifically, it includes the following steps:
[0090] (3.1) Initialize the page - characteristic calculation formula, and define the redundancy R and the popularity H;
[0091] (3.2) Collect page - ratio information, including the zero - page ratio Rz, the duplicate ratio Rd, the access frequency Fac, and the write ratio Rw;
[0092] (3.3) Calculate the access frequency Fac and write ratio Rw of the page, and comprehensively calculate the page heat H. For large pages: Heat H = (write ratio + access frequency / 2) * number of sampled sub - pages. Since the number of overlapping pages in two scans for large pages is half of the number of sampled pages, the maximum write ratio is 0.5 and the maximum access frequency is 1. Therefore, the access frequency needs to be divided by 2. The zero - page ratio and the duplicate ratio are mutually exclusive in the detection. Therefore, the sum of the zero - page ratio and the duplicate ratio is at most 1. For small pages: Heat H = ((write ratio + access frequency) / 2) * number of small pages within the large - page range (i.e., the number of sampled sub - pages of small pages). For small pages, the maximum values of both the write ratio and the access frequency are 1. Therefore, the overall value needs to be divided by 2. The zero - page ratio and the duplicate ratio are still mutually exclusive in the calculation.
[0093] (3.4) Statistically calculate the duplicate ratio Rd and zero - page ratio Rz of the page, and comprehensively calculate the page redundancy R. Redundancy R = (duplicate ratio + zero - page ratio) * number of sampled sub - pages. For small pages: Redundancy R = (duplicate ratio + zero - page ratio) * number of small pages within the large - page range (i.e., the number of sampled sub - pages of small pages). Through the above coefficients, the redundancy R and the heat H can be compared.
[0094] (3.5) Compare the redundancy and heat of the page to determine whether its characteristic is more inclined to redundancy or hot access. If the page is inclined to redundancy, that is, the redundancy is greater than or equals the heat, then continue. Otherwise (i.e., the heat is greater than the redundancy), go to step (3.7).
[0095] (3.6) Compare the redundancy R and the redundancy threshold ThR to determine whether the redundancy R is greater than the redundancy threshold ThR. If so, enter (3.8) for further judgment. Otherwise, jump to step (3.10) to enter the waiting area.
[0096] (3.7) Compare the heat H and the heat threshold ThH to determine whether the heat H is greater than the heat threshold ThH. If so, enter (3.8) for further judgment. Otherwise, jump to step (3.10) to enter the waiting area.
[0097] (3.8) Calculate the distinctness Dv of the page access characteristics. Distinctness Dv = |heat - redundancy|.
[0098] (3.9) Compare the distinctness Dv and the distinctness threshold ThDv to determine whether the characteristic distinctness Dv is lower than the distinctness threshold ThDv. If so, allocate the page to the waiting area. Otherwise, enter the deduplication area and jump to (3.11) or enter the promotion area and jump to (3.13). Specifically, in the case where the redundancy is greater than or equal to the heat and greater than the redundancy threshold, enter the deduplication area and jump to (3.11). Or in the case where the heat is greater than the redundancy and greater than the heat threshold, enter the promotion area and jump to (3.13).
[0099] (3.10) For the pages in the waiting area, after waiting for several rounds, rescan and update their characteristics, that is, reduce their waiting count. When the waiting count reaches 0, leave the waiting area and proceed with processing, then jump to step (3.15).
[0100] (3.11) Check whether the page is a large page. If so, split it; otherwise (i.e., if it is a small page), directly perform deduplication.
[0101] (3.12) Allocate the split pages that meet the conditions in (3.11) to the deduplication area, screen out the cold pages for deduplication, and then jump to (3.16).
[0102] (3.13) For small pages, determine whether they are suitable for promotion and perform corresponding processing. For example: if a page is pinned by the kernel, it cannot be promoted, and there are also many other types of pages that cannot be promoted, which is related to the specific implementation of this function in the kernel.
[0103] If it is a large page, enter the waiting area.
[0104] (3.14) Ensure that the pages entering the promotion area have the least sharing, that is, break all the shared pages in the page to be promoted, and break all the shared pages through copy-on-write, that is, ensure that there are no shared pages in the promoted page, so as to optimize the effect and consistency of the promotion operation, and then jump to step (3.16).
[0105] (3.15) Process the fuzzy feature pages that are frequently accessed but have few writes, and wait in the waiting area for further processing, that is, wait for the next round of scanning.
[0106] (3.16) The processing of the distributed pages is completed.
[0107] (4) Monitor the pages after distribution processing, and perform quick conversion on the pages with unreasonable processing.
[0108] (4.1) Detect whether the page needs quick conversion. For example, after a large page is deduplicated through the above process, maybe only one or two sub-pages are deduplicated. At this time, it indicates that the previous decision is inappropriate and needs to be quickly converted back to a large page. That is, if after the above steps of deduplication or promotion, the expected effect is not achieved, then it needs to be processed through quick conversion.
[0109] Therefore, select to bypass the sequential scan according to the current page state.
[0110] (4.2) Collect and analyze the pages that do not achieve the expected effect after deduplication and promotion processing.
[0111] (4.3) Add the page to be quickly converted to the front end of the quick queue plist (i.e., the above-mentioned communication linked list), ensure priority processing, and limit each page to only obtain one priority conversion opportunity per round to avoid resource contention;
[0112] (4.4) Capture the free pages released by deduplication, and perform synchronous zeroing operations when they are added to the free list freelist to reduce the risk of memory pollution;
[0113] (4.5) Disable the "dead store" optimization of the compiler, and use non-compiler optimization instructions to ensure that the released pages are correctly zeroed to avoid the invalidation of the cleaning operation caused by compiler optimization;
[0114] (4.6) Regularly scan the cold shared pages (such as scanning the cold shared pages after each round of deduplication ksmd is completed), and merge them into continuous memory areas through compact to alleviate the memory fragmentation problem and create continuous space for the large page allocation of page_alloc.
[0115] 1. The page detection method based on fine-grained large page awareness proposed by the present invention can, without splitting large pages, obtain accurate page characteristics (i.e., access frequency, write ratio, repetition ratio, and zero page ratio) through sliding hash comparison of pages, perform deduplication on highly redundant pages suitable for deduplication, and promote hot pages suitable for promotion. While maximizing memory savings, it retains high performance by avoiding unnecessary large page splitting, thereby increasing the number of large pages in the system and improving the overall memory access performance of the system.
[0116] 2. The coordinated distribution method based on page detection proposed by the present invention ensures that pages are reasonably deduplicated or promoted according to their characteristics through precise detection and identification of different types of pages. This method can reduce the waste of system resources caused by unreasonable page distribution. At the same time, through the setting of the waiting area, the coordinated distribution module improves the effective utilization of large pages, avoids incorrect processing of pages with fuzzy characteristics, optimizes the system's memory allocation strategy, improves memory utilization and reduces memory access latency.
[0117] 3. The adaptation conversion method based on coordinated distribution proposed by the present invention can preferentially process specific pages by combining fine-grained page awareness and coordinated distribution mechanisms, avoiding the processing delay caused by sequential scanning in traditional methods. This method ensures that pages that need to be quickly converted can bypass the long waiting time and complete the conversion quickly, improving the system's adaptability to large pages.
[0118] The descriptions of the foregoing specification and embodiments are used to explain the scope of protection of the present invention, but do not constitute a limitation to the scope of protection of the present invention. Modifications, equivalent replacements or other improvements to the embodiments of the present invention or some of its technical features obtained by those of ordinary skill in the art through logical analysis, reasoning or limited experiments in combination with common general knowledge, ordinary technical knowledge in the art and / or the prior art under the inspiration of the present invention or the foregoing embodiments shall be included within the scope of protection of the present invention.
Claims
1. A page processing method, characterized in that: Including the following steps: When the redundancy is greater than the redundancy threshold and greater than the heat, and the distinctness is greater than the distinctness threshold, and it is a large page, split the large page; when the redundancy is greater than the redundancy threshold and greater than the heat, and the distinctness is greater than the distinctness threshold, and it is a small page, perform deduplication; When the heat is greater than the heat threshold and greater than the redundancy, and the distinctness is greater than the distinctness threshold, and it is a small page, enter the promotion step, wherein, the heat is calculated based on the access frequency and write ratio of the corresponding page, the redundancy is calculated based on the repetition ratio and zero-page ratio, and the distinctness is calculated based on the redundancy and heat.
2. A page processing method according to claim 1, wherein: For large pages: Heat = (write ratio + access frequency / 2) * number of sampled sub-pages, redundancy = (repetition ratio + zero-page ratio) * number of sampled sub-pages, distinctness = |heat - redundancy|; For small pages: Heat = ((write ratio + access frequency) / 2) * number of small pages within the large page range, redundancy = (repetition ratio + zero-page ratio) * number of small pages within the large page range, distinctness = |heat - redundancy|.
3. A page processing method according to claim 2, wherein: For large pages: Calculate the access frequency of the large page = access count / total scanning period according to the total scanning period and the access count of the access counter; Determine the total number of sampled pages according to the interval of initialized sliding sampling and the number of fragment pages; Move each sampled fragment according to the step size, and determine the sampled sub-pages that need to be hashed according to the moved sampled fragment; Determine the sampled sub-pages, and perform hashing calculation on the sampled sub-pages inside the sampled large page; compare the second half of each sampled fragment with the first half of the sub-pages of the previous sampled fragment. If it is the first sampling, no comparison is made. Determine whether the content of the sub-page has changed according to whether the hash value has changed; record the number of sub-pages whose content has changed, and combine this value with the number of sampled sub-pages to calculate and obtain the write ratio of the content that has changed = number of sub-pages whose content has changed / number of sampled sub-pages; Hash-record the first half of each sampled fragment; compare the hash value of the sampled sub-page with the zero page. If they are equal, increase the number of zero pages; record the number of zero pages, and combine it with the number of sampled sub-pages to calculate and obtain the zero-page ratio = number of zero pages / number of sampled sub-pages; Perform bitwise operations of simple hashing on the hash value of the sampled sub-page to obtain multiple secondary hash values generated by this hash value; Put these simple hashes into a Bloom filter; Record the repeated sub-pages calculated by the Bloom filter, and combine it with the number of sampled sub-pages to calculate and obtain the repetition ratio = number of repeated sub-pages / number of sampled sub-pages.
4. A page processing method according to claim 2, wherein: For small pages: Obtain the number of small pages within the currently split large page that have been shared by deduplication, record the number of these shared pages, and calculate the duplication ratio = the number of shared pages / the number of small pages within a large page range; by calculating the small page hash and comparing it with the stored old hash value, obtain the write ratio of content changes during two scans = the number of content changes / the number of small pages within a large page range; by counting the access frequency of each small page, obtain the access frequency of the overall split large page = the number of small pages with the access bit set to 1 within a large page range / the number of small pages within a large page range; by comparing the calculated small page hash with the zero page hash, obtain the zero page ratio = the number of zero pages within a large page range / the number of small pages within a large page range.
5. A page processing method as claimed in claim 1, wherein: Calculate the heat and redundancy of the page; and compare the redundancy and heat of the page; If the redundancy is greater than or equal to the heat, then compare the redundancy with the redundancy threshold; if the redundancy is greater than the redundancy threshold, then calculate the distinctiveness of the page access characteristics and compare the distinctiveness with the distinctiveness threshold. If the distinctiveness is greater than the distinctiveness threshold, then check whether the page is a large page. If so, perform splitting to allocate the eligible split pages to the deduplication area, screen out the cold pages for deduplication, and the distributed page processing is completed; If it is a small page, directly perform deduplication, and the distributed page processing is completed; If the heat is greater than the redundancy, then compare the heat with the heat threshold; If the heat is greater than the heat threshold, then calculate the distinctiveness of the page access characteristics and compare the distinctiveness with the distinctiveness threshold. If the distinctiveness is greater than the distinctiveness threshold, if it is a small page, determine whether it is suitable for further promotion and perform corresponding processing, and ensure that the pages entering the promotion area have the least sharing, and the distributed page processing is completed.
6. A page processing method as claimed in claim 5, wherein: If the distinctiveness is less than the distinctiveness threshold, or the redundancy is greater than the heat and the redundancy is less than the redundancy threshold, or the heat is greater than the redundancy and the heat is less than the heat threshold, or the heat is greater than the redundancy and the heat is greater than the heat threshold and the distinctiveness is greater than the distinctiveness threshold and it is a large page, then allocate the page to the waiting area. For the pages in the waiting area, rescan after waiting for several rounds and reduce their waiting count. When the waiting count reaches 0, leave the waiting area and perform processing; For the fuzzy characteristic pages with frequent access but few writes, wait in the waiting area for further processing, that is, wait for the next round of scanning.
7. A page processing method as claimed in claim 5, wherein: It further includes the step of monitoring the pages after distribution processing and performing quick conversion for the unreasonably processed pages: Detect whether the page needs quick conversion, and for the pages that need quick conversion, select to bypass the sequential scan; Collect and analyze the pages that have not achieved the expected effect after deduplication and promotion processing, and collect the pages; Add the pages for quick conversion to the front of the quick queue to ensure priority processing, and limit each page to only obtain one priority conversion opportunity per round. Capture the free pages released by deduplication, perform synchronous zeroing operations when they are added to the free list, and use non-compiler optimization instructions to ensure that the released pages are correctly zeroed; Periodically scan cold shared pages and merge them into contiguous memory regions.
8. A page processing method according to claim 3 or 4, characterized in that: It further includes an initialization step: Initialize the scan threads for in-memory deduplication and large page promotion, set the initial number of pages scanned each time and the sleep interval after reaching that number of pages; initialize the Bloom filter and the intervals for sliding sampling hashing, the number of sampled subpages, the step size, and the number of segments of the sampled pages, and calculate the hash values of zero pages; Initialize the access counter for storing the page access frequency and the data structure for the sampled hash values; initialize the communication linked list for deduplication and promotion.
9. A page processing system, characterized in that: It is adapted to execute a page processing method according to any one of claims 1-8.