Memory page migration method and device and related equipment

By obtaining the shared and read-write properties of memory pages and optimizing the memory migration strategy, the system performance issues caused by TLB Shootdown are resolved, achieving more efficient memory page migration and more stable system operation.

CN120670339APending Publication Date: 2025-09-19CHINA TELECOM CORP LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510706114.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In existing memory migration solutions under non-uniform memory access architectures, frequent TLB shootdowns lead to large-scale IPI interrupt storms, increase the synchronization delay of multi-core processors, and affect system performance.

Method used

By obtaining the sharing and read-write properties of memory pages, the migration priority is determined, and the migration task is added to the corresponding migration queue. The page table with low overhead TLB Shootdown is processed first, reducing the negative performance impact of the migration operation on the running load.

Benefits of technology

Improves memory page migration efficiency, reduces TLB Shootdown overhead, and improves overall system performance and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670339A_ABST
    Figure CN120670339A_ABST
Patent Text Reader

Abstract

The invention provides a memory page migration method and device and related equipment, and is applied to the technical field of computers. The method comprises the following steps: acquiring shared attribute information of a memory page; determining a migration priority of the memory page according to the shared attribute information of the memory page; and adding the migration task of the memory page into a migration queue corresponding to the migration priority based on the migration priority of the memory page. By reducing the TLB (Transport Layer Block) Shaotdown overhead of memory page migration, the negative performance influence on the running load is reduced, and the memory page migration efficiency and the overall migration performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a memory page migration method, apparatus, and related equipment. Background Art

[0002] In modern computer architecture, memory page migration technology is an important means to optimize system performance and improve resource utilization. Especially in the non-uniform memory access (NUMA) architecture, the demand for reducing remote access latency by dynamically adjusting the physical location of memory pages is becoming increasingly prominent.

[0003] However, existing publicly available memory migration solutions are generally based on static policy designs, typically targeting only load balancing or locality optimization, without fully considering the impact of the overhead introduced by the migration operation on overall system performance. For example, during memory migration, to ensure memory consistency, stale page table mappings between processor cores must be cleared through a TLB (Translation Lookaside Buffer) flush mechanism, a process known as a TLB shootdown. Traditional solutions, whether asynchronous or synchronous, rely on inter-processor interrupts (IPIs) to trigger TLB flushes. In workloads across NUMA nodes, frequent TLB shootdowns can lead to large-scale IPI interrupt storms, significantly increasing synchronization latency in multi-core processors. While asynchronous migration can reduce migration blocking time by delaying flushes, the accumulated unhandled TLB invalid entries can trigger cascading flushes in subsequent migration phases, exacerbating system jitter. Therefore, improving memory page migration efficiency while minimizing the negative performance impact on the running workload is a pressing technical challenge.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0005] The present disclosure provides a memory page migration method, apparatus, and related devices, which can improve the efficiency of memory page migration and reduce the negative performance impact on the running load.

[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0007] According to one aspect of the present disclosure, a memory page migration method is provided, which includes: obtaining shared attribute information of a memory page; determining a migration priority of the memory page based on the shared attribute information of the memory page; and adding the migration task of the memory page to a migration queue corresponding to the migration priority based on the migration priority of the memory page.

[0008] In some possible embodiments, the method further includes: obtaining the read and write attributes of the memory page; determining the migration priority of the memory page based on the shared attribute information of the memory page includes: determining the migration priority of the memory page based on the shared attribute information and read and write attributes of the memory page.

[0009] In some possible embodiments, obtaining the read and write properties of the memory page includes: sampling the last-level cache miss event to obtain the read and write operations of the virtual address during the last-level cache miss event; determining the read and write ratio of the memory page based on the read and write operations of the virtual address during the last-level cache miss event; determining the read and write properties of the memory page based on the read and write ratio of the memory page; or periodically scanning the global virtual page table to obtain the access frequency of the access bit and the dirty bit in the global virtual page table, and determining the read and write ratio of the memory page based on the access frequency of the access bit and the dirty bit in the global virtual page table; determining the read and write properties of the memory page based on the read and write ratio of the memory page; wherein, when the read ratio is greater than or equal to the write ratio, the read and write properties of the memory page are determined to be more read and less write; when the read ratio is less than the write ratio, the read and write properties of the memory page are determined to be more write and less read.

[0010] In some possible embodiments, the sharing attributes include shared and private, and the migration priority of the memory page is determined based on the sharing attribute information and read-write attributes of the memory page, including: when the sharing attribute information of the memory page is shared and the read-write attribute is more writes than reads, determining the migration priority of the memory page to be the first priority; when the sharing attribute information of the memory page is private and the read-write attribute is more writes than reads, determining the migration priority of the memory page to be the second priority, and the second priority is greater than the first priority; when the sharing attribute information of the memory page is shared and the read-write attribute is more reads than writes, determining the migration priority of the memory page to be the third priority, and the third priority is greater than the second priority; when the sharing attribute information of the memory page is private and the read-write attribute is more reads than writes, determining the migration priority of the memory page to be the fourth priority, and the fourth priority is greater than the third priority.

[0011] In some possible embodiments, the method further includes: obtaining a first heat index of the memory page; and adding the migration task of the memory page to a migration queue corresponding to the migration priority based on the migration priority of the memory page, including: adding the migration task of the memory page to a migration queue corresponding to the migration priority based on the first heat index of the memory page and the migration priority of the memory page.

[0012] In some possible embodiments, the method further includes: obtaining a second heat index of the memory page; when the second heat index of the memory page is greater than the first heat index, updating the migration priority of the memory page; and transferring the migration task of the memory page to a migration queue corresponding to the updated migration priority according to the updated migration priority of the memory page.

[0013] In some possible embodiments, before obtaining the shared attribute information of the memory page, the method further includes: in response to creating a thread in the process, establishing a first memory page table for the thread.

[0014] In some possible embodiments, the first memory page table includes a private memory page table and a shared memory page table, and the shared memory page table is a last-level page table shared with other threads in the process.

[0015] In some possible embodiments, the method further includes: establishing a global virtual page table shared by multiple threads for the process, wherein the global virtual page table contains a private memory page table entry placeholder and a shared memory page table entry in the process, and the private memory page table entry placeholder is used to indicate a private relationship between the private memory page table entry and the corresponding thread.

[0016] In some possible embodiments, the private memory page table is pointed to by a page global directory pointer in a process descriptor.

[0017] In some possible embodiments, the method further includes: when a thread in a process accesses an unallocated memory page in a private memory page table, in response to the CPU where the thread is located failing to find a mapping relationship between the virtual address of the unallocated memory page and the physical address, triggering a cache miss; in response to triggering the cache miss, performing a page table traversal in the private memory page table; and when there is no mapping relationship between the virtual address of the unallocated memory page and the physical address in the private memory page table, triggering a page fault interrupt.

[0018] In some possible embodiments, the method further includes: in response to triggering a page fault interrupt, traversing the global virtual page table to search for the virtual address corresponding to the page fault interrupt; when the virtual address corresponding to the page fault interrupt exists in the global virtual page table, obtaining the memory page corresponding to the page fault interrupt; allocating a page table page for the memory page corresponding to the page fault interrupt in the private memory page table corresponding to the thread, and sharing the last-level page table of the private memory page table with the global virtual page table; when the virtual address corresponding to the page fault interrupt does not exist in the global virtual page table, allocating a page table page for the virtual address corresponding to the page fault interrupt in the private memory page table corresponding to the thread, and creating a page table page for the virtual address corresponding to the page fault interrupt in the global virtual page table.

[0019] In some possible embodiments, the method further includes: marking the memory page corresponding to the page fault interrupt as a multi-thread shared item in the global virtual page table.

[0020] In some possible embodiments, the method further includes: recording the thread identifier of the thread in the page table entry corresponding to the page table page in the global virtual page table, wherein the thread identifier is used to indicate that the mapping relationship contained in the page table entry is a private item of the thread.

[0021] According to another aspect of the present disclosure, a memory page migration device is also provided, which includes: an acquisition module for acquiring shared attribute information of a memory page; a determination module for determining the migration priority of the memory page based on the shared attribute information of the memory page; and a adding module for adding the migration task of the memory page to a migration queue corresponding to the migration priority based on the migration priority of the memory page.

[0022] According to another aspect of the present disclosure, an electronic device is also provided, which includes: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned memory page migration methods by executing the executable instructions.

[0023] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the memory page migration method described above is implemented.

[0024] According to another aspect of the present disclosure, a computer program product is provided, including: a computer program or instructions, wherein when the computer program or instructions are executed by a processor, any one of the above-mentioned memory page migration methods is implemented.

[0025] A memory page migration method, apparatus and related equipment are provided in the embodiments of the present disclosure. The method includes: obtaining the shared attribute information of the memory page; determining the migration priority of the memory page according to the shared attribute information of the memory page; and adding the migration task of the memory page to the migration queue corresponding to the migration priority based on the migration priority of the memory page. The TLB Shootdown overheads corresponding to different shared attributes of the present disclosure are different. For example, if the shared attribute is a page table shared by multiple threads, the TLB Shootdown overhead is relatively large, and if the shared attribute is a page table shared by private threads, the TLB Shootdown overhead is relatively small. The migration priority is determined based on the shared attribute information, and the page tables to be migrated with low-overhead TLB Shootdown are screened out and processed first. Then, the migration task is added to the corresponding queue based on the migration priority. The TLB Shootdown overhead of memory page migration is reduced, the negative performance impact on the running load is reduced, and the memory page migration efficiency and overall migration performance are improved.

[0026] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0028] Figure 1 A schematic diagram showing the system architecture of a memory page migration method according to an embodiment of the present disclosure is shown;

[0029] Figure 2A A flow chart of a memory page migration method according to an embodiment of the present disclosure is shown;

[0030] Figure 2B A schematic diagram illustrating the relationship between a thread-level page table and a global virtual page table of a thread in an embodiment of the present disclosure is shown.

[0031] Figure 3 A flow chart of another memory page migration method according to an embodiment of the present disclosure is shown;

[0032] Figure 4 A flow chart of a method for adding a migration task of a memory page to a migration queue corresponding to a migration priority according to an embodiment of the present disclosure is shown;

[0033] Figure 5A flow chart showing another method for adding a migration task of a memory page to a migration queue corresponding to a migration priority according to an embodiment of the present disclosure is shown;

[0034] Figure 6 A flow chart of a method for accessing a memory page according to an embodiment of the present disclosure is shown;

[0035] Figure 7 A flow chart of another method for accessing a memory page according to an embodiment of the present disclosure is shown;

[0036] Figure 8 A flowchart of a specific method for memory page migration according to an embodiment of the present disclosure is shown;

[0037] Figure 9 A schematic diagram of a memory page migration device according to an embodiment of the present disclosure is shown;

[0038] Figure 10 A structural block diagram of an electronic device in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0039] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0040] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0041] For ease of understanding, before introducing the embodiments of the present disclosure, several terms involved in the embodiments of the present disclosure are first explained as follows:

[0042] Dynamic Random Access Memory (DRAM): A type of computer memory used to temporarily store running applications and data. When the CPU needs to process a task, the relevant program and data are transferred to DRAM, allowing the CPU to quickly access this data and efficiently execute instructions.

[0043] Memory paging technology: Currently, memory paging technology (paging) basically uses a set of logic that cooperates with software and hardware to convert process virtual addresses to actual physical memory addresses. Taking the x86 platform as an example, the memory paging mechanism divides the virtual address space and physical memory into fixed-size blocks, called pages and page frames, and uses page tables to manage the mapping relationship between them. Each process has an independent virtual address space, divided into pages of equal size (such as 4KB), and the physical memory is also divided into page frames of the same size. The page table data structure records the physical page frame address corresponding to each virtual page.

[0044] Memory pages are the basic unit of memory management in the operating system. They divide the virtual memory space and physical memory space of a process into blocks of fixed size. Pages in virtual memory correspond one-to-one with page frames in physical memory. The operating system uses page tables to record the mapping relationship between virtual memory pages and physical memory page frames. When a process accesses memory, if the required page is not in physical memory (page missing), the operating system will load the corresponding page from external storage (such as disk) into physical memory. The memory page mechanism implements virtual memory, allowing processes to use a larger address space than physical memory, thereby improving memory utilization. At the same time, the swapping of pages allows multiple processes to share limited physical memory resources, improving the system's concurrent processing capabilities.

[0045] The memory page table is a key data structure in the virtual memory system, managing the translation of virtual addresses to physical addresses. Each application has its own page table to ensure memory isolation and security. The operating system is responsible for maintaining the page table to ensure the accuracy and efficiency of address translation. To achieve efficient lookups, page tables often use a radix tree data structure. Taking a four-level page table as an example, its structure is divided into: PGD (Page Global Directory), also known as PML4 (Page Map Level-4), serving as the top-level directory, pointing to the next level; PUD (Page Upper Directory), also known as PDP (Page Directory Pointer), further refining the address mapping; PMD (Page Middle Directory), also known as PD (Page Directory), further decomposing the address; and PT (Page Table), ultimately pointing to the physical memory page. Each level stores partial address information, and through level-by-level lookups, the virtual address is ultimately translated into a physical address. This hierarchical design saves space and improves lookup speed.

[0046] Radix Tree: A multi-branch search tree. The leaf nodes of the tree store the actual data entries, that is, the content to be searched or stored in the end. Each non-leaf node has a certain number of pointers pointing to child nodes, and this number depends on how the key is divided. Its core feature is to use the common prefix of the key to reduce storage space and search time. For example, when storing strings that begin with "apple" or "appetite", the first few characters can share a node, and each string does not need to be stored in full. When searching, starting from the root node, the pointers are matched in sequence according to the characters of the key to quickly locate the target leaf node, improving search efficiency.

[0047] Last-Level Cache (LLC): This is the last level of cache in the processor core. To improve data access speed and reduce main memory access time, the processor divides the cache into multiple levels, represented by L1, L2, and L3. L1 is the fastest and closest to the processor core, but has a smaller capacity; L2 is slightly slower and has a larger capacity than L1; L3, as the last level of cache, is the slowest compared to L1 and L2, but has the largest capacity.

[0048] Page-Table Entry (PTE): A key data structure used in operating systems to manage virtual memory. Each PTE records the mapping between a virtual page in a process's virtual memory space and a specific physical page in physical memory. In addition, PTEs contain several control bits that manage the access permissions and state of physical memory. For example, the Access bit indicates whether the physical page has been accessed, while the Dirty bit indicates whether the contents of the physical page have been modified.

[0049] Memory Management Unit (MMU): It is computer hardware responsible for processing the CPU's memory access requests and converting the virtual address given by the CPU into a physical address.

[0050] Translation Lookaside Buffer (TLB): It is a key component of the memory management unit (MMU). The function of the TLB is to cache the mapping relationship between recently used virtual addresses and physical addresses. When the MMU performs address translation, it first checks the TLB. If a matching mapping is found, the physical address can be obtained directly without accessing the page table, which greatly improves the translation speed. If it is not found, that is, TLB Miss, it is necessary to traverse the page table in the memory to perform address translation. The page table records the mapping of all virtual addresses to physical addresses. After the translation is completed, the new mapping relationship will be refilled into the TLB for subsequent quick query, improving system performance. In this way, the TLB effectively reduces the number of page table accesses and improves the overall system performance.

[0051] TLB Miss: When the CPU accesses memory, it first checks the TLB. If it hits, the physical address is quickly obtained. If it misses (TLB Miss), it needs to traverse the page table through the MMU (memory management unit) to find the corresponding page table entry in the memory to obtain the physical address. Nowadays, the demand for large memory resources has increased, and the capacity of the hardware TLB is limited, making it difficult to cover all memory address mappings, resulting in frequent TLB Miss. Each TLB Miss requires multiple memory accesses to complete the page table traversal, which increases the memory access latency and reduces the load performance. In short, the TLB accelerates address translation like a cache. With large memory, the TLB is easily insufficient, and TLBMiss will slow down address search, affecting program performance.

[0052] Page table traversal: The virtual address is divided into a page number and an offset within the page. The page number is used to find the corresponding entry in the page table, while the offset within the page is used to calculate the specific physical address. If the accessed virtual address has not yet been mapped to physical memory, a page fault interrupt will be triggered. The operating system needs to handle this interrupt by allocating a new page frame and updating the page table to ensure that the virtual address is correctly mapped to the physical address.

[0053] Inter-Processor Interrupts (IPI): An interrupt sent by one processor to another processor.

[0054] TLB Shootdown: When the operating system or system components modify the process page table, in order to ensure the correctness of address translation, the relevant entries in the TLB cached by the CPU need to be invalidated. For example, the x86 architecture does not provide consistency guarantees for TLB table entries across multiple CPUs at the hardware level, and this needs to be implemented in software. When the page table is modified, the relevant CPUs are notified through IPI (Inter-Processor Interrupt) to mark the entries related to the page table in their own TLB as invalid, preventing these CPUs from using outdated TLB table entries for address translation, thereby ensuring the consistency and correctness of address translation in a multi-CPU environment.

[0055] Non-Uniform Memory Access (NUMA) is a memory architecture design for multi-processor systems. In a NUMA architecture, the physical location of the processor and memory significantly affects memory access latency. Each processor has its own local memory. When a processor accesses local memory, the data transmission path is short, latency is low, and speed is high. However, when accessing non-local memory (i.e., memory allocated to other processors), data must travel a more complex transmission path, resulting in increased latency and relatively slow access speeds.

[0056] The specific implementation of the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.

[0057] Figure 1 FIG. 1 shows an exemplary application system architecture diagram to which the memory page migration method according to the embodiment of the present disclosure can be applied. Figure 1 As shown, the system architecture may include a terminal device 101 , a network 102 and a server 103 .

[0058] The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103 , and can be a wired network or a wireless network.

[0059] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPSec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.

[0060] The terminal device 101 can be various electronic devices, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, wearable devices, augmented reality devices, virtual reality devices, etc.

[0061] Optionally, the client of the application installed in different terminal devices 101 is the same, or the client of the same type of application based on different operating systems. Based on the different terminal platforms, the specific form of the client of the application can also be different, for example, the application client can be a mobile phone client, a PC client, etc.

[0062] The server 103 may be a server that provides various services, such as a background management server that provides support for the devices operated by the user using the terminal device 101. The background management server may analyze and process the received request and other data, and feed back the processing results to the terminal device.

[0063] Optionally, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0064] Those skilled in the art will know that Figure 1 The number of terminal devices, networks, and servers in the embodiment is merely illustrative, and any number of terminal devices, networks, and servers may be provided based on actual needs. This embodiment of the present disclosure does not limit this.

[0065] Under the above system architecture, an embodiment of the present disclosure provides a memory page migration method, which can be executed by any electronic device with computing and processing capabilities.

[0066] In some embodiments, the memory page migration method provided in the embodiments of the present disclosure can be executed by a terminal device of the above-mentioned system architecture; in other embodiments, the memory page migration method provided in the embodiments of the present disclosure can be executed by a server in the above-mentioned system architecture; in other embodiments, the memory page migration method provided in the embodiments of the present disclosure can be implemented by the terminal device and the server in the above-mentioned system architecture through interaction.

[0067] Figure 2A A flow chart of a memory page migration method according to an embodiment of the present disclosure is shown as follows: Figure 2A As shown, the memory page migration method provided in the embodiment of the present disclosure includes the following steps:

[0068] S202: Acquire sharing attribute information of a memory page.

[0069] In some embodiments, multiple threads sharing the same memory page can save resources, but thread operations may interfere with each other, especially when the memory page is frequently accessed. If migration operations hinder normal business processes, this can cause a certain degree of performance degradation. Therefore, unlike the current situation where multiple threads share the same memory page, this embodiment, before obtaining the shared attribute information of the memory page, further includes: in response to creating a thread in the process, establishing a first memory page table for the thread.

[0070] In this embodiment, a process is the basic unit of resource allocation and has its own independent address space and other resources. A thread is the basic unit of CPU scheduling, and multiple threads can share process resources. When a new thread is created, a relatively independent page table, namely the first memory page table, is maintained for the new thread. By allowing each thread to have its own independent first memory page table, this embodiment reduces mutual interference, ensures the normal operation of business processes, and improves overall system performance and stability.

[0071] In some embodiments, the first memory page table includes a private memory page table and a shared memory page table. The shared memory page table is the last-level page table shared with other threads in the process. For example, using a four-level page table as an example, a new thread has its own independent PML4, PDP, and PD, ensuring a certain degree of independence and reducing interference caused by page table operations between threads. However, the new thread shares the last-level page table (such as PT) with other threads, which can save memory space to a certain extent.

[0072] In some embodiments, a global virtual page table shared by multiple threads is also established for the process, wherein the global virtual page table contains a placeholder for private memory page table entries and a shared memory page table entry in the process, and the placeholder for the private memory page table entry is used to indicate a private relationship between the private memory page table entry and the corresponding thread.

[0073] In this embodiment, taking the Linux system as an example, task_struct is the Linux kernel process descriptor, representing a process. mm_struct is used to represent the address space of the process, and each process has an independent address space. pgd is the page global directory pointer. The task_struct originally associated the process address space information through mm_struct. The pgd pointer in mm_struct points to the page global directory. The page global directory is the first level of the multi-level page table. It stores pointers to the next level page table. Through these multi-level page tables, the virtual address of the process can be mapped to the physical memory address, thereby realizing the process's access and management of the memory. Simply put, the pgd pointer of mm_struct is the key starting point for mapping the process address space to the physical memory.

[0074] Specifically, when the Linux system establishes a global virtual page table shared by multiple threads for a process, it adds a new PML4 pseudo-pointer (denoted as GMM, Global Memory Manager) to the Linux task_struct structure, which represents the shared page table structure between threads. The GMM is a global page table information summary for the multiple threads contained in the process. The GMM is also manifested as a page table hierarchy, namely the global virtual page table, which contains shared last-level page table entries for all threads and placeholders for non-shared (i.e., private) page table entries. These placeholders are used to indicate the thread-private relationship of the private memory page table entries represented, reflecting the global virtual page table structure from the perspective of a process.

[0075] In this embodiment, taking the Linux system as an example, when establishing a global virtual page table shared by multiple threads, the concept of GMM (Global Memory Manager) is introduced. PML4 (Page Map Level 4) is the top-level page table in the page table hierarchy under the x86-64 architecture. What is mentioned here is that the task_struct is newly added with a PML4 pseudo pointer (GMM) representing the shared page table structure between threads, which means adding a pointer type member variable to the task_struct structure. The shared page table structure (GMM) pointed to by this pointer allows multiple threads to share the same set of virtual address to physical address mapping relationships. In this way, when accessing memory, multiple threads can quickly locate memory pages through this shared page table structure, reducing the overhead of memory management and improving the execution efficiency of multi-threaded programs.

[0076] Figure 2B A schematic diagram of the relationship between a thread-level page table and a global virtual page table for a thread provided in an embodiment of the present disclosure. Figure 2BAs shown in the figure, the structure of the virtual address is as follows: bits 47:39: used to index the PML4 (Page Map Level 4) table. Bits 38:30: used to index the PDP (Page Directory Pointer) table. Bits 29:21: used to index the PD (Page Directory) table. Bits 20:12: used to index the PT (Page Table) table. Bits 11:0: offset, directly mapped to the lower 12 bits of the physical address. CPU 0 and CPU 1: represent two different threads, each with its own CR3 register pointing to its own PML4 table (i.e., PGD, Page Global Directory). The global virtual page table has its own PML4E table, which is used for process-level shared page table queries. It should be noted that, considering the physical address mapping relationship of a virtual address, the offsets of the virtual address with respect to PML4, PDP, and PD in its private thread-level page table and the global virtual page table are the same, but actually exist in two physical memory spaces, and PT is shared by the thread-level page table and the global virtual page table.

[0077] Furthermore, the ignored field in the x86 architecture, located at bits 52:58 in the PTE corresponding to the virtual address, is used in this embodiment to maintain the thread sharing flag for the PTE, reflecting the inter-thread sharing of the mapping relationship, given that the hardware does not use this segment's contents. This bit segment has 7 bits and can accommodate a total of 127 thread IDs (excluding all 1s). When the number of threads in a process exceeds 127, the contents of the flag field become invalid. For example, 0x1 indicates private to thread 1, while 0x7F indicates shared to threads. When a process forks, both the global virtual page table and the thread-private page table are copied for subsequent copy-on-write. Fork is a system call that creates a new process. When a process calls the fork function, the system creates a child process that is nearly identical to the parent process. In this way, in a multi-threaded environment, the ignored field in the PTE can be used to mark inter-thread sharing and handle page table copying and sharing issues when forking between processes. This design ensures independence between threads while enabling necessary sharing, allowing for rapid access to shared attribute information for memory pages, improving system performance and resource utilization.

[0078] In some embodiments, the private memory page table is pointed to by the page global directory pointer in the process descriptor. The process descriptor is a data structure used by the operating system to manage processes, of which the page global directory pointer is an important component. The page global directory is a key data structure for implementing virtual address to physical address conversion, which points to page tables at all levels. The private memory page table is unique to each process and is used to manage the mapping relationship of the private memory pages of the process. The private memory page table is pointed to by the page global directory pointer in the process descriptor, indicating the access path to the private memory page table, that is, it is pointed to by the page global directory pointer in the process descriptor. When performing memory page migration, the system finds the corresponding private memory page table based on this path, so that it can accurately obtain and modify the relevant mapping information of the process memory page, ensuring that the memory page migration operation can accurately locate the memory page of the target process and perform effective processing. Exemplarily, the private memory page table of the thread is pointed to by the pgd pointer in the mm_struct in the task_struct, which is originally used to represent the address space.

[0079] In this embodiment, for a page table entry with thread-private permissions, when other threads in the same process attempt to modify the page table entry (including mapping, permission modification, etc.), it is only necessary to initiate a TLB Shootdown operation to the CPU where the private thread of the memory page is located, without initiating it to other CPUs, thereby reducing TLB Shootdown overhead. That is, through the collaborative design of hierarchical page table privatization and global virtual page table, while ensuring thread independence, the overhead of TLB Shootdown and lock contention is minimized. It inherits the resource efficiency of the shared memory model and significantly improves system performance and stability in a multi-threaded environment.

[0080] In this embodiment, the global virtual page table (GVT) is used to record information about memory pages. It can be used to obtain the sharing attribute information of memory pages, that is, to determine whether the memory page is shared or private. If it is private, it means that it is exclusively used by a specific process; if it is shared, it can be shared by multiple processes.

[0081] S204: Determine the migration priority of the memory page according to the shared attribute information of the memory page.

[0082] In this embodiment, private memory pages are used by only one process, so the TLB Shootdown scope involved in modification is small and the overhead is low. Shared memory pages are shared by multiple processes, and when modified, multiple cores need to be notified to invalidate the corresponding TLB entries, which results in high overhead. This embodiment determines the migration priority based on the shared attribute information of the memory page. Because the TLB Shootdown overhead of private memory pages is small, most of the memory migratable page tables with low-overhead TLB Shootdown are private memory pages. The system will prioritize migrating such page tables to reduce the performance loss caused by TLB Shootdown. That is, the migration priority of private memory pages is higher than that of shared memory pages.

[0083] S206: Add the migration task of the memory page to a migration queue corresponding to the migration priority based on the migration priority of the memory page.

[0084] In this embodiment, after determining the memory page migration priority, the memory page migration task is added to the corresponding migration queue according to the determined memory page migration priority. High-priority private memory page migration tasks are added to the corresponding high-priority queue, and low-priority shared memory page migration tasks are added to the low-priority queue. Subsequent migration is performed in the order of the queues.

[0085] In this embodiment, by considering the shared state of memory pages, migration overhead-aware memory page migration priority is formulated, thereby improving the efficiency of memory page migration and reducing the negative performance impact of memory page migration on the load.

[0086] In some embodiments, it is considered that the read-write properties of memory pages directly affect the migration success rate and resource consumption. For "hot pages" (pages that are frequently modified) with a high overwrite rate, the asynchronous migration strategy is prone to fall into a "migration-failure-retry" cycle due to continuous dirtying of the pages, which not only wastes memory bandwidth, but also aggravates the I / O pressure of the target node due to repeated copying. Although synchronous migration can ensure the atomicity of migration by freezing page writes, it will directly block the business process, resulting in a surge in tail latency. In addition, the migration of shared pages requires coordination of data consistency between multiple access nodes. If the migration strategy is not dynamically adjusted according to the degree of sharing, it may cause large-scale lock contention or redundant data transmission, further amplifying performance overhead. In view of this, the present disclosure also takes into account the read-write properties of memory pages when performing memory page migration, so as to optimize the migration strategy, avoid the above problems, improve migration efficiency, and reduce resource consumption and performance loss.

[0087] In some embodiments, Figure 3 A flow chart of another memory page migration method provided by the embodiment of the present disclosure. Figure 3 As shown, the memory page migration method provided by the embodiment of the present disclosure also includes the following steps:

[0088] S302: Obtain the read and write attributes of the memory page.

[0089] In this embodiment, there are multiple possible implementations for obtaining the read / write attributes of a memory page.

[0090] In some embodiments, by sampling the last-level cache miss event, the read and write operations of the virtual address during the last-level cache miss event can be obtained; based on the read and write operations of the virtual address during the last-level cache miss event, the read and write ratio of the memory page can be determined; and based on the read and write ratio of the memory page, the read and write properties of the memory page can be determined.

[0091] For example, the last-level cache (LLC) miss events can be sampled with the help of PEBS (Precise Event Sampling) technology. When an LLC Miss caused by Store (write operation) or Read (read operation) occurs, the corresponding virtual address is recorded. By counting the read and write operations of these virtual addresses, the read and write ratio of the memory page is calculated. Finally, the read and write properties of the memory page are determined based on this read-write ratio, so as to understand the usage tendency of the memory page and provide a basis for the subsequent formulation of memory page migration strategies. Specifically, if the number of read operations is greater than or equal to the number of write operations, the memory page is classified as a "read more, write less" mode; conversely, if the number of read operations is less than the number of read operations, it is classified as a "write more, read less" mode.

[0092] In some embodiments, the global virtual page table can also be periodically scanned to obtain the access frequency of the access bit and dirty bit in the global virtual page table, and the read-write ratio of the memory page can be determined based on the access frequency of the access bit and dirty bit in the global virtual page table; the read-write attributes of the memory page can be determined based on the read-write ratio of the memory page.

[0093] For example, the global virtual page table is periodically scanned to obtain the access frequencies of the access bit and dirty bit. The access bit indicates whether the page has been accessed, while the dirty bit indicates whether the page data has been modified. Based on the access frequencies of these two bits, the read / write ratio of the memory page can be determined.

[0094] For example, if the access bit has a high access frequency and the dirty bit has a low access frequency, the read-write ratio tends to be more read; conversely, it may be more write-intensive. The read-write properties of the memory page are then determined based on the read-write ratio. If reads outnumber writes, meaning the number of read operations is greater than or equal to the number of write operations, the memory page is primarily read, with fewer writes. If writes outnumber reads, meaning the number of read operations is less than the number of read operations, then writes are frequent and reads are relatively rare.

[0095] In this embodiment, by using this periodic page table scanning and setting method, the read and write attributes of the memory pages can be clarified and classified, providing a basis for the subsequent formulation of memory page migration strategies.

[0096] S304: Determine the migration priority of the memory page according to the sharing attribute information and the read / write attribute of the memory page.

[0097] In this embodiment, considering the different read and write overheads of memory pages, read operations generally have lower overhead than write operations. Read operations generally only require retrieving data from memory, while write operations may involve additional operations such as data modification and synchronization. To optimize system performance, by obtaining the read and write properties of memory pages, priority migration can be set for memory pages that are read more frequently and written less frequently; for memory pages that are written frequently, the migration level can be lowered, thereby improving the overall system operating efficiency.

[0098] In some embodiments, the sharing attributes include shared and private, and the migration priority of the memory page is determined based on the sharing attribute information and read-write attributes of the memory page, including: when the sharing attribute information of the memory page is shared and the read-write attributes are write-more-read-less, determining the migration priority of the memory page as the first priority; when the sharing attribute information of the memory page is private and the read-write attributes are write-more-read-less, determining the migration priority of the memory page as the second priority, and the second priority is greater than the first priority; when the sharing attribute information of the memory page is shared and the read-write attributes are read-more-write-less, determining the migration priority of the memory page as the third priority, and the third priority is greater than the second priority; when the sharing attribute information of the memory page is private and the read-write attributes are read-more-write-less, determining the migration priority of the memory page as the fourth priority, and the fourth priority is greater than the third priority.

[0099] For example, Table 1 is a schematic table showing how the migration priority of a memory page can be determined according to the shared attribute information and the read / write attributes of the memory page in an embodiment of the present disclosure.

[0100] Table 1

[0101] Shared properties Read and write properties Migration Priority Migration method shared Read more and write less 3 asynchronous shared Write more and read less 1 synchronous private Read more and write less 4 asynchronous private Write more and read less 2 synchronous

[0102] As shown in Table 1, four corresponding migration queues can be maintained and sampled memory pages can be added to the corresponding migration queues in sequence based on the migration priority determined by the memory page's shared attribute information and read / write attributes. Memory page migration operations are prioritized from queues with higher priorities (i.e., those with higher numbers) according to the queue's default migration policy. The basis for this migration priority is that private memory pages with a migration priority of 4, which are read-heavy and write-less, have lower TLB shootdown overhead and can maximize the performance benefits of asynchronous migration operations. Therefore, these pages are migrated first. Given that page copy overhead is greater than TLB shootdown overhead, shared memory pages with read-heavy and write-less performance are migrated first over private memory pages with write-heavy and read-less performance. Finally, shared memory pages with write-heavy and read-less performance have the lowest migration priority due to their higher memory copy overhead and TLB shootdown overhead.

[0103] In addition, as shown in Table 1, the migration method of the memory page can also be determined based on the read and write attributes of the memory page. When the read and write attributes of the memory page are more reads than writes, the asynchronous migration method is adopted; asynchronous migration means that during the migration process, the migration operation and other system operations can be performed in parallel. Because there are many read operations, the requirements for real-time consistency are relatively low, which can improve the overall efficiency of the system. When the read and write attributes of the memory page are more writes than reads, the synchronous migration method is adopted. Synchronous migration requires that the migration operation must be completed immediately to ensure data consistency. Because there are many write operations, if they are not synchronized, data inconsistency problems may occur, affecting the normal operation of the system. In this embodiment, different migration methods are selected according to the read and write attributes to achieve a balance between performance and data consistency.

[0104] In this embodiment, different migration strategies (synchronous / asynchronous) are selected based on the read and write properties of the memory pages to reduce the memory copy operation overhead, and the memory pages are divided into four categories in combination with the TLB Shootdown overhead. Migration queues are maintained for different memory pages to provide multi-queue scheduling based on memory migration priority.

[0105] In some embodiments, since different memory pages have different hotness, the heat of the memory page can also be considered when adding the migration task of the memory page to the migration queue corresponding to the migration priority. Here, "the heat of the memory page" refers to the frequency with which the page table is accessed. When different programs are running, the access patterns to the memory are different. Some page tables may be frequently accessed, that is, they have high heat, such as the page tables related to the program being executed; while some are rarely accessed and have low heat. Figure 4 A flowchart of a method for adding a migration task of a memory page to a migration queue corresponding to a migration priority is provided in an embodiment of the present disclosure, Figure 4 As shown, adding the migration task of the memory page to the migration queue corresponding to the migration priority may also include the following steps:

[0106] S402: Obtain a first heat index of a memory page.

[0107] In this embodiment, the first heat index refers to the frequency of page table access. By collecting and calculating various heat data of memory pages (such as the number of accesses, the number of modifications, etc.), the frequency of use can be quantified to obtain the first heat index.

[0108] S404 : Based on the first heat index of the memory page and the migration priority of the memory page, add the migration task of the memory page to a migration queue corresponding to the migration priority.

[0109] In this embodiment, each migration queue updates the order of memory page migration tasks in the migration queue according to the first heat index of the memory page. For example, each migration queue is arranged in descending priority based on the first heat index of the memory page, that is, page table migration tasks with high access frequency have higher priority and are higher in the queue. This can reasonably allocate resources, give priority to frequently used page table migration tasks, and improve the overall performance of the system.

[0110] In some embodiments, to prevent memory pages with continuously high heat from being unable to be migrated due to being in a suboptimal migration queue, memory pages that were not migrated in the previous round and whose heat has further increased are upgraded during each sampling. Figure 5 A flowchart of another method for adding a migration task of a memory page to a migration queue corresponding to a migration priority is provided in an embodiment of the present disclosure, combined with Figure 5 As shown, adding the migration task of the memory page to the migration queue corresponding to the migration priority may also include the following steps:

[0111] S502: Obtain a second heat index of the memory page.

[0112] In this embodiment, the second heat index is an index obtained during the sampling period for memory pages that have not completed migration in the previous round and whose heat has further increased. It reflects the change in page table heat within a specific time period.

[0113] S504: When the second heat index of the memory page is greater than the first heat index, update the migration priority of the memory page.

[0114] In this embodiment, when the second heat index of a memory page is higher than its first heat index, it indicates that the page table has become more active during the sampling period. The system will accordingly increase its migration priority, allowing it to be migrated to a more appropriate location (e.g., from a low-priority queue to a high-priority queue) more quickly to optimize performance.

[0115] For example, in the migration queue shown in Table 1, assume that memory page A is initially in the priority 1 queue. If it has not been migrated since the last sampling, and this sampling finds that its access frequency has increased significantly (the second heat index has increased), it will be moved to the priority 2 queue to ensure that frequently accessed pages are processed faster.

[0116] S506 , according to the updated migration priority of the memory page, transferring the migration task of the memory page to a migration queue corresponding to the updated migration priority.

[0117] It should be noted that, in this embodiment, the updated migration queue is sorted according to its latest popularity, but since the migration mode of the memory page is determined by the read-write attribute, the migration mode of the migration queue corresponding to the updated migration priority should be the same as the original queue.

[0118] In this embodiment, in order to prevent the memory pages of the low-priority queue from not being scheduled in time, when it is found that the popularity of such memory pages continues to increase, the queue is automatically switched for them, their migration priority is adjusted, the migration operation is accelerated, and the migration efficiency of the memory pages is further improved.

[0119] In some embodiments, in response to creating a thread in a process and establishing a first memory page for the thread, the context switch in the system scheduling is different from the existing scheme. Context switching refers to switching from one thread / process to another. The global directory pointer (PGD) is a key data structure used to convert virtual addresses into physical addresses in paged memory management. Usually, different processes have their own independent page tables, and when switching processes, the PGD pointer must be replaced to point to the page table of the new process. However, it is mentioned here that the PGD pointer is still used for assignment, but it points to the private memory page of the current thread.

[0120] A thread is part of a process, and threads within the same process can share some of the process's resources. In this embodiment, each thread has a private memory page. PGD pointing to it means that when a thread accesses memory, address mapping is completed through its own private memory page, achieving memory isolation and protection. Compared to using process page tables, this embodiment uses thread-private memory pages to give threads more independent memory space, reduce interference between threads, and improve concurrency performance. This isolates memory access between threads to a certain extent and reduces the overhead of thread switching. This is because there is no need to completely replace the page table as in traditional methods. Instead, the relevant content of the private memory page needs to be adjusted as needed, which can improve system scheduling efficiency.

[0121] During context switching, this embodiment still uses the global directory pointer in the process for assignment, but the global directory pointer points to the private memory page of the current thread instead of the process page table. By maintaining the PGD pointer assignment method and only changing the content it points to, the original address conversion process is maintained and the thread-level memory management optimization is achieved.

[0122] Figure 6 A flow chart of a method for accessing a memory page provided by an embodiment of the present disclosure. Figure 6 As shown, the method for accessing a memory page provided by an embodiment of the present disclosure includes the following steps:

[0123] S602: When a thread in a process accesses an unallocated memory page in a private memory page table, a cache miss is triggered in response to the CPU where the thread is located failing to find a mapping relationship between a virtual address and a physical address of the unallocated memory page.

[0124] In this embodiment, when a thread in a process accesses an unallocated memory page in a private memory page table, it will first look up the mapping relationship between the virtual address of the unallocated page and the physical address in the cache of the CPU where the thread is located. The cache is a high-speed cache in the CPU that is used to store the mapping relationship between the most recently used virtual address and the physical address. Its function is to speed up the address translation process and reduce the number of times the page table in the main memory is accessed. When a thread accesses memory, the CPU will first look up the corresponding mapping relationship in the cache. If the CPU where the thread is located cannot find the mapping relationship between the virtual address of the unallocated page and the physical address, it means that the corresponding entry is not found in the cache of the CPU where the thread is located, thereby triggering a cache miss TLB.

[0125] S604 , in response to triggering a cache miss, performing a page table traversal in the private memory page table.

[0126] In this embodiment, the page table is a data structure maintained by the operating system that records the complete mapping relationship between virtual addresses and physical addresses. Each process has a separate page table, stored in main memory. When the required mapping cannot be found in the cache, the CPU will search the page table instead. The MMU automatically performs page table traversal. That is, when a cache miss is triggered, the MMU automatically searches the current thread's private page table layer by layer according to the rules to find the correct page table entry, establish the mapping between virtual addresses and physical addresses, and enable the CPU to correctly access memory data.

[0127] S606: When there is no mapping relationship between the virtual address and the physical address of the unallocated memory page in the private memory page table, a page fault interrupt is triggered.

[0128] In this embodiment, the private memory page table is used to record the mapping relationship between the thread virtual address and the physical address. When a thread in a process accesses a virtual address, it first checks the private memory page table. If the physical address corresponding to the virtual address does not exist in the thread's private memory page table (which means that the page corresponding to this virtual address has not been allocated physical memory), a page fault will be triggered. A page fault is an interrupt signal. It notifies the operating system that the page the thread is trying to access is not in physical memory. At this time, the operating system will perform corresponding processing, such as transferring the required page from the disk swap area to the physical memory, and updating the page table to establish a mapping relationship between the virtual address and the newly allocated physical address. Only then can the thread continue to access the storage space corresponding to the virtual address normally.

[0129] Figure 7 A flow chart of another method for accessing a memory page provided by an embodiment of the present disclosure. Figure 7 As shown, the method for accessing a memory page provided by an embodiment of the present disclosure includes the following steps:

[0130] S702 : In response to triggering a page fault interrupt, traverse the global virtual page table to search for a virtual address corresponding to the page fault interrupt.

[0131] In this embodiment, when the operating system (OS) handles a page fault, since the global virtual page table records the virtual address mapping relationship of all threads in the process, when a page fault occurs in a thread, the virtual page table entry corresponding to the page fault can be traversed in the global virtual page table to find the virtual address that caused the page fault.

[0132] S704 , when a virtual address corresponding to a page fault exists in the global virtual page table, a memory page corresponding to the page fault is obtained.

[0133] In this embodiment, if a page table entry for the virtual address corresponding to the page fault is found in the global virtual page table, it indicates that the virtual address mapping corresponding to the page fault has been completed by other threads in the same process. The system can then use the page table entry information to retrieve the memory page corresponding to the page fault, including operations such as loading the page from disk swap into physical memory, to resolve the page fault and ensure normal process operation.

[0134] S706 , allocating a page table page for the memory page corresponding to the page fault interrupt in the private memory page table corresponding to the thread, and sharing the last-level page table of the private memory page table with the global virtual page table.

[0135] In this embodiment, when there is a virtual address corresponding to a page fault interrupt in the global virtual page table, it means that the virtual address mapping has been completed by other threads of the current process. At this time, there is no need to repeatedly establish the mapping relationship. It is only necessary to apply for page table pages of various levels on demand in the page table corresponding to the thread that currently triggers the page fault interrupt. Page table pages of various levels are multi-level page table structures adopted by the operating system, such as the common second-level or third-level page tables, to save memory space. Applying on demand means that page table pages of various levels are allocated only when needed to avoid waste. For example, only when the first-level page table is not enough will the second-level page table page be applied for. In this way, memory utilization can be improved while meeting the address mapping requirements. It can be understood that since the page table pages of various levels are used to gradually convert virtual addresses into physical addresses, there are mapping records in the global table. It is only necessary to improve the thread's own page table to complete the address conversion, so that the process can normally access the required memory pages.

[0136] It should be noted that, in this embodiment, sharing the last-level page table of the private memory page table with the global virtual page table means that when a thread triggers a page fault interrupt, if the mapping of the target virtual address already exists in the global page table (indicating that other threads have already established the mapping), the current thread does not need to repeatedly build the complete page table structure, but only allocates the necessary intermediate page table pages (such as PUD / PMD) and directly references the last-level page table entry (PTE) of the global page table. This approach avoids repeated mapping of the same physical page, saving memory overhead, and at the same time, through the write-time copy mechanism, ensures that the thread's modification of the shared page will not affect other threads. That is, only when the thread attempts to write, the kernel allocates a private copy for it, thus taking into account both efficiency and isolation.

[0137] In some embodiments, the method further includes marking a memory page corresponding to a page fault interrupt in a global virtual page table as a multi-thread shared item.

[0138] In this embodiment, the ignored field in the page table entry corresponding to the global virtual page table can be modified to all 1 bits to indicate that the page table entry is a multi-thread shared item. It can be understood that the page table is used to map virtual addresses to physical addresses, and the page table entry is each item in the page table, which records information such as the mapping relationship. The ignored field is a field in the page table entry. Setting the ignored field in the page table entry corresponding to the global virtual page table to all 1 bits means setting all bits of the ignored field to 1. In a multi-threaded environment, multiple threads may need to access certain memory areas at the same time. When the ignored field is set to all 1s, it is equivalent to marking this page table entry with a "multi-thread shared" mark. Based on the corresponding mark, the system determines that the memory page corresponding to the page table entry can be used by multiple threads, that is, the sharing attribute of the memory page is shared, and thus adopts corresponding strategies in memory management and access control to ensure the correctness and consistency of multi-threaded access.

[0139] S708, when the virtual address corresponding to the page fault does not exist in the global virtual page table, allocate a page table page for the virtual address corresponding to the page fault in the private memory page table corresponding to the thread, and create a page table page for the virtual address corresponding to the page fault in the global virtual page table.

[0140] In this embodiment, when a page fault occurs, if there is no mapping corresponding to the virtual address in the global virtual page table, it means that the virtual address has not yet been allocated physical memory. At this time, a page table page will be allocated for the virtual address in the private memory page table corresponding to the thread that triggered the page fault. The private memory page table is unique to the thread and records the mapping of its own virtual address to the physical address. It should be noted that allocating a page table page for the virtual address corresponding to the page fault interrupt may include calling the physical memory allocator to apply for a physical page (such as 4KB in size), that is, a memory page. If large pages (such as 2MB / 1GB) are supported, it is necessary to decide whether to allocate large pages based on the mapping requirements. Clear the newly allocated physical page to ensure that the initial state of all table entries is invalid. Set the upper page table entry to point to the physical address of the page and mark it as valid (valid=1) and intermediate level (such as PTE_TABLE).

[0141] At the same time, a virtual address mapping corresponding to the page fault is created in the global virtual page table. This creates a dedicated page table page for the page fault in the global virtual page table, establishing a mapping relationship between its virtual address and its physical address. This allows the system to uniformly manage the virtual address space and ensure subsequent normal access. This ensures that threads have their own independent address management space, while also enabling overall address coordination and resource allocation through the global virtual page table, ensuring normal program operation.

[0142] In some embodiments, a shared private memory page table can be linked to the global virtual page table by sharing the last-level page table. Sharing the last-level page table allows thread-private page tables and the global virtual page table to share some page table pages. This ensures that threads have their own independent address space management while enabling unified physical memory management with the help of the global page table, improving memory utilization and management efficiency.

[0143] In some embodiments, the method further includes: recording a thread identifier of the thread in a page table entry corresponding to a page table page of a global virtual page table, where the thread identifier is used to indicate that the mapping relationship contained in the page table entry is a private item of the thread.

[0144] In this embodiment, a thread identifier is introduced to distinguish the private memory of different threads. The specific method is to set the "ignored" field in the page table entry to the thread ID. This ID is private within the process and increases in order of creation starting from 0. When the "ignored" field of a page table entry is set to a specific thread ID, it means that the mapping relationship contained in the page table entry is private to the thread, that is, the corresponding memory page sharing attribute is private, and other threads cannot access it at will. This can clearly distinguish the memory usage of different threads, which is convenient for memory management and permission control.

[0145] Figure 8 A flow chart of a specific method for memory page migration in an embodiment of the present disclosure is shown.

[0146] S801: Sampling memory page access heat of the memory page to be migrated.

[0147] S802: Arrange the memory pages in descending order of access heat.

[0148] S803: Select memory pages to be migrated in sequence and classify their attributes.

[0149] In this embodiment, the attribute classification includes sharing attribute and read-write attribute. The specific method of obtaining the sharing attribute and read-write attribute of the memory page is referred to above and will not be described in detail here.

[0150] S804, determine whether the memory page to be migrated is already in the migration queue, if so, go to S807; if not, go to S805.

[0151] S805: Add the memory page to be migrated to a corresponding migration queue according to the attribute of the memory page to be migrated.

[0152] S806: Execute the memory page migration task in each migration queue in the order of the migration queue priority.

[0153] S807: Obtain the heat information of the memory pages to be migrated in the migration queue.

[0154] In this embodiment, the heat information of the memory page to be migrated is the second memory heat index mentioned above.

[0155] S808 , judging whether the heat of the memory page to be migrated increases according to the heat information of the memory page to be migrated, if so, go to S809 , if not, go to S806 .

[0156] S809: Transfer the memory page to be migrated to a migration queue with a higher priority, and then go to S806.

[0157] Based on the same inventive concept, the present disclosure also provides a memory page migration device, as described in the following embodiments. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.

[0158] Figure 9 A schematic diagram of a memory page migration device according to an embodiment of the present disclosure is shown. Figure 9 As shown, the device includes: an acquisition module 91, a determination module 92 and an adding module 93;

[0159] An acquisition module 91 is used to obtain the shared attribute information of a memory page; a determination module 92 is used to determine the migration priority of the memory page based on the shared attribute information of the memory page; and a adding module 93 is used to add the migration task of the memory page to a migration queue corresponding to the migration priority based on the migration priority of the memory page.

[0160] In some possible embodiments, the acquisition module is further configured to: acquire the read / write attributes of the memory page; and determine the migration priority of the memory page according to the shared attribute information and the read / write attributes of the memory page.

[0161] In some possible embodiments, the acquisition module is specifically used to: sample the last-level cache miss event to obtain the read and write operations of the virtual address during the last-level cache miss event; determine the read and write ratio of the memory page based on the read and write operations of the virtual address during the last-level cache miss event; determine the read and write attributes of the memory page based on the read and write ratio of the memory page; or periodically scan the global virtual page table to obtain the access frequency of the access bit and the dirty bit in the global virtual page table, and determine the read and write ratio of the memory page based on the access frequency of the access bit and the dirty bit in the global virtual page table; determine the read and write attributes of the memory page based on the read and write ratio of the memory page; wherein, when the read ratio is greater than or equal to the write ratio, the read and write attributes of the memory page are determined to be more read and less write; when the read ratio is less than the write ratio, the read and write attributes of the memory page are determined to be more write and less read.

[0162] In some possible embodiments, the sharing attributes include shared and private, and the determination module is specifically used to: when the sharing attribute information of the memory page is shared and the read-write attribute is more writes than reads, determine the migration priority of the memory page to be the first priority; when the sharing attribute information of the memory page is private and the read-write attribute is more writes than reads, determine the migration priority of the memory page to be the second priority, and the second priority is greater than the first priority; when the sharing attribute information of the memory page is shared and the read-write attribute is more reads than writes, determine the migration priority of the memory page to be the third priority, and the third priority is greater than the second priority; when the sharing attribute information of the memory page is private and the read-write attribute is more reads than writes, determine the migration priority of the memory page to be the fourth priority, and the fourth priority is greater than the third priority.

[0163] In some possible embodiments, the acquisition module is also used to: obtain the first heat index of the memory page; the adding module is specifically used to: based on the first heat index of the memory page and the migration priority of the memory page, add the migration task of the memory page to the migration queue corresponding to the migration priority.

[0164] In some possible embodiments, the acquisition module is also used to: obtain a second heat index of the memory page; when the second heat index of the memory page is greater than the first heat index, update the migration priority of the memory page; and transfer the migration task of the memory page to the migration queue corresponding to the updated migration priority according to the updated migration priority of the memory page.

[0165] In some possible embodiments, before obtaining the shared attribute information of the memory page, a creation module is further included, which is used to: in response to creating a thread in the process, establish a first memory page table for the thread.

[0166] In some possible embodiments, the first memory page table includes a private memory page table and a shared memory page table, and the shared memory page table is a last-level page table shared with other threads in the process.

[0167] In some possible embodiments, the creation module is also used to: establish a global virtual page table shared by multiple threads for the process, wherein the global virtual page table contains a private memory page table entry placeholder and a shared memory page table entry in the process, and the private memory page table entry placeholder is used to indicate a private relationship between the private memory page table entry and the corresponding thread.

[0168] In some possible embodiments, the private memory page table is pointed to by a page global directory pointer in a process descriptor.

[0169] In some possible embodiments, the access module is used to trigger a cache miss in response to the CPU where the thread is located failing to find a mapping relationship between the virtual address and the physical address of the unallocated memory page when a thread in the process accesses an unallocated memory page in a private memory page table; in response to triggering the cache miss, perform a page table traversal in the private memory page table; and trigger a page fault interrupt when a mapping relationship between the virtual address and the physical address of the unallocated memory page does not exist in the private memory page table.

[0170] In some possible embodiments, the access module is further used to: in response to triggering a page fault interrupt, traverse the global virtual page table to search for the virtual address corresponding to the page fault interrupt; when the virtual address corresponding to the page fault interrupt exists in the global virtual page table, obtain the memory page corresponding to the page fault interrupt; allocate a page table page for the memory page corresponding to the page fault interrupt in the private memory page table corresponding to the thread, and share the last-level page table of the private memory page table with the global virtual page table; when the virtual address corresponding to the page fault interrupt does not exist in the global virtual page table, allocate a page table page for the virtual address corresponding to the page fault interrupt in the private memory page table corresponding to the thread, and create a page table page for the virtual address corresponding to the page fault interrupt in the global virtual page table.

[0171] In some possible embodiments, the access module is further configured to mark, in the global virtual page table, the memory page corresponding to the page fault interrupt as a multi-thread shared item.

[0172] In some possible embodiments, the access module is further used to: record the thread identifier of the thread in the page table entry corresponding to the page table page in the global virtual page table, and the thread identifier is used to indicate that the mapping relationship contained in the page table entry is a private item of the thread.

[0173] It should be noted that the examples and application scenarios implemented by the modules in the above-mentioned apparatus embodiment are the same as those implemented by the corresponding steps in the method embodiment, but are not limited to the contents disclosed in the above-mentioned method embodiment. It should be noted that the above-mentioned modules, as part of the apparatus, can be executed in a computer system, such as a set of computer-executable instructions.

[0174] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."

[0175] Based on the same inventive concept, an embodiment of the present disclosure further provides an electronic device, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any of the above-mentioned memory page migration methods by executing the executable instructions. Since the principles for solving the problem in this electronic device embodiment are similar to those in the above-mentioned method embodiment, the implementation of this electronic device embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated here.

[0176] Refer to the following Figure 10 1000 according to this embodiment of the present disclosure will be described. Figure 10 The electronic device 1000 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0177] like Figure 10 As shown, electronic device 1000 is implemented as a general-purpose computing device. Components of electronic device 1000 may include, but are not limited to, the aforementioned at least one processing unit 1010, the aforementioned at least one storage unit 1020, and a bus 1030 connecting various system components (including storage unit 1020 and processing unit 1010).

[0178] The storage unit stores program code, which can be executed by the processing unit 1010, so that the processing unit 1010 performs the steps described in the "Exemplary Method" section of this specification according to various exemplary embodiments of the present disclosure. For example, the processing unit 1010 can perform the following steps of the above-mentioned method embodiment: obtaining shared attribute information of a memory page; determining the migration priority of the memory page based on the shared attribute information of the memory page; and adding the migration task of the memory page to the migration queue corresponding to the migration priority based on the migration priority of the memory page.

[0179] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 10201 and / or a cache memory unit 10202 , and may further include a read-only memory unit (ROM) 10203 .

[0180] The storage unit 1020 may also include a program / utility 10204 having a set (at least one) of program modules 10205, such program modules 10205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0181] Bus 1030 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0182] The electronic device 1000 may also communicate with one or more external devices 1040 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 1000, and / or any device that enables the electronic device 1000 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication may occur via an input / output (I / O) interface 1050. Furthermore, the electronic device 1000 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 1060. As shown, the network adapter 1060 communicates with other modules of the electronic device 1000 via the bus 1030. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0183] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0184] Based on the same inventive concept, embodiments of the present disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements any of the aforementioned memory page migration methods. Because the principles underlying the problem solved by this computer-readable storage medium embodiment are similar to those of the aforementioned method embodiment, the implementation of this computer-readable storage medium embodiment can be referenced to the implementation of the aforementioned method embodiment, and any repetitions will not be repeated.

[0185] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0186] In the present disclosure, a computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0187] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0188] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0189] Based on the same inventive concept, embodiments of the present disclosure further provide a computer program product, including a computer program or instructions, which, when executed by a processor, implements the memory page migration method of any one of the above-described method embodiments. Because the principles for solving the problems in this computer program product embodiment are similar to those in the above-described method embodiments, the implementation of this computer program product embodiment can refer to the implementation of the above-described method embodiments, and any repetitions will not be repeated.

[0190] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0191] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0192] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0193] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.

Claims

1. A memory page migration method, characterized in that: The method comprises: Get the shared attribute information of the memory page; Determining the migration priority of the memory page according to the shared attribute information of the memory page; Based on the migration priority of the memory page, the migration task of the memory page is added to a migration queue corresponding to the migration priority.

2. The memory page migration method according to claim 1, characterized in that: The method further comprises: Obtaining the read and write attributes of the memory page; The determining the migration priority of the memory page according to the shared attribute information of the memory page includes: The migration priority of the memory page is determined according to the sharing attribute information and the read-write attribute of the memory page.

3. The memory page migration method according to claim 1, wherein: The obtaining of the read and write attributes of the memory page includes: Sampling last-level cache miss events to obtain virtual address read and write operations during last-level cache miss events; Determining a read-write ratio of a memory page according to the read and write operations of the virtual address during the last-level cache miss event; Determining the read and write attributes of the memory page according to the read and write ratio of the memory page; or, Periodically scan the global virtual page table to obtain access frequencies of the access bit and the dirty bit in the global virtual page table, and determine the read and write ratio of the memory page according to the access frequencies of the access bit and the dirty bit in the global virtual page table; Determining the read and write attributes of the memory page according to the read and write ratio of the memory page; When the read ratio is greater than or equal to the write ratio, the read / write attribute of the memory page is determined to be more read than write; when the read ratio is less than the write ratio, the read / write attribute of the memory page is determined to be more write than read.

4. The memory page migration method according to claim 3, characterized in that: The sharing attribute includes shared and private, and determining the migration priority of the memory page according to the sharing attribute information and the read / write attribute of the memory page includes: When the sharing attribute information of the memory page is shared and the read / write attribute is write-mostly-read-less, determining the migration priority of the memory page to be the first priority; When the shared attribute information of the memory page is private and the read / write attribute is write-mostly-read-less, determining that the migration priority of the memory page is a second priority, where the second priority is greater than the first priority; When the sharing attribute information of the memory page is shared and the read / write attribute is read-mostly-write-less, determining that the migration priority of the memory page is a third priority, the third priority being greater than the second priority; When the shared attribute information of the memory page is private and the read / write attribute is read more and write less, the migration priority of the memory page is determined to be the fourth priority, and the fourth priority is greater than the third priority.

5. The memory page migration method according to any one of claims 1 to 4, characterized in that: The method further comprises: Obtaining a first heat index of the memory page; The adding the migration task of the memory page to a migration queue corresponding to the migration priority based on the migration priority of the memory page includes: Based on the first heat index of the memory page and the migration priority of the memory page, the migration task of the memory page is added to a migration queue corresponding to the migration priority.

6. The memory page migration method according to claim 5, characterized in that: The method further comprises: Obtaining a second heat index of the memory page; When the second heat index of the memory page is greater than the first heat index, updating the migration priority of the memory page; According to the updated migration priority of the memory page, the migration task of the memory page is transferred to a migration queue corresponding to the updated migration priority.

7. The memory page migration method according to claim 1, characterized in that: Before obtaining the shared attribute information of the memory page, the method further includes: In response to creating a thread in a process, a first memory page table is established for the thread.

8. The memory page migration method according to claim 7, characterized in that: The first memory page table includes a private memory page table and a shared memory page table. The shared memory page table is a last-level page table shared with other threads in the process.

9. The memory page migration method according to claim 8, characterized in that: The method further comprises: A global virtual page table shared by multiple threads is established for a process, wherein the global virtual page table contains a private memory page table entry placeholder and a shared memory page table entry in the process, and the private memory page table entry placeholder is used to indicate a private relationship between the private memory page table entry and the corresponding thread.

10. The memory page migration method according to claim 8, characterized in that: The private memory page table is pointed to by the page global directory pointer in the process descriptor.

11. The memory page migration method according to claim 9, characterized in that: The method further comprises: When a thread in a process accesses an unallocated memory page in a private memory page table, in response to the CPU where the thread is located not finding a mapping relationship between a virtual address and a physical address of the unallocated memory page, a cache miss is triggered; In response to triggering the cache miss, performing a page table traversal in the private memory page table; When there is no mapping relationship from the virtual address to the physical address of the unallocated memory page in the private memory page table, a page fault interrupt is triggered.

12. The memory page migration method according to claim 11, characterized in that: The method further comprises: In response to triggering a page fault interrupt, searching the global virtual page table for a virtual address corresponding to the page fault interrupt; When the virtual address corresponding to the page fault interrupt exists in the global virtual page table, obtaining the memory page corresponding to the page fault interrupt; Allocate a page table page for the memory page corresponding to the page fault interrupt in the private memory page table corresponding to the thread, and share the last level page table of the private memory page table with the global virtual page table; When the virtual address corresponding to the page fault interrupt does not exist in the global virtual page table, a page table page is allocated for the virtual address corresponding to the page fault interrupt in the private memory page table corresponding to the thread, and a page table page for the virtual address corresponding to the page fault interrupt is created in the global virtual page table.

13. The memory page migration method according to claim 12, characterized in that: The method further comprises: The memory page corresponding to the page fault interrupt is marked in the global virtual page table as a multi-thread shared item.

14. The memory page migration method according to claim 12, wherein: The method further comprises: In the page table entry corresponding to the page table page in the global virtual page table, the thread identifier of the thread is recorded, and the thread identifier is used to indicate that the mapping relationship included in the page table entry is a private item of the thread.

15. A memory page migration device, characterized in that: The device comprises: The acquisition module is used to obtain the shared attribute information of the memory page; a determination module, configured to determine the migration priority of the memory page according to the shared attribute information of the memory page; The adding module is used to add the migration task of the memory page to a migration queue corresponding to the migration priority based on the migration priority of the memory page.

16. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the memory page migration method according to any one of claims 1 to 14 by executing the executable instructions.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the memory page migration method according to any one of claims 1 to 14 is implemented.

18. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, it implements the memory page migration method described in any one of claims 1-14.

Citation Information

Cited By

  • Data processing method and device for distributed database, equipment, storage medium and program product

    CN121542356A