A Fork memory support method based on heterogeneous processors
By defining a new management data structure in Linux and improving the implementation of Fork, the problem of memory replacement during Fork in heterogeneous CPU architecture is solved, and more efficient memory management and performance improvement is achieved.
Patent Information
- Application Number
- CN202110381659.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-09
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2041-04-09
AI Technical Summary
In heterogeneous CPU architecture, the problem of memory being replaced during Fork system calls has caused some coprocessors to fail to fix physical pages during the calculation process, affecting performance.
Define the new management data structures Struct child_pte and struct Fork_page_info, add buddy_page pointer to the struct page structure, improve the implementation of Fork and page missing processing flow, and avoid page replacement.
It effectively solves the copy_on_write problem in Fork, reduces page copy time and memory consumption, and is suitable for heterogeneous CPU architectures.
Smart Images

Figure CN114218125B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a Fork memory support method based on heterogeneous processors, and belongs to the technical field of high-performance computing. Background Art
[0002] With the development of high-performance computing technology, it has gradually become the mainstream to improve the processing performance of the CPU by using heterogeneous hardware acceleration methods. However, various coprocessors are adopted, and Linux, currently the mainstream operating system for high-performance computing, has to continuously adapt to new heterogeneous architectures. However, due to the relatively early birth time of Linux, many of its design concepts cannot well adapt to new heterogeneous architectures. For this reason, many high-performance computing systems use customized Linux kernels, and only in this way can the performance of heterogeneity be maximized. Fork, as a very important system call in the Linux system, is also widely used in large-scale application projects. Therefore, in order to minimize the porting difficulty of projects on heterogeneous CPUs, it is also necessary to support Fork for some heterogeneous CPUs.
[0003] The current architectures of heterogeneous CPUs mainly include coprocessors such as GPUs, FGPAs, and many-core processors. The main processor of the chip needs to uniformly and coordinately manage the computing resources of its coprocessors. Memory management, as an important computing resource, requires key support. In addition, the memory used by some coprocessors needs to be fixed during the calculation process and cannot be replaced midway. In this case, there will be great problems in supporting the Fork system call under Linux. Because an important design concept adopted by Linux is the "lazy" design concept, especially for the management of memory resources, it will only truly allocate physical memory at the last moment when the memory is needed. Therefore, when Linux implements the Fork system call, at the time of Fork, only the page table entries of the parent process are completely copied to the child process and the page table entries of the parent / child processes are all set to read-only attributes. When the parent / child process writes to a certain page, a page fault exception will occur. When the kernel processes the page fault exception, it will apply for a new page and copy the old page, thus completing the final page replacement. From the implementation mechanism of Fork, once the main process writes, the final page is the newly allocated page, which is fatal to coprocessors that require fixed physical pages during the calculation process.
[0004] The current memory management process in the Fork implementation technology adopted by Linux is as follows:
[0005] 1) The parent process calls Fork to enter the operating system kernel;
[0006] 2) The kernel creates a new child process, allocates a new memory management structure `mm_struct` for the child process, and a new page table;
[0007] 3) The kernel traverses the `vma` of the parent process and decides whether the child process needs to copy the corresponding virtual memory area (`vma`) of the parent process according to the attributes of the `vma`. If copying is required, go to step 4); otherwise, continue step 3) to traverse the next `vma`;
[0008] 4) Create a `vma` for the child process and copy the content of the `vma` of the parent process that needs to be copied into the new `vma`;
[0009] 5) Traverse the page table entries corresponding to this `vma` in the parent process page table. If the page table entry exists, copy the page table entry to the corresponding page table of the child process, and set both of these (parent / child process) page table entries to read-only attributes (cannot be written);
[0010] 6) Return to step 3).
[0011] As can be seen from this process, after Fork is completed, the page table entries in the page table of the child process are directly copied from the parent process, and both the parent / child processes have only read-only permissions for these spaces. When the parent / child processes need to modify these spaces, due to permission issues, it will enter the operating system page fault handling. During page fault handling, the kernel will allocate a new physical page, copy the content of the page to the new page, and then update the new physical page to the page tables of the parent / child processes. Therefore, for the parent process, as long as it writes to the space, it will ultimately cause the physical page to be replaced, which is not feasible for some coprocessors. Summary of the Invention
[0012] The purpose of the present invention is to provide a Fork memory support method based on heterogeneous processors to solve the problem of memory replacement during Fork.
[0013] To achieve the above object, the technical solution adopted by the present invention is: provide a Fork memory support method based on heterogeneous processors, define 2 new management data structures `Struct child_pte` and `struct Fork_page_info`, and add a pointer `buddy_page` to the `struct page` structure, specifically as follows:
[0014] `Struct child_pte{`
[0015] `struct mm_struct* mm;`
[0016] `pmd_t* pmd;`
[0017] `pte_t* pte;`
[0018] };
[0019] struct Fork_page_info{
[0020] unsigned long vaddr;
[0021] struct child_pte cp[CHILD_NUM];
[0022] };
[0023] Struct child_pte data item description:
[0024] mm: mm_struct of the child process;
[0025] pmd: pmd item corresponding to this page table entry;
[0026] pte: page table entry corresponding to the page;
[0027] struct Fork_page_info data item description:
[0028] vaddr: virtual address corresponding to the page;
[0029] cp: array of management structures related to the child processes to which the page Fork belongs;
[0030] struct page {
[0031] …
[0032] Struct Fork_page_info* buddy_page;
[0033] …
[0034] };
[0035] Include the following steps:
[0036] S1. The parent process calls Fork to enter the operating system kernel;
[0037] S2. The kernel creates a new child process and allocates a new memory management structure mm_struct and a new page table for the child process;
[0038] S3. The kernel traverses the vmas of the parent process and decides whether the child process needs to copy the corresponding virtual memory area (vma) of the parent process according to the attributes of the vma. If copying is required, go to S4; otherwise, continue to traverse the next vma in S3;
[0039] S4. Create a vma for the child process and copy the content of the vma that needs to be copied from the parent process to the new vma;
[0040] S5. Traverse the page table entries corresponding to this vma in the parent process page table. If the page table entry exists, copy the page table entry to the corresponding page table of the child process, and set the read-only attribute for both the parent / child process page table entries. Then obtain the management structure struct page of the physical page corresponding to this page table entry, initialize the data structure child_pte using the mm_struct, pmd, and pte information of the child process, and copy the data structure child_pte to an empty array element in the Fork_page_info corresponding to the page;
[0041] S6. Return to S3;
[0042] When the parent process writes to the Fork page and enters the kernel page fault handling, bypass the core standard page fault handling. The specific steps are as follows:
[0043] S11. Check whether the reason for the page fault is caused by write permission. If so, enter S12; otherwise, enter the core standard page fault handling process;
[0044] S12. Obtain the management structure struct page of the physical page corresponding to the page table entry and retrieve the Fork_page_info information from it;
[0045] S13. Traverse the array struct child_pte cp related to the child process in the Fork_page_info information. If there is relevant child process information for this page, apply for a new physical page new_page, copy the content of the Fork physical page to the new physical page, and then go to S14; otherwise, go to S17;
[0046] S14. Retrieve the relevant child process information, check whether the page in the page table entry of the process is still a Fork page. If so, go to S15; otherwise, go back to S13 to continue traversing;
[0047] S15. Update the physical address corresponding to new_page to the page table entry of the child process and refresh the tlb corresponding to the child process;
[0048] S16. Modify the relevant counters of the management structure struct page corresponding to the Fork page and go back to S13 to continue traversing;
[0049] S17. Modify the permission of the page table entry corresponding to the main process to add write permission, complete the page fault handling, and return to the user.
[0050] The further improved solutions in the above technical solutions are as follows:
[0051] 1. In the above solution, the child process information described in S14 is mainly the content of struct child_pte.
[0052] 2. In the above solution, the counters described in S16 include _count and _map_count.
[0053] Due to the application of the above technical solutions, the present invention has the following advantages compared with the prior art:
[0054] The present invention provides a Fork memory support method based on heterogeneous processors. For a heterogeneous CPU architecture, it can effectively solve the copy_on_write problem during Fork. And compared with the method of copying all pages during Fork, it not only eliminates the time for page copying but also greatly saves memory. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Attached Figure 1 is a schematic diagram of the main process of Fork in the present invention;
[0056] Attached Figure 2 is a schematic diagram of the process of handling page faults in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] Embodiment: The present invention provides a Fork memory support method based on heterogeneous processors, defines two new management data structures Struct child_pte and struct Fork_page_info, and adds a pointer buddy_page to the struct page structure, specifically as follows:
[0058] Struct child_pte{
[0059] struct mm_struct* mm;
[0060] pmd_t* pmd;
[0061] pte_t* pte;
[0062] };
[0063] struct Fork_page_info{
[0064] unsigned long vaddr;
[0065] struct child_pte cp[CHILD_NUM];
[0066] };
[0067] Struct child_pte data item description:
[0068] mm: mm_struct of the child process;
[0069] pmd: pmd item corresponding to this page table entry;
[0070] pte: page table entry corresponding to the page;
[0071] Struct Fork_page_info data item description:
[0072] vaddr: virtual address corresponding to the page;
[0073] cp: array of management structures related to the child process to which the page Fork belongs;
[0074] struct page {
[0075] …
[0076] Struct Fork_page_info* buddy_page;
[0077] …
[0078] };
[0079] The following steps are included:
[0080] S1. The parent process calls Fork to enter the operating system kernel;
[0081] S2. The kernel creates a new child process and allocates a new memory management structure mm_struct and a new page table for the child process;
[0082] S3. The kernel traverses the vma of the parent process and decides whether the child process needs to copy the corresponding virtual memory area (vma) of the parent process according to the attributes of the vma. If copying is required, go to S4; otherwise, continue to traverse the next vma in S3;
[0083] S4. Create a vma for the child process and copy the content of the vma that needs to be copied from the parent process into the new vma;
[0084] S5. Traverse the page table entries corresponding to this vma in the parent process's page table. If the page table entry exists, copy the page table entry to the corresponding page table of the child process, and at the same time set both the parent / child process page table entries to read-only attributes. Obtain the management structure struct page of the physical page corresponding to this page table entry, initialize the data structure child_pte using the mm_struct, pmd, and pte information of the child process, and copy the data structure child_pte to an idle array element in the Fork_page_info corresponding to the page.
[0085] S6. Return to S3;
[0086] When the parent process writes to the Fork page and enters the kernel's page fault handling, it does not follow the core standard page fault handling but bypasses the page fault handling for the page, specifically as follows:
[0087] S11. Check whether the reason for the page fault is caused by write permission. If so, enter S12; otherwise, enter the core standard page fault handling process.
[0088] S12. Obtain the management structure struct page of the physical page corresponding to the page table entry and extract the Fork_page_info information from it.
[0089] S13. Traverse the array struct child_pte cp related to the child process in the Fork_page_info information. If there is relevant child process information for this page, apply for a new physical page new_page, copy the content of the Fork physical page to the new physical page, and then go to S14; otherwise, go to S17.
[0090] S14. Extract the relevant child process information and check whether the page in the page table entry of the process is still the Fork page. If so, go to S15; otherwise, go back to S13 to continue traversing.
[0091] S15. Update the physical address corresponding to new_page to the page table entry of this child process and refresh the tlb corresponding to this child process.
[0092] S16. Modify the relevant counters of the management structure struct page corresponding to the Fork page and go back to S13 to continue traversing.
[0093] S17. Modify the permission of the page table entry corresponding to the main process to add write permission, complete the page fault handling, and return to the user.
[0094] The child process information described in S14 is mainly the content of struct child_pte.
[0095] The counters described in S16 include _count and _map_count.
[0096] The further explanations of the above embodiments are as follows:
[0097] The present invention solves the problem that the page will be replaced when the main process writes to the Fork page by actively updating the process page table entry, which mainly includes three aspects: 1) adding a data structure; 2) modifying Fork; 3) handling page faults in the bypass. The specific technologies are as follows:
[0098] 1. Preparation of data structure:
[0099] 1) Define 2 new management data structures
[0100] struct child_pte{
[0101] struct mm_struct* mm;
[0102] pmd_t* pmd;
[0103] pte_t* pte;
[0104] };
[0105] struct Fork_page_info{
[0106] unsigned long vaddr;
[0107] struct child_pte cp[CHILD_NUM];
[0108] };
[0109] Explanation of the data items of Struct child_pte:
[0110] a) mm: mm_struct of the child process
[0111] b) pmd: pmd item corresponding to this page table entry
[0112] c) pte: page table entry corresponding to the page
[0113] Explanation of the data items of struct Fork_page_info:
[0114] a) vaddr: virtual address corresponding to the page
[0115] b) cp: array of management structures related to the child processes to which the page Fork belongs
[0116] 2) Modify the struct page structure and add a pointer buddy_page to the structure
[0117] struct page {
[0118] …
[0119] Struct Fork_page_info* buddy_page;
[0120] …
[0121] };
[0122] 2. The main process of Fork is as follows:
[0123] 1) The parent process calls Fork to enter the operating system kernel;
[0124] 2) The kernel creates a new child process and allocates a new memory management structure mm_struct and a new page table for the child process;
[0125] 3) The kernel traverses the vmas of the parent process and decides whether the child process needs to copy the corresponding virtual memory area (vma) of the parent process according to the attributes of the vma. If copying is required, go to step 4), otherwise continue to step 3) to traverse the next vma;
[0126] 4) Create a vma for the child process and copy the content of the vma that needs to be copied from the parent process to the new vma.
[0127] 5) Traverse the page table entries corresponding to this vma in the parent process page table. If the page table entry exists, copy the page table entry to the corresponding page table of the child process, and set both of these (parent / child process) page table entries to read-only attributes (cannot be written). In addition, obtain the management structure struct page of the physical page corresponding to this page table entry, initialize the data structure child_pte with the relevant information (mm_struct, pmd, and pte) of the child process, and copy this structure to an empty array element in the Fork_page_info corresponding to the page.
[0128] 6) Return to step 3)
[0129] An operation of recording the relevant information of the child process in the physical page management structure being Forked is added in step 5) for the page fault handling of the main process.
[0130] 3. The page fault handling process is as follows:
[0131] When the parent process writes to a Fork page, due to write permission issues, it will enter the kernel page fault handling. At this time, the page fault handling for this type of page is bypassed, and it does not follow the core standard page fault handling. The specific bypass process is as follows:
[0132] 1) Check whether the reason for the page fault is caused by write permission. If so, enter 2); otherwise, enter the core standard page fault process;
[0133] 2) Obtain the physical page management structure struct page corresponding to the page table entry, and extract the Fork_page_info information from it
[0134] 3) Traverse the child process-related array struct child_pte cp in the Fork_page_info information. If there is relevant child process information for this page, apply for a new physical page new_page, copy the content of the Fork physical page to the new physical page, and then go to step 4); otherwise, go to step 7)
[0135] 4) Extract the relevant child process information (mainly the content of struct child_pte), and check whether the page in the process's page table entry is still a Fork page. If so, go to step 5); otherwise, go back to step 3) to continue traversing;
[0136] 5) Update the physical address corresponding to new_page to the page table entry of this child process, and refresh the tlb corresponding to this child process;
[0137] 6) Modify the relevant counters of the physical page management structure struct page corresponding to the Fork page, such as _count, _map_count, etc., and go back to step 3) to continue traversing;
[0138] 7) Modify the permissions of the page table entry corresponding to the main process, add write permissions, complete the page fault handling, and return to the user.
[0139] By actively updating the page table entries of the child processes related to the Fork page, the process itself becomes the only process that owns the Fork page. When the main process writes, it also replaces the page table entries of the child processes, retains the physical page of the parent process, and finally obtains the write permission for this page.
[0140] When adopting the above-mentioned Fork memory support method based on heterogeneous processors, for the heterogeneous CPU architecture, it can effectively solve the copy_on_write problem during Fork, and compared with the method of copying all pages during Fork, it not only eliminates the time for page copying but also greatly saves memory.
[0141] To facilitate a better understanding of the present invention, the terms used herein will be briefly explained below:
[0142] Heterogeneous computing: Through hardware acceleration, a heterogeneous computing method using a dedicated coprocessor is adopted to improve the processing performance of the CPU.
[0143] Fork: A very important system call in Unix or Unix-like systems, used to create a process that is almost exactly the same as the original process. After a process calls Fork, a new memory space will be allocated to the new process, and most of the values of the original process will be copied to the new process.
[0144] Parent process: The process that calls Fork is the parent process.
[0145] Child process: The new process Forked out is called the child process.
[0146] Memory management: The technology of allocating and using computer memory resources during software operation.
[0147] Page table entry: To facilitate memory management, the operating system divides the memory into several pages. Each page table entry represents the specific physical address of a page, that is, a physical page.
[0148] Physical page: When the operating system manages physical memory, it also divides the memory into several pages, and each page represents a physical page.
[0149] VMA: A user process runs in a virtual address space, and these virtual address spaces are divided into several different regions according to their different attributes, such as stack space, file mapping space, code space, global variable space, etc. Linux calls such a region a virtual memory area, abbreviated as VMA.
[0150] Page table: In the operating system, a data structure that stores the correspondence between logical pages and physical pages (that is, the mapping between virtual addresses and physical addresses), and each process has its own page table.
[0151] Page fault handling: When a process is executing, the CPU accesses the virtual address of the process, and this address needs to be converted into a physical address through the hardware mmu (i.e., tlb substitution). The substitution relationship between such virtual addresses and physical addresses on the mmu needs to be created, and the mmu can also set permissions such as whether the physical page is writable. When the mmu cannot substitute out the physical address or the permissions of the substituted physical page cannot meet the permissions required for the current memory access instruction execution, a page fault exception will occur. This exception requires the intervention of the operating system, and the intervention action of the operating system is called page fault handling.
[0152] The above embodiments are only used to illustrate the technical concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. It is not intended to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit of the present invention should be covered within the protection scope of the present invention.
Claims
1. A Fork memory support method based on heterogeneous processors, characterized in that: Define two new management data structures, struct child_pte and struct Fork_page_info, and add a pointer buddy_page in the struct page structure, including the following steps: S1, the parent process calls Fork to enter the operating system kernel; S2, the kernel creates a new child process and allocates a new memory management structure mm_struct and a new page table for the child process; S3, the kernel traverses the vma of the parent process, and decides whether the child process needs to copy vma according to the attributes of vma. If necessary, go to S4, otherwise continue to S3; S4. Create a vma for the child process and copy the content of the vma that the parent process needs to copy to the new vma; S5. Traverse the page table entry corresponding to the vma in the parent process page table. If the page table entry exists, copy the page table entry to the page table corresponding to the child process. Set both the parent and child process page table entries to read-only attributes, obtain the management structure struct page of the physical page corresponding to the page table entry, use the mm_struct, pmd and pte information of the child process to initialize the data structure child_pte, and copy the data structure child_pte to an idle array element in Fork_page_info corresponding to the page. S6, return to S3; When the parent process writes the Forked page and enters the kernel's page fault processing, the details are as follows: S11, check whether the cause of the page fault is caused by the write permission, if so, enter S12, otherwise enter the core standard page fault processing flow; S12, obtain the physical page management structure struct page corresponding to the page table entry, and extract the Fork_page_info information therein; S13, traverse the child process related array struct child_pte cp in the Fork_page_info information. If the page contains related child process information, apply for a new physical page new_page, and copy the physical page content of the Fork to the new physical page, then go to S14, otherwise go to S17; S14, take out the relevant child process information, check whether the page in the page table entry of the process is still the Fork page, if so, go to S15, otherwise go to S13 to continue traversal; S15, update the physical address corresponding to new_page to the page table entry of the child process, and refresh the tlb corresponding to the child process; S16, modify the physical page management structure struct page related counters corresponding to the Fork page, and go to S13 to continue traversal; S17. Modify the permissions of the page table entry corresponding to the main process, add writable permissions, complete the page fault processing, and return to the user.
2. The Fork memory support method based on heterogeneous processors according to claim 1, characterized in that: The child process information in S14 is mainly the content of struct child_pte.
3. The Fork memory support method based on heterogeneous processors according to claim 1, characterized in that: The counters in S16 include _count and _map_count.
Citation Information
Patent Citations
Tagging images with emotional state information
CN105830061A
A method and system for in-process data isolation protection based on a user-level page table
CN109002706A