Method, system and equipment for creating NUMA (Non Uniform Memory Access)-level memory copy
By creating a memory copy of the dynamic library code segment for each NUMA node in the server of the NUMA architecture, the problem of high latency of cross-node memory access caused by read-only file mapping of code segments in multiple NUMA node systems is solved, which significantly improves the performance of the application and the scalability of the system.
Patent Information
- Application Number
- CN202510240095.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-23
AI Technical Summary
In multi-NUMA node systems, read-only file mapping of code segments results in high latency of memory access across NUMA nodes, reducing application performance.
In a server in NUMA architecture, when creating a process, it detects whether the dynamic library code snippet exists in local physical memory. If it does not exist, physical memory is allocated on the local NUMA node and a NUMA-level memory copy is created to ensure that the process accesses local memory.
By creating independent code segment copies on each NUMA node, the memory access delay across NUMA nodes is significantly reduced, the execution efficiency and overall performance of the program are improved, and the scalability and memory resource utilization of the system are enhanced.
Smart Images

Figure CN120029787A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of memory copies, and in particular to a method, system and device for creating a NUMA-level memory copy. Background Art
[0002] In modern computer systems, the operating system is responsible for managing the execution process of applications, including program loading, memory allocation, and instruction execution. When the program is run, the operating system and loader will load the application and its dynamic dependency libraries into the memory, and then the CPU will read the instructions from the memory and execute them. Taking the Linux operating system as an example, it usually divides the program into executable programs and dynamic libraries. The operating system kernel will load the executable program and the dynamic loader into the memory, and then hand over control to the dynamic loader. The dynamic loader parses the dynamic library dependencies of the executable program, loads the dependent libraries into the memory, and finally hands over control to the main program of the application. This process is usually implemented by mapping the application on the disk to the memory address space of the process. When accessed for the first time, the page fault handling function will allocate physical memory.
[0003] Non-Uniform Memory Access (NUMA) is a computer memory design for multiprocessing, characterized by memory access time that depends on the location of the memory relative to the processor. Under the NUMA architecture, a processor can access its own local memory faster than non-local memory (the local memory of another processor or the memory shared between processors). The I / O bus is also associated with NUMA. Assume two NUMA nodes, each NUMA node has local memory and CPU. NUMA0 processes can access NUMA1 memory, but in the absence of mandatory configuration of memory allocation policies, the operating system kernel usually allocates local memory to application processes to optimize memory access performance.
[0004] Under the NUMA architecture, the operating system loads executable programs and the dynamic loader loads dynamic libraries as follows: When an application 1 is executed on NUMA0, the kernel and dynamic loader will load the data segment and code segment into memory, and the operating system will prioritize the allocation of NUMA0 local memory. The mapping mode of the data segment is read-write, while the mapping mode of the code segment is read-only. Similarly, when an application 2 is executed on NUMA1, the kernel and dynamic loader will also load the data segment and code segment into memory. Since the mapping mode of the data segment is read-write, memory will be allocated on the local memory of NUMA1. The mapping mode of the code segment is read-only, and the operating system will perform memory optimization for the file mapped in this read-only mode. When a read page fault interrupt occurs, the operating system will point the virtual address of process 2 to the physical address of the code segment of NUMA0. In order to optimize the mapping of read-only files, the operating system kernel only has one copy in physical memory.
[0005] However, the existing technical solutions have some disadvantages in the case of multiple NUMA. Executable files and code segments such as dynamic libraries have read-only file mapping memory accesses across NUMA memory, which will lead to higher memory access latency, thereby reducing application performance. Moreover, this performance degradation is positively correlated with the cross-NUMA memory access latency.
[0006] Therefore, how to solve the high latency problem of memory access across NUMA nodes in existing solutions and improve program performance has become a technical problem that needs to be solved urgently. Summary of the invention
[0007] In view of this, in order to overcome the deficiencies of the prior art, the present application aims to provide a method, system and device for creating a NUMA-level memory copy.
[0008] According to a first aspect of the present application, a method for creating a NUMA-level memory copy is provided, the method comprising the following steps: In a NUMA architecture server, when a process is created on a NUMA node, it is detected whether the dynamic library code segment of the process exists in the local physical memory of the NUMA node; If the dynamic library code segment of the process does not exist in the local physical memory of the NUMA node, allocate physical memory for the dynamic library code segment of the process on the NUMA node, and create a NUMA-level memory copy of the dynamic library code segment; If the dynamic library code segment of the process exists in the local physical memory of the NUMA node, a mapping relationship between the physical address and the virtual address is directly created for the process, and the reference count of the dynamic library code segment is increased.
[0009] Optionally, in the method for creating a NUMA-level memory copy of the present application, the NUMA architecture includes multiple NUMA nodes, each NUMA node includes local memory and a CPU, and the local memory of each NUMA node is used to store a code segment copy of the process on the node.
[0010] Optionally, in the method for creating a NUMA-level memory copy of the present application, the mapping mode of the dynamic library code segment is read-only, and under the NUMA architecture, the operating system kernel will optimize the read-only file mapping so that there is only one copy of the code segment in the physical memory.
[0011] Optionally, in the method for creating a NUMA-level memory copy of the present application, a kernel switch is introduced to control the process of creating a NUMA-level memory copy. When the kernel switch is turned on, the process of creating a NUMA-level memory copy is executed, and when the kernel switch is turned off, the copy function is canceled.
[0012] Optionally, in the NUMA-level memory copy creation method of the present application, the steps of detecting whether the dynamic library code segment exists in the local physical memory, allocating physical memory for the dynamic library code segment, creating a physical address and virtual address mapping relationship, and increasing the reference count are all completed in the page fault processing stage during the process creation process.
[0013] Optionally, in the method for creating a NUMA-level memory copy of the present application, in the step of allocating physical memory for the dynamic library code segment, the operating system will preferentially allocate the local memory of the NUMA node to optimize memory access performance.
[0014] Optionally, in the method for creating a NUMA-level memory copy of the present application, in the step of creating a mapping relationship between a physical address and a virtual address, the operating system creates an independent mapping relationship for a process on each NUMA node to ensure that the process accesses a local memory copy.
[0015] Optionally, in the method for creating a NUMA-level memory copy of the present application, in the step of increasing the reference count, the operating system tracks the usage of the code segment copy on each NUMA node, and releases the physical memory occupied by the copy when the copy is no longer used by any process.
[0016] According to a second aspect of the present application, a system for creating a NUMA-level memory replica is provided, the system comprising a creation server, the creation server comprising: A detection module is used to detect whether a dynamic library code segment of a process exists in a local physical memory of the NUMA node when a process is created on a NUMA node in a server of a NUMA architecture; A memory copy creation module, used for allocating physical memory for the dynamic library code segment of the process on the NUMA node and creating a NUMA-level memory copy of the dynamic library code segment if the dynamic library code segment of the process does not exist in the local physical memory of the NUMA node; The mapping relationship and reference count creation module is used to directly create a mapping relationship between the physical address and the virtual address for the process if the dynamic library code segment of the process exists in the local physical memory of the NUMA node, and increase the reference count of the dynamic library code segment.
[0017] According to a third aspect of the present application, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect of the present application when executing the program.
[0018] The method, system and device for creating a NUMA-level memory copy according to the present application have the following beneficial technical effects: 1. Significantly reduce the memory access latency across NUMA nodes. In the prior art, the read-only file mapping of the code segment has only one copy in the physical memory, which may cause the process in the multi-NUMA node system to access the code segment across nodes, thereby increasing the memory access latency. This application creates an independent copy of the code segment on each NUMA node to ensure that the process accesses the code segment in the local memory, avoiding cross-node access, significantly reducing the memory access latency, and improving the execution efficiency of the program.
[0019] 2. Improve program performance. In a multi-NUMA node system, local memory access speed is much higher than cross-node access. This application optimizes the memory mapping strategy of code segments so that processes on each NUMA node can quickly access code segments in local memory, reducing CPU waiting time caused by memory access latency, thereby significantly improving the overall performance of the application, especially in multi-threaded and multi-process computing-intensive application scenarios.
[0020] 3. Enhance the scalability of the system. As hardware capabilities continue to increase, memory costs gradually decrease. This application makes full use of modern hardware resources and avoids performance degradation caused by memory access bottlenecks by creating a copy of the code segment for each NUMA node. This enables the system to better support large-scale parallel computing and distributed processing tasks, and improves the overall scalability of the system.
[0021] 4. Reduce memory fragmentation. In the prior art, a single mapping of a code segment may lead to fragmentation problems during memory allocation and release. This application reduces the complexity of memory allocation and the risk of memory fragmentation by independently managing code segment copies on each NUMA node, thereby improving the utilization of memory resources.
[0022] 5. Compatible with existing operating systems and applications. The implementation mechanism of this application is compatible with the memory management mechanism of existing operating systems, and there is no need to make large-scale modifications to the existing operating system kernel or applications. By introducing a kernel switch, users can enable or disable the NUMA-level memory copy function as needed, so that this application can be seamlessly integrated into the existing computing environment, reducing the difficulty and cost of technical implementation.
[0023] 6. Improve the flexibility and configuration capabilities of the system. This application supports system configuration functions. Users can flexibly choose whether to enable the NUMA-level memory copy function according to the needs of actual application scenarios. This flexibility enables the system to better adapt to different workloads and hardware configurations, further improving the overall performance and resource utilization efficiency of the system.
[0024] 7. Applicable to a wide range of read-only file mapping scenarios. This application is not only applicable to read-only file mapping of executable program code segments, but also to other read-only file mapping scenarios. By creating NUMA-level memory copies for these scenarios, memory access latency can be effectively reduced, program performance can be improved, and it has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0026] Figure 1 An example diagram of the architecture of a system for creating a NUMA-level memory copy according to an embodiment of the present application; Figure 2 This is an example diagram of the architecture of a creation server of a NUMA-level memory copy creation system according to an embodiment of the present application; Figure 3 A flowchart of a method for creating a NUMA-level memory copy according to an embodiment of the present application; Figure 4 This is an application example diagram of a method for creating a NUMA-level memory copy according to an embodiment of the present application; Figure 5 A schematic diagram of the structure of the device provided in this application. DETAILED DESCRIPTION
[0027] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0028] It should be noted that the following embodiments and features in the embodiments may be combined with each other in the absence of conflict; and, based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making any creative work are within the scope of protection of the present disclosure.
[0029] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.
[0030] Figure 1 FIG. 1 is an example diagram of the architecture of a system for creating a NUMA-level memory copy according to an embodiment of the present application, such as Figure 1 As shown, the system may include a creation server 101, a communication network 102 and / or one or more creation clients 103. Figure 1 The example in FIG. 1 is a multiple creation client 103 .
[0031] The creation server 101 can be any appropriate server for storing information, data, programs and / or any other suitable type of content. In some embodiments, the creation server 101 can perform appropriate functions. For example, in some embodiments, the creation server 101 can be used to: in a server with a NUMA architecture, when a process is created on a NUMA node, detect whether the dynamic library code segment of the process exists in the local physical memory of the NUMA node; if the dynamic library code segment of the process does not exist in the local physical memory of the NUMA node, allocate physical memory for the dynamic library code segment of the process on the NUMA node, and create a NUMA-level memory copy of the dynamic library code segment; if the dynamic library code segment of the process exists in the local physical memory of the NUMA node, directly create a mapping relationship between the physical address and the virtual address for the process, and increase the reference count of the dynamic library code segment.
[0032] Figure 2 The following is an example diagram of the architecture of a creation server of a NUMA-level memory copy creation system according to an embodiment of the present application, such as Figure 2 As shown, the creation server of this embodiment includes: A detection module is used to detect whether a dynamic library code segment of a process exists in a local physical memory of the NUMA node when a process is created on a NUMA node in a server of a NUMA architecture; A memory copy creation module, used for allocating physical memory for the dynamic library code segment of the process on the NUMA node and creating a NUMA-level memory copy of the dynamic library code segment if the dynamic library code segment of the process does not exist in the local physical memory of the NUMA node; The mapping relationship and reference count creation module is used to directly create a mapping relationship between the physical address and the virtual address for the process if the dynamic library code segment of the process exists in the local physical memory of the NUMA node, and increase the reference count of the dynamic library code segment.
[0033] As another example, in some embodiments, the creation server 101 may send a method for creating a NUMA-level memory copy to the creation client 103 for user use based on a request from the creation client 103 .
[0034] As an optional example, in some embodiments, the creation client 103 is used to provide a visual creation interface, which is used to receive a user's selection input operation to create a NUMA-level memory copy, and, in response to the selection input operation, obtain a creation interface corresponding to the option selected by the selection input operation from the creation server 101 and display the creation interface, wherein the creation interface at least displays information on creating a NUMA-level memory copy and operation options for the information on creating a NUMA-level memory copy.
[0035] In some embodiments, the communication network 102 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 102 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The creation client 103 can be connected to the communication network 102 via one or more communication links (e.g., communication link 104), and the communication network 102 can be linked to the creation service end 101 via one or more communication links (e.g., communication link 105). The communication link can be any communication link suitable for transmitting data between the creation client 103 and the creation service end 101, such as a network link, a dial-up link, a wireless link, a hard-wired link, any other suitable communication link, or any suitable combination of such links.
[0036] The creation client 103 may include any one or more clients that present an interface related to creating a NUMA-level memory copy in an appropriate form for use and operation by a user. In some embodiments, the creation client 103 may include any suitable type of device. For example, in some embodiments, the creation client 103 may include a mobile device, a tablet computer, a laptop computer, a desktop computer, and / or any other suitable type of client device.
[0037] Although the creation server 101 is illustrated as one device, in some embodiments, any suitable number of devices may be used to perform the functions performed by the creation server 101. For example, in some embodiments, multiple devices may be used to implement the functions performed by the creation server 101. Alternatively, the functions of the creation server 101 may be implemented using a cloud service.
[0038] Based on the above system, an embodiment of the present application provides a method for creating a NUMA-level memory copy, which is illustrated by the following embodiment.
[0039] Figure 3 The following is a flowchart of a method for creating a NUMA-level memory copy according to an embodiment of the present application. The method for creating a NUMA-level memory copy in this embodiment can be executed on a creation server, and the method for creating a NUMA-level memory copy includes the following steps: Step S201: In a server with a NUMA architecture, when a process is created on a NUMA node, it is detected whether a dynamic library code segment of the process exists in the local physical memory of the NUMA node.
[0040] As an optional example, in this embodiment, the NUMA architecture includes multiple NUMA nodes, each NUMA node includes local memory and CPU, and the local memory of each NUMA node is used to store a copy of the code segment of the process on the node. The mapping mode of the dynamic library code segment is read-only, and under the NUMA architecture, the operating system kernel will optimize the read-only file mapping so that there is only one copy of the code segment in the physical memory.
[0041] Step S202: If the dynamic library code segment of the process does not exist in the local physical memory of the NUMA node, physical memory is allocated for the dynamic library code segment of the process on the NUMA node, and a NUMA-level memory copy of the dynamic library code segment is created.
[0042] As an optional example, in this embodiment, a kernel switch is introduced to control the process of creating a NUMA-level memory copy. When the kernel switch is turned on, the process of creating a NUMA-level memory copy is executed, and when the kernel switch is turned off, the copy function is canceled.
[0043] In the step of allocating physical memory for the dynamic library code segment, the operating system will give priority to allocating the local memory of the NUMA node to optimize memory access performance.
[0044] Step S203: If the dynamic library code segment of the process exists in the local physical memory of the NUMA node, a mapping relationship between the physical address and the virtual address is directly created for the process, and a reference count of the dynamic library code segment is increased.
[0045] As an optional example, in this embodiment, in the step of creating a mapping relationship between physical addresses and virtual addresses, the operating system will create an independent mapping relationship for each process on the NUMA node to ensure that the process accesses the local memory copy. In the step of increasing the reference count, the operating system will track the usage of the code segment copy on each NUMA node, and when the copy is no longer used by any process, the physical memory occupied by the copy is released.
[0046] It should be noted that, in this embodiment, the steps of detecting whether the dynamic library code segment exists in the local physical memory, allocating physical memory for the dynamic library code segment, creating a mapping relationship between the physical address and the virtual address, and increasing the reference count are all completed in the page fault processing stage during the process creation process.
[0047] The following further describes in detail the method for creating a NUMA-level memory copy of this embodiment in a specific scenario.
[0048] Figure 4 FIG. 1 is an application example diagram of a method for creating a NUMA-level memory copy according to an embodiment of the present application, such as Figure 4 As shown in the figure, in this scenario, the system NUMA-level memory copy is to create a code segment copy for each NUMA, and each process in the NUMA only accesses its own local memory copy. Each process has its own data segment physical memory. Specifically, follow the steps below to create a NUMA-level memory copy: 1. Create process 1 on NUMA0. During the creation process, the dynamic library code segment does not exist in the local physical memory. After the page fault is processed, physical memory is allocated for it on NUMA0. 2. Create process 2 on NUMA0. After the process is created, it will be found in the page fault processing that the read-only file mapping already exists in the local physical memory of NUMA0. The mapping relationship between the physical address and the virtual machine address is directly created and the reference count is increased. 3. Create process 3 on NUMA1. During the creation process, the dynamic library code segment does not exist in the local physical memory. After the page fault is processed, physical memory is allocated for it on NUMA1. 4. Create process 4 on NUMA1. After the process is created, it will be found in the page fault processing that the read-only file mapping already exists in the local physical memory of NUMA1. The mapping relationship between the physical address and the virtual machine address is directly created and the reference count is increased. Through the above process, it is possible to create a read-only mapping of the file code segment on each NUMA node. The above process of creating a NUMA physical memory copy can be controlled by introducing a kernel switch. When it is turned on, the above process is executed. When it is turned off, the copy function is canceled.
[0049] The method for creating a NUMA-level memory copy according to the embodiment of the present application has the following beneficial technical effects: 1. Significantly reduce the memory access latency across NUMA nodes. In the prior art, the read-only file mapping of the code segment has only one copy in the physical memory, which may cause the process in the multi-NUMA node system to access the code segment across nodes, thereby increasing the memory access latency. This application creates an independent copy of the code segment on each NUMA node to ensure that the process accesses the code segment in the local memory, avoiding cross-node access, significantly reducing the memory access latency, and improving the execution efficiency of the program.
[0050] 2. Improve program performance. In a multi-NUMA node system, local memory access speed is much higher than cross-node access. This application optimizes the memory mapping strategy of code segments so that processes on each NUMA node can quickly access code segments in local memory, reducing CPU waiting time caused by memory access latency, thereby significantly improving the overall performance of the application, especially in multi-threaded and multi-process computing-intensive application scenarios.
[0051] 3. Enhance the scalability of the system. As hardware capabilities continue to increase, memory costs gradually decrease. This application makes full use of modern hardware resources and avoids performance degradation caused by memory access bottlenecks by creating a copy of the code segment for each NUMA node. This enables the system to better support large-scale parallel computing and distributed processing tasks, and improves the overall scalability of the system.
[0052] 4. Reduce memory fragmentation. In the prior art, a single mapping of a code segment may lead to fragmentation problems during memory allocation and release. This application reduces the complexity of memory allocation and the risk of memory fragmentation by independently managing code segment copies on each NUMA node, thereby improving the utilization of memory resources.
[0053] 5. Compatible with existing operating systems and applications. The implementation mechanism of this application is compatible with the memory management mechanism of existing operating systems, and there is no need to make large-scale modifications to the existing operating system kernel or applications. By introducing a kernel switch, users can enable or disable the NUMA-level memory copy function as needed, so that this application can be seamlessly integrated into the existing computing environment, reducing the difficulty and cost of technical implementation.
[0054] 6. Improve the flexibility and configuration capabilities of the system. This application supports system configuration functions. Users can flexibly choose whether to enable the NUMA-level memory copy function according to the needs of actual application scenarios. This flexibility enables the system to better adapt to different workloads and hardware configurations, further improving the overall performance and resource utilization efficiency of the system.
[0055] 7. Applicable to a wide range of read-only file mapping scenarios. This application is not only applicable to read-only file mapping of executable program code segments, but also to other read-only file mapping scenarios. By creating NUMA-level memory copies for these scenarios, memory access latency can be effectively reduced, program performance can be improved, and it has broad application prospects.
[0056] like Figure 5 As shown, the present application also provides a device, including a processor 310, a communication interface 320, a memory 330 for storing a processor executable computer program, and a communication bus 340. The processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 implements the above-mentioned method for creating a NUMA-level memory copy by running an executable computer program.
[0057] Among them, the computer program in the memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0058] The system embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected based on actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art may understand and implement it without creative effort.
[0059] Through the description of the above implementation modes, those skilled in the art can clearly understand that each implementation mode can be implemented by means of software plus a necessary general hardware platform, or of course by hardware. Based on such an understanding, the above technical solution can essentially or in other words be embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiment.
[0060] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A method for creating a NUMA-level memory replica, characterized in that: The method comprises the following steps: In a NUMA architecture server, when a process is created on a NUMA node, it is detected whether the dynamic library code segment of the process exists in the local physical memory of the NUMA node; If the dynamic library code segment of the process does not exist in the local physical memory of the NUMA node, allocate physical memory for the dynamic library code segment of the process on the NUMA node, and create a NUMA-level memory copy of the dynamic library code segment; If the dynamic library code segment of the process exists in the local physical memory of the NUMA node, a mapping relationship between the physical address and the virtual address is directly created for the process, and the reference count of the dynamic library code segment is increased.
2. The method for creating a NUMA-level memory copy according to claim 1, characterized in that: The NUMA architecture includes multiple NUMA nodes, each NUMA node includes local memory and a CPU, and the local memory of each NUMA node is used to store a copy of the code segment of the process on the node.
3. The method for creating a NUMA-level memory copy according to claim 1, characterized in that: The mapping mode of the dynamic library code segment is read-only, and under the NUMA architecture, the operating system kernel will optimize the read-only file mapping so that there is only one copy of the code segment in the physical memory.
4. The method for creating a NUMA-level memory copy according to claim 1, characterized in that: The process of creating a NUMA-level memory copy is controlled by introducing a kernel switch. When the kernel switch is turned on, the process of creating a NUMA-level memory copy is executed. When the kernel switch is turned off, the copy function is canceled.
5. The method for creating a NUMA-level memory copy according to claim 1, characterized in that: The steps of detecting whether the dynamic library code segment exists in the local physical memory, allocating physical memory for the dynamic library code segment, creating a mapping relationship between the physical address and the virtual address, and increasing the reference count are all completed in the page fault processing stage during the process creation process.
6. The method for creating a NUMA-level memory replica according to claim 1, characterized in that: In the step of allocating physical memory for the dynamic library code segment, the operating system will give priority to allocating the local memory of the NUMA node to optimize memory access performance.
7. The method for creating a NUMA-level memory copy according to claim 1, characterized in that: In the step of creating a mapping relationship between physical addresses and virtual addresses, the operating system creates an independent mapping relationship for each process on the NUMA node to ensure that the process accesses the local memory copy.
8. The method for creating a NUMA-level memory replica according to claim 1, characterized in that: During the step of increasing the reference count, the operating system tracks the usage of the code segment copy on each NUMA node, and releases the physical memory occupied by the copy when it is no longer used by any process.
9. A system for creating a NUMA-level memory replica, characterized in that: The system includes a creation server, and the creation server includes: A detection module is used to detect whether a dynamic library code segment of a process exists in a local physical memory of the NUMA node when a process is created on a NUMA node in a server of a NUMA architecture; A memory copy creation module, used for allocating physical memory for the dynamic library code segment of the process on the NUMA node and creating a NUMA-level memory copy of the dynamic library code segment if the dynamic library code segment of the process does not exist in the local physical memory of the NUMA node; The mapping relationship and reference count creation module is used to directly create a mapping relationship between the physical address and the virtual address for the process if the dynamic library code segment of the process exists in the local physical memory of the NUMA node, and increase the reference count of the dynamic library code segment.
10. A computer device, characterized in that: The computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the method according to any one of claims 1 to 8 when executing the program.
Citation Information
Cited By
Method for optimizing user mode program processing speed under NUMA architecture
CN120631803A
NUMA (Non Uniform Memory Access) system running state detection method, electronic equipment and medium
CN122132253A
NUMA system operational status detection methods, electronic devices and media
CN122132253B