Large page memory configuration method and device of NUMA node, electronic equipment and communication board card
By obtaining the large page pooling parameters and the number of NUMA node memory, marking independent large page memory pools and reserved memory, mapping and pooling, the problem of low efficiency of large page memory configuration of NUMA nodes is solved, and efficient management and compatibility with diversified needs is achieved.
Patent Information
- Application Number
- CN202311868716.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
The existing NUMA node large-page memory pool configuration scheme is inefficient, cannot adapt to the differences in CPU physical address space layout of different manufacturers, and cannot configure some large-page memory for the specified NUMA node, and cannot adapt to other process scenarios on the same NUMA node that need to manage large-page memory by itself.
By obtaining the pre-configured large page pooling parameters and the number of large page memory of NUMA nodes, marking the NUMA nodes that require independent large page memory pools and whether large page memory needs to be reserved, perform large page memory mapping, create independent large page memory pools, and convert large page memory blocks that are not independently pooled to form the default large page memory pool.
It improves the efficiency and management efficiency of large-page memory configuration of NUMA nodes, supports independent large-page memory pooling of designated NUMA nodes, adapts to diversified large-page memory requirements scenarios, and improves software portability and compatibility of high-performance communication software.
Smart Images

Figure CN120234263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of communication computer boards, and in particular, to a method and device for configuring large page memory of a NUMA node, an electronic device, and a communication board. Background Art
[0002] Currently, the field of network communication faces an increasing demand for high-traffic data forwarding. To meet higher performance standards, more and more devices or boards are starting to adopt a super-large-scale multi-core CPU architecture to cope with the increasing data communication volume. In a device or the communication board of a device, the application of a multi-core CPU with a NUMA (Non-Uniform Memory Access) architecture will become more and more common. For a traditional CPU with a UMA (Uniform Memory Access) architecture, some software technical details are no longer applicable and need to be improved. At the same time, new software also needs to be compatible with both UMA and NUMA processor architectures, thus posing higher requirements for software design.
[0003] Compared with UMA, a multi-core CPU with a NUMA architecture is divided into multiple NUMA nodes. Each NUMA node consists of several CPU cores, a DDR controller, and other CPU peripherals. The entire system is connected by establishing a high-speed connection bus from chip to chip between NUMA nodes. Each NUMA node has its own independent DDR controller and large page memory configuration. However, due to the existence of the chip-to-chip interconnection bus, the CPU cores of this NUMA node are allowed to access the large page memory of other NUMA nodes, which is completely different from the traditional system with one CPU and one DDR controller. Therefore, the NUMA architecture also poses certain improvement requirements for software.
[0004] Under the traditional UMA architecture, devices or boards often use a large page memory pool to minimize the number of system calls for memory allocation and achieve the highest memory access performance, which is particularly important for high-speed data processing. However, under the NUMA architecture, this method faces challenges because of the NUMA memory access characteristics. High-performance communication software needs to always maintain a state of not accessing the large page memory of other NUMA nodes for a specific NUMA node, so as to achieve the same high performance as UMA. On the other hand, other similar management service modules require a larger capacity of large page memory access capabilities. Therefore, the ability of NUMA to allow cross-node access becomes an advantage. How to combine these demands to achieve the highest efficiency has become a key challenge.
[0005] The typical solutions for configuring the large page memory pool of NUMA nodes are as Figure 1 shown Figure 1It is a schematic flowchart of the configuration method for the large page memory pool of the existing NUMA node, which specifically includes the following steps:
[0006] S1. Configure the number of large page memories of each NUMA node in the system;
[0007] S2. The memory pool management module obtains the large page memory pages and the number of large page memories of each NUMA node;
[0008] S3. Map to obtain the virtual address according to the large page memory block size, and convert the virtual address to the physical address PA1 through the kernel state call;
[0009] S4. Obtain the node number of the NUMA node corresponding to the physical address PA1 in S3 according to the address space distribution of the CPU;
[0010] S5. Loop and execute S3 and S4 until all the physical addresses PA1 of the large page memories, their corresponding virtual addresses and node numbers are obtained;
[0011] S6. Configure the result obtained in S5 to the memory pool management module;
[0012] S7. When the memory applicant applies for large page memory, determine the memory block address of the large page memory block returned from which NUMA memory pool according to the result obtained in S4.
[0013] This solution has some deficiencies: First, it needs to map all large page memory blocks one by one and infer its NUMA node number through the physical address. Such a process is very inefficient, and the physical address space layout of CPUs from different manufacturers may be different, which may lead to poor portability of some code. Second, it cannot perform partial large page memory mapping configuration only for a specified NUMA node, that is, it cannot adapt to the scenario where there are other processes that need to manage large page memory by themselves on the same NUMA node. Summary of the Invention
[0014] The objectives of the present invention include, for example, providing a method, device, electronic device, and communication board for configuring large page memory of a NUMA node to improve the configuration efficiency and large page memory management efficiency. The embodiments of the present invention can be implemented as follows:
[0015] In a first aspect, the present invention provides a method for configuring large page memory of a NUMA node, the method comprising: obtaining pre-configured large page pooling parameters and the large page memory quantities of each NUMA node; wherein the large page pooling parameters are used to mark the NUMA nodes that require independent large page memory pools and whether large page memory reservation is required for the NUMA nodes; performing large page memory mapping according to the large page pooling parameters and the large page memory quantities of each NUMA node to obtain a set of large page memory blocks corresponding to each NUMA node; for the NUMA nodes that require independent large page memory pools in the large page pooling parameters, creating independent large page memory pools according to the set of large page memory blocks, and forming a default large page memory pool with the remaining large page memory blocks that are not independently pooled.
[0016] In an alternative embodiment, performing large page memory mapping according to the large page pooling parameters and the large page memory quantities of each NUMA node to obtain a set of large page memory blocks corresponding to each NUMA node comprises: for each NUMA node, if there is no reserved large page memory corresponding to the NUMA node in the large page pooling parameters, directly performing large page memory mapping according to the large page memory quantity to obtain the set of large page memory blocks; otherwise, after configuring the specified quantity of reserved large page memory in the large page pooling parameters during the process of performing large page memory mapping according to the large page memory quantity, obtaining the set of large page memory blocks.
[0017] In an alternative embodiment, after configuring the specified reserved large page memory in the large page pooling parameters during the process of performing large page memory mapping according to the large page memory quantity, obtaining the set of large page memory blocks comprises: determining the difference between the large page memory quantity and the specified quantity of the reserved large page memory as the number of mapping times; performing large page memory mapping on the large page memory page size of the NUMA node according to the number of mapping times to obtain the set of large page memory blocks.
[0018] In an alternative embodiment, directly performing large page memory mapping according to the large page memory quantity to obtain the set of large page memory blocks comprises: taking the large page memory quantity as the number of mapping times; performing large page memory mapping on the large page memory page size of the NUMA node according to the number of mapping times to obtain the set of large page memory blocks;
[0019] In an alternative embodiment, after performing large page memory mapping on the large page memory page size of the NUMA node according to the number of mapping times to obtain the set of large page memory blocks, the method further comprises: establishing a correspondence between the node identifier of each NUMA node and the large page memory address of the large page memory blocks configured for the NUMA node, and generating a large page memory table according to all the correspondences.
[0020] In an alternative embodiment, before obtaining the pre-configured large page pooling parameters and the large page memory amounts of each NUMA node, the method further includes: configuring the large page pooling parameters and the large page memory amounts of each of the NUMA nodes.
[0021] In an alternative embodiment, configuring the large page memory amounts of each of the NUMA nodes includes: determining whether the NUMA node needs to perform large page memory configuration; if so, configuring the large page memory amount according to the large page memory page size of the NUMA node, otherwise, configuring the large page memory amount to zero.
[0022] In an alternative embodiment, before performing large page memory mapping according to the large page pooling parameters and the large page memory amounts of each of the NUMA nodes to obtain a set of large page memory blocks corresponding to each of the NUMA nodes, the method further includes: binding the thread for performing the large page memory mapping task to one of the CPU cores of the NUMA node.
[0023] In a second aspect, the present invention provides a large page memory configuration device for a NUMA node, and a large page memory management module, configured to: obtain the pre-configured large page pooling parameters and the large page memory amounts of each NUMA node; wherein, the large page pooling parameters are used to mark the NUMA nodes that require independent large page memory pools and whether the NUMA nodes need to perform large page memory reservation; perform large page memory mapping according to the large page pooling parameters and the large page memory amounts of each of the NUMA nodes to obtain a set of large page memory blocks corresponding to each of the NUMA nodes; for the NUMA nodes that require independent large page memory pools in the large page pooling parameters, create independent large page memory pools according to the set of large page memory blocks, and form a default large page memory pool with the remaining un-pooled large page memory blocks.
[0024] In a third aspect, the present invention provides an electronic device, including: a memory for storing a computer program; a processor for executing the computer program to implement the method for configuring the large page memory of a NUMA node as described in the first aspect above.
[0025] In a fourth aspect, the present invention provides a communication board, including: a memory for storing a computer program; a processor for executing the computer program to implement the method for configuring the large page memory of a NUMA node as described in the first aspect above.
[0026] The present invention provides a large page memory configuration method, device, electronic device, and communication board for NUMA nodes, and the method includes: marking NUMA nodes that need independent large page memory pools through large page pooling parameters, and marking whether large page memory reservation is required for these NUMA nodes. This plays a guiding role in the subsequent system-level large page memory configuration and reservation for the specified NUMA node. When performing large page memory configuration, first obtain the large page pooling parameters and the large page memory quantity for configuration guidance, and then complete the large page memory mapping task, without first converting the virtual addresses of the large page memory blocks into physical addresses one by one and then judging which NUMA nodes should be configured, thereby improving the configuration efficiency, and being able to reserve large page memory during the configuration process, and then creating an independent large page memory pool for the marked NUMA node so that the application process can use the large page memory block of the specified NUMA node, if a NUMA node does not specify the need for independent pooling through the large page pooling parameters, its large page memory block is returned to the default large page memory pool of the system default for management, thereby improving the large page memory management efficiency and achieving the compatibility of the requirements of different large page memories of various application processes. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0028] Figure 1 A schematic flow chart of a method for configuring a large page memory pool for an existing NUMA node;
[0029] Figure 2 A schematic diagram showing an example of an application scenario of the large page memory configuration method for a NUMA node provided in an embodiment of the present invention;
[0030] Figure 3 A schematic flow chart of a large page memory configuration method for a NUMA node provided by an embodiment of the present invention;
[0031] Figure 4 A functional module diagram of a large page memory configuration device 400 for a NUMA node provided in an embodiment of the present invention;
[0032] Figure 5 This is a structural block diagram of an electronic device / communication board provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0033] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.
[0034] Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0035] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not require further definition and explanation in subsequent drawings.
[0036] In the description of the present invention, it should be noted that if terms such as "upper", "lower", "inner", "outer", etc. indicate orientations or positional relationships based on the orientations or positional relationships shown in the drawings, or the orientations or positional relationships in which the inventive product is customarily placed during use, it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0037] In addition, terms such as "first", "second", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance.
[0038] It should be noted that the features in the embodiments of the present invention may be combined with each other without conflict.
[0039] Please refer to Figure 2 , Figure 2 which is a schematic diagram showing an example of the application scenario of the large page memory configuration method for NUMA nodes provided by the embodiments of the present invention. The system in this application scenario example includes 4 NUMA nodes, namely Node0, Node1, Node2, and Node3. At the same time, three types of application processes, P0, P1, and P2, are running in the system. Each NUMA node contains its corresponding DDR (Double Data Rate) memory.
[0040] Among them, P0, P1, and P2 represent three types of application processes with different functions. For example, P0 is a system management process that does not care about memory access performance; P1 is an application process that requires the highest memory access performance and uses a large page memory pool. Assume that P1 runs on one of the CPU cores of Node1; P2 is an application process that requires the highest memory access performance and manages large pages by itself. Assume that P2 runs on one of the CPU cores of Node2.
[0041] Among them, T1 and T2 are two threads in the application process P1. For example, both T1 and T2 are application threads that require the highest memory access performance. For example, T1 uses the large page memory in Node1, and T2 uses the large page memory in Node2. M0 is a large page memory pool management module in the process P0 that supports configuring the large page memory pool according to NUMA nodes, that is, the large page memory management module, which is used to allocate and manage large page memory for the entire system. Other processes can interact with M0 through inter-process communication to obtain the large page memory they need.
[0042] For example, P1 and P2 can send the node identifier of the NUMA node where the process is located and the memory block size requirement to M0 through inter-process communication. After receiving the application, M0 returns the physical address of the large page memory block that has been allocated to P1 and P2. P1 and P2 can directly obtain the virtual address of the process by using the map system call with the physical address of the large page memory obtained by the application, and then they can directly use it. The purpose of the large page memory configuration method provided by the embodiments of the present invention is to be able to Figure 2 In this system scenario, the large page memory mapping and fixed allocation of specific NUMA nodes can be directly performed without converting the virtual addresses of all large page memory blocks to physical addresses and then determining the NUMA nodes for allocation. As Figure 2 shown, the embodiments of the present invention can configure the large page memory for Node0, Node1, and Node2 as needed, and support Node2 to use reserved large page memory (reserved large page memory is represented by Reservation, abbreviated as RSV).
[0043] It should be noted that Figure 2 The 4 NUMA nodes (Node0 to Node3) and their required large page memory configuration and process thread setting methods in the shown scenario are only for illustrative purposes and are not intended to limit the present invention.
[0044] Figure 3 shows a schematic flowchart of the large page memory configuration method for NUMA nodes provided by the embodiments of the present invention, which may include the following steps:
[0045] S301. Obtain the pre-configured large page pooling parameters and the large page memory quantities of each NUMA node;
[0046] In the embodiments of the present invention, the large page pooling parameter is used to mark the NUMA nodes that require an independent large page memory pool and whether large page memory reservation is required for the NUMA nodes. The node identifier of the NUMA node can take a non-zero value.
[0047] S302. Perform large page memory mapping according to the large page pooling parameter and the large page memory quantity of each NUMA node to obtain a set of large page memory blocks corresponding to each NUMA node;
[0048] S303. For the NUMA nodes that require an independent large page memory pool in the large page pooling parameter, create an independent large page memory pool according to the set of large page memory blocks, and form a default large page memory pool with the remaining large page memory blocks that have not been independently pooled.
[0049] In the above steps S301 to S303, in the embodiments of the present invention, the NUMA nodes that require an independent large page memory pool are marked by the large page pooling parameter, and it is marked whether large page memory reservation is required for these NUMA nodes. This plays a guiding role in subsequent system-level large page memory configuration and reservation for specified NUMA nodes. When performing large page memory configuration, only need to first obtain the large page pooling parameter and the large page memory quantity for configuration guidance, and then complete the large page memory mapping task, without first converting the virtual addresses of large page memory blocks to physical addresses one by one and then determining which NUMA nodes should be configured, improving the configuration efficiency, and large page memory reservation can be performed during the configuration process to support other processes on the same NUMA node to directly use the allocated large page memory. Immediately afterwards, create an independent large page memory pool for the marked NUMA nodes so that the application process can use the large page memory blocks of the specified NUMA node. If a NUMA node is not specified to require independent pooling through the large page pooling parameter, then the large page memory of this node will not establish its own exclusive pool, but will be managed in the system default large page memory pool, thereby improving the large page memory management efficiency and realizing the compatibility of different large page memory requirements of various application processes.
[0050] The large page memory configuration method provided above will be introduced in detail below.
[0051] First, in step S301, it is necessary to obtain the pre-configured large page pooling parameter and the large page memory quantity. Therefore, before executing step S301, the pre-configuration task can be completed first, that is:
[0052] Configure the large page pooling parameter and the large page memory quantity of each NUMA node.
[0053] In the embodiments of the present invention, the large page pooling parameter supports specifying whether a NUMA node needs independent pooling and supports partitioning a part of the large page system configuration of the NUMA node as a "reserved" part dedicated to a certain application software. Therefore, the large page pooling parameter in the embodiments of the present invention is in the form of:
[0054] pool@node-x...hugepgrsv@node-x:q.
[0055] Among them, pool means to configure an independent large page memory pool for the NUMA node corresponding to the node identifier node-x, hugepgrsv means to reserve large page memory for this NUMA node, and q means the size of the reserved large page memory.
[0056] In the actual implementation process, the large page pooling parameter can be obtained from the Linux kernel cmdline (command line parameter).
[0057] Continuing with the combination of Figure 2 As shown, assuming that only the NUMA node 0, NUMA node 1, and NUMA node 2 in the system are configured with large page memory, and the node identifiers are respectively represented as Node0, Node1, and Node2, it can be configured in the Linux kernel startup parameter cmdline as:
[0058] pool@node0 pool@node1 pool@node2 hugepgrsv@node2:16.
[0059] Among them, pool@node0 means to configure an independent large page memory pool for Node0; pool@node1 means to configure an independent large page memory pool for Node1; pool@node2 means to configure an independent large page memory pool for Node2; hugepgrsv@node2:16 means to configure 16 reserved large page memories for Node2. Then, after the large page memory is configured, NUMA node 0 and NUMA node 1 will each have an independently managed large page memory pool, NUMA node 2 will also have an independently managed large page memory pool, and 16 large page memories will be reserved from the total number of large page memories of NUMA node 2. The NUMA node 3 in the large page pooling parameter is not configured, indicating that large page memory is not required.
[0060] Before configuring the number of huge pages for each NUMA node, the total number of NUMA nodes n of the system can be obtained from Linux / sys first. If n is 1, it indicates that there is only NUMA node 0, and it is regarded as a CPU computer system of UMA. If n is greater than 1, then according to the system requirements, configure the array N = {C0, … Cn-1} of the number of huge pages for n NUMA nodes, where the elements in the array are the number of huge pages required for each NUMA node. If a certain NUMA node x does not need to configure huge pages, then Cx = 0. Continuing to combine with Figure 2 As shown, the total number of NUMA nodes in the system is 4, and configure the array N = {C0, C1, C2, C3} of the number of huge pages for 4 NUMA nodes, where C0, C1, C2, C3 respectively correspond to the number of huge pages required for Node0, Node1, Node2. Since Node3 does not need to configure huge pages, then C3 = 0.
[0061] In an alternative embodiment, the method of configuring the number of huge pages may include the following steps:
[0062] Step a1: Determine whether the NUMA node needs to configure huge pages;
[0063] If so, execute step a2; otherwise, execute step a3;
[0064] Step a2: Configure the number of huge pages according to the huge page configuration of the NUMA node;
[0065] Step a3: Configure the number of huge pages to zero.
[0066] In the embodiment of the present invention, for the NUMA nodes that need to configure huge pages, a most suitable huge page can be selected from multiple huge pages of different sizes configured in the system according to the actual needs, and then the number of huge pages can be configured according to this huge page; while for the NUMA nodes that do not need to configure huge pages, the number of huge pages can be directly configured to zero.
[0067] Therefore, before configuring the number of huge pages, the huge page number array P = {P1, … Pn} under each NUMA node can also be obtained from Linux / sys. If a certain NUMA node x does not need to configure huge pages, then Px = 0.
[0068] After completing the configuration of the huge page pooling parameters and the huge page number array through the above configuration method, the huge page configuration of each NUMA node can be directly called subsequently, that is, execute step S302.
[0069] In step S302, since some NUMA nodes need to reserve large-page memory during the large-page memory configuration process, for step S302, the implementation manner provided by the embodiments of the present invention is as follows:
[0070] The first case: If there is no reserved large-page memory corresponding to the NUMA node in the large-page pooling parameter, directly perform large-page memory mapping according to the large-page memory quantity to obtain a set of large-page memory blocks.
[0071] The second case: If there is reserved large-page memory corresponding to the NUMA node in the large-page pooling parameter, after configuring the specified quantity of reserved large-page memory in the large-page pooling parameter during the process of performing large-page memory mapping according to the large-page memory quantity, obtain a set of large-page memory blocks.
[0072] That is to say, if a certain NUMA node needs to reserve large-page memory, then the reserved large-page memory can be configured first according to the specified quantity of reserved large-page memory in the large-page pooling parameter, and then the remaining large-page memory is mapped. If there is no need to reserve large-page memory, directly configure according to the pre-configured large-page memory quantity.
[0073] In an alternative implementation manner, for the first case, the following steps may be executed:
[0074] Step b1: Use the large-page memory quantity as the mapping times;
[0075] Step b2: Perform large-page memory mapping on the large-page memory pages of the NUMA node according to the mapping times to obtain a set of large-page memory blocks.
[0076] Combined with the above-configured large-page memory quantity array N = {C1,..., Cn}, the large-page memory mapping times m in the first case is equal to C.
[0077] In an alternative implementation manner, for the above second case, the following steps may be executed:
[0078] Step c1: Determine the difference between the large-page memory quantity and the specified quantity of reserved large-page memory as the mapping times;
[0079] Step c2: Perform large-page memory mapping on the large-page memory pages of the NUMA node according to the mapping times to obtain a set of large-page memory blocks.
[0080] Combined with the above large page memory quantity array N, the number of large page memory mappings in the second case can be expressed as m = C – q, where q is the value of the hugepgrsv parameter in the large page pooling parameter. It can be seen that if the hugepgrsv parameter exists in NUMA node x, then there must be m < Cx. In this case, the large pages of this NUMA node x are not fully pooled, but a part is "reserved", and the reserved part is allowed to be used by application processes that need to manage large page memory by themselves under this NUMA node.
[0081] Through the above method, the large page memory page size of each NUMA node is mapped m times. Due to the memory access characteristics of the NUMA architecture, the large page memory blocks (virtual addresses VA) obtained by mapping are naturally the large page memory spaces of this NUMA node. If Cx corresponding to a certain NUMA node x is zero, then the number of large page memory mappings is also 0, that is, the set of large page memory blocks is empty.
[0082] In an alternative embodiment, before performing the above large page memory mapping task, for those NUMA nodes that need to perform mapping, the threads used to perform the large page memory mapping task need to be bound to a CPU core under this NUMA node first. For example, according to the total number of NUMA nodes n, n threads T11, …, T1n used to perform the large page memory mapping task are started first. Among them, the T1x thread is sequentially bound to one of the CPU cores of the xth NUMA node to complete the large page memory mapping task.
[0083] It can be understood that binding the threads for large page memory mapping to the CPU core first and then applying for mapping the large page memory, based on the memory access policy characteristics of the NUMA architecture, and applying for large page memory according to the system configuration number, will not result in remote memory access, and the large page memory blocks obtained by mapping will definitely remain under this NUMA node. After each thread performs a large page memory mapping, a virtual address VA is obtained and added to the set. After performing m mappings, the set V of large page memory blocks of this NUMA node is obtained.
[0084] In an alternative embodiment, after obtaining the large page memory blocks configured for each NUMA node, a correspondence relationship can also be established between the node identifier of this NUMA node and the large page memory addresses of the configured large page memory blocks. All the correspondence relationships can generate a large page memory table.
[0085] For example, for the xth NUMA node, assuming that the virtual address mapped to it is VA-x, then a correspondence relationship can also be established between the node identifier x and each VA-x and the corresponding physical address. The correspondence relationships of all NUMA nodes can also form a large page memory table for the large page thread large page memory pool management module (such as the above Figure 2For the use of M0 shown, if other threads need to apply for large page memory of a specified NUMA node, they provide the node identifier x of the specified NUMA node and the size of the required large page memory. The large page memory pool management module can then return the starting addresses of the large page memory blocks of the corresponding size from the large page memory table according to the large page memory block corresponding to the node identifier x, including the virtual address and the physical address.
[0086] As can be seen from the above embodiments, the embodiments of the present invention have the following advantages compared with the prior art:
[0087] First, by defining a set of large page memory configuration parameter tags suitable for the NUMA architecture, large page memory management control is achieved for threads sensitive to high-performance large page memory access performance (pooled independently by specified nodes), threads insensitive to access performance (pooled publicly), and processes that need to manage large page memory by themselves, thus meeting the complex demand combinations of system software.
[0088] Second, the large page memory configuration method provided by the embodiments of the present invention does not need to use the exhaustive method to determine the NUMA node affiliation relationship, improves the large page memory configuration efficiency, and can flexibly support the coexistence of multiple threads and multiple processes in the software system, and can meet the requirements of the software architecture to support both NUMA and UMA at the same time. At the same time, the embodiments of the present invention can apply for large page memory according to NUMA nodes subsequently, with higher efficiency and software portability.
[0089] Based on the Figure 3 same inventive concept, the embodiments of the present invention also provide a large page memory configuration device 400 for NUMA nodes. Please refer to Figure 4 , Figure 4 which is the functional module diagram of the large page memory configuration device 400 for NUMA nodes provided by the embodiments of the present invention, including an acquisition module 410, a mapping module 420, and a configuration module 430;
[0090] The acquisition module 410 is used to acquire the pre-configured large page pooling parameters and the large page memory quantities of each NUMA node; wherein, the large page pooling parameters are used to mark the NUMA nodes that require independent large page memory pools and whether large page memory reservation is required for the NUMA nodes;
[0091] The mapping module 420 is used to perform large page memory mapping according to the large page pooling parameters and the large page memory quantities of each NUMA node to obtain a set of large page memory blocks corresponding to each NUMA node;
[0092] The configuration module 430 is configured to create an independent large page memory pool for each NUMA node that requires an independent large page memory pool among the large page pooling parameters, and form a default large page memory pool with the remaining large page memory blocks that are not independently pooled.
[0093] It can be understood that the acquisition module 410, the mapping module 420, and the configuration module 430 can cooperate to execute Figure 3 each step in to achieve the corresponding technical effects.
[0094] In an optional implementation, the mapping module 420 is specifically configured to: for each NUMA node, if there is no reserved large page memory corresponding to the NUMA node in the large page pooling parameters, directly perform large page memory mapping according to the large page memory quantity to obtain the large page memory block set; otherwise, after configuring the specified quantity of reserved large page memory in the large page memory pooling parameters during the large page memory mapping according to the large page memory quantity, obtain the large page memory block set.
[0095] In an optional implementation, the mapping module 420 is further specifically configured to: determine the difference between the large page memory quantity and the specified quantity of the reserved large page memory as the mapping times; perform large page memory mapping on the large page memory pages of the NUMA node according to the mapping times to obtain the large page memory block set.
[0096] In an optional implementation, the mapping module 420 is further specifically configured to: use the large page memory quantity as the mapping times; perform large page memory mapping on the large page memory pages of the NUMA node according to the mapping times to obtain the large page memory block set;
[0097] In an optional implementation, the large page memory configuration device 400 for the NUMA node may further include an establishment module, configured to: establish a correspondence between the node identifier of each NUMA node and the large page memory address of the large page memory block configured for the NUMA node, and generate a large page memory table according to all the correspondences.
[0098] In an optional implementation, the configuration module 430 may further be configured to configure the large page pooling parameters and the large page memory quantity of each NUMA node.
[0099] In an optional implementation, the configuration module 430 may further be specifically configured to determine whether a NUMA node needs to perform large page memory configuration; if so, configure the large page memory quantity according to the large page memory page size of the NUMA node, otherwise, configure the large page memory quantity to zero.
[0100] In an alternative embodiment, the large page memory configuration device 400 of the NUMA node may further include a binding module for binding a thread that executes a large page memory mapping task to one of the CPU cores of the NUMA node.
[0101] It should be noted that the division of modules in the above embodiments of the present application is illustrative. It is only a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, may exist independently physically, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0102] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or an electronic device, etc.) or a processor to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0103] Based on the above embodiments, the embodiments of the present application further provide an electronic device / communication board 500. Please refer to Figure 5 , Figure 5 which is a structural block diagram of the electronic device / communication board 500 provided by the embodiments of the present invention. The electronic device / communication board is used to execute the large page memory configuration method of the NUMA node provided by the embodiments of the present invention. The electronic device / communication board 500 includes: a memory 501, a processor 502, a communication interface 503, and a bus. The memory 501, the processor 502, and the communication interface 503 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these components may be electrically connected to each other through one or more communication buses or signal lines.
[0104] Optionally, the bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0105] In the embodiment of the present application, the processor 502 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiment of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor. The software module may be located in the memory 501, and the processor 502 reads the program instructions in the memory 501, and completes the steps of the above method in combination with its hardware.
[0106] In the embodiment of the present application, the memory 501 may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), etc., or a volatile memory (volatile memory), such as RAM. The memory may also be any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the embodiment of the present application may also be a circuit or any other device that can implement a storage function, for storing instructions and / or data.
[0107] The memory 501 can be used to store software programs and modules, such as the instructions / modules of the large page memory configuration device 400 of the NUMA node provided in the embodiment of the present invention, which can be stored in the memory 501 in the form of software or firmware or solidified in the operating system (OS) of the electronic device / communication board 500. The processor 502 executes various functional applications and data processing by executing the software programs and modules stored in the memory 501. The communication interface 503 can be used to communicate signaling or data with other node devices.
[0108] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0109] It can be understood that Figure 5 The structure shown is only schematic, and the electronic device / communication board 500 may further include more or fewer components than those shown in Figure 5 or have a different configuration from that shown in Figure 5 shown. Figure 5 Each component shown can be implemented by hardware, software, or a combination thereof.
[0110] Based on the above embodiments, the present application also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a computer, the computer is caused to execute the method for configuring large page memory of NUMA nodes provided in the above embodiments.
[0111] Based on the above embodiments, the embodiments of the present application also provide a computer program. When the computer program runs on an electronic device / communication board, the electronic device / communication board is caused to execute the method for configuring large page memory of NUMA nodes provided in the above embodiments.
[0112] Based on the above embodiments, the embodiments of the present application also provide a chip. The chip is used to read the computer program stored in the memory and is used to execute the method for configuring large page memory of NUMA nodes provided in the above embodiments.
[0113] The embodiments of the present application also provide a computer program product, including instructions. When it runs on an electronic device / communication board, the electronic device / communication board is caused to execute the method for configuring large page memory of the body NUMA nodes provided in the above embodiments.
[0114] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by instructions. These instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0115] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process or Figure 1 one or more processes and / or blocks Figure 1 specified in a block or blocks.
[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or Figure 1 one or more processes and / or blocks Figure 1 specified in a block or blocks.
[0117] The above is only a specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for configuring large page memory of a NUMA node, characterized in that The method includes: Obtaining pre-configured large page pooling parameters and the large page memory quantities of each NUMA node; wherein, the large page pooling parameters are used to mark the NUMA nodes that require independent large page memory pools and whether large page memory reservation is required for the NUMA nodes; Performing large page memory mapping according to the large page pooling parameters and the large page memory quantities of each NUMA node to obtain a set of large page memory blocks corresponding to each NUMA node; For the NUMA nodes that require independent large page memory pools in the large page pooling parameters, creating independent large page memory pools according to the set of large page memory blocks, and forming a default large page memory pool with the remaining un-pooled large page memory blocks.
2. The method for configuring large page memory of a NUMA node according to claim 1, characterized in that Performing large page memory mapping according to the large page pooling parameters and the large page memory quantities of each NUMA node to obtain a set of large page memory blocks corresponding to each NUMA node, including: For each NUMA node, if there is no reserved large page memory corresponding to the NUMA node in the large page pooling parameters, directly performing large page memory mapping according to the large page memory quantity to obtain the set of large page memory blocks; Otherwise, after configuring the specified quantity of reserved large page memory in the large page pooling parameters during the process of performing large page memory mapping according to the large page memory quantity, obtaining the set of large page memory blocks.
3. The large page memory configuration method for the NUMA node according to claim 2, wherein After configuring the specified quantity of reserved large page memory in the large page pooling parameters during the process of performing large page memory mapping according to the large page memory quantity, obtaining the set of large page memory blocks, including: Determining the number of mapping times as the difference between the large page memory quantity and the specified quantity of the reserved large page memory; Performing large page memory mapping on the large page memory pages of the NUMA node according to the number of mapping times to obtain the set of large page memory blocks.
4. The large page memory configuration method for a NUMA node according to claim 2, wherein Directly performing large page memory mapping according to the large page memory quantity to obtain the set of large page memory blocks, including: Taking the large page memory quantity as the number of mapping times; Performing large page memory mapping on the large page memory pages of the NUMA node according to the number of mapping times to obtain the set of large page memory blocks.
5. The method for configuring large page memory of a NUMA node according to claim 3 or 4, characterized in that After performing large page memory mapping on the large page memory page size of the NUMA node according to the number of mapping times to obtain the set of large page memory blocks, the method further includes: Establishing a correspondence between the node identifier of each NUMA node and the large page memory address of the large page memory blocks configured for the NUMA node, and generating a large page memory table according to all the correspondences.
6. The large page memory configuration method for NUMA nodes according to claim 1, wherein Before obtaining the pre-configured large page pooling parameters and the large page memory quantities of each NUMA node, the method further includes: Configuring the large page pooling parameters and the large page memory quantities of each NUMA node.
7. The method for configuring large page memory of a NUMA node according to claim 6, wherein Configuring the large page memory quantities of each NUMA node includes: Determining whether large page memory configuration is required for the NUMA node; If so, configuring the large page memory quantity according to the large page memory page size of the NUMA node, otherwise, configuring the large page memory quantity to zero.
8. The large page memory configuration method for a NUMA node according to claim 1, characterized in that Before performing large page memory mapping according to the large page pooling parameters and the large page memory amounts of the respective NUMA nodes to obtain the large page memory block sets corresponding to the respective NUMA nodes, the method further includes: Binding the thread for performing the large page memory mapping task to one of the CPU cores of the NUMA node.
9. A large page memory configuration device for a NUMA node, characterized in that, An acquisition module, a mapping module, and a configuration module; The acquisition module is configured to acquire the pre-configured large page pooling parameters and the large page memory amounts of the respective NUMA nodes; wherein the large page pooling parameters are used to mark the NUMA nodes that require independent large page memory pools and whether large page memory reservation is required for the NUMA nodes; The mapping module is configured to perform large page memory mapping according to the large page pooling parameters and the large page memory amounts of the respective NUMA nodes to obtain the large page memory block sets corresponding to the respective NUMA nodes; The configuration module is configured to, for the NUMA nodes that require independent large page memory pools in the large page pooling parameters, create independent large page memory pools according to the large page memory block sets, and form a default large page memory pool with the remaining large page memory blocks that are not independently pooled.
10. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for executing the computer program to implement the method for configuring large page memory of a NUMA node according to any one of claims 1 to 8.
11. A communication board, characterized in that, Includes: A memory for storing a computer program; A processor for executing the computer program to implement the method for configuring large page memory of a NUMA node according to any one of claims 1 to 8.