Process configuration method, multiprocessor system and computing equipment
By monitoring the httpd program creation process in a multiprocessor system and binding it to multiple NUMA nodes, and configuring NUMA nodes that specifically handle network card interrupt requests, the httpd program has solved the problem of low performance and uneven load operation in the NUMA architecture system, and more efficient system performance and network processing capabilities are achieved.
Patent Information
- Application Number
- CN202510272625.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
In multiprocessor systems using NUMA architecture, the process configuration method of httpd programs leads to low operating performance and unbalanced load.
A process configuration method is proposed to monitor the operations of httpd program creation process, and bind each httpd process to multiple NUMA nodes to ensure that the process is evenly distributed. In addition, a second NUMA node is configured to process network interrupt requests to improve system network performance.
By decentralizing the httpd process to multiple NUMA nodes, we can make full use of the parallel processing performance and memory access efficiency of multi-processor systems, improve the operating performance of httpd programs and make the load more balanced. At the same time, by specifically handling network card interrupt requests, the system's network processing efficiency is improved.
Smart Images

Figure CN120216174A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a process configuration method, a multi-processor system, and a computing device. Background Art
[0002] Apache HTTP Server (abbreviated as httpd) is an open-source and widely used web server software, known for its high configurability and flexibility. Httpd uses a multi-process or multi-threaded architecture to handle client requests, mainly implementing the processing of concurrent connections through the "process model". Common process models include Prefork, Worker, and Event, etc., and these models allocate processes or threads in different ways to handle client requests.
[0003] NUMA (Non-Uniform Memory Access) is a memory architecture design for multi-processor systems, aiming to improve the parallel processing performance and memory access efficiency of the system. In a multi-processor system adopting the numa architecture, the httpd parent process usually creates each httpd child process in its numa node where it is located, and the running performance of the httpd program under this httpd process allocation method is relatively low. Summary of the Invention
[0004] Based on the above technical status quo, this application proposes a process configuration method, a multi-processor system, and a computing device, which can improve the running performance of the httpd program in a multi-processor system adopting the numa architecture.
[0005] The first aspect of this application proposes a process configuration method, which is applied to a multi-processor system adopting the numa architecture. The multi-processor system includes multiple numa nodes, and the method includes:
[0006] When the httpd program starts, start monitoring the operation of the httpd program to create httpd processes;
[0007] Whenever it is monitored that the httpd program creates an httpd process, bind the created httpd process to a first numa node, so as to scatter and bind each httpd process created by the httpd program to each first numa node of the multi-processor system.
[0008] In some implementation manners, the method further includes:
[0009] Configure the second numa node in the multi-processor system to be used for processing network card interrupt requests.
[0010] In some implementations, when it is monitored that the httpd program creates an httpd process, binding the created httpd process to a first NUMA node includes:
[0011] Determining a plurality of first NUMA nodes from among the multiple NUMA nodes of the processor system;
[0012] When it is monitored that the httpd program creates an httpd process, selecting a first NUMA node from the plurality of first NUMA nodes and binding the created httpd process to the selected first NUMA node.
[0013] In some implementations, when it is monitored that the httpd program creates an httpd process, selecting a first NUMA node from the plurality of first NUMA nodes and binding the created httpd process to the selected first NUMA node includes:
[0014] During the process of the httpd program creating an httpd process, traversing the plurality of first NUMA nodes in a loop. When each first NUMA node is traversed, binding an httpd process created by the httpd program to that first NUMA node and then traversing the next first NUMA node after the binding is completed, until the creation of the httpd process by the httpd program ends, at which point the loop traversal of the plurality of first NUMA nodes stops.
[0015] In some implementations, the method further includes:
[0016] In the first NUMA node, creating a copy file of the memory file required for the target httpd process bound to that first NUMA node;
[0017] Mapping the memory access address of the target httpd process to the copy file and loading the copy file into the memory of that first NUMA node.
[0018] In some implementations, in the first NUMA node, creating a copy file of the memory file required for the target httpd process bound to that first NUMA node includes:
[0019] Reading the maps file of the target httpd process bound to that first NUMA node and, by parsing the maps file, determining the starting address of the memory to be accessed by the target httpd process and the memory file name stored at that memory starting address; the maps file is used to record the mapping information of the virtual memory regions related to the target httpd process;
[0020] Determine the file name of the copy file to be created based on the memory file name according to the set copy file name naming rule;
[0021] In the case where it is determined that the file name of the copy file does not exist in the first NUMA node, create a copy file according to the file name, and write the file content stored at the memory start address into the copy file.
[0022] In some implementation manners, by parsing the maps file, determine the memory start address required by the target httpd process and the memory file name stored at the memory start address, including:
[0023] Parse the maps file line by line, and search for set fields in each line of the maps file, where the set fields include a first field corresponding to the executable code segment required to be accessed by the target httpd process and / or a second field corresponding to the read-only data segment required to be accessed by the target httpd process;
[0024] In the case where the set field is found, determine the memory start address corresponding to the set field and the memory file name stored at the memory start address.
[0025] A second aspect of the present application proposes a multi-processor system. The multi-processor system adopts a NUMA architecture. In the multi-processor system, there are multiple NUMA nodes. In the multi-processor system, whenever the httpd program creates an httpd process, the created httpd process is bound to a first NUMA node, so that each httpd process created by the httpd program is scattered and bound to each first NUMA node of the multi-processor system.
[0026] In some implementation manners, the second NUMA node in the multi-processor system is configured to process network card interrupt requests.
[0027] In some implementation manners, in the memory of the first NUMA node, there is also a copy file of the memory file required by the target httpd process bound to the first NUMA node, and the memory access address of the target httpd process is mapped to the copy file.
[0028] A third aspect of the present application proposes a computing device, including the above multi-processor system.
[0029] The process configuration method proposed in this application is applied to a multi-processor system adopting the NUMA architecture. Based on this multi-processor system, when the httpd program starts, each httpd process created by the httpd program is scattered and bound to each first NUMA node of the multi-processor system. The above method can disperse the httpd processes to multiple NUMA nodes, thereby making full use of the characteristics of strong parallel processing performance and high memory access efficiency of the multi-processor system with the NUMA architecture, improving the running performance of the httpd program. At the same time, this method can also make the load of the httpd program on the multi-processor system with the NUMA architecture more balanced, which is beneficial to improving the performance of the processor system. Description of the Drawings
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0031] Figure 1 It is a schematic structural diagram of a multi-processor system adopting the NUMA architecture provided by an embodiment of the present application.
[0032] Figure 2 It is a schematic structural diagram of a multi-processor system provided by an embodiment of the present application.
[0033] Figure 3 It is a schematic flow diagram of the process configuration method provided by an embodiment of the present application.
[0034] Figure 4 It is a schematic structural diagram of another multi-processor system provided by an embodiment of the present application.
[0035] Figure 5 It is a schematic structural diagram of yet another multi-processor system provided by an embodiment of the present application. Detailed Embodiments
[0036] NUMA (Non-Uniform Memory Access) is a computer architecture that is derived with the development of multi-processor or multi-processor core technology and is used to optimize the memory access performance of the processor.
[0037] In the NUMA architecture, processors or processor cores are divided into multiple NUMA nodes. Each node has its own independent memory space and bus system, and the various NUMA nodes are also interconnected through the bus.
[0038] In one embodiment, a NUMA architecture includes a processor which comprises a plurality of processor cores. The plurality of processor cores are divided into different groups, each group including at least one processor core, and each group serves as a NUMA node.
[0039] In another embodiment, as Figure 1 shown, a NUMA architecture includes a plurality of processors, each processor comprising a plurality of processor cores. The processor cores of each processor are divided into different NUMA nodes.
[0040] Regardless of the adopted method, for each NUMA node, an independent memory controller is also configured. The processor cores in the same NUMA node can share this memory controller. Meanwhile, corresponding physical memory is configured for each NUMA node. Each NUMA node can access the memory space of its own configured local memory through its corresponding memory controller, and at the same time, can also access the memory space of other NUMA nodes.
[0041] In a NUMA architecture, each NUMA node corresponds to its own memory controller and memory. Any NUMA node can access its own corresponding memory and can also access the memory of any other node in the architecture. According to the relationship between NUMA nodes, NUMA nodes can be divided into local nodes, neighbor nodes, and remote nodes. The speed at which a processor core accesses the memory of different types of nodes is different. The speed of accessing a local node is the fastest, and the speed of accessing a remote node is the slowest.
[0042] Apache HTTP Server (abbreviation: httpd) is an open-source and widely used Web server software, known for its high configurability and flexibility. httpd uses a multi-process or multi-threaded architecture to handle client requests and mainly implements the processing of concurrent connections through the "process model". Common process models include Prefork, Worker, and Event, etc. These models allocate processes or threads in different ways to handle client requests. Among them, the Worker model processes client requests through threads, and improves the system's concurrent processing ability and resource utilization efficiency by generating multiple threads in each child process. This model also well balances the utilization rate of system resources and the response speed, but requires a more complex thread synchronization mechanism.
[0043] When running the httpd program in a multi-processor system adopting a NUMA architecture, the memory allocation of the httpd process is managed by the operating system and the memory allocation is carried out through the fork mechanism. As Figure 2As shown in the figure, the operating system first allocates a block of memory space for the httpd parent process on a certain NUMA node. The httpd parent process then creates each httpd child process, and then the httpd child process uses the fork system call to copy the entire address space of the parent process, including the code segment, data segment, heap, stack, etc. Therefore, when the httpd program is just started, all httpd child processes and the parent process are on the same NUMA node, or the operating system allocates the httpd child processes to the idle CPUs on each NUMA node. At this time, all httpd processes have the same virtual memory address space and point to the same physical memory area.
[0044] Under this httpd process allocation method, each httpd process is on the same NUMA node or is allocated to an idle CPU. These methods will make the distribution of httpd processes in a multi-processor system more concentrated. It is actually measured that the running performance of the httpd program under this process configuration method is relatively low.
[0045] In view of the above technical problems, this application proposes a process configuration scheme for running the httpd program in a multi-processor system using the NUMA architecture. Based on this scheme, the performance advantages of the NUMA architecture can be fully utilized, thereby improving the performance of the httpd program during operation.
[0046] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts shall fall within the scope of protection of this application.
[0047] The embodiments of this application first propose a process configuration method. This method is applied to a multi-processor system using the NUMA architecture, as described in the above embodiments and Figure 1 As shown in the figure, in this multi-processor system using the NUMA architecture, there are multiple NUMA nodes, and each NUMA node has local memory and one or more processors or processor cores.
[0048] The process configuration method proposed in the embodiments of this application is executed by the above multi-processor system using the NUMA architecture. Specifically, it can be executed by the operating system running on this multi-processor system, or by a specific application program running on this multi-processor system.
[0049] In some embodiments, the processing flow of the process configuration method proposed in this application is written into a program file, such as being written into an httpd_optimize.c file, and then the httpd_optimize.c file is compiled into an httpd_optimize.so dynamic library file by using the GCC compiler. When the httpd process starts, the multi-processor system preloads the compiled httpd_optimize.so file in the LD_PRELOAD manner, and by running this file, the processing flow of the process configuration method proposed in this application is executed to implement the configuration of the httpd process in the multi-processing system.
[0050] The following introduces the specific processing flow of the process configuration method proposed in this application.
[0051] See Figure 3 As shown, the process configuration method proposed in this application includes:
[0052] S101. When the httpd program starts, start monitoring the operation of the httpd program to create an httpd process.
[0053] Specifically, when the httpd program starts, it will first create an httpd parent process, and then the httpd parent process creates a child process by calling the fork function.
[0054] Therefore, when the httpd program starts, based on the start signal of the httpd program, it can be determined that the httpd program has created an httpd parent process. Then, when the system fork function is called, by detecting the operation of the system fork function being called, the operation of the httpd parent process creating an httpd child process can be identified. In the above manner, the real-time monitoring of the operation of the httpd program to create an httpd process can be realized, and the operation of the httpd program to create an httpd process can be identified in a timely manner.
[0055] S102. Whenever it is monitored that the httpd program creates an httpd process, bind the created httpd process to a first numa node, so as to disperse and bind each httpd process created by the httpd program to each first numa node of the multi-processor system.
[0056] Specifically, the above-mentioned first NUMA nodes refer to multiple NUMA nodes selected from multiple NUMA nodes of a multi-processor system. That is to say, in this embodiment, multiple NUMA nodes are selected from multiple NUMA nodes of the multi-processor system as multiple first NUMA nodes, and the specific number of selections can be flexibly set. In this embodiment, it is assumed that the multi-processor system has a total of K NUMA nodes, then K - 1 NUMA nodes are selected as the first NUMA nodes, that is, K - 1 first NUMA nodes are obtained.
[0057] During the process of monitoring the operation of the httpd program to create an httpd process in the manner of step S101, whenever it is monitored that the httpd program creates an httpd process, the created httpd process is bound to a first NUMA node. When it is monitored that the httpd program creates another httpd process, the created httpd process is bound to another first NUMA node. When it is monitored that the httpd program creates yet another httpd process, the created httpd process is bound to yet another first NUMA node, and so on. The httpd processes are bound to each first NUMA node in a polling manner, that is, each newly created httpd process is bound to a first NUMA node each time, so as to disperse and bind each httpd process created by the httpd program to each first NUMA node.
[0058] Among them, binding the httpd process to the first NUMA node means making the httpd process run on the first NUMA node. Exemplarily, by calling two system functions, CPU_SET() and sched_setaffinity() of the multi-processor system, the binding of the httpd process to the first NUMA node can be achieved.
[0059] After the above processing, the distribution of httpd processes in the multi-processor system is as Figure 4 shown. Among them, die0 to die6 are the first NUMA nodes respectively, and die7 is the other NUMA node. Refer to Figure 4 It can be seen that the number of httpd processes bound to each first NUMA node is balanced.
[0060] The process configuration method proposed in the embodiments of this application starts monitoring the operation of the httpd program to create an httpd process when the httpd program starts. Whenever it monitors that the httpd program creates an httpd process, the created httpd process is bound to a first NUMA node. This method can disperse and bind each httpd process created by the httpd program to each first NUMA node of the multi-processor system, so that the distribution of httpd processes on each first NUMA node used to run the httpd process is more uniform, avoiding the impact on the operation of httpd processes due to the concentration of httpd processes on a certain NUMA node. This httpd process configuration solution enables the httpd program to run with the hardware resources of each first NUMA node, thereby making more full use of the performance advantages of the NUMA architecture's processor system, ensuring the smooth operation of the httpd program, and facilitating the improvement of the running performance of the httpd program.
[0061] In another embodiment, in addition to dispersing and binding each httpd process created by the httpd program to each first NUMA node of the multi-processor system according to the processing described in the above embodiment, each second NUMA node of the multi-processor system is also configured to process network card interrupt requests.
[0062] The above-mentioned second NUMA node is the other NUMA nodes among the multiple NUMA nodes of the above-mentioned multi-processor system that are determined to be outside the first NUMA node. For example, in this embodiment, assuming that the multi-processor system has a total of K NUMA nodes, then K - 1 NUMA nodes are selected as the first NUMA nodes, that is, K - 1 first NUMA nodes are obtained, and the remaining 1 NUMA node is used as the second NUMA node.
[0063] In a modern high-performance network system, a network interface card (NIC) is a key component for data reception and transmission. When a data packet enters the system through the network card, it triggers an interrupt request (IRQ) to notify the CPU to process the data. With the continuous improvement of network speed and data throughput, the impact of interrupt processing on system performance becomes more and more significant. The network card interrupt binding (IRQ Affinity) technology optimizes interrupt processing and improves the overall system performance by binding network card interrupts to specific CPU cores.
[0064] When running the httpd program in a multi-processor system with a NUMA architecture, it is also necessary to perform the network card interrupt binding operation, that is, bind the network card interrupt to a specific CPU or CPU core among multiple NUMA nodes.
[0065] In this embodiment, one or more second NUMA nodes are selected from multiple NUMA nodes of a multi-processor system and configured to process network card interrupt requests, that is, bind the network card interrupts to one or more second NUMA nodes.
[0066] As Figure 5 shown, in a multi-processor system including multiple NUMA nodes, die0 to die6 are the first NUMA nodes respectively for binding the httpd process, and die7 is the second NUMA node to which all network card interrupts are bound.
[0067] Through the above configuration, the httpd process and network card interrupts can run through different NUMA nodes respectively. In this way, the multi-processor system can be divided into an httpd process domain and a network card interrupt binding domain. The httpd process domain composed of each first NUMA node is responsible for running the httpd process and does not process network card interrupt requests, while the network card interrupt domain composed of the second NUMA nodes is specifically responsible for processing network card interrupt requests and is not responsible for running the httpd process. Such a configuration method can avoid the impact of network card interrupt requests on the running of the httpd process. At the same time, since the network card interrupt requests are processed by dedicated second NUMA nodes, the processing efficiency of network card interrupt requests can be improved, and thus the overall network performance of the system can be improved.
[0068] In another embodiment, to achieve dispersedly binding each httpd process created by the httpd program to each first NUMA node, the following processing can be implemented:
[0069] First, multiple first NUMA nodes are determined from multiple NUMA nodes of the processor system. For example, K - 1 first NUMA nodes are selected from K NUMA nodes to obtain K - 1 first NUMA nodes.
[0070] Then, whenever it is monitored that the httpd program creates a new httpd process, a first NUMA node is selected from multiple first NUMA nodes, and the newly created httpd process is bound to the selected first NUMA node.
[0071] In some embodiments, each time a first NUMA node is selected from multiple first NUMA nodes for binding the httpd process, the first NUMA node with the lowest selection times is preferentially selected from multiple first NUMA nodes. When there are multiple first NUMA nodes with the lowest selection times, a first NUMA node is randomly selected from these multiple first NUMA nodes with the lowest selection times.
[0072] In some other embodiments, in order to evenly bind each httpd process created by the httpd program to each first NUMA node, during the process of the httpd program creating an httpd process, multiple first NUMA nodes can be traversed in a loop.
[0073] During the above loop traversal process, whenever a first NUMA node is traversed, wait for the monitoring result of the operation of the httpd program creating an httpd process. When it is confirmed that the httpd program has created an httpd process, bind the created httpd process to the traversed first NUMA node.
[0074] After the binding is completed, traverse the next first NUMA node. At this time, wait for the monitoring result of the operation of the httpd program creating an httpd process. When it is confirmed that the httpd program has created an httpd process, bind the created httpd process to the traversed first NUMA node.
[0075] And so on, continuously traverse each first NUMA node in a loop, and implement the binding of each httpd process to each first NUMA node during this process. Stop the loop traversal of each first NUMA node until it is confirmed that the httpd program has finished creating httpd processes.
[0076] For example, when the number of httpd processes created by the httpd program is fixed, when it is counted and determined that the httpd program has created all the httpd processes, the loop traversal of each first NUMA node can be stopped.
[0077] The above solution adopts the method of traversing each first NUMA node while binding httpd processes, achieving an efficient and uniform decentralized binding of each httpd process to each first NUMA node.
[0078] In the above solution, each httpd process is decentralized and bound to multiple first NUMA nodes. For example Figure 2As shown in the figure, when the operating system allocates memory space for the httpd program, it only allocates a block of memory space for the httpd parent process on a certain NUMA node. Then, other child processes of httpd use the fork system call to copy the entire memory space of the parent process, including the code segment, data segment, heap, stack, etc. The copying process does not immediately generate actual physical memory, but instead uses the Copy-On-Write (COW) technology. The child process and the parent process share the same memory pages until one of the processes attempts to modify a certain page, at which point the operating system allocates a new physical page for that process. Therefore, when the httpd program is just started, all httpd processes have the same virtual memory address space and point to the same physical memory area.
[0079] In this case, after dispersedly binding each httpd process to multiple first NUMA nodes through the solution of the above embodiment of the present application, since each child process accesses the same memory area, a situation where most httpd processes access memory across NUMA nodes will occur, which will generate a relatively large memory access latency, thereby affecting system performance.
[0080] To address the above problem, the process configuration method proposed in the embodiment of the present application further copies the memory files required to be accessed by each httpd process, thereby avoiding the situation where httpd processes access memory across NUMA nodes.
[0081] To achieve the above objective, in addition to dispersedly binding each httpd process to each first NUMA node, the process configuration method proposed in the embodiment of the present application also creates a copy file of the memory file required by the target httpd process bound to the first NUMA node in the first NUMA node for each first NUMA node.
[0082] On this basis, map the memory access address of the target httpd process bound to the first NUMA node to the copy file, and load the copy file into the memory of the first NUMA node, so that the target httpd process bound to the first NUMA node can access the memory file it needs in the local memory of the first NUMA node.
[0083] In some embodiments, identify the memory files accessed by each httpd process from the memory space allocated by the operating system for the httpd program, that is, identify the memory files required by each httpd process, and then copy the memory files required by each httpd process to the first NUMA node where each httpd process is located to generate copy files.
[0084] By calling the mmap function of the system, the memory access address of the httpd process can be mapped to the copy file in the first NUMA node where it is located, and then the copy file can be loaded into the memory of the first NUMA node in units of pages.
[0085] In another embodiment, when creating a copy file of a memory file, the following processing steps A1 to A3 can be executed to achieve it:
[0086] A1. Read the maps file of the target httpd process bound to the first NUMA node, and by parsing the maps file, determine the memory start address of each shared file that the target httpd process needs to access and the memory file name stored at the memory start address.
[0087] Among them, the above-mentioned maps file is used to record the mapping information of the virtual memory area related to the target httpd process. When the pid number of the httpd process is obtained, the maps file of the httpd process can be read through the path / proc / httpd process pid number / maps.
[0088] After reading the maps file of the target httpd process bound to the first NUMA node, parse the maps file line by line, and identify the memory start address that the target httpd process needs to access and the memory file name stored at the memory start address from it until all lines of the maps file are scanned.
[0089] In some embodiments, when parsing the maps file of the httpd process, parse the maps file line by line and search for a set field from each line of the maps file.
[0090] The above-mentioned set field includes the first field "r-xp" corresponding to the executable code segment that the target httpd process needs to access and / or the second field "r--p" corresponding to the read-only data segment that the target httpd process needs to access.
[0091] When the set field is searched, determine the memory start address corresponding to the set field and the memory file name stored at the memory start address.
[0092] For example, assume that the content of a certain line of the maps file is:
[0093] 00400000-004a1000r-xp 00000000fd:00 6293196 / usr / local / apache2 / bin / httpd
[0094] By parsing this line, the first field "r-xp" can be identified from it. Based on this, the starting memory address "00400000 - 004a1000" corresponding to the first field "r-xp" can be further determined from the content of this line, as well as the memory file name "httpd" stored at this starting memory address.
[0095] A2. According to the set naming rule for the copy file name, based on the memory file name, determine the file name of the copy file to be created.
[0096] Specifically, after determining the starting memory address that the target httpd process needs to access and the memory file name stored at this starting memory address through the above step A1, according to the set naming rule for the copy file name, use the memory file name of the memory file that the target httpd process needs to access to determine the file name of the copy file to be created.
[0097] For example, assuming the naming rule of "memory file name_numa node id" is adopted, and the memory file name of the memory file that the target httpd process needs to access is nameSo, then determine that the file name of the copy file to be created on the first numa node with "numa node id" is "nameSo_numa node id".
[0098] A3. In the case where it is determined that the file name of the copy file does not exist in the first numa node, create a copy file according to the file name, and write the file content stored at the starting memory address into the copy file.
[0099] Specifically, since multiple httpd processes will be bound on the same first numa node, and some httpd processes will access the same memory file, and other httpd processes will create copy files on this first numa node before the target httpd process, therefore, to avoid the situation of repeatedly creating the same copy file, when determining the file name of the copy file to be created, first judge whether the copy file exists in the first numa node where the copy file is to be created according to the file name of the copy file to be created.
[0100] If it exists, do not create the copy file repeatedly.
[0101] If it does not exist, create a copy file according to the file name of the copy file, and write the file content stored at the starting memory address into the copy file.
[0102] As a more specific example, in Figure 5In the multi-processor system with 8 NUMA nodes shown, die0 to die6 are used as the first NUMA node, and die7 is used as the second NUMA node. Through the following steps, the binding of each httpd process to each first NUMA node can be achieved, and the backup of the memory files required by the httpd process in the first NUMA node where the httpd process is located can be realized:
[0103] Step 1: Set the global variable NUMAId of the first NUMA node ID to 0, and start the httpd program.
[0104] Step 2: Intercept the fork system function called when the httpd program creates an httpd process, and read the pid number of the httpd process.
[0105] Step 3: Bind the httpd process to the first NUMA node with the ID of NUMAId.
[0106] Step 4: Read the maps file of the httpd process line by line, and the specific path is / proc / pid number of the httpd process / maps.
[0107] Step 5: Determine whether the current line of the maps file contains the fields "r--p" or "r-xp". If not, read the next line and execute Step 5 again; if so, execute Step 6.
[0108] Step 6: Read the start and end addresses start--end of the shared memory corresponding to the "r--p" or "r-xp" field, and the file name nameSo.
[0109] Step 7: Determine the file name nameSo_NUMAId of the copy file to be created according to the set naming rule.
[0110] Step 8: Determine whether the nameSo_NUMAId file exists in the first NUMA node with the ID of NUMAId. If not, execute Step 9; if so, read the next line of the maps file and return to execute Step 5.
[0111] Step 9: Copy the file content stored in the start address of the shared memory area start--end to the local memory tmpcode through the memcpy function, and then write the tmpcode memory to the nameSo_NUMAId file.
[0112] Step 10: Map the memory mapping relationship of the httpd process originally referring to the content of the start--end address to the nameSo_NUMAId file through the mmap function.
[0113] Step 11: Load the nameSo_NUMAId file into the memory of the first numa node with the Id of NUMAId page by page.
[0114] Step 12: Update the value of the NUMAId variable, NUMAId = (NUMAId + 1) % 7.
[0115] Step 13: Read the pid number of the next httpd process. If there is one, jump to Step 2; if not, terminate.
[0116] Through the process configuration method of the above embodiments of the present application, when the httpd program starts, each httpd process can be quickly and dispersedly bound to multiple first numa nodes, and a copy file of the memory file required for the httpd process is created locally on the first numa node where the httpd process is located, so that each httpd process can access the required memory file locally on the first numa node where it is located. At the same time, the above method also binds the system network card interrupt to a dedicated first numa node, thereby improving the processing efficiency of the network card interrupt request. The above method can improve the processing efficiency of the httpd process on the multi-processor system for processing network httpd requests.
[0117] Corresponding to the above process configuration method, an embodiment of the present application also proposes a multi-processor system. The multi-processor system adopts a numa architecture and includes multiple numa nodes. In the multi-processor system, whenever the httpd program creates an httpd process, the created httpd process is bound to a first numa node, so that each httpd process created by the httpd program is dispersedly bound to each first numa node of the multi-processor system.
[0118] The structure of this multi-processor system can be seen Figure 4 As shown, in this multi-processor system, whenever the httpd program creates an httpd process, the specific implementation process of binding the created httpd process to a first numa node can refer to the corresponding processing process of the process configuration method introduced in the above embodiments, which will not be repeated here.
[0119] In the multi-processor system described above, each httpd process created by the httpd program is scattered and bound to each of the first numa nodes of the multi-processor system, so that the distribution of httpd processes on each of the first numa nodes used to run the httpd processes is more uniform, avoiding the impact on the operation of the httpd process due to the concentration of httpd processes on a certain numa node. This httpd process configuration scheme enables the httpd program to run with the hardware resources of each of the first numa nodes, thus making more full use of the performance advantages of the processor system with the numa architecture, ensuring the smooth operation of the httpd program, and facilitating the improvement of the operation performance of the httpd program.
[0120] In some other embodiments, the second numa node in the multi-processor system described above is configured to handle network card interrupt requests.
[0121] At this time, the structure of the multi-processor system is as Figure 5 shown. In this multi-processor system, it is possible to make the httpd process and the network card interrupt run through different numa nodes respectively. In this way, the multi-processor system can be divided into an httpd process domain and a network card interrupt binding domain. The httpd process domain composed of each of the first numa nodes is responsible for the operation of the httpd process and does not handle network card interrupt requests, while the network card interrupt domain composed of the second numa nodes is specifically responsible for handling network card interrupt requests and is not responsible for the operation of the httpd process. Such a configuration method can avoid the impact of network card interrupt requests on the operation of the httpd process. At the same time, since the network card interrupt requests are processed by the dedicated second numa node, the processing efficiency of the network card interrupt requests can be improved, and thus the overall network performance of the system can be improved.
[0122] In some other embodiments, in the memory of the first numa node in the multi-processor system, a copy file of the memory file required by the target httpd process bound to the first numa node is further set, and the memory access address of the target httpd process is mapped to the copy file.
[0123] In the multi-processor system of this embodiment, a copy file of the memory file required for access by the httpd process is created locally on the first numa node where the httpd process is located, so that each httpd process can access the memory file it needs locally on the first numa node where it is located. In this way, the processing efficiency of the httpd process on the multi-processor system for handling network httpd requests can be improved.
[0124] The multi-processor system provided in this embodiment belongs to the same inventive concept as the process configuration method provided in the foregoing embodiments of the present application, and can execute the process configuration methods provided in any of the foregoing embodiments of the present application, and has the corresponding functional modules and beneficial effects of the execution method. For the technical details not described in detail in this embodiment, reference may be made to the specific processing content of the process configuration method provided in the foregoing embodiments of the present application, and details will not be repeated here.
[0125] Another embodiment of the present application further provides a computing device, which includes the multi-processor system described in the foregoing embodiments of the present application. The computing device can be a computer, a server, a workstation, an intelligent terminal, a handheld terminal, a wearable device, etc.
[0126] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0127] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0128] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0129] The modules and sub-modules in the devices and terminals in the embodiments of the present application can be combined, divided, and deleted according to actual needs.
[0130] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in an electrical, mechanical or other form.
[0131] A module or sub-module described as a separate component may or may not be physically separated. A component as a module or sub-module may or may not be a physical module or sub-module, that is, it may be located in one place or may be distributed across multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0132] In addition, each functional module or sub-module in various embodiments of the present application can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware or in the form of software functional modules or sub-modules.
[0133] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0134] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software unit executed by a processor, or a combination of the two. The software unit can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0135] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0136] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A process configuration method, characterized in that: Applied to a multiprocessor system adopting the NUMA architecture, the multiprocessor system includes a plurality of NUMA nodes, and the method includes: When the httpd program is started, start monitoring the operation of the httpd program to create the httpd process; Whenever it is detected that the httpd program creates an httpd process, the created httpd process is bound to a first numa node, so that each httpd process created by the httpd program is dispersedly bound to each first numa node of the multi-processor system.
2. The method according to claim 1, characterized in that The method further comprises: The second numa node in the multiprocessor system is configured to process a network card interrupt request.
3. The method according to claim 1, characterized in that Whenever it is detected that the httpd program creates an httpd process, the created httpd process is bound to a first numa node, including: Determining a plurality of first numa nodes from the plurality of numa nodes of the processor system; Whenever it is monitored that the httpd program creates an httpd process, a first numa node is selected from the plurality of first numa nodes, and the created httpd process is bound to the selected first numa node.
4. The method according to claim 3, characterized in that Whenever it is detected that the httpd program creates an httpd process, a first numa node is selected from the plurality of first numa nodes, and the created httpd process is bound to the selected first numa node, including: In the process of the httpd program creating an httpd process, the multiple first numa nodes are traversed in a loop. When each first numa node is traversed, an httpd process created by the httpd program is bound to the first numa node and the next first numa node is traversed after the binding is completed, until the httpd program finishes creating the httpd process, and the traversal of the multiple first numa nodes is stopped.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: In the first numa node, creating a copy file of a memory file required by a target httpd process bound to the first numa node; The memory access address of the target httpd process is mapped to the copy file, and the copy file is loaded into the memory of the first numa node.
6. The method according to claim 5, characterized in that In the first NUMA node, a copy file of a memory file required by a target httpd process bound to the first NUMA node is created, including: Reading a maps file of a target httpd process bound to the first numa node, and determining a memory start address that the target httpd process needs to access and a memory file name stored at the memory start address by parsing the maps file; the maps file is used to record mapping information of a virtual memory area related to the target httpd process; According to the set copy file name naming rule, based on the memory file name, determine the file name of the copy file to be created; When it is determined that the file name of the copy file does not exist in the first numa node, a copy file is created according to the file name, and the file content stored in the memory start address is written into the copy file.
7. The method according to claim 6, characterized in that By parsing the maps file, determining the memory start address that the target httpd process needs to access and the memory file name stored at the memory start address, including: Parsing the maps file line by line, and searching for a setting field from each line of the maps file, wherein the setting field includes a first field corresponding to an executable code segment that the target httpd process needs to access and / or a second field corresponding to a read-only data segment that the target httpd process needs to access; When the setting field is found, a memory start address corresponding to the setting field and a memory file name stored at the memory start address are determined.
8. A multiprocessor system, characterized in that: The multi-processor system adopts the NUMA architecture and includes multiple NUMA nodes. In the multi-processor system, whenever the httpd program creates an httpd process, the created httpd process is bound to a first NUMA node, so that each httpd process created by the httpd program is dispersedly bound to each first NUMA node of the multi-processor system.
9. The multiprocessor system according to claim 8, characterized in that: The second numa node in the multiprocessor system is configured to process a network card interrupt request.
10. The multiprocessor system according to claim 8 or 9, characterized in that: In the memory of the first numa node, a copy file of the memory file required by the target httpd process bound to the first numa node is also provided, and the memory access address of the target httpd process is mapped to the copy file.
11. A computing device, characterized in that: A multi-processor system comprising the process described in any one of claims 8 to 10.