NUMA Architecture-based Pipeline Communication Optimization Method and System

By allocating memory for pipelines in NUMA architecture servers and migrating processes according to communication rates, the performance problems caused by pipeline communication across node memory access under NUMA architecture are solved, and more efficient pipeline communication is achieved.

CN119883953BActive Publication Date: 2025-06-20GALAXY UNICORN SOFTWARE (CHANGSHA) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510334110.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-20
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The performance problems caused by cross-node memory access during pipeline communication under NUMA architecture server cannot be effectively solved by the existing technology.

Method used

When creating a pipeline, the pipeline communication memory is allocated to the specified node node, the communication data rate of each pipeline is detected, and the threshold is calculated based on the process type and priority, and the important process or process with a communication data rate greater than the threshold is migrated to the specified node node to run.

Benefits of technology

By optimizing pipeline communication, avoiding memory access across node nodes, improving pipeline communication bandwidth and reducing latency, giving full play to the hardware characteristics of the NUMA architecture, and improving system operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883953B_ABST
    Figure CN119883953B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for optimizing pipeline communication under a NUMA architecture. The method includes the following steps: when creating a pipeline, allocate pipeline communication memory on a specified node; detect the communication data rate of each pipeline; calculate the threshold of each pipeline according to the type and priority of the process, compare the communication data rate of each pipeline with the corresponding threshold, migrate important processes or processes with a communication data rate greater than the threshold of the corresponding pipeline to the specified node for running, and retain unimportant processes or processes with a communication data rate less than the threshold of the corresponding pipeline to run on the corresponding original node. The present invention can give full play to the characteristics of the NUMA hardware architecture, thereby improving the system operation efficiency and solving the performance problem caused by cross-node memory access of pipeline communication in NUMA architecture servers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer technology, and in particular to a method and system for optimizing pipeline communication under a NUMA architecture. Background Art

[0002] The NUMA architecture is suitable for large-scale multi-processor systems. Especially when the performance improvement of a single processor is limited by physical limits, the overall computing power of the hardware can be improved by increasing the number of processors. Currently, domestic mainstream CPU manufacturers basically design server CPUs based on the NUMA architecture. For example, Feiteng S5000C, Huawei HiSilicon Kunpeng 920, Haiguang 4th, etc. are all NUMA architectures. Because the memory access time of the NUMA architecture depends on the relative position between the processor and the memory, the processor can access the local memory faster than accessing the non-local memory. Software design needs to consider the characteristics of NUMA to give full play to the hardware characteristics and optimize the system performance.

[0003] As an important inter-process communication method in the Linux system, pipeline communication allows different processes to share data and transmit messages, and is one of the important mechanisms for realizing inter-process synchronization and communication. The communication efficiency will directly affect the running efficiency of the process. However, the current design and implementation of pipeline communication in the Linux kernel do not fully consider the characteristics of the NUMA hardware architecture. For example, when an important process communicates across node nodes through a pipeline, the communication bandwidth will be limited by the cross-node communication bandwidth, and the pipeline communication delay will also increase, thus affecting the running efficiency of the important process.

[0004] The Chinese invention patent application "A Pipeline Communication Method and Device" (application number CN201310034616.6) discloses a pipeline communication method and device, including establishing a data storage space, storing data in the data storage space, assigning characteristic parameters corresponding to the data stored in the data storage space, reading the corresponding data from the data storage space according to the characteristic parameters, connecting to the input and output pipelines through a common space, uniformly managing and storing the data volume input by the input pipeline, and any output pipeline of the processes in the system can read the required data volume from the common space, thus greatly reducing the scale and complexity of the pipeline communication system and improving the reliability of the system. This solution mainly aims to improve the system reliability by reducing the scale and complexity of the pipeline communication system, and cannot solve the performance problems caused by cross-node memory access during pipeline communication in a NUMA architecture server.

[0005] The Chinese invention patent application "Memory Access Method, Device and System" (Application No. CN201310257057.5) discloses a memory access method, device and system. The memory access method includes that a node controller receives monitoring information sent by an operating system. The monitoring information carries information about the monitored memory in the first node to which the node controller belongs. The monitored memory is the memory resource occupied by a target process on the first node. The target process is a process that runs on the central processing unit (CPU) of the first node and accesses the memory of an accessed node other than the first node in the server system. If it is monitored that the frequency of the target process occupying the monitored memory accessing the memory of the accessed node is greater than or equal to a threshold, the information of the accessed node is sent to the operating system to migrate the target process to the accessed node according to the information of the accessed node; convert remote memory access into local memory access or adjacent memory access, so as to reduce the time for the target process to access memory and effectively improve the performance of the server system. This solution is mainly implemented through the hardware of the node controller NC chip. The operating system sends the monitoring information required to the NC chip hardware, and then the NC chip hardware monitors. When the NC chip hardware monitors that the frequency of remote memory access is too high, it sends information to the operating system for process migration. The implementation of the solution highly depends on the hardware node controller NC chip and is not a general implementation solution of the operating system itself, and cannot solve the performance problem caused by cross-node memory access during the operating system pipe communication. Summary of the Invention

[0006] The technical problem to be solved by the present invention: Aiming at the above problems of the prior art, a method and system for optimizing pipe communication under the NUMA architecture are provided, which can give full play to the characteristics of the NUMA hardware architecture, thereby improving the system operation efficiency and solving the performance problem caused by cross-node memory access in the pipe communication of the NUMA architecture server.

[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0008] A method for optimizing pipe communication under the NUMA architecture includes the following steps:

[0009] When creating a pipe, allocate pipe communication memory on a specified node.

[0010] Detect the communication data rate of each pipe.

[0011] Calculate the threshold of each pipe according to the type and priority of the process, compare the communication data rate of each pipe with the corresponding threshold, migrate important processes or processes with the communication data rate of the corresponding pipe greater than the threshold to the specified node to run, and retain unimportant processes or processes with the communication data rate of the corresponding pipe less than the threshold on the corresponding original node to run.

[0012] Further, the specified node is specifically the node specified by the user or the node where the process creating the pipeline runs.

[0013] Further, when detecting the communication data rate of each pipeline, it includes:

[0014] Obtain the historical communication data rate of the current pipeline, and wait for the process corresponding to the current pipeline to write data to the current pipeline or read data from the current pipeline;

[0015] When the process writes data to the current pipeline or reads data from the current pipeline, accumulate and record the read data or write data in the current pipeline during the current time period to obtain the communication data volume during the current time period;

[0016] Divide the communication data volume during the current time period by the duration of the current time period to obtain the current communication data rate of the current pipeline;

[0017] Smooth the current communication data rate using the historical communication data rate and a preset amplification factor.

[0018] Further, when smoothing the current communication data rate using the historical communication data rate and a preset amplification factor, the expression is as follows:

[0019] S = Cur_S * GF + Pre_S * (1 - GF)

[0020] Among them, S represents the communication data rate after smoothing, Cur_S represents the current communication data rate, Pre_S represents the historical communication data rate, GF represents the amplification factor, and the value range is (0, 1].

[0021] Further, after waiting for the process corresponding to the current pipeline to write data to the current pipeline or read data from the current pipeline, it further includes:

[0022] Determine the decay time according to the priority of the process corresponding to the current pipeline. If the waiting time exceeds the decay time, decay the historical communication data rate using a preset decay factor to obtain the current communication data rate, and then continue to wait for the process corresponding to the current pipeline to write data to the current pipeline or read data from the current pipeline.

[0023] Further, when decaying the historical communication data rate using a preset decay factor, the expression is as follows:

[0024] S = Pre_S * AF

[0025] Wherein, S represents the attenuated communication data rate, Pre_S represents the historical communication data rate, and AF represents the attenuation factor, with a value range of (0, 1].

[0026] Further, when calculating the threshold of each pipeline according to the type and priority of the process, it includes:

[0027] Obtain the scheduling policy of the current process;

[0028] If the scheduling policy of the current process is a real-time scheduling policy, set the threshold of the pipeline corresponding to the current process to 0;

[0029] If the scheduling policy of the current process is not a real-time scheduling policy, obtain the priority of the current process, substitute the priority into the threshold calculation function, and obtain the threshold of the pipeline corresponding to the current process.

[0030] Further, the threshold calculation function is:

[0031] pipe->threshold = PIPE_THRESHOLD * pow(a, task_nice(current))

[0032] Wherein, PIPE_THRESHOLD represents the threshold when the process priority is 0, the pow(a, task_nice(current)) function implements the function of the power operation logic of a, task_nice() is the function to obtain the process priority, and current is the struct task_struct structure variable representing the current process.

[0033] Further, before obtaining the scheduling policy of the current process, it also includes:

[0034] If the current process is a specified important process, set the threshold of the pipeline corresponding to the current process to 0;

[0035] If the current process is a specified unimportant process, set the threshold of the pipeline corresponding to the current process to the maximum value of the storage type.

[0036] The present invention also proposes a pipeline communication optimization system under the NUMA architecture, including a microprocessor and a computer-readable storage medium connected to each other, and the microprocessor is programmed or configured to execute any one of the pipeline communication optimization methods under the NUMA architecture.

[0037] Compared with the prior art, the advantages of the present invention are:

[0038] In view of the characteristics of NUMA architecture hardware, the present invention realizes the detection of the communication bandwidth of each pipeline in the system by calculating the communication data rate of the pipeline, and at the same time differentiates the importance of system processes. For processes with a relatively small pipeline communication bandwidth or unimportant processes, cross-node memory communication is allowed. For processes with a relatively large pipeline communication bandwidth or important processes, they are scheduled to communicate on the same node to improve their pipeline communication bandwidth and reduce pipeline communication latency, giving full play to the hardware characteristics of the NUMA architecture and improving the system operation efficiency. Brief Description of the Drawings

[0039] Figure 1 It is a brief flowchart of an embodiment of the present invention.

[0040] Figure 2 It is a flowchart of the communication data rate calculation of an embodiment of the present invention. Detailed Embodiment

[0041] The present invention will be further described below in conjunction with the accompanying drawings of the specification and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.

[0042] Before introducing the specific embodiments of the present invention, relevant concepts or terms will be explained first.

[0043] NUMA: NUMA (Non Uniform Memory Access) non-uniform memory access is a design for multi-processors, and its memory access time depends on the memory location of the processor. In the NUMA architecture, the concept of node is added. Each node has its own internal CPU, bus, memory, and I / O slots, etc. At the same time, the internal CPU of each node can also access the memory and I / O slots in other nodes through the interconnection bus. Therefore, each CPU can access the memory of the entire system. However, under NUMA, a processor accessing its own local memory (that is, the CPU and the memory are in the same node) is much faster in terms of speed and latency than accessing non-local memory (also called remote memory, that is, the CPU and the memory are not in the same node), which is the origin of non-uniform memory access NUMA. Due to this characteristic, in order to better exert the system performance, it is necessary to minimize the information interaction between different nodes when developing application programs.

[0044] Pipe: Pipe is a Linux inter-process communication mechanism that allows one process to pass data to another process. There are two main types of pipes: anonymous pipes and named pipes. Anonymous pipes are usually used for communication between parent and child processes or sibling processes, while named pipes allow communication between unrelated processes.

[0045] Embodiment 1

[0046] An optimization method for pipeline communication under the NUMA architecture. By detecting the pipeline communication bandwidth, processes with larger communication bandwidth or important processes are scheduled to run on the same node to prevent cross-node memory communication access. For non-important processes with smaller communication bandwidth, cross-node memory communication is allowed, giving full play to the characteristics of the NUMA hardware architecture, thereby improving the system operation efficiency.

[0047] As Figure 1 shown, the method of this embodiment includes the following steps:

[0048] S1) Declare and define the required data structures and initialize them;

[0049] S2) When creating a pipeline, allocate pipeline communication memory on the specified node;

[0050] S3) Detect the communication data rate of each pipeline;

[0051] S4) Calculate the threshold for each pipeline according to the type and priority of the process, compare the communication data rate of each pipeline with the corresponding threshold, and migrate important processes or processes with a communication data rate greater than the threshold of the corresponding pipeline to the specified node to run to avoid cross-node memory access and improve the efficiency of its pipeline communication. Retain non-important processes or processes with a communication data rate less than the threshold of the corresponding pipeline to run on the corresponding original node to achieve cross-node memory communication.

[0052] The following is a specific description of each step.

[0053] In this embodiment, step S1 is to define and initialize the data and variables to be used in the subsequent steps, specifically including:

[0054] S11) In the `struct pipe_inode_info` structure in the `include / linux / pipe_fs_i.h` file, define the variable `int nid` to specify the node where the memory for pipe communication is allocated, define the variable `unsigned long traffic_num` to record the communication data volume of the pipe for a period of time, define the variable `unsigned long pipe_time` to calculate the time point of the pipe communication data rate, define the variable `int pipe_speed` to store the calculated pipe communication data rate, define the variable `int threshold` to store the threshold calculated through the process priority, define the variable `int decay_time` to store the decay time of the process pipe communication, define the variable `struct timer_list speed_timer` for the pipe communication rate timer, define the pipe communication rate increase factor GF constant as 0.6, and define the pipe communication rate decay factor AF constant as 0.3;

[0055] S12) After allocating the memory of the `struct pipe_inode_info` structure in the kernel function `alloc_pipe_info()`, initialize `traffic_num` and `pipe_time` in the structure to 0, and call `timer_setup(&pipe->speed_timer, speed_timer_function, 0)` to initialize the timer function as `speed_timer_function`. The `speed_timer_function` function completes the decay of the pipe communication data rate, and the decay factor is the constant AF.

[0056] In step S2 of this embodiment, when the process creates a pipe, the memory for pipe communication is allocated to a specific node. In step S2, the specified node can be the node specified by the user. If the user does not specify, it can also be defaulted to the node where the process creating the pipe runs.

[0057] In the default case, the specific implementation code is as follows:

[0058] S21) Set the node for allocating the pipe communication memory as the running node where the current process creating the pipe is located: `pipe->nid = cpu_to_node(task_cpu(current))`; where `task_cpu(current)` is to obtain the CPU number where the current process creating the pipe is located, and `cpu_to_node()` is to obtain the node to which the specified CPU number belongs;

[0059] S22) Allocate pipe communication memory on the node specified in step S21: pipe->bufs = kmalloc_node(pipe_bufs * sizeof(struct pipe_buffer), GFP_KERNEL_ACCOUNT|__GFP_ZERO, pipe->nid); where pipe->nid represents the specified node, pipe_bufs * sizeof(struct pipe_buffer) represents the size of the allocated memory, pipe_bufs is an integer representing the number of pipe buffers to be allocated, sizeof(struct pipe_buffer) is the size of each pipe buffer, and struct pipe_buffer is a structure used to represent the buffer in the pipe. GFP_KERNEL_ACCOUNT is a flag for kernel memory allocation, indicating that memory is allocated in kernel space and the allocated memory will be counted in the process's memory usage statistics; __GFP_ZERO means the allocated memory will be cleared. That is, after the allocation is completed, the allocated memory area will be initialized to all zeros.

[0060] In the case specified by the user, only step S22 is executed. Memory is allocated according to the node specified by the user. The allocated memory will be cleared and the allocated memory will be counted in the process's memory usage statistics.

[0061] Step S3 of this embodiment aims to detect the communication data rate of its own communication link for each pipe. Using the passive measurement method, when the process writes data to the pipe or reads data from the pipe, the total data volume for a period of time is counted, and then the communication data rate can be measured by dividing the total data volume written or read by the time value for counting the data volume written or read. It includes the following steps:

[0062] S31) Obtain the historical communication data rate of the current pipe. Wait for the process corresponding to the current pipe to write data to the current pipe or read data from the current pipe. Determine the decay time according to the priority of the process corresponding to the current pipe. If the waiting time exceeds the decay time, use a preset decay factor to decay the historical communication data rate to obtain the current communication data rate, and then continue to wait for the process corresponding to the current pipe to write data to the current pipe or read data from the current pipe;

[0063] Specifically, when the pipeline has not communicated for a period of time, it is necessary to attenuate the measured pipeline communication data rate. Assume that the historical communication data rate of the current pipeline, that is, the communication data rate measured last time, is Pre_S, and the attenuation factor is AF. The value range of AF is [0, 1), that is, AF is greater than or equal to 0 and less than 1. Assume that the attenuation time is DT. Then, when the pipeline does not detect communication data every DT time period, use the preset attenuation factor to attenuate the historical communication data rate. The expression is as follows:

[0064] S = Pre_S *AF;

[0065] It can be seen from the calculation formula that the smaller the attenuation factor AF, the faster the pipeline communication rate S attenuates. When the attenuation factor AF is 0, it means that when the pipeline has not communicated for a period of time, the pipeline communication rate S directly attenuates to 0;

[0066] In this embodiment, the attenuation time DT can be uniformly set to a default empirical value or determined by the process priority. When the attenuation time DT is determined by the process priority, the higher the process priority, the larger the attenuation time DT, and the lower the process priority, the smaller the attenuation time DT;

[0067] S32) When the process writes data to the current pipeline or reads data from the current pipeline, accumulate and record the read data or write data in the current time period of the current pipeline to obtain the communication data volume in the current time period;

[0068] S33) Divide the communication data volume in the current time period by the duration of the current time period to obtain the current communication data rate of the current pipeline, that is, S = N / T, where S is the pipeline communication data rate, N is the amount of data written to the pipeline or read from the pipeline within time T. Then, use the historical communication data rate and the preset amplification factor to smooth the current communication data rate;

[0069] Specifically, in order to prevent the instantaneous measurement of a very high pipeline communication data rate from causing frequent scheduling migrations of the process, an amplification factor GF is introduced in this embodiment to smooth the measured pipeline communication data rate. The value range of GF is (0, 1], that is, GF is greater than 0 and less than or equal to 1. Assume that the historical pipeline communication data rate, that is, the communication data rate measured last time, is Pre_S, and the currently measured pipeline communication data rate is Cur_S. Then, when using the historical communication data rate and the preset amplification factor to smooth the current communication data rate, the expression is as follows:

[0070] S = Cur_S * GF + Pre_S * (1 - GF);

[0071] As can be seen from the calculation formula, the smaller the amplification factor GF, the smaller the contribution of the currently measured pipeline communication data rate to the pipeline communication data rate S. When the pipeline communication data rate is increasing, the growth rate of the calculated pipeline communication data rate S is smaller. When the pipeline communication data rate is decreasing, the deceleration of the calculated pipeline communication data rate S is also smaller. That is to say, the smaller the amplification factor GF, the smoother the calculated pipeline communication rate. When the amplification factor GF is 1, it means that no smoothing process is performed.

[0072] The specific process of the above steps is as Figure 2 shown, specifically including;

[0073] 1) Obtain the current system time point: unsigned long cur_time = jiffies;

[0074] 2) Reset the timer at 1-second intervals: mod_timer(&pipe->speed_timer, jiffies + msecs_to_jiffies(1000)). That is, if the pipeline communication data is 0 within a 1-second time period, it means that the waiting process has exceeded the decay time when writing data to or reading data from the current pipeline, and the pipeline communication rate will be decayed. The decay amplitude each time is AF, and the calculation method is pipe_speed = pipe_speed * AF, where pipe_speed is the pipeline communication data rate. Among them, mod_timer() is a function implemented by the kernel to modify the timeout time of an existing timer, and msecs_to_jiffies() is a macro implemented by the kernel to convert milliseconds (ms) to the kernel's timing unit jiffies;

[0075] 3) When the system calls the kernel function pipe_write() to send data to the pipeline or calls the kernel function pipe_read() to read data from the pipeline, accumulate and record the amount of data sent or read: pipe->traffic_num = pipe->traffic_num + total_len, where total_len is the amount of data sent or read, in bytes B;

[0076] 4) Determine the difference between cur_time and pipe->pipe_time: int pass_time = jiffies_to_msecs(cur_time - pipe->pipe_time). If pass_time is greater than or equal to 1000, jump to step 5); otherwise, jump to step 3). Here, jiffies_to_msecs() is a macro implemented in the kernel to convert jiffies (the kernel's timing unit) to milliseconds (ms).

[0077] 5) Calculate the current pipeline communication data rate: int pipe_speed = pipe->traffic_num / pass_time;

[0078] 6) Smooth the pipeline communication data rate: pipe->pipe_speed = pipe_speed * GF + pipe->pipe_speed * (1 - GF);

[0079] 7) Reset the pipeline communication data volume: pipe->traffic_num = 0;

[0080] 8) Set pipe->pipe_time to the current system time point: pipe->pipe_time = cur_time.

[0081] In the above steps, step 1 and step 2 correspond to step S31, step 3 and step 4 correspond to step S32, and step 5 to step 8 correspond to step S33. By executing one round of the above steps, the communication data rate of a pipeline for a period (1 second in this embodiment) is detected. The detection result of this round can be used for subsequent threshold comparison and as the historical communication data rate for the next round. It should be noted that if the statistical period for detecting the communication data rate of each pipeline is greater than the period for executing one round of the above steps, multiple rounds of the above steps can be executed, and the detection result of the last round is output for subsequent threshold comparison.

[0082] In step S4 of this embodiment, by comparing the communication data rate of the pipeline communication link with the threshold, it is determined whether the process that creates and uses the pipeline communication needs to be migrated to the node node that is the same as the pipeline communication memory. Therefore, when calculating the threshold for each pipeline according to the type and priority of the process, the threshold for the corresponding pipeline of all processes using pipeline communication is calculated, including the following situations:

[0083] If the current process is a specified important process, set the threshold of the pipeline corresponding to the current process to 0. When the threshold is set to 0, it indicates that this is a critical communication pipeline, such as the communication pipeline of a real-time process. Communication efficiency is very important. Therefore, regardless of the pipeline communication rate, the communication process needs to be scheduled to run on the node with the same memory as the pipeline communication to improve the pipeline communication efficiency;

[0084] If the current process is a specified unimportant process, set the threshold of the pipeline corresponding to the current process to the maximum value of the storage type. When the threshold is set to the maximum value of the storage type, it indicates that this is an unimportant communication pipeline and communication efficiency is not important. Therefore, regardless of the pipeline communication rate, there is no need to schedule the communication process to run on the node with the same memory as the pipeline communication;

[0085] If the current process does not belong to the above two cases, the threshold of the corresponding pipeline can be set to a default empirical value or determined by the process priority. When the threshold is determined by the process priority, the higher the process priority, the lower the threshold, and the lower the process priority, the higher the threshold.

[0086] Specifically, step S4 can be implemented through the following steps:

[0087] S41) Define the constant PIPE_THRESHOLD as 100. This constant represents the pipeline communication threshold when the process priority nice is 0, with the unit B / ms. That is, when the pipeline communication rate of a process with a process priority nice of 0 is greater than 100 B / ms, the process needs to be migrated to the node with the same memory as the pipeline communication;

[0088] S42) Obtain the scheduling policy of the current process through current->policy to determine whether the current process is a real-time process. If the scheduling policy of the current process is SCHED_FIFO or SCHED_RR, the scheduling policy of the current process is a real-time scheduling policy. Set the threshold pipe->threshold of the pipeline corresponding to the current process to 0 and then jump to step S44. Otherwise, the current process is a normal process and step S43 is executed. Here, current is a struct task_struct structure variable defined by the kernel to represent the current process, and its member policy sets the type of scheduling policy. SCHED_FIFO and SCHED_RR are real-time scheduling policy constants defined by the kernel;

[0089] S43) If the scheduling policy of the current process is not a real-time scheduling policy, obtain the priority of the current process through task_nice(current), substitute the priority into the threshold calculation function, and obtain the threshold of the pipeline corresponding to the current process. The threshold calculation function is:

[0090] pipe->threshold = PIPE_THRESHOLD * pow(a, task_nice(current))

[0091] Among them, PIPE_THRESHOLD represents the threshold when the process priority is 0. The pow(a, b) function implements the function of calculating the b-th power of a. In this embodiment, a takes 1.25, b is task_nice(current), which is the process priority. task_nice() is a function to obtain the process priority, and current is a struct task_struct structure variable representing the current process;

[0092] S44) Determine whether pipe->pipe_speed is greater than pipe->threshold. If so, execute step S45; otherwise, end step S44 and exit to keep the current process running on its original node node;

[0093] S45) Schedule the current process to run on the node where the pipe communication memory is located. The execution statement is: sched_setaffinity(current->pid, cpumask_of_node(pipe->nid)); where sched_setaffinity() is a function implemented by the kernel to specify the cpu set for a process to run, current->pid is the pid of the current process, and cpumask_of_node() is a function implemented by the kernel to obtain the CPU mask on a specified node node.

[0094] Embodiment 2

[0095] The present invention also proposes a pipeline communication optimization system under the NUMA architecture, including a microprocessor and a computer-readable storage medium connected to each other. The microprocessor is programmed or configured to execute any one of the pipeline communication optimization methods under the NUMA architecture.

[0096] In summary, the present invention discloses a pipeline communication optimization method and system under the NUMA architecture. First, when creating a communication pipeline, allocate the memory for pipeline communication to a specific node node. Secondly, each pipeline communication link in the system needs to implement the detection of its own communication data rate. Finally, when the communication data rate of the pipeline communication link reaches the threshold or it is detected that it is an important process in the system, migrate the process using the pipeline communication to the same node node as the pipeline communication memory to avoid cross-node memory access and improve the efficiency of its pipeline communication. Compared with the prior art, the advantages of the present invention are:

[0097] (1)Autonomous controllability. Since the software optimization and improvement are realized through independent design and research and development, it has complete intellectual property rights.

[0098] (2)The implementation effect is obvious. In view of the characteristics of NUMA architecture hardware, the present invention proposes to detect the communication bandwidth of each pipeline in the system, and at the same time distinguish the importance of system processes. For non-important processes with small pipeline communication bandwidth, cross-node memory communication is allowed, and for processes with large pipeline communication bandwidth or important processes, they are scheduled to communicate on the same node to improve their pipeline communication bandwidth and reduce pipeline communication latency. This method gives full play to the hardware characteristics of the NUMA architecture and improves the system operation efficiency.

[0099] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A pipeline communication optimization method under NUMA architecture, characterized in that: The following steps are involved: When creating a pipeline, allocate pipeline communication memory on the specified node; Detect the communication data rate of each pipeline; The threshold for each pipeline is calculated based on the type and priority of the process, including: Get the scheduling policy of the current process. If the scheduling policy of the current process is not a real-time scheduling policy, get the priority of the current process, substitute the priority into the threshold calculation function, and get the threshold of the pipeline corresponding to the current process. The threshold calculation function is: pipe->threshold = PIPE_THRESHOLD * pow(a, task_nice(current)) Among them, PIPE_THRESHOLD represents the threshold when the process priority is 0, the pow(a, task_nice(current)) function implements the power operation logic of the constant a, task_nice() is a function for obtaining the process priority, and current is a struct task_struct structure variable representing the current process; The communication data rate of each pipeline is compared with the corresponding threshold, and the important processes or processes whose corresponding pipeline communication data rate is greater than the threshold are migrated to the specified node node for execution, and the unimportant processes or processes whose corresponding pipeline communication data rate is less than the threshold are retained in the corresponding original node node for execution. The important processes are processes that need to improve the pipeline communication efficiency, and the unimportant processes are processes that do not need to improve the pipeline communication efficiency.

2. The pipeline communication optimization method under the NUMA architecture according to claim 1, characterized in that: The specified node is specifically a node specified by a user, or a node running by a process for creating a pipeline.

3. The pipeline communication optimization method under the NUMA architecture according to claim 1, characterized in that: When detecting the communication data rate of each pipe, include: Get the historical communication data rate of the current pipeline, and wait for the process corresponding to the current pipeline to write data to the current pipeline or read data from the current pipeline; When a process writes data to or reads data from the current pipe, the read data or write data in the current time period in the current pipe is accumulated and recorded to obtain the communication data volume in the current time period; The communication data volume of the current time period is divided by the duration of the current time period to obtain the current communication data rate of the current pipeline; The current communication data rate is smoothed using the historical communication data rate and a preset amplification factor.

4. The pipeline communication optimization method under the NUMA architecture according to claim 3, characterized in that: When the current communication data rate is smoothed using the historical communication data rate and the preset amplification factor, the expression is as follows: S = Cur_S * GF + Pre_S * (1 - GF) Where S represents the communication data rate after smoothing, Cur_S represents the current communication data rate, Pre_S represents the historical communication data rate, and GF represents the amplification factor, which ranges from (0, 1].

5. The pipeline communication optimization method under NUMA architecture according to claim 3, characterized in that: After waiting for the process corresponding to the current pipe to write data to the current pipe or read data from the current pipe, it also includes: The decay time is determined according to the priority of the process corresponding to the current pipeline. If the waiting time exceeds the decay time, the historical communication data rate is decayed using the preset decay factor to obtain the current communication data rate, and then the process corresponding to the current pipeline continues to write data to the current pipeline or read data from the current pipeline.

6. The pipeline communication optimization method under NUMA architecture according to claim 5, characterized in that: When the preset attenuation factor is used to attenuate the historical communication data rate, the expression is as follows: S = Pre_S *AF Where S represents the communication data rate after attenuation, Pre_S represents the historical communication data rate, and AF represents the attenuation factor, which ranges from (0, 1].

7. The pipeline communication optimization method under NUMA architecture according to claim 1, characterized in that: When calculating the threshold for each pipeline based on the type and priority of the process, it also includes: Get the scheduling policy of the current process; If the scheduling policy of the current process is the real-time scheduling policy, the threshold of the pipeline corresponding to the current process is set to 0.

8. The pipeline communication optimization method under NUMA architecture according to claim 7, characterized in that: Before getting the scheduling policy of the current process, it also includes: If the current process is a designated important process, the threshold of the pipeline corresponding to the current process is set to 0; If the current process is a designated unimportant process, the threshold of the pipeline corresponding to the current process is set to the maximum value of the storage type.

9. A pipeline communication optimization system under NUMA architecture, comprising a microprocessor and a computer-readable storage medium connected to each other, characterized in that: The microprocessor is programmed or configured to execute the pipeline communication optimization method under the NUMA architecture described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • A pipeline communication method and device

    CN103164359B

  • Memory access method, device and system

    CN103365717A

  • Scheduling method and system for Direct IO intensive tasks under NUMA architecture

    CN119473564A