Multi-task parallel computing method and system for NUMA (Non Uniform Memory Access) perception
Through the NUMA-aware multi-task parallel computing method, task execution is reasonably scheduled, which solves the cross-node memory access and nested task locality problems of multi-task parallel programs under the NUMA architecture, and achieves low-overhead, efficient computing performance and programming transparency.
Patent Information
- Application Number
- CN202510865395.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-28
AI Technical Summary
In the existing technology, multi-task parallel programs lack NUMA awareness under the NUMA architecture, resulting in high cross-node memory access overhead and lack of locality of nested tasks. Existing optimization solutions have problems of high overhead and programming complexity.
A NUMA-aware multi-task parallel computing method is adopted. NUMA node numbers are recorded through thread binding, task queues, and thread local storage to reasonably schedule task execution, ensure that nested tasks are executed on the same node, and reduce remote memory access.
It achieves low-overhead, automated task scheduling, reduces remote memory access frequency, improves computing performance, adapts to large-scale nested task scenarios, and has high programming transparency.
Smart Images

Figure CN120849095A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of high-performance computing technology, specifically relating to a NUMA-aware multi-task parallel computing method and system. Background Technology
[0002] Modern servers commonly employ NUMA (Non-Uniform Memory Access) architecture, especially in multi-core, high-performance computing. NUMA architecture enhances server scalability through distributed memory design, but its non-uniform memory access characteristics significantly increase cross-node access latency. Although NUMA-based servers have increasingly more CPUs (Center Processing Units), such as a certain advanced domestic ARM server with 8 NUMA nodes and 72 cores per node, theoretically supporting 576 parallel tasks, existing software cannot fully utilize the machine's performance; in fact, performance may even decrease as the number of parallel tasks increases. This is mainly because in existing technologies, multi-task parallel programs often lack NUMA awareness in thread scheduling and memory allocation, leading to the following problems: 1) High cross-node memory access overhead: different tasks in the computation process may be assigned to different NUMA nodes, causing tasks to frequently access remote memory; 2) Nested tasks lack locality guarantees: existing task scheduling methods do not consider the NUMA correlation between task levels, making it difficult to ensure that nested tasks are executed on the same node.
[0003] Existing optimization solutions, such as dynamic thread binding and thread migration based on statistics, can optimize locality, but still have the following drawbacks: 1) Dynamic adjustment is costly, and frequent migration of threads or data increases additional computation and synchronization costs; 2) They cannot adapt to nested tasks, and hierarchical tasks require explicit specification of nodes, increasing programming complexity.
[0004] Therefore, those skilled in the art are dedicated to developing a NUMA-aware multi-task parallel computing method and system. Summary of the Invention
[0005] Therefore, in order to address at least one of the above-mentioned defects or improvement needs of the existing technology and to solve the performance degradation problem caused by cross-node memory access in multi-task under NUMA architecture, this invention proposes a NUMA-aware multi-task parallel computing method and system. It provides a low-overhead, automated task scheduling method that rationally schedules the task execution process according to the task dependencies, ensuring that multi-level nested tasks with the same data dependencies are executed on the same NUMA node, guaranteeing the locality of memory access, and reducing remote memory access.
[0006] This invention discloses a NUMA-aware multi-task parallel computing method, characterized in that the method includes the following steps:
[0007] S1: Based on the number and ID of available CPU cores specified by the user, create an equal number of threads and bind the threads to the corresponding CPU cores;
[0008] S2: Based on NUMA topology information, threads within the same NUMA node are divided into execution groups, and a task queue is created and associated with each execution group;
[0009] S3: Users submit the tasks involved in the calculation process to the task queue. When a top-level task is created for the first time, a round-robin strategy is used to allocate it to the task queue of each execution group. When a top-level task is created dynamically thereafter, a least recently used strategy is used to allocate it to the task queue.
[0010] S4: After the thread obtains a task from the corresponding queue and starts execution, it obtains the NUMA node number according to the CPU core in which the thread is running, and records the NUMA node number through the thread's local storage.
[0011] S5: If the current task creates a subtask, the subtask automatically obtains the NUMA node number of the parent task through the thread local storage and submits it to the task queue of the corresponding node.
[0012] S6: If a subtask continues to create the next level of subtasks, repeat steps S4-S5 until all tasks are completed;
[0013] Furthermore, the step of dividing threads within the same NUMA node into execution groups specifically involves dividing threads corresponding to CPU cores of the same NUMA node into execution groups according to the CPU core allocation of the NUMA node, with each execution group corresponding to one NUMA node.
[0014] Furthermore, the polling strategy specifically involves taking the modulo of the top-level task number with the number of execution groups to obtain the task queue number, and then submitting the top-level task to the corresponding task queue.
[0015] Furthermore, the least recently used strategy specifically means: prioritizing the allocation of dynamically created top-level tasks to the least recently used execution group task queue;
[0016] Furthermore, the task queue adopts a lock-free design with atomic operations, and the task queue supports a multi-producer, multi-consumer mode.
[0017] Furthermore, the computation process includes single-level tasks and multi-level nested tasks. Each subtask of a multi-level nested task inherits the NUMA node number of the parent task through the thread-local storage, ensuring that the parent and child tasks are executed on the same node.
[0018] This invention also discloses a NUMA-aware multi-task parallel computing system, comprising:
[0019] The execution group module is used to create an equal number of threads and bind them to the corresponding cores based on the number and ID of available CPU cores specified by the user. Based on the NUMA topology, threads within the same node are divided into execution groups.
[0020] Task queue module: Creates a lock-free task queue for each execution group;
[0021] Task scheduling module: Configured to use a round-robin strategy to allocate top-level tasks to the task queue for newly created tasks; and to use a least recently used strategy to allocate top-level tasks to the task queue for dynamically created tasks.
[0022] Node number recording module: used to record the NUMA node number to which the thread belongs through thread-local storage when the thread is executing a task;
[0023] Subtask scheduling module: When a task creates a subtask, it is configured to obtain the NUMA node number of the parent task through the thread local storage and submit the subtask to the task queue of the corresponding node;
[0024] Furthermore, the execution group module fixes the thread on the CPU core of the corresponding NUMA node through a pro-core setting;
[0025] Furthermore, the task queue created by the task queue module adopts a lock-free design with atomic operations;
[0026] Furthermore, the polling strategy specifically involves taking the modulo of the top-level task number with the number of execution groups to obtain the task queue number, and then submitting the top-level task to the corresponding task queue; the least recently used strategy specifically involves prioritizing the allocation of dynamically created top-level tasks to the least recently used execution group task queue.
[0027] In summary, compared with the prior art, the above-described technical solutions conceived in this invention can achieve the following beneficial effects:
[0028] 1) Reduce cross-node memory access: Through the TLS inheritance mechanism, nested tasks are ensured to be executed on the same NUMA node, reducing the frequency of remote memory access;
[0029] 2) Low-overhead scheduling: Combining static grouping with dynamic load balancing avoids frequent thread migration;
[0030] 3) Programming transparency: Users do not need to explicitly specify task nodes; the system automatically maintains NUMA locality.
[0031] 4) High scalability: The lock-free queue design and load balancing strategy are adapted to large-scale nested task scenarios. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of the method architecture of the present invention;
[0033] Figure 2 This is a system execution flowchart of the present invention;
[0034] Figure 3 This is a schematic diagram of the system structure of the present invention. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0036] like Figure 1 As shown in the figure, the present invention proposes a NUMA-aware multi-task parallel computing method, which includes the following steps:
[0037] S1: Based on the number and ID of available CPU cores specified by the user, create an equal number of threads and bind the threads to the corresponding CPU cores;
[0038] S2: Based on NUMA topology information, threads within the same NUMA node are divided into execution groups, and a task queue is created and associated with each execution group;
[0039] S3: Users submit the tasks involved in the calculation process to the task queue. When a top-level task is created for the first time, a round-robin strategy is used to allocate it to the task queue of each execution group. When a top-level task is created dynamically thereafter, a least recently used strategy is used to allocate it to the task queue.
[0040] S4: After the thread obtains a task from the corresponding queue and starts execution, it obtains the NUMA node number according to the CPU core in which the thread is running, and records the NUMA node number through the thread's local storage.
[0041] S5: If the current task creates a subtask, the subtask automatically obtains the NUMA node number of the parent task through the thread local storage and submits it to the task queue of the corresponding node.
[0042] S6: If a subtask continues to create the next level of subtasks, repeat steps S4-S5 until all tasks are completed;
[0043] In one embodiment, such as Figure 3 As shown, the execution flow of a NUMA-aware multi-task parallel computing method disclosed in this invention is as follows:
[0044] Step S101: Initialize threads and bind to cores. The user specifies the available CPU cores (e.g., cores 0-7), and the system creates the same number (8) threads and binds them to the CPU cores.
[0045] Step S102: Create task queues. Based on NUMA topology information (e.g., node 0 contains cores 0-3, node 1 contains cores 4-7), group the tasks into several (2) execution groups. Create and associate each group with a task queue. Threads continuously retrieve tasks from their respective queues for execution. The task queues employ a lock-free design with atomic operations, supporting multiple producers and multiple consumers.
[0046] Step S103: The user executes the computation process, which is completed in parallel through a series of fine-grained computation tasks. These tasks may be single-layered or multi-layered nested.
[0047] Step S104: Create and deliver top-level tasks, and use a round-robin strategy to allocate them to the task queues of each execution group. For example, create 3 top-level tasks with task numbers 0 to 2, and use the task number modulo the number of execution groups to obtain the task queue number for delivery.
[0048] Step 105: The thread retrieves a task from the corresponding task queue. It is blocked when there is no task and is woken up to execute when there is a task.
[0049] Step 106: Execute the task and record the NUMA node number. After execution begins, obtain the NUMA node number of the CPU core in which the thread is running, and record it to the current thread via TLS.
[0050] Step 107: Create and submit subtasks. If the subtask is a top-level task, it is assigned to the task queue using the least recently used strategy. If the subtask is not a top-level task, the NUMA node number of the parent task is automatically obtained via TLS, and it is submitted to the corresponding node's task queue.
[0051] Steps 105-107 are executed repeatedly until all tasks are completed.
[0052] In a preferred embodiment, the present invention also provides a NUMA-aware multi-task parallel computing system, such as... Figure 3 As shown, the system includes:
[0053] Server: A server using a NUMA architecture, with multiple NUMA nodes, each with multiple CPU cores.
[0054] The execution group module is used to manage threads and core affinity settings. It creates an equal number of threads and binds them to the corresponding cores based on the number and ID of available CPU cores specified by the user. Based on NUMA topology, it divides threads within the same node into execution groups.
[0055] Task queue module: Creates a lock-free queue with atomic operations for each execution group, supports multiple producers and multiple consumers, the computation process submits tasks to the queue, and the execution group retrieves tasks from the queue;
[0056] Task scheduling module: Uses different scheduling strategies to deliver computing tasks to appropriate queues, ensuring that related tasks run on the same NUMA node, avoiding cross-node memory access, and ensuring system load balancing. Specifically, for newly created top-level tasks, a round-robin strategy is used to allocate them to the task queue; for dynamically created top-level tasks, a least recently used strategy is used to allocate them to the task queue.
[0057] Node number recording module: used to record the NUMA node number to which the thread belongs through thread-local storage when the thread is executing a task;
[0058] Subtask scheduling module: When a task creates a subtask, it is configured to obtain the NUMA node number of the parent task through the thread local storage and submit the subtask to the task queue of the corresponding node.
[0059] The description in this specification is merely illustrative of the invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the content of this specification or exceed the scope defined in the claims, they should all fall within the protection scope of this invention.
Claims
1. A NUMA-aware multi-task parallel computing method, characterized in that, The above method includes the following steps: S1: Based on the number and ID of available CPU cores specified by the user, create an equal number of threads and bind the threads to the corresponding CPU cores; S2: Based on NUMA topology information, threads within the same NUMA node are divided into execution groups, and a task queue is created and associated with each execution group; S3: Users submit the tasks involved in the calculation process to the task queue. When a top-level task is created for the first time, a round-robin strategy is used to allocate it to the task queue of each execution group. When a top-level task is created dynamically thereafter, a least recently used strategy is used to allocate it to the task queue. S4: After the thread obtains a task from the corresponding queue and starts execution, it obtains the NUMA node number according to the CPU core in which the thread is running, and records the NUMA node number through the thread's local storage. S5: If the current task creates a subtask, the subtask automatically obtains the NUMA node number of the parent task through the thread local storage and submits it to the task queue of the corresponding node. S6: If a subtask continues to create the next level of subtasks, repeat steps S4-S5 until all tasks are completed.
2. The NUMA-aware multi-task parallel computing method according to claim 1, characterized in that, The process of dividing threads within the same NUMA node into execution groups specifically involves dividing threads belonging to the same NUMA node into execution groups based on the CPU core allocation of the NUMA node, with each execution group corresponding to one NUMA node.
3. The NUMA-aware multi-task parallel computing method according to claim 1, characterized in that, The polling strategy is as follows: take the modulo of the top-level task number with the number of execution groups to obtain the task queue number, and then submit the top-level task to the corresponding task queue.
4. The NUMA-aware multi-task parallel computing method according to claim 1, characterized in that, The Least Recently Used strategy specifically means that dynamically created top-level tasks are preferentially assigned to the least recently used execution group task queue.
5. The NUMA-aware multi-task parallel computing method according to claim 1, characterized in that, The task queue adopts a lock-free design with atomic operations and supports a multi-producer, multi-consumer model.
6. The NUMA-aware multi-task parallel computing method according to claim 1, characterized in that, The computation process includes single-level tasks and multi-level nested tasks. Each subtask of a multi-level nested task inherits the NUMA node number of the parent task through the thread-local storage, ensuring that the parent and child tasks are executed on the same node.
7. A NUMA-aware multi-task parallel computing system, characterized in that, include: The execution group module is used to create an equal number of threads and bind them to the corresponding cores based on the number and ID of available CPU cores specified by the user. Based on the NUMA topology, threads within the same node are divided into execution groups. Task queue module: Creates a lock-free task queue for each execution group; Task scheduling module: Configured to use a round-robin strategy to allocate top-level tasks to the task queue for newly created tasks; and to use a least recently used strategy to allocate top-level tasks to the task queue for dynamically created tasks. Node number recording module: used to record the NUMA node number to which the thread belongs through thread-local storage when the thread is executing a task; Subtask scheduling module: When a task creates a subtask, it is configured to obtain the NUMA node number of the parent task through the thread local storage and submit the subtask to the task queue of the corresponding node.
8. The NUMA-aware multi-task parallel computing system as described in claim 7, characterized in that, The execution group module fixes the thread on the CPU core of the corresponding NUMA node through a pro-core setting.
9. The NUMA-aware multi-task parallel computing system as described in claim 7, characterized in that, The task queue created by the task queue module adopts a lock-free design with atomic operations.
10. The NUMA-aware multi-task parallel computing system as described in claim 7, characterized in that, The polling strategy is as follows: take the modulo of the top-level task number with the number of execution groups to obtain the task queue number, and then submit the top-level task to the corresponding task queue. The Least Recently Used strategy specifically means that dynamically created top-level tasks are preferentially assigned to the least recently used execution group task queue.