Multi-task scheduling method based on multiple NPUs
By monitoring resource status and assigning task priority in a multi-NPU environment, efficient scheduling of multi-task processing is achieved, solving the problems of low resource utilization and long task waiting time in multi-task processing scenarios, and improving overall performance and resource utilization.
Patent Information
- Application Number
- CN202411749908.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-13
AI Technical Summary
In complex multitasking scenarios, how to efficiently and fairly schedule multiple tasks to be executed on multiple neural network processors (NPUs) to maximize resource utilization and reduce task waiting time has become an urgent problem.
By acquiring and monitoring the resource status of each NPU, NPU resources are allocated reasonably based on task scale, demand core resources and task priorities, ensuring that all NPUs are maximized, and efficient tasks are implemented and efficient resource recycling is achieved.
It improves the multi-task processing efficiency and resource utilization in multi-NPU environments, reduces task waiting time and system energy consumption, and provides strong support for building a high-performance, low-latency artificial intelligence application platform.
Smart Images

Figure CN119987958A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer computing and deep learning hardware acceleration, and in particular to an efficient scheduling method for multiple neural network processors (NPUs) in a multi-tasking environment. Background Art
[0002] With the rapid development of computer vision and artificial intelligence technology, image and video processing technology based on deep learning has been widely used in target detection, image classification, semantic segmentation, scene understanding and other fields. The training and reasoning process of neural network models usually requires a lot of computing resources.
[0003] The neural network processor (NPU) is a processor specifically designed for artificial intelligence (AI) algorithms. It provides powerful computing power to support the efficient operation of various AI algorithms. With the rapid development of AI technology, NPU has been widely used in many fields such as smart cameras, autonomous driving, and smart robots.
[0004] In complex multi-tasking scenarios, multiple NPUs may be required to execute a large number of AI algorithms at the same time. These tasks usually have different running times and execution frequencies, and are difficult to schedule with a single configuration. How to efficiently and fairly schedule multiple tasks to be executed on multiple NPUs to maximize resource utilization and reduce task waiting time has become an urgent problem to be solved. Summary of the invention
[0005] In order to solve the technical problems existing in the background technology, the present invention provides a multi-task scheduling method based on multi-NPU, which improves the multi-task processing efficiency and resource utilization in a multi-NPU environment by reasonably allocating NPU resources and task configuration, reduces task waiting time and system energy consumption, and provides strong support for building a high-performance, low-latency artificial intelligence application platform.
[0006] The technical solution of the present invention is: the present invention is a multi-task scheduling method based on multiple NPUs, and its special feature is that the method comprises the following steps:
[0007] 1) Obtain and monitor the NPU resource status such as the core load, data transmission path status, memory usage, etc. of each NPU;
[0008] 2) Determine the priority based on the algorithm task size, required core resources, and task issuance order;
[0009] 3) Allocate relevant resources to appropriate tasks and execute them according to the current resource status of the NPU and the algorithm task priority;
[0010] 4) After the tasks are assigned, continue to assign appropriate tasks based on the number of remaining resources and priority to ensure that all NPUs can be maximized;
[0011] 5) After the algorithm task is completed, the relevant resources are immediately recovered and steps 3) and 4) are repeated until all algorithm tasks are completed.
[0012] Furthermore, the specific steps of step 1) include:
[0013] 1.1) Obtain the status of all NPU cores, including whether the core is working, whether the core is abnormal, whether the core is bound to a task, and other core load information, and continuously monitor relevant information;
[0014] 1.2) Obtain the data transmission channel status, including whether the data channel is in the transmission state, the number of working channels and idle channels, the remaining amount of data transmitted by the working channels, and continuously monitor relevant information;
[0015] 1.3) Obtain and manage NPU free space.
[0016] Furthermore, the specific steps of step 2) include:
[0017] 2.1) The algorithm scales of different algorithm tasks vary greatly. Depending on the needs, the same algorithm task may be assigned different numbers of core resources for execution. Therefore, it is necessary to briefly evaluate each task and roughly determine its required execution time. Then, determine its priority based on the order in which it is issued and the task level to ensure that the task can be executed in a balanced and efficient manner.
[0018] 2.2) First, tasks with higher levels are prioritized. Second, tasks with the same level are prioritized based on the number of cores required. Tasks with more cores required are prioritized. Third, tasks that are issued first are prioritized.
[0019] Furthermore, the specific steps of step 3) are: according to the current NPU status information obtained in step 1), especially the number of available cores, select the task with the highest algorithm task priority in step 2) and the current core resources and memory resources meet the requirements, allocate corresponding resources to it and execute the task.
[0020] Furthermore, the specific steps of step 4) are: after the task in step 3) is executed, according to the number of idle cores on each NPU, the tasks with higher algorithm priority and the required number of cores are obtained from step 2), and the corresponding resources are allocated to them and the tasks are executed, thereby ensuring that each core on each NPU can be fully utilized, maximizing the computing power of the NPU, and achieving the purpose of improving overall performance.
[0021] Furthermore, the specific steps of step 5) are: according to the current NPU status information obtained in step 1), analyze and judge the completion status of the task. When the NPU completes the calculation of the algorithm task, the result is moved to the specified address through the transmission path, and then the core occupancy of the NPU, memory space usage and other information are updated, all related resources are recovered, and then steps 3) and 4) are repeated until all tasks are completed, and then wait for the next round of tasks to arrive.
[0022] The present invention provides a multi-task scheduling method based on multiple NPUs, which improves the multi-task processing efficiency and resource utilization in a multi-NPU environment by reasonably allocating NPU resources and task configuration, reduces task waiting time and system energy consumption, and provides strong support for building a high-performance, low-latency artificial intelligence application platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION
[0024] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0025] See also Figure 1 , the method steps of the specific embodiment of the present invention are as follows:
[0026] 1) Obtain and monitor the NPU resource status such as the core load, data transmission path status, memory usage, etc. of each NPU;
[0027] 1.1) Obtain the status of all NPU cores, including whether the core is working, whether the core is abnormal, whether the core is bound to a task, and other core load information, and continuously monitor relevant information;
[0028] 1.2) Obtain the data transmission channel status, including whether the data channel is in the transmission state, the number of working channels and idle channels, the remaining amount of data transmitted by the working channels, and continuously monitor relevant information;
[0029] 1.3) Obtain and manage NPU free space.
[0030] 2) Determine the priority based on the algorithm task size, required core resources, and task issuance order;
[0031] 2.1) The algorithm scales of different algorithm tasks vary greatly. Depending on the needs, the same algorithm task may be assigned different numbers of core resources for execution. Therefore, it is necessary to briefly evaluate each task and roughly determine its required execution time. Then, determine its priority based on the order in which it is issued and the task level to ensure that the task can be executed in a balanced and efficient manner.
[0032] 2.2) First, tasks with higher levels are prioritized. Second, tasks with the same level are prioritized based on the number of cores required. Tasks with more cores required are prioritized. Third, tasks that are issued first are prioritized.
[0033] 3) Allocate relevant resources to appropriate tasks and execute them according to the current resource status of the NPU and the algorithm task priority; specifically:
[0034] According to the current status information of each NPU obtained in step 1), especially the number of available cores, select the task with the highest algorithm task priority in step 2) and the current core resources and memory resources that meet the requirements, allocate corresponding resources to it and execute the task.
[0035] 4) After the tasks are assigned, continue to assign appropriate tasks based on the number of remaining resources and priority to ensure that all NPUs can be maximized; specifically:
[0036] After the task is executed in step 3), according to the number of idle cores on each NPU, the tasks with higher algorithm priority and the required number of cores are obtained from step 2), and the corresponding resources are allocated to them and the tasks are executed. In this way, it is ensured that each core on each NPU can be fully utilized, maximizing the computing power of the NPU and achieving the purpose of improving overall performance.
[0037] 5) After the algorithm task is completed, immediately recycle the relevant resources and repeat steps 3) and 4) until all algorithm tasks are completed.
[0038] According to the current status information of each NPU obtained in step 1), the completion status of the task is analyzed and judged. When the NPU completes the calculation of the algorithm task, the result is moved to the specified address through the transmission path, and then the core occupancy of the NPU, memory space usage and other information are updated, all related resources are recovered, and then steps 3) and 4) are repeated until all tasks are completed, and then wait for the next round of tasks to arrive.
[0039] The above are only specific embodiments disclosed in the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.
[0040] The content of the present invention and the technical content not specifically described in the above embodiments are the same as the prior art.
[0041] The present invention is not limited to the above embodiments, and all of the contents of the present invention can be implemented and have the above good effects.
Claims
1. A multi-task scheduling method based on multiple NPUs, characterized in that: The method comprises the following steps: 1) Obtain and monitor the NPU resource status such as the core load, data transmission path status, memory usage, etc. of each NPU; 2) Determine the priority based on the algorithm task size, required core resources, and task issuance order; 3) Allocate relevant resources to appropriate tasks and execute them according to the current resource status of the NPU and the algorithm task priority; 4) After the tasks are assigned, continue to assign appropriate tasks based on the number of remaining resources and priority to ensure that all NPUs can be maximized; 5) After the algorithm task is completed, the relevant resources are immediately recovered and steps 3) and 4) are repeated until all algorithm tasks are completed.
2. The multi-task scheduling method based on multiple NPUs according to claim 1, characterized in that: The specific steps of step 1) include: 1).1 Obtain the status of all NPU cores, including whether the core is working, whether the core is abnormal, whether the core is bound to a task, and other core load information, and continuously monitor relevant information; 1).2 Obtain the data transmission channel status, including whether the data channel is in the transmission state, the number of working channels and idle channels, the remaining amount of data transmitted by the working channels, and continuously monitor relevant information; 1).3 Obtain and manage NPU free space.
3. The multi-task scheduling method based on multiple NPUs according to claim 2, characterized in that: The specific steps of step 2) include: 2) The algorithm scales of different algorithm tasks vary greatly. Depending on the needs, the same algorithm task may be assigned different numbers of core resources for execution. Therefore, it is necessary to briefly evaluate each task and roughly determine its required execution time. Then, determine its priority according to the order in which it is issued and the task level to ensure that the task can be executed in a balanced and efficient manner. 2).2First, the tasks with higher levels are prioritized. Secondly, the tasks with the same level are prioritized by the number of cores required. The tasks with more cores required are prioritized. Thirdly, the tasks that are issued first are prioritized.
4. The multi-task scheduling method based on multiple NPUs according to claim 3 is characterized in that: The specific steps of step 3) are: according to the current NPU status information obtained in step 1), especially the number of available cores, select the task with the highest algorithm task priority in step 2) and the current core resources and memory resources meet the requirements, allocate corresponding resources to it and execute the task.
5. The multi-task scheduling method based on multiple NPUs according to claim 4 is characterized in that: The specific steps of step 4) are: after the task in step 3) is executed, according to the number of idle cores on each NPU, the tasks with higher algorithm priority and the required number of cores are obtained from step 2), and corresponding resources are allocated to them and the tasks are executed, thereby ensuring that each core on each NPU can be fully utilized, maximizing the computing power of the NPU, and achieving the purpose of improving overall performance.
6. The multi-task scheduling method based on multiple NPUs according to claim 5, characterized in that: The specific steps of step 5) are: according to the current NPU status information obtained in step 1), analyze and judge the completion status of the task. When the NPU completes the calculation of the algorithm task, the result is moved to the specified address through the transmission path, and then the core occupancy of the NPU, the memory space usage and other information are updated, all related resources are recovered, and then steps 3) and 4) are repeated until all tasks are completed, and then wait for the next round of tasks to arrive.
Citation Information
Cited By
Scheduling method and device of NPU computing task, artificial intelligence equipment and medium
CN120492133A
Dynamic data management method and system based on AI
CN120541115A