Multi-task scheduling method based on multiple NPUs

By monitoring resource status and assigning task priority in a multi-NPU environment, efficient scheduling of multi-task processing is achieved, solving the problems of low resource utilization and long task waiting time in multi-task processing scenarios, and improving overall performance and resource utilization.

CN119987958APending Publication Date: 2025-05-13西安翔腾微电子科技有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411749908.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In complex multitasking scenarios, how to efficiently and fairly schedule multiple tasks to be executed on multiple neural network processors (NPUs) to maximize resource utilization and reduce task waiting time has become an urgent problem.

Method used

By acquiring and monitoring the resource status of each NPU, NPU resources are allocated reasonably based on task scale, demand core resources and task priorities, ensuring that all NPUs are maximized, and efficient tasks are implemented and efficient resource recycling is achieved.

Benefits of technology

It improves the multi-task processing efficiency and resource utilization in multi-NPU environments, reduces task waiting time and system energy consumption, and provides strong support for building a high-performance, low-latency artificial intelligence application platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987958A_ABST
    Figure CN119987958A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-task scheduling method based on multiple NPUs. The method comprises the following steps: 1) acquiring and monitoring NPU resource states such as a core load condition, a data transmission path state, memory occupation and the like of each NPU; 2) based on an algorithm task scale, demand core resources, a task issuing sequence and the like, the priority is confirmed; 3) according to the current resource state of the NPU and the priority of the algorithm task, allocating related resources for the appropriate task and executing the resources; 4) after task allocation, continuing to allocate proper tasks according to the number of residual resources and the priority, and ensuring that all NPUs can be utilized to the maximum extent; and 5) after the execution of the algorithm tasks is completed, immediately recovering related resources, repeating the step 3) and the step 4), and continuing to execute until all the algorithm tasks are completed. According to the method, NPU resources and task configuration are reasonably allocated, so that the multi-task processing efficiency and the resource utilization rate in a multi-NPU environment are improved, the task waiting time and the system energy consumption are reduced, and powerful support is provided for constructing a high-performance and low-delay artificial intelligence application platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer computing and deep learning hardware acceleration, and in particular to an efficient scheduling method for multiple neural network processors (NPUs) in a multi-tasking environment. Background Art

[0002] With the rapid development of computer vision and artificial intelligence technology, image and video processing technology based on deep learning has been widely used in target detection, image classification, semantic segmentation, scene understanding and other fields. The training and reasoning process of neural network models usually requires a lot of computing resources.

[0003] The neural network processor (NPU) is a processor specifically designed for artificial intelligence (AI) algorithms. It provides powerful computing power to support the efficient operation of various AI algorithms. With the rapid development of AI technology, NPU has been widely used in many fields such as smart cameras, autonomous driving, and smart robots.

[0004] In complex multi-tasking scenarios, multiple NPUs may be required to execute a large number of AI algorithms at the same time. These tasks usually have different running times and execution frequencies, and are difficult to schedule with a single configuration. How to efficiently and fairly schedule multiple tasks to be executed on multiple NPUs to maximize resource utilization and reduce task waiting time has become an urgent problem to be solved. Summary of the invention

[0005] In order to solve the technical problems existing in the background technology, the present invention provides a multi-task scheduling method based on multi-NPU, which improves the multi-task processing efficiency and resource utilization in a multi-NPU environment by reasonably allocating NPU resources and task configuration, reduces task waiting time and system energy consumption, and provides strong support for building a high-performance, low-latency artificial intelligence application platform.

[0006] The technical solution of the present invention is: the present invention is a multi-task scheduling method based on multiple NPUs, and its special feature is that the method comprises the following steps:

[0007] 1) Obtain and monitor the NPU resource status such as the core load, data transmission path status, memory usage, etc. of each NPU;

[0008] 2) Determine the priority based on the algorithm task size, required core resources, and task issuance order;

[0009] 3) Allocate relevant resources to appropriate tasks and execute them according to the current resource status of the NPU and the algorithm task priority;

[0010] 4) After the tasks are assigned, continue to assign appropriate tasks based on the number of remaining resources and priority to ensure that all NPUs can be maximized;

[0011] 5) After the algorithm task is completed, the relevant resources are immediately recovered and steps 3) and 4) are repeated until all algorithm tasks are completed.

[0012] Furthermore, the specific steps of step 1) include:

[0013] 1.1) Obtain the status of all NPU cores, including whether the core is working, whether the core is abnormal, whether the core is bound to a task, and other core load information, and continuously monitor relevant information;

[0014] 1.2) Obtain the data transmission channel status, including whether the data channel is in the transmission state, the number of working channels and idle channels, the remaining amount of data transmitted by the working channels, and continuously monitor relevant information;

[0015] 1.3) Obtain and manage NPU free space.

[0016] Furthermore, the specific steps of step 2) include:

[0017] 2.1) The algorithm scales of different algorithm tasks vary greatly. Depending on the needs, the same algorithm task may be assigned different numbers of core resources for execution. Therefore, it is necessary to briefly evaluate each task and roughly determine its required execution time. Then, determine its priority based on the order in which it is issued and the task level to ensure that the task can be executed in a balanced and efficient manner.

[0018] 2.2) First, tasks with higher levels are prioritized. Second, tasks with the same level are prioritized based on the number of cores required. Tasks with more cores required are prioritized. Third, tasks that are issued first are prioritized.

[0019] Furthermore, the specific steps of step 3) are: according to the current NPU status information obtained in step 1), especially the number of available cores, select the task with the highest algorithm task priority in step 2) and the current core resources and memory resources meet the requirements, allocate corresponding resources to it and execute the task.

[0020] Furthermore, the specific steps of step 4) are: after the task in step 3) is executed, according to the number of idle cores on each NPU, the tasks with higher algorithm priority and the required number of cores are obtained from step 2), and the corresponding resources are allocated to them and the tasks are executed, thereby ensuring that each core on each NPU can be fully utilized, maximizing the computing power of the NPU, and achieving the purpose of improving overall performance.

[0021] Furthermore, the specific steps of step 5) are: according to the current NPU status information obtained in step 1), analyze and judge the completion status of the task. When the NPU completes the calculation of the algorithm task, the result is moved to the specified address through the transmission path, and then the core occupancy of the NPU, memory space usage and other information are updated, all related resources are recovered, and then steps 3) and 4) are repeated until all tasks are completed, and then wait for the next round of tasks to arrive.

[0022] The present invention provides a multi-task scheduling method based on multiple NPUs, which improves the multi-task processing efficiency and resource utilization in a multi-NPU environment by reasonably allocating NPU resources and task configuration, reduces task waiting time and system energy consumption, and provides strong support for building a high-performance, low-latency artificial intelligence application platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION

[0024] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] See also Figure 1 , the method steps of the specific embodiment of the present invention are as follows:

[0026] 1) Obtain and monitor the NPU resource status such as the core load, data transmission path status, memory usage, etc. of each NPU;

[0027] 1.1) Obtain the status of all NPU cores, including whether the core is working, whether the core is abnormal, whether the core is bound to a task, and other core load information, and continuously monitor relevant information;

[0028] 1.2) Obtain the data transmission channel status, including whether the data channel is in the transmission state, the number of working channels and idle channels, the remaining amount of data transmitted by the working channels, and continuously monitor relevant information;

[0029] 1.3) Obtain and manage NPU free space.

[0030] 2) Determine the priority based on the algorithm task size, required core resources, and task issuance order;

[0031] 2.1) The algorithm scales of different algorithm tasks vary greatly. Depending on the needs, the same algorithm task may be assigned different numbers of core resources for execution. Therefore, it is necessary to briefly evaluate each task and roughly determine its required execution time. Then, determine its priority based on the order in which it is issued and the task level to ensure that the task can be executed in a balanced and efficient manner.

[0032] 2.2) First, tasks with higher levels are prioritized. Second, tasks with the same level are prioritized based on the number of cores required. Tasks with more cores required are prioritized. Third, tasks that are issued first are prioritized.

[0033] 3) Allocate relevant resources to appropriate tasks and execute them according to the current resource status of the NPU and the algorithm task priority; specifically:

[0034] According to the current status information of each NPU obtained in step 1), especially the number of available cores, select the task with the highest algorithm task priority in step 2) and the current core resources and memory resources that meet the requirements, allocate corresponding resources to it and execute the task.

[0035] 4) After the tasks are assigned, continue to assign appropriate tasks based on the number of remaining resources and priority to ensure that all NPUs can be maximized; specifically:

[0036] After the task is executed in step 3), according to the number of idle cores on each NPU, the tasks with higher algorithm priority and the required number of cores are obtained from step 2), and the corresponding resources are allocated to them and the tasks are executed. In this way, it is ensured that each core on each NPU can be fully utilized, maximizing the computing power of the NPU and achieving the purpose of improving overall performance.

[0037] 5) After the algorithm task is completed, immediately recycle the relevant resources and repeat steps 3) and 4) until all algorithm tasks are completed.

[0038] According to the current status information of each NPU obtained in step 1), the completion status of the task is analyzed and judged. When the NPU completes the calculation of the algorithm task, the result is moved to the specified address through the transmission path, and then the core occupancy of the NPU, memory space usage and other information are updated, all related resources are recovered, and then steps 3) and 4) are repeated until all tasks are completed, and then wait for the next round of tasks to arrive.

[0039] The above are only specific embodiments disclosed in the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.

[0040] The content of the present invention and the technical content not specifically described in the above embodiments are the same as the prior art.

[0041] The present invention is not limited to the above embodiments, and all of the contents of the present invention can be implemented and have the above good effects.

Claims

1. A multi-task scheduling method based on multiple NPUs, characterized in that: The method comprises the following steps: 1) Obtain and monitor the NPU resource status such as the core load, data transmission path status, memory usage, etc. of each NPU; 2) Determine the priority based on the algorithm task size, required core resources, and task issuance order; 3) Allocate relevant resources to appropriate tasks and execute them according to the current resource status of the NPU and the algorithm task priority; 4) After the tasks are assigned, continue to assign appropriate tasks based on the number of remaining resources and priority to ensure that all NPUs can be maximized; 5) After the algorithm task is completed, the relevant resources are immediately recovered and steps 3) and 4) are repeated until all algorithm tasks are completed.

2. The multi-task scheduling method based on multiple NPUs according to claim 1, characterized in that: The specific steps of step 1) include: 1).1 Obtain the status of all NPU cores, including whether the core is working, whether the core is abnormal, whether the core is bound to a task, and other core load information, and continuously monitor relevant information; 1).2 Obtain the data transmission channel status, including whether the data channel is in the transmission state, the number of working channels and idle channels, the remaining amount of data transmitted by the working channels, and continuously monitor relevant information; 1).3 Obtain and manage NPU free space.

3. The multi-task scheduling method based on multiple NPUs according to claim 2, characterized in that: The specific steps of step 2) include: 2) The algorithm scales of different algorithm tasks vary greatly. Depending on the needs, the same algorithm task may be assigned different numbers of core resources for execution. Therefore, it is necessary to briefly evaluate each task and roughly determine its required execution time. Then, determine its priority according to the order in which it is issued and the task level to ensure that the task can be executed in a balanced and efficient manner. 2).2First, the tasks with higher levels are prioritized. Secondly, the tasks with the same level are prioritized by the number of cores required. The tasks with more cores required are prioritized. Thirdly, the tasks that are issued first are prioritized.

4. The multi-task scheduling method based on multiple NPUs according to claim 3 is characterized in that: The specific steps of step 3) are: according to the current NPU status information obtained in step 1), especially the number of available cores, select the task with the highest algorithm task priority in step 2) and the current core resources and memory resources meet the requirements, allocate corresponding resources to it and execute the task.

5. The multi-task scheduling method based on multiple NPUs according to claim 4 is characterized in that: The specific steps of step 4) are: after the task in step 3) is executed, according to the number of idle cores on each NPU, the tasks with higher algorithm priority and the required number of cores are obtained from step 2), and corresponding resources are allocated to them and the tasks are executed, thereby ensuring that each core on each NPU can be fully utilized, maximizing the computing power of the NPU, and achieving the purpose of improving overall performance.

6. The multi-task scheduling method based on multiple NPUs according to claim 5, characterized in that: The specific steps of step 5) are: according to the current NPU status information obtained in step 1), analyze and judge the completion status of the task. When the NPU completes the calculation of the algorithm task, the result is moved to the specified address through the transmission path, and then the core occupancy of the NPU, the memory space usage and other information are updated, all related resources are recovered, and then steps 3) and 4) are repeated until all tasks are completed, and then wait for the next round of tasks to arrive.

Citation Information

Cited By

  • Scheduling method and device of NPU computing task, artificial intelligence equipment and medium

    CN120492133A

  • Dynamic data management method and system based on AI

    CN120541115A