A Real-time Inference System and Scheduling Method with Time-sharing Multiplexing of NPU

By adopting the real-time inference system and scheduling method of NPU time-sharing multiplexing on edge intelligent computing service devices, fine-grained model segmentation and improved scheduling methods are realized, and the real-time and efficiency of NPU inference are solved, which solves the problem of difficult to guarantee the real-time nature of multi-model inference tasks in the prior art, and improves the real-time and efficiency of NPU inference.

CN119668811BActive Publication Date: 2025-06-27HARBIN INST OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411809599.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-06-27
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

The prior art is difficult to realize the real-time nature of multi-model inference tasks on edge intelligent computing service devices, especially in resource-constrained environments. The scheduling granularity of the existing NPU time-sharing multiplexing scheme is too coarse, making it difficult to ensure the real-time nature of the inference tasks.

Method used

A real-time inference system and scheduling method for NPU time-sharing multiplexing is adopted. Through a heterogeneous computing system and a multi-model serial real-time inference controller, fine-grained model segmentation and improved scheduling methods are realized, and the time-sharing multiplexing efficiency of NPU resources is improved. The specific steps include: the prepared pre-segmenter segments the intelligent model, the runtime planner divides the task into the inference request, and the runtime executor calls the corresponding model blocks to infer according to the task block.

Benefits of technology

Through fine-grained model segmentation and improved scheduling methods, the real-time NPU inference on edge intelligent computing service devices is improved, and the inference tasks can be effectively executed in scenarios where different task cycles vary greatly, reducing the risk that the real-time nature of periodic tasks is affected by burst tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119668811B_ABST
    Figure CN119668811B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time inference system and scheduling method for NPU time-sharing multiplexing, belonging to the field of edge intelligent computing technology, and solves the problem that traditional edge intelligent computing service devices and inference scheduling methods in the prior art are difficult to ensure the real-time performance of multi-inference tasks; in the preprocessing stage, the present invention converts the intelligent inference model through a ready-state pre-segmenter, that is, divides the intelligent inference model into chunks of different granularities in combination with the unit granularity, and obtains the running attribute information of the model and its chunks; in the execution stage, the runtime planner receives the task request of the remote procedure call, then determines the optimal scheduling granularity through nonlinear optimization, and generates the corresponding task scheduling sequence according to the NPU real-time scheduling algorithm with low segmentation, and the runtime executor obtains the model chunks required by the task according to the job sequence and executes them. The present invention effectively improves the real-time performance of time-sharing multiplexing NPU resources for edge intelligent inference computing tasks in a multi-task scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a real-time inference system and a scheduling method, and in particular to a real-time inference system and a scheduling method for time-sharing multiplexing of an NPU, belonging to the field of edge intelligent computing technology. Background Art

[0002] Edge intelligence is a new computing paradigm that directly deploys artificial intelligence technology to the source or near the source of data generation, and can provide intelligent computing services with fast response for numerous terminal devices, significantly reducing the latency of intelligent computing. However, edge computing devices are usually used in resource-constrained environments. Therefore, the neural processing unit (NPU) has gradually become a popular choice in the field of edge intelligence due to its efficient intelligent processing ability and low power consumption characteristics. However, edge intelligent computing service devices based on NPU often need to process multiple time-sensitive task requests simultaneously. For example, in application scenarios such as unmanned equipment and autonomous driving, the device needs to concurrently process multiple real-time model inference tasks such as target recognition, automatic obstacle avoidance, and path planning. Due to resource limitations in the edge computing scenario, the concurrent execution of the above real-time inference tasks usually needs to be completed on a single NPU, posing challenges to the system resource management and scheduling capabilities.

[0003] Existing mainstream solutions solve the problem of real-time concurrent execution of model inference through technologies such as hardware virtualization and time-sharing multiplexing. However, the application of existing mainstream solutions on edge intelligent computing service devices is limited. This is because: ① The hardware virtualization solution requires the support of specific hardware design, but the commonly used low-power inference-type NPU in edge intelligent computing service devices does not have the hardware virtualization ability; ② Time-sharing multiplexing can be used as a general resource scheduling technology to provide the ability of multi-task concurrency for NPU. However, the current NPU design can only execute the model inference as a whole and does not support switching to other tasks during the execution process. This results in too coarse-grained inference task granularity in the existing NPU time-sharing multiplexing solution and cannot guarantee the real-time performance of the inference task.

[0004] Prior Art 1: In hardware virtualization technology, by dividing physical Graphics Processing Unit (GPU) or NPU resources into multiple virtual devices, multiple users or applications can share the same physical GPU resource and independently execute their respective models. This technology not only significantly improves the utilization rate of GPU resources, but also helps reduce costs and enhance system reliability. The implementation of mainstream virtualization technologies usually relies on dedicated support at the hardware level. For example, NVIDIA's vGPU (Virtual Graphics Processing Unit) technology allows multiple virtual machines to bypass the host and directly connect to the physical GPU, thus realizing the sharing of the physical GPU among multiple virtual machines. However, the above technology relies on the support of technologies such as virtualization (SR-IOV, Single Root I / O Virtualization) and access control services (ACS, Access Control Services), and is only available on certain server-level or professional-level graphics cards of NVIDIA; Huawei's Atlas series of NPUs provides higher performance and lower latency by directly implementing the scheduling and management of virtual machines at the hardware level. Resource isolation is achieved by dividing the physical NPU resources into multiple independent AI Cores and AI CPUs, but it depends on specific hardware designs and is only applicable to specific NPU models. Therefore, the disadvantages of hardware virtualization technology are that it is difficult to apply to edge intelligent computing service scenarios. Existing hardware virtualization technologies are restricted by the hardware designs of NPUs and GPUs and are only applicable to high-end training and inference integrated processors, while low-power NPUs suitable for edge intelligent computing services are difficult to share hardware resources through hardware virtualization solutions.

[0005] Prior Art Two: Time-division multiplexing technology divides the processing time of a GPU or NPU into discrete time slices. Each time slice allocates a portion of the processor's computing and memory resources to different applications. Specifically, the time-division multiplexing method typically places the arriving tasks in a queue managed by a scheduler and executes them sequentially according to a user-selected scheduling policy. When the time slice allocated to a task ends or the task is completed, the scheduler performs a context switch to load the state of the next task. In this way, time-division multiplexing technology can concurrently execute multiple tasks on a single processor, thereby improving resource utilization and scheduling flexibility. This technology can be applied to various types of processors, including low-power NPUs, but it can usually only be scheduled after the model processing is completed, and it is difficult to perform task switching during the model processing, resulting in a too coarse scheduling granularity for existing time-division multiplexing NPU solutions and making it difficult to ensure the real-time performance of execution. At the same time, frequent context switches between different workloads will bring performance overhead and increase the latency of task execution, thereby reducing the overall system efficiency. When a low-power NPU performs inference, it can usually only schedule and execute the next inference task after the model processing corresponding to a certain task is completed, resulting in the situation that a task with a longer running time may cause other tasks queuing up to wait for a long time and miss the deadline. Therefore, the time-division multiplexing NPU solution is difficult to meet the real-time requirements of edge intelligence scenarios.

[0006] Prior Art Three: The patent document with the publication (announcement) number CN112348172B discloses a GPU resource scheduling method, device, electronic device, and storage medium, which is mainly used for the deployment and execution of large-scale neural network models. The implementation method includes: first determining the service type of the target large-scale neural network model, then obtaining the corresponding target preset model segmentation method according to the service type, dividing the large-scale neural network model into multiple sub-models, and finally loading the sub-models onto the corresponding GPUs to complete the calculation; the model segmentation method can ensure that each sub-model can independently complete its corresponding task part, and the outputs of the above task parts can be directly combined into the final result. Through this method of automatically segmenting the neural network, effective resource scheduling is achieved, avoiding the complex process of manually segmenting the model and summarizing the results of the model block processing; however, it neither considers the need to execute multiple models on the same device in edge devices nor guarantees the computational real-time performance. Therefore, it is not suitable for the application scenario of real-time inference in edge devices.

[0007] Prior Art Four: The patent document with the publication (announcement) number CN114780240A discloses a workflow scheduling method, device, and storage medium based on GPU time-sharing multiplexing. By dividing all GPU video memory resources in the cluster into multiple time segments and allocating and marking them on demand, it ensures the efficient use of resources and solves the problem of performance degradation caused by the inability to dynamically allocate GPU resources in the prior art. Its innovations are as follows: real-time monitoring of the utilization rate of GPU resources by the workflow and dynamically allocating or recycling resources, thereby improving the utilization efficiency of GPU resources; setting time-segment-based marks for GPU resources to avoid workflow failures caused by resource conflicts. However, the above method focuses on optimizing the dynamic allocation of video memory resources in the GPU cluster rather than the computing resource sharing of a single NPU. On the other hand, it also does not consider the real-time scheduling of tasks. Therefore, it is not suitable for real-time inference application scenarios.

[0008] Prior Art Five: The patent document with the publication (announcement) number CN114004730A discloses a multi-model parallel inference method based on a graphics processing unit. This method constructs multiple inference engines according to the network models to be inferred, binds the inference engines with the corresponding input addresses and GPU stream objects, then creates inference groups and binds them with CPU thread objects, and finally adds the inference engines to the corresponding inference groups and constructs an inference manager to initiate multi-threaded operations. The above method realizes data sharing during the multi-model inference process, thereby enabling the more efficient utilization of the parallel processing power of the GPU and effectively performing thread synchronization operations during multi-model inference calculations. However, the above method is essentially a GPU parallel inference program flow. Although the efficiency of data sharing and thread synchronization is improved by introducing the multi-threaded inference group method, it is difficult to construct a real-time scheduling strategy for parallel inference threads. At the same time, the implementation details of the above method are highly bound to the GPU structure and processing flow and are not suitable for NPUs.

[0009] Prior Art Six: The patent document with the publication (announcement) number CN112348172B discloses a method for collaborative inference of deep neural networks based on an edge-cloud-end architecture. This method first aggregates the transmission bandwidth, data volume, and inference energy consumption of the edge side and the cloud side to the end side, and then, based on the above information, through a convex optimization or reinforcement learning algorithm with latency / energy consumption as the goal, divides the computing tasks in the model inference process into three parts according to the network environment, resource quotas, and usage conditions of the three parties of the edge, cloud, and end. Then, the divided computing tasks are sent to the corresponding end, edge, and cloud for execution in sequence and the inference results are returned. The above method can accelerate the inference speed of the end side through the collaboration of the edge, cloud, and end, improve the inference real-time performance in business scenarios, and reduce the energy consumption of the resource side. However, the above method does not consider the situation where resource-constrained edge devices may cause tasks to not respond in a timely manner due to resource competition when processing multiple computing requests, thus affecting real-time performance. Therefore, it is impossible to maintain the real-time performance of inference tasks under resource-constrained conditions.

[0010] In summary, most of the existing related methods are oriented towards GPU clusters or processor internal caches, and no method specifically for ensuring the real-time performance of multi-model inference tasks has been proposed, so it is difficult to meet the requirements of real-time edge intelligent computing. The mainstream solutions for solving the problem of multi-model execution of NPU currently, namely sharing NPU hardware resources and time-sharing multiplexing of NPU, have the following problems: (1) Sharing NPU hardware resources requires resource isolation at the hardware level through NPU virtualization technology and allocating it for parallel use by multiple inference tasks, which can achieve the real-time performance of multi-tasks to a certain extent, but the cost and power consumption are relatively high. Therefore, the solution of sharing NPU hardware resources cannot be used. (2) Although the existing NPU time-sharing multiplexing technology can be applied to low-power inference-type NPUs commonly used in edge intelligent scenarios, the existing technology schedules with the entire single inference task as a unit. When there are multiple tasks, the next inference task cannot be executed until the previous task ends. At the same time, due to the relatively coarse model granularity, other queued tasks may wait for a long time and miss the deadline, making it difficult to ensure the real-time performance of multiple inference tasks.

[0011] Therefore, a real-time inference system and scheduling method for NPU time-sharing multiplexing in a real-time edge intelligent computing scenario are needed. Summary of the Invention

[0012] A brief overview of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description to follow.

[0013] In view of this, to solve the problem that traditional edge intelligent computing service devices and inference scheduling methods in the existing technology are difficult to ensure the real-time performance of multi-inference task inference, the present invention provides a real-time inference system and scheduling method for NPU time-sharing multiplexing.

[0014] The first technical solution is as follows: A real-time inference system for NPU time-sharing multiplexing, including a heterogeneous computing system and a multi-model serial real-time inference controller;

[0015] The heterogeneous computing system includes an NPU module, a CPU module, and a DDR memory. The NPU module is connected to the CPU module through a PCIe bus, and the CPU module is connected to the DDR memory;

[0016] The CPU module is used to control the NPU module to complete the intelligent computing process;

[0017] The DDR memory is used to cache model data and service data to be processed;

[0018] The multi-model serial real-time inference controller includes a ready-state pre-segmenter, a runtime planner, and a runtime executor connected in sequence. The ready-state pre-segmenter provides model blocks for the runtime executor. The runtime planner divides the inference request into task blocks, and the runtime executor calls the corresponding model blocks according to the task blocks until the entire model inference is completed.

[0019] Further, the NPU module includes an on-chip cache, a matrix calculation unit, a vector calculation unit, and a data transfer unit. The data transfer unit is connected to the on-chip cache, and transfers data from the DDR into the on-chip cache, and the cache is then connected to the matrix calculation unit and the vector calculation unit.

[0020] Further, the ready-state pre-segmenter includes a model evaluation component and a model division component. The model evaluation component is connected to the model division component. The ready-state pre-segmenter is used to segment the model and place it into the model block repository in the preprocessing stage. The runtime planner includes an RPC server, an inference task pool, an inference task scheduler, and an inference queue connected in sequence. The runtime planner is used to receive the model inference calculation request task submitted by the terminal in the execution stage, decompose the task, and fill it into the inference queue containing task blocks. The runtime executor is used to sequentially take out tasks from the inference queue in the execution stage and call the corresponding model blocks for execution. The model division component is connected to the inference task scheduler, and the inference task scheduler (and the inference queue therein) is connected to the runtime executor.

[0021] The second technical solution is as follows: A real-time inference scheduling method for NPU time-sharing multiplexing, used for a real-time inference system for NPU time-sharing multiplexing described in the first technical solution, includes the following steps:

[0022] S1. Use a preparatory pre - splitter to convert and split the intelligent model, that is, convert the intelligent model into a format file executable by the NPU model, test and evaluate the execution time of each layer of the model, and combine the unit granularity to split the intelligent model into model blocks of different granularities;

[0023] S2. Use a runtime planner to discretize the total execution time and task cycle of each intelligent model, that is, discretize the total execution time and task cycle of each intelligent model into integers to reduce the computational overhead of the scheduling algorithm. Before all tasks start, determine the scheduling granularity through the unit granularity and the user - given load threshold using non - linear optimization, and use the scheduling granularity to solve the execution time and the discretized values of the task cycle of the intelligent model blocks, obtaining the discretized execution time and task cycle used by the scheduling algorithm;

[0024] S3. Based on the discretized execution time and task cycle, use a low - sliced NPU real - time scheduling algorithm with a reuse mechanism to obtain the task scheduling sequence, and use a runtime executor to execute according to the task scheduling sequence.

[0025] Furthermore, in S1, it specifically includes the following steps:

[0026] S11. Obtain the intelligent model required for the task from the model repository;

[0027] S12. Use a model evaluation component to obtain the weight files of all intelligent models, and convert the format of the intelligent model into an NPU operator graph file;

[0028] S13. Statistically measure the total execution time and the execution time of each layer of the operator graph file of all intelligent models based on the NPU runtime;

[0029] S14. According to the execution time of each layer of the intelligent model, unit granularity and other conditions, use a model partitioning component to recursively split each intelligent model to obtain model blocks of different granularities;

[0030] In S12, during the conversion process, automatically quantize and fuse the intelligent model, and compile and generate an intelligent model file executable by the NPU module;

[0031] In S13, establish a time - consuming list τ, the total execution time of the intelligent model of the i - th task is τ i , and the execution time of the j - th layer of the intelligent model is Take the task cycle of the task of the corresponding type of the intelligent model as the task cycle t of the intelligent model i ;

[0032] In S14, taking the layer as a unit, each time the intelligent model or the model block is divided into two parts with equal execution time granularity until further splitting will result in the execution time of the model block being less than the unit granularity g0 or the model block only contains one layer. All model blocks with different granularities generated by the division are stored in the model block warehouse.

[0033] Further, in S2, obtain the user task attributes, and discretize the execution time and task period of the task. The specific method is as follows: For task i, its execution time is τ i , and the task period is t i . The discretized execution time using the scheduling granularity g is ceil(τ i / g), and the discretized task period is floor(t i / g). Here, the ceil() and floor() functions represent rounding up and rounding down respectively. The task load of task i after discretization is U i (g), and U i (g) = ceil(τ i / g) / floor(t i / g). For a system with I tasks, the total load is the sum of the loads of all tasks, that is where I is the total number of tasks;

[0034] To reduce the scheduling overhead, that is, to reduce the scheduling frequency and the increase in computational complexity caused by model splitting, find a larger scheduling granularity g on the premise of at least ensuring that the total load does not exceed the preset threshold U max . Model this problem as a non-linear multi-objective constrained optimization problem, where the independent variable is the scheduling granularity g, and the optimization objectives are: 1. Minimize the number of scheduling operations required within the observation time, where the observation time is the least common multiple of the scheduling periods of each task; 2. Minimize the sum of the gaps generated when all model blocks are scheduled onto consecutive time slices. The constraint conditions are: 1. The total period of the discretized tasks does not exceed the preset load threshold U max ; 2. The scheduling granularity g is in the interval [g0, τ min , where g0 is the unit granularity, and τ min is the minimum value of the task execution time among the I tasks (where I is the total number of tasks). This optimization problem uses an optimization method to search for the optimal solution set and selects the optimal scheduling granularity g from it, and uses this scheduling granularity to discretize the execution time and task period of all tasks.

[0035] Further, in S3, the low-segmented NPU real-time scheduling algorithm uses the branch and bound method to generate a scheduling scheme within the observation time. The algorithm starts from the initial state and repeatedly calculates the jobs to be scheduled for the task state in the current time slice, thereby recursively advancing the task state to the next time slice. When it is detected during the generation process that situations such as preemption are about to occur, two choices of executing the original job and executing the new job are recursively processed, and the losses caused by the two choices are calculated respectively: 1. The loss of not being able to execute the most prioritized task, and 2. The additional loss brought about by model segmentation. When the number of recursive generations of the scheduling scheme reaches the number of observation time slices given by the user, the algorithm terminates the generation, outputs the scheduling process corresponding to a series of choices with the minimum cumulative loss, and the segmentation granularity of the corresponding intelligent model as the final result, and updates the inference queue. The above tasks include but are not limited to actual inference tasks such as target recognition and edge detection required by the user, and a job is the specific execution situation of a task at a specific time point;

[0036] In the process of calculating the loss, the loss function is defined as Loss = cω + K. Where c is the number of task preemption times experienced from the initial state to the current process; ω is the weight parameter defined by the user; K is the degree of violating the principle of preferentially scheduling the most urgent task during the process, and is defined as where n is the number of time slices experienced from the initial state to the current state, N is the total number of time slices, and k n represents whether the scheduling situation of the nth time slice violates the principle of preferentially scheduling the most urgent task. When the job executed in the nth time slice belongs to the most urgent task at this time, k n = 0, otherwise, k n = 1;

[0037] Prune the scheduling branches according to the accumulated loss. The process is as follows: If multiple branches reach the same state at a certain moment, that is, the remaining execution time of the task in different branches, the execution duration of the most recent job in all branches, and the task to which the most recent job belongs in all branches are exactly the same, then only keep the branch with the lowest loss function value among the branches with the same state. If there are multiple scheduling branches with the same state and the loss function values are equally the lowest, then arbitrarily keep one of them;

[0038] The low-segmented NPU real-time scheduling algorithm is compatible with aperiodic tasks by introducing a time slice multiplexing mechanism. After allocating tasks for each time slice, the low-segmented NPU real-time scheduling algorithm generates the execution order of model segmentation through two steps: merging and multiplexing, as follows:

[0039] 1. Merge consecutive time slices assigned to the same task, find the model segment in the model segment warehouse that is not greater than and closest to the length of this consecutive time slice for allocation. If the idle time slot of this time slice is too large after allocation, then find the model segment that matches the idle time slot again for allocation;

[0040] 2. Without changing the execution order of the model chunks, the time - slice multiplexing mechanism is used to advance the execution of the model chunks to the larger value between the end time of the previous model chunk and the arrival time of the current task. Then, the idle time slots generated in each time - slice are merged to create continuous gaps, and the appropriate model chunks of bursty aperiodic tasks or other transactions are inserted into the inference queue for execution using these continuous gaps, thereby reducing the impact of bursty tasks on the real - time performance of periodic task execution.

[0041] The beneficial effects of the present invention are as follows:

[0042] 1. Through fine - grained model segmentation and improved scheduling methods, the present invention realizes the time - sharing multiplexing of NPU resources on edge intelligent computing service devices, thereby effectively improving the real - time performance of NPU inference in task scenarios with large differences in different task cycles;

[0043] 2. The present invention is an NPU time - sharing multiplexing real - time inference system, which consists of hardware modules such as CPU, NPU, and DDR memory. It adapts to the scheduling method based on fine - grained model chunks, provides an acceleration mechanism for pre - buffering data to be processed for the next model chunk inference using the on - chip high - speed cache of the NPU, and prepares and buffers the data required for the next inference in advance through the management of the on - chip high - speed cache structure of the NPU, thereby achieving the effect of accelerating execution;

[0044] 3. The present invention proposes a real - time inference scheduling method for NPU time - sharing multiplexing. In the preparation stage, the model is divided into a series of chunks with different granularities by a preparation - state pre - splitter; before task execution, the scheduling granularity is determined according to task attributes; during task execution, the corresponding inference job sequence is generated according to the real - time scheduling algorithm; finally, the model chunks with corresponding granularities are called and executed according to the job sequence;

[0045] 4. The present invention proposes a scheme for recursively splitting the model to divide the model into model chunks with different granularities and save them separately. On the one hand, the granularity of the model is refined; on the other hand, the model chunks with different granularities ensure that the subsequent scheduling algorithm can adopt a dynamic strategy. The recursive splitting process of the model does not depend on task attributes and can dynamically adapt to task scenarios according to changes in task attributes.

[0046] 5. The present invention can select the optimal scheduling granularity, i.e., the scheduling time - slice, according to the task execution time and period, and can discretize the specific execution time, task period, and total load of all tasks' models and then select the optimal granularity;

[0047] 6. The present invention proposes a low-segmentation NPU inference real-time scheduling algorithm, which makes a trade-off between scheduling more prioritized tasks as much as possible and avoiding model segmentation as much as possible through the branch and bound method, so as to give consideration to both preferentially scheduling urgent tasks and improving the execution efficiency of the algorithm on the NPU; at the same time, a time slice multiplexing mechanism is used to execute the idle time slots generated during the model block execution process as early as possible by concentrating the model blocks, and the model blocks of aperiodic tasks are inserted into the inference queue for execution by using continuous gaps, thereby reducing the impact on the real-time performance of periodic tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0049] Figure 1 is a system architecture diagram of a real-time inference system with NPU time-sharing multiplexing;

[0050] Figure 2 is an example system architecture diagram of a real-time inference system with NPU time-sharing multiplexing;

[0051] Figure 3 is a schematic structural diagram of a heterogeneous computing system;

[0052] Figure 4 is a schematic diagram of a multi-model serial real-time inference controller;

[0053] Figure 5 is a flow overview diagram of a real-time inference scheduling method with NPU time-sharing multiplexing;

[0054] Figure 6 is a flowchart of an example of a model segmentation and scheduling granularity acquisition algorithm;

[0055] Figure 7 is a flowchart of an example of a low-segmentation NPU real-time scheduling algorithm.

[0056] BRIEF DESCRIPTION OF THE DRAWINGS: 1. NPU module; 2. NPU module; 3. NPU module; 4. Preparation state pre-segmenter; 5. Runtime planner; 6. Runtime executor. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] In order to make the technical solutions and advantages in the embodiments of the present invention clearer and more understandable, the following further details the exemplary embodiments of the present invention with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0058] Embodiment 1: Reference Figures 1 - 3 This embodiment will be described in detail. A real-time inference system with time-sharing multiplexing of NPUs, including a heterogeneous computing system as the hardware part and a multi-model serial real-time inference controller as the software part;

[0059] The heterogeneous computing system includes an NPU module 1, a CPU module 2, and a DDR memory 3. The NPU module 1 is connected to the CPU module 2 through a PCIe bus, and the CPU module 2 is connected to the DDR memory 3;

[0060] The CPU module 2 is used to control the NPU module 1 to complete the intelligent computing process;

[0061] The DDR memory 3 is used to cache model data and service data to be processed;

[0062] The multi-model serial real-time inference controller includes a ready-state pre-segmenter 4, a runtime planner 5, and a runtime executor 6 connected in sequence. The ready-state pre-segmenter 4 provides model chunks for the runtime executor 6. The runtime planner 5 divides the inference request into task chunks, and the runtime executor 6 calls the corresponding model chunks according to the task chunks until the entire model inference is completed.

[0063] Furthermore, the NPU module 1 includes an on-chip cache, a matrix calculation unit, a vector calculation unit, and a data transfer unit. The data transfer unit is connected to the on-chip cache, transfers data from the DDR into the on-chip cache, and the cache is then connected to the matrix calculation unit and the vector calculation unit;

[0064] Specifically, refer to Figure 2 and Figure 3 , the CPU module is used to schedule algorithms and other CPU algorithms, the NPU module is used as a coprocessor for DNN model inference acceleration, the DDR memory is used to cache inference process data, and the data transfer unit transfers data between the DDR memory and the on-chip cache;

[0065] Furthermore, the ready-state pre-segmenter 4 includes a model evaluation component and a model partitioning component. The model evaluation component is connected to the model partitioning component. The ready-state pre-segmenter 4 is used to segment the model and place it in the model chunk repository during the preprocessing stage. The runtime planner 5 includes an RPC server, an inference task pool, an inference task scheduler, and an inference queue connected in sequence. The runtime planner 5 is used to receive the model inference calculation request task submitted by the terminal during the execution stage, decompose the task, and fill it into the inference queue containing task chunks. The runtime executor 6 is used to sequentially take out tasks from the inference queue during the execution stage and call the corresponding model chunks for execution. The model partitioning component is connected to the inference task scheduler, and the inference queue is connected to the runtime executor 6;

[0066] Specifically, referring to Figure 4 , the preparation-state pre-segmenter is used to segment the model and place it in the model block repository during the preprocessing stage. It includes a model evaluation component and a model partitioning component. The model evaluation component is used to convert the model format, obtain the model and the running time of each layer therein, and calculate the model segmentation granularity. The model partitioning component is used to divide the model into model blocks according to the segmentation granularity;

[0067] The runtime planner is used to receive the model inference calculation request task submitted by the terminal during the execution stage, and decompose and fill the task into the inference queue containing task blocks according to the scheduling algorithm. During runtime, the RPC server is used to receive the task request initiated by the intelligent terminal through RPC and save the task in the inference task pool. The inference task scheduler is used to generate corresponding jobs according to the inference tasks in the inference task pool, adjust the job order and fill it into the inference queue. The inference queue is used to save the order of scheduling the NPU to execute the block inference job.

[0068] The runtime executor is used to sequentially fetch the inference jobs from the inference queue, then obtain the corresponding model blocks from the model block repository for calculation, and output the inference result when the calculation of the last model block of the task is completed.

[0069] The runtime executor will use the on-chip high-speed cache of the NPU to pre-buffer the data to be processed for the next model block inference. When the runtime executor is executing, it sequentially fetches jobs from the head of the inference queue and obtains the corresponding model blocks from the model block repository for serial calculation until all the inference tasks in the inference queue are completed. 1. If the calculation result is an intermediate result of a certain inference task, when the on-chip high-speed cache has sufficient capacity, it is preferably temporarily stored in the on-chip high-speed cache instead of being moved into the DDR memory for subsequent use of the task; when the on-chip high-speed cache is full, according to the queuing order of the jobs in the inference queue, the on-chip high-speed cache is adjusted, and the buffered data required by the jobs queued later is moved to the DDR memory. 2. If the calculation result is the final result, it is output through the corresponding interface.

[0070] Embodiment 2: Referring to Figures 4 - 6 This embodiment is described in detail. A real-time inference scheduling method for NPU time-sharing multiplexing is used for the real-time inference system for NPU time-sharing multiplexing described in Embodiment 1, and specifically includes the following steps:

[0071] S1. Use a preparatory pre - splitter to convert and split the intelligent model. First, convert the intelligent model into a format file executable by the NPU, then test and evaluate the execution time of each layer of the model, and finally, combine the unit granularity to split the model into model chunks of different granularities. If the computing time of a model chunk of a certain granularity is equal on the CPU module and the NPU module respectively, then this computing time is defined as the unit granularity, denoted as g0. This granularity can be obtained based on testing the target heterogeneous computing system.

[0072] S2. Use a runtime planner to discretize the total execution time and task cycle of each intelligent model, that is, discretize the total execution time and task cycle of each intelligent model into integers to reduce the computational overhead of the scheduling algorithm, specifically including: 1. By discretizing the execution time of the intelligent model, the scheduling time slice (generally the greatest common divisor of the execution times of I intelligent models) is not too small; 2. By discretizing the task cycle, the observation time of the scheduling algorithm (generally the least common multiple of the task cycles of I inferences) is not too large (where I is the total number of tasks);

[0073] Before all tasks start, first through the unit granularity g0 and the user - given load threshold U max Determine the scheduling granularity g, that is, the length of the scheduling time slice, through non - linear optimization. Then use this scheduling granularity to discretize the execution time and task cycle of the intelligent model chunks, so as to obtain the discretized execution time and task cycle used by the scheduling algorithm.

[0074] S3. Based on the discretized execution time and task cycle, use a low - cut NPU real - time scheduling algorithm with a reuse mechanism to obtain the task scheduling sequence, and use a runtime executor to execute according to the task scheduling sequence.

[0075] Furthermore, in S1, it specifically includes the following steps:

[0076] S11. Obtain the intelligent model required for the task from the model repository;

[0077] S12. Use a model evaluation component to obtain the weight files of all intelligent models, and convert the format of the intelligent models into NPU operator graph files;

[0078] S13. Statistically measure the total execution time and the execution time of each layer of the operator graph files of all intelligent models based on NPU runtime;

[0079] S14. According to the execution time of each layer of the intelligent model, unit granularity and other conditions, use a model partitioning component to recursively split each intelligent model to obtain model chunks of different granularities;

[0080] In S12, during the conversion process, the intelligent model is automatically quantified and fused, and an intelligent model file executable by the NPU module is compiled and generated;

[0081] In S13, a time-consuming list τ is established, and the total execution time (or the corresponding task execution time) of the intelligent model of the i-th task is τ i , and the execution time of the j-th layer of the intelligent model is τ j i , and the task period of the intelligent model corresponding to the type of task is used as the task period t of the intelligent model i ;

[0082] In S14, each time the intelligent model or the model block is divided into two parts with equal execution time granularity until further splitting will result in the execution time of the model block being less than the unit granularity g0 or the model block only contains one layer, and all model blocks with different granularities generated by the division are stored in the model block warehouse;

[0083] Specifically, the unit granularity g0 means that only when the execution time of the intelligent model is not less than the computing time corresponding to g0, the execution of the model block on the NPU module can obtain acceleration compared with the execution on the CPU module.

[0084] Furthermore, in S2, the user task attributes are obtained, and the execution time and task period of the task are discretized. The specific method is as follows: for task i, its execution time is τ i and the task period is t i , the execution time after discretization using the scheduling granularity g is ceil(τ i / g), and the discretized task period is floor(t i / g), where the ceil() and floor() functions represent rounding up and rounding down respectively, and the task load of task i after discretization is U i (g), U i (g) = ceil(τ i / g) / floor(t i / g). For a system with I tasks, the total load is the sum of the loads of all tasks, that is

[0085] To reduce the scheduling overhead, that is, to reduce the scheduling frequency and the increase in the amount of computation caused by model splitting, at least ensuring that the total load does not exceed the preset threshold U maxFind a larger scheduling granularity \(g\) on the premise of this. Model this problem as a non-linear multi-objective constrained optimization problem, where the independent variable is the scheduling granularity \(g\); the optimization objectives are: 1. Minimize the number of times of scheduling required within the observation time, where the observation time is the least common multiple of the scheduling periods of each task, 2. Minimize the sum of the gaps generated when all model blocks are scheduled onto consecutive time slices; the constraint conditions are: 1. The total period of the discretized tasks does not exceed the preset load threshold \(U\). max , 2. The scheduling granularity is in the interval \([g_0, \tau min \), where \(g_0\) is the unit granularity and \(\tau min \) is the minimum of the execution times among the \(I\) tasks. This optimization problem uses an optimization method to search for the optimal solution set and selects the optimal scheduling granularity \(g\) from it. Use this scheduling granularity to discretize the execution times and task periods of all tasks.

[0086] Furthermore, in the above S3, the low-cut NPU real-time scheduling algorithm uses the branch and bound method to generate a scheduling scheme within the observation time. The algorithm starts from the initial state and repeatedly calculates the jobs that should be scheduled for the task state at the current time slice, thereby recursively advancing the task state to the next time slice. When it is detected during the generation process that situations such as task preemption are about to occur, the two options of executing the original job and executing the new job are recursively processed respectively, and the losses caused by the two options are calculated respectively: 1. The loss of not being able to execute the most urgent task, 2. The additional loss brought by model segmentation. When the number of recursive generations of the scheduling scheme reaches the number of observation time slices given by the user, the algorithm terminates the generation, outputs the scheduling process corresponding to a series of selections with the minimum cumulative loss, and the segmentation granularity of the corresponding intelligent model as the final result, and updates the inference queue. The above tasks include but are not limited to actual inference tasks such as target recognition and edge detection required by the user, and the job is the specific execution situation of the task at a specific time point;

[0087] During the process of calculating the loss, the loss function is defined as Loss = \(c\omega+K\). Where \(c\) is the number of task preemption times experienced from the initial state to the current process; \(\omega\) is the weight parameter defined by the user; \(K\) is the degree of violating the principle of preferentially scheduling the most urgent task during the process, and is defined as where \(n\) is the number of time slices experienced from the initial state to the current state, \(N\) is the total number of time slices, and \(k n represents whether the scheduling situation of the \(n\)th time slice violates the principle of preferentially scheduling the most urgent task. When the job executed in the \(n\)th time slice belongs to the most urgent task at this time, \(k n = 0, otherwise, \(k n = 1;

[0088] After that, pruning is performed on the scheduling branches according to the accumulated losses, and the process is as follows: If multiple branches reach the same state at a certain moment, that is, the remaining execution time of the task in different branches, the execution duration of the most recent job in all branches, and the task to which the most recent job belongs in all branches are exactly the same, then only the branch with the lowest loss function value is retained among the branches with the same state. If there are multiple scheduling branches with the same state and the loss function values are equally the lowest, then any one of them is retained randomly;

[0089] The low-partitioning NPU real-time scheduling algorithm is compatible with aperiodic tasks by introducing a time-slice multiplexing mechanism. After tasks are allocated to each time-slice, the low-partitioning NPU real-time scheduling algorithm generates the execution order of the model blocks through two steps: merging and multiplexing, as follows:

[0090] 1. Merge consecutive time-slices assigned the same task, and then find the model block in the model block repository that is not greater than and closest to the time-slice length for allocation. If the idle time-slot of the time-slice is too large after allocation, then search for a model block that matches the idle time-slot again for allocation;

[0091] 2. Without changing the execution order of the model blocks, use the time-slice multiplexing mechanism to advance the execution of the model blocks to the larger value between the end time of the previous model block and the arrival time of the current task. Then, merge the idle time-slots generated by each time-slice to create continuous gaps, and use the continuous gaps to insert appropriate model blocks of burst aperiodic tasks or other transactions into the inference queue for execution, thereby reducing the impact of burst tasks on the real-time performance of periodic task execution.

[0092] Reference Figure 6 In this embodiment, the branch and bound method is used as the implementation method. The initial state is modeled as the root node of the solution space tree. After that, a first-in-first-out (FIFO) queue is used to maintain all leaf nodes in the solution space tree, and their states are recursively pushed to the states of the next time-slice in sequence. The above FIFO queue is also called the active node queue;

[0093] Referring to Table 1, for the data domain of the low-partitioning NPU real-time scheduling algorithm, for the convenience of algorithm design, the active node queue may also include a sentinel node with the same state time as the state node but without other information to facilitate determining whether all state nodes have been recursively pushed to the next time-slice. The state node may include a data domain;

[0094] Table 1 Data Domain of the Low-Partitioning NPU Real-Time Scheduling Algorithm

[0095]

[0096]

[0097] Based on the above data domain, the implementation method includes the following steps:

[0098] 1. Initialize the search tree:

[0099] Initialize the search time. If the least common multiple of the task cycles is less than the maximum search time T specified by the user max , then set the search time to the least common multiple of the task cycles;

[0100] Create a root node. Initialize data fields such as the current state time, task execution status, job sequence, forced execution flag, and last executed task of a node, and use it as the root node of the search tree;

[0101] Add the root node and a sentinel node with the same current state time as the root node to the active node queue;

[0102] 2. Process the nodes in the active node queue in sequence. Take out a node from the head of the active node queue, and then process it according to the type of the node:

[0103] (1) If the node is a sentinel node:

[0104] a. If the state time reaches the search time, the search process ends, and backtracking is performed to obtain the job scheduling sequence;

[0105] b. Otherwise, traverse the active node queue, compare the cumulative losses of the nodes with the same task execution table content, and then, among the nodes with the same task execution table and counter, only retain the node with the smallest cumulative loss, and prune the remaining nodes;

[0106] (2) If the node is a state node, then:

[0107] First, check the deadline and addition status of the tasks, including:

[0108] a. Check whether there are tasks that have not been completed beyond the deadline. If so, prune the current node;

[0109] b. Determine whether there is a task whose cycle start time plus the preprocessing time is equal to the current state time. If so, add the task to the task execution table;

[0110] (3) Then, select one of the following steps to execute according to whether the forced execution flag is empty:

[0111] a. If the reference task is empty or the same as the task selected by the EDF algorithm, then task preemption will not occur in the current time slice. At this time, if the forced execution flag is empty, set the reference task and the task to be executed as the job selected by the EDF algorithm; if the forced execution flag is not empty, set the task to be executed as the same as the forced execution flag. In particular, if the chunks with the smallest granularity of some models are larger than a single time slice, then multiple consecutive time slices need to be allocated for the chunks of this model during the scheduling phase. For the above situation, when one of the following conditions is met, that is, 1) the sum of the execution time of the model chunk and the current value of the counter exceeds the preset threshold, 2) a more urgent task arrives during the execution time of the model chunk, then refer to b to create two child nodes for the current node to deduce the two alternative situations of "allocating these consecutive time slices at the beginning of the current time slice" and "scheduling other jobs in the current time slice" respectively, for subsequent evaluation of the cumulative loss and pruning; otherwise, directly allocate consecutive time slices starting from the current time slice that match the time required to execute the model chunk;

[0112] b. If the reference task is not empty and different from the task selected by the EDF algorithm, then task preemption will occur in the current time slice. At this time, create two child nodes for the current node to deduce the two alternative situations of executing the original task and executing the new task respectively. The above child nodes inherit the state time, task execution table, and cumulative loss of the current node, but clear the job sequence. Among them, the child node for executing the new task sets the forced execution flag to empty, sets the counter to zero, sets the reference task and the job to be executed as the task selected by the EDF algorithm. The child node for continuing to execute the original task sets the forced execution flag to the value of the reference task if the forced execution flag is empty, and does not change if the forced execution flag is not empty. After that, the child node for continuing to execute the original task sets the task to be executed as the same as the forced execution flag, sets the counter to the same as the parent node, and sets the reference task as the task selected by the EDF algorithm;

[0113] (4) Finally, update the node state through the following steps. If child nodes are generated, update the states of all child nodes:

[0114] a. Update the cumulative loss of the node. If this node is a newly generated child node for executing a new task, then include the loss caused by task preemption in the loss tuple. After that, if the currently executing task is not the task selected by the EDF algorithm, then include the loss of not executing the most urgent task in the cumulative loss;

[0115] b. Update the remaining execution time of the current task and the counter. If the remaining execution time of the current task becomes zero after the update, then remove the currently executing task from the task execution table, and clear the reference task and the forced execution flag, set the counter to zero. If the counter reaches the preset threshold, then set the counter to zero, also clear the reference task and the forced execution flag, and update the cumulative loss;

[0116] c. Place the currently executing task at the end of the task scheduling queue;

[0117] d. Increment the status time of this node by one;

[0118] e. Place the current node at the end of the active node queue;

[0119] 3. Generate a task scheduling sequence:

[0120] (1) Take the state node with the smallest cumulative loss in the active node queue and start backtracking until reaching the root node, and then splice the task scheduling sequences of each node on the path from the root node to this state node to obtain a task scheduling sequence sliced by time slices;

[0121] (2) Use the greedy algorithm to convert the task scheduling sequence into a job scheduling sequence for execution model chunks, that is, regard consecutive time slices assigned to the same task as a longer complete time slice, and arrange the largest possible model chunks in turn for each time slice under the premise of not exceeding the maximum remaining time of the time slice until no more can be arranged;

[0122] (3) Change the execution time of the model chunk to its execution time before discretization and try to advance the execution of the chunk as much as possible. Since this execution time is shorter than the execution time of the model chunk, the idle time generated by this process can be used to advance the execution of the model chunk as much as possible without changing the execution order, so as to generate as long a continuous idle time as possible for the scheduling of burst tasks, thereby avoiding the influence of burst tasks on the execution of periodic tasks to a certain extent, and further avoiding the impact on the real-time performance of periodic tasks;

[0123] (4) Place the generated job scheduling sequence in the inference queue.

[0124] Although the present invention has been described with reference to a limited number of embodiments, those skilled in the art in this technical field will understand, based on the above description, that other embodiments can be conceived within the scope of the present invention thus described. In addition, it should be noted that the language used in this specification is mainly selected for readability and teaching purposes, rather than for the purpose of explaining or limiting the subject matter of the present invention. Therefore, many modifications and variations are obvious to those of ordinary skill in the art in this technical field without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure of the present invention is illustrative rather than restrictive, and the scope of the present invention is defined by the appended claims.

Claims

1. A real-time reasoning scheduling method for NPU time-sharing multiplexing, characterized in that: The following steps are involved: S1. Use the preparatory pre-segmenter to convert and segment the intelligent model, that is, convert the intelligent model into a format file executable by the NPU model, test and evaluate the execution time of each layer of the model, and split the intelligent model into model blocks of different granularities based on the unit granularity; S2. The total execution time and task cycle of each intelligent model are discretized by the runtime planner, that is, the total execution time and task cycle of each intelligent model are discretized into integers to reduce the computational overhead of the scheduling algorithm. Before all tasks start, the scheduling granularity is determined by nonlinear optimization based on the unit granularity and the user-given load threshold, and the scheduling granularity is used to solve the execution time and task cycle discretization values ​​of the intelligent model blocks to obtain the discretized execution time and task cycle used by the scheduling algorithm; S3. Based on the discrete execution time and task cycle, the low-cut NPU real-time scheduling algorithm with a multiplexing mechanism is used to obtain the task scheduling sequence, and the runtime executor is used to execute according to the task scheduling sequence; The S1 specifically includes the following steps: S11. Obtain the intelligent model required for the task from the model warehouse; S12. Use the model evaluation component to obtain the weight files of all intelligent models and convert the format of the intelligent model into an NPU operator graph file; S13. Statistically measure the total execution time and each layer execution time of the operator graph files of all intelligent models based on the NPU runtime; S14. According to the execution time and unit granularity conditions of each layer of the intelligent model, each intelligent model is recursively split using the model partitioning component to obtain model blocks of different granularities; In S12, the intelligent model is automatically quantized and integrated during the conversion process, and the intelligent model file executable by the NPU module is compiled and generated; In S13, a time-consuming list τ is established, and the total execution time of the intelligent model of the i-th task is τ i , the execution time of the jth layer of the intelligent model is The task period of the corresponding type of task of the intelligent model is taken as the task period t of the intelligent model. i ; In S14, the intelligent model or model block is divided into two parts with equal execution time granularity each time, until further splitting will cause the execution time of the model block to be less than the unit granularity g0 or the model block contains only one layer, and all model blocks of different granularities generated by the division are stored in the model block warehouse.

2. According to the real-time inference scheduling method of NPU time-sharing multiplexing according to claim 1, it is characterized in that: In S2, the user task attributes are obtained, and the task execution time and task cycle are discretized. The specific method is: for task i, its execution time is τ i , the task cycle is t i , the execution time after discretization using the scheduling granularity g is ceil(τ i / g), the discretized task period is floor(t i / g), where ceil() and floor() functions represent rounding up and rounding down respectively. The task load after task i is discretized is U i (g), U i (g) = ceil(τ i / g) / floor(t i / g), for a system with I tasks, the total load U(g) is the sum of the loads of all tasks, that is, Where I is the total number of tasks; It will at least ensure that the total load does not exceed the preset threshold U max The problem of finding a larger scheduling granularity g under the premise of is modeled as a nonlinear objective constraint optimization problem, where the independent variable is the scheduling granularity g, and the optimization objectives are: 1) minimize the number of schedulings required within the observation time, where the observation time is the least common multiple of the scheduling period of each task, and 2) minimize the gaps generated by all model block scheduling on continuous time slices; the constraints are: 1) the total period of the discretized task does not exceed the preset load threshold U max , 2) The scheduling granularity g is in [g0,τ min ] interval, where g0 is the unit particle size, τ min The minimum execution time of a task in a task is obtained by using an optimization method to search for the optimal solution set, and the optimal scheduling granularity g is selected from it. This scheduling granularity is used to discretize the execution time and task cycle of all tasks.

3. The real-time inference scheduling method for NPU time-division multiplexing according to claim 2 is characterized in that: In S3, the low-cut NPU real-time scheduling algorithm uses the branch and bound method to generate a scheduling plan within the observation time. It repeatedly calculates the task scheduling state of the task state in the time slice from the initial state, thereby recursively pushing the task state to the next time slice. When it is detected during the generation process that task preemption or scheduling of a larger model block is about to occur, the two options of executing the original job and executing the new job are recursively pushed, and the losses caused by the two options are calculated: 1) the loss of failing to execute the highest priority task; 2) Failure to avoid the loss of the segmentation model. When the number of recursive generation of the scheduling scheme reaches the number of observation time slices given by the user, the algorithm terminates the generation and outputs a series of corresponding scheduling processes with the smallest cumulative loss. The block granularity of the corresponding intelligent model is used as the final result to update the inference queue. The above tasks include the actual inference tasks of target recognition and edge detection required by the user, and the job is the specific execution of the task at a specific time point; In the process of calculating the loss, the loss function is defined as Loss = cω + K, where c is the number of task preemptions experienced from the initial state to the current process; ω is a user-defined weight parameter; K is the degree of violation of the principle of prioritizing the most urgent tasks in the process, defined as Among them, n is the number of time slices from the initial state to the current state, N is the total number of time slices, and k is the total number of time slices. n Indicates whether the scheduling of the nth time slice violates the principle of prioritizing the most urgent tasks. When the job executed in the nth time slice belongs to the most urgent task at this time, k n =0, otherwise, k n =1; The scheduling branches are pruned according to the accumulated loss. The process is as follows: if multiple branches reach the same state at a certain moment, that is, the remaining execution time of tasks in different branches, the execution time of the most recent job in all branches, and the tasks to which the most recent job belongs in all branches are exactly the same, then only the branch with the lowest loss function value is retained among the branches with the same state. If there are multiple scheduling branches with the same state and the loss function values ​​are the same as the lowest, any one of them is retained. The low-cut NPU real-time scheduling algorithm is compatible with non-periodic tasks by introducing a time slice reuse mechanism. After assigning tasks to each time slice, the low-cut NPU real-time scheduling algorithm generates the execution order of model blocks through two steps of merging and reuse, namely: 1) Merge consecutive time slices assigned the same task, find a model block in the model block warehouse that is not greater than and closest to the time slice length and assign it to the time slice. If the idle time slot of the allocated time slice is too large, find a model block that matches the idle time slot again for allocation. 2) Without changing the execution order of the model blocks, the time slice reuse mechanism is used to advance the execution of the model blocks to the larger value of the end time of the previous model block and the arrival time of the current task. The idle time slots generated by each time slice are then merged to produce continuous gaps. The continuous gaps are used to insert appropriate model blocks of sudden non-periodic tasks or other transactions into the inference queue for execution, thereby reducing the impact of sudden tasks on the real-time execution of periodic tasks.

4. A real-time reasoning system with NPU time-sharing multiplexing, characterized in that: A real-time inference scheduling method for executing a NPU time-sharing multiplexing method as described in any one of claims 1 to 3, comprising a heterogeneous computing system and a multi-model serial real-time inference controller; The heterogeneous computing system comprises an NPU module (1), a CPU module (2) and a DDR memory (3), wherein the NPU (1) is connected to the CPU module (2) via a PCIe bus, and the CPU module (2) is connected to the DDR memory (3); The CPU module (2) is used to control the NPU module (1) to complete the intelligent computing process; The DDR memory (3) is used to cache model data and service data to be processed; The multi-model serial real-time reasoning controller comprises a ready state pre-segmenter (4), a runtime planner (5), and a runtime executor (6) which are connected in sequence. The ready state pre-segmenter (4) provides model blocks for the runtime executor (6), the runtime planner (5) divides the reasoning request into task blocks, and the runtime executor (6) calls the corresponding model block according to the task block until the entire model reasoning is completed.

5. The real-time reasoning system of NPU time-sharing multiplexing according to claim 4, characterized in that: The NPU module (1) comprises an on-chip cache, a matrix calculation unit, a vector calculation unit and a data handling unit. The data handling unit is connected to the on-chip cache to move data from the DDR into the on-chip cache. The cache is then connected to the matrix calculation unit and the vector calculation unit.

6. The real-time reasoning system of NPU time-sharing multiplexing according to claim 5, characterized in that: The preparation state pre-segmenter (4) includes a model evaluation component and a model partitioning component, the model evaluation component is connected to the model partitioning component, the preparation state pre-segmenter (4) is used to partition the model in the preprocessing stage and place it into the model block warehouse, the runtime planner (5) includes an RPC server, an inference task pool, an inference task scheduler and an inference queue connected in sequence, the runtime planner (5) is used to receive the model inference calculation request task submitted by the terminal in the execution stage, and decompose the task and fill it into the inference queue containing the task blocks, the runtime executor (6) is used to take out the tasks from the inference queue in sequence in the execution stage, and call the corresponding model block for execution, the model partitioning component is connected to the inference task scheduler, and the inference queue is connected to the runtime executor (6).

Citation Information

Patent Citations

  • A Deep Neural Network Collaborative Inference Method Based on Edge-Cloud Architecture

    CN112348172B

  • Deep neural network multi-model parallel reasoning method based on graphics processor

    CN114004730A

  • Workflow scheduling method and device based on GPU time division multiplexing and storage medium

    CN114780240A

  • Lightweight improved target detection method and device based on Rockchip micro platform

    CN113065555A

  • Cloud edge end DNN collaborative reasoning acceleration method for edge intelligence

    CN113592077A