Intelligent application control system, intelligent application task processing method and electronic equipment

By managing the operating status of the functional components of the intelligent application on the first processor and sending the model tasks to the second processor for execution, the processor overhead problem caused by frequent switching of intelligent applications in heterogeneous systems is solved, and the effect of reducing processor overhead and improving processor utilization is achieved.

CN119987976AActive Publication Date: 2025-05-13HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510459053.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

In heterogeneous systems, each time the intelligent application calls the processor 2 API, it is necessary to switch between user-state and kernel-state, resulting in an increase in processor overhead.

Method used

By running the component scheduling module and the model scheduling module on the first processor, the operating status of functional components in the intelligent application is managed, and the model tasks are sent to the second processor for execution, avoiding calling the user state API and reducing switching between the kernel state and the user state.

Benefits of technology

Reduces processor overhead, improves utilization of the second processor, and reduces bandwidth overhead by batching multiple model tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987976A_ABST
    Figure CN119987976A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of edge computing, and provides an intelligent application control system, an intelligent application task processing method and electronic equipment. The system comprises a first processor and a second processor, the first processor is configured with a component scheduling module and a model scheduling module which run in a kernel mode; the component scheduling module manages the running state of each functional component in the intelligent application, and sends a first task set to the model scheduling module when the functional components run; under the condition that multiple target model tasks exist in the first task set, the model scheduling module partially or completely merges the multiple target model tasks to obtain a second task set meeting the batch processing condition, and sends a model driving instruction to a second processor; and the second processor is used for triggering the related neural network model to execute the model task in the second task set based on the model driving instruction, and feeding back a result obtained by task processing to the first processor. According to the invention, processor overhead can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of edge computing technology, and in particular relates to an intelligent application control system, an intelligent application task processing method and an electronic device. Background Art

[0002] A heterogeneous system includes many different types of processors, such as processor 1 and processor 2. Processor 1 can control processor 2. The operating system (such as Linux) distinguishes between kernel state and user state. Processor 2 has a driver in kernel state and an application programming interface (API) in user state. When an intelligent application involving processor 2 calls a functional component in user state to complete the application process, the functional component needs to call the driver of processor 2 through the API of processor 2 to control processor 2. In this working mode, each time an intelligent application calls the API of processor 2, a switch between user state and kernel state will occur, and this state switching will generate processor overhead. Summary of the invention

[0003] The embodiments of the present application provide an intelligent application control system, an intelligent application task processing method and an electronic device, which can reduce processor overhead.

[0004] In a first aspect, an embodiment of the present application provides an intelligent application control system, including: a first processor and a second processor; The first processor is configured with: a component scheduling module and a model scheduling module running in kernel state; The component scheduling module is used to manage the running status of each functional component in the intelligent application, and send a first task set to the model scheduling module when the functional component is running, wherein the first task set includes: model tasks triggered when the functional component is running, and the model tasks are tasks that depend on the execution of the neural network model; The model scheduling module is used to: when there are multiple target model tasks in the first task set, partially or completely merge the multiple target model tasks in the first task set to obtain a second task set that meets the batch processing condition, wherein the multiple target model tasks are model tasks corresponding to the same neural network model, and the batch processing condition includes: the execution time of each model task in the second task set does not exceed the execution time limit corresponding to each model task; based on the model tasks included in the second task set, send a model driving instruction to the second processor; The second processor is used to: trigger the relevant neural network model to execute the model tasks in the second task set based on the model-driven instructions, and feed back the results of task processing to the first processor.

[0005] In an embodiment of the present application, the first processor manages the running status of each functional component in the intelligent application through a component scheduling module running in kernel state, and can send the model task involving the second processor to the second processor through the model scheduling module running in kernel state to trigger the relevant neural network model to execute the model task, and feedback the result of task processing to the first processor. In the control process of the above-mentioned intelligent application, there is no need to call the API of the second processor running in user state, so it can reduce the repeated switching between kernel state and user state and reduce the overhead of the first processor. And when there are multiple target model tasks in the first task set, the model scheduling module can batch process some or all of the target model tasks in the first task set by partially or completely merging the multiple target model tasks in the first task set, thereby reducing bandwidth overhead and improving the utilization rate of the second processor.

[0006] In a second aspect, an embodiment of the present application provides a method for processing intelligent application tasks, which is applied to a first processor, wherein the first processor is configured with: a component scheduling module and a model scheduling module running in a kernel state; the component scheduling module is used to manage the running status of each functional component in the intelligent application; The method comprises: When the functional component is running, triggering the component scheduling module to send a first task set to the model scheduling module, the first task set includes: model tasks triggered when the functional component is running, the model tasks are tasks that depend on the execution of the neural network model; In the case where there are multiple target model tasks in the first task set, the model scheduling module is triggered to partially or completely merge the multiple target model tasks in the first task set to obtain a second task set that meets a batch processing condition, wherein the multiple target model tasks are model tasks corresponding to the same neural network model, and the batch processing condition includes: the execution time of each model task in the second task set does not exceed the execution time limit corresponding to each model task; The model scheduling module is triggered to send a model-driven instruction to the second processor based on the tasks included in the second task set, and the model-driven instruction is used to trigger the second processor to execute the tasks in the second task set and feedback the results of task processing to the first processor. The second processor is configured with task processing capabilities based on a neural network model.

[0007] In the embodiment of the present application, the first processor manages the running state of each functional component in the intelligent application through the component scheduling module running in the kernel state, which can reduce the repeated switching between the kernel state and the user state, thereby reducing the overhead of the first processor. And when there are multiple target model tasks in the first task set, the first processor can batch process some or all of the target model tasks in the first task set by triggering the model scheduling module to merge part or all of the multiple target model tasks in the first task set, thereby reducing bandwidth overhead and improving the utilization rate of the second processor.

[0008] In a third aspect, an embodiment of the present application provides an intelligent application task processing method, which is applied to a second processor, wherein the second processor is configured with an operator scheduling module; The method comprises: In the case where the model tasks included in the second task set are split into a plurality of operator tasks based on the model-driven instruction, triggering the operator scheduling module to schedule the plurality of operator tasks to obtain scheduling strategies for the plurality of operator tasks; The operator scheduling module is triggered to trigger the operator of the relevant neural network model to execute the corresponding operator task according to the scheduling strategy, and the result obtained by task processing is fed back to the first processor.

[0009] In an embodiment of the present application, the second processor is configured with an operator scheduling module. When the second processor splits the model tasks included in the second task set into multiple operator tasks based on the model-driven instructions, the second processor schedules the multiple operator tasks by triggering the operator scheduler, and triggers the operators of the relevant neural network model to execute the corresponding operator tasks according to the obtained scheduling strategy. The resources of the second processor can be reasonably allocated to the operator tasks, thereby improving the utilization rate of the second processor.

[0010] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the method described in the second aspect or the third aspect above.

[0011] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method described in the second aspect or the third aspect is implemented.

[0012] In a sixth aspect, an embodiment of the present application provides a computer program product, including a computer program, wherein when the computer program is executed, the method described in the second aspect or the third aspect is executed.

[0013] It can be understood that the beneficial effects of the fourth to sixth aspects can be found in the relevant descriptions of the first, second and third aspects, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0015] Figure 1 It is an example diagram of the operation mode of the intelligent application; Figure 2 This is an example diagram of running multiple intelligent applications simultaneously; Figure 3 is a structural diagram of an intelligent application control system provided by an embodiment of the present application; Figure 4 This is an example diagram of an existing model task in the model queue provided in an embodiment of the present application; Figure 5 The embodiment of the present application provides Figure 4 An example diagram of model task arrangement after each model task is scheduled; Figure 6 This is an example diagram of the actual execution start and end time and execution time limit of each model task provided in the embodiment of the present application; Figure 7 This is an example diagram of an existing operator task in an operator queue provided in an embodiment of the present application; Figure 8 This is an example diagram of a scheduling strategy for operator tasks in an operator queue provided in an embodiment of the present application; Fig. 9 is another structural schematic diagram of the intelligent application control system provided by an embodiment of the present application; Fig.10 It is a flowchart of the intelligent application task processing method provided by an embodiment of the present application; Fig.11 is another flowchart of the intelligent application task processing method provided in an embodiment of the present application; Fig.12 It is a structural diagram of an intelligent application task processing device provided in an embodiment of the present application; Fig.13 is another structural schematic diagram of the intelligent application task processing device provided in an embodiment of the present application; Fig.14 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0016] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0017] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0018] It should be understood that a heterogeneous system usually includes a first processor and at least one second processor. The first processor can be understood as the core of the heterogeneous system, responsible for assigning different tasks to the second processors, coordinating the work between the second processors, and ensuring that the second processors work together to complete complex tasks. As an example and not a limitation, the first processor is a central processing unit (CPU), and the second processor is a neural network processor (NPU), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a graphics processing unit (GPU) and at least one of the multiple processors. NPU is a hardware module specially designed to run neural network models, aiming to improve the processing efficiency and performance of machine learning and deep learning tasks.

[0019] For ease of understanding, this application is described by taking the first processor as a CPU and the second processor as an NPU as an example, but this does not constitute a limitation on the types of the first processor and the second processor.

[0020] At present, the operation mode of intelligent applications is as follows Figure 1As shown, the intelligent application sequentially calls each functional component in the user state, such as component A1, component A2 and component A3, most of which need to call the NPU. The current way to call the NPU is to call the NPU driver running in the kernel state through the NPU API running in the user state to control the NPU. Each time the intelligent application calls the NPU API, a state switch between the user state and the kernel state will occur, thereby generating CPU overhead. In order to reduce the CPU overhead generated by the state switching, the embodiment of the present application provides an intelligent application control system, which includes a first processor and a second processor. The first processor manages the running state of each functional component in the intelligent application through a component scheduling module running in the kernel state, and can send the model task involving the second processor to the second processor through the model scheduling module running in the kernel state to trigger the relevant neural network model to execute the model task, and feedback the result of the task processing to the first processor. In the control process of the above-mentioned intelligent application, there is no need to call the API of the second processor running in the user state, so it can reduce the repeated switching between the kernel state and the user state, and reduce the overhead of the first processor (for example, reduce the CPU overhead). The API of the second processor is a general term for the user-mode program and its interface that controls the operation of the second processor. The API of the second processor can call the driver of the second processor through input / output control (IOCTL) and other methods to control the second processor.

[0021] Figure 1 The operation mode of the intelligent application shown in the figure requires loading the operation instructions and coefficients each time the NPU is called to run the neural network model. The computing unit starts to run only after the operation instructions and coefficients are loaded. The computing unit is not run during the process of loading the operation instructions and the system, which reduces the utilization rate of the NPU. The utilization rate of the NPU can refer to the efficiency of the use of its computing power units when the NPU performs tasks, usually expressed as a percentage. High utilization means that the NPU is performing neural network calculations most of the time, while low utilization indicates that the computing power units of the NPU may be idle. For example, when the NPU is loading operation instructions and coefficients, the computing power units may be idle. Moreover, with the development of electronic devices, there will be more and more demands for running multiple intelligent applications at the same time, such as the requirement to run multiple perimeter protection, multiple target recognition and other intelligent applications at the same time. Figure 2The figure shows an example of running multiple intelligent applications at the same time. Running multiple intelligent applications at the same time not only strengthens the challenge to the utilization rate of the NPU, but also increases the pressure on resource overhead such as the CPU and bandwidth. That is, the NPU needs to spend more time loading operation instructions and coefficients, and the frequent switching between the user state and the kernel state will also significantly increase the CPU overhead. The embodiment of the present application provides a method for processing intelligent application tasks, which is applied to a first processor. The first processor manages the running state of each functional component in the intelligent application through a component scheduling module running in the kernel state, which can reduce the repeated switching between the kernel state and the user state, thereby reducing the overhead of the first processor. And when there are multiple target model tasks in the first task set, the first processor can merge the multiple target model tasks in the first task set partially or completely by triggering the model scheduling module, and can batch process some or all of the target model tasks in the first task set, thereby reducing bandwidth overhead and improving the utilization rate of the second processor. The above process of partially or completely merging multiple target model tasks in the first task set is automatically executed by the model scheduling module at the bottom layer, and each intelligent application does not need to directly intervene or pay attention, and each intelligent application remains decoupled from each other.

[0022] As an example but not a limitation, for N inference tasks corresponding to the same neural network model that are running simultaneously, by partially or completely merging these N inference tasks, these N inference tasks can be batch processed, and the NPU only needs to load the operation instructions and coefficients once, and then perform N inferences, reducing the operation of loading the operation instructions and coefficients N-1 times, so that the computing unit or computing unit of the NPU can have more time for calculation, that is, improve the utilization rate of the NPU, and reduce the number of times the operation instructions and coefficients are read from the memory, that is, reduce the bandwidth overhead. N is an integer greater than 1.

[0023] Among them, the neural network model is a mathematical model or computing model that imitates the structure and function of a biological neural network (brain), which can be used to perform tasks such as target detection and classification. In edge computing, the neural network model usually runs on the NPU, and the neural network model usually consists of multiple operators, each of which is calculated by different sub-intellectual properties (IPs) within the NPU. The sub-IP of the NPU can be understood as the computing power unit of the NPU. An operation instruction may refer to an instruction to perform a specific operation. As an example and not a limitation, the above-mentioned operation instructions may include operation instructions such as addition, subtraction, multiplication, and division, and may also include convolution operation instructions, normalization instructions, matrix operation instructions, etc. This application does not limit the specific types of the above-mentioned operation instructions. The above-mentioned coefficients may refer to weight parameters in the neural network model, which are learned during the model training process and are used to determine the connection strength between neurons.

[0024] In addition, the NPU usually includes multiple computing units for processing different operators in the neural network model. In multi-channel intelligent applications, there are usually multiple operator tasks of multiple neural network models waiting to be executed. The embodiment of the present application schedules multiple operator tasks by triggering the operator scheduler, and triggers the operators of the relevant neural network models to execute the corresponding operator tasks according to the obtained scheduling strategy. It can reasonably allocate NPU resources to operator tasks and improve the utilization rate of the NPU.

[0025] See also Figure 3 , Figure 3 A schematic diagram of the structure of the intelligent application control system provided by the embodiment of the present application is shown. Figure 3 As shown, the intelligent application control system includes: a first processor 301 and a second processor 302 .

[0026] The first processor is configured with: a component scheduling module and a model scheduling module running in kernel state; The component scheduling module is used to manage the running status of each functional component in the intelligent application, and send a first task set to the model scheduling module when the functional component is running, the first task set includes: model tasks triggered when the functional component is running, and the model tasks are tasks that depend on the execution of the neural network model; The model scheduling module is used to: when there are multiple target model tasks in the first task set, partially or completely merge the multiple target model tasks in the first task set to obtain a second task set that meets the batch processing condition, wherein the multiple target model tasks are model tasks corresponding to the same neural network model, and the batch processing condition includes: the execution time of each model task in the second task set does not exceed the execution time limit corresponding to each model task; based on the model tasks included in the second task set, send a model driving instruction to the second processor; The second processor is used to: trigger the relevant neural network model to execute the model tasks in the second task set based on the model-driven instruction, and feed back the results of the task processing to the first processor.

[0027] Among them, intelligent applications are functions with practical uses, such as target recognition systems, perimeter defense systems, etc. An intelligent application consists of several components, such as components that implement detection, classification, etc. In intelligent applications, these components are usually completed by neural network models.

[0028] The target recognition system mainly authenticates or identifies the target by analyzing and comparing the visual feature information of the target. It is widely used in security systems, mobile device unlocking, monitoring systems, social media tagging and other scenarios.

[0029] Perimeter protection systems are protective measures set up around a specific area or facility to prevent unauthorized people or objects from entering or leaving the area. Such protection measures usually include physical barriers and electronic monitoring systems, aiming to provide multi-level security. Perimeter protection systems can be applied in a variety of occasions, such as residential areas, commercial facilities, industrial sites, some special bases, etc.

[0030] The functional components in the intelligent application may refer to the functional modules constituting the intelligent application. Taking the perimeter protection system as an example, the perimeter protection system may include a functional component for human body detection, a functional component for face detection, a functional component for classification, and the like.

[0031] The running state of the functional component includes but is not limited to the creation, startup, operation, termination and destruction of the functional component. In this embodiment, each functional component runs in the kernel state, and the running state of each functional component can be managed by the component scheduling module running in the kernel state. On this basis, the model task is executed, which can reduce the repeated switching between the kernel state and the user state and reduce the CPU overhead. Among them, the user state is an execution mode of the operating system. The program in the user state is subject to strict permission restrictions and cannot directly access hardware resources or execute privileged instructions. The kernel state is an execution mode of the operating system. The program in the kernel state has the highest authority and can access all hardware resources and execute all privileged instructions.

[0032] The first processor can run one intelligent application or multiple intelligent applications at the same time. When multiple intelligent applications are run at the same time, the component scheduling module can send the corresponding model task to the model scheduling module when any functional component in the multiple intelligent applications is running. Among them, the multi-channel intelligent application can be a variety of different intelligent applications, or can be multiple copies of the same intelligent application from different sources.

[0033] As an example and not a limitation, IP cameras (IP CAMERA, IPC), network video recorders (Network Video Recorder, NVR), mobile phones, tablet computers, IoT devices and other electronic devices can run a variety of different smart applications at the same time, or they can run multiple copies of the same smart application at the same time, but the video sources processed by these multiple smart applications are different. IPC is a new generation of cameras that combines traditional cameras with network technology. It can transmit video images to a remote end through the network, and the remote viewer can view its video images through a web browser. NVR is the storage and forwarding part of the network video surveillance system. NVR works with the video encoder or network camera to complete the video recording, storage and forwarding functions.

[0034] It should be understood that in the case of multi-channel intelligent applications, there are cases where different intelligent applications use the same neural network model. For example, if 8 perimeter protection systems are running simultaneously in an NVR, the neural network models used by these 8 perimeter protection systems are usually the same.

[0035] It should be understood that the first processor and the second processor in the intelligent application control system can be deployed on the same electronic device or on different devices, and this application does not limit this. As an example but not a limitation, the first processor is deployed on the electronic device and the second processor is deployed on the server, or the first processor is deployed on one electronic device and the second processor is deployed on another electronic device.

[0036] In one embodiment, the component scheduling module may add the model task to the model queue when each functional component triggers the corresponding model task, and the model scheduling module obtains the first task set from the model queue.

[0037] In one embodiment, the model scheduling module may schedule the model tasks in the model queue each time a new model task is queued in the model queue.

[0038] In one embodiment, the model scheduling module may send the model tasks to the second processor in sequence according to the execution order of each model task in the second task set. Optionally, multiple model tasks may be sent to the second processor in sequence until the number of sent model tasks reaches a task number threshold of the second processor. The task number threshold is used to limit the number of model tasks that the second processor can execute simultaneously.

[0039] In one embodiment, the model scheduling module can send a model driving instruction to the second processor through a driver of the second processor running in kernel state, and the model driving instruction carries a second task set. The driver of the second processor is a program that controls the operation of the second processor, and generally includes functions such as configuring registers, controlling the startup of the second processor, and receiving interrupt signals of the second processor.

[0040] The target model tasks in the above-mentioned multiple target model tasks may be model tasks corresponding to different functional components of one intelligent application, or model tasks corresponding to different functional components of multiple intelligent applications. The model tasks corresponding to the functional components refer to the model tasks triggered when the functional components are running.

[0041] As an example but not limitation, intelligent application 1 includes functional component 11, functional component 12 and functional component 13, the model task corresponding to functional component 11 and the model task corresponding to functional component 13 correspond to the same neural network model 1 and these two functional components can be executed simultaneously, then when running intelligent application 1, the model task corresponding to functional component 11 and the model task corresponding to functional component 13 can be determined as the target model task. Intelligent application 2 includes functional component 21, functional component 22 and functional component 23, the model task corresponding to functional component 22 and the model task corresponding to functional component 12 in intelligent application 1 correspond to the same neural network model 2, then when running intelligent application 1 and intelligent application 2 simultaneously, the model task corresponding to functional component 22 and the model task corresponding to functional component 12 can be determined as the target model task.

[0042] After partially or completely merging multiple target model tasks, batch tasks can be obtained, and the second task set includes the batch tasks. If the first task set also includes remaining tasks, it is determined that the second task set also includes remaining tasks. The remaining tasks may refer to model tasks in the first task set other than the target model tasks to be merged.

[0043] like Figure 4 The figure shows an example of the model tasks in the model queue. Assume that the current time is 7, when model task 5 enters the model queue. The model queue has received five model tasks, namely model task 1, model task 2, model task 3, model task 4 and model task 5, from the component scheduling module in sequence, waiting for execution. Model task 2 and model task 4 correspond to the same neural network model. Then the model scheduling module can merge model task 2 and model task 4 together for batch processing while ensuring that the execution time of each model task does not exceed the corresponding execution time limit.

[0044] When there are multiple target model tasks in the first task set, the task merging unit can enable the second processor to batch process the batch tasks in the second task set by partially or fully merging the multiple target model tasks in the first task set, thereby reducing bandwidth overhead, improving the utilization rate of the second processor, and increasing the number of intelligent applications.

[0045] In one possible implementation, the model scheduling module includes a task scheduling unit, a sequential execution unit, and a target determination unit; The task scheduling unit is used to: schedule each model task in the first task set based on the priority, execution time limit, estimated running time and queue entry time of each model task in the first task set, obtain a first execution order of each model task in the first task set, and store each model task in the first task set in a model queue; The sequential execution unit is used to: determine a second execution order of the plurality of target model tasks based on the first execution order when there are a plurality of target model tasks in the first task set; The target determination unit is used to partially or completely merge the multiple target model tasks based on the second execution order and the estimated running time of the multiple target model tasks.

[0046] The execution time limit of the model task may refer to the limit set on the execution time of the model task during the execution of the model task.

[0047] This embodiment schedules each model task in the first task set based on the priority, execution time limit, estimated running time and queue entry time of each model task in the first task set, thereby ensuring that important tasks can be executed in a timely manner.

[0048] The first execution order is the execution order of each model task in the first task set. The multiple target model tasks are model tasks corresponding to the same neural network in the first task set. Therefore, the sequential execution unit can determine the execution order of the multiple target model tasks (i.e., the second execution order) based on the first execution order. The target determination unit partially or fully merges the multiple target model tasks based on the second execution order and the estimated running time of the multiple target model tasks, thereby ensuring that the execution time of each model task in the second task set does not exceed the corresponding execution time limit of each model task.

[0049] As an example but not limitation, the first execution order of model task 1, model task 2, model task 3, model task 4 and model task 5 is model task 1, model task 5, model task 4, model task 2 and model task 3. If after merging model task 2 and model task 4, it can be ensured that the execution time of model task 1, model task 5, batch task and model task 3 does not exceed the corresponding execution time limit, and model task 2 and model task 4 correspond to the model of the same neural network, then model task 2 and model task 4 can be merged. Therefore, the execution order of the model tasks in the model queue after the model scheduling module schedules them is model task 1, model task 5, batch task and model task 3.

[0050] It should be understood that the estimated running time of multiple target model tasks corresponding to the same neural network model is usually the same. For a given neural network model and NPU, the running time of this neural network model on this NPU can be theoretically calculated, so the estimated running time can be set according to the theoretical value.

[0051] In one embodiment, before scheduling each model task in the first task set based on the priority, execution deadline, estimated running time and queue entry time of each model task in the first task set, the task scheduling unit may first determine whether there is a model task in the first task set whose execution deadline and the time interval between the current moment are less than the first preset time length. If there is a model task in the first task set whose execution deadline and the time interval between the current moment are less than the first preset time length, the priority of the model task may be increased, that is, as time goes by, the priority of the model task approaching the execution deadline is gradually increased. Optionally, for any model task, a first preset time length may be set based on the estimated running time of the model task, such as setting the first preset time length to be slightly greater than or equal to the estimated running time of the model task, for example, the sum of the estimated running time of the model task and the preset first target value is determined as the first preset time length.

[0052] Optionally, when the first preset duration is greater than the estimated running time of the corresponding model task, for any model task that needs to increase its priority, if the time interval between the execution deadline of the model task and the current moment is less than the first preset duration and is greater than or equal to the estimated running time of the model task, the priority of the model task is increased by P levels, where P is an integer greater than zero; if the time interval between the execution deadline of the model task and the current moment is less than the estimated running time, the priority of the model task is increased by M levels, where M is an integer greater than P.

[0053] In a possible implementation, the task scheduling unit is specifically used to: The model tasks in the first task set are sorted in descending order of priority, ascending order of execution time limit, ascending order of estimated running time, and ascending order of time of entering the queue.

[0054] Among them, the priority descending order may refer to the order of priority from high to low. The execution time ascending order may refer to the order of execution time from first to last. The estimated running time ascending order may refer to the order of estimated running time from short to long. The queue entry time ascending order may refer to the order of queue entry time from first to last.

[0055] Specifically, the task scheduling module can sort the model tasks in the first task set in descending order of priority; when there are multiple model tasks with the same priority in the first task set, sort the multiple model tasks with the same priority in ascending order of execution time limit; when there are not multiple model tasks with the same priority in the first task set, it means that the sorting of the model tasks in the first task set can be completed in descending order of priority, and there is no need to sort them in ascending order of execution time limit, ascending order of estimated running time and ascending order of time of entering the queue. When there are multiple model tasks with the same execution time limit among multiple model tasks with the same priority, sort the multiple model tasks with the same execution time limit in ascending order of estimated running time; when there are not multiple model tasks with the same execution time limit among multiple model tasks with the same priority, it means that the sorting of the model tasks in the first task set can be completed in descending order of priority and ascending order of execution time limit, and there is no need to sort them in ascending order of estimated running time and ascending order of time of entering the queue. When there are multiple model tasks with the same estimated running time among multiple model tasks with the same execution time (that is, multiple model tasks with the same execution time among multiple model tasks with the same priority), the multiple model tasks with the same estimated running time are sorted in ascending order of the time of joining the queue; when there are not multiple model tasks with the same estimated running time among multiple model tasks with the same execution time, it means that the model tasks in the first task set can be sorted in descending order of priority, ascending order of execution time and ascending order of estimated running time, and there is no need to sort in ascending order of the time of joining the queue.

[0056] This embodiment sorts the model tasks in the first task set in descending order of priority, ascending order of execution deadline, ascending order of estimated running time, and ascending order of time to join the queue, thereby ensuring that important tasks with higher priority, tighter execution deadline, shorter estimated running time, and earlier time to join the queue can be executed in time.

[0057] It should be understood that when sorting the model tasks in the first task set, the execution order of descending priority, ascending execution time, ascending estimated running time and ascending entry time can also be set according to scenario requirements. This application does not limit this.

[0058] In a possible implementation manner, the target determination unit is specifically configured to: The first model task in the second execution order is determined as the first model task, and the target model task next to the first model task in the second execution order is determined as the second model task; the first model task is the target model task that is in the first position in the second execution order; Based on the estimated running time of the first model task and the estimated running time of the second model task, determine whether the execution time of the third model task exceeds the corresponding execution time limit; the third model task is a model task that is located after the first model task in the first execution order and is other than the first model task and the second model task; When the execution time of the third model task does not exceed the corresponding execution time limit, the first model task and the second model task are merged to obtain a batch processing task; Determine the batch task as the first model task, determine the target model task that is second only to the second model task in the second execution order as the second model task, return the estimated running time based on the first model task and the estimated running time of the second model task, and determine whether the execution time of the third model task exceeds the corresponding execution time limit and subsequent steps until the execution time of the third model task exceeds the corresponding execution time limit or multiple target model tasks are traversed.

[0059] The execution time of the first model task can be determined based on the estimated running time of the model task that precedes the first model task in the first execution order. The execution time of the third model task can be determined based on the execution time of the first model task, the estimated running time of the first model task, and the estimated running time of the second model task, so as to judge whether the execution time of the third model task exceeds the corresponding execution time limit. If the execution time of the third model task does not exceed the corresponding execution time limit, the first model task can be merged with the second model task, and the traversal of multiple target model tasks can continue. If the execution time of the third model task does not exceed the corresponding execution time limit, it is determined that the first model task and the second model task cannot be merged, and the traversal of multiple target model tasks is stopped.

[0060] In a possible implementation, the model scheduling module further includes a time determination unit; The time determination unit is used to determine the execution time of the first model task as the execution time of the batch processing task.

[0061] This embodiment determines the execution time of the first model task as the execution time of the batch task, so that other model tasks in the batch task can be batch processed together with the first model task. Other model tasks refer to target model tasks other than the first model task in the batch task.

[0062] like Figure 5 The shown is Figure 4 The following is an example of the model task arrangement after the model tasks in are scheduled. Assume that the current time is 7, at which time model task 5 enters the model queue. The model queue contains model tasks 1, 2, 3, 4, and 5. Figure 4If the model tasks 2 and 4 in the example are executed separately, the operation instructions and coefficients of the corresponding neural network models need to be loaded twice. Figure 5 The batch processing task in only needs to load the operation instructions and coefficients once, which saves the loading process of the operation instructions and coefficients of the corresponding neural network model, and can make the running time of the batch processing task less than the total running time of model task 2 and model task 4. Figure 4 The estimated running time of model task 2 and model task 4 is 4 seconds each, so the total running time of these two model tasks is 8 seconds. Figure 5 The batch processing task takes 6 seconds to run.

[0063] like Figure 6 The following is an example of the actual execution start and end time and the required execution time limit of each model task after scheduling. Figure 6 The actual execution start time of the model task is the execution time of the model task. Figure 6 It can be seen that the execution time of each model task in the second task set does not exceed the corresponding execution time limit.

[0064] In one embodiment, the first processor is further configured with an application scheduling module. The application scheduling module runs in user mode, is responsible for the creation, startup, termination and destruction of each intelligent application, and does not participate in the scheduling of functional components during the operation of the intelligent application.

[0065] In a possible implementation, the second processor is configured with an operator scheduling module; The operator scheduling module is used to: when the model tasks included in the second task set are split into multiple operator tasks based on the model-driven instructions, schedule the multiple operator tasks, obtain the scheduling strategies of the multiple operator tasks, trigger the operators of the relevant neural network model to execute the corresponding operator tasks according to the scheduling strategies, and feed back the results of task processing to the first processor.

[0066] The operator scheduling module schedules multiple operator tasks and triggers the operators of the relevant neural network model to execute corresponding operator tasks according to the obtained scheduling strategy. It can reasonably allocate the resources of the second processor to the operator tasks and improve the utilization rate of the second processor.

[0067] In a possible implementation, the operator scheduling module includes a parallel determination unit, a first processing unit, and a second processing unit; The parallel determination unit is used to: when there are multiple operator task teams that need to be executed in parallel among the multiple operator tasks and the computing power units corresponding to the operator tasks in the multiple operator task teams are in an idle state, determine the parallel execution as a scheduling strategy for the multiple operator task teams; each operator task team includes at least one operator task; The first processing unit is used for: for an operator task team including more than two operator tasks, scheduling each operator task in the operator task team based on the priority, execution time limit, estimated running time and entry time of each operator task in the operator task team, obtaining the execution order of each operator task in the operator task team, and determining the execution order of each operator task in the operator task team as the scheduling strategy of each operator task in the operator task team, and storing each operator task in the operator task team in the operator queue; The second processing unit is used to: when there is no operator task queue among the multiple operator tasks and the computing power units corresponding to the multiple operator tasks are in an idle state, schedule the multiple operator tasks based on the priorities, execution deadlines, estimated running times and queue entry times of the multiple operator tasks, obtain the execution order of the multiple operator tasks, and determine the execution order of the multiple operator tasks as the scheduling strategy of the multiple operator tasks, and the multiple operator tasks are stored in the operator queue.

[0068] The operators of the neural network model usually run on the computing unit of the second processor, and the multiple operator task teams executed in parallel correspond to different computing units. The operator tasks in the same operator task team can be related operator tasks or operator tasks corresponding to the same computing unit.

[0069] The execution time limit of an operator task may refer to a limit set on the execution time of the operator task during the execution of the operator task.

[0070] In one embodiment, before scheduling each operator task based on the priority, execution deadline, estimated running time and queue entry time of each operator task, the first processing unit or the second processing unit may first determine whether there is an operator task whose execution deadline and the time interval between the current moment are less than the second preset time length among each operator task. If there is an operator task whose execution deadline and the time interval between the current moment are less than the second preset time length among each operator task, the priority of the operator task may be increased, that is, as time goes by, the priority of the operator task approaching the execution deadline is gradually increased. Optionally, for any operator task, a second preset time length may be set based on the estimated running time length of the operator task, such as setting the second preset time length to be slightly greater than or equal to the estimated running time length of the operator task, such as determining the sum of the estimated running time length of the operator task and the preset second target value as the second preset time length.

[0071] Optionally, when the second preset duration is greater than the estimated running time of the corresponding operator task, for any operator task that needs to increase its priority, if the time interval between the execution time limit of the operator task and the current moment is less than the second preset duration and greater than or equal to the estimated running time of the operator task, the priority of the operator task is increased by L levels, where L is an integer greater than zero; if the time interval between the execution time limit of the operator task and the current moment is less than the estimated running time, the priority of the operator task is increased by S levels, where S is an integer greater than L. This embodiment schedules each operator task according to its priority, execution time limit, estimated running time, and time of entering the queue, thereby ensuring that important tasks can be executed in a timely manner, and by executing multiple operator task teams in parallel, the total time required to complete the operator task can be reduced, thereby improving the utilization rate of the second processor.

[0072] In one embodiment, after splitting each model task into multiple operator tasks, the second processor can add the operator task to the operator queue. The operator scheduling module schedules each operator task in the operator queue and determines the execution order or parallel execution of each operator task based on the priority, execution time limit, estimated running time, entry time and idle status of the computing unit in the second processor.

[0073] In a possible implementation, the first processing unit is specifically used to: sort the operator tasks in the operator task team in descending order of priority, ascending order of execution time limit, ascending order of estimated running time, and ascending order of time of entering the team; The second processing unit is specifically used to sort the multiple operator tasks in descending order of priority, ascending order of execution time limit, ascending order of estimated running time and ascending order of entry time.

[0074] Specifically, the first processing unit can sort the operator tasks in the operator task team in descending order of priority; when there are multiple operator tasks with the same priority in the operator task team, the multiple operator tasks with the same priority are sorted in ascending order of execution time limit; when there are not multiple operator tasks with the same priority in the operator task team, it means that the sorting of the operator tasks in the operator task team can be completed in descending order of priority, and there is no need to sort them in ascending order of execution time limit, ascending order of estimated running time and ascending order of time of entering the team. When there are multiple operator tasks with the same execution time limit among multiple operator tasks with the same priority, the multiple operator tasks with the same execution time limit are sorted in ascending order of estimated running time; when there are not multiple operator tasks with the same execution time limit among multiple operator tasks with the same priority, it means that the sorting of the operator tasks in the operator task team can be completed in descending order of priority and ascending order of execution time limit, and there is no need to sort them in ascending order of estimated running time and ascending order of time of entering the team. When there are multiple operator tasks with the same estimated running time among multiple operator tasks with the same execution time (that is, multiple operator tasks with the same execution time among multiple operator tasks with the same priority), the multiple operator tasks with the same estimated running time are sorted in ascending order of the time they enter the queue; when there are not multiple operator tasks with the same estimated running time among multiple operator tasks with the same execution time, it means that the sorting of the operator tasks in the operator task team can be completed in descending order of priority, ascending order of execution time, and ascending order of estimated running time, and there is no need to sort them in ascending order of the time they enter the queue.

[0075] The second processing unit can sort each of the multiple operator tasks in descending order of priority; in the case where there are multiple operator tasks with the same priority among the multiple operator tasks, the multiple operator tasks with the same priority are sorted in ascending order of execution time limit; in the case where there are not multiple operator tasks with the same priority among the multiple operator tasks, it means that the sorting of each of the multiple operator tasks can be completed in descending order of priority, and there is no need to sort them in ascending order of execution time limit, ascending order of estimated running time and ascending order of time of entering the queue. In the case where there are multiple operator tasks with the same execution time limit among the multiple operator tasks with the same priority, the multiple operator tasks with the same execution time limit are sorted in ascending order of estimated running time; in the case where there are not multiple operator tasks with the same execution time limit among the multiple operator tasks with the same priority, it means that the sorting of each of the multiple operator tasks can be completed in descending order of priority and ascending order of execution time limit, and there is no need to sort them in ascending order of estimated running time and ascending order of time of entering the queue. When there are multiple operator tasks with the same estimated running time among multiple operator tasks with the same execution time limit (that is, multiple operator tasks with the same execution time limit among multiple operator tasks with the same priority), the multiple operator tasks with the same estimated running time are sorted in ascending order of the time they enter the queue; when there are no multiple operator tasks with the same estimated running time among multiple operator tasks with the same execution time limit, it means that the sorting of each operator task in the multiple operator tasks can be completed in descending order of priority, ascending order of execution time limit and ascending order of estimated running time, and there is no need to sort them in ascending order of the time they enter the queue.

[0076] This embodiment sorts the operator tasks in descending order of priority, ascending order of execution time, ascending order of estimated running time and ascending order of time of joining the queue, thereby ensuring that important tasks with higher priority, tighter execution time, shorter estimated running time and earlier time of joining the queue can be executed in time.

[0077] It should be understood that when the first processing unit sorts the operator tasks in the operator task team and the second processing unit sorts multiple operator tasks, the execution order of descending priority, ascending execution time, ascending estimated running time and ascending entry time can be set according to scenario requirements. This application does not limit this.

[0078] like Figure 7 The following is an example of an operator task in the operator queue. Figure 7The operator queue in includes operator task 1, operator task 2, operator task 3, operator task 4 and operator task 5. The operators corresponding to operator task 1, operator task 2 and operator task 4 are all running on computing unit 1. The operators corresponding to operator task 3 and operator task 5 are all running on computing unit 2. And computing unit 1 and computing unit 2 are in idle state. Therefore, operator task team 1 including operator task 1, operator task 2 and operator task 4 and operator task team 2 including operator task 3 and operator task 5 can be executed in parallel. For operator task team 1, operator task 1 has the highest priority, and operator task 2 and operator task 4 have the same priority, but the execution deadline of operator task 2 is earlier than the execution deadline of operator task 4, so if Figure 8 The execution order of each operator task in operator task team 1 is shown as operator task 1, operator task 2 and operator task 4. For operator task team 2, the priority of operator task 5 is higher than that of operator task 3, so Figure 8 The execution order of the operator tasks in the operator task team 2 is shown as operator task 5 and operator task .

[0079] like Fig. 9 FIG. 1 is another schematic diagram of the structure of the intelligent application control system provided by an embodiment of the present application. Fig. 9 The multi-level scheduling modules such as application scheduling module, component scheduling module, model scheduling module and operator scheduling module can realize the control of intelligent applications, reduce CPU overhead and bandwidth overhead, improve NPU utilization, and thus increase the number of intelligent applications.

[0080] See also Fig.10 , Fig.10 A flow chart of a method for processing intelligent application tasks provided by an embodiment of the present application is shown. As an example but not a limitation, the method is applied to a first processor, and the first processor is configured with: a component scheduling module and a model scheduling module running in kernel state; the component scheduling module is used to manage the running status of each functional component in the intelligent application. The method includes the following steps: Step 1001, when the functional component is running, the component scheduling module is triggered to send a first task set to the model scheduling module.

[0081] Among them, the first task set includes: model tasks triggered when the functional components are running, and the model tasks are tasks that depend on the execution of the neural network model.

[0082] Step 1002, when there are multiple target model tasks in the first task set, trigger the model scheduling module to partially or completely merge the multiple target model tasks in the first task set to obtain a second task set that meets the batch processing condition.

[0083] Among them, the multiple target model tasks are model tasks corresponding to the same neural network model, and the batch processing conditions include: the execution time of each model task in the second task set does not exceed the execution time limit corresponding to each model task.

[0084] Step 1003, triggering the model scheduling module to send a model driving instruction to the second processor based on the tasks included in the second task set, the model driving instruction is used to trigger the second processor to execute the tasks in the second task set and feed back the results of task processing to the first processor.

[0085] The second processor is configured with task processing capability based on a neural network model.

[0086] In a possible implementation, before the model scheduling module is triggered to partially or completely merge the multiple target model tasks in the first task set, the method further includes: The trigger model scheduling module schedules each model task in the first task set based on the priority, execution time limit, estimated running time and queue entry time of each model task in the first task set to obtain a first execution order of each model task in the first task set. Each model task in the first task set is stored in the model queue; The triggering model scheduling module determines a second execution order of the plurality of target model tasks based on the first execution order when there are a plurality of target model tasks in the first task set; The triggering model scheduling module partially or completely merges multiple target model tasks in the first task set, including: The trigger model scheduling module partially or completely merges the multiple target model tasks based on the second execution order and the estimated running time of the multiple target model tasks.

[0087] In a possible implementation, the triggering model scheduling module schedules each model task in the first task set based on the priority, execution time limit, estimated running time and queue entry time of each model task in the first task set, including: The trigger model scheduling module sorts the model tasks in the first task set in descending order of priority, ascending order of execution time limit, ascending order of estimated running time and ascending order of entry time.

[0088] In a possible implementation, the trigger model scheduling module partially or completely merges the multiple target model tasks based on the second execution order and the estimated running time of the multiple target model tasks, including: The trigger model scheduling module determines the first model task in the second execution order as the first model task, and determines the target model task next to the first model task in the second execution order as the second model task; the first model task is the target model task that is first in the second execution order; The trigger model scheduling module determines whether the execution time of the third model task exceeds the corresponding execution time limit based on the estimated running time of the first model task and the estimated running time of the second model task; the third model task is a model task that is located after the first model task in the first execution order and is other than the first model task and the second model task; The triggering model scheduling module merges the first model task with the second model task to obtain a batch processing task when the execution time of the third model task does not exceed the corresponding execution time limit; The trigger model scheduling module determines the batch task as the first model task, determines the target model task that is second only to the second model task in the second execution order as the second model task, and returns the steps of executing the estimated running time based on the first model task and the estimated running time of the second model task, and judging whether the execution time of the third model task exceeds the corresponding execution time limit and subsequent steps until the execution time of the third model task exceeds the corresponding execution time limit or multiple target model tasks are traversed.

[0089] In a possible implementation, after obtaining the batch processing task, the following is further included: The execution time of the first model task is determined as the execution time of the batch task.

[0090] In the embodiment of the present application, the first processor manages the running state of each functional component in the intelligent application through the component scheduling module running in the kernel state, which can reduce the repeated switching between the kernel state and the user state, thereby reducing the overhead of the first processor. And when there are multiple target model tasks in the first task set, the first processor can batch process some or all of the target model tasks in the first task set by triggering the model scheduling module to merge part or all of the multiple target model tasks in the first task set, thereby reducing bandwidth overhead and improving the utilization rate of the second processor.

[0091] It should be noted that Fig.10 The method embodiment shown is Figure 2 The function of the first processor in the control system embodiment shown is based on the same concept. The relevant contents of the method embodiment can be found in the system embodiment part, which will not be repeated here.

[0092] See also Fig.11 , Fig.11 Another flow chart of the intelligent application task processing method provided in the embodiment of the present application is shown. As an example but not a limitation, the method is applied to a second processor, and the second processor is configured with an operator scheduling module. The method comprises the following steps: Step 1101 , when the model task included in the second task set is split into multiple operator tasks based on the model driving instruction, triggering the operator scheduling module to schedule the multiple operator tasks to obtain the scheduling strategies of the multiple operator tasks.

[0093] Step 1102, triggering the operator scheduling module to trigger the operator of the relevant neural network model to execute the corresponding operator task according to the scheduling strategy, and feeding back the result of the task processing to the first processor.

[0094] In a possible implementation, the operator scheduling module is triggered to schedule multiple operator tasks, and the scheduling strategies of the multiple operator tasks are obtained, including: When there are multiple operator task teams that need to be executed in parallel among the multiple operator tasks and the computing power units corresponding to the operator tasks in the multiple operator task teams are in an idle state, the operator scheduling module is triggered to determine the parallel execution as a scheduling strategy for the multiple operator task teams; each operator task team includes at least one operator task; For an operator task team including more than two operator tasks, the operator scheduling module is triggered to schedule each operator task in the operator task team based on the priority, execution time limit, estimated running time and entry time of each operator task in the operator task team, obtain the execution order of each operator task in the operator task team, and determine the execution order of each operator task in the operator task team as the scheduling strategy of each operator task in the operator task team. Each operator task in the operator task team is stored in the operator queue; When there is no operator task queue among multiple operator tasks and the computing power units corresponding to the multiple operator tasks are in an idle state, the operator scheduling module is triggered to schedule the multiple operator tasks based on the priorities, execution deadlines, estimated running times and queue entry times of the multiple operator tasks, obtain the execution order of the multiple operator tasks, and determine the execution order of the multiple operator tasks as the scheduling strategy of the multiple operator tasks, and the multiple operator tasks are stored in the operator queue; The computing power unit is used to run the operator that executes the corresponding operator task.

[0095] In a possible implementation, the triggering operator scheduling module schedules each operator task in the operator task team based on the priority, execution time limit, estimated running time and entry time of each operator task in the operator task team, including: The trigger operator scheduling module sorts the operator tasks in the operator task queue in descending order of priority, ascending order of execution time limit, ascending order of estimated running time, and ascending order of entry time; The trigger operator scheduling module schedules multiple operator tasks based on their priorities, execution deadlines, estimated running times, and queue entry times, including: The trigger operator scheduling module sorts multiple operator tasks in descending order of priority, ascending order of execution time limit, ascending order of estimated running time, and ascending order of enqueue time.

[0096] In an embodiment of the present application, the second processor is configured with an operator scheduling module. When the second processor splits the model tasks included in the second task set into multiple operator tasks based on the model-driven instructions, the second processor schedules the multiple operator tasks by triggering the operator scheduler, and triggers the operators of the relevant neural network model to execute the corresponding operator tasks according to the obtained scheduling strategy. The resources of the second processor can be reasonably allocated to the operator tasks, thereby improving the utilization rate of the second processor.

[0097] It should be noted that Fig.11 The method embodiment shown is Figure 2 The function of the second processor in the control system embodiment shown is based on the same concept. For the relevant contents of the method embodiment, please refer to the system embodiment part, which will not be repeated here.

[0098] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0099] Corresponding to the above Fig.10 The intelligent application task processing method described in the embodiment, Fig.12 A schematic diagram of the structure of an intelligent application task processing device provided by an embodiment of the present application is shown, the device is applied to a first processor, and the first processor is configured with: a component scheduling module and a model scheduling module running in kernel state; the component scheduling module is used to manage the running status of each functional component in the intelligent application. For ease of explanation, only the part related to the embodiment of the present application is shown.

[0100] Reference Fig.12 , the device comprises: The first model triggering module 1201 is used to trigger the component scheduling module to send a first task set to the model scheduling module when the functional component is running, and the first task set includes: model tasks triggered when the functional component is running, and the model tasks are tasks that depend on the execution of the neural network model; The second model triggering module 1202 is used to trigger the model scheduling module to partially or completely merge the multiple target model tasks in the first task set to obtain a second task set that meets the batch processing condition when there are multiple target model tasks in the first task set, and the multiple target model tasks are model tasks corresponding to the same neural network model. The batch processing condition includes: the execution time of each model task in the second task set does not exceed the execution time limit corresponding to each model task; The third model triggering module 1203 is used to trigger the model scheduling module to send a model-driven instruction to the second processor based on the tasks included in the second task set. The model-driven instruction is used to trigger the second processor to execute the tasks in the second task set and feedback the results of task processing to the first processor. The second processor is configured with task processing capabilities based on a neural network model.

[0101] Optionally, the above device further comprises: The fourth model triggering module is used to trigger the model scheduling module to schedule each model task in the first task set based on the priority, execution time limit, estimated running time and queue entry time of each model task in the first task set, and obtain the first execution order of each model task in the first task set. Each model task in the first task set is stored in the model queue; A fifth model triggering module, configured to trigger the model scheduling module to determine a second execution order of the plurality of target model tasks based on the first execution order when there are a plurality of target model tasks in the first task set; The second model triggering module 1202 is specifically used for: The trigger model scheduling module partially or completely merges the multiple target model tasks based on the second execution order and the estimated running time of the multiple target model tasks.

[0102] Optionally, the fourth model triggering module is specifically used for: The trigger model scheduling module sorts the model tasks in the first task set in descending order of priority, ascending order of execution time limit, ascending order of estimated running time and ascending order of entry time.

[0103] Optionally, the second model triggering module 1202 is specifically used for: The trigger model scheduling module determines the first model task in the second execution order as the first model task, and determines the target model task next to the first model task in the second execution order as the second model task; the first model task is the target model task that is first in the second execution order; The trigger model scheduling module determines whether the execution time of the third model task exceeds the corresponding execution time limit based on the estimated running time of the first model task and the estimated running time of the second model task; the third model task is a model task that is located after the first model task in the first execution order and is other than the first model task and the second model task; The triggering model scheduling module merges the first model task with the second model task to obtain a batch processing task when the execution time of the third model task does not exceed the corresponding execution time limit; The trigger model scheduling module determines the batch task as the first model task, determines the target model task that is second only to the second model task in the second execution order as the second model task, and returns the steps of executing the estimated running time based on the first model task and the estimated running time of the second model task, and judging whether the execution time of the third model task exceeds the corresponding execution time limit and subsequent steps until the execution time of the third model task exceeds the corresponding execution time limit or multiple target model tasks are traversed.

[0104] Optionally, the second model triggering module 1202 is further used for: The execution time of the first model task is determined as the execution time of the batch task.

[0105] It should be noted that the information interaction and execution process between the above modules are Figure 2 The function of the first processor in the control system embodiment shown is based on the same concept. Its specific functions and technical effects can be found in the system embodiment section and will not be repeated here.

[0106] Corresponding to the above Fig.11 The intelligent application task processing method of the embodiment, Fig.13 Another structural diagram of the intelligent application task processing device provided in the embodiment of the present application is shown, and the device is applied to the second processor, and the second processor is configured with an operator scheduling module. For ease of description, only the part related to the embodiment of the present application is shown.

[0107] Reference Fig.13 , the device comprises: The first operator triggering module 1301 is used to trigger the operator scheduling module to schedule the multiple operator tasks and obtain the scheduling strategies of the multiple operator tasks when the model task included in the second task set is split into multiple operator tasks based on the model driving instruction; The second operator triggering module 1302 is used to trigger the operator scheduling module to trigger the operator of the relevant neural network model to execute the corresponding operator task according to the scheduling strategy, and feed back the result of task processing to the first processor.

[0108] Optionally, the first operator triggering module 1301 includes: A first triggering unit is used to trigger the operator scheduling module to determine parallel execution as a scheduling strategy for multiple operator task teams when there are multiple operator task teams that need to be executed in parallel among the multiple operator tasks and the computing power units corresponding to the operator tasks in the multiple operator task teams are in an idle state, each of which includes at least one operator task; The second triggering unit is used for triggering the operator scheduling module to schedule each operator task in the operator task team based on the priority, execution time limit, estimated running time and entry time of each operator task in the operator task team for an operator task team including more than two operator tasks, obtain the execution order of each operator task in the operator task team, and determine the execution order of each operator task in the operator task team as the scheduling strategy of each operator task in the operator task team, and each operator task in the operator task team is stored in the operator queue; The third triggering unit is used to trigger the operator scheduling module to schedule the multiple operator tasks based on the priorities, execution time limits, estimated running times and queue entry times of the multiple operator tasks when there is no operator task queue among the multiple operator tasks and the computing power units corresponding to the multiple operator tasks are in an idle state, obtain the execution order of the multiple operator tasks, and determine the execution order of the multiple operator tasks as the scheduling strategy of the multiple operator tasks, and store the multiple operator tasks in the operator queue; The computing power unit is used to run the operator that executes the corresponding operator task.

[0109] Optionally, the second trigger unit is specifically used for: The trigger operator scheduling module sorts the operator tasks in the operator task queue in descending order of priority, ascending order of execution time limit, ascending order of estimated running time, and ascending order of entry time; Optionally, the third trigger unit is specifically used for: The trigger operator scheduling module sorts multiple operator tasks in descending order of priority, ascending order of execution time limit, ascending order of estimated running time, and ascending order of enqueue time.

[0110] It should be noted that the information interaction and execution process between the above modules are Figure 2 The function of the second processor in the control system embodiment shown is based on the same concept. Its specific functions and technical effects can be found in the system embodiment section and will not be repeated here.

[0111] Fig.14 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Fig.14 As shown, the electronic device 14 of this embodiment includes: at least one processor 1400 ( Fig.14 Only one is shown in the figure), a memory 1401, and a computer program 1402 stored in the memory 1401 and executable on the at least one processor 1400, wherein the processor 1400 implements the steps in any of the above-mentioned method embodiments when executing the computer program 1402. The electronic device may include, but is not limited to, a processor 1400 and a memory 1401. Those skilled in the art will appreciate that Fig.14It is only an example of the electronic device 14 and does not constitute a limitation on the electronic device 14. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.

[0112] The processor 1400 may be a CPU, or other general-purpose processors, NPU, DSP, ASIC, FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The memory 1401 may be an internal storage unit of the electronic device 14 in some embodiments, such as a hard disk or memory of the electronic device 14. The memory 1401 may also be an external storage device of the electronic device 14 in other embodiments, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card, etc. equipped on the electronic device 14. Further, the memory 1401 may also include both an internal storage unit and an external storage device of the electronic device 14. The memory 1401 is used to store an operating system, an application program, a boot loader (BootLoader), data and other programs, such as the program code of the computer program, etc. The memory 1401 may also be used to temporarily store data that has been output or is to be output.

[0113] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit, and the above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application.

[0114] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

Claims

1. An intelligent application control system, characterized in that: include: a first processor and a second processor; The first processor is configured with: a component scheduling module and a model scheduling module running in kernel state; The component scheduling module is used to manage the running status of each functional component in the intelligent application, and send a first task set to the model scheduling module when the functional component is running, wherein the first task set includes: model tasks triggered when the functional component is running, and the model tasks are tasks that depend on the execution of the neural network model; The model scheduling module is used to: when there are multiple target model tasks in the first task set, partially or completely merge the multiple target model tasks in the first task set to obtain a second task set that meets the batch processing condition, wherein the multiple target model tasks are model tasks corresponding to the same neural network model, and the batch processing condition includes: the execution time of each model task in the second task set does not exceed the execution time limit corresponding to each model task; based on the model tasks included in the second task set, send a model driving instruction to the second processor; The second processor is used to: trigger the relevant neural network model to execute the model tasks in the second task set based on the model-driven instructions, and feed back the results of task processing to the first processor.

2. The intelligent application control system according to claim 1, characterized in that: The model scheduling module includes a task scheduling unit, a sequential execution unit and a target determination unit; The task scheduling unit is used to schedule each model task in the first task set based on the priority, execution time limit, estimated running time and queue entry time of each model task in the first task set to obtain a first execution order of each model task in the first task set, and each model task in the first task set is stored in a model queue; The sequential execution unit is used to: determine a second execution order of the plurality of target model tasks based on the first execution order when the plurality of target model tasks exist in the first task set; The target determination unit is used to partially or completely merge the multiple target model tasks based on the second execution order and the estimated running time of the multiple target model tasks.

3. The intelligent application control system according to claim 2, characterized in that: The task scheduling unit is specifically used for: The model tasks in the first task set are sorted in descending order of priority, ascending order of execution time limit, ascending order of estimated running time and ascending order of time of entering the queue.

4. The intelligent application control system according to claim 2, characterized in that: The target determination unit is specifically used for: Determine the first model task in the second execution order as the first model task, and determine the target model task that is second only to the first model task in the second execution order as the second model task; the first model task is the target model task that is first in the second execution order; Based on the estimated running time of the first model task and the estimated running time of the second model task, determining whether the execution time of the third model task exceeds the corresponding execution time limit; The third model task is a model task that is located after the first model task in the first execution order and is other than the first model task and the second model task; When the execution time of the third model task does not exceed the corresponding execution time limit, merging the first model task with the second model task to obtain a batch processing task; Determine the batch task as the first model task, determine the target model task that is second only to the second model task in the second execution order as the second model task, return the steps of executing the estimated running time based on the first model task and the estimated running time of the second model task, and determine whether the execution time of the third model task exceeds the corresponding execution time limit and subsequent steps until the execution time of the third model task exceeds the corresponding execution time limit or the multiple target model tasks are traversed.

5. The intelligent application control system according to claim 4, characterized in that: The model scheduling module also includes a time determination unit; The time determination unit is used to determine the execution time of the first model task as the execution time of the batch processing task.

6. The intelligent application control system according to any one of claims 1 to 5, characterized in that: The second processor is configured with an operator scheduling module; The operator scheduling module is used to: when the model tasks included in the second task set are split into multiple operator tasks based on the model driving instructions, schedule the multiple operator tasks, obtain the scheduling strategies of the multiple operator tasks, trigger the operators of the relevant neural network model to execute the corresponding operator tasks according to the scheduling strategies, and feed back the results of task processing to the first processor.

7. The intelligent application control system according to claim 6, characterized in that: The operator scheduling module includes a parallel determination unit, a first processing unit and a second processing unit; The parallel determination unit is used to: when there are multiple operator task teams that need to be executed in parallel among the multiple operator tasks and the computing power units corresponding to the operator tasks in the multiple operator task teams are in an idle state, determine parallel execution as the scheduling strategy of the multiple operator task teams; each operator task team includes at least one operator task; The first processing unit is used for: for the operator task team including more than two operator tasks, scheduling each operator task in the operator task team based on the priority, execution time limit, estimated running time and entry time of each operator task in the operator task team, obtaining the execution order of each operator task in the operator task team, and determining the execution order of each operator task in the operator task team as the scheduling strategy of each operator task in the operator task team, and storing each operator task in the operator task team in the operator queue; The second processing unit is used to: when the operator task team does not exist in the multiple operator tasks and the computing power units corresponding to the multiple operator tasks are in an idle state, schedule the multiple operator tasks based on the priorities, execution time limits, estimated running times and queue entry times of the multiple operator tasks to obtain the execution order of the multiple operator tasks, and determine the execution order of the multiple operator tasks as the scheduling strategy of the multiple operator tasks, and the multiple operator tasks are stored in the operator queue; The computing unit is used to run the operator that executes the corresponding operator task.

8. The intelligent application control system according to claim 7, characterized in that: The first processing unit is specifically used to: sort the operator tasks in the operator task team in descending order of priority, ascending order of execution time limit, ascending order of estimated running time and ascending order of time of entering the team; The second processing unit is specifically used to sort the multiple operator tasks in descending order of priority, ascending order of execution time limit, ascending order of estimated running time and ascending order of entry time.

9. A method for processing intelligent application tasks, characterized in that: Applied to a first processor, the first processor is configured with: a component scheduling module and a model scheduling module running in kernel state; The component scheduling module is used to manage the operating status of each functional component in the intelligent application; The method comprises: When the functional component is running, triggering the component scheduling module to send a first task set to the model scheduling module, the first task set includes: model tasks triggered when the functional component is running, the model tasks are tasks that depend on the execution of the neural network model; In the case where there are multiple target model tasks in the first task set, the model scheduling module is triggered to partially or completely merge the multiple target model tasks in the first task set to obtain a second task set that meets a batch processing condition, wherein the multiple target model tasks are model tasks corresponding to the same neural network model, and the batch processing condition includes: the execution time of each model task in the second task set does not exceed the execution time limit corresponding to each model task; The model scheduling module is triggered to send a model-driven instruction to the second processor based on the tasks included in the second task set, and the model-driven instruction is used to trigger the second processor to execute the tasks in the second task set and feedback the results of task processing to the first processor. The second processor is configured with task processing capabilities based on a neural network model.

10. The method according to claim 9, characterized in that Before triggering the model scheduling module to partially or completely merge the multiple target model tasks in the first task set, the method further includes: The model scheduling module is triggered to schedule each model task in the first task set based on the priority, execution time limit, estimated running time and queue entry time of each model task in the first task set to obtain a first execution order of each model task in the first task set, and each model task in the first task set is stored in a model queue; triggering the model scheduling module to determine a second execution order of the plurality of target model tasks based on the first execution order when there are a plurality of target model tasks in the first task set; The triggering the model scheduling module to partially or completely merge the multiple target model tasks in the first task set includes: The model scheduling module is triggered to partially or completely merge the multiple target model tasks based on the second execution order and the estimated running time of the multiple target model tasks.

11. The method according to claim 10, characterized in that The triggering of the model scheduling module to schedule each model task in the first task set based on the priority, execution time limit, estimated running time and queue entry time of each model task in the first task set includes: The model scheduling module is triggered to sort the model tasks in the first task set in descending order of priority, ascending order of execution time limit, ascending order of estimated running time and ascending order of entry time.

12. The method according to claim 10, characterized in that The triggering of the model scheduling module to partially or completely merge the multiple target model tasks based on the second execution order and the estimated running time of the multiple target model tasks includes: The model scheduling module is triggered to determine the first model task in the second execution order as the first model task, and to determine the target model task that is second only to the first model task in the second execution order as the second model task; the first model task is the target model task that is first in the second execution order; The model scheduling module is triggered to determine whether the execution time of a third model task exceeds the corresponding execution time limit based on the estimated running time of the first model task and the estimated running time of the second model task; the third model task is a model task that is located after the first model task in the first execution order and is other than the first model task and the second model task; triggering the model scheduling module to merge the first model task with the second model task to obtain a batch task when the execution time of the third model task does not exceed the corresponding execution time limit; Trigger the model scheduling module to determine the batch task as the first model task, determine the target model task that is second only to the second model task in the second execution order as the second model task, and return to the steps of executing the estimated running time based on the first model task and the estimated running time of the second model task, and judging whether the execution time of the third model task exceeds the corresponding execution time limit and subsequent steps until the execution time of the third model task exceeds the corresponding execution time limit or the multiple target model tasks are traversed.

13. The method according to claim 12, characterized in that After getting the batch processing task, it also includes: The execution time of the first model task is determined as the execution time of the batch processing task.

14. A method for processing intelligent application tasks, characterized in that: Applied to a second processor, the second processor being configured with an operator scheduling module; The method comprises: In the case where the model tasks included in the second task set are split into a plurality of operator tasks based on the model-driven instruction, triggering the operator scheduling module to schedule the plurality of operator tasks to obtain scheduling strategies for the plurality of operator tasks; The operator scheduling module is triggered to trigger the operator of the relevant neural network model to execute the corresponding operator task according to the scheduling strategy, and the result obtained by task processing is fed back to the first processor.

15. The method according to claim 14, characterized in that The triggering the operator scheduling module to schedule the multiple operator tasks to obtain the scheduling strategies of the multiple operator tasks includes: When there are multiple operator task teams that need to be executed in parallel among the multiple operator tasks and the computing power units corresponding to the operator tasks in the multiple operator task teams are in an idle state, the operator scheduling module is triggered to determine parallel execution as a scheduling strategy for the multiple operator task teams, each of which includes at least one operator task; For the operator task team including more than two operator tasks, trigger the operator scheduling module to schedule each operator task in the operator task team based on the priority, execution time limit, estimated running time and entry time of each operator task in the operator task team, obtain the execution order of each operator task in the operator task team, and determine the execution order of each operator task in the operator task team as the scheduling strategy of each operator task in the operator task team, and each operator task in the operator task team is stored in the operator queue; In the case that the operator task team does not exist in the multiple operator tasks and the computing power units corresponding to the multiple operator tasks are in an idle state, the operator scheduling module is triggered to schedule the multiple operator tasks based on the priorities, execution deadlines, estimated running times and queue entry times of the multiple operator tasks, obtain the execution order of the multiple operator tasks, and determine the execution order of the multiple operator tasks as the scheduling strategy of the multiple operator tasks, and the multiple operator tasks are stored in the operator queue; The computing unit is used to run the operator that executes the corresponding operator task.

16. The method according to claim 15, characterized in that The triggering of the operator scheduling module schedules each operator task in the operator task team based on the priority, execution time limit, estimated running time and entry time of each operator task in the operator task team, including: Triggering the operator scheduling module to sort the operator tasks in the operator task team in descending order of priority, ascending order of execution time limit, ascending order of estimated running time and ascending order of time of entering the team; The triggering of the operator scheduling module to schedule the multiple operator tasks based on the priorities, execution deadlines, estimated running times and queue entry times of the multiple operator tasks includes: The operator scheduling module is triggered to sort the multiple operator tasks in descending order of priority, ascending order of execution time limit, ascending order of estimated running time and ascending order of enqueue time.

17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the electronic device implements the method according to any one of claims 9 to 13 or the method according to any one of claims 14 to 16.

Citation Information

Patent Citations

  • Computing task processing device and method and electronic equipment

    CN116848509A

  • Heterogeneous core scheduling method and device, storage medium and electronic equipment

    CN116932168A

  • Large model scheduling method and device based on NPU computing power

    CN119336457A

  • Server resource scheduling system integrating AI and edge computing

    CN119718682A

  • Latency and dependency-aware task scheduling workloads on multicore platforms using for energy efficiency

    US20220114033A1