Task scheduling execution method, task scheduling execution instruction generation method and device
By determining and executing parallel task groups based on occupancy rates and state information, the method enhances neural network accelerator efficiency and resource utilization.
Patent Information
- Application Number
- JP2025529848
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-22
- Filing Date
- 2023-09-22
- Publication Date
- 2026-02-04
AI Technical Summary
Neural network accelerators typically execute tasks sequentially, limiting computational efficiency and resource utilization.
Determine a first target task occupying a predetermined occupancy rate and form a first task group based on parallel execution conditions, executing these tasks in parallel using state information and expected occupancy rates.
Improves computational efficiency by enabling parallel processing of multiple tasks, optimizing resource utilization and meeting actual needs.
Smart Images

Figure 2026504251000001_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims priority to a Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on November 22, 2022, bearing application number CN202211467576.X and entitled "Task scheduling execution method, task scheduling execution instruction generation method and device," the entire contents of which are incorporated herein by reference. [Technical Field]
[0002] The present disclosure relates to chip technology, and in particular to a task scheduling execution method, a task scheduling execution instruction generation method and apparatus. [Background technology]
[0003] The chip may include a neural network accelerator, such as a brain processing unit (BPU). The neural network accelerator may have multiple tasks to process, and typically executes these tasks sequentially in the order in which they occur. Summary of the Invention [Problem to be solved by the invention]
[0004] The embodiments of the present disclosure provide a task scheduling execution method, a method and apparatus for generating a task scheduling execution instruction. [Means for solving the problem]
[0005] According to one aspect of the embodiment of the present disclosure, determining that there exists a first target task corresponding to a first version model file that occupies a predetermined occupancy rate of the computational resources of the neural network accelerator in an execution state; determining, as a first task group, a set of the first target tasks that satisfy a preset parallel execution condition based on a state information group of the first version model file, which includes states corresponding to each functional unit of the neural network accelerator in an execution state of the first version model file, and the expected occupancy rate; and executing in parallel the first version model files corresponding to the first target tasks in the first task group, respectively.
[0006] According to another aspect of the present disclosure, a step of generating a first version model file corresponding to a first group of operators through a compilation process, wherein the first version model file occupies a predetermined occupancy rate of the computational resources of the neural network accelerator in an execution state; generating a set of state information of the first version model file based on a set of functional units corresponding to the first set of operators, the set of functional units corresponding to the first set of operators including each functional unit of the neural network accelerator for operating the first set of operators, and the set of state information of the first version model file including a state corresponding to each functional unit of the neural network accelerator in an execution state of the first version model file; generating a task scheduling execution instruction for executing the task scheduling execution method based on the first version model file, a set of status information of the first version model file, and the scheduled occupancy rate.
[0007] According to yet another aspect of the present disclosure, a first determination module for determining that there exists a first target task corresponding to a first version model file that occupies a computational resource of the neural network accelerator at a predetermined occupancy rate in an execution state; a second determination module for determining, as a first task group, a set of first target tasks that satisfy a preset parallel execution condition and that are determined by the first determination module, based on a state information group of the first version model file, the state information group including states corresponding to each functional unit of the neural network accelerator in an execution state of the first version model file, and the expected occupancy rate; a first execution module for parallel executing the first version model files corresponding to the first target tasks in the first task group determined by the second determination module.
[0008] According to yet another aspect of the present disclosure, a first generation module for generating a first version model file corresponding to a first group of operators through a compilation process, the first version model file occupying a predetermined occupancy rate of the computational resources of the neural network accelerator in an execution state; a second generating module for generating a set of state information of the first version model file generated by the first generating module based on a set of functional units corresponding to the first set of operators, wherein the set of functional units corresponding to the first set of operators includes each functional unit of the neural network accelerator for operating the first set of operators, and the set of state information of the first version model file includes a state corresponding to each functional unit of the neural network accelerator in an execution state of the first version model file; and a third generation module for generating a task scheduling execution command for executing the task scheduling execution method based on the first version model file generated by the first generation module, a set of state information of the first version model file generated by the second generation module, and the scheduled occupancy rate.
[0009] According to yet another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program for executing the above-described task scheduling execution method or the above-described method for generating a task scheduling execution instruction.
[0010] According to yet another aspect of the present disclosure, a processor; a memory for storing instructions executable by said processor; An electronic device is provided in which the processor is used to read the executable instructions from the memory and execute the instructions to implement the task scheduling execution method or the task scheduling execution instruction generation method.
[0011] According to yet another aspect of an embodiment of the present disclosure, there is provided a computer program product that, when instructions in the computer program product are executed by a processor, implements the task scheduling execution method or the method for generating a task scheduling execution instruction. [Effects of the Invention]
[0012] According to the task scheduling execution method, task scheduling execution instruction generation method, device, computer-readable storage medium, electronic device, and product provided in the above embodiments of the present disclosure, a first task group can be determined based on the state information group of the first version model file and the expected occupancy rate of the computing resources of the neural network accelerator in the execution state of the first version model file, and the first version model files corresponding to each first target task in the first task group can be executed in parallel. In this way, the task scheduling mechanism can realize parallel processing of multiple tasks by the neural network accelerator, thereby improving the computational efficiency of the neural network accelerator and better meeting actual needs. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a schematic diagram of a chip structure in one exemplary embodiment of the present disclosure. [Figure 2] FIG. 1 is a diagram illustrating the principle of realizing parallel processing of multiple tasks by a neural network accelerator in an embodiment of the present disclosure. [Figure 3] 1 is a flowchart of a task scheduling execution method provided in one exemplary embodiment of the present disclosure. [Figure 4] 10 is a flowchart of a task scheduling execution method provided in another exemplary embodiment of the present disclosure. [Figure 5-1] FIG. 2 is a schematic diagram of a task queue and a task scheduling table in a task scheduling execution method provided in an exemplary embodiment of the present disclosure. [Figure 5-2] FIG. 2 is a schematic diagram of task decomposition in a task scheduling execution method provided in one exemplary embodiment of the present disclosure. [Figure 6] 10 is a flowchart of a task scheduling execution method provided in yet another exemplary embodiment of the present disclosure. [Figure 7]1 is a flowchart of a method for generating a task scheduling execution instruction provided in an exemplary embodiment of the present disclosure. [Figure 8] 10 is a flowchart of a method for generating a task scheduling execution instruction provided in another exemplary embodiment of the present disclosure. [Figure 9] 10 is a flowchart of a method for generating a task scheduling execution instruction provided in yet another exemplary embodiment of the present disclosure. [Figure 10] 10 is a flowchart of a method for generating a task scheduling execution instruction provided in yet another exemplary embodiment of the present disclosure. [Figure 11] FIG. 2 is a schematic diagram of the structure of a task scheduling execution device provided in one exemplary embodiment of the present disclosure; [Figure 12] FIG. 10 is a schematic diagram of the structure of a task scheduling execution device provided in another exemplary embodiment of the present disclosure; [Figure 13] FIG. 2 is a schematic diagram of the structure of a task scheduling execution instruction generating device provided in one exemplary embodiment of the present disclosure; [Figure 14] FIG. 10 is a schematic diagram of a structure of a task scheduling execution instruction generating device provided in another exemplary embodiment of the present disclosure; [Figure 15] 1 is a structural diagram of an electronic device provided in one exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0014] In order to understand the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the drawings. Obviously, it should be understood that the described embodiments are only a part of the embodiments of the present disclosure, not all of the embodiments, and the present disclosure is not limited to the exemplary embodiments.
[0015] It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, formulas and numerical values described in these examples do not limit the scope of the present disclosure. Summary of the application
[0016] Some chips may include a neural network accelerator, for example, an artificial intelligence (AI) chip may include a BPU. A neural network accelerator may have multiple tasks to process, and typically executes these tasks sequentially in the order in which the tasks occurred, and the neural network accelerator executes only one task at a time. Exemplary System
[0017] The neural network accelerator in the chip may include a computational element and multiple functional units, where the L1 SRAM (Static Random-Access Memory) in FIG. 1 may be a computational element, and the Tensor Core, Vector Core, Scalar Core, and DSU (Domain Specific Unit) in FIG. 1 may each be a functional unit.
[0018] Optionally, the chip may include other components in addition to the neural network accelerator, such as a graphics processing unit (GPU), a digital signal processing unit (DSP), etc.
[0019] As shown in FIG. 2, in the embodiments of the present disclosure, the compilation stage and the execution stage can be improved, respectively. In the compilation stage, a task scheduling execution instruction can be generated. In the execution stage, the task scheduling execution instruction generated in the compilation stage is executed, thereby realizing parallel processing of multiple tasks by the neural network accelerator, thereby improving the computational efficiency of the neural network accelerator. Exemplary Methods
[0020] 3 is a flowchart of a task scheduling execution method provided in one exemplary embodiment of the present disclosure. The method shown in FIG. 3 includes steps 310, 320, and 330, each of which will be described below.
[0021] In step 310, it is determined that there is a first target task corresponding to a first version model file in a running state that occupies the computational resources of the neural network accelerator at a predetermined occupancy rate.
[0022] In alternative examples, step 310 may be performed by a processor invoking corresponding instructions stored in a memory, or may be performed by a first decision module executed by a processor.
[0023] In step 310, all tasks to be processed by the neural network accelerator are determined, and these tasks are traversed to determine which of these tasks have corresponding first version model files. Then, each task having a corresponding first version model file can be selected as a first target task. In this way, by performing step 310, several first target tasks can be determined (for convenience of explanation, it is subsequently assumed that the number of first target tasks is N, where N may be an integer greater than or equal to 2).
[0024] The relationship between any first target task and the corresponding first version model file can be understood as the first target task being completed by executing the first version model file, and the occupancy rate of the first version model file in the computational resources of the neural network accelerator during the execution of the first version model file is a predetermined occupancy rate.
[0025] Alternatively, any predetermined occupancy rate may be any rate greater than 0% and less than 100%, such as 30%, 40%, 60%, etc. The predetermined occupancy rates corresponding to different first version model files may be the same or different. The computational resources of the neural network accelerator may refer to the computational resources of the computing elements in the neural network accelerator, such as the computational resources of the L1 SRAM in FIG. 1.
[0026] In step 320, a set of first target tasks that satisfy a preset concurrent execution condition is determined as a first task group based on a set of state information of the first version model file, which includes states corresponding to each functional unit of the neural network accelerator in the execution state of the first version model file, and the expected occupancy rate.
[0027] In an alternative example, step 320 may be performed by a processor invoking corresponding instructions stored in a memory, or may be performed by a second decision module executed by the processor.
[0028] After determining the N first target tasks by performing step 310, the corresponding state information groups and scheduled occupancy rates are determined for the first version model files corresponding to each of the N first target tasks, thereby obtaining the N state information groups and the N scheduled occupancy rates, where a state corresponding to any functional unit in the state information group corresponding to any first version model file can be used to characterize whether the functional unit is used in the execution state of the first version model file.By referring to the N state information groups and the N scheduled occupancy rates, it can be determined which first target tasks among the N first target tasks satisfy the preset parallel execution condition, thereby determining the first task group.
[0029] In one example, the N first target tasks may specifically be four first target tasks, respectively designated as task 1 to task 4, and only task 1, task 3, and task 4 satisfy a preset parallel execution condition. In this case, a set consisting of task 1, task 3, and task 4 can be determined as one first task group.
[0030] In another example, the N first target tasks may specifically be six first target tasks, respectively designated as task 1 to task 6, where task 1 to task 3 satisfy a preset parallel execution condition, and task 4 to task 6 satisfy a preset parallel execution condition. In this case, a set of three tasks, task 1 to task 3, can be determined as one first task group, and a set of three tasks, task 4 to task 6, can be determined as another first task group.
[0031] In step 330, the first version model files corresponding to the first target tasks in the first task group are executed in parallel.
[0032] In alternative examples, step 330 may be performed by the processor invoking corresponding instructions stored in memory, or may be performed by a first execution module executed by the processor.
[0033] If the number of first task groups is one and this first task group is a set of three tasks, namely, task 1, task 3, and task 4, then parallel processing of tasks 1, 3, and 4 by the neural network accelerator can be realized by executing the first version model files corresponding to tasks 1, 3, and 4 in parallel using the neural network accelerator.
[0034] If there are two first task groups, one of which is a set of three tasks, Task 1 to Task 3, and the other is a set of three tasks, Task 4 to Task 6, then the neural network accelerator can first execute in parallel the first version model files corresponding to Task 1, Task 2, and Task 3, respectively, thereby achieving parallel processing of Task 1, Task 2, and Task 3 by the neural network accelerator, and then the neural network accelerator can execute in parallel the first version model files corresponding to Task 4, Task 5, and Task 6, respectively, thereby achieving parallel processing of Task 4, Task 5, and Task 6 by the neural network accelerator. Naturally, depending on the actual situation, the neural network accelerator can first execute in parallel the first version model files corresponding to Task 4, Task 5, and Task 6, respectively, and then the neural network accelerator can execute in parallel the first version model files corresponding to Task 1, Task 2, and Task 3, respectively.
[0035] In the embodiment of the present disclosure, a first task group is determined based on the state information group of the first version model file and the expected occupancy rate of the computing resources of the neural network accelerator in the execution state of the first version model file, and the first version model file corresponding to each first target task in the first task group can be executed in parallel. In this way, the task scheduling mechanism is equivalent to realizing parallel processing of multiple tasks by the neural network accelerator, thereby improving the computing efficiency of the neural network accelerator and better meeting actual needs.
[0036] In one alternative example, the plurality of first target tasks satisfying the preset parallel execution condition is: (1) In an information set consisting of a group of state information corresponding to each of the first target tasks among a plurality of first target tasks, among all states corresponding to any of the functional units, one state is a used state and the remaining states are all idle states, or all of the states are idle states; (2) The sum of the scheduled occupancy rates corresponding to each of the plurality of first target tasks is smaller than a preset rate.
[0037] Alternatively, the preset percentage may be 100%. Alternatively, the preset percentage may be smaller than 100% but close to 100%. For ease of understanding, the embodiments of the present disclosure will be described taking the case where the preset percentage is 100% as an example.
[0038] In addition, in the state information group corresponding to any of the first target tasks, the state corresponding to any of the functional units may be either a used state or an idle state, where the used state may be represented by shared and the idle state may be represented by available.
[0039] In one example, the neural network accelerator includes three functional units, which are a Tensor Core, a Vector Core, and a DSU, and the N first target tasks are specifically four first target tasks, which are Task 1 to Task 4, respectively, and the state information groups corresponding to Task 1 to Task 4 are as follows: Task 1: Tensor Core-shared, Vector core-available, DSU-available Task 2: Tensor Core-shared, Vector core-shared, DSU-shared Task 3: Tensor Core-available, Vector core-shared, DSU-available Task 4: Tensor Core-available, Vector core-available, DSU-shared Here, a format such as "A-shared" indicates that the state corresponding to the functional unit A is in use, and a format such as "B-available" indicates that the state corresponding to the functional unit B is in idle.
[0040] As can be easily seen, of the three tasks, task 1, task 3, and task 4, only task 1 is in a busy state when it corresponds to a Tensor Core, while the states corresponding to the other two are both idle when it corresponds to a Tensor Core; of the three tasks, task 1, task 3, and task 4, only task 3 is in a busy state when it corresponds to a Vector Core, while the states corresponding to the other two are both idle when it corresponds to a Vector Core; and of the three tasks, task 1, task 3, and task 4, only task 4 is in a busy state when it corresponds to a DSU, while the states corresponding to the other two are both idle when it corresponds to a DSU. Therefore, the condition defined in (1) above is met for task 1, task 3, and task 4. In addition, assuming that the scheduled occupancy rates corresponding to Task 1 to Task 4 are 30%, 30%, 25%, and 35%, respectively, the sum of the scheduled occupancy rates corresponding to Task 1, Task 3, and Task 4 is obviously less than 100%, and therefore, for Task 1, Task 3, and Task 4, the condition defined in (2) above is also satisfied, and it can be determined that Task 1, Task 3, and Task 4 satisfy the preset parallel execution condition. In this way, Task 1 can be realized using Tensor Cores, while Task 3 can be realized using Vector Cores and Task 4 can be realized using DSUs. In other words, the neural network accelerator can simultaneously execute Task 1, Task 3, and Task 4.
[0041] Furthermore, for task 1, task 3, and task 4, if the condition defined in (1) above is satisfied, and the schedule occupancy rates corresponding to tasks 1 to 4 are assumed to be 40%, 25%, 50%, and 65%, respectively, then it is clear that the total of the schedule occupancy rates for task 1, task 3, and task 4 is greater than 100%, the total of the schedule occupancy rates for task 1 and task 4 is greater than 100%, the total of the schedule occupancy rates for task 3 and task 4 is greater than 100%, and the total of the schedule occupancy rates for task 1 and task 3 is less than 100%. In other words, for task 1 and task 3, the condition defined in (2) above is satisfied, and it can be determined that task 1 and task 3 satisfy the preset parallel execution condition. In this way, Task 1 can be achieved by using the Tensor Cores, while Task 3 can be achieved by using the Vector Cores. In other words, the neural network accelerator can execute both Task 1 and Task 3 at the same time.
[0042] Assume that the N first target tasks in the above example are not four first target tasks but five first target tasks, and further include task 5 in addition to tasks 1 to 4 above, and the planned occupancy rate corresponding to task 5 is 30%, and the state information group corresponding to task 5 is as follows: Task 5: Tensor Core-shared, Vector core-shared, DSU-available
[0043] Assume that the scheduled occupancy rates corresponding to tasks 1 to 4 are 40%, 25%, 50%, and 65%, respectively. Obviously, tasks 1 and 3 satisfy the preset parallel execution conditions. Furthermore, only one of tasks 4 and 5 is in a busy state corresponding to a Tensor core, only one of tasks 4 and 5 is in a busy state corresponding to a Vector core, and only one of tasks 4 and 5 is in a busy state corresponding to a DSU. Furthermore, the sum of the scheduled occupancy rates corresponding to tasks 4 and 5 is 95%, which is less than 100%, so it can be determined that tasks 4 and 5 also satisfy the preset parallel execution conditions. In this way, tasks 1 and 3 can be divided into one first task group, and tasks 4 and 5 can be divided into another first task group, thereby determining two first task groups.
[0044] In the embodiment of the present disclosure, the condition defined in (1) above ensures that for each functional unit in the neural network accelerator, there can be at most one first target task in the first task group that uses the functional unit at the same time, thereby preventing different first target tasks in the first task group from using the same functional unit simultaneously and thereby avoiding usage conflicts of the functional unit.The condition defined in (2) above ensures that the computing resources of the neural network accelerator can support parallel processing of each first target task in the first task group, thereby ensuring that each first target task in the first task group is successfully completed through parallel processing, thereby improving the utilization rate of the computing components in the neural network accelerator.
[0045] Based on the embodiment shown in FIG. 3, the method further includes step 301, step 303, step 305 and step 307, as shown in FIG.
[0046] In step 301, a task queue is obtained, and each task in the task queue corresponds to a neural network model.
[0047] In alternative examples, step 301 may be performed by a processor invoking corresponding instructions stored in a memory, or may be performed by an acquisition module executed by a processor.
[0048] Alternatively, the task queue may be the BPU Task Queue of Figure 5-1.
[0049] A neural network model may be considered to be a sequence of operator units, i.e., the neural network model may include a plurality of operator units (e.g., 40, 50, 100, etc.) arranged according to a particular order, where the plurality of operator units include, but are not limited to, a convolution (Conv) operator unit, a pooling (Pool) operator unit, a deconvolution operator unit, a rectified linear unit (ReLU) operator unit, a batch normalization (BN) operator unit, etc.
[0050] In addition, the relationship between any task in the task queue and the corresponding neural network model may be understood as the need for the task to be completed depending on the neural network model. For example, if the task is a target detection task and the neural network model is a model used for target detection, the task can be considered to be completed by providing an image of the target to be detected as input to the neural network model, performing calculations, and obtaining the target detection result output by the neural network model.
[0051] In step 303, it is determined whether there is segmentation scheme information corresponding to a target neural network model corresponding to a second target task, where the second target task is any task in the task queue. In response to the segmentation scheme information corresponding to the target neural network model being present, step 305 is executed, and in response to the segmentation scheme information not being present, step 307 is executed.
[0052] In an alternative example, step 303 may be performed by a processor invoking corresponding instructions stored in a memory, or may be performed by a third decision module executed by the processor.
[0053] Alternatively, the target memory area may store a correspondence between the neural network model and the division method information. The origin of the correspondence stored in the target memory area may refer to the relevant description of the compilation stage below, and will not be explained in detail here.
[0054] In step 303, the correspondence relationships stored in the target memory area can be traversed. If it is determined by traversing the correspondence relationships stored in the target memory area that the division method information corresponding to the target neural network model exists in the correspondence relationships stored in the target memory area, step 305 can be executed; if it is determined by traversing the correspondence relationships stored in the target memory area that the division method information corresponding to the target neural network model does not exist in the correspondence relationships stored in the target memory area, step 307 can be executed.
[0055] In step 305, the second target task is divided to obtain K divided tasks, and the K divided tasks are added to the task scheduling table. The K divided tasks correspond to K operator groups obtained by dividing the target neural network model according to the division method information corresponding to the target neural network model.
[0056] In alternative examples, step 305 may be performed by a processor invoking corresponding instructions stored in a memory, or may be performed by a first processing module executed by a processor.
[0057] Alternatively, the task scheduling table may be the Task Scheduler of FIG. 5-1.
[0058] It should be noted that the division method information corresponding to the target neural network model can be used to divide the target neural network model into K operator groups, each operator group including at least one operator unit in the target neural network model. Thus, the second target task can be divided based on the division method information corresponding to the target neural network model to obtain K divided tasks that correspond one-to-one to the K operator groups. Optionally, K can be 2, 3, 4, or an integer greater than 4, and is not listed here.
[0059] In step 307, the second target task is added to the task scheduling table.
[0060] In alternative examples, step 307 may be performed by the processor invoking corresponding instructions stored in memory, or may be performed by a second processing module executed by the processor.
[0061] Step 310 includes step 3101 .
[0062] In step 3101, it is determined from the task scheduling table that there is a first target task corresponding to the first version model file.
[0063] In step 3101, all tasks in the task scheduling table can be traversed to determine which of these tasks have a corresponding first version model file, and then each task having a corresponding first version model file can be set as a first target task.
[0064] In one example, the second target task is Task1 in Figure 5-2, and the target neural network model includes five operator units, which are Convolution1, Pooling1, Convolution2, Pooling2, and Convolution3, in this case, Task1 can be divided to obtain five divided tasks, which are Task1.1, Task1.2, Task1.3, Task1.4, and Task1.5, respectively, where Task1.1 corresponds to Convolution1, Task1.2 corresponds to Pooling1, Task1.3 corresponds to Convolution2, Task1.4 corresponds to Pooling2, and Task1.5 corresponds to Convolution3. Assuming that each convolution is executed on a Tensor core and each pooling is executed on a Vector core, when executing Task1, it is necessary to use a Tensor core and a Vector core, when executing Task1.1, it is necessary to use a Tensor core, when executing Task1.2, it is necessary to use a Vector core, when executing Task1.3, it is necessary to use a Tensor core, when executing Task1.4, it is necessary to use a Vector core, when executing Task1.5, it is necessary to use a Vector core, and clearly, the number of functional units that need to be used when executing any of Task1.1 to Task1.5 is fewer than the number of functional units that need to be used when executing Task1.
[0065] In this way, Task1.1 and Task1.2 can each be a first target task, and if the sum of the occupancy rates of the two schedules corresponding to Task1.1 and Task1.2 is less than 100%, the neural network accelerator can process Task1.1 and Task1.2 in parallel to improve the computational efficiency of the neural network accelerator. Similarly, Task1.3 and Task1.4 can each be a subsequent first target task, and if the sum of the occupancy rates of the two schedules corresponding to Task1.3 and Task1.4 is less than 100%, the neural network accelerator can process Task1.3 and Task1.4 in parallel.
[0066] In the embodiments of the present disclosure, it is possible to determine whether to add the second target task directly to the task scheduling table or to add K split tasks obtained by splitting the second target task to the task scheduling table based on whether there is splitting method information corresponding to the neural network model corresponding to the second target task in the task queue. In this way, for a task for which there is corresponding splitting method information, several split tasks with finer granularity than the task can be obtained through the split process, and fewer functional units need to be used when executing each split task. This increases the probability that different tasks in the task scheduling table can be processed in parallel, which is advantageous for improving the computational efficiency of the neural network accelerator.
[0067] Based on the embodiment shown in FIG. 4, the method further includes step 340 and step 350, as shown in FIG.
[0068] In step 340, a second task group is determined in the task scheduling table, the second task group including a set of third target tasks other than each first target task in the first task group.
[0069] In an alternative example, step 340 may be performed by a processor invoking corresponding instructions stored in a memory, or may be performed by a fourth decision module executed by the processor.
[0070] In step 340, all tasks in the task scheduling table may be traversed to determine which tasks in the task scheduling table are not in the first group of tasks, each of which may be a first third target task, and the collection of all third target tasks may be a second group of tasks.
[0071] In step 350, the second version model files respectively corresponding to the third target tasks in the second task group are executed according to a preset order, and the second version model files fully occupy the computing resources in the running state.
[0072] In alternative examples, step 350 may be performed by the processor invoking corresponding instructions stored in memory, or may be performed by a second execution module executed by the processor.
[0073] In addition, the relationship between any third target task and the corresponding second version model file can be understood as follows: the third target task can be completed by executing the second version model file; and when the second version model file is executed, the proportion of the computing resources of the neural network accelerator occupied by the second version model file is 100%.
[0074] Optionally, any of the second version model files may have a set of state information, and the set of state information of any of the second version models includes a state corresponding to each functional unit of the neural network accelerator in the execution state of the second version model file, and each state in the set of state information is an exclusive state, where the exclusive state can be represented by "exclusive."
[0075] In step 350, the time at which each third target task in the second task group is added to the task scheduling table is determined, and the second version model files corresponding to each third target task in the second task group are sequentially executed by the neural network accelerator in order of earliest addition time, thereby realizing sequential processing of each third target task in the second task group.
[0076] In the embodiments of the present disclosure, for tasks in the task scheduling table that cannot be processed in parallel, these tasks can be processed sequentially by the neural network accelerator, and in this way, all tasks in the task scheduling table can be processed normally and will not cause task leakage.
[0077] In one example, as shown in FIG. 5-1, it is assumed that there are three tasks, Task1, Task2, and Task3, in the task queue, the neural network model corresponding to Task1 is model1, the neural network model corresponding to Task2 is model2, and the neural network model corresponding to Task3 is model3, there is no division method information corresponding to model1, there is division method information corresponding to model2, and this division method information is used to divide model2 into operator group1 and operator group2, and there is division method information corresponding to model3, and this division method information is used to divide model3 into operator group3 and operator group4. In this case, it is not necessary to divide Task1, and Task2 may be divided to obtain Task2.1 and Task2.2, and Task3 may be divided to obtain Task3.1 and Task3.2. Here, Task2.1 corresponds to operator group1, Task2.2 corresponds to operator group2, Task3.1 corresponds to operator group3, and Task3.2 corresponds to operator group4. Task1, Task2.1, Task2.2, Task3.1, and Task3.2 can all be added to the Task Scheduler.
[0078] Assume that a neural network accelerator includes two functional units, a Tensor Core and a Vector Core, and that for Task1, Task2.1, and Task3.1, corresponding first-version model files do not exist, but only corresponding second-version model files exist, and that for Task2.2 and Task3.2, corresponding second-version model files exist. A set of status information for the second-version model files corresponding to Task1, Task2.1, and Task3.1, respectively, and a set of status information for the first-version model files corresponding to Task2.2 and Task3.2, respectively, may be added to the Task Scheduler. The contents added to the Task Scheduler can be seen in Figure 5-1. Here, a format such as "A_exclusive" indicates that the state corresponding to functional unit A is exclusive, a format such as "B_available" indicates that the state corresponding to functional unit B is idle, and a format such as "C_shared" indicates that the state corresponding to functional unit C is in use.
[0079] Since there are no corresponding first version model files for Task1, Task2.1, and Task3.1, none of Task1, Task2.1, and Task3.1 can be the first target task, and each can only be one third target task. As such, none of Task1, Task2.1, and Task3.1 can be processed in parallel with other tasks, and each can only be processed independently. Only one of Task2.2 and Task3.2 is in a busy state corresponding to a Tensor Core, and only one of Task2.2 and Task3.2 is in a busy state corresponding to a Vector Core. Therefore, if the total value of the expected occupancy rates corresponding to the first version model files corresponding to Task2.2 and Task3.2 is less than 100%, it can be determined that Task2.2 and Task3.2 satisfy the preset parallel execution condition. In this case, the neural network accelerator may first process Task2.2 and Task3.2 in parallel, and then process Task1, Task2.1, and Task3.1 sequentially, thereby realizing the processing of all tasks in the Task Scheduler.
[0080] Any of the task scheduling execution methods provided in the embodiments of the present disclosure may be executed by any suitable device having data processing capabilities, including, but not limited to, a terminal device, a server, etc. Alternatively, any of the task scheduling execution methods provided in the embodiments of the present disclosure may be executed by a processor, for example, the processor executes any of the task scheduling execution methods according to the embodiments of the present disclosure by calling corresponding instructions stored in a memory. Hereinafter, redundant description will be omitted.
[0081] 7 is a flowchart of a method for generating a task scheduling execution instruction provided in an exemplary embodiment of the present disclosure. The method shown in FIG. 7 includes steps 710, 720, and 730, each of which will be described below.
[0082] In step 710, a first version model file corresponding to the first group of operators is generated through a compilation process, and the first version model file occupies the computational resources of the neural network accelerator at a predetermined occupancy rate in an execution state.
[0083] In alternative examples, step 710 may be performed by a processor invoking corresponding instructions stored in memory, or may be performed by a first generation module executed by a processor.
[0084] Alternatively, the first operator group may be a complete neural network model, or a set of consecutive operator units in the complete neural network model. The number of first operator groups may be multiple, and each first operator group in the multiple first operator groups may correspond to one of the above-mentioned first target tasks.
[0085] In step 710, a compilation process can be performed by a compiler to generate a first version model file corresponding to each first operator group. The specific compilation process method can be any feasible method according to actual needs, and this disclosure will not describe it.
[0086] In step 720, a set of state information of a first version model file is generated based on a set of functional units corresponding to the first set of operators, where the set of functional units corresponding to the first set of operators includes each functional unit of a neural network accelerator for operating the first set of operators, and the set of state information of the first version model file includes a state corresponding to each functional unit of the neural network accelerator in the execution state of the first version model file.
[0087] In alternative examples, step 720 may be performed by a processor invoking corresponding instructions stored in memory, or may be performed by a second generation module executed by a processor.
[0088] In step 720, a set of functional units corresponding to the first set of operators may be determined. Assuming that the neural network accelerator includes three functional units, namely, a Tensor Core, a Vector Core, and a DSU, and that the Tensor Core and the Vector Core need to be used when executing the first set of operators, the set of functional units corresponding to the first set of operators includes the Tensor Core and the Vector Core. Next, a set of state information for the first version model file may be generated by referring to the set of functional units corresponding to the first set of operators. In the set of state information for the first version model file, a state corresponding to any functional unit can be used to characterize whether the functional unit is used in the execution state of the first version model file.
[0089] In step 730, a task scheduling execution command is generated based on the first version model file, the state information group of the first version model file, and the scheduled occupancy rate, the task scheduling execution command being used in the task scheduling execution method (specifically, the task scheduling execution method may be the task scheduling execution method in the embodiment shown in FIG. 3).
[0090] In an alternative example, step 730 may be performed by a processor invoking corresponding instructions stored in memory, or may be performed by a third generation module executed by a processor.
[0091] In an embodiment of the present disclosure, the compilation stage sequentially executes a step of generating a first version model file and a step of generating a state information group for the first version model file, and further generates a task scheduling execution instruction according to the expected occupancy ratio corresponding to the first version model file. Thus, the execution stage executes the task scheduling execution instruction generated in the compilation stage to sequentially determine a first target task and a first task group, and executes the first version model files corresponding to each first target task in the first task group in parallel. This task scheduling mechanism is equivalent to realizing parallel processing of multiple tasks by the neural network accelerator, thereby improving the computational efficiency of the neural network accelerator and better meeting practical needs.
[0092] Based on the embodiment shown in FIG. 7, any functional unit of the neural network accelerator can be the target functional unit, and as shown in FIG. 8, step 720 includes step 7201 and step 7203.
[0093] In step 7201, in response to the target functional unit being present in the functional unit group corresponding to the first operator group, it is determined that the state corresponding to the target functional unit in the state information group of the first version model file is in a used state.
[0094] In an alternative example, step 7201 may be performed by a processor invoking corresponding instructions stored in a memory, or may be performed by a first determination sub-module executed by a processor.
[0095] In step 7203, in response to the target functional unit not being present in the functional unit group corresponding to the first operator group, it is determined that the state corresponding to the target functional unit in the state information group of the first version model file is idle.
[0096] In an alternative example, step 7203 may be performed by the processor invoking corresponding instructions stored in memory, or may be performed by a second determination sub-module executed by the processor.
[0097] In one example, the neural network accelerator includes three functional units: a Tensor Core, a Vector Core, and a DSU. The functional unit group corresponding to the first operator group includes a Tensor Core and a Vector Core. Since the Tensor Core and the Vector Core are both present in the functional unit group corresponding to the first operator group, the states corresponding to the Tensor Core and the Vector Core in the state information group of the first version model file may both be in used states. Furthermore, since the DSU is not present in the functional unit group corresponding to the first operator group, the state corresponding to the DSU in the state information group of the first version model file may be idle. Thus, the state information group of the first version model file can be expressed in the format of Tensor Core-shared, Vector core-shared, and DSU-available.
[0098] In an embodiment of the present disclosure, by referring to whether the target functional unit exists in the functional unit group corresponding to the first operator group, it is possible to efficiently and reliably determine the state of each functional unit in the neural network accelerator corresponding to the target functional unit, thereby efficiently and reliably generating a state information group of the first version model file.
[0099] Based on the embodiment shown in FIG. 7, before step 710, the method further includes step 701 and step 703, as shown in FIG.
[0100] In step 701, if the functional unit groups corresponding to the operator units in the neural network model are not completely the same, divide the neural network model into K operator groups corresponding to different functional unit groups, record the corresponding division method information, and the functional unit group corresponding to any of the operator groups includes each functional unit in the neural network accelerator for operating the operator group.
[0101] In an alternative example, step 701 may be performed by a processor invoking corresponding instructions stored in a memory, or may be performed by a third processing module executed by the processor.
[0102] Alternatively, K may be 2, 3, 4, or an integer greater than 4, not all of which are listed here.
[0103] In step 701, a set of functional units corresponding to each operator unit in the neural network model may be determined, where the set of functional units corresponding to any operator unit includes functional units for operating the set of operators in the neural network accelerator, and then the sets of functional units corresponding to each operator unit in the neural network model may be compared to determine whether the sets of functional units corresponding to each operator unit in the neural network model are exactly the same.
[0104] If the functional unit groups corresponding to the operator units of the neural network model are exactly the same, there is no need to divide the neural network model, and naturally, there is no division method information corresponding to the neural network model.
[0105] In one example, the neural network accelerator includes three functional units: a Tensor Core, a Vector Core, and a DSU. The neural network model corresponds to Task 2 in Figure 5-1. That is, the functional units corresponding to each operator unit included in the neural network model all include only a Tensor Core. In this case, the neural network model does not need to be divided. Thus, in the execution phase, Task 2 can be entirely executed on the Tensor Core. Similarly to Task 2, Task 3 in Figure 5-1 can be entirely executed on the Vector Core.
[0106] If the functional unit groups corresponding to each operator unit in the neural network model are not exactly the same, the neural network model may be divided into K operator groups each corresponding to a different functional unit group, and the corresponding division method information may be recorded.
[0107] In one example, a neural network accelerator includes three functional units: a Tensor Core, a Vector Core, and a DSU; the neural network model includes 30 operator units, of which the functional unit group corresponding to the first 10 operator units includes a Tensor Core and a Vector Core, the functional unit group corresponding to the middle 10 operator units includes a Vector Core and a DSU, and the functional unit group corresponding to the last 10 operator units includes a Tensor Core, a Vector Core, and a DSU; in this case, the neural network model may be divided into three operator groups. Here, the first operator group includes the first 10 operator units of the 30 operator units included in the neural network model, the second operator group includes the middle 10 operator units of the 30 operator units included in the neural network model, and the third operator group includes the last 10 operator units of the 30 operator units included in the neural network model, and the division scheme information recorded for the neural network model can be used to characterize the 30 operator units included in the neural network model as being divided into three equal parts. Alternatively, the neural network model can be divided into two operator groups, where the first operator group includes the first 10 operator units of the 30 operator units included in the neural network model and the second operator group includes the remaining 20 operator units of the 30 operator units included in the neural network model, and the division scheme information recorded for the neural network model can be used to characterize the 30 operator units included in the neural network model as being divided in a 1:2 ratio. After the division method information is recorded for the neural network model, the correspondence between the neural network model and the division method information may be recorded in the target storage area.
[0108] In step 703, each operator group in at least some of the K operator groups is set as a first operator group.
[0109] In an alternative example, step 703 may be performed by a processor invoking corresponding instructions stored in a memory, or may be performed by a fifth decision module executed by the processor.
[0110] Alternatively, each operator group of the K operator groups may be a first operator group.
[0111] Step 730 The method includes a step 7301 of generating, for each first operator group, a task scheduling execution command for executing a task scheduling execution method (specifically, the task scheduling execution method may be the task scheduling execution method in the embodiment shown in FIG. 4 ) based on a first version model file corresponding to the first operator group, a state information group of the first version model file corresponding to the first operator group, a planned occupancy rate corresponding to the first operator group, and division method information.
[0112] By executing step 7301, it is possible to generate corresponding task scheduling execution instructions for each first operator group.
[0113] In the embodiments of the present disclosure, when the functional unit groups corresponding to each operator unit in the neural network model are not completely the same, the neural network model can be divided by referring to the functional unit groups corresponding to each operator unit, so that different operator groups obtained by the division correspond to different functional unit groups, and the division method information can be recorded. In this way, in the execution stage, the correspondence stored in the target memory area can be referred to to determine whether to add the second target task directly to the task scheduling table or to divide the second target task and then add it to the task scheduling table. The task division process is advantageous to increase the probability that different tasks in the task scheduling table can be processed in parallel.
[0114] Based on the embodiment shown in FIG. 9, as shown in FIG. 10, the method further includes step 711.
[0115] In step 711, a second version model file corresponding to each first operator group is generated through a compilation process, and the second version model file fully occupies the computing resources in the running state.
[0116] In an alternative example, step 711 may be performed by a processor invoking corresponding instructions stored in a memory, or may be performed by a fourth generation module executed by the processor.
[0117] In step 711, a compilation process can be performed by a compiler to generate second version model files corresponding to each first operator group, and the specific compilation process method can be any feasible method according to actual needs, and this disclosure will not describe it.
[0118] Step 7301 includes step 73011.
[0119] In step 73011, for each first operator group, a task scheduling execution command is generated to execute a task scheduling execution method (specifically, it may be the task scheduling execution method in the embodiment shown in Figure 6) based on the first version model file corresponding to the first operator group, the state information group of the first version model file corresponding to the first operator group, the expected occupancy rate corresponding to the first operator group, the division method information, and the second version model file corresponding to the first operator group.
[0120] In the embodiment of the present disclosure, a second version model file corresponding to each first operator group is generated, and the generated second version model file is used to generate a task scheduling execution command. In the execution stage, for tasks that cannot be processed in parallel in the task scheduling table, the neural network accelerator can process these tasks sequentially to ensure that these tasks can be processed normally. In this way, each task in the task scheduling table can be processed normally.
[0121] In one optional example, during the compilation stage, for each operator group in the plurality of operator groups (e.g., each first operator group described above), a multi-version model file can be compiled to generate a first version model file and a second version model file for realizing the same function, where the first version model file occupies a predetermined occupancy rate of the L1 SRAM, and the second version model file occupies the entire L1 SRAM.
[0122] In addition, the first version model file and the second version model file may both have corresponding state information groups, each of which includes a state corresponding to each functional unit of the neural network accelerator. The state corresponding to any functional unit has three possible states: exclusive, shared, and available. Among them, exclusive indicates that the operator group needs to exclusively occupy the L1 SRAM. Available indicates that the operator group does not need to use the functional unit and only occupies a portion of the L1 SRAM. Shared indicates that the operator group needs to use the functional unit and can share functional units other than the functional unit that it needs to use, and only occupies a portion of the L1 SRAM.
[0123] In this way, in the execution stage, by referring to the state information group of the first version model file and the expected occupancy rate corresponding to the first version model file, it is possible to efficiently and reliably determine which tasks satisfy the preset parallel execution conditions, and then these tasks can be processed in parallel, thereby improving the computational efficiency of the neural network accelerator, and for tasks that cannot be processed in parallel, these tasks can be processed sequentially.
[0124] Any of the methods for generating task scheduling execution instructions provided in the embodiments of the present disclosure may be executed by any suitable device having data processing capabilities, including, but not limited to, a terminal device, a server, etc. Alternatively, any of the methods for generating task scheduling execution instructions provided in the embodiments of the present disclosure may be executed by a processor, for example, the processor executes any of the methods for generating task scheduling execution instructions in the embodiments of the present disclosure by calling corresponding instructions stored in a memory. Hereinafter, redundant explanations will be omitted.
[0125] As can be understood by those skilled in the art, all or part of the steps of the above method embodiments can be completed by hardware associated with program instructions, and the program can be stored in a computer-readable storage medium, which, when executed, performs the steps comprising the above method embodiments, and the storage medium includes various media capable of storing program code, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. Exemplary Apparatus
[0126] 11 is a schematic diagram of the structure of a task scheduling execution device provided in one exemplary embodiment of the present disclosure. The device shown in FIG. 11 can be used to implement any of the above-mentioned task scheduling execution method embodiments of the present disclosure. The device shown in FIG. 11 includes a first determination module 1110, a second determination module 1120, and a first execution module 1130.
[0127] The first determination module 1110 is used to determine that there exists a first target task corresponding to a first version model file that, in a running state, occupies the computational resources of the neural network accelerator at a predetermined occupancy rate.
[0128] The second determination module 1120 is used to determine a set of first target tasks determined by the first determination module 1110 as a first task group, which satisfies a predetermined parallel execution condition based on a set of state information of the first version model file including states corresponding to each functional unit of the neural network accelerator in the execution state of the first version model file and the expected occupancy rate.
[0129] The first execution module 1130 is used to execute in parallel the first version model files respectively corresponding to the first target tasks in the first task group determined by the second determination module 1120 .
[0130] In one alternative example, the plurality of first target tasks satisfying the preset parallel execution condition is: In an information set consisting of state information groups respectively corresponding to a plurality of first target tasks, among all states corresponding to any of the functional units, one state is a used state and the remaining states are all idle states, or all of the states are idle states; and the total value of the scheduled occupancy rates corresponding to each of the plurality of first target tasks is smaller than a preset rate.
[0131] In one alternative example, as shown in FIG. 12, the device further comprises: an acquisition module 1101 for acquiring a task queue, where each task in the task queue corresponds to a neural network model; a third determination module 1103 for determining whether there is division scheme information corresponding to a target neural network model corresponding to a second target task, where the second target task is any task in the task queue acquired by the acquisition module 1101; a first processing module 1105 for dividing the second target task, obtaining K divided tasks, and adding the K divided tasks to a task scheduling table in response to the existence of division scheme information corresponding to the target neural network model determined by the third determination module 1103, wherein the K divided tasks correspond to K operator groups obtained by dividing the target neural network model according to the division scheme information corresponding to the target neural network model; a second processing module 1107 for adding a second target task to the task scheduling table in response to the absence of division scheme information corresponding to the target neural network model determined by the third determination module 1103; Specifically, the first determining module 1110 is for determining from the task scheduling table that a first target task corresponding to the first version model file exists.
[0132] In one alternative example, as shown in FIG. 12, the device further comprises: a fourth determination module 1140 for determining a second task group in the task scheduling table, the second task group including a set of third target tasks other than each first target task in the first task group determined by the second determination module 1120; a second execution module 1150 for executing second version model files corresponding to each third target task in the second task group determined by the fourth determination module 1140 in a predetermined order, wherein the second version model files fully occupy the computing resources in the execution state of the second execution module 1150.
[0133] 13 is a schematic diagram of the structure of a task scheduling execution instruction generating device provided in one exemplary embodiment of the present disclosure. The device shown in FIG. 13 can be used to realize any of the above-mentioned task scheduling execution instruction generating method embodiments of the present disclosure. The device shown in FIG. 13 includes a first generating module 1310, a second generating module 1320, and a third generating module 1330.
[0134] The first generation module 1310 is used to generate a first version model file corresponding to the first group of operators through a compilation process, and the first version model file occupies the computing resources of the neural network accelerator at a predetermined occupancy rate in an execution state.
[0135] The second generation module 1320 is used to generate a set of state information of the first version model file generated by the first generation module 1310 based on a set of functional units corresponding to the first set of operators, where the set of functional units corresponding to the first set of operators includes each functional unit of the neural network accelerator for operating the first set of operators, and the set of state information of the first version model file includes a state respectively corresponding to each functional unit of the neural network accelerator in the execution state of the first version model file.
[0136] The third generation module 1320 is used to generate a task scheduling execution command for executing the task scheduling execution method in the embodiment shown in FIG. 3 according to the first version model file generated by the first generation module 1310, the state information group of the first version model file generated by the second generation module 1320, and the scheduled occupancy rate.
[0137] In one alternative example, any functional unit of the neural network accelerator is taken as the target functional unit, and as shown in FIG. 14, the second generating module 1320: a first determination sub-module 13201 for determining, in response to the target functional unit being present in the functional unit group corresponding to the first operator group, that a state corresponding to the target functional unit in the state information group of the first version model file generated by the first generation module 1310 is a used state; and a second determination submodule 13203 for determining, in response to the target functional unit not being present in the functional unit group corresponding to the first operator group, that the state corresponding to the target functional unit in the state information group of the first version model file generated by the first generation module 1310 is an idle state.
[0138] In one alternative example, as shown in FIG. 14, the device further comprises: a third processing module 1301 for dividing the neural network model into K operator groups corresponding to different functional unit groups and recording corresponding division method information when functional unit groups corresponding to each operator unit in the neural network model are not completely the same before generating a first version model file corresponding to the first operator group through compilation processing, wherein the functional unit group corresponding to any one of the operator groups includes each functional unit in the neural network accelerator for operating the operator group; a fifth determination module 1303 for determining each operator group in at least a portion of the K operator groups obtained by division by the third processing module 1301 as one first operator group; Specifically, the fifth determination module 1303 includes a third generation module 1320 for generating, for each first operator group determined by the fifth determination module 1303, a task scheduling execution command for executing the task scheduling execution method in the embodiment shown in FIG. 4 based on the first version model file corresponding to the first operator group, the state information group of the first version model file corresponding to the first operator group, the expected occupancy rate corresponding to the first operator group, and the division method information.
[0139] In one alternative example, as shown in FIG. 14, the device further comprises: a fourth generating module 1311 for generating second version model files corresponding to each first operator group determined by the fifth determining module 1303 through a compiling process, wherein the second version model files fully occupy the computing resources in an execution state; Specifically, for each first operator group determined by the fifth determination module 1303, a third generation module 1320 is included for generating a task scheduling execution command for executing the task scheduling execution method in the embodiment shown in FIG. 6 based on the first version model file corresponding to the first operator group, the state information group of the first version model file corresponding to the first operator group, the expected occupancy rate corresponding to the first operator group, the division method information, and the second version model file corresponding to the first operator group.
[0140] In the device of the present disclosure, the various selective examples, selective embodiments, and selective examples disclosed above can all be flexibly selected and combined as needed to achieve corresponding functions and effects, and will not be listed one by one in the present disclosure. Exemplary Electronic Devices
[0141] 15 shows a block diagram of an electronic device 1500 according to an embodiment of the present disclosure. The electronic device 1500 includes one or more processors 1510 and a memory 1520.
[0142] Processor 1510 may be a central processing unit (CPU) or other form of processing device having data processing and / or instruction execution capabilities, and may control other components in electronic device 1500 to perform desired functions.
[0143] The memory 1520 may include one or more computer program products, which may include various forms of computer-readable storage media, such as, for example, volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, and the like. The computer-readable storage medium may store one or more computer program instructions. The processor 1510 may execute the one or more computer program instructions to realize any of the above-described embodiments of the method of the present disclosure, for example, determining that there is a first target task corresponding to a first version model file that occupies a predetermined occupancy rate of the computing resources of the neural network accelerator in an execution state; determining a set of first target tasks that satisfy a predetermined parallel execution condition as a first task group based on a set of state information of the first version model file, the set including states corresponding to each functional unit of the neural network accelerator in the execution state of the first version model file, and the predetermined occupancy rate; and executing in parallel the first version model files corresponding to each first target task in the first task group.
[0144] In one optional example, the electronic device 1500 may further include input devices 1530 and output devices 1540, these components being interconnected via a bus system and / or other form of connection (not shown).
[0145] The input device 1530 may further include, for example, a keyboard, a mouse, etc. The output device 1540 can output various types of information to the outside, and may include, for example, a display, a speaker, a printer, a communication network, and a remote output device connected thereto.
[0146] Of course, for simplicity, Figure 15 shows only some of the components in electronic device 1500 that are relevant to the present disclosure, omitting components such as buses, input / output interfaces, etc. Additionally, electronic device 1500 may further include any other appropriate components depending on the particular situation. Exemplary Computer Program Products and Computer-Readable Storage Media
[0147] In addition to the above-described methods and apparatuses, an embodiment of the present disclosure may also be a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform any of the steps of the embodiment of the method of the present disclosure described in the "exemplary method" section of this specification, such as determining that, in an execution state, there is a first target task corresponding to a first version model file that occupies a predetermined occupancy rate of the computational resources of a neural network accelerator; determining, as a first task group, a set of first target tasks that satisfy a predetermined parallel execution condition based on a set of state information of the first version model file, including states corresponding to each functional unit of the neural network accelerator, in the execution state of the first version model file, and the predetermined occupancy rate; and executing in parallel the first version model files corresponding to each first target task in the first task group.
[0148] The computer program product may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on a user's computing device, partially on a user's device, as separate software packages, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or a server.
[0149] Furthermore, an embodiment of the present disclosure may be a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, cause the processor to perform the steps of any of the embodiment of the method according to the present disclosure described in the "Exemplary Method" section above of this specification.
[0150] The computer-readable storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a readable signal medium or a computer-readable storage medium. The computer-readable storage medium may include, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include an electrical connection having one or more leads, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a compact disc read-only memory (CD-ROM), an optical storage element, a magnetic storage element, or any suitable combination of the above.
[0151] Although the basic principles of the present disclosure have been described above with reference to specific embodiments, the benefits, advantages, effects, etc. mentioned in the present disclosure are not limiting but merely illustrative, and not all embodiments of the present disclosure necessarily possess these benefits, advantages, effects, etc. Furthermore, the specific details disclosed above are not limiting but merely serve to illustrate and facilitate understanding, and the above details do not necessarily limit the present disclosure to be realized using the above specific details.
[0152] Those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure also intends to include these modifications and variations.
Claims
1. determining that there exists a first target task corresponding to a first version model file that occupies a predetermined occupancy rate of the computational resources of the neural network accelerator in an execution state; determining, as a first task group, a set of the first target tasks that satisfy a preset parallel execution condition based on a state information group of the first version model file, the state information group including states corresponding to each functional unit of the neural network accelerator in an execution state of the first version model file, and the expected occupancy rate; and executing in parallel the first version model files corresponding to the first target tasks in the first task group.
2. The plurality of first target tasks satisfying the preset parallel execution condition is In an information set consisting of the state information group corresponding to each of the first target tasks in the plurality of first target tasks, one state among all states corresponding to any of the functional units is a used state and the remaining states are all idle states, or all of the states are idle states; The method of claim 1, wherein the total value of the scheduled occupancy ratios corresponding to each of the first target tasks in the plurality of first target tasks is smaller than a predetermined ratio.
3. obtaining a task queue, each task in the task queue corresponding to a neural network model; determining whether there is segmentation scheme information corresponding to a target neural network model corresponding to a second target task, the second target task being any task in the task queue; In response to the existence of division scheme information corresponding to the target neural network model, dividing the second target task to obtain K division tasks, and adding the K division tasks to a task scheduling table, wherein the K division tasks correspond to K operator groups obtained by dividing the target neural network model according to the division scheme information corresponding to the target neural network model; adding the second target task to the task scheduling table in response to the absence of partitioning scheme information corresponding to the target neural network model; The step of determining that a first target task corresponding to the first version model file exists includes: The method of claim 1 or 2, further comprising determining from the task scheduling table that the first target task corresponding to a first version model file exists.
4. determining a second task group in the task scheduling table, the second task group including a set of third target tasks other than each of the first target tasks in the first task group; 4. The method of claim 3, further comprising: executing second version model files corresponding to the third target tasks in the second task group in a predetermined order, wherein the second version model files fully occupy the computing resources in an execution state.
5. a step of generating a first version model file corresponding to a first group of operators through a compilation process, wherein the first version model file occupies a predetermined occupancy rate of the computational resources of the neural network accelerator in an execution state; generating a set of state information of the first version model file based on a set of functional units corresponding to the first set of operators, the set of functional units corresponding to the first set of operators including each functional unit of the neural network accelerator for operating the first set of operators, and the set of state information of the first version model file including a state corresponding to each functional unit of the neural network accelerator in an execution state of the first version model file; and generating a task scheduling execution command for executing the task scheduling execution method according to claim 1 or 2, based on the first version model file, a set of status information of the first version model file, and the scheduled occupancy rate.
6. generating a set of state information of the first version model file based on a set of functional units corresponding to the first set of operators, using any of the functional units of the neural network accelerator as a target functional unit, determining, in response to the target functional unit being present in the functional unit group corresponding to the first operator group, that a state corresponding to the target functional unit in the state information group of the first version model file is in a used state; and determining, in response to the target functional unit not being present in the functional unit group corresponding to the first operator group, that a state corresponding to the target functional unit in the state information group of the first version model file is an idle state.
7. Before generating a first version model file corresponding to the first group of operators by a compilation process, When the functional unit groups corresponding to the operator units in the neural network model are not completely the same, dividing the neural network model into K operator groups corresponding to different functional unit groups, and recording corresponding division method information, wherein the functional unit group corresponding to any one of the operator groups includes each functional unit for operating the operator group in the neural network accelerator; and further comprising a step of setting each of the operator groups in at least some of the K operator groups as one of the first operator groups, a step of generating a task scheduling execution command for executing the task scheduling execution method according to claim 1 or 2 based on the first version model file, the state information group of the first version model file, and the scheduled occupancy rate, the step comprising:
6. The method of claim 5, further comprising: generating, for each of the first operator groups, a task scheduling execution instruction for executing the task scheduling execution method of claim 3, based on the first version model file corresponding to the first operator group, a state information group of the first version model file corresponding to the first operator group, the expected occupancy rate corresponding to the first operator group, and the division method information.
8. generating second version model files corresponding to the first operator groups through a compiling process, the second version model files fully occupying the computing resources in an execution state; generating a task scheduling execution command for executing the task scheduling execution method according to claim 3 based on the first version model file corresponding to each of the first operator groups, a state information group of the first version model file corresponding to each of the first operator groups, the expected occupancy rate corresponding to each of the first operator groups, and the division method information, the step of:
8. The method of claim 7, further comprising: generating, for each of the first operator groups, a task scheduling execution instruction for executing the task scheduling execution method of claim 4, based on the first version model file corresponding to the first operator group, a state information group of the first version model file corresponding to the first operator group, the expected occupancy rate corresponding to the first operator group, the division method information, and the second version model file corresponding to the first operator group.
9. a first determination module for determining that there exists a first target task corresponding to a first version model file that occupies a predetermined occupancy rate of a computational resource of the neural network accelerator in an execution state; a second determination module for determining, as a first task group, a set of the first target tasks determined by the first determination module that satisfies a preset parallel execution condition, based on a state information group of the first version model file, the state information group including states corresponding to each functional unit of the neural network accelerator in an execution state of the first version model file, and the expected occupancy rate; a first execution module for executing in parallel the first version model files corresponding to the first target tasks in the first task group determined by the second determination module.
10. a first generation module for generating a first version model file corresponding to a first group of operators through a compilation process, the first version model file occupying a predetermined occupancy rate of a computation resource of a neural network accelerator in an execution state; a second generation module for generating a set of state information of the first version model file generated by the first generation module based on a set of functional units corresponding to the first set of operators, the set of functional units corresponding to the first set of operators including each functional unit of the neural network accelerator for operating the first set of operators, and the set of state information of the first version model file including a state corresponding to each functional unit of the neural network accelerator in an execution state of the first version model file; a third generation module for generating a task scheduling execution command for executing the task scheduling execution method according to claim 1 or 2, based on the first version model file generated by the first generation module, a set of status information of the first version model file generated by the second generation module, and the scheduled occupancy rate; A task scheduling execution instruction generation device.
11. A computer-readable storage medium storing a computer program used to execute the task scheduling execution method according to any one of claims 1 to 4, or to execute the task scheduling execution instruction generation method according to any one of claims 5 to 8.
12. a processor; a memory for storing instructions executable by said processor; The processor reads the executable instructions from the memory and executes the instructions to realize the task scheduling execution method described in any one of claims 1 to 4, or to realize the task scheduling execution instruction generation method described in any one of claims 5 to 8. An electronic device used to realize this method.
13. A computer program product, the computer program product realizing the task scheduling execution method according to any one of claims 1 to 4, or the method for generating a task scheduling execution instruction according to any one of claims 5 to 8, when instructions in the computer program product are executed by a processor.