Task Scheduling Method, Device, Electronic Device and Storage Medium

By combining parameters such as the number of available computing units for task scheduling, the problems of unbalanced computing load and low efficiency in artificial intelligence chips are solved, more efficient load balancing and computing efficiency are achieved, and the applicable scenarios of task scheduling are expanded.

CN118963967BActive Publication Date: 2025-06-17SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411225259.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2025-06-17
Estimated Expiration
2044-09-02

AI Technical Summary

Technical Problem

During the task scheduling process, artificial intelligence chips have problems such as unbalanced computing load and low computing efficiency, resulting in the processor being idle or the computing efficiency is not high.

Method used

By combining the number of available calculation units and other parameters for task scheduling, the coordinates of the target data block that matches the number of currently available calculation units are determined, and the corresponding sub-computation tasks are performed. This method provides data block scheduling configuration parameters by the main processor to ensure that each processor has better computing efficiency.

Benefits of technology

It achieves better load balancing effect, avoids the processor being idle, improves the computing efficiency of artificial intelligence operators, and expands the applicable scenarios for operator task scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118963967B_ABST
    Figure CN118963967B_ABST
Patent Text Reader

Abstract

The present invention provides a task scheduling method, device, electronic device and storage medium, wherein the method includes: in response to a task scheduling instruction, obtaining different data block scheduling configuration parameters after preprocessing of the target computing task of the artificial intelligence operator; each data block scheduling configuration parameter is provided by a main control processor; based on each data block scheduling configuration parameter, determining multiple target data block coordinates that match the number of currently available computing units, and executing sub-computing tasks that process multiple target data block coordinates corresponding to the target data blocks. The present invention combines the number of available computing units and other parameters for scheduling, which has a better load balancing effect, and can also avoid the processor from being idle as much as possible, ensuring that each processor has better computing efficiency, improving the performance of the artificial intelligence operator, and at the same time, combining the heterogeneous framework scenario of interaction between the main control processor at the front end and the artificial intelligence chip at the back end, it also expands the applicable scenarios of operator task scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a task scheduling method, device, electronic device, and storage medium. Background Art

[0002] Currently, operators with large computational workloads such as matrix multiplication (Matrix matmul) and convolution are common operator types on artificial intelligence chips. And when an artificial intelligence chip runs the training and inference processes of an artificial intelligence model, it usually needs to process the computational tasks of relevant operators. Therefore, in order to accelerate the training efficiency or inference efficiency of the artificial intelligence model, a task scheduling system can be pre-set on the artificial intelligence chip, and the task scheduling system is used to schedule the computational tasks of the operators called in model inference or model training. Therefore, how to achieve efficient and load-balanced task scheduling has become a key problem to be solved urgently.

[0003] In related technologies, an artificial intelligence chip first divides a computational task row by row or column by column to obtain multiple data blocks each containing sub-computational tasks, and then uses each computing unit to calculate the corresponding data block respectively; due to the simple way of dividing the computational task, the task amounts in the multiple data blocks obtained by the division are unevenly distributed. In this way, when both the task division process and the task scheduling process are implemented on the artificial intelligence chip, there are technical problems that the computational load of each computing unit is unbalanced and the computational efficiency is very low during the process of calculating the sub-computational tasks in the corresponding data block. Summary of the Invention

[0004] The present invention provides a task scheduling method, device, electronic device, and storage medium, which are used to solve the defects of unbalanced computational load and very low computational efficiency in the task scheduling process of existing artificial intelligence operators. Scheduling in combination with parameters such as the number of available computing units has a better load balancing effect, and at the same time can also avoid the processor being idle as much as possible, ensure that each processor has better computational efficiency, improve the performance of artificial intelligence operators, and at the same time expand the applicable scenarios of operator task scheduling in combination with the heterogeneous framework scenario of the interaction between the front-end main control processor and the back-end artificial intelligence chip.

[0005] The present invention provides a task scheduling method, which is applied to an artificial intelligence chip running an artificial intelligence algorithm, and the method includes the following steps.

[0006] In response to a task scheduling instruction, obtain different data block scheduling configuration parameters after preprocessing the target computational task of the artificial intelligence operator; each of the data block scheduling configuration parameters is provided by the main control processor.

[0007] Based on each of the data block scheduling configuration parameters, determine a plurality of target data block coordinates that match the number of currently available computing units, and execute sub-computation tasks for the target data blocks corresponding to each of the plurality of target data block coordinates.

[0008] According to a task scheduling method provided by the present invention, the determining, based on each of the data block scheduling configuration parameters, a plurality of target data block coordinates that match the number of currently available computing units includes: performing the following steps for each thread block having the same number as the currently available computing units: determining a thread block family number of the thread block family based on a linear index of the thread block and a thread block family size of a thread block family corresponding to the thread block; determining a scheduling area size corresponding to the thread block family based on the thread block family number and each of the data block scheduling configuration parameters; and determining the target data block coordinates required to be processed by the available computing units corresponding to the thread block based on the scheduling area size and a position offset of the thread block in the corresponding thread block family.

[0009] According to a task scheduling method provided by the present invention, the determining the thread block family number of the thread block family based on the linear index of the thread block and the thread block family size of the thread block family corresponding to the thread block includes: determining the thread block family size based on a row size of the thread block family and a column size of the thread block family; and determining the thread block family number based on the linear index and the thread block family size.

[0010] According to a task scheduling method provided by the present invention, the determining the scheduling area size corresponding to the thread block family based on the thread block family number and each of the data block scheduling configuration parameters includes: in a case where each of the data block scheduling configuration parameters includes the number of thread block families included in each row when a plurality of the thread blocks are scheduled in a row direction, determining a row number where the thread block family is located based on a division operation result of the thread block family number and the number of thread block families included in each row; determining a column number where the thread block family is located based on a remainder operation result of the thread block family number and the number of thread block families included in each row; and determining the scheduling area size based on the row number and the column number; or, in a case where each of the data block scheduling configuration parameters includes the number of thread block families included in each column when a plurality of the thread blocks are scheduled in a column direction, determining a column number where the thread block family is located based on a division operation result of the thread block family number and the number of thread block families included in each column; determining a row number where the thread block family is located based on a remainder operation result of the thread block family number and the number of thread block families included in each column; and determining the scheduling area size based on the row number and the column number.

[0011] A task scheduling method provided by the present invention, determining the target data block coordinates that the thread block needs to process corresponding to the available computing units based on the size of the scheduling area and the position offset of the thread block in the corresponding thread block family, includes: determining the row coordinates of the data block that the thread block needs to process corresponding to the available computing units based on the row size of the data block that the thread block needs to process corresponding to the available computing units, the row size of the thread block family, the row number of the scheduling area size, and the row offset of the thread block in the corresponding thread block family; determining the column coordinates of the data block that the thread block needs to process corresponding to the available computing units based on the column size of the data block that the thread block needs to process corresponding to the available computing units, the column size of the thread block family, the column number of the scheduling area size, and the column offset of the thread block in the corresponding thread block family; and determining the data block row coordinates and the data block column coordinates as the target data block coordinates that the thread block needs to process corresponding to the available computing units.

[0012] A task scheduling method provided by the present invention, responding to a task scheduling instruction to obtain different data block scheduling configuration parameters after preprocessing the target computing task of an artificial intelligence operator, includes:

[0013] Responding to the task scheduling instruction, based on the target task scheduling parameters and the target data storage instruction carried in the task scheduling instruction, and the mapping relationship between the task scheduling parameter - data storage instruction - data block scheduling configuration parameter data packet, obtaining the target data block scheduling configuration parameter data packet, where the target data block scheduling configuration parameter data packet contains the different data block scheduling configuration parameters; wherein, the mapping relationship is established based on the different data block scheduling configuration parameter data packets and different data storage instructions determined by the main control processor after preprocessing the computing tasks of different types of artificial intelligence operators in advance.

[0014] A main control processor in a task scheduling method provided by the present invention is used to perform the following steps for each computing task: aligning the computing task based on the data block size, the thread block family size, and the data block scheduling parameters in sequence to obtain an updated computing task; splitting the updated computing task based on the data block size, the data block scheduling direction, and the number of current available computing units to obtain the data block scheduling configuration parameter data packet corresponding to the computing task; and storing the data block scheduling configuration parameter data packet in an immediate number manner or a register manner to generate the data storage instruction corresponding to the computing task.

[0015] A task scheduling method provided by the present invention, the sub-computation tasks of executing and processing the target data blocks corresponding to the respective multiple target data block coordinates include: when all of the multiple target data block coordinates belong to the target data block coordinates of the current batch and do not exceed the target scale of the target computation task, processing each of the sub-computation tasks; based on a processing completion flag, re-determining the target data block coordinates of the next batch among the multiple data blocks into which the target computation task is segmented; taking all the data block coordinates included in the target data block coordinates of the next batch as new target data block coordinates, and repeating the above steps; until there is at least one invalid data block exceeding the target scale among the multiple newly re-determined target data block coordinates.

[0016] The present invention also provides a task scheduling device, which is applied to an artificial intelligence chip running an artificial intelligence algorithm. The device includes: a parameter acquisition unit and a task scheduling unit.

[0017] The parameter acquisition unit is configured to, in response to a task scheduling instruction, acquire different data block scheduling configuration parameters after preprocessing of a target computation task of an artificial intelligence operator; each of the data block scheduling configuration parameters is provided by a main control processor.

[0018] The task scheduling unit is configured to, based on each of the data block scheduling configuration parameters, determine a plurality of target data block coordinates matching the number of currently available computing units, and execute and process the sub-computation tasks of the target data blocks corresponding to the respective multiple target data block coordinates.

[0019] The present invention also provides an electronic device, including a main control processor and an artificial intelligence chip running an artificial intelligence algorithm connected in a heterogeneous framework form; the artificial intelligence chip is configured to, in response to a task scheduling instruction, acquire different data block scheduling configuration parameters after preprocessing of a target computation task of an artificial intelligence operator; based on each of the data block scheduling configuration parameters, determine a plurality of target data block coordinates matching the number of currently available computing units, and execute and process the sub-computation tasks of the target data blocks corresponding to the respective multiple target data block coordinates; the main control processor is configured to provide each of the data block scheduling configuration parameters to the artificial intelligence chip.

[0020] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the task scheduling method described in any one of the above is implemented.

[0021] The task scheduling method, device, electronic device, and storage medium provided by the present invention. In the task scheduling method, when the artificial intelligence chip responds to a task scheduling instruction, it first obtains the scheduling configuration parameters of different data blocks after preprocessing the target computing task of the artificial intelligence operator, and then further determines multiple target data block coordinates that match the number of currently available computing units based on the scheduling configuration parameters of each data block, and executes the sub-computing tasks for processing the target data blocks corresponding to the multiple target data block coordinates respectively. Since the scheduling configuration parameters of different data blocks after preprocessing the target computing task of the artificial intelligence operator are provided by the main control processor, the task scheduling purpose of the artificial intelligence operator task is quickly and efficiently achieved by calculating the data block coordinates of the data blocks to be processed for each available computing unit in the artificial intelligence chip based on the different data block scheduling configuration parameters provided by the main control processor at the front end for the target computing task of the artificial intelligence operator and executing the corresponding sub-computing tasks. The entire task scheduling process is scheduled in combination with parameters such as the number of available computing units, having a better load balancing effect, and also being able to avoid the processor being idle as much as possible, ensuring that each processor has better computing efficiency, improving the performance of the artificial intelligence operator. At the same time, the heterogeneous framework scenario of the interaction between the main control processor at the front end and the artificial intelligence chip at the back end also expands the applicable scenarios of the operator task scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0023] Figure 1 is a schematic flowchart of the task scheduling method provided by the present invention.

[0024] Figure 2A is one of the schematic diagrams of the target scheduling direction provided by the present invention.

[0025] Figure 2B is the second schematic diagram of the target scheduling direction provided by the present invention.

[0026] Figure 3 is a schematic diagram of the relationship between the target data block scheduling parameters provided by the present invention and the linear index of the thread block and the data block coordinates respectively.

[0027] Figure 4 is a schematic diagram of the block division of the matrix multiplication computing task provided by the present invention.

[0028] Figure 5 is a schematic diagram of the dynamic scheduling provided by the present invention.

[0029] Figure 6 It is a schematic structural diagram of the task scheduling device provided by the present invention.

[0030] Figure 7 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments

[0031] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0032] In the embodiments of the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B may be singular or plural. In the written description of the present invention, the character " / " generally represents an "or" relationship between the associated objects before and after. In addition, it should be noted that the serial numbers assigned to the objects described in the present invention itself, such as "first", "second", etc., are only used to distinguish the described objects and do not have any sequential or technical meanings.

[0033] Currently, operators with large computational amounts such as Matrix matmul and Convolution are common operator types on artificial intelligence chips. Moreover, during the training and inference processes of running artificial intelligence models on artificial intelligence chips, it is usually necessary to process the computational tasks of relevant operators. Therefore, in order to accelerate the training efficiency or inference efficiency of artificial intelligence models, a task scheduling system can be preset on the artificial intelligence chip, and the task scheduling system is used to schedule the computational tasks of the operators called during model inference or model training. Therefore, how to achieve efficient and load-balanced task scheduling has become a key problem that needs to be solved urgently.

[0034] In the related art, an AI chip first divides a computing task row by row or column by column to obtain multiple data blocks each containing a sub-computing task, and then uses each computing unit to calculate the corresponding data block respectively; due to the simple way of dividing the computing task, the task amounts in the multiple obtained data blocks are unevenly distributed. In this way, when both the task division process and the task scheduling process are implemented on the AI chip, there are technical problems that the computing loads of each computing unit are unbalanced and the computing efficiency is very low during the process of calculating the sub-computing tasks in the corresponding data block. For example, there is a situation where a certain computing unit only calculates some rows or columns of a matrix, resulting in idle computing resources, or in order to avoid idle computing resources, the computing within the data block is not efficient enough.

[0035] To solve the above technical problems, the present invention provides a task scheduling method, device, electronic device and storage medium, which perform scheduling in combination with parameters such as the number of available computing units to obtain a better load balancing effect, and at the same time can also avoid the processor being in an idle state as much as possible, ensure that each processor has better computing efficiency, improve the performance of AI operators, and at the same time, in combination with the heterogeneous framework scenario of the interaction between the front-end main control processor and the back-end AI chip, also expand the applicable scenarios of operator task scheduling.

[0036] The following combines Figures 1 - 7 Describe the task scheduling method, device, electronic device and storage medium of the present invention. The execution subject of the task scheduling method is the back-end AI chip. This AI chip has at least the task scheduling function and can be called an AI accelerator or a computing card, that is, a module specifically used to process a large number of computing tasks in AI applications, such as GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), NPU (Neural network Processing Unit), DPU (Deeplearning Processing Unit), APU (Accelerated Processing Unit), and GPGPU (General-Purpose computing on Graphics Processing Unit), etc. Further, the execution subject of the task scheduling method can also be applied to the task scheduling device set in the AI chip, and this task scheduling device can be implemented by software, hardware or a combination of both. The following takes the execution subject of the task scheduling method as an AI chip with at least the task scheduling function as an example to describe the task scheduling method.

[0037] To facilitate the understanding of the task scheduling method provided by the embodiments of the present invention, the following will detail the task scheduling method provided by the present invention through several exemplary embodiments below. It can be understood that these exemplary embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0038] Referring to Figure 1 , which is a schematic flowchart of the task scheduling method provided by the present invention. As Figure 1 shown, the task scheduling method includes the following steps 110 and 120.

[0039] Step 110: In response to a task scheduling instruction, obtain the scheduling configuration parameters of different data blocks after preprocessing of the target computing task of the artificial intelligence operator; each data block scheduling configuration parameter is provided by the main control processor.

[0040] Among them, the task scheduling instruction can be an instruction automatically generated when the artificial intelligence operator is called during the artificial intelligence model training or inference process.

[0041] The artificial intelligence operator can be an operator with a large amount of task calculation, such as a matrix multiplication operator, a convolution operator, and an activation operator, etc.

[0042] The target computing task is determined by the artificial intelligence operator. For example, when the artificial intelligence operator is a matrix multiplication operator, the target computing task is a matrix multiplication task; or when the artificial intelligence operator is a convolution operator, the target computing task is a convolution task.

[0043] The number of artificial intelligence operators is the same as and corresponds one-to-one with the number of preprocessing times.

[0044] The target computing task of each artificial intelligence operator is preprocessed to obtain a set of data block scheduling configuration parameters respectively, and the preprocessing process of each target computing task runs in the front-end main control processor. The purpose is to simplify the task scheduling process in the artificial intelligence chip and assist the artificial intelligence chip to quickly and accurately calculate the data block coordinate results obtained by splitting the target computing task.

[0045] Each preprocessing can include but is not limited to alignment operations, task splitting operations, and storage instruction generation operations.

[0046] The main control processor can be a processor with other functions such as at least alignment operations, task splitting operations, and storage instruction generation operations, such as a main control chip, a main control module, a Central Processing Unit (CPU), an Advanced RISC Machine (ARM) processor, and an X86 processor, etc.

[0047] The different data block scheduling configuration parameters may include, but are not limited to, the number of data blocks in each row, the number of data blocks in each column, and the number of data blocks in the surface area after the target computing task of the artificial intelligence operator is segmented based on the currently available computing units, as well as the number of thread block families in each row and the number of thread blocks in each column in different scheduling directions of multiple thread blocks; the number of thread blocks is the same as the number of currently available computing units.

[0048] Specifically, when the artificial intelligence chip responds to a task scheduling instruction, it first obtains the different data block scheduling configuration parameters after preprocessing the target computing task of the artificial intelligence operator. The obtaining method can directly read a set of data block scheduling configuration parameters corresponding to the target computing task from its memory when a set of data block scheduling configuration parameters corresponding to different computing tasks provided by the main control processor are pre-stored in the memory of the artificial intelligence chip; or, when a set of data block scheduling configuration parameters corresponding to different computing tasks provided by the main control processor are pre-stored in the cloud server, it can also call a set of data block scheduling configuration parameters corresponding to the target computing task from the cloud server; or, the artificial intelligence chip can also feedback a parameter obtaining instruction carrying the target computing task to the front-end main control processor through a communication transmission method or a copy method, so that the main control processor preprocesses the target computing task to obtain a corresponding set of data block scheduling configuration parameters and sends this set of data block scheduling configuration parameters to the artificial intelligence chip. The present invention does not make specific limitations on this.

[0049] Step 120: Based on each data block scheduling configuration parameter, determine multiple target data block coordinates that match the number of currently available computing units, and execute sub-computing tasks for processing the target data blocks corresponding to each of the multiple target data block coordinates.

[0050] Among them, the multiple target data block coordinates may be partial data block coordinates among the multiple data block coordinates obtained after task segmentation of the target computing task.

[0051] The artificial intelligence chip is pre-configured with a Graphics Processing Unit (GPU) architecture. The GPU architecture includes a series of Streaming Multiprocessors (SMs), and each SM is composed of multiple streaming processors, cores or threads; for example, when the GPU architecture contains 132 SMs and each SM has 64 streaming processors, the total number of streaming processors can be as high as 8448; this ensures that the artificial intelligence chip has the ability to process massive data and perform parallel computing.

[0052] Specifically, the artificial intelligence chip can calculate the coordinates of multiple target data blocks that need to be processed synchronously in each cycle based on the scheduling configuration parameters of different data blocks after preprocessing of the target computing tasks; and the number of target data block coordinates that need to be processed synchronously in each cycle is the same as the number of currently available computing units (streaming multiprocessors) in the artificial intelligence chip. For example, when the number of currently available computing units is 16, each cycle requires synchronous calculation of the sub-computing tasks of each target data block of the 16 target data block coordinates.

[0053] It should be noted that, considering that a large number of data blocks divided by the target computing task can form a data block area, and multiple target data blocks can be selected in sequence by row, column or other specified methods in the data block area for synchronous processing, until at least one of the multiple target data blocks selected at a certain time does not belong to the data block area, only some target data blocks belonging to the data block area need to be synchronously processed; thereby completing the task scheduling process of the artificial intelligence operator.

[0054] The task scheduling method provided by the present invention, when responding to the task scheduling instruction, the artificial intelligence chip first obtains the different data block scheduling configuration parameters after preprocessing of the target computing task of the artificial intelligence operator, and then further determines multiple target data block coordinates matching the number of currently available computing units based on the scheduling configuration parameters of each data block, and executes the sub-computation tasks of processing the target data blocks corresponding to the multiple target data block coordinates. Since the different data block scheduling configuration parameters after preprocessing of the target computing task of the artificial intelligence operator are provided by the main control processor, the main control processor at the front end calculates the data block coordinates of the required processing data block for each available computing unit in the artificial intelligence chip and executes the corresponding sub-computation tasks, so as to quickly and efficiently achieve the task scheduling purpose of the artificial intelligence operator task, and the whole task scheduling process is combined with the number of available computing units and other parameters for scheduling, which has a better load balancing effect, and can also avoid the processor from being in an idle state as much as possible, ensuring that each processor has better computing efficiency, improving the performance of the artificial intelligence operator, and at the same time, the heterogeneous framework scenario of the interaction between the main control processor at the front end and the artificial intelligence chip at the back end also expands the applicable scenario of operator task scheduling.

[0055] Based on the above Figure 1 In the task scheduling method shown, in an exemplary embodiment, in step 120, the artificial intelligence chip determines multiple target data block coordinates that match the number of currently available computing units based on the scheduling configuration parameters of each data block, and the specific determination process can be implemented through the following steps.

[0056] The following steps are performed for each thread block with the same number of available compute units:

[0057] First, based on the linear index of the thread block and the thread block family size of the thread block in the corresponding thread block family, determine the thread block family number of the thread block in the corresponding thread block family; further, based on the thread block family number of the thread block in the corresponding thread block family and each data block scheduling configuration parameter, determine the scheduling area size corresponding to the thread block family; then, based on the scheduling area size and the position offset of the thread block in the corresponding thread block family, determine the target data block coordinates that the available computing units corresponding to the thread block need to process.

[0058] Specifically, the artificial intelligence chip can initiate the same number of thread blocks according to the number of currently available computing units, and multiple thread blocks can be presented in a matrix manner; for the convenience of subsequent task scheduling to be more convenient and efficient, the multiple thread blocks in matrix form can be flattened into a one-dimensional thread block array. Therefore, each thread block in this one-dimensional thread block array has a unique linear index (linear_index), and each thread block resides on an available computing unit inside the artificial intelligence chip and is used to execute the sub-computation task corresponding to a target data block.

[0059] The number of thread block families can be multiple, and each thread block family can be composed of at least 1 thread block; for example, the number of thread blocks in each thread block family can be 1, 2, 4, 8, etc. Therefore, when the linear index of each thread block is determined and the size is known, the size of each thread block family is also determined.

[0060] Based on this, for multiple thread blocks with the same number as the number of currently available computing units, when the linear index of one of the thread blocks and the thread block family size of the thread block family corresponding to the thread block are known, the thread block family number of the thread block family can be determined, and then combined with each data block scheduling configuration parameter, determine the scheduling area size corresponding to the thread block family. The scheduling area size can be an area composed of the size of the row and the size of the column when scheduling in the row direction. At this time, the size of the row is determined based on the data block scheduling configuration parameter of the number of thread block families contained in each row; or, it can also be an area composed of the size of the column and the size of the row when scheduling in the column direction. At this time, the size of the column is determined based on the data block scheduling configuration parameter of the number of thread block families contained in each column.

[0061] Since the position offset of each thread block in its respective thread block family is determined, based on the position offset of one of the thread blocks in the corresponding thread block family and the scheduling area size determined based on the thread block, the target data block coordinates that the available computing units corresponding to the thread block need to process can be determined.

[0062] In an exemplary embodiment, the artificial intelligence chip determines the thread block family number of a thread block family based on the linear index of the thread block and the thread block family size of the thread block family corresponding to the thread block. The specific determination process can be implemented through the following steps.

[0063] First, determine the thread block family size based on the row size and column size of the thread block family; then, determine the thread block family number based on the linear index and the thread block family size.

[0064] Specifically, the thread block family size is calculated by Equation (1).

[0065] cluster_size = cluster_row * cluster_col (1).

[0066] In Equation (1), cluster_size represents the thread block family size, cluster_row represents the row size of the thread block family, cluster_col represents the column size of the thread block family, and * represents the multiplication operation.

[0067] The thread block family number is calculated by Equation (2).

[0068] cluster_id = linear_index / cluster_size (2).

[0069] In Equation (1), cluster_id represents the thread block family number, linear_index represents the linear index, and / represents the integer division operation.

[0070] In an exemplary embodiment, the artificial intelligence chip determines the scheduling area size corresponding to a thread block family based on the thread block family number and each data block scheduling configuration parameter. The determination process is implemented through the following steps.

[0071] When each data block scheduling configuration parameter includes the number of thread block families contained in each row when multiple thread blocks are scheduled in the row direction, first determine the row number where the thread block family is located based on the integer division operation result of the thread block family number and the number of thread block families contained in each row; further, determine the column number where the thread block family is located based on the remainder operation result of the thread block family number and the number of thread block families contained in each row; then, determine the scheduling area size based on the row number and the column number; or, when each data block scheduling configuration parameter includes the number of thread block families contained in each column when multiple thread blocks are scheduled in the column direction, first determine the column number where the thread block family is located based on the integer division operation result of the thread block family number and the number of thread block families contained in each column; then further determine the row number where the thread block family is located based on the remainder operation result of the thread block family number and the number of thread block families contained in each column; then, determine the scheduling area size based on the row number and the column number.

[0072] Specifically, for multiple thread blocks equal in number to the currently available computing units, the different data block scheduling configuration parameters after preprocessing the target computing tasks of the artificial intelligence operator may include the number of thread block families contained in each row and the number of thread blocks contained in each column in the case of scheduling multiple thread blocks in the row direction, or may also include the number of thread block families contained in each column and the number of thread blocks contained in each row in the case of scheduling multiple thread blocks in the column direction.

[0073] Taking the number of thread block families per row cluster_number_per_row in the case of scheduling multiple thread blocks in the row direction as an example below, the row number where the thread block family is located and the column number where the thread block family is located are determined.

[0074] Exemplarily, in the manner of dividing the thread block family number by the number of thread block families per row according to Equation (3), the row number where the thread block family is located is calculated, and the remainder after dividing the thread block family number by the number of thread block families per row is calculated according to Equation (4) to obtain the column number where the thread block family is located.

[0075] row_id = cluster_id / cluster_number_per_row (3).

[0076] col_id = cluster_id % cluster_number_per_row (4).

[0077] In Equation (3) and Equation (4), row_id represents the row number where the thread block family is located, col_id represents the column number where the thread block family is located, cluster_id represents the thread block family number, / represents the integer division operation; % represents the remainder operator, which is used to calculate the remainder after dividing two numbers.

[0078] It should be noted that when taking the number of thread block families per column cluster_number_per_col in the case of scheduling multiple thread blocks in the column direction as an example, the column number where the thread block family is located and the row number where the thread block family is located can also be calculated according to Equation (5) and Equation (6).

[0079] col_id = cluster_id / cluster_number_per_col (5).

[0080] row_id = cluster_id % cluster_number_per_col (6).

[0081] At this time, based on the row number where the thread block family is located and the column number where the thread block family is located, the scheduling area size corresponding to the thread block family can be determined, that is, the area size corresponding to the thread block family where the thread block is located can be determined. Specifically, the area obtained by multiplying row_id and col_id can be used as the scheduling area size.

[0082] In an exemplary embodiment, the artificial intelligence chip determines the target data block coordinates to be processed by the available computing units corresponding to the thread block based on the scheduling area size and the position offset of the thread block in the corresponding thread block family. The specific determination process can be determined through the following steps.

[0083] First, based on the data block row size required to be processed by the available computing units corresponding to the thread block, the row size of the thread block family, the row number of the scheduling area size, and the row offset of the thread block in the corresponding thread block family, determine the data block row coordinates required to be processed by the available computing units corresponding to the thread block; further, based on the data block column size required to be processed by the available computing units corresponding to the thread block, the column number of the thread block family, the column size of the scheduling area size, and the column offset of the thread block in the corresponding thread block family, determine the data block column coordinates required to be processed by the available computing units corresponding to the thread block; then, determine the data block row coordinates and the data block column coordinates as the target data block coordinates required to be processed by the available computing units corresponding to the thread block.

[0084] Specifically, calculate the data block row coordinates tile_row_coord required to be processed by the available computing units corresponding to the thread block according to Equation (7), and calculate the data block column coordinates tile_col_coord required to be processed by the available computing units corresponding to the thread block according to Equation (8).

[0085] tile_row_coord=tile_row*(cluster_row*row_id+row_offset) (7).

[0086] tile_col_coord=tile_col*(cluster_col*col_id+col_offset) (8).

[0087] In formulas (7) and (8), tile_row represents the row size of the data block to be processed by the available computing units corresponding to the thread block, tile_col represents the column size of the data block to be processed by the available computing units corresponding to the thread block, cluster_row represents the row size of the thread block family where the thread block is located, cluster_col represents the column size of the thread block family where the thread block is located; row_id represents the row number of the scheduling area size, that is, the row number where the thread block family is located; col_id represents the column number of the scheduling area size, that is, the column number where the thread block family is located; row_offset represents the row offset of the thread block in the corresponding thread block family, and col_offset represents the column offset of the thread block in the corresponding thread block family.

[0088] Based on the above Figure 1 In an exemplary embodiment of the task scheduling method shown above, step 110 can be specifically implemented by the following steps.

[0089] In response to the task scheduling instruction, based on the target task scheduling parameters and the target data storage instruction carried in the task scheduling instruction, as well as the mapping relationship between the task scheduling parameter - data storage instruction - data block scheduling configuration parameter data packet, obtain the target data block scheduling configuration parameter data packet, and different data block scheduling configuration parameters are included in the target data block scheduling configuration parameter data packet.

[0090] Among them, the mapping relationship is established based on different data block scheduling configuration parameter data packets and different data storage instructions that are pre - processed by the main control processor for the computing tasks of different types of artificial intelligence operators respectively.

[0091] The target task scheduling parameters include the target scheduling direction and the target data block scheduling parameters.

[0092] The target scheduling direction can be the scheduling direction of multiple target data blocks to be processed in each loop among the multiple data blocks obtained by splitting the target computing task; for example, row - by - row scheduling or column - by - column scheduling; and, the target scheduling direction can be specified manually or automatically determined by the task scheduling system built into the artificial intelligence chip based on the scale of the target computing task.

[0093] Exemplarily, referring to Figure 2A and Figure 2B the schematic diagram of the target scheduling direction shown, as Figure 2A and Figure 2B shown, when 16 target data blocks need to be processed in each loop, Figure 2A Among them, AlongN is to schedule 16 target data blocks row - by - row, and wave0~wave5 are the coordinates of 16 target data blocks scheduled row - by - row in 6 loops; Figure 2BAmong them, AlongM schedules 16 target data blocks column by column, and the coordinates of 16 target data blocks scheduled column by column in each of the 6 loops from wave0 to wave5.

[0094] The scheduling parameters of the target data blocks can be represented by the Swizzle size, which is used to characterize the mapping relationship between the linear index of the thread blocks in multiple thread blocks equal to the number of currently available computing units and the coordinates of the data blocks, and the specific value of the scheduling parameters of the target data blocks can be specified manually.

[0095] Exemplarily, referring to Figure 3 the schematic diagram showing the relationship between the scheduling parameters of the target data blocks and the linear index of the thread blocks and the coordinates of the data blocks respectively shown in Figure 3 wherein, cluster shape represents the thread block family shape, wave 0 tile~wave3 tile represents multiple data blocks that need to be synchronously processed in each of the 4 loops, problem block shape represents the data block shape, oob represents out-of-band data, as Figure 3 shown, when 16 target data blocks need to be processed in each loop, if the Swizzle size is 1, then 1 target data block is scheduled column by column first in each loop, and then scheduled row by row until 16 target data blocks are scheduled; if the Swizzle size is 2, then 2 target data blocks are scheduled column by column first in each loop, and then scheduled row by row until 16 target data blocks are scheduled; if the Swizzle size is 4, then 4 target data blocks are scheduled column by column first in each loop, and then scheduled row by row until 16 target data blocks are scheduled; if the Swizzle size is 8, then 8 target data blocks are scheduled column by column first in each loop, and then scheduled row by row until 16 target data blocks are scheduled.

[0096] The target data storage instruction can be a storage instruction generated when storing different data block scheduling configuration parameters with an immediate number, or can also be a storage instruction generated when storing different data block scheduling configuration parameters with a register; specifically, the target data storage instruction can be a computer program instruction generated by the manually specified immediate number storage method, denoted as a static kernel program; or can also be a computer program instruction automatically generated by the manually specified register storage method, denoted as a dynamic kernel program.

[0097] It should be noted that during the generation process of the program running on the artificial intelligence chip, a dynamic kernel program or a static kernel program can be generated according to the configuration. The dynamic kernel program can be compatible with different computing scales and has strong versatility or compatibility; the static kernel program can pre-compute the required scheduling configuration parameters of different data blocks in the central processing unit and reflect them in the corresponding computer program instructions in the form of immediate numbers, thereby avoiding certain calculations on some artificial intelligence chips and having better performance.

[0098] Specifically, when the artificial intelligence chip calls an artificial intelligence operator during the process of training or inferring an artificial intelligence model, the corresponding target task scheduling parameters and target data storage instructions can be determined for the artificial intelligence operator first, and then the task scheduling instructions can be automatically generated accordingly.

[0099] When the artificial intelligence chip responds to the task scheduling instruction, the mapping relationship between the task scheduling parameter - data storage instruction - data block scheduling configuration parameter data packet can be called. This mapping relationship can be stored in the memory of the main control processor and obtained by means of information interaction between the artificial intelligence chip and the main control processor. For example, the artificial intelligence chip sends the target task scheduling parameters and target data storage instructions to the main control processor and then receives the target data block scheduling configuration parameter data packet fed back by the main control processor; alternatively, the above mapping relationship can also be established by the main control processor and then communicated and transmitted or copied to the memory of the artificial intelligence chip for storage, so that the artificial intelligence chip can obtain the corresponding target data block scheduling configuration parameter data packet by looking up the above mapping relationship based on the target task scheduling parameters and target data storage instructions. The present invention does not make specific limitations on this.

[0100] Based on the above Figure 1 For the main control processor in an example embodiment of the task scheduling method shown above, it should be noted that:

[0101] The main control processor is used to perform the following steps for each computing task: First, align the computing task based on the data block size, thread block family size, and data block scheduling parameters in sequence to obtain the updated computing task; further, divide the updated computing task based on the data block size, data block scheduling direction, and the number of currently available computing units to obtain the data block scheduling configuration parameter data packet corresponding to the computing task; then, store the data block scheduling configuration parameter data packet in the form of an immediate number or a register to generate the data storage instruction corresponding to the computing task.

[0102] Among them, the data block scheduling configuration parameter data packet may include, but is not limited to, the number of data blocks contained in each row, the number of data blocks contained in each column, and the number of data blocks contained in the plane area after the calculation task corresponding to the artificial intelligence operator is segmented based on the currently available computing units, as well as the number of thread block families contained in each row and the number of thread blocks contained in each column in different scheduling directions of multiple thread blocks; the number of thread blocks is the same as the number of currently available computing units.

[0103] Specifically, for each calculation task of different types of artificial intelligence operators, the calculation task is first aligned using the data block size (such as the size of a tile), then secondarily aligned using the thread block family size (such as the size of a cluster), and finally thirdly aligned using the data block scheduling parameter (such as the Swizzle size) to obtain the updated calculation task.

[0104] At this time, for the updated calculation task, it is segmented according to the data block size, the data block scheduling direction, and the number of currently available computing units. For example, when the calculation task is a matrix multiplication calculation task and the matrix multiplication calculation task is A m×k ×B k×n =C m×n when, the corresponding segmented data block can be Figure 4 the dashed box in the C matrix of, and the scale of the matrix multiplication calculation task is the size of m, the size of k, and the size of n; thus, the data block scheduling configuration parameter data packet corresponding to the calculation task is obtained, and then all the data in the data block scheduling configuration parameter data packet is used to generate the data storage instruction corresponding to the calculation task in the form of an immediate number or a register. Traverse all calculation tasks in the above manner until the data block scheduling configuration parameter data packet and the data storage instruction corresponding to each calculation task are obtained respectively.

[0105] Based on the above Figure 1 shown task scheduling method, in an exemplary embodiment, in step 130, the artificial intelligence chip executes and processes the sub-calculation tasks of multiple target data blocks corresponding to their respective target data blocks, and its specific process is implemented through the following steps.

[0106] When multiple target data block coordinates all belong to the current batch of target data block coordinates and do not exceed the target scale of the target calculation task, process each sub-calculation task; further, based on the processing completion flag, re-determine the next batch of target data block coordinates among the multiple data blocks into which the target calculation task is segmented; then, use all the data block coordinates contained in the next batch of target data block coordinates as the new target data block coordinates, and repeat the above steps; until there is at least one invalid data block exceeding the target scale among the re-determined multiple new target data block coordinates.

[0107] Specifically, continue to refer to Figure 2AThe schematic diagram of the target scheduling direction shown. When 16 target data blocks need to be synchronously processed in each loop, each batch of target data block coordinates includes 16 target data block coordinates. Through Figure 2A It can be seen from the row scheduling of 16 target data blocks represented by AlongN in Figure 2A that wave0 to wave5 are 16 target data block coordinates scheduled row by row in each of the 6 loops, and none of them contain invalid data blocks exceeding the target scale. Figure 2A In Figure 2A , wave6 indicates that among the 16 target data block coordinates scheduled row by row in the 7th loop, only 3 target data blocks are included, and the remaining 13 are invalid data blocks exceeding the target scale, that is, 13 data blocks all exceed the target scale and are all illegal. At this time, after the 3 target data blocks are executed in the 7th loop, the loop can be exited to complete the calculation process of the target calculation task.

[0108] Exemplarily, referring to Figure 5 the second schematic diagram of the process of the task scheduling method shown. As Figure 5 shown, when the main control processor is a central processing unit, the central processing unit can calculate different data block scheduling configuration parameters after preprocessing the target calculation task according to the target scale of the target calculation task, the data block size, the number of currently available computing units in the artificial intelligence chip, the target task scheduling parameters, and the target data storage instruction, and then copy the different data block scheduling configuration parameters to the artificial intelligence chip in the form of a target data block scheduling configuration parameter data packet.

[0109] The artificial intelligence chip calculates the coordinates of each target data block that needs to be synchronously processed in each loop. Each target data block coordinate can be calculated based on the linear index of the corresponding thread block, and the next target data block coordinate is calculated based on the updated linear index until the coordinates of multiple target data blocks are calculated and multiple sub-computation tasks are synchronously executed; then, the coordinates of each target data block that needs to be synchronously processed in the next loop are calculated in the above manner. Until the calculation process of the target calculation task is completed.

[0110] Next, the task scheduling device provided by the present invention will be described. The task scheduling device described below can be mutually corresponded and referred to the task scheduling method described above.

[0111] Referring to Figure 6 the schematic diagram of the structure of the task scheduling device shown. As Figure 6 shown, the task scheduling device 600 is applied to an artificial intelligence chip running an artificial intelligence algorithm, and includes: a parameter acquisition unit 610 and a task scheduling unit 620.

[0112] The parameter acquisition unit 610 is configured to obtain different data block scheduling configuration parameters after preprocessing the target calculation task of the artificial intelligence operator in response to a task scheduling instruction; each data block scheduling configuration parameter is provided by the main control processor.

[0113] A task scheduling unit 620, configured to determine multiple target data block coordinates matching the number of currently available computing units based on each data block scheduling configuration parameter, and execute sub-computation tasks for processing the target data blocks corresponding to the multiple target data block coordinates respectively.

[0114] Optionally, the task scheduling unit 620 is specifically configured to execute the following steps for each thread block having the same number as the number of currently available computing units: determine the thread block family number of the thread block family based on the linear index of the thread block and the thread block family size of the thread block family corresponding to the thread block; determine the scheduling area size corresponding to the thread block family based on the thread block family number and each data block scheduling configuration parameter; determine the target data block coordinates required to be processed by the available computing units corresponding to the thread block based on the scheduling area size and the position offset of the thread block in the corresponding thread block family.

[0115] Optionally, the task scheduling unit 620 is specifically configured to determine the thread block family size based on the row size of the thread block family and the column size of the thread block family; determine the thread block family number based on the linear index and the thread block family size.

[0116] Optionally, when each data block scheduling configuration parameter includes the number of thread block families contained in each row when multiple thread blocks are scheduled in the row direction, the task scheduling unit 620 is specifically configured to determine the row number where the thread block family is located based on the integer division operation result of the thread block family number and the number of thread block families contained in each row; determine the column number where the thread block family is located based on the remainder operation result of the thread block family number and the number of thread block families contained in each row; determine the scheduling area size based on the row number and the column number; or, when each data block scheduling configuration parameter includes the number of thread block families contained in each column when multiple thread blocks are scheduled in the column direction, determine the column number where the thread block family is located based on the integer division operation result of the thread block family number and the number of thread block families contained in each column; determine the row number where the thread block family is located based on the remainder operation result of the thread block family number and the number of thread block families contained in each column; determine the scheduling area size based on the row number and the column number.

[0117] Optionally, the task scheduling unit 620 is specifically configured to determine the data block row coordinates required to be processed by the available computing units corresponding to the thread block based on the data block row size required to be processed by the available computing units corresponding to the thread block, the row size of the thread block family, the row number of the scheduling area size, and the row offset of the thread block in the corresponding thread block family; determine the data block column coordinates required to be processed by the available computing units corresponding to the thread block based on the data block column size required to be processed by the available computing units corresponding to the thread block, the column size of the thread block family, the column number of the scheduling area size, and the column offset of the thread block in the corresponding thread block family; determine the data block row coordinates and the data block column coordinates as the target data block coordinates required to be processed by the available computing units corresponding to the thread block.

[0118] Optionally, the parameter acquisition unit 610 is specifically configured to, in response to a task scheduling instruction, obtain a target data block scheduling configuration parameter data packet based on the target task scheduling parameters and target data storage instructions carried in the task scheduling instruction, and the mapping relationship between the task scheduling parameter - data storage instruction - data block scheduling configuration parameter data packet; wherein, the mapping relationship is established based on different data block scheduling configuration parameter data packets and different data storage instructions determined by the main control processor after preprocessing the computing tasks of different types of artificial intelligence operators respectively.

[0119] Optionally, the main control processor in the parameter acquisition unit 610 is configured to perform the following steps for each computing task: align the computing task in sequence based on the data block size, thread block family size, and data block scheduling parameters to obtain an updated computing task; split the updated computing task based on the data block size, data block scheduling direction, and the number of currently available computing units to obtain a data block scheduling configuration parameter data packet corresponding to the computing task; store the data block scheduling configuration parameter data packet in the form of an immediate number or a register to generate a data storage instruction corresponding to the computing task.

[0120] Optionally, the task scheduling unit 620 is specifically configured to process each sub-computing task when multiple target data block coordinates all belong to the current batch of target data block coordinates and do not exceed the target scale of the target computing task; re-determine the next batch of target data block coordinates among the multiple data blocks into which the target computing task is split based on the processing completion flag; use all the data block coordinates included in the next batch of target data block coordinates as new target data block coordinates, and repeat the above steps; until there is at least one invalid data block exceeding the target scale among the multiple newly determined target data block coordinates.

[0121] Figure 7 An example of the entity structure diagram of an electronic device is shown in Figure 7As shown, the electronic device may include: a main control processor (not shown in the figure) connected in the form of a heterogeneous framework, an artificial intelligence (AI) chip 710 that runs artificial intelligence algorithms, a communication interface 720, a memory 730, and a communication bus 740. Among them, the AI chip 710, the communication interface 720, and the memory 730 complete mutual communication through the communication bus 740. The AI chip 710 can call the logical instructions in the memory 730 to execute a task scheduling method, which includes: in response to a task scheduling instruction, obtaining different data block scheduling configuration parameters after preprocessing the target computing tasks of the artificial intelligence operator; each data block scheduling configuration parameter is provided by the main control processor; based on each data block scheduling configuration parameter, determining a plurality of target data block coordinates that match the number of currently available computing units, and executing sub-computing tasks for processing the target data blocks corresponding to each of the plurality of target data block coordinates; the main control processor is used to provide each data block scheduling configuration parameter to the AI chip 710.

[0122] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0123] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the task scheduling method provided by the above-mentioned various methods. The method includes: in response to a task scheduling instruction, obtaining different data block scheduling configuration parameters after preprocessing the target computing tasks of the artificial intelligence operator; each data block scheduling configuration parameter is provided by the main control processor; based on each data block scheduling configuration parameter, determining a plurality of target data block coordinates that match the number of currently available computing units, and executing sub-computing tasks for processing the target data blocks corresponding to each of the plurality of target data block coordinates.

[0124] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a task scheduling method provided by the above-mentioned various methods. The method includes: in response to a task scheduling instruction, obtaining different data block scheduling configuration parameters after preprocessing the target computing tasks of artificial intelligence operators; each data block scheduling configuration parameter is provided by a main control processor; based on each data block scheduling configuration parameter, determining a plurality of target data block coordinates matching the number of currently available computing units, and executing sub-computing tasks for processing the target data blocks corresponding to each of the plurality of target data block coordinates.

[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A task scheduling method, characterized in that: An artificial intelligence chip for running an artificial intelligence algorithm, the method comprising: In response to the task scheduling instruction, different data block scheduling configuration parameters after preprocessing of the target computing task of the artificial intelligence operator are obtained; each of the data block scheduling configuration parameters is provided by the main control processor, and each of the data block scheduling configuration parameters includes the number of data blocks after the target computing task of the artificial intelligence operator is divided based on the currently available computing units, and the number of thread blocks of multiple thread blocks under different scheduling directions; Based on each of the data block scheduling configuration parameters, multiple target data block coordinates that match the number of currently available computing units are determined, and sub-computing tasks for processing the target data blocks corresponding to each of the multiple target data block coordinates are executed; the multiple target data block coordinates are part of the multiple data block coordinates obtained after task segmentation of the target computing task.

2. The task scheduling method according to claim 1, characterized in that: The step of determining a plurality of target data block coordinates that match the number of currently available computing units based on each of the data block scheduling configuration parameters comprises: The following steps are performed for each thread block having the same number as the currently available computing units: Determine a thread block family number of the thread block family based on the linear index of the thread block and the thread block family size of the thread block family corresponding to the thread block; Determine the scheduling region size corresponding to the thread block family based on the thread block family number and each of the data block scheduling configuration parameters; the scheduling region size is the region size formed when scheduling according to the scheduling direction, and the scheduling region size is determined based on the number of thread block families included in the scheduling direction; Determine the target data block coordinates required to be processed by the available computing unit corresponding to the thread block based on the scheduling region size and the position offset of the thread block in the corresponding thread block cluster; The determining the thread block family number of the thread block family based on the linear index of the thread block and the thread block family size of the thread block family corresponding to the thread block includes: Determining a size of the thread block cluster based on a row size of the thread block cluster and a column size of the thread block cluster; Determining the thread block group number based on the linear index and the thread block group size; The determining, based on the scheduling region size and the position offset of the thread block in the corresponding thread block cluster, the target data block coordinates required to be processed by the thread block corresponding to the available computing unit includes: Determine the row coordinates of the data blocks that the thread block needs to process corresponding to the available computing unit based on the row size of the data blocks that the thread block needs to process corresponding to the available computing unit, the row size of the thread block family, the row number of the scheduling region size, and the row offset of the thread block in the corresponding thread block family; Determine the data block column coordinates that the thread block needs to process corresponding to the available computing unit based on the data block column size that the thread block needs to process corresponding to the available computing unit, the column size of the thread block family, the column number of the scheduling region size, and the column offset of the thread block in the corresponding thread block family; The data block row coordinates and the data block column coordinates are determined as the target data block coordinates that the thread block needs to process corresponding to the available computing unit.

3. The task scheduling method according to claim 2, characterized in that: The determining, based on the thread block family number and each of the data block scheduling configuration parameters, a scheduling region size corresponding to the thread block family includes: When each of the data block scheduling configuration parameters includes the number of thread block clusters contained in each row when the plurality of thread blocks are scheduled in a row direction, Determine the row number where the thread block family is located based on the integer division result of the thread block family number and the number of thread block families contained in each row; Determine the column number where the thread block family is located based on a modulo operation result of the thread block family number and the number of thread block families contained in each row; Determining the scheduling region size based on the row number and the column number; Alternatively, when each of the data block scheduling configuration parameters includes the number of thread block clusters contained in each column when the plurality of thread blocks are scheduled in a column direction, Determine the column number where the thread block family is located based on the integer division result of the thread block family number and the number of thread block families contained in each column; Determine the row number where the thread block family is located based on a modulo operation result of the thread block family number and the number of thread block families contained in each column; The scheduling region size is determined based on the row number and the column number.

4. The task scheduling method according to any one of claims 1 to 3, characterized in that: The step of obtaining, in response to the task scheduling instruction, different data block scheduling configuration parameters after preprocessing of the target computing task of the artificial intelligence operator includes: In response to the task scheduling instruction, based on the target task scheduling parameters and the target data storage instruction carried in the task scheduling instruction, and the mapping relationship between the task scheduling parameters, the data storage instruction, and the data block scheduling configuration parameter data packet, a target data block scheduling configuration parameter data packet is obtained, wherein the target data block scheduling configuration parameter data packet contains the different data block scheduling configuration parameters; The mapping relationship is established based on different data block scheduling configuration parameter data packets and different data storage instructions determined by the main control processor after preprocessing the computing tasks of different types of artificial intelligence operators.

5. The task scheduling method according to claim 4, characterized in that: The main control processor is used to perform the following steps for each computing task: Aligning the computing tasks in sequence based on the data block size, the thread block family size, and the data block scheduling parameters to obtain updated computing tasks; Dividing the updated computing task based on the data block size, the data block scheduling direction and the number of currently available computing units to obtain a data block scheduling configuration parameter data packet corresponding to the computing task; The data block scheduling configuration parameter data packet is stored in an immediate value manner or a register manner, and a data storage instruction corresponding to the computing task is generated.

6. The task scheduling method according to any one of claims 1 to 3, characterized in that: The executing and processing the sub-computing tasks of the target data blocks corresponding to the coordinates of the plurality of target data blocks respectively comprises: S1, processing each of the sub-computing tasks when the coordinates of the multiple target data blocks all belong to the coordinates of the target data blocks of the current batch and do not exceed the target scale of the target computing task; S2, based on the processing completion mark, re-determining the coordinates of the next batch of target data blocks in the multiple data blocks into which the target computing task is divided; S3, taking all data block coordinates contained in the next batch of target data block coordinates as new target data block coordinates; Steps S1 to S3 are repeatedly executed until at least one invalid data block that exceeds the target size exists in the newly determined multiple target data block coordinates.

7. A task scheduling device, characterized in that: An artificial intelligence chip for running an artificial intelligence algorithm, the device comprising: A parameter acquisition unit, for obtaining, in response to a task scheduling instruction, different data block scheduling configuration parameters after preprocessing of a target computing task of an artificial intelligence operator; each of the data block scheduling configuration parameters is provided by a main control processor, and each of the data block scheduling configuration parameters includes the number of data blocks after the target computing task of the artificial intelligence operator is divided based on currently available computing units, and the number of thread blocks of multiple thread blocks under different scheduling directions; A task scheduling unit is used to determine multiple target data block coordinates that match the number of currently available computing units based on the data block scheduling configuration parameters, and execute sub-computing tasks for processing the target data blocks corresponding to each of the multiple target data block coordinates; the multiple target data block coordinates are part of the multiple data block coordinates obtained after task segmentation of the target computing task.

8. An electronic device, characterized in that: It includes a main control processor and an AI chip that runs an AI algorithm connected in a heterogeneous framework; The artificial intelligence chip is used to respond to the task scheduling instruction and obtain the scheduling configuration parameters of different data blocks after preprocessing of the target computing task of the artificial intelligence operator; Each of the data block scheduling configuration parameters includes the number of data blocks after the target computing task of the artificial intelligence operator is divided based on the currently available computing units, and the number of thread blocks of multiple thread blocks under different scheduling directions; based on each of the data block scheduling configuration parameters, multiple target data block coordinates matching the number of currently available computing units are determined, and sub-computing tasks of processing the target data blocks corresponding to the multiple target data block coordinates are executed; the multiple target data block coordinates are partial data block coordinates of the multiple data block coordinates obtained after the target computing task is divided; The main control processor is used to provide the artificial intelligence chip with scheduling configuration parameters for each of the data blocks.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the task scheduling method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Operator execution method and device, equipment and storage medium

    CN117808048A

  • Task scheduling method and device based on multiple operators, computer equipment and storage medium

    CN118550674A