Instruction Scheduling Method and Apparatus, Electronic Device, and Medium
By integrating the first instruction scheduling unit in the graphics processing unit for instruction scheduling in unit of working groups, the problem of poor adaptability of the hardware layer and the triton programming language in the prior art is solved, user programming is simplified and resource utilization is improved.
Patent Information
- Application Number
- CN202510088774.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-01-21
AI Technical Summary
In the prior art, the instruction scheduling mode of the graphics processing unit has poor adaptability to block-based GPU programming languages such as triton, resulting in the user needing additional programming to manage thread bundle operations and setting synchronization mechanisms, increasing programming complexity.
By integrating the first instruction scheduling unit in the computing unit, target instructions are directly identified and scheduled in work groups, the creation and synchronization mechanism of thread bundles are avoided, and instructions are directly scheduled to the target work group and executed.
It improves the adaptability of the hardware layer to block-based GPU programming languages such as triton, simplifies user programming, reduces programming complexity, avoids the management of thread bundle operations, and improves the resource utilization of computing units.
Smart Images

Figure CN119536819B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and particularly to an instruction scheduling method and apparatus, an electronic device, and a medium. Background Art
[0002] Traditional graphics processing units often have multiple computing units, and each computing unit is respectively provided with multiple execution units. The data processing process of the graphics processing unit is completed by the execution units. With the development of artificial intelligence technology, in order to accelerate the efficiency of neural network model training and inference, some dedicated hardware modules for performing specific data processing tasks, such as TensorCore, Copy Engine, etc., are integrated in the computing units of the graphics processing unit in the prior art. When certain specific tasks need to be executed, these hardware modules can be directly used to complete them, and these dedicated hardwares can support executing corresponding data processing tasks with a relatively high degree of parallelism. At the same time, in order to simplify the GPU programming difficulty, there are some block-based GPU programming languages in the prior art, such as the triton programming language. Based on these programming languages, users can flexibly decompose the tensor data to be processed into multiple tensor blocks of a specified size, and then allocate these tensors to each computing unit for execution.
[0003] In the prior art, the computing unit schedules instructions to corresponding threads through the instruction scheduler in the execution unit, and then schedules the threads carrying the instructions to the corresponding hardware modules for execution to complete the corresponding data processing tasks. The execution unit often schedules instructions in units of warps, that is, in each instruction scheduling, a single instruction is scheduled to each thread in a single warp, and then the warp is scheduled to the hardware unit for execution. That is, in the prior art, although the upper-layer application decomposes the data processing task and the corresponding tensor data into the granularity specified by the user and then distributes them to the computing unit for execution; at the hardware layer, the computing unit still executes the tasks issued by the upper-layer application in units of warps, and the number of threads included in a single warp is often fixed, which makes the granularity of instruction scheduling at the hardware layer often inconsistent with the task decomposition granularity specified by the user. At this time, the hardware layer often needs to schedule the same instruction to multiple warps to cooperate to complete the execution of a single subtask assigned by the upper-layer application. Since the instruction scheduling and execution processes of different warps are independent of each other, this leads to the need for additional programming by the user when writing the kernel to manage the operations of the warps. For example, the user needs to set a synchronization mechanism between the warps to prevent unpredictable situations such as race conditions between the warps; when different hardware modules need to execute different instructions on different tensor blocks simultaneously, the instruction scheduler needs to schedule multiple instructions to multiple warps respectively. At this time, the user needs to perform Warp specialization to indicate which instruction each warp is used to execute so that the execution unit can accurately schedule the instructions to the corresponding warps. This instruction scheduling mode has a poor compatibility with block-based GPU programming languages such as triton. Summary of the Invention
[0004] Embodiments of the present disclosure provide an instruction scheduling method, an electronic device, and a storage medium, which can improve the compatibility between the hardware layer and block-based GPU programming languages such as triton.
[0005] According to one aspect of the present disclosure, an instruction scheduling method is proposed, which is applied to a computing unit. The computing unit includes a first instruction scheduling unit, at least one execution unit, and at least one first functional unit. The first functional unit includes at least one of a copy engine, a tensor core, and a scalar processing unit. The method includes:
[0006] Obtain a target instruction through the first instruction scheduling unit;
[0007] Identify a target unit for executing the target instruction, where the target unit is one of the execution unit or the first functional unit;
[0008] In the case where the target unit is the first functional unit, determine a target tensor block corresponding to the target instruction and a target workgroup corresponding to the target tensor block, where the target workgroup is used to indicate the memory address range of the target tensor block;
[0009] Dispatch the target instruction to the target workgroup through the first instruction scheduler, and dispatch the target workgroup to the target unit, so that the target unit reads the target tensor block based on the memory address range and executes the target instruction on the target tensor block.
[0010] Optionally, the first instruction scheduler is provided with a corresponding instruction buffer, and the computing unit further includes a command processor. Before obtaining the target instruction through the first instruction scheduler, the method further includes:
[0011] Obtain a target command, where the target command is used to instruct the computing unit to start a target kernel corresponding to the target command, and the target kernel is composed of multiple instructions;
[0012] Parse the target command through the command processor to determine the target kernel to be started, determine multiple instructions to be executed based on the target kernel, and write the instructions to be executed into the instruction buffer of the first instruction scheduler;
[0013] The obtaining the target instruction through the first instruction scheduler includes:
[0014] Obtain the target instruction from the corresponding instruction buffer through the first instruction scheduler.
[0015] Optionally, the target kernel at least includes a main function kernel part, a scalar kernel part, and a vector kernel part. The scalar kernel part includes multiple preset scalar operation kernels, the vector kernel part includes multiple preset vector operation kernels, and the main function kernel part calls the scalar operation kernels and the vector operation kernels through call instructions.
[0016] Optionally, the identifying the target unit for executing the target instruction includes:
[0017] Identify the instruction type corresponding to the target instruction, where the instruction type is used to characterize the type of data processing task corresponding to the target instruction, and the execution unit and each of the first functional units respectively correspond to different instruction types;
[0018] Determine the target unit for executing the target instruction based on the instruction type.
[0019] Optionally, before the identifying the target unit for executing the target instruction, the method further includes:
[0020] Create a first instruction set, where the first instruction set includes a plurality of preset instructions, and the target instruction is one of the preset instructions;
[0021] Divide the preset instructions into a plurality of instruction subsets, each instruction subset includes at least one of the preset instructions, and the preset instructions in a single instruction subset correspond to the same instruction type;
[0022] The identifying the instruction type corresponding to the target instruction includes:
[0023] Identify the instruction subset corresponding to the target instruction to determine the instruction type corresponding to the target instruction.
[0024] Optionally, the instruction type includes at least one of the following:
[0025] Tensor operation instruction, which indicates an instruction executed by the tensor core;
[0026] Scalar operation instruction, which indicates an instruction executed by the scalar processing unit;
[0027] Vector operation instruction, which indicates an instruction executed by the execution unit;
[0028] Data copy instruction, which is an instruction executed by the copy engine.
[0029] Optionally, before determining the target tensor block corresponding to the target instruction and the target workgroup corresponding to the target tensor block, the method further includes:
[0030] Obtain workgroup creation parameters, where the workgroup creation parameters include workgroup size and workgroup number parameters, and the workgroup size characterizes the size of the tensor block corresponding to a single workgroup in each dimension;
[0031] Create a plurality of workgroups based on the workgroup number parameter and assign workgroup identifiers to each of the workgroups, and divide the tensor to be processed into a plurality of tensor blocks based on the workgroup size, where the workgroups and the tensor blocks correspond one by one.
[0032] Optionally, each of the execution units includes a second instruction scheduling unit, and the method further includes:
[0033] If the target unit is the execution unit, determine the target workgroup assigned to the target unit, and create a plurality of threads corresponding to the target workgroup, where each of the threads corresponds to the data in the tensor block corresponding to the target workgroup;
[0034] The first instruction scheduling unit sends the target instruction to the second instruction scheduling unit in the execution unit;
[0035] In the execution unit, the created threads are determined as a plurality of warps, where each warp includes a second number of threads;
[0036] In the execution unit, the second instruction scheduling unit schedules the target instruction to the threads in the plurality of warps respectively.
[0037] According to an aspect of the present disclosure, there is provided an instruction scheduling device, the device includes:
[0038] An instruction acquisition unit, configured to acquire a target instruction through a first instruction scheduling unit;
[0039] A first recognition unit, configured to recognize a target unit for executing the target instruction, where the target unit is one of an execution unit or a first functional unit;
[0040] A target workgroup determination unit, configured to, when the target unit is the first functional unit, determine a target tensor block corresponding to the target instruction and a target workgroup corresponding to the target tensor block, where the target workgroup is used to indicate a memory address range of the target tensor block;
[0041] A first scheduling unit, configured to schedule the target instruction to the target workgroup through the first instruction scheduling unit, and schedule the target workgroup to the target unit, so that the target unit reads the target tensor block based on the memory address range and executes the target instruction on the target tensor block.
[0042] According to an aspect of the present disclosure, there is provided an electronic device, the electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory, and when the program is run by the processor, it implements the instruction scheduling method as described in any one of the above.
[0043] According to an aspect of the present disclosure, there is provided a computer-readable storage medium, the computer-readable storage medium stores one or more programs, and the one or more programs can be run by one or more processors to implement the instruction scheduling method as described in any one of the above.
[0044] According to one aspect of the present disclosure, there is provided an instruction scheduling method, an electronic device, and a storage medium. A first instruction scheduling unit and at least one first functional unit for performing specific data processing tasks are directly integrated in a computing unit. Then, when it is necessary to use the computing unit to execute a data processing task, first, the target instruction to be executed is obtained through the first instruction scheduling unit, and the target unit for executing the target instruction is identified. Then, if the target unit is a first functional unit, the target tensor block to be used for executing the target instruction and the target workgroup corresponding to the target tensor block are determined. Then, the target instruction is scheduled to the target workgroup, and the target workgroup carrying the target instruction is scheduled to the target unit, so that the target unit can determine the address of the target tensor block in the memory based on the target workgroup, thereby reading the target tensor block and performing the data processing task corresponding to the target instruction on the target tensor block. Based on this, in this embodiment, when the target instruction is an instruction executed by a first functional unit such as a tensor core or a copy engine, the instruction is directly scheduled in units of workgroups through the first instruction scheduling unit, so that the target unit reads the target tensor block based on the memory address range indicated by the target workgroup and performs the target instruction on the target tensor block. In this way, the granularity of the instruction scheduling of the computing unit as the underlying hardware is consistent with the task decomposition granularity set by the user in the kernel function, and the first instruction scheduling unit directly schedules the target instruction to the target workgroup instead of a thread or a warp. At this time, there is no need to create multithreads and warps, so there is no need to set up a synchronization mechanism to manage the operations of the warps, nor is it necessary to perform Warp specialization to indicate what instructions each warp is used to execute. In this way, the instruction scheduling mode of the computing unit can be more adapted to the kernels written in programming languages such as Triton.
[0045] Other features and advantages of the present disclosure will be described in the following specification, and in part, will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be realized and obtained by the structures specifically pointed out in the specification, the claims, and the drawings. Brief Description of the Drawings
[0046] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. They are used together with the embodiments of the present disclosure to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.
[0047] Figure 1 is an architecture diagram of a computing unit applying the embodiment of the present disclosure;
[0048] Figure 2 is a main flowchart of an instruction scheduling method according to an embodiment of the present disclosure;
[0049] Figure 3 is Figure 2 The flowchart of step S201 and its associated steps in
[0050] Figure 4 The schematic structural diagram of the target kernel of an embodiment of the present disclosure;
[0051] Figure 5 is Figure 2 A sub - flowchart of step S202 in
[0052] Figure 6 is Figure 5 The flowchart of step S501 and its associated steps in
[0053] Figure 7 According to an embodiment of the present disclosure, it is the flowchart of creating a workgroup and dividing a tensor into tensor blocks corresponding to the workgroup;
[0054] Figure 8 is the flowchart of an instruction scheduling method of another embodiment of the present disclosure;
[0055] Figure 9 is the architecture diagram of an instruction scheduling device of an embodiment of the present disclosure.
[0056] Figure 10 is the architecture diagram of an electronic device of an embodiment of the present disclosure. Detailed implementation manners
[0057] In order to make the purpose, technical solutions and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0058] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:
[0059] Thread: It is the smallest execution unit when the graphics processing unit executes data processing tasks. Each thread can independently perform the same - mode processing on different data.
[0060] Work Group: It is a sub-task entity corresponding to a tensor block of a user-specified size in block-based programming languages such as Triton. Specifically, in programming languages such as Triton, when a user writes a kernel function, they can divide the tensor to be processed into multiple tensor blocks by setting meta-parameters such as block_size. Block_size is used to indicate the number of elements contained in a single tensor block, and each tensor block corresponds to a work group. For example, if the tensor to be processed is a three-dimensional tensor with a shape of (32, 256, 128), and block_size is (2, 32, 128), then the tensor to be processed is divided into 128 tensor blocks with a shape of (2, 32, 128). Each tensor block contains 2 * 32 * 128 elements, and at the same time, 128 work groups are created, with each work group corresponding to a tensor block. Specifically, each work group itself can carry information indicating the address of the corresponding tensor block in memory, such as the address index or address offset of the tensor block in memory.
[0061] Triton: A programming language used to write kernel functions for graphics processing units. Different from CUDA, in the Triton language, users can more flexibly and actively decompose the tensor to be processed into tensor blocks of a specified size, so as to decompose the required data processing tasks into sub-tasks of the user-specified size, and allocate each tensor block to each computing unit to execute each sub-task on each computing unit.
[0062] Warp: It is the basic unit for the execution unit to process data. Generally, each warp includes 32 fixed threads. Of course, in some graphics processing units, the number of threads in a single warp can also be 16, 64, or 128. However, for any given graphics processing unit, the number of threads contained in a single warp is fixed. The same thread often executes in the single instruction multiple threads (SIMT) mode, that is, the threads in the same warp are often used to execute the same instruction on different data in parallel.
[0063] Synchronization mechanism: A mechanism for managing multi-threaded operations. In a multi-threaded environment, multiple threads are often executed independently. Some threads may be executed by the hardware unit before other threads. In the instruction scheduling mode in the related art, the resources occupied by the threads that complete execution first will be released in advance and used to execute the next instruction, which may lead to unpredictable situations such as race conditions. Therefore, in a multi-threaded environment, it is often necessary to manage multi-threaded operations through a synchronization mechanism. Specifically, the synchronization mechanism sets a synchronization node for multiple threads. When some of the threads first execute to the synchronization node, these threads will be suspended and wait for other threads to also execute to this synchronization node. After all the threads with the synchronization mechanism set have executed to the synchronization node, they will continue to execute the subsequent processing. In the graphics processing unit of the related art, since the instruction scheduler in the execution unit executes instruction scheduling in units of warps, when multiple warps are used to cooperate to complete a data processing task, the user needs to set the synchronization mechanism at the level of different threads in the same warp and also at the level of different warps during programming. The synchronization mechanism becomes extremely complex. If the synchronization node of the synchronization mechanism is set improperly, some threads need to wait for other threads in the same warp or other warps to execute to the synchronization node for a long time, resulting in a significant reduction in computing efficiency, which greatly increases the setting of the synchronization mechanism.
[0064] Race Condition: It refers to multiple threads attempting to modify the data at the same location simultaneously. When multiple warps are used to complete the same data processing task, since each warp cannot access the memory occupied by other warps, these warps need to cooperate through shared memory. However, when performing instruction scheduling in units of warps, since the process of scheduling instructions to different warps is independent of each other, and the execution process of each warp is also independent of each other, this may cause some warps to be executed and completed before other warps. The instruction scheduler will directly schedule the next instruction to these warps that have completed the previous instruction. When these warps execute the next instruction, they may need to access the shared memory to modify the data generated by other warps when they executed the previous instruction. At this time, the other warps have not actually completed the previous instruction yet, which may cause the threads of the two warps to write data to the same location simultaneously, thus causing a race condition.
[0065] System architecture description of the application of the embodiments of the present disclosure
[0066] Figure 1 It is the system architecture diagram of the computing unit to which the instruction scheduling method of the embodiments of the present disclosure is applied, and it includes: a first instruction scheduling unit 110, at least one execution unit 120, and at least one first functional unit.
[0067] The first functional unit may include at least one of a tensor core 131, a replication engine 132, and a scalar processing unit 133, and each first functional unit is respectively used to execute a certain specific type of instruction. It can be understood that, as a general-purpose hardware unit, the execution unit 120 itself can execute various different instructions. The difference between the execution unit 120 and the replication engine 132, the tensor core 131, and the scalar processing unit 133 lies in the efficiency when executing specific types of tasks.
[0068] In this embodiment, the first functional unit integrated in the computing unit may include only one of the tensor core 131, the replication engine 132, or the scalar processing unit 133. For example, the computing unit only integrates the tensor core 131 as the first functional unit, without integrating the scalar processing unit 133 and the replication engine 132, and the instructions for data replication and scalar operations are executed by the execution unit 120. Similarly, the computing unit may also only integrate the replication engine 132 as the first functional unit, without integrating the tensor core 131 and the scalar processing unit 133. Of course, the computing unit may also integrate the tensor core 131, the replication engine 132, and the scalar processing unit 133 at the same time, respectively as the first functional units for executing different instructions. In addition, the computing unit may also integrate a dedicated hardware module for executing certain specific data processing tasks in addition to the tensor core 131, the replication engine 132, and the scalar processing unit 133, and this dedicated hardware module also serves as the first functional unit. Specifically, the first functional unit in the computing unit needs to be determined according to the actual hardware architecture of the computing unit.
[0069] The first instruction scheduling unit 110 is used to obtain a target instruction to be executed and identify a target unit for executing the target instruction, where the target unit may be one of the execution unit 120 or the first functional unit.
[0070] The first instruction scheduling unit 110 is further used to, when the target unit is the first functional unit, determine a target tensor block corresponding to the target instruction and a target workgroup corresponding to the target tensor block, and schedule the target instruction into the target workgroup, where the target tensor block is a tensor block formed by the data for executing the target instruction. For example, if the target instruction is a matrix multiplication instruction, the target tensor block is the matrix block in the two matrices required for matrix multiplication. In this way, when the target unit is other dedicated hardware except the execution unit, the first instruction scheduling unit integrated outside the execution unit can perform instruction scheduling in units of workgroups, rather than performing instruction scheduling in units of warps through the instruction scheduler in the execution unit.
[0071] Of course, with reference to Figure 1, a local memory may also be integrated in the computing unit. The computing unit may be coupled to an external on-chip memory and off-chip memory. When a copy engine 132 is integrated in the computing unit, the copy engine 132 will be integrated at a position closer to the local memory, on-chip memory, and off-chip memory than the execution unit, so as to shorten the physical path length when copying the data stored in one memory space to another memory space through the copy engine, and improve the efficiency when executing the corresponding data copy instruction using the copy engine.
[0072] Overall implementation manner of the instruction scheduling method of the embodiments of the present disclosure
[0073] Embodiments of the present disclosure propose an instruction scheduling method, which is applied to a computing unit as shown in Figure 1 Referring to Figure 2 , the instruction scheduling method includes:
[0074] Step S201, obtaining a target instruction through a first instruction scheduling unit;
[0075] Step S202, identifying a target unit for executing the target instruction, where the target unit is one of an execution unit or a first functional unit;
[0076] Step S203, when the target unit is a first functional unit, determining a target tensor block corresponding to the target instruction and a target workgroup corresponding to the target tensor block, where the target workgroup is used to indicate the memory address range of the target tensor block;
[0077] Step S204, scheduling the target instruction to the target workgroup through the first instruction scheduling unit, and scheduling the target workgroup to the target unit, so that the target unit reads the target tensor block based on the memory address range and executes the target instruction on the target tensor block.
[0078] In this embodiment, a Tensor Core is a dedicated hardware module in a computing unit that is specifically used to execute tensor operations. In some architectures, it is also referred to as a tensor engine, which is used to process tensors with two or more dimensions. That is, for the instructions executed by the Tensor Core, at least one of the data participating in the operation is in the form of a tensor with two or more dimensions. It should be noted that in this embodiment, two or more dimensions refer to the dimensions of the data when it is used for operation, rather than the dimensions in which the data is organized when stored. For example, when normalizing each row of a matrix separately, although the matrix itself is in a two-dimensional form, the actual operation is to normalize each row of the matrix separately, that is, the normalization process is actually performed in the form of a one-dimensional vector. Accordingly, the instructions for normalizing each row of the matrix are not the instructions executed by the Tensor Core; another example is that when adding two matrices, in fact, the elements at the same positions in the two matrices are added separately, and each addition operation actually only involves two elements at the same position in the two matrices. This process is actually performed in the form of a zero-dimensional scalar. Accordingly, the instructions for summing the two matrices are not the instructions executed by the Tensor Core.
[0079] Specifically, in this embodiment, the Tensor Core can at least be used to execute matrix multiplication instructions and matrix multiply accumulate (MMA) instructions. The Tensor Core can include a General Matrix Multiplication (GEMM) unit and an accumulation buffer. The General Matrix Multiplication unit is used to perform matrix multiplication on the input matrices, and the accumulation buffer can be used to accumulate the matrix multiplication results of the General Matrix Multiplication unit. Exemplarily, when executing the matrix multiply accumulate instruction, the two matrices for which matrix multiplication operations need to be performed are divided into multiple matrix blocks according to a preset rule, and then these matrix blocks are input into the Tensor Core in a certain order. The General Matrix Multiplication unit is used to perform matrix multiplication operations on the input matrix blocks, and the accumulation buffer accumulates the matrix multiplication results of the matrix blocks output by the General Matrix Multiplication unit at a certain period to execute the matrix multiply accumulate operation.
[0080] The Copy Engine is a hardware module dedicated to data copying and transfer in the computing unit. The graphics processing unit often has three levels of memory: global memory, local memory, and registers, and its memory structure is relatively complex. The global memory is set outside the computing unit and can be accessed by all computing units, but the access speed of the global memory is slower than that of the local memory and registers; the registers are set inside the execution units in the computing unit and can only be accessed by the threads in the corresponding execution units, and the access speed of the registers is faster than that of the global memory and local memory. Due to this complex memory structure, when the graphics processing unit performs the training and inference of neural network models, in order to improve efficiency, a large number of data copying operations will be performed to pre-copy the data required by the hardware modules in each computing unit when performing the corresponding computing tasks to the memory space that can be quickly accessed. For example, copying the data required by the execution unit when performing the corresponding computing task from the global memory to the registers in the corresponding execution unit. The Copy Engine is a hardware unit dedicated to executing such data copying instructions.
[0081] The Scalar Unit is a dedicated hardware module in the graphics processing unit for scalar operations. It can be used to execute basic arithmetic operation instructions and logical operation instructions on data. For example, relatively simple basic arithmetic operation instructions such as addition, subtraction, multiplication, and division of elements in two matrices, exponentiation and logarithmization of each element in a single matrix respectively, as well as basic logical operation instructions such as logical AND, logical OR, bitwise AND, bitwise OR, and exclusive OR. The Scalar Unit can be an Arithmetic and Logical Unit (ALU).
[0082] The execution unit is a general-purpose data processing module in the computing unit. The hardware structure of the execution unit itself is not optimized for certain specific data processing tasks, and the efficiency of the execution unit when performing these specific tasks is lower than that of dedicated hardware such as tensor cores and copy engines. However, as a general-purpose data processing module, the execution unit can execute most of the instructions in the neural network model, and for some instructions that are not suitable for execution by dedicated hardware such as tensor cores and copy engines, they will be assigned to the execution unit for execution.
[0083] In step S201, the target instruction is an instruction that the computing unit needs to execute subsequently, which indicates that the computing unit needs to perform a data processing task, such as performing arithmetic or logical operations on the corresponding data, accessing specific memory, etc.
[0084] It can be understood that when the graphics processing unit performs the training or inference task of the neural network model, each hardware module in the computing unit is driven by instructions. The computing unit will control each execution unit or the first functional unit therein to execute the data processing tasks corresponding to these instructions according to these instructions, so as to complete the inference or calculation task of the neural network model. For example, when the graphics processing unit needs to perform convolution processing on the intermediate data generated by a certain neural network layer, at this time, the graphics processing unit needs to distribute corresponding tasks to each computing unit therein, so that each computing unit is respectively used to perform convolution processing on a part of the data. Each computing unit will obtain an instruction for instructing the execution unit or a certain first functional unit in the computing unit to perform convolution processing on a part of the data allocated to the computing unit.
[0085] Specifically, the first instruction scheduling unit often maintains an instruction queue, and the instruction queue includes multiple to-be-executed instructions that need to be executed by the computing unit. The target instruction is actually an instruction in the instruction queue. When the first instruction scheduling unit detects that the execution condition of a certain instruction in the instruction queue is met, it will obtain the instruction from the instruction scheduling as the target instruction for subsequent instruction scheduling operations. Specifically, the first instruction scheduling unit can monitor whether the hardware units used to execute each instruction in the instruction queue are idle, and whether the data required to execute the instruction has been prepared. After the corresponding hardware unit is idle and the corresponding data has been prepared, the first instruction scheduling unit obtains the corresponding instruction from the instruction queue as the target instruction. For example, the first instruction scheduling unit can use the scoreboard mechanism to determine whether the data required for each instruction in the instruction queue is ready and whether the hardware unit for executing the instruction is idle.
[0086] In step S202, the target instruction indicates the data processing task that the computing unit needs to perform on the data subsequently. For example, it indicates that data needs to be processed by convolution, matrix multiplication, normalization, etc. In this embodiment, each first functional unit is used to execute a specific type of data processing task. For example, the tensor core is used to perform operations on tensors of two or more dimensions; the copy engine is used to perform data copy operations. In addition to the specific type of data processing tasks executed by these first functional units, the data processing tasks are executed by the execution units as general processing units. And the target instruction, as the one used to drive the hardware modules in the graphics processing unit, is actually one of multiple preset instructions, and these preset instructions are respectively used to instruct the hardware in the graphics processing unit to execute different data processing tasks. Based on this, in this embodiment, the mapping relationship between each preset instruction and the execution unit or the first functional unit can be preset. When the target instruction is obtained, the target unit for executing the target instruction can be identified based on this mapping relationship. The details of step S202 will be described later and will not be elaborated here.
[0087] In step S203, the target tensor block is the tensor block that needs to be used to perform the data processing task corresponding to the target instruction. The target workgroup is the workgroup corresponding to the target tensor block, and the target workgroup itself carries the memory address range used to characterize the target tensor block. Specifically, in a graphics processing unit, the elements in a tensor are stored at consecutive addresses, each address corresponding to an element in the tensor, and a tensor block often contains several consecutive elements in the tensor, that is, each tensor block is actually composed of multiple elements stored at a consecutive address range in memory. Based on this, in this embodiment, the target workgroup can indicate the memory address of the target tensor block by indicating the range formed by a consecutive address range corresponding to multiple elements constituting the tensor block.
[0088] Specifically, when writing a kernel function using a block-based programming language such as Triton, the user can set the relevant parameters of the workgroup in the kernel function through specific code statements to indicate how the graphics processing unit divides the tensor to be processed into multiple tensor blocks, so as to decompose the calculation tasks that the graphics processing unit needs to perform into multiple subtasks and distribute them to the computing units for parallel execution. For example, in Triton, the size of a single tensor block, that is, the number of elements in a single tensor block, can be set through parameters such as block_size and grid_size. At this time, when the graphics processing unit starts the kernel function written in a programming language such as Triton, multiple workgroups will be created based on the grid_size parameter set by the user. These workgroups will be organized into an N-dimensional grid, and each workgroup will be assigned corresponding N-dimensional coordinates. The block_size parameter indicates the stride of the workgroup in each dimension. Based on the coordinates of the workgroup, the block_size parameter, and the base address corresponding to the first element, the addresses of the elements in the tensor block corresponding to the workgroup can be determined, thereby dividing the tensor to be processed into tensor blocks corresponding to each workgroup. Exemplarily, the base address is A, the coordinates corresponding to a certain workgroup are (4, 0), and the strides in the two dimensions are determined to be (128, 1) based on the block_size parameter. Then, the elements stored at addresses A + 512 to A + 639 constitute the tensor block corresponding to this workgroup.
[0089] It should be noted that the workgroup itself is only used to indicate the address range of a tensor block. The workgroup itself can contain no threads or only one thread, and the workgroup works in the SPMD (Single-Program, Multiple-Data) mode, that is, each workgroup can be used to execute the same program (i.e., instruction) on multiple elements included in the corresponding tensor block.
[0090] It can be understood that the number of workgroups is determined by the kernel functions written in programming languages such as Triton, while the number of computing units in a single image processing unit is fixed. Based on this, when allocating the workgroups created based on the kernel functions to the computing units, each computing unit can correspond to multiple workgroups, that is, the tensor blocks allocated to each computing unit can be multiple.
[0091] In one embodiment, after dividing the tensor to be processed into multiple tensor blocks and creating multiple workgroups, these workgroups can be directly allocated to each computing unit so that each computing unit is used to process a fixed number of tensor blocks.
[0092] In addition, referring to step S201, it can be known that the first instruction scheduling unit itself maintains an instruction queue, and the first instruction scheduling unit can obtain the target instruction from the instruction queue for instruction scheduling after monitoring that the hardware unit for executing a certain instruction in the instruction queue is idle and the data required for executing the instruction is ready. That is, in this embodiment, the first instruction scheduling unit can take out the target instruction from the instruction queue for instruction scheduling after determining the target tensor block and the target unit.
[0093] In step S204, the first instruction scheduling unit is integrated outside the execution unit. After determining the target workgroup corresponding to the target unit, the first instruction scheduling unit can directly perform instruction scheduling in units of workgroups and schedule the target instruction to the corresponding target workgroup. Specifically, in this embodiment, after creating multiple workgroups, each workgroup is set with a corresponding workgroup identifier, that is, block id. Based on this, after determining the target workgroup allocated to the target unit for execution, the first instruction scheduling unit can directly schedule the target instruction to the target workgroup based on the workgroup identifier corresponding to the target workgroup. In this way, the first instruction scheduling unit can perform instruction scheduling in units of workgroups. In this instruction scheduling mode, referring to step S203, a single workgroup itself corresponds to at most one thread and the workgroup works in the single-program multiple-data mode. That is, in this embodiment, actually at most one thread is used to instruct the target unit to process all elements in the tensor block corresponding to a workgroup, and the situation of multiple threads collaborating to complete the same computing task is not involved. Therefore, there is no need to set a synchronization mechanism to manage the operations of multiple threads.
[0094] In addition, during the operation of a neural network model, a graphics processing unit often needs to execute multiple instructions simultaneously. For example, under flash attention, each computing unit needs to simultaneously perform data copying tasks and matrix multiplication computing tasks on different data respectively. In related technologies, when the instruction scheduler in the execution unit performs instruction scheduling in units of warps, since each warp only includes 32 threads, a data processing task may require multiple warps to complete. When the computing unit needs to simultaneously execute multiple data processing tasks, at this time, the instruction scheduler of the same execution unit needs to separately schedule different instructions to several different warps, and the number of warps to which each instruction needs to be scheduled is different depending on the number of threads for executing a single data processing task. The instruction scheduler itself cannot intelligently identify how many warps the instructions corresponding to each data processing task need to be scheduled to for execution. This requires the user to perform Warp specialization on each warp during programming, that is, to allocate an additional warp identifier to each warp and preset the instructions executed by each warp. When the instruction scheduler performs instruction scheduling, it determines which data processing task each warp is used to complete by detecting the warp identifier, so as to determine which instruction needs to be scheduled to that warp. When performing Warp specialization, it is often necessary to combine some characteristics of the underlying hardware of the computing unit to maintain the efficiency of instruction scheduling and warp execution, which greatly increases the complexity of user programming. In the instruction scheduling method proposed in this embodiment, the first instruction scheduling unit performs instruction scheduling in units of workgroups. The user does not need to perform additional Warp specialization during programming to guide the first instruction scheduling unit on which warps each instruction needs to be separately scheduled to, thereby greatly reducing the complexity of user programming.
[0095] In addition, in the related art, since instruction scheduling is performed in units of warps, the warp needs to follow the SIMT mode to execute, that is, the threads in the same warp can only be used to execute the same instruction on different data. When a user programs using a thread-block-based programming framework such as Triton, the user will specify the size of a single workgroup by setting block_size in the kernel function, and this size represents the number of elements corresponding to the single workgroup. If the value of block_size is not an integer multiple of the number of threads in a warp, when the data corresponding to the workgroup is allocated to multiple warps, at least one warp will be allocated elements that cannot be evenly distributed to all the threads in that warp, that is, at least one thread in that warp will not be allocated the data to be processed. However, since a single warp needs to follow the SIMT mode to execute, that is, the threads in a warp can only execute the same instruction on different data simultaneously, these threads that are not allocated corresponding elements cannot be used to execute other instructions either. At this time, some threads in some warps will be wasted. For example, if block_size is set to 158, that is, a workgroup contains 158 unit data. When instruction scheduling is performed in units of warps, the instruction scheduler in the execution unit will allocate these 158 unit data to different threads for execution, divide them into 5 warps and schedule the instructions to these 5 warps respectively. However, in the 5th warp, actually only 30 threads belong to these 158 threads. Due to the limitation of SIMT, the remaining two threads in this warp cannot be used to execute other instructions either and are idle, and this part of the computing resources will be wasted. Moreover, due to the idleness of these two threads, the synchronization process between the threads within this warp will become more complex. Therefore, when the user writes the kernel function, it is necessary to ensure that the data corresponding to a single workgroup can be evenly divided into several warps to ensure the utilization rate of the computing resources of the graphics processing unit, which limits the flexibility of the user when setting block_size to a certain extent. In addition, when the user writes operators based on programming frameworks such as Triton, the computing resource allocation of the graphics processing unit is often controlled in units of workgroups. However, in the computing unit, this instruction scheduling mode in units of warps is not friendly to thread-block-based programming frameworks such as Triton.In the method disclosed in this embodiment, the first instruction scheduling unit integrated outside the execution unit is used to perform instruction scheduling in units of workgroups. When performing instruction scheduling, the target instruction is directly scheduled to the workgroup corresponding to the tensor block that needs to be used to execute the target instruction, and the workgroup is scheduled to the corresponding first functional unit for execution, without creating threads through the instruction scheduler in the execution unit and then performing instruction scheduling in units of warps. Even if the data contained in a single workgroup is not an integer multiple of 32, it will not cause some threads to be idle, thus improving the flexibility of the user when setting block_size and being more friendly to programming frameworks based on thread blocks such as Triton.
[0096] In the embodiments disclosed in steps S201 to S204, by integrating the first instruction scheduling unit outside the execution unit, then, when performing instruction scheduling, first identify the target unit for executing the target instruction. When the target unit is a dedicated hardware such as a tensor core, a copy engine, or a scalar processing unit, determine the target tensor block corresponding to the target instruction and the target workgroup corresponding to the target tensor block. The target workgroup indicates the memory address range of the target tensor block. Then, the first instruction scheduling unit performs instruction scheduling in units of workgroups, directly schedules the target instruction to the target workgroup, and schedules the target workgroup loaded with the target instruction to the target unit, so that the target unit reads the target tensor block based on the memory address range indicated by the target workgroup and executes the target instruction on the target tensor block. In this way, the granularity of the instruction scheduling of the computing unit as the underlying hardware is consistent with the task decomposition granularity set by the user in the kernel function, and the first instruction scheduling unit directly schedules the target instruction to the target workgroup instead of threads or warps. At this time, there is no need to create multi-threads and warps, so there is no need to set a synchronization mechanism to manage the operations of warps, nor is it necessary to perform Warpspecialization to indicate what instructions each warp is used to execute. In this way, the instruction scheduling mode of the computing unit can be more adapted to the kernels written in programming languages such as Triton.
[0097] Detailed description of step S201
[0098] In one embodiment, before step S201, the method further includes:
[0099] Step S301, obtain a target command, where the target command is used to instruct the computing unit to start a target kernel corresponding to the target command, and the target kernel is composed of multiple instructions;
[0100] Step S302, parse the target command through a command processor to determine the target kernel to be started, determine multiple instructions to be executed based on the target kernel, and write the instructions to be executed into the instruction buffer of the first instruction scheduling unit;
[0101] Step S301 includes:
[0102] In step S303, the target instruction is obtained from the corresponding instruction buffer through the first instruction scheduling unit.
[0103] In step S301, the target command is a command stream generated and sent to the graphics unit by the host when executing the program code of the neural network model. The host can be a central processing unit. It can be understood that the graphics processing unit itself is composed of a large number of relatively simple computing units. It is suitable for tasks that require large-scale parallel computing, but it cannot handle some relatively complex control logics and computing tasks. In related technologies, the operation process of the neural network model is jointly completed by the central processing unit and the graphics processing unit. The central processing unit is used to run the program code of the neural network model and execute the relatively complex processing in it, while some relatively simple but large-scale parallel execution tasks will be scheduled to the graphics processing unit for execution. At this time, the central processing unit will generate a corresponding command stream according to the program code of the neural network model being executed and send it to the computing unit. This command stream contains multiple commands, and the target command is one of them. The target command indicates that the graphics processing unit needs to execute a data processing task.
[0104] Specifically, the program code of the neural network model is often composed of multiple operators that respectively complete specific computing tasks. Each operator can actually be regarded as a packaged kernel function. The kernel functions of these operators can be converted into a series of multiple instructions that need to be executed in order to complete the computing tasks corresponding to the operators. When the central processing unit executes these operators in the program code of the neural network model, it will generate corresponding target commands and send them to the graphics processing unit to indicate which operator's kernel function the graphics processing unit needs to start.
[0105] Exemplarily, the neural network model is provided with an average pooling layer. When the central processing unit executes the program code of this average pooling layer and sends a command to perform average pooling on a matrix to the graphics processing unit. The essence of average pooling is to calculate the average value of each region in the corresponding matrix according to the given pooling window. Based on this, after the command parser parses this command, it generates a data copy instruction for copying the data of each region of the matrix stored in the global memory to the local memory of each computing unit according to the pooling window, and a mean calculation instruction for instructing the computing unit to calculate the mean value of the data in each region of the matrix.
[0106] In step S302, the command processor is a module used to parse the target commands issued by the central processing unit into instructions that can be understood and executed by the hardware modules in the graphics processing unit. It can be understood that the underlying hardware in the graphics processing unit can only understand machine language and cannot directly understand the program code of the neural network model written by the user. Therefore, the graphics processing unit needs to parse the commands in the command stream issued by the central processing unit through the command processor and convert these commands into instructions that can be understood and executed by the hardware in the graphics processing unit. Specifically, by parsing the target command, the command processor determines the target kernel corresponding to the target command, that is, the kernel function to be started when executing the target command. Referring to the detailed description of step S301, this kernel function consists of a series of instructions required to complete the computational task of the target command. After determining the kernel function to be started, that is, determining the multiple instructions that the computational unit needs to execute subsequently, namely the instructions to be executed. It can be understood that these instructions often need to be executed in an orderly manner, and some of the instructions need to wait for other instructions to be executed before starting to execute. Based on this, after obtaining multiple instructions to be executed, these instructions to be executed are written into the instruction buffer corresponding to the first instruction scheduling unit to form an instruction queue, so that the first instruction scheduling unit can subsequently read each instruction to be executed from the corresponding instruction buffer in an orderly manner as the target instruction for instruction scheduling.
[0107] In step S303, it can be understood that referring to the relevant description of step S201, the first instruction scheduling unit monitors in real time whether the execution conditions of each instruction to be executed in the instruction buffer are met, that is, whether there are available idle workgroups and whether the data required to execute each instruction has been prepared. When the execution conditions of a certain instruction in the instruction buffer are met, the first instruction scheduling unit reads the corresponding instruction to be executed from the instruction buffer as the target instruction to be scheduled.
[0108] In the embodiments disclosed in steps S301 to S303, by obtaining the target commands issued by the central processing unit and parsing the target commands through the command processor, the target kernel to be started is determined. Then, the target kernel corresponding to the target command is started, and each instruction included in the target kernel is written as an instruction to be executed into the instruction buffer corresponding to the first instruction scheduling unit, thereby converting the target commands issued by the central processing unit into machine code instructions that can be understood and executed by the hardware modules in the computational unit. In this way, the first instruction scheduling unit can schedule by sequentially reading each instruction to be executed from the instruction buffer as the target instruction and scheduling it to the corresponding workgroups.
[0109] In one embodiment, the target kernel at least includes a main function kernel part, a scalar kernel part, and a vector kernel part. The scalar kernel part includes multiple preset scalar operation kernels, the vector kernel part includes multiple preset vector operation kernels, and the main function kernel part includes call instructions for calling at least one of the scalar operation kernels and the vector operation kernels. Each scalar operation kernel is composed of at least one scalar operation instruction, and the vector operation kernel is composed of at least one vector operation instruction.
[0110] Exemplarily, referring to Figure 4 , Figure 4 is a schematic diagram of an operator kernel according to an embodiment of the present disclosure. The operator kernel includes a main function kernel part, a scalar kernel part, and a vector kernel part. Among them, in the scalar kernel part, the 0th to 12th lines constitute the first scalar operation kernel of the scalar kernel part, that is, kernel0. Among them, the ret instruction on the 12th line indicates that kernel0 has been executed and returns to the main function kernel part, and so on. Specifically, as Figure 4 shows, the instruction numbered 38 in the main function kernel, that is, launch vector 100, is a call instruction for calling the instruction numbered 100 in the vector kernel part. When the main function kernel part executes the instruction numbered 38, it will jump to execute the instruction numbered 100 in the vector kernel part and sequentially execute each subsequent instruction in the vector kernel part until it executes the ret instruction, indicating that kernel1 in the vector kernel part has been executed and returns to the main function kernel; similarly, when it executes the instruction numbered 70 in the main function kernel, it will jump to execute the instruction numbered 13 in the scalar kernel part (that is, Figure 4 13 kernel1 in
[0111] to call kernel1 in the scalar kernel part and sequentially execute each subsequent instruction in the scalar kernel part until it executes the ret instruction. At this time, it indicates that kernel1 in the scalar kernel part has been executed and returns to the main function kernel part to execute the instruction numbered 71 in the main function kernel.
[0112] Meanwhile, when it is necessary to perform repeated scalar operations or vector operations on multiple nodes, only call instructions for jumping to the corresponding positions in the scalar kernel part or vector kernel part need to be added at the corresponding positions in the main function kernel respectively, without repeating the instructions for performing the same type of operations in the main function kernel. For example, in Figure 5 In the example shown, the 38th instruction in the main function kernel is used to call kernel1 of the vector operation kernel. If in this operator, after executing the 48th instruction of the main function kernel, it is necessary to perform the vector operation task corresponding to kernel1 of the corresponding operation kernel again, the 49th instruction of the main function kernel can be written as launch vector 100 to call the vector operation kernel kernel1 again when executing the 49th instruction of the main function kernel. In this way, the complexity of the kernel function of the operator can be further reduced, so as to reduce the difficulty of the user in writing the kernel function of the operator.
[0113] In addition, it can be understood that since the graphics processing unit itself is often used to execute some relatively simple but computationally intensive tasks that require large-scale parallel execution, and the operations of tensors with more than two dimensions are already relatively complex, and the operations on tensor data with more than two dimensions in the neural network model are mainly matrix multiplications. Therefore, the hardware structure of the tensor core itself is optimized and designed with the goal of accelerating matrix multiplication. For example, referring to the above embodiments, the tensor core is composed of a general matrix multiplication unit and an accumulation buffer. At this time, the tensor core itself is often only used to execute matrix multiplication and matrix multiplication accumulation calculation tasks, that is, the instructions executed by the tensor core actually often only include matrix multiplication instructions and matrix multiplication accumulation instructions, and the instructions executed by the tensor core are actually relatively single. Based on this, in this embodiment, there is no need to specifically set a part for filling in the tensor operation kernel in the target kernel; similarly, the copy engine itself is used to perform data copy instructions, and the types of instructions it executes are also relatively single. Based on this, in this embodiment, there is no need to set a kernel function part corresponding to the copy engine in the target kernel.
[0114] Detailed description of step S202
[0115] In one embodiment, referring to Figure 5 , step S202 includes:
[0116] Step S501, identifying the instruction type corresponding to the target instruction;
[0117] Step S502, determining the target unit for executing the target instruction based on the instruction type.
[0118] In step S501, the instruction type is used to characterize the type of data processing task corresponding to the target instruction. It can be understood that each first functional unit is actually a hardware structure, which is a low-level hardware designed for certain specific types of data processing tasks. For example, a tensor core is provided with a general matrix multiplication unit and an accumulation cache, and the tensor core is designed to execute matrix multiplication and matrix multiplication accumulation calculation tasks; each first functional unit actually only executes instructions corresponding to one or more specific types of data processing tasks. For example, a scalar processing unit is used to execute basic arithmetic operation instructions such as addition, subtraction, multiplication, and division, as well as basic logic operation instructions such as AND, OR, and NOT. Based on this, in this embodiment, multiple preset instructions can be divided into different instruction types based on the instructions executed by each first functional unit, and each instruction belongs to one of the instruction types.
[0119] Specifically, in one embodiment, the instruction type at least includes at least one of: tensor operation instructions, scalar operation instructions, vector operation instructions, and data copy instructions. Tensor operation instructions are instructions that need to be scheduled to the tensor core for execution. Tensor operation instructions refer to instructions used for operating on tensors with two or more dimensions. In this embodiment, tensor operation instructions at least include matrix multiplication instructions or matrix accumulation multiplication instructions. Scalar operation instructions are instructions that need to be scheduled to the scalar processing unit for execution. They are instructions used for operating on zero-dimensional scalar data, such as basic arithmetic instructions such as addition, subtraction, multiplication, division, exponentiation, and mean calculation, as well as basic logic operation instructions such as AND, OR, and NOT. Data copy instructions are instructions that need to be scheduled to the copy engine, and are used to instruct the copy engine to copy the data stored in one memory address to another memory address. For example, instructions to copy the data that the computing unit needs to process subsequently from the global memory to the local register or local memory of the computing unit, or to copy the data locally calculated by the computing unit to the global memory for access by other computing units in the graphics processing unit. Vector operation instructions are instructions that need to be scheduled to the execution unit for execution, and are instructions used for operating on one-dimensional vector data, such as normalization, vector dot product, etc. Based on this, for the types of first functional units integrated in the computing unit, the preset instructions are divided into multiple instruction types corresponding to the execution unit and each first functional unit respectively. Thus, the target unit for executing the target instruction can be determined based on the instruction type of the target instruction.
[0120] In a possible embodiment, it can be understood that each first functional unit can only be used to execute specific types of instructions, and the first functional units in the computing unit include at least one of a tensor core, a copy engine, and a scalar processing unit. That is, depending on the hardware structure of the computing unit itself, the integrated first functional units in the computing unit are different. Based on this, in this embodiment, the division of instruction types is determined based on the integrated first functional units in the computing unit. Depending on the different hardware structures of the computing unit itself, the instruction types may include only one or more of the above-mentioned tensor operation instructions, scalar operation instructions, vector operation instructions, and data copy instructions. For example, in addition to the execution unit in the computing unit, only the tensor core is integrated as the first functional unit, and the scalar processing unit and the copy engine are not integrated. At this time, the instruction types only include tensor operation instructions and vector operation instructions. When and only when the instruction type of the target instruction is a tensor operation instruction, the tensor core is determined as the corresponding computing unit. If the target instruction is a vector operation instruction or other instructions whose instruction types are not divided, the execution unit is determined as the corresponding target unit. Similarly, when the tensor core, the scalar processing unit, and the copy engine are integrated in the computing unit at the same time, the instruction types may include tensor operation instructions, scalar operation instructions, vector operation instructions, and data copy instructions. Of course, if other dedicated hardware is integrated in the computing unit in addition to the tensor core, the copy engine, and the scalar processing unit, the instruction type corresponding to the dedicated hardware can also be determined based on the data processing tasks executed by the dedicated hardware, and the instructions that need to be scheduled to the dedicated hardware for execution are divided into the instruction types corresponding to the dedicated hardware.
[0121] In another possible embodiment, when setting the instruction types, the types of the first functional units integrated in the computing unit may not be considered, but the preset instructions are directly divided into multiple instruction types such as tensor operation instructions, scalar operation instructions, vector operation instructions, and data copy instructions. When the first instruction scheduling unit performs instruction scheduling, after identifying the instruction type of the target instruction, the target unit corresponding to the target instruction is determined according to the situation of the first functional units integrated in the computing unit itself. That is, when the first functional unit corresponding to the instruction type is integrated in the computing unit, the first functional unit is determined as the corresponding target unit. If the first functional unit corresponding to the instruction type is not integrated in the computing unit, the execution unit is determined as the corresponding target unit. Exemplarily, only the tensor core and the copy engine are integrated as the first functional units in the computing unit, and the preset instruction types include tensor operation instructions, scalar operation instructions, vector operation instructions, and data copy instructions. At this time, if the instruction type corresponding to the target instruction is a scalar operation instruction, since the corresponding first functional unit is not integrated in the computing unit, the execution unit is determined as the corresponding target unit.
[0122] In step S502, referring to step S501, it can be known that the execution unit and each first functional unit correspond to different instruction types respectively. After determining the instruction type corresponding to the target instruction, the corresponding target unit can be determined based on the instruction type.
[0123] Exemplarily, the target instruction is a basic arithmetic operation instruction for adding two data, and its instruction type is a scalar operation instruction. When a scalar operation unit is integrated in the computing unit, the scalar operation instruction will be scheduled to the scalar processing unit for processing, that is, the target unit corresponding to the target instruction is the scalar processing unit.
[0124] In the embodiments disclosed in steps S501 to S502, by presetting the instruction types corresponding to the execution unit and each first functional unit, when the first instruction scheduling unit obtains the target instruction that needs to be scheduled and executed, it identifies the instruction type corresponding to the target instruction, and based on this instruction type, determines the target unit for executing the target instruction. Based on this, after the first instruction scheduling unit obtains the target instruction, it can automatically determine which hardware module in the computing unit the target instruction needs to be scheduled to for execution.
[0125] In one embodiment, referring to Figure 6 , before step S501, the method further includes:
[0126] Step S601, creating a first instruction set;
[0127] Step S602, dividing the preset instructions into multiple instruction subsets, each instruction subset including at least one preset instruction, and the preset instructions in a single instruction subset correspond to the same instruction type;
[0128] Step S501 includes:
[0129] Step S603, identifying the instruction subset corresponding to the target instruction to determine the instruction type corresponding to the target instruction.
[0130] In step S601, the first instruction set is a preset instruction set, and the first instruction set includes multiple preset instructions; the preset instructions are preset by the user and are machine code instructions that can be understood and executed by the underlying hardware of the graphics processing unit. It can be understood that each hardware module in the computing unit can only understand and execute machine code instructions and cannot directly understand the program code written by the user. Based on this, in this embodiment, it is necessary to pre-create a first instruction set composed of machine code instructions that can be understood and executed by each hardware module. In this way, by subsequently converting the program code written by the user into a series of corresponding instructions and scheduling them to the corresponding hardware modules for execution, the corresponding data processing tasks can be completed by each hardware module in the computing unit.
[0131] In step S602, the preset instructions include various data processing instructions that the computing unit can execute. Referring to the description of step S501 above, when the first functional unit is integrated in the computing unit, it is necessary to identify the instruction type corresponding to the instruction to determine which working group corresponding to which hardware module in the computing unit each instruction needs to be scheduled to. Based on this, in this embodiment, after creating the first instruction set, the preset instructions in the first instruction set can be divided based on different instruction types to obtain instruction subsets corresponding to each instruction type respectively.
[0132] In step S603, since the first instruction set is divided into instruction subsets corresponding to each instruction type respectively, each preset instruction will be divided into an instruction subset. Based on the instruction subset to which the target instruction belongs and the correspondence between each instruction subset and the instruction type, the instruction type corresponding to the target instruction can be determined.
[0133] In the embodiments disclosed in steps S601 to S603, after creating the first instruction set, the preset instructions in the first instruction set are divided into instruction subsets corresponding to different instruction types. In this way, when the target instruction needs to be instruction-scheduled subsequently, the instruction type corresponding to the target instruction can be determined based on the instruction subset to which the target instruction belongs.
[0134] Detailed description of step S204
[0135] In one embodiment, referring to Figure 7 , before step S204, the method further includes:
[0136] Step S701, obtaining working group creation parameters, where the working group creation parameters include a working group size and a working group number parameter, and the working group size represents the size of the tensor block corresponding to a single working group in each dimension;
[0137] Step S702, creating multiple working groups based on the working group number parameter and assigning working group identifiers to each working group, and dividing the tensor to be processed into multiple tensor blocks based on the working group size, where the working groups and the tensor blocks are in one-to-one correspondence.
[0138] In step S701, the workgroup creates parameters for indicating how the graphics processing unit divides the tensors to be processed and creates workgroups corresponding to each tensor block, so as to allocate the computing resources of the graphics processing unit to each workgroup, thereby executing the same instruction or different instructions in parallel for different data. In this embodiment, the workgroup creation parameters at least include the workgroup size and the workgroup number parameter. Among them, the workgroup size is used to indicate the size of the tensor block corresponding to a single workgroup in each dimension, that is, the number of elements included in a single workgroup in each dimension. Based on the workgroup size, the number of elements included in a single workgroup can be determined, and the workgroup number parameter is used to indicate the number of workgroups. In one embodiment, the workgroup number parameter can be a value directly representing the number of workgroups to be created, or it can also be an array of the sizes of the N-dimensional grid formed by the workgroups. For example, the workgroup number parameter can be (M, N, K), which means that M * N * K workgroups need to be created, and these workgroups are organized into a three-dimensional grid with a shape of (M, N, K). It should be noted that in this embodiment, the product of the workgroup size and the workgroup number parameter can correspond to the total number of elements in the tensor to be processed, so as to ensure that each tensor block obtained by dividing the tensor to be processed corresponds to a workgroup.
[0139] In another embodiment, the workgroup number parameter can be a set of parameters used to characterize the shape of an N-dimensional grid formed by the workgroups. Specifically, referring to step S203, the user can control the shape of the grid formed by the workgroups by setting grid_size in the kernel function, so as to organize each workgroup into a workgroup grid. Exemplarily, the user can set gridsize to (a2, b2, c2) in the kernel function. At this time, it indicates that the workgroup grid includes a2 workgroups along the x dimension, b2 workgroups along the y dimension, and c2 workgroups along the z dimension. At this time, the number of workgroups to be created is a2 * b2 * c2.
[0140] In step S702, the workgroup number parameter is used to characterize the number of workgroups to be created, and the workgroup size characterizes the number of elements included in the tensor block corresponding to a single workgroup in each dimension. Based on the workgroup number parameter, workgroups are created, and at the same time, based on the workgroup size, the tensors to be processed are processed, so as to obtain multiple workgroups and the tensor blocks corresponding to each workgroup.
[0141] It can be understood that after creating multiple workgroups, each workgroup corresponds to a different tensor block. It is necessary to assign corresponding workgroup identifiers to each workgroup to distinguish each workgroup, and the identifiers corresponding to each workgroup are different. For example, a serial number can be assigned to each workgroup as the identifier corresponding to the workgroup.
[0142] In another possible implementation, multiple working groups are organized into an N-dimensional grid of working groups, and each working group serves as a grid point in this grid of working groups. At this time, each working group has corresponding coordinates to represent its position in the grid. Since the coordinates of each working group in the grid of working groups are different, based on this, the coordinates corresponding to each working group can be used as the working group identifier corresponding to this working group. For example, if the working groups form a three-dimensional grid, the working group identifier corresponding to each working group can be expressed as (x1, y1, z1).
[0143] At this time, referring to the relevant description of step S203 above, subsequently, based on the identifier corresponding to the working group, the address of the first element in the tensor to be processed, and the working group size, the address range of the elements corresponding to the tensor block corresponding to this working group in the memory can be indicated.
[0144] It should be noted that, referring to the above step S301, the target kernel often includes multiple instructions that need to be executed in sequence. When executing the kernel function corresponding to each operator, only one creation of a workgroup is required. Each workgroup corresponds to a tensor block in the original input tensor of the operator and the intermediate data generated after each instruction is executed for this tensor block. These workgroups are allocated to execute multiple instructions in the kernel function. Specifically, that is, each workgroup will be allocated to the corresponding hardware module after the scheduled instruction is executed. When the corresponding hardware module finishes executing the workgroup, the workgroup will enter the idle state until the next instruction to be executed is scheduled to this workgroup. At this time, if in the related art, the instruction scheduling is performed in units of warps, the data corresponding to a workgroup will be divided into multiple warps for instruction scheduling respectively. At this time, once a warp finishes executing the previous instruction earlier than other warps, the computing resources of this warp will be released in advance, and the instruction scheduler will schedule the next instruction to this warp in advance, so that this warp is used to execute the next instruction. This will cause the threads in a workgroup to execute different instructions simultaneously, resulting in race conditions or other unpredictable situations. Therefore, when using the method in the related art to perform instruction scheduling in units of warps, it is necessary to set up a complex synchronization mechanism between warps to restrict each warp to prevent some warps from being used to execute the next instruction in advance. When using the instruction scheduling method of the present disclosure to perform instruction scheduling in units of workgroups by using the first instruction scheduling unit, there is no need to create multiple threads and warps. Instead, the instruction is directly scheduled to this workgroup, and then the workgroup is scheduled to the target unit for execution. Before this workgroup is executed completely, this workgroup will not enter the idle state and the occupied computing resources will not be released, and the first instruction scheduling unit will not schedule the next instruction to this workgroup. Only when this workgroup is executed completely, that is, all elements in the corresponding tensor block have been executed for the data processing task corresponding to the target instruction, at this time, the first instruction scheduling unit will schedule the next instruction to this workgroup. The whole process neither involves the processing of multiple threads nor needs to consider the situation of race conditions caused by scheduling the next instruction to some warps that finish executing earlier in advance. Thus, there is no need to set up a relatively complex synchronization mechanism to restrict the behavior of each thread.
[0145] In the embodiments disclosed in steps S701 to S703, by obtaining the workgroup creation parameters, creating multiple workgroups based on these parameters, then dividing the tensor to be processed into tensor blocks corresponding to each workgroup according to the workgroup size, and determining the workgroup identifier corresponding to each workgroup, in this way, the memory address range of the elements in the tensor blocks corresponding to each workgroup can be determined according to the workgroup identifier subsequently.
[0146] Detailed description of the instruction scheduling of the compatible execution unit in units of warps
[0147] In one embodiment, a second instruction scheduling unit is further provided in each execution unit. Referring to Figure 8 , the method further includes:
[0148] Step S801, if the target unit is an execution unit, determine the target workgroup assigned to the target unit, and create a plurality of threads corresponding to the target workgroup, where each thread corresponds to the data in the tensor block corresponding to the target workgroup;
[0149] Step S802, the first instruction scheduling unit sends the target instruction to the second instruction scheduling unit in the execution unit;
[0150] Step S803, in the execution unit, determine the created threads as a plurality of warps, where each warp includes a second number of threads;
[0151] Step S804, in the execution unit, the second instruction scheduling unit schedules the target instruction to the threads in the plurality of warps respectively.
[0152] In step S801, if the target unit is an execution unit, the hardware structure of the execution unit is designed based on the mode of instruction scheduling and warp concurrent execution in units of warps for the execution unit, that is, the hardware structure of the execution unit itself is more suitable for instruction scheduling and concurrent execution of threads in units of warps. At the same time, the memory resources and computing resources of a single execution unit are relatively limited, and the size of the workgroup is set by the user when writing the kernel function, and the size of a single workgroup is variable. When the size of a single workgroup is large, that is, when the number of elements in the tensor block corresponding to the workgroup is large, the computing resources and memory resources of the execution unit itself may not be sufficient to concurrently execute the target instruction for all the elements in the tensor block corresponding to the workgroup, that is, the execution unit itself may not support instruction scheduling and execution in units of workgroups. Based on this, in this embodiment, if the target unit is an execution unit, the instruction scheduling is not performed through the first instruction scheduling unit, but a plurality of threads are created in the execution unit to perform instruction scheduling in units of warps through the second instruction scheduling unit in the execution unit subsequently. Specifically, the number of created threads can be determined based on the number of elements in the target tensor block. For example, each thread corresponds to one element in the target tensor block.
[0153] In step S802, as described in step S801, the execution unit itself is more suitable for executing data processing tasks in units of warps, while the first instruction scheduling unit schedules instructions in units of workgroups. If the target instruction is directly scheduled by the first instruction scheduling unit to the workgroup composed of the threads allocated to the execution unit, and then the workgroup loaded with the target instruction is scheduled to the execution unit for execution, this may instead cause the execution unit to be difficult to execute and manage these threads in units of warps. Based on this, in this embodiment, when the target unit is the execution unit, the first instruction scheduling unit can directly send the target instruction to the second instruction scheduling unit in the execution unit, and the second instruction scheduling unit schedules the target instruction to each thread in units of warps, so that it can better adapt to the mode in which the execution unit itself manages and executes threads in units of threads.
[0154] In step S803, the second number refers to the number of threads in a single warp, and the second number can be 16, 32, 64, etc. Specifically, the second number is determined by the graphics processing unit itself, and on the premise that the graphics processing unit is determined, the second number is fixed and unchanged. As described in steps S801 and S802, the execution unit itself is more suitable for managing and executing threads in units of warps. Therefore, in this embodiment, after multiple threads are scheduled to the execution unit, the threads allocated to the execution unit are divided into multiple warps, and then the second instruction scheduling unit in the execution unit can schedule instructions in units of warps.
[0155] In step S804, the second instruction scheduling unit can be a warp scheduler, which schedules instructions in units of warps, that is, it schedules one instruction to 32 threads in the same warp each time. Specifically, after the threads allocated to the execution unit are divided into multiple warps, the process of the second instruction scheduling unit scheduling instructions in units of warps can refer to the process of scheduling instructions in units of warps in the related art, which will not be elaborated here. It should be emphasized that in the embodiments of the present disclosure, only when the target instruction is an instruction that needs to be allocated to the execution unit to complete, will the second instruction scheduling unit in the execution unit schedule instructions in units of warps. When the target instruction is an instruction executed by a tensor core, a copy engine, or a scalar processing unit, the second instruction scheduling unit will not schedule instructions in units of warps, but the first instruction scheduling unit integrated outside the execution unit will schedule instructions in units of workgroups.
[0156] In the embodiments disclosed in steps S801 to S804, when the target unit is an execution unit, instead of using the first instruction scheduling unit to perform instruction scheduling in units of workgroups, the threads assigned to the target unit are directly scheduled to the execution unit, and the second instruction scheduling unit in the execution unit is used to perform instruction scheduling in units of warps, so as to be more adapted to the data processing mode of the execution unit itself, enabling the instruction scheduling method of the present disclosure to better be compatible with the hardware characteristics of the execution unit itself.
[0157] Description of the apparatus and device according to the embodiments of the present disclosure
[0158] Referring to Figure 9 , Figure 9 is a schematic structural diagram of an instruction scheduling apparatus 900 proposed by the present disclosure. The apparatus includes:
[0159] An instruction acquisition unit 910, configured to acquire a target instruction through the first instruction scheduling unit;
[0160] A first identification unit 920, configured to identify a target unit for executing the target instruction, where the target unit is one of an execution unit or a first functional unit;
[0161] A target workgroup determination unit 930, configured to, when the target unit is a first functional unit, determine a target tensor block corresponding to the target instruction and a target workgroup corresponding to the target tensor block, where the target workgroup is used to indicate the memory address range of the target tensor block;
[0162] A first scheduling unit 940, configured to schedule the target instruction to the target workgroup through the first instruction scheduling unit, and schedule the target workgroup to the target unit, so that the target unit reads the target tensor block based on the memory address range and executes the target instruction on the target tensor block.
[0163] The instruction scheduling apparatus of the present disclosure is used to execute the instruction scheduling method in the above embodiments. The specific processing process is the same as that of the instruction scheduling method in the above embodiments and will not be described herein again.
[0164] The embodiments of the present disclosure further provide an electronic device 1000, including:
[0165] At least one processor, and,
[0166] A memory communicatively connected to the at least one processor; wherein,
[0167] The memory stores instructions, and the instructions are executed by the at least one processor, so that when the at least one processor executes the instructions, the method in any one of the above embodiments of the present application is implemented.
[0168] Next, in conjunction with Figure 10A detailed description of the hardware structure of the electronic device is provided. The electronic device includes: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050.
[0169] The processor 1010 can be implemented in ways such as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present disclosure;
[0170] The memory 1020 can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called by the processor 1010 to execute the instruction scheduling method of the embodiments of the present disclosure;
[0171] The input / output interface 1030 is used to implement information input and output;
[0172] The communication interface 1040 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); and
[0173] The bus 1050 transmits information between the various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040);
[0174] Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 achieve communication connections with each other inside the device through the bus 1050.
[0175] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the instruction scheduling method of the above embodiments, which will not be elaborated here.
[0176] The computing unit of the present disclosure is used to execute the instruction scheduling method of the above embodiments, and its specific processing process is the same as that of the instruction scheduling method of the above embodiments, which will not be elaborated here.
[0177] In the description of the present disclosure and the above accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprise" and "include" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0178] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0179] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.
[0180] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0181] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0182] In addition, each functional unit in various embodiments of the present disclosure may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0183] It should also be understood that the various embodiments provided by the present disclosure can be combined arbitrarily to achieve different technical effects.
[0184] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.
Claims
1. An instruction scheduling method, characterized in that, Applied to a computing unit, the computing unit includes a first instruction scheduling unit, at least one execution unit, and at least one first functional unit, and the first functional unit includes at least one of a copy engine, a tensor core, and a scalar processing unit. The method includes: Obtain a target instruction through the first instruction scheduling unit; Identify a target unit for executing the target instruction, where the target unit is one of the execution unit or the first functional unit; When the target unit is the first functional unit, determine a target tensor block corresponding to the target instruction and a target workgroup corresponding to the target tensor block, where the target workgroup is used to indicate the memory address range of the target tensor block; Schedule the target instruction to the target workgroup through the first instruction scheduling unit, and schedule the target workgroup to the target unit, so that the target unit reads the target tensor block based on the memory address range and executes the target instruction on the target tensor block.
2. The instruction scheduling method according to claim 1, wherein The first instruction scheduling unit is provided with a corresponding instruction buffer area, and the computing unit further includes a command processor. Before obtaining the target instruction through the first instruction scheduling unit, the method further includes: Obtain a target command, where the target command is used to instruct the computing unit to start a target kernel corresponding to the target command, and the target kernel is composed of multiple instructions; Parse the target command through the command processor to determine the target kernel to be started, determine multiple instructions to be executed based on the target kernel, and write the instructions to be executed into the instruction buffer area of the first instruction scheduling unit; The obtaining the target instruction through the first instruction scheduling unit includes: Obtain the target instruction from the corresponding instruction buffer area through the first instruction scheduling unit.
3. The instruction scheduling method according to claim 2, wherein The target kernel includes at least a main function kernel part, a scalar kernel part, and a vector kernel part. The scalar kernel part includes multiple preset scalar operation kernels, the vector kernel part includes multiple preset vector operation kernels, and the main function kernel part calls the scalar operation kernels and the vector operation kernels through call instructions.
4. The instruction scheduling method according to claim 1, characterized in that, The identifying the target unit for executing the target instruction includes: Identify the instruction type corresponding to the target instruction, where the instruction type is used to characterize the type of data processing task corresponding to the target instruction, and the execution unit and each of the first functional units respectively correspond to different instruction types; Determine the target unit for executing the target instruction based on the instruction type.
5. The instruction scheduling method according to claim 4, wherein Before the identifying the target unit for executing the target instruction, the method further includes: Create a first instruction set, where the first instruction set includes multiple preset instructions, and the target instruction is one of the preset instructions; Divide the preset instructions into multiple instruction subsets, each instruction subset includes at least one of the preset instructions, and the preset instructions in a single instruction subset correspond to the same instruction type; The identifying the instruction type corresponding to the target instruction includes: Identify an instruction subset corresponding to the target instruction to determine an instruction type corresponding to the target instruction.
6. The instruction scheduling method according to claim 4, wherein The instruction type includes at least one of the following: A tensor operation instruction, which indicates an instruction to be executed by the tensor core; A scalar operation instruction, which indicates an instruction to be executed by the scalar processing unit; A vector operation instruction, which indicates an instruction to be executed by the execution unit; A data copy instruction, which is an instruction to be executed by the copy engine.
7. The instruction scheduling method according to claim 1, wherein Before determining a target tensor block corresponding to the target instruction and a target workgroup corresponding to the target tensor block, the method further includes: Obtaining workgroup creation parameters, where the workgroup creation parameters include a workgroup size and a workgroup number parameter, and the workgroup size characterizes the size of a tensor block corresponding to a single workgroup in each dimension; Creating a plurality of workgroups based on the workgroup number parameter, assigning a workgroup identifier to each of the workgroups, and dividing a tensor to be processed into a plurality of tensor blocks based on the workgroup size, where the workgroups and the tensor blocks are in one-to-one correspondence.
8. The instruction scheduling method according to claim 1, wherein Each of the execution units includes a second instruction scheduling unit, and the method further includes: If the target unit is the execution unit, determining a target workgroup assigned to the target unit, and creating a plurality of threads corresponding to the target workgroup, where each of the threads corresponds to data in a tensor block corresponding to the target workgroup; The first instruction scheduling unit sends the target instruction to the second instruction scheduling unit in the execution unit; In the execution unit, determining the created threads as a plurality of warps, where each warp includes a second number of threads; In the execution unit, scheduling the target instruction to the threads in the plurality of warps respectively through the second instruction scheduling unit.
9. An instruction scheduling device, characterized in that, The apparatus includes: An instruction acquisition unit, configured to acquire a target instruction through a first instruction scheduling unit; A first identification unit, configured to identify a target unit for executing the target instruction, where the target unit is one of an execution unit or a first functional unit; A target workgroup determination unit, configured to determine a target tensor block corresponding to the target instruction and a target workgroup corresponding to the target tensor block when the target unit is the first functional unit, where the target workgroup is used to indicate a memory address range of the target tensor block; A first scheduling unit, configured to schedule the target instruction to the target workgroup through the first instruction scheduling unit, and schedule the target workgroup to the target unit, so that the target unit reads the target tensor block based on the memory address range and executes the target instruction on the target tensor block.
10. An electronic device, characterized in that, The electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for implementing connection communication between the processor and the memory. When the program is run by the processor, it implements the instruction scheduling method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be run by one or more processors to implement the instruction scheduling method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Data processing method and device, storage medium and electronic device
CN110780921A