Neural network operator acceleration method and device and graphics processor
By identifying neural network operators with insufficient parallelism on the GPU, performing reduction operation division and subtask fusion, the problem of hardware resource waste is solved and the performance of reduction operations and hardware utilization are improved.
Patent Information
- Application Number
- CN202410382011.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-09-30
AI Technical Summary
On hardware platforms with powerful computing capabilities such as GPUs, large-scale reduction operations cannot be calculated in parallel on multiple computing units, resulting in a waste of hardware resources.
By identifying neural network operators with insufficient parallelism, we divide the operation into reduction operations, assign tasks to multiple computing units, and perform subtask fusion on the reduction operation subtasks to improve the utilization of parallelism.
It improves the performance of reduction operations, reduces the memory access overhead of intermediate variables, and improves the utilization of hardware computing units.
Smart Images

Figure CN120725075A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and in particular to the field of neural network operator acceleration. Background Art
[0002] In large language models, there is a type of operator that lacks parallelism on hardware platforms with strong computing power, such as GPUs, which wastes some hardware resources.
[0003] With the rapid development of large language model technology, optimizing operators within these models is becoming an increasingly important task. Matrix-vector multiplication (such as GEMV and GEMM) and convolution (CONV) operations play a crucial role in inference tasks involving large language models. GEMV, GEMM, and convolution (CONV) operators are also fundamental or core components of many deep learning operations, often appearing in fully connected layers, recurrent neural networks, and the attention layer commonly used in large language models.
[0004] Existing loop transformations for reducing operations primarily utilize loop splitting and loop reordering to improve data locality and optimize memory access patterns. By reducing the size of the loop body, the data processed within the loop can be made more cache-friendly, reducing cache misses. By changing the order of the loops, data access patterns can be made more consistent with the hardware's memory hierarchy, while enhancing both temporal and spatial locality, allowing the program to access the same data more frequently and improving cache hit rates.
[0005] On powerful hardware like GPUs, existing optimizations for matrix-vector multiplication operations fail to fully utilize the hardware's computing performance. For large-scale reduction operations, regardless of loop splitting and reordering, the operations can only be performed on a single hardware unit, not across multiple units. This wastes idle computing resources. Summary of the Invention
[0006] In order to solve the current problem that large-scale reduction operations cannot be calculated in parallel on multiple computing units at the same time, the present invention discloses a neural network operator acceleration method, comprising the following steps:
[0007] Identify neural network operators that are insufficiently parallelized in the computational units of a graphics processor;
[0008] Performing reduction operations on the neural network operators with insufficient parallelism to obtain reduction operation subtasks;
[0009] Allocating the reduction operation subtask to multiple computing units of the graphics processor to execute the reduction operation subtask;
[0010] Perform subtask fusion on the reduction operation subtask to obtain the reduction operation result.
[0011] In one embodiment of the above method of the present invention, the identification of neural network operators with insufficient parallelism in the computing units of the graphics processor is obtained by the operation dimensions that the operators need to perform outside the reduced dimensions.
[0012] In one embodiment of the above method of the present invention, the step of identifying a neural network operator having insufficient parallelism in a computing unit of a graphics processor, obtained by calculating the dimensions of operations that the operator needs to perform outside the reduced dimensions, further includes:
[0013] Calculating the degree of parallelism utilization of the neural network operators in the computing units of the graphics processing unit;
[0014] When the parallelism utilization degree is less than a division threshold, dividing the neural network operator;
[0015] When the parallelism utilization degree is greater than or equal to the division threshold, an attempt is made to divide the neural network operator.
[0016] In one embodiment of the above method of the present invention, the step of calculating the parallelism utilization degree of the neural network operators in the computing units of the graphics processor further comprises:
[0017] The degree of parallelism utilization is calculated by P = α * S / H, where S represents the operation dimension that the neural network operator needs to perform outside the specification dimension, H represents the number of computing units on the graphics processor, and α is a parallelism correction parameter related to the specific hardware platform and specific task, 0 < α ≤ 1.
[0018] In one embodiment of the above method of the present invention, the reduction operation division of the neural network operator with insufficient parallelism is performed in the reduction dimension division.
[0019] In one embodiment of the above method of the present invention, the step of performing reduction operation division on the neural network operators with insufficient parallelism in the reduction dimension division further includes:
[0020] The first operator and the second operator of the neural network operator with insufficient parallelism are divided according to a first division quantity on the specification dimension respectively.
[0021] In an embodiment of the above method of the present invention, the dividing according to a first number of divisions is uniformly dividing according to a first number of divisions.
[0022] In one embodiment of the above method of the present invention, the reduction operation subtask further includes:
[0023] A first reduction operation subtask is to perform the first-partition-number reduction operations on the first operator and the second operator after being divided according to the first partition number;
[0024] The second reduction operation subtask is used to perform a reduction operation of a first partition quantity dimension on the result of the first reduction operation subtask.
[0025] In one embodiment of the above method of the present invention, the subtask fusion of the reduction operation subtask is performed by an operator fusion method or a non-operator fusion method.
[0026] In one embodiment of the above method of the present invention, the step of fusing the subtasks of the reduction operation subtasks in a non-operator fusion manner further includes:
[0027] The operators of the reduction operation subtask are mapped to the computing units of the graphics processor according to the reduction dimension and the non-reduction dimension for execution.
[0028] In one embodiment of the above method of the present invention, the step of fusing the execution reduction operation subtask and the execution subtask by means of operator fusion further includes:
[0029] The operators of the reduction operation subtask are mapped to the computing units of the graphics processor according to the non-reduction dimensions for execution.
[0030] In one embodiment of the above method of the present invention, the neural network operator is a matrix multiplication operator and / or a convolution operator.
[0031] The present invention also discloses a neural network operator acceleration device for implementing any of the above methods, comprising:
[0032] an identification module for identifying neural network operators with insufficient parallelism in a computational unit of a graphics processor;
[0033] A reduction operation partitioning module is used to perform reduction operation partitioning on the neural network operator with insufficient parallelism to obtain reduction operation subtasks;
[0034] a subtask execution module, configured to distribute the reduction operation subtask to a plurality of computing units of the graphics processor to execute the reduction operation subtask;
[0035] The subtask fusion module is used to perform subtask fusion on the reduction operation subtasks to obtain the reduction operation results.
[0036] The present invention also discloses a graphics processor, including a video memory, a video memory controller and an input and output unit connected to a computing unit, and also includes the neural network operator acceleration device as described above.
[0037] The present invention also discloses a storage medium for storing a computer control program, wherein the computer control program is used to execute the steps of any of the above methods.
[0038] The present invention discloses a neural network operator acceleration method, device, and graphics processor that identifies large-scale reduction operations and divides the identified operators. These divided subtasks are then assigned to multiple computing units for parallel computation. These subtasks are then fused, reducing the memory access overhead of intermediate variables to obtain the final reduction result, thereby improving the performance of the reduction operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 The figure is a flow chart of a neural network operator acceleration method according to one embodiment of the present invention.
[0040] Figure 2 Schematic diagram of the reduction operation division of neural network operators in one embodiment of the present invention.
[0041] Figure 3 Schematic diagram of subtask fusion for reduction operation subtasks in one embodiment of the present invention.
[0042] Figure 4 Schematic diagram of subtask fusion for reduction operation subtasks in another embodiment of the present invention.
[0043] Figure 5 Schematic diagram of performing reduction operations on neural network operators in another embodiment of the present invention.
[0044] Figure 6 This is a block diagram of the composition of a neural network operator acceleration device in one embodiment of the present invention.
[0045] Figure 7 FIG. 4 is a block diagram of a graphics processor according to an embodiment of the present invention.
[0046] Wherein, the reference numerals:
[0047] 1: First operator
[0048] 2: Second operator
[0049] 1': first child operator
[0050] 2': Second child operator
[0051] 1”: The third child operator
[0052] 2”: The fourth operator
[0053] 3: Reduced operation result
[0054] 10: Neural Network Operator Accelerator
[0055] 11: Identification module
[0056] 12: Reduction operation partition module
[0057] 13: Subtask execution module
[0058] 14: Subtask Fusion Module
[0059] 100: Graphics Processor
[0060] 110: Computing unit
[0061] 120: Video memory
[0062] 130: Video memory controller
[0063] 140: Input and output unit
[0064] T: the number of first divisions
[0065] T1: Second division quantity
[0066] T2: The third division number
[0067] Re1: First reduction operation subtask
[0068] Re2: Second reduction operation subtask DETAILED DESCRIPTION
[0069] The technical solution of the present invention is described in detail below with reference to the accompanying drawings and specific embodiments to further understand the purpose, solution and beneficial technical effects of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0070] It should be noted that, in this specification, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, article or device. In the absence of further restrictions, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.
[0071] Certain terms are used in this specification and the appended claims to refer to specific components or parts. Persons skilled in the art will understand that technology users or manufacturers may use different nouns or terms to refer to the same component or part. This specification and the appended claims do not distinguish components or parts based on differences in name, but rather on differences in their functions.
[0072] In the present invention, terms such as "upper," "lower," "left," "right," "front," "back," "top," "bottom," "inner," "outer," "center," "vertical," "horizontal," "transverse," and "longitudinal" indicate positions or locations based on the positions or locations shown in the accompanying drawings. These terms are primarily intended to better describe the present invention and its embodiments and are not intended to limit the devices, elements, or components indicated to having a specific orientation, or to being constructed or operated in a specific orientation.
[0073] Furthermore, some of the above terms may be used to express other meanings besides indicating a position or location. For example, the term "on" may also be used to indicate a dependency or connection in certain circumstances. Those skilled in the art will understand the specific meanings of these terms in the present invention based on the specific circumstances.
[0074] Furthermore, the terms "installed," "disposed," "provided with," "connected," and "connected" should be interpreted broadly. For example, they can refer to fixed connections, removable connections, or integral structures; mechanical connections or electrical connections; direct connections or indirect connections through an intermediary; or internal communication between two devices, elements, or components. Those skilled in the art will understand the specific meanings of these terms in the present invention based on specific circumstances.
[0075] About Deep Learning Automatic Tuning:
[0076] In the field of deep learning, automatic tuning plays a crucial role. The performance of deep learning models varies significantly across different hardware. Automatic tuning can help optimize models to fully utilize the capabilities of specific hardware, thereby improving the model's efficiency and performance on that hardware. This optimization typically involves adjusting the underlying computation and data storage methods of the deep learning model to maximize the characteristics of the target hardware.
[0077] TVM (Tensor Virtual Machine) is a representative work in the field of automatic tuning. It is an open source end-to-end tensor compiler framework for optimizing the performance of deep learning models on various hardware. TVM uses loop transformation as the main means of automatic tuning. TVM uses scheduling primitives to perform loop transformation operations. TVM mainly uses the following loop transformation methods:
[0078] 1. Loop unrolling: Reducing the number of loop iterations by duplicating the code within the loop body can reduce loop control overhead and help improve instruction-level parallelism.
[0079] 2. Loop splitting: Splitting a loop into two nested loops helps reduce the size of the loop body, thereby improving cache utilization.
[0080] 3. Loop fusion: Merging multiple independent loops into one loop helps reduce the number of memory accesses and improve data locality.
[0081] 4. Loop reordering: Changing the order of nested loops can help better utilize data locality, especially in multidimensional array operations.
[0082] 5. Loop tiling: Breaking the loop into smaller blocks or "tiles" that can fit into cache more efficiently. This is particularly useful for processing large arrays and can significantly improve cache hit rates.
[0083] Optimizing operators in large language models is one of the automatic tuning steps of TVM.
[0084] Reduction operation: The GEMV operator needs to multiply a matrix A of shape M*K and a vector x of length K. Its calculation expression is as follows:
[0085] y i =a i1 x1+a i2 x2+…+a ik x k ,i=1,2,…,m among them, a ij is the element in the i-th row and j-th column of matrix A, x j is the jth element in vector x.
[0086] The accumulation operation performed to obtain the element y can be regarded as an operation performed on the dimension with a length of K, which is usually called a reduction operation.
[0087] Please refer to Figure 1 and Figure 2 The present invention discloses a neural network operator acceleration method, comprising the following steps:
[0088] Step S1: Identifying neural network operators with insufficient parallelism in the computational units of the graphics processor;
[0089] Step S2: performing reduction operations on the neural network operators with insufficient parallelism to obtain reduction operation subtasks;
[0090] Step S3: allocating the reduction operation subtask to multiple computing units of the graphics processor to execute the reduction operation subtask;
[0091] Step S4: perform subtask fusion on the reduction operation subtask to obtain the reduction operation result 3.
[0092] In one embodiment of the above method of the present invention, the identification of neural network operators with insufficient parallelism in the computing units of the graphics processor is obtained by the operation dimensions that the operators need to perform outside the reduced dimensions.
[0093] In one embodiment of the above method of the present invention, the step of identifying a neural network operator having insufficient parallelism in a computing unit of a graphics processor, obtained by calculating the dimensions of operations that the operator needs to perform outside the reduced dimensions, further includes:
[0094] Calculating the degree of parallelism utilization of the neural network operators in the computing units of the graphics processing unit;
[0095] When the parallelism utilization degree is less than a division threshold, dividing the neural network operator;
[0096] When the parallelism utilization degree is greater than or equal to the division threshold, an attempt is made to divide the neural network operator.
[0097] In one embodiment of the above method of the present invention, the step of calculating the parallelism utilization degree of the neural network operators in the computing units of the graphics processor further comprises:
[0098] The degree of parallelism utilization is calculated using P = α * S / H, 0 < α ≤ 1, where S represents the dimension of operations that the neural network operator needs to perform outside the reduced dimension, and H represents the number of computing units on the graphics processor. During the execution of a GPU program, a larger number of threads than computing units are often required for latency hiding (for example, if thread bundle 1 needs to read data from memory, the computing unit can be released, allowing thread bundle 2 to occupy the computing unit and continue computing). Therefore, α needs to be increased to correct the calculation of parallelism. This parameter is related to the specific hardware platform and the specific task. Therefore, even if P ≥ 1, it is necessary to try to partition.
[0099] Specifically, for the identification of operators with insufficient parallelism, in deep learning, computationally intensive operators refer to those operations that require a large amount of computing resources during model training and inference. These operators are usually the key to the performance bottleneck of deep learning models. The most important operators include matrix multiplication and convolution.
[0100] The most common hardware platform for running deep learning programs is the GPU. GPUs contain multiple cores (called "stream processors" or "shader cores"), each of which can independently execute instructions and process data in parallel. GPUs can execute thousands of threads simultaneously, making them ideal for tasks that can be parallelized. GPUs include several types of memory, such as global memory, shared memory, constant memory, and texture memory. Shared memory is used to share data between processors within the same group and is fast but limited in capacity. Global memory has a larger capacity but is slower to access.
[0101] For compute-intensive operators, when their scale is large in the specified dimension but small in other dimensions, they may encounter insufficient parallelism on the GPU. Only a portion of the hardware's computing units will be used for calculations, while the remaining units will remain idle. In this case, a formula can be used to express the degree of parallelism utilization.
[0102] P = α * S / H, where, in one embodiment, α = 1, S represents the dimensional product of the operation that the operator needs to perform outside the specification dimension, and H represents the computing units available on the hardware. When P < 1, it means that the parallelism of the operator cannot meet the requirements, so the utilization of the computing units can be increased by dividing the specification dimension. When P > = 1, it means that the operator can fully utilize the computing units of the hardware, but it is also necessary to try to divide the operation.
[0103] Take the GEMV operator as an example: Assume that the scale of the GEMV operator is M*K*1, that is, its input is a matrix of shape M*K and a vector of shape K*1. Now we need to multiply these two inputs, specifically reducing them along the dimension of length K, to obtain an output vector of shape M*1. Here, the dimensions that the operator needs to operate on outside the reduced dimension are:
[0104] S=M*1=M
[0105] Taking the NVIDIA GPU A100 architecture as an example, A100 contains a total of 108 streaming multiprocessors, each of which contains 64 CUDA cores. The GPU schedules warps consisting of 32 threads as the basic unit, and each warp is operated by 16 CUDA cores. Therefore, the number of parallel computing units H of A100 can be estimated as:
[0106] H = 108*(64 / 16*32) = 13824
[0107] The degree of parallelism utilization can be estimated as:
[0108] P=S / H=M / 13824
[0109] The scale of the GEMV operator commonly used in large models is 4096*4096*1. At this time, P=8 / 27<1. It can be seen that its parallelism is obviously insufficient.
[0110] like Figure 2 As shown, in one embodiment of the above method of the present invention, the reduction operation division of the neural network operator with insufficient parallelism is performed in the reduction dimension division.
[0111] In one embodiment of the above method of the present invention, the step of performing reduction operation division on the neural network operators with insufficient parallelism in the reduction dimension division further includes:
[0112] The first operator 1 and the second operator 2 of the neural network operator with insufficient parallelism are divided according to a first division quantity T on the reduced dimension.
[0113] In an embodiment of the above method of the present invention, the dividing according to a first division number T is uniform division according to the first division number T.
[0114] In one embodiment of the above method of the present invention, the reduction operation subtask further includes:
[0115] A first reduction operation subtask Re1 is to perform the first number T of reduction operations on the first operator 1 and the second operator 2 after being divided according to the first number T of divisions;
[0116] The second reduction operation subtask Re2 is used to perform a reduction operation of a first partitioning number T dimensions on the result of the first reduction operation subtask Re1.
[0117] In one embodiment of the present invention, after identifying an operator with insufficient parallelism, it is necessary to divide its reduction operation so as to assign subtasks to various computing units. Taking the most typical GEMV operator as an example, the matrix shape is M*K and the vector shape is K*1. Each reduction operation is divided into a first number of T subtasks, so the scale of each reduction operation is reduced to one-T of the original first number of divisions. A reduction operation in each subtask can utilize a computing unit alone (each subtask consists of M reduction operations and will be assigned to different computing units during runtime), thereby improving the utilization of computing units on the hardware.
[0118] It should be noted that the assignment of subtasks to different compute units is guaranteed by the GPU's scheduling policy. The GPU will dispatch subtasks that need to be run to idle compute units. When all compute units are used for computation, the remaining subtasks will wait for the dispatched subtasks to complete. Therefore, not all subtasks will be assigned to the same compute unit.
[0119] Specifically, continuing with the example of the GEMV operator and execution environment, the length of the reduction operation of the GEMV operator is K = 4096. If it is divided into 4 parts, the size of the reduction operation in each subtask is reduced to K' = 4096 / 4 = 1024. At this time:
[0120] P'=4*P=32 / 27
[0121] For performance verification after the reduction operation is divided:
[0122] The final operating performance of an operator is affected by various factors, including data access, thread bundle scheduling, and other factors in addition to the degree of parallelism. Therefore, a higher degree of parallelism does not necessarily mean better final performance. The degree of partitioning cannot be determined by a simple formula.
[0123] Reduction partitioning has similar characteristics to existing scheduling methods such as loop transformation in terms of their effects on operators. The performance of operators cannot be determined by directly derived formulas. Therefore, reduction partitioning can be added to existing scheduling rules for automatic tuning of operators.
[0124] After determining a set of transformation configurations for an operator (including the size of the partitioning), run the transformed operator directly on the corresponding target hardware machine and measure its runtime to determine the impact of this configuration on the operator's performance. A common method is to run the transformed operator 1000 times, measure the total runtime, and then calculate the average.
[0125] Therefore, in actual operation, we will choose to directly divide the calculation task into a smaller one and compare the running performance with the original task. If the effect is better, we can further divide the task.
[0126] like Figure 5 As shown, in one embodiment of the present invention, another method of reducing the operation division of the neural network operator with insufficient parallelism is shown:
[0127] The first operator 1 is divided into a first sub-operator 1' and a third sub-operator 1", and the two second operators 2 correspond to the first sub-operator 1' and the third sub-operator 1", respectively, and are recorded as the second sub-operator 2' and the fourth sub-operator 2".
[0128] The first sub-operator 1' and the second sub-operator 2' are respectively evenly divided on the reduced dimension according to a second division quantity T1;
[0129] The third sub-operator 1 ″ and the fourth sub-operator 2 ″ are evenly divided on the reduced dimension according to a third division quantity T2 .
[0130] The first reduction operation subtask Re1 corresponding to this partitioning method is: performing the second partitioning number T1 reduction operations on the first sub-operator 1' and the second sub-operator 2' after being evenly partitioned according to the second partitioning number T1, and performing the third sub-operator 1" and the fourth sub-operator 2" after being evenly partitioned according to the third partitioning number T2 reduction operations;
[0131] The second reduction operation subtask Re2 is to perform a reduction operation on the result of the first reduction operation subtask Re1 in the dimensions of the second number of divisions T1 and the third number of divisions T2.
[0132] In one embodiment of the above method of the present invention, the subtask fusion of the reduction operation subtask is performed by an operator fusion method or a non-operator fusion method.
[0133] In one embodiment of the above method of the present invention, the step of fusing the subtasks of the reduction operation subtasks in a non-operator fusion manner further includes:
[0134] The operators of the reduction operation subtask are mapped to the computing units of the graphics processor according to the reduction dimension and the non-reduction dimension for execution.
[0135] In one embodiment of the above method of the present invention, the step of fusing the execution reduction operation subtask and the execution subtask by means of operator fusion further includes:
[0136] The operators of the reduction operation subtask are mapped to the computing units of the graphics processor according to the non-reduction dimensions for execution.
[0137] In one embodiment of the above method of the present invention, the neural network operator is a matrix multiplication operator and / or a convolution operator.
[0138] Specifically, in CUDA, a kernel is a function executed on the GPU, which can be understood as a block of code executed in parallel on the GPU. In deep learning operations, an operator is typically mapped to a kernel function for computation. Therefore, the act of assigning all operations to a kernel function can be called operator fusion. The act of performing the first reduction operation subtask Re1 and the second reduction operation subtask Re2 on two separate kernel functions can be seen as dividing the overall operation into two operators.
[0139] like Figure 3 As shown in the figure, in one embodiment of the present invention, when no fusion is performed, both the unreduced and reduced dimensions of the input matrix can be mapped to the GPU's streaming multiprocessor level. That is, the first reduction operation subtask Re1 is first completed in the first kernel function, and then the second reduction operation subtask Re2 is completed in the second kernel function. As can be seen from the figure, during the first reduction operation subtask Re1, the black block portion can be allocated to a GPU streaming multiprocessor.
[0140] like Figure 4 As shown, in one embodiment of the present invention, when performing fusion, the unreduced dimensions of the input matrix need to be mapped to the GPU's streaming multiprocessor level, while the reduced dimensions cannot be mapped to the GPU's streaming multiprocessor level. Therefore, within the same streaming multiprocessor of the GPU, the calculation of the black block portion in the first reduction operation subtask Re1 must be completed first, and then the calculation of the black block portion in the second reduction operation subtask Re2 must be completed.
[0141] The two processing methods, operator fusion or non-operator fusion, have their own advantages and disadvantages, and need to be analyzed according to the actual situation. Specifically, treating the operation as two operators, that is, not performing fusion, can enable the first reduction operation subtask Re1 to obtain a larger search space during automatic tuning and obtain better performance, but it needs to store the intermediate results of the operation in the global memory of the GPU, and at the beginning of the second reduction operation subtask Re2, the data must be taken out of the global memory. Reading data from the global memory of the GPU is relatively time-consuming. Fusing the two operations into one operator, that is, fusing them, will reduce the search space of the first operation during optimization, reduce the performance of the first operation, but can save the very time-consuming operation of reading data from the global memory of the GPU.
[0142] Therefore, the two treatment methods mentioned above need to be analyzed according to the actual situation and the final method should be selected.
[0143] like Figure 6 As shown, the present invention further discloses a neural network operator acceleration device 10, which is used to implement any of the above methods, including:
[0144] an identification module 11 for identifying neural network operators with insufficient parallelism in a computing unit of a graphics processor;
[0145] A reduction operation division module 12 is used to perform reduction operation division on the neural network operator with insufficient parallelism to obtain reduction operation subtasks;
[0146] a subtask execution module 13, configured to distribute the reduction operation subtask to a plurality of computing units of the graphics processor to execute the reduction operation subtask;
[0147] The subtask fusion module 14 is used to perform subtask fusion on the reduction operation subtask to obtain the reduction operation result 3.
[0148] like Figure 7 As shown, the present invention also discloses a graphics processor 100, including a video memory 120 connected to a computing unit 110, a video memory controller 130 and an input / output unit 140, and also includes the neural network operator acceleration device 10 as described above.
[0149] The computing unit 140 may include a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The computing unit may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0150] The present invention discloses a storage medium for storing a computer control program, wherein the computer control program is used to execute the steps of any one of the above methods.
[0151] The computer program that can be executed by the processor can be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the technical field.
[0152] The present invention discloses a neural network operator acceleration method, device, and graphics processor that identifies large-scale reduction operations and divides the identified operators. These divided subtasks are then assigned to multiple computing units for parallel computation. These subtasks are then fused, reducing the memory access overhead of intermediate variables to obtain the final reduction result, thereby improving the performance of the reduction operation.
[0153] This invention enables the partitioning and parallelization of large-scale reduction operations. Compared to the existing TVM, when running on the A100, it can improve the performance of commonly used GEMV operators on large language models by 20% to 35%. For some GEMM operators with smaller N, this invention can also achieve performance improvements similar to GEMV.
[0154] In summary, the present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may evolve various corresponding changes and deformations based on the present invention, but these corresponding changes and deformations should all fall within the scope of protection of the patent applied for the present invention.
Claims
1. A neural network operator acceleration method, characterized in that: The following steps are involved: Identify neural network operators that are insufficiently parallelized in the computational units of a graphics processor; Performing reduction operations on the neural network operators with insufficient parallelism to obtain reduction operation subtasks; Allocating the reduction operation subtask to multiple computing units of the graphics processor to execute the reduction operation subtask; Perform subtask fusion on the reduction operation subtask to obtain the reduction operation result.
2. The method according to claim 1, wherein The identification of neural network operators with insufficient parallelism in the computing units of the graphics processor is obtained by the operation dimensions that the operators need to perform outside the specification dimensions.
3. The method according to claim 2, wherein The step of identifying a neural network operator with insufficient parallelism in a computing unit of a graphics processor, obtained by the operation dimension that the operator needs to perform outside the reduced dimension, further includes: Calculating the degree of parallelism utilization of the neural network operators in the computing units of the graphics processing unit; When the parallelism utilization degree is less than a division threshold, dividing the neural network operator; When the parallelism utilization degree is greater than or equal to the division threshold, an attempt is made to divide the neural network operator.
4. The method according to claim 3, wherein The step of calculating the parallelism utilization degree of the neural network operators in the computing units of the graphics processor further includes: The degree of parallelism utilization is calculated by P = α * S / H, where S represents the operation dimension that the neural network operator needs to perform outside the specification dimension, H represents the number of computing units on the graphics processor, and α is a parallelism correction parameter related to the specific hardware platform and specific task, 0 < α ≤ 1.
5. The method according to claim 1, wherein The reduction operation division of the neural network operators with insufficient parallelism is performed in the reduction dimension division.
6. The method according to claim 5, wherein The step of performing reduction operation division on the neural network operators with insufficient parallelism in the reduction dimension division further includes: The first operator and the second operator of the neural network operator with insufficient parallelism are divided according to a first division quantity on the specification dimension respectively.
7. The method according to claim 6, wherein The dividing according to a first division number is uniformly dividing according to a first division number.
8. The method according to claim 7, wherein The reduction operation subtask further includes: A first reduction operation subtask is to perform the first-partition-number reduction operations on the first operator and the second operator after being divided according to the first partition number; The second reduction operation subtask is used to perform a reduction operation of a first partition quantity dimension on the result of the first reduction operation subtask.
9. The method according to claim 1, wherein The subtask fusion of the reduction operation subtask is performed through an operator fusion method or a non-operator fusion method.
10. The method according to claim 9, wherein The step of fusing the subtasks of the reduction operation subtasks in a non-operator fusion manner further includes: The operators of the reduction operation subtask are mapped to the computing units of the graphics processor according to the reduction dimension and the non-reduction dimension for execution.
11. The method according to claim 9, wherein The step of fusing the execution reduction operation subtask and the execution subtask by means of operator fusion further includes: The operators of the reduction operation subtask are mapped to the computing units of the graphics processor according to the non-reduction dimensions for execution.
12. The method according to any one of claims 1 to 6, characterized in that The neural network operator is a matrix multiplication operator and / or a convolution operator.
13. A neural network operator acceleration device for implementing the method according to any one of claims 1 to 12, characterized in that: include: an identification module for identifying neural network operators with insufficient parallelism in a computational unit of a graphics processing unit; A reduction operation partitioning module is used to perform reduction operation partitioning on the neural network operator with insufficient parallelism to obtain reduction operation subtasks; a subtask execution module, configured to distribute the reduction operation subtask to a plurality of computing units of the graphics processor to execute the reduction operation subtask; The subtask fusion module is used to perform subtask fusion on the reduction operation subtasks to obtain the reduction operation results.
14. A graphics processor comprising a video memory, a video memory controller, and an input / output unit connected to a computing unit, characterized in that: It also includes the neural network operator acceleration device as described in claim 13.
15. A storage medium for storing a computer control program, characterized in that: The computer control program is used to execute the steps of the method according to any one of claims 1 to 12.