Neural network operation acceleration method and device
By adding the maximum value calculation unit to the matrix multiplication array, the maximum value of the matrix multiplication result is calculated in advance, the problems of waste of hardware resources and the impact of accuracy in the prior art are solved, and efficient acceleration of neural network operations is achieved.
Patent Information
- Application Number
- CN202510199254.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-08
AI Technical Summary
The schemes used to accelerate BERT neural network computing in the prior art have problems with wasting hardware resources and model accuracy, especially in softmax computing, which cannot be integrated with other computing in on-chip memory, and data distribution needs to be analyzed before model deployment.
The maximum value calculation unit is added to the nonlinear calculation unit of the matrix multiplication array, calculate the maximum value of the matrix multiplication result in advance, reduce the data loading overhead in the softmax calculation, and perform softmax subsequent calculation in the decoder-based large-scale model network.
It effectively reduces the data loading overhead in softmax calculations, while maintaining the accuracy of the original algorithm, avoiding the waste of hardware resources and the impact of accuracy.
Smart Images

Figure CN120277309A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural networks and deep learning, and in particular to an acceleration method and device for neural network operations. Background Art
[0002] After retrieval, a Chinese patent application with the publication number CN111062471A discloses a deep learning accelerator for accelerating BERT neural network operations, including three matrix multiplication arrays, a softmax and dot product calculation unit, three feature memories, two weight memories, a controller, and an on-chip and off-chip interface. This solution optimizes the branch network structure in the neural network, effectively reducing the storage space of intermediate data, reducing the number of on-chip and off-chip data interactions, and reducing power consumption; at the same time, by configuring the reconfigurable data interconnection between the storage unit and the calculation unit, it meets the calculation requirements of the branch network structure in the BERT neural network and can be used for end-to-end neural network calculations. However, this solution has two disadvantages: (1) Using a dedicated softmax calculation unit, other calculations cannot reuse this calculation unit, resulting in waste of hardware resources; (2) Although the calculation speed of softmax is accelerated, its output cannot be fused with other calculations in on-chip memory.
[0003] In addition, the prior art also uses the data distribution of the statistical model to estimate the maximum value of each dimension in a softmax input. Similarly, this solution also has certain disadvantages: (1) It is necessary to analyze the data distribution of different models before model deployment to find a suitable maximum value; (2) In this solution, the maximum value is an empirical value and is inconsistent with the actual maximum value, which will affect the overall accuracy of the model. Summary of the Invention
[0004] Aiming at the technical defects mentioned in the background art, the purpose of the embodiments of the present invention is to provide an acceleration method and device for neural network operations.
[0005] To achieve the above object, in the first aspect, the embodiments of the present invention provide an acceleration method for neural network operations, which is applicable to a large model network based on a decoder; the acceleration method includes:
[0006] Obtain the first output result of the matrix multiplication array; the first output result includes a plurality of matrix multiplication results;
[0007] Obtain a second output result from the plurality of matrix multiplication results; wherein, the second output result is the maximum value of each row of the matrix multiplication results;
[0008] Perform subsequent softmax calculations on the first output result and the second output result, and output the target calculation result.
[0009] As a preferred implementation manner of the present application, before obtaining the first output result of the matrix multiplication array, the method further includes:
[0010] Adding a maximum value calculation unit; the maximum value calculation unit is used to select the maximum value of each row of the matrix multiplication result.
[0011] In a second aspect, an embodiment of the present invention provides an acceleration device for neural network operations, including:
[0012] A first unit, configured to obtain a first output result of a matrix multiplication array; the first output result includes a plurality of matrix multiplication results;
[0013] A second unit, configured to select the maximum value of each row from the plurality of matrix multiplication results as a second output result;
[0014] A third unit, configured to perform subsequent softmax calculations on the first output result and the second output result, and output a target calculation result.
[0015] In a third aspect, an embodiment of the present invention further provides another acceleration device for neural network operations, including a processor, an input device, an output device, and a memory. The processor, the input device, the output device, and the memory are interconnected. Wherein, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method of the first aspect above.
[0016] In a fourth aspect, an embodiment of the present invention further provides another acceleration device for neural network operations, including:
[0017] A matrix multiplication array, configured to perform multiplication calculation on a matrix to obtain a plurality of matrix multiplication results;
[0018] A maximum value calculation unit, configured to select the maximum value of each row from the plurality of matrix multiplication results;
[0019] A vector accelerator, configured to receive the matrix multiplication result and the maximum value of each row, perform subsequent softmax calculations, and output a target calculation result.
[0020] As an implementation manner, the maximum value calculation unit is integrated in the matrix multiplication array.
[0021] By implementing the embodiment of the present invention, by adding a maximum value calculation unit to the non-linear calculation unit of the matrix multiplication array, the maximum value calculation originally in the softmax calculation is advanced to the non-linear calculation unit in the matrix multiplication calculation array. Compared with the existing solution, the overhead of one data loading in the subsequent softmax calculation is reduced, and the accuracy of the original algorithm is not changed. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art.
[0023] Figure 1 and Figure 2 is the block diagram of the acceleration calculation principle of neural network operations provided by the embodiments of the present invention;
[0024] Figure 3 is the flowchart of the acceleration method for neural network operations provided by the embodiments of the present invention;
[0025] Figure 4 is the structural diagram of the acceleration device for neural network operations provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0028] Please refer to Figure 1 and Figure 2 , the inventive concept of the present invention is:
[0029] In the decoder-based large model network, attention is an indispensable operator. An attention operator is as Figure 2 shown, mainly composed of two matrix multiplications and one softmax. Among them, the shape of the Q matrix is (M, N), M is the length of the sequence, and N is the length of the word vector. When the first multiplication of Q and the transpose of K generates a matrix in the form of (M, M) as the input of softmax, where M is a variable that changes according to the length of the sequence.
[0030]
[0031] As the length of the sequence increases and is limited by the size of the vector accelerator's memory space, there will be M data that cannot be stored in the vector accelerator's memory space simultaneously. As can be seen from the softmax formula, each step of the entire calculation depends on the result of the max calculation. And since the amount of data is greater than the overall storage space of the vector accelerator, all the data needs to be loaded in batches once to obtain its maximum value.
[0032] In this embodiment, a maximum value extraction functional unit is added behind the conventional matrix array. After the data passes through the matrix multiplication array, the result of the matrix multiplication and the maximum value of each row of the matrix multiplication result are output.
[0033] Please refer to Figure 3 , the acceleration method for neural network operations provided by the embodiments of the present invention is applicable to large model networks based on decoders, and includes the following steps:
[0034] S1, obtain the first output result of the matrix multiplication array.
[0035] Specifically, in implementation, as described above, by calculating two matrices Q and K through the matrix multiplication array, multiple matrix multiplication results are obtained, that is, the first output result.
[0036] S2, obtain the second output result from the multiple matrix multiplication results.
[0037] Specifically, in implementation, a maximum value calculation unit is integrated in the matrix multiplication array. By this maximum value calculation unit, the maximum value of each row of the matrix multiplication result is selected and used as the second output result.
[0038] S3, perform subsequent softmax calculations on the first output result and the second output result, and output the target calculation result.
[0039] Combined with Figure 2 , taking the first output result (matrix multiplication result) and the second output result (the maximum value of each row of the matrix multiplication result) as inputs, and inputting them into the vector accelerator for subsequent softmax calculations, and finally outputting the final result of the attention calculation.
[0040] From the above description, it can be known that by implementing the acceleration calculation method of the embodiments of the present invention, by adding a maximum value calculation unit in the non-linear calculation unit of the matrix multiplication array, the maximum value calculation in the original softmax calculation is advanced to the non-linear calculation unit in the matrix multiplication calculation array. Compared with the existing solutions, the overhead of one data loading in the subsequent softmax calculation is reduced, and the accuracy of the original algorithm is not changed.
[0041] Based on the same inventive concept, an embodiment of the present invention provides an acceleration device for neural network operations, which is applicable to large model networks based on decoders and includes:
[0042] A first unit for obtaining a first output result of a matrix multiplication array; the first output result includes a plurality of matrix multiplication results;
[0043] A second unit for selecting the maximum value of each row from the plurality of matrix multiplication results as a second output result;
[0044] A third unit for performing subsequent softmax calculations on the first output result and the second output result and outputting a target calculation result.
[0045] Optionally, as Figure 4 shown, an embodiment of the present invention also provides another acceleration device for neural network operations, which may include: one or more processors 101, one or more input devices 102, one or more output devices 103, and a memory 104. The above-mentioned processors 101, input devices 102, output devices 103, and memory 104 are interconnected through a bus 105. The memory 104 is used to store a computer program, the computer program includes program instructions, and the processor 101 is configured to call the program instructions to execute the methods in the method embodiment part above.
[0046] It should be understood that in the embodiment of the present invention, the so-called processor 101 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0047] The input device 102 may include a keyboard, etc., and the output device 103 may include a display (such as an LCD), a speaker, etc.
[0048] The memory 104 may include a read-only memory and a random access memory, and provide instructions and data to the processor 101. A part of the memory 104 may also include a non-volatile random access memory. For example, the memory 104 may also store information about the device type.
[0049] In a specific implementation, the processor 101, input device 102, and output device 103 described in the embodiments of the present invention may execute the implementation manners described in the embodiments of the acceleration method for neural network operations provided by the embodiments of the present invention, which will not be elaborated herein.
[0050] Optionally, please refer to Figure 2 , the present invention further provides an acceleration device for neural network operations, including:
[0051] A matrix multiplication array for performing multiplication calculations on matrices to obtain a plurality of matrix multiplication results;
[0052] A maximum value calculation unit for selecting the maximum value of each row from the plurality of matrix multiplication results;
[0053] A vector accelerator for receiving the matrix multiplication results and the maximum value of each row, performing subsequent softmax calculations, and outputting a target calculation result.
[0054] Among them, the maximum value calculation unit is integrated within the matrix multiplication array.
[0055] It should be noted that for a more specific working process of the acceleration calculation device, please refer to the foregoing method embodiment part, which will not be elaborated herein.
[0056] As described above, the above are only specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. An acceleration method for neural network operations, characterized in that The acceleration method is applicable to a large model network based on a decoder; the acceleration method includes: Obtain a first output result of a matrix multiplication array; the first output result includes a plurality of matrix multiplication results; Obtain a second output result from the plurality of matrix multiplication results; Perform subsequent softmax calculations on the first output result and the second output result, and output a target calculation result.
2. The acceleration method according to claim 1, wherein The second output result is the maximum value of each row of the matrix multiplication results.
3. The acceleration method according to claim 1, wherein Before obtaining the first output result of the matrix multiplication array, the method further includes: Add a maximum value calculation unit; the maximum value calculation unit is used to select the maximum value of each row of the matrix multiplication results.
4. An acceleration device for neural network operations, characterized in that, Includes: A first unit for obtaining a first output result of a matrix multiplication array; The first output result includes a plurality of matrix multiplication results; A second unit for selecting the maximum value of each row from the plurality of matrix multiplication results as the second output result; A third unit for performing subsequent softmax calculations on the first output result and the second output result, and outputting a target calculation result.
5. The acceleration device according to claim 4, characterized in that The acceleration device is applicable to a large model network based on a decoder.
6. An acceleration device for neural network operations, characterized in that, Includes a processor, an input device, an output device, and a memory, the processor, the input device, the output device, and the memory are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1-3.
7. An acceleration device for neural network operations, characterized in that, Includes: A matrix multiplication array for performing matrix multiplication calculations to obtain a plurality of matrix multiplication results; A maximum value calculation unit for selecting the maximum value of each row from the plurality of matrix multiplication results; A vector accelerator for receiving the matrix multiplication results and the maximum value of each row, performing subsequent softmax calculations, and outputting a target calculation result.
8. The acceleration device according to claim 7, wherein, The maximum value calculation unit is integrated within the matrix multiplication array.
9. The acceleration device according to claim 7 or 8, characterized in that The acceleration device is applicable to a large model network based on a decoder.
Citation Information
Patent Citations
Deep learning accelerator for accelerating BERT neural network operation
CN111062471A
Cited By
Tensor core component of processor and processor
CN121029685A