A Transformer Accelerator Based on Offset Diagonal Matrix
By adopting an accelerator based on offset diagonal matrix in the Transformer model, the problem of low computing and storage efficiency in edge-end deployment is solved, efficient data multiplexing and load balancing is achieved, and overall power consumption and area are reduced.
Patent Information
- Application Number
- CN202210839287.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-18
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-07-18
AI Technical Summary
When the existing Transformer model is deployed at the edge end, due to the large amount of model parameters and calculations, it is difficult to efficiently deploy in environments with limited resources and power. The structured pruning method has low data reuse rate and large index overhead in local sparseness.
Using a Transformer accelerator based on an offset diagonal matrix, the non-zero value and offset of each sub-matrix are divided into sub-matrixes of the same size are obtained, and matrix multiplication and addition operations are performed in the operation array, and data allocation is used to achieve efficient data multiplexing and load balancing.
Improves data multiplexing, reduces index overhead, reduces non-zero value movement and storage overhead, reduces overall power consumption and area, and meets resource and power limitations for edge deployment.
Smart Images

Figure CN115329260B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of integrated circuits, and more particularly to a Transformer accelerator based on an offset diagonal matrix. Background Art
[0002] Natural language processing technology gives machines the ability to interact with people using human language. With the rapid development of deep learning, various models for natural language processing have emerged one after another. Recently, the Transformer model based on the self-attention mechanism and its variants have shown superior performance in various natural language processing tasks, far exceeding traditional models based on convolutional neural networks and recurrent neural networks. Therefore, deploying the Transformer model to the edge has aroused increasing interest.
[0003] However, the price of excellent performance is the rapid growth of model parameters and computation. The large amount of computation and memory access during model inference limits the deployment of the complete Transformer model on the edge with limited power and resources. Recent studies have shown that the Transformer model has considerable redundancy. Therefore, many model compression methods such as quantization and pruning have been proposed to compress the model. In order to accelerate the deployment of the Transformer model on the edge with limited resources and power, it is necessary to first compress the Transformer model and then design a dedicated hardware accelerator for the calculation and storage of the compressed model.
[0004] Among them, pruning is the main method of model compression, which can be roughly divided into structured pruning and unstructured pruning. At present, the pruning of Transformer models is mostly unstructured pruning. Although the irregular sparse weights after pruning are more sparse than those after structured pruning, their irregularity leads to low storage and computing efficiency when mapped to hardware, and the randomness of non-zero values also leads to low data reuse efficiency. Therefore, researchers have devoted their efforts to more regular structured pruning, such as using rows, columns or blocks as pruning units. The pruned model weights are structural as a whole. When mapped to the operation array and storage unit, their structured form makes the load of the operation array relatively balanced, and the data reuse efficiency is higher than that of the matrix obtained by unstructured pruning. However, this structural pruning only focuses on the overall structure of the weight distribution, and the distribution of local non-zero values is uneven, which increases additional index overhead when mapped to the operation array and storage unit. The local uneven distribution makes data reuse relatively difficult, the data movement overhead is large, and the structure of the data is not fully utilized. Therefore, there is an urgent need for a matrix format that can balance the overall and local sparsity of weights and to design a Transformer accelerator that can fully utilize this structure to reduce the mobile storage and indexing overhead of non-zero values and reduce the overall power consumption and area of the chip to meet the resource and power constraints of deploying Transformer models at the edge. Summary of the invention
[0005] In order to overcome the shortcomings and deficiencies in the prior art, the purpose of the present invention is to provide a Transformer accelerator based on an offset diagonal matrix; the accelerator can meet the requirements of Transformer model acceleration based on an offset diagonal structured sparse matrix, with high data reuse rate, load balancing and low index overhead.
[0006] In order to achieve the above object, the present invention is implemented by the following technical scheme: a Transformer accelerator based on an offset diagonal matrix, characterized in that it includes a top-level control module, an operation array, a nonlinear function unit and an on-chip cache module;
[0007] The on-chip cache module is used to store input data, a weight matrix, intermediate operation results and an output matrix; the weight matrix is stored in the on-chip cache module in an offset diagonal matrix manner; the offset diagonal matrix includes non-zero values and offsets; the non-zero values and offsets are obtained by dividing the weight matrix into a plurality of sub-matrices of the same size, obtaining non-zero values of each sub-matrix, and obtaining the offset of each sub-matrix from the degree of offset of each sub-matrix from the diagonal line;
[0008] The operation array is used for the operation array to read input data and weight matrix from the on-chip cache module to perform matrix multiplication and addition operations; when the operation array performs matrix multiplication and addition operations, the operation array simultaneously reads the non-zero values of the offset diagonal matrix and the offset to perform operation distribution on the non-zero values according to the offset;
[0009] The nonlinear function unit is used to perform nonlinear function calculation on the output matrix; after the matrix multiplication and addition operation is completed, the operation result is written back to the on-chip cache, or the operation result is first input into the nonlinear function for processing and then written back to the on-chip cache.
[0010] Preferably, the method for performing matrix multiplication and addition operations on the operation array is: in the operation array, matrix multiplication and addition operations are performed on the input data and the weight matrix cycle by cycle, and the intermediate results are transmitted to the next cycle to achieve accumulation; the input data is input horizontally from the operation array and moved to the right cycle by cycle; the weight matrix is input vertically from the operation array and transmitted downward cycle by cycle, thereby achieving input and weight reuse between arrays.
[0011] Preferably, the operation array is composed of a plurality of operation units; each operation unit includes a plurality of multipliers and adders, a data distributor, a unit output buffer, a quantization module, a Relu function module and a control module; the number of multipliers and adders in each operation unit is the same as the size of the submatrix; each multiplier and adder is composed of a connected multiplier and an adder;
[0012] The data distributor distributes the non-zero values of the offset diagonal matrix of the input operation unit to the corresponding multipliers and adders according to the offset size, and distributes them to the multipliers in a multicast form to achieve input multiplexing; the multipliers and adders are responsible for multiplication and addition operations; the unit output cache is responsible for storing the multiplication and addition results;
[0013] The control module is responsible for controlling the output data path and the calculation mode signal; according to the calculation mode signal output by the control module, the unit output cache is directly connected to the quantization module to directly quantize the multiplication and addition operation results and then output, or the unit output cache is connected to the quantization module through the Relu function module to quantize the multiplication and addition operation results after the nonlinear activation of the Relu function module and then output; the output results of each operation unit are accumulated to obtain the output of the operation array.
[0014] Preferably, the operation array is a reconfigurable operation matrix; in the reconfigurable operation matrix, the arrangement of each operation unit is set according to the length of input data.
[0015] Preferably, in the reconfigurable operation matrix, there are 64 operation units; according to the length of the input array, the arrangement of the operation units is set to any one of 2x32, 4x16, and 8x8.
[0016] Preferably, the on-chip cache module includes an input cache, an output cache, a weight cache, an intermediate result cache and an output cache; the input cache is used to store input data; the weight cache is used to store a weight matrix; the intermediate result cache is used to cache intermediate results of calculations; and the output cache is used to store a final output matrix.
[0017] Preferably, the nonlinear function unit includes a Layernorm calculation module and a Softmax calculation module; the Layernorm calculation module is responsible for calculating the layer normalized output of each sublayer of the Transformer model, and the Softmax calculation module is responsible for calculating the attention score value of the multiplication of the query matrix Q and the key matrix K of the multi-head attention sublayer in the Transformer model.
[0018] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0019] (1) The present invention focuses on matrix multiplication, which accounts for the vast majority of calculations in the Transformer model. In view of the problems of load imbalance and low computational efficiency caused by random reading, writing and calculation caused by ordinary sparse matrices, as well as the relatively low data reuse rate caused by the uneven local sparsity of current structured matrices, the present invention proposes the use of a Transformer model based on offset diagonal structured sparseness, which has the characteristics of high data reuse rate, load balance and low index overhead.
[0020] (2) The present invention proposes a reconfigurable operation array, which uses the input data length as a reference to reconstruct the array arrangement mapping of each layer of matrix operation, which can minimize the problem of operation unit pauses caused by changes in input data length and different model codec calculation modes, and improve the utilization rate of operation units. At the same time, input and weights are transmitted between arrays in both horizontal and vertical directions, which improves the multiplexing of input and weights and reduces data movement power consumption.
[0021] (3) According to the weight distribution characteristics of the designed offset diagonal matrix, the present invention designs a computing unit block with the number of multipliers and adders consistent with the size of the sub-matrix, distributes the computational load of the input weights and inputs according to the input offset, and implements accumulation within the computing unit, thereby improving input and output multiplexing, reducing the index overhead of the sparse matrix, and reducing overall power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the overall architecture diagram of the Transformer accelerator of the present invention;
[0023] Figure 2 is a diagram of the computing array architecture of the Transformer accelerator of the present invention;
[0024] Figure 3Schematic diagram of the computing unit of the Transformer accelerator of the present invention.
[0025] Figure 4 is a schematic diagram of an example of an offset diagonal matrix in an embodiment;
[0026] Figure 5 is the storage method of the offset diagonal matrix in the embodiment;
[0027] Figure 6 is the data transmission mode of the operation array in the embodiment;
[0028] Figure 7 This is the data allocation method of the computing unit in the embodiment. DETAILED DESCRIPTION
[0029] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0030] Example
[0031] Compared with the existing structured sparse matrix, this embodiment proposes an offset diagonal sparse matrix with more balanced non-zero values of model weights both overall and locally as the weight matrix form, and designs a dedicated hardware acceleration architecture based on the proposed offset diagonal sparse matrix data distribution characteristics to improve data reuse rate and reduce index overhead. The accelerator can meet the acceleration of the Transformer model based on the offset diagonal structured sparse matrix, can reuse data from multiple angles of input, weight and output, and has the characteristics of low power consumption of data movement and low power consumption overhead of index matching area.
[0032] The present embodiment provides a Transformer accelerator based on an offset diagonal matrix, including a top-level control module, a computing array, a nonlinear function unit, and an on-chip cache module, such as Figure 1 When working, input data is read into the on-chip cache module, and the on-chip cache module stores a weight matrix; the weight matrix is stored in the on-chip cache module in the form of an offset diagonal matrix; the offset diagonal matrix includes non-zero values and offsets; the non-zero values and offsets are obtained by dividing the weight matrix into multiple sub-matrices of the same size, obtaining the non-zero values of each sub-matrix, and obtaining the offset of each sub-matrix from the degree of offset of each sub-matrix from the diagonal line;
[0033] The operation array reads input data and a weight matrix from an on-chip cache module to perform a matrix multiplication and addition operation; when the operation array performs the matrix multiplication and addition operation, the operation array simultaneously reads the non-zero value of the offset diagonal matrix and the offset to perform operation allocation on the non-zero value according to the offset;
[0034] After the matrix multiplication and addition operation is completed, the operation result is written back to the on-chip cache, or the operation result is first input into the nonlinear function for processing and then written back to the on-chip cache.
[0035] Specifically, the top-level control module controls the data flow and computing method of the entire accelerator, allocates computing load and storage strategy to the different computing methods of the Transformer model encoder and decoder, and controls the data that needs to be pre-read in the current layer and the intermediate results of the current operation array calculation.
[0036] The on-chip cache module includes an input cache, an output cache, a weight cache, an intermediate result cache and an output cache; the input cache is used to store input data; the weight cache is used to store weight matrices; the intermediate result cache is used to cache intermediate results; and the output cache is used to store the final output matrix. The offset diagonal matrix stores non-zero values in a dense form, and the offsets are stored separately in corresponding positions according to the storage order of the non-zero values.
[0037] The method for performing matrix multiplication and addition operations on the operation array is as follows: in the operation array, matrix multiplication and addition operations are performed on input data and weight matrices period by period, and the intermediate results obtained are transmitted to the next period to achieve accumulation; input data is input horizontally from the operation array and moved to the right period by period; weight matrices are input vertically from the operation array and transmitted downward period by period, thereby achieving input and weight reuse between arrays.
[0038] The accelerator of the present invention uses a structured sparse matrix weight model based on an offset diagonal, divides the entire weight matrix into multiple sub-matrices of the same size, each sub-matrix has a specific offset, and has local and overall sparse uniformity. During storage and calculation, data is allocated according to the offset of the offset diagonal matrix, so that the sparse model has the reuse space of a dense matrix and uniform computing load distribution.
[0039] The operation array is composed of several operation units; it is responsible for the operation of sparse and dense matrices in the Transformer model, accounting for more than 95% of the computational load in the entire model operation. The intermediate results are output to the on-chip cache module or nonlinear function unit for the next step of calculation.
[0040] Each operation unit includes several multipliers, data distributors, unit output buffers, quantization modules, ReLU function modules and control modules, such as Figure 3 As shown in FIG. 1 , the number of multipliers and adders in each operation unit is the same as the size of the submatrix; each multiplier and adder is composed of a multiplier and an adder connected to each other. In this embodiment, one operation unit has 16 multipliers and 16 adders.
[0041] The operation array is preferably a reconfigurable operation matrix, such as Figure 2 As shown; in the reconfigurable operation matrix, the arrangement of each operation unit is set according to the length of the input data (such as a sentence).
[0042] Based on the offset diagonal structured sparse matrix, a reconfigurable operation array is used, which consists of operation units of multipliers and adders with the same size as the submatrix. The data flow uses fixed output, and the input data and weight matrix are transmitted in a pulsating form. In the operation unit, when all parts are accumulated, the results are written back to the on-chip cache module or nonlinear function unit for the next calculation.
[0043] The preferred solution is: in the operation array, there are 64 operation units; according to the length of the input array, the arrangement of the operation units is set to any one of 2x32, 4x16, and 8x8. Therefore, when the length of the input array changes, the arrangement with the highest utilization rate of the reconfigurable operation array can be selected to minimize the pause of the operation unit.
[0044] The data distributor distributes the non-zero values of the offset diagonal matrix of the input operation unit to the corresponding multipliers and adders according to the offset size, and distributes them to the multipliers in a multicast form to achieve input multiplexing; the multipliers and adders are responsible for multiplication and addition operations; the unit output cache is responsible for storing the multiplication and addition results;
[0045] The control module is responsible for controlling the output data path and the calculation mode signal; according to the calculation mode signal output by the control module, the unit output buffer is directly connected to the quantization module to directly quantize the multiplication and addition operation result and then output it, or the unit output buffer is connected to the quantization module through the Relu function module to quantize the multiplication and addition operation result after the nonlinear activation of the Relu function module and then output it; the output results of each operation unit are accumulated to obtain the output of the operation array.
[0046] The nonlinear function unit includes a Layernorm calculation module and a Softmax calculation module; the Layernorm calculation module is responsible for calculating the layer normalization output of each sublayer of the Transformer model, and the Softmax calculation module is responsible for calculating the attention score value of the multiplication of the query matrix Q and the key matrix K of the multi-head attention sublayer in the Transformer model.
[0047] The present invention is based on a more structurally stronger offset diagonal sparse matrix, and adopts a reconfigurable operation array composed of operation units with the same size as the sub-matrix block. The input and weight data are transmitted and reused between the operation arrays in each cycle, which reduces the bandwidth demand and the power consumption of data movement. The reconfigurable operation array design makes the operation array still have extremely high utilization when the input sentence length and the calculation mode of different layers change. The data is allocated by offset in the operation unit to efficiently reuse the input and output. The present invention performs efficient hardware mapping on the Transformer model based on the offset diagonal matrix, alleviating the energy consumption and resource limitations of the Transformer model deployed at the edge.
[0048] The following is an example of the calculation of the multi-head attention sub-layer of the encoder in the Transformer model, which includes the following steps:
[0049] S1, read the weight matrix W from off-chip DRAM Q , weight matrix W K , weight matrix W V and bias B Q , B K , B V Load it into the weight cache of the on-chip cache module, and read the input data X into the input cache of the on-chip cache module. Read the weight matrix and input matrix of the on-chip cache module into the operation array to calculate the matrix Q, matrix K and matrix V, and store the matrix Q, matrix K and matrix V into the intermediate result cache of the on-chip cache module:
[0050] Q=W Q X+B Q
[0051] K=W K X+B K
[0052] V=W V X+B V
[0053] Read the weight matrix W from off-chip DRAM O and bias B O into the weight cache, and read the matrix Q and matrix K in the intermediate result cache into the operation array. After matrix multiplication, the result is input into the Softmax calculation module in the nonlinear calculation unit to calculate the attention score matrix S, and the result is stored back in the intermediate result cache:
[0054] P=QK T
[0055]
[0056] Among them, d k The number of columns of matrix Q and K corresponds to the dimension of word vector.
[0057] Read the matrix V and the attention score matrix S in the intermediate result cache into the operation array to calculate the output matrix Z i , and then the output matrix Z of each head of the multi-head attention module i After splicing, calculate the final output matrix Z of the multi-head attention module:
[0058] Z i =SV
[0059] Z=WO (Z 1 …Z i )+B o
[0060] Write the final output matrix Z back to the intermediate result buffer.
[0061] The weight matrices to be calculated are all offset diagonal matrices. The entire weight matrix is divided into multiple sub-matrices of a certain size. Each sub-matrix has a specific offset and has local and overall sparse uniformity. During storage and calculation, data is allocated according to the offset of the offset diagonal matrix, so that the offset diagonal sparse matrix has the reuse space of a dense matrix and uniform calculation load distribution. Figure 4 is a graphical representation of an offset diagonal matrix, Figure 5 The offset diagonal matrices are stored in dense form with the corresponding offsets in the same positions of different storage units.
[0062] The offset diagonal matrix includes non-zero values and offsets, and divides the entire weight matrix into multiple sub-matrices. Each sub-matrix has its own unique offset, and the offset indicates the degree of deviation of the sub-matrix from the diagonal.
[0063] The offset diagonal matrix needs to read non-zero values and offsets simultaneously when calculating the operation array, and the input is distributed to the corresponding calculation load by the data distributor according to the offset in each operation unit.
[0064] Based on the calculation of the offset diagonal structured matrix, efficient mapping is performed in a reconfigurable operation array composed of operation units with the same number of multipliers and adders as the sub-matrix size. The data flow uses an operation array with fixed output parts and inputs and weights transmitted in a pulsating form. After waiting for all parts and accumulations of the multiplication and addition blocks, the results of the operation units are written back to the on-chip cache or nonlinear function for the next calculation. The calculation data flow method is as follows Figure 6 shown.
[0065] The operation array consists of 64 operation units. According to the change of the input sentence length, the arrangement of the operation units is reconstructed into 2x32, 4x16, and 8x8 operation array modes, respectively. In this way, when the sentence length changes, the arrangement mode with the highest utilization rate of the operation array can be selected to minimize the pause of the operation unit.
[0066] The operation array uses a fixed output data stream to perform matrix multiplication in the form of inner product. The input matrix is input horizontally from the operation array and moves to the right cycle by cycle. The weight matrix is input vertically from the operation array and transmitted downward cycle by cycle, so that input and weight reuse can be achieved between operation arrays.
[0067] The operation of the weight matrix and the input is divided into multiple sub-matrices and calculated in the operation unit. The multiplier and adder of the operation unit are consistent with the size of the sub-matrix, including data distributor, multiplier, adder, quantization module and Relu function module, etc. The input is reused in part and in the operation unit.
[0068] The data distributor is responsible for distributing the non-zero values of the offset diagonal matrix of the input operation unit to the corresponding multipliers and adders according to the offset size, and distributes the input to the multiplier in the form of multicast to achieve input multiplexing, and accumulates the partial sums in the adding unit connected to the multiplier to achieve output multiplexing.
[0069] The quantization module and the Relu function module determine whether the multiplication and addition calculation results are directly quantized and output according to the calculation mode signal output by the control module, or quantized and output after nonlinear activation of the Relu function. The control module is responsible for controlling the output data path and the calculation mode, such as Figure 7 shown.
[0070] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A Transformer accelerator based on an offset diagonal matrix, characterized in that: It includes a top-level control module, an on-chip cache module, an operation array, and a nonlinear function unit; The on-chip cache module is used to store input data, weight matrix, intermediate operation results and output matrix; the weight matrix is stored in the on-chip cache module in the form of offset diagonal matrix; The offset diagonal matrix includes non-zero values and offsets; the non-zero values and offsets are obtained by dividing the weight matrix into a plurality of sub-matrices of the same size, obtaining the non-zero values of each sub-matrix, and obtaining the offset of each sub-matrix from the degree of offset of each sub-matrix from the diagonal line; The operation array is used to read input data and a weight matrix from an on-chip cache module to perform matrix multiplication and addition operations; when the operation array performs matrix multiplication and addition operations, the operation array simultaneously reads non-zero values of the offset diagonal matrix and an offset to perform operation allocation on the non-zero values according to the offset; The nonlinear function unit is used to perform nonlinear function calculation on the output matrix; after the matrix multiplication and addition operation is completed, the operation result is written back to the on-chip cache, or the operation result is first input into the nonlinear function for processing and then written back to the on-chip cache; The operation array is composed of a number of operation units; each operation unit includes a number of multipliers and adders, a data distributor, a unit output buffer, a quantization module, a Relu function module and a control module; the number of multipliers and adders in each operation unit is the same as the size of the submatrix; each multiplier and adder is composed of a connected multiplier and adder; The data distributor distributes the non-zero values of the offset diagonal matrix of the input operation unit to the corresponding multipliers and adders according to the offset size, and distributes them to the multipliers in a multicast form; the multipliers and adders are responsible for performing multiplication and addition operations; the unit output cache is responsible for storing the multiplication and addition operation results; The control module is responsible for controlling the output data path and the calculation mode signal; according to the calculation mode signal output by the control module, the unit output cache is directly connected to the quantization module to directly quantize the multiplication and addition operation results and then output, or the unit output cache is connected to the quantization module through the Relu function module to quantize the multiplication and addition operation results after the nonlinear activation of the Relu function module and then output; the output results of each operation unit are accumulated to obtain the output of the operation array.
2. The Transformer accelerator based on offset diagonal matrix according to claim 1, characterized in that: The method for performing matrix multiplication and addition operations in the operation array is: in the operation array, matrix multiplication and addition operations are performed on the input data and the weight matrix cycle by cycle, and the intermediate results are transmitted to the next cycle to achieve accumulation; the input data is input horizontally from the operation array and moved to the right cycle by cycle; the weight matrix is input vertically from the operation array and transmitted downward cycle by cycle.
3. The Transformer accelerator based on offset diagonal matrix according to claim 1, characterized in that: The operation array is a reconfigurable operation matrix; in the reconfigurable operation matrix, the arrangement of each operation unit is set according to the length of input data.
4. The Transformer accelerator based on offset diagonal matrix according to claim 3, characterized in that: In the reconfigurable operation matrix, there are 64 operation units; according to the length of the input array, the arrangement of the operation units is set to any one of 2x32, 4x16, and 8x8.
5. The Transformer accelerator based on offset diagonal matrix according to claim 1, characterized in that: The on-chip cache module includes an input cache, an output cache, a weight cache, an intermediate result cache and an output cache; the input cache is used to store input data; the weight cache is used to store a weight matrix; the intermediate result cache is used to cache intermediate results; The output buffer is used to store the final output matrix.
6. The Transformer accelerator based on offset diagonal matrix according to claim 1, characterized in that: The nonlinear function unit includes a Layernorm calculation module and a Softmax calculation module; the Layernorm calculation module is responsible for calculating the layer normalization output of each sublayer of the Transformer model, and the Softmax calculation module is responsible for calculating the attention score value of the multiplication of the query matrix Q and the key matrix K of the multi-head attention sublayer in the Transformer model.
Citation Information
Patent Citations
Method and device for realizing sparse matrix multiplication on reconfigurable processor array
CN112507284A
Low-power-consumption neural network accelerator storage architecture based on NAND flash memory
CN113159309A