A Low Bitwidth Quantization Compression LSTM Accelerator
Through technologies such as low bit width quantization and matrix slicing, the LSTM model is optimized, which solves the power consumption and inference delay problems of the LSTM model on the FPGA accelerator in the prior art, and realizes an efficient and low-power LSTM accelerator.
Patent Information
- Application Number
- CN202211473669.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-22
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-11-22
AI Technical Summary
In the prior art, when designing FPGA accelerators for recurrent neural network LSTMs, it is difficult to reduce power consumption and inference delay time while ensuring model accuracy.
The low-bit width quantization compression technology is used to quantify the weight and activation value of the LSTM model, quantize it from 32-bit floating point to 8-bit integer, and optimize the matrix vector multiplication and activation function calculation through matrix slicing and block loop compression technology.
While ensuring model accuracy, the power consumption and inference delay time are significantly reduced, and the implementation efficiency of the LSTM model on FPGA is improved.
Smart Images

Figure CN115730648B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of neural network hardware acceleration, and particularly relates to an LSTM accelerator with low-bitwidth quantization compression. Background Art
[0002] With the development of science and technology, deep neural networks have been widely applied in fields such as image recognition, speech recognition, natural language processing, and machine translation. By adding a loop structure to the hidden layer, the RNN enables the neural network model to have "memory". The network structure of the RNN determines its natural advantage in processing sequential information, so it is often used in the training and inference of neural network models with sequential information such as audio and text as inputs. However, the RNN has problems of gradient disappearance and gradient explosion, making it impossible for the neural network to remember information for a long time. The LSTM modifies the network structure of the traditional RNN, enabling the neural network to remember long-term information.
[0003] The LSTM cell uses the input gate i t , forget gate f t , and output gate o t These three structures to control the process of discarding and transmitting the input information. In an LSTM cell, the following operations are performed. For each LSTM cell, the inputs are x t , h t-1 , c t-1 , and the outputs are h t and c t . The LSTM algorithm can be expressed by the following formula:
[0004] i t = sigmoid(W ii x t + b ii + W hi h t-1 + b hi )
[0005] f t = sigmoid(W if x t + b if + W hf h t-1 + b hf )
[0006] g t = tanh(W ig x t + b ig + W hg h t-1 + b hg )
[0007] Ot = sigmoid(W io x t + b io + W ho h t-1 + b ho )
[0008] c t = f t ⊙ c t-1 + i t ⊙ g t
[0009] h t = o t ⊙ tanh(c t )
[0010] In the above formula, x t represents the input data of the LSTM cell at time t, and h t-1 represents the output value of the hidden layer at time t - 1. i t , f t , o t represent the output results of the input gate, forget gate, and output gate of the LSTM cell respectively. c t represents the cell state value of the LSTM cell at time t. g t represents the candidate cell state at time t. The ⊙ operation represents the accumulated result after multiplying the corresponding elements in two vectors. W ii , W if , W ig , W io , W hi , W hf , W hg , W ho are all weight matrices. W ii represents the relationship between the input node x t and the input gate i t . W hi represents the relationship between the output h t-1 of the hidden layer at the previous time and the input gate i t . The rest can be deduced by analogy. b ii , b if , b ig , b io , b hi , b hf , b hg , b ho are bias vectors. There are two activation functions in the LSTM cell, the sigmoid function and the tanh function. The output range of the sigmoid function is (0, 1). The output range of the tanh function is (-1, 1).
[0011] As the accuracy of model prediction increases, the number of model layers and parameters of the neural network also increases. In the ImageNet competition, with the successive introduction of deep neural network models such as AlexNet, VGG, GoogLeNet, and ResNet, the image recognition accuracy has been rising, indicating that the structure of the deep neural network is becoming more complex. While ensuring the accuracy of the deep neural network model, the model is compressed to make the model structure simpler, the number of parameters fewer, the amount of computation less, and the power consumption reduced. Common model compression methods include structured or unstructured pruning of the neural network, low-bitwidth quantization of the parameters in the model, etc.
[0012] For deep neural networks, common hardware acceleration platforms include GPUs, ASICs, and FPGAs. FPGAs are reconfigurable, have good flexibility, and low power consumption. For different neural networks, FPGAs can design different hardware accelerators for them. FPGAs have rich on-chip resources and support parallel and pipelined operations, and are widely used in the acceleration of the inference stage of neural networks.
[0013] The hardware implementation and model compression of recurrent neural networks have challenges and difficulties different from those of other neural networks. For the long short-term memory network LSTM, this type of recurrent neural network has special activation functions and complex calculation processes, thus posing higher requirements on the hardware. In view of the characteristics of the recurrent neural network algorithm, an FPGA accelerator is designed under the constraint of ensuring low or even lossless model accuracy. Summary of the Invention
[0014] In view of the technical problems existing in the prior art, the present invention provides an LSTM accelerator with low-bitwidth quantization compression, which reduces power consumption and greatly reduces the inference delay time on the premise of ensuring low loss of model accuracy.
[0015] The technical solution adopted by the present invention is: an LSTM accelerator with low-bitwidth quantization compression, including a storage module, a matrix-vector multiplication calculation module, an activation function module, and a dot product operation module. The storage module is respectively connected to the matrix-vector multiplication calculation module, the activation function module, and the dot product operation module, and the matrix-vector multiplication calculation module, the activation function module, and the dot product operation module are connected in sequence;
[0016] The storage module is used to cache the weight matrix, bias vector, input data, output value, and cell state value;
[0017] The matrix-vector multiplication calculation module is used to perform matrix-vector multiplication calculation of the weight matrix and the input data or the output value of the previous moment, and then add the bias vector; the weight matrix, the input data, and the output value of the previous moment are quantized from 32-bit floating-point numbers to 8-bit integers for calculation in the matrix-vector multiplication calculation module, and the calculation result of the matrix-vector multiplication calculation module is dequantized to 32-bit floating-point numbers;
[0018] The activation function module uses a piecewise linear function to calculate the calculation result of the matrix-vector multiplication calculation module;
[0019] The dot product operation module performs dot product operations on the vectors of the calculation results of the activation function module to calculate the output value and the cell state value.
[0020] Furthermore, the formulas for quantization and dequantization operations are as follows:
[0021]
[0022] q r = q * scale
[0023]
[0024]
[0025] In the formula, r represents the true value before quantization; scale represents the scaling factor; n represents the number of bits after quantization, n = 8; q represents the fixed-point value after quantization, which is an integer; q r represents the value after dequantization, which is a floating-point number; round represents the integer obtained by rounding; (a, b) represents the quantization range.
[0026] Furthermore, when backpropagating in the LSTM neural network, a straight-through estimator is used for processing, and the formula is q r = Q(r)
[0027]
[0028] In the formula, Q represents the quantization function, and c represents the objective function.
[0029] Furthermore, the weight matrix and the input data are stored in the storage module in the form of 8-bit integers.
[0030] Furthermore, the weight matrix is stored in the storage module in the form of 8-bit integers after block cyclic compression.
[0031] Further, during the matrix-vector multiplication calculation, the weight matrix is divided into several slice matrices of size a1*b1, and the input vector is divided into several slice input vectors of length b1. During the matrix-vector multiplication calculation between the slice matrix and the slice input vector, a1 accumulation calculations are completed in parallel, and the multiplication calculations of the b1 multiplication calculation results corresponding to each accumulation calculation are completed in parallel. The weight matrix is sliced and divided, and two-layer loops are fully expanded inside the slice to increase the parallelism of the program. The inner loop is b multiplication calculations, and the outer loop is a accumulation calculations. The input vector is input data or an output value.
[0032] Further, if the weight matrix cannot be divided evenly by the size of the slice matrix, then padding with 0s is performed around the weight matrix.
[0033] Further, the sigmoid function and the tanh function are represented by 20-segment piecewise linear functions.
[0034] Further, the dot product operation module adopts a loop pipelining pragma instruction.
[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0036] 1. The present invention uses the QAT method to quantize the weights and activation values of the LSTM neural network, quantizing 32-bit floating-point weights and activations into 8-bit integers. During the forward propagation process of model training, simulated quantization inference is performed. All weights and activation values are floating-point numbers, and backpropagation can still continue after processing, so that updated parameters can be obtained. Simulated quantization is to restore the accuracy of the LSTM model. When performing model inference, 8-bit fixed-point calculations can be used to replace the floating-point calculations in the neural network, which is very beneficial for the implementation of LSTM on FPGA. Quantizing activation values and weights with low bit widths significantly reduces the consumption of computing resources.
[0037] 2. The weight matrix of the present invention is a matrix after block cyclic compression. The weight matrix after block cyclic compression is dense and structured, which is beneficial for the implementation of FPGA hardware, can reduce the number of model parameters, and the storage space is also correspondingly reduced without loss of model accuracy.
[0038] 3. The present invention slices the weight matrix, and the parallel operations inside the slice, at the cost of consuming the hardware resources on the FPGA, greatly speeds up the execution speed of the LSTM model. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a structural block diagram of an embodiment of the present invention;
[0040] Figure 2 It is a schematic diagram of loop unrolling inside the slice of an embodiment of the present invention. Detailed implementation manners
[0041] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0042] An embodiment of the present invention provides a low-bitwidth quantization-compressed LSTM accelerator, as Figure 1 shown, which includes a storage module, a matrix-vector multiplication calculation module, an activation function module, and a dot product operation module. The storage module is respectively connected to the matrix-vector multiplication calculation module, the activation function module, and the dot product operation module, and the matrix-vector multiplication calculation module, the activation function module, and the dot product operation module are connected in sequence.
[0043] The storage module is used to cache the weight matrix, bias vector, input data, output value, and cell state value.
[0044] The matrix-vector multiplication calculation module is used to perform matrix-vector multiplication calculation of the weight matrix and the input data or the output value of the previous moment, and then add the bias vector. The weight matrix, input data, and output value of the previous moment are quantized from 32-bit floating-point numbers to 8-bit integers and calculated in the matrix-vector multiplication calculation module, and the calculation result of the matrix-vector multiplication calculation module is dequantized to 32-bit floating-point numbers.
[0045] The quantization and dequantization operations adopt the following formulas:
[0046]
[0047] q r = round(r / scale) * scale (1 - 2)
[0048]
[0049]
[0050] In the formula, r represents the true value before quantization; scale represents the scaling factor, generally a floating-point number; n represents the number of bits after quantization, n = 8; q represents the fixed-point value after quantization, which is an integer; q r represents the value after dequantization, which is a floating-point number; round represents the integer after rounding; (a, b) represents the quantization range.
[0051] Formula (1 - 1) is used for quantization operation, and formula (1 - 2) is used for dequantization operation. It is precisely the quantization operation and dequantization operation that enable the LSTM model to simulate quantization during training. Formula (1 - 3) maps the quantization range to a specified representation range to obtain the scaling factor. Formula (1 - 4) is a truncation function.
[0052] During the forward propagation of the already quantized LSTM neural network, since the round function is used for rounding, the round function is not differentiable.
[0053] During the backpropagation of the LSTM neural network, the Straight-Through Estimator (STE) is used to approximately process the gradients between layers during the backpropagation process. The formula
[0054] q r = Q(r) (1-5)
[0055]
[0056] In the formula, Q represents the quantization function, and c represents the objective function. Formula (1-5) represents the relationship between q r and r. Formula (1-6) represents the STE algorithm.
[0057] According to Formula (1-5) and Formula (1-6), during the backpropagation of the LSTM neural network, STE directly passes the gradient of the previous layer to the next layer. Using STE to solve the problem of uncomputable gradients in the quantized LSTM neural network enables the quantized neural network to continue training, so that the model can learn the loss brought by quantization and improve the model accuracy.
[0058] For the weights, the selection of the quantization range is very simple, a = max(w), b = min(w). For symmetric quantization, take a = max(w), b = -max(w).
[0059] For the activation values, the input of each LSTM cell has x t and h t-1 , and the activation values of the LSTM network are determined by the input. By analyzing the activation values during training, the ranges (a, b) of the activations of x t and h t-1 are determined respectively. After symmetric quantization, for x t , the selection of the quantization range: a = max(x), b = -max(x). For h tSelection of quantization range: a = max(h), b = -max(h). The data distributions of the inputs for each batch are different. If the dataset in each batch is small, the maximum and minimum values obtained for this batch often differ significantly from the actual maximum and minimum values of the entire training set. During the training of the LSTM model, the batch value is generally large, and many images are input in the same batch to speed up the model training. However, during the inference process of the LSTM model, it is hoped that the model can support testing with a single image, that is, the batch is set to 1. Through formulas (1-7) and (1-8), a relatively fixed temp_max value of the input x can be obtained by sliding.
[0060] temp_max = momentum * temp_max + (1 - momentum) * max_x (1-7)
[0061]
[0062] batch_num represents the number of batches during the training of the model and can be understood as dividing the training dataset into batch_num parts. max_x represents the maximum value of the input for this batch, and 1 - momentum represents the proportion of the maximum value x of the input for this batch when calculating the temp_max value for this batch. As the number of batches increases, the value of momentum gradually decreases, and the proportion of the maximum value max_x of the input for this batch gradually increases. momentum is a variable that changes with the number of input batches. temp_max represents the relatively fixed maximum value that conforms to the real input for all inputs after training.
[0063] When performing LSTM model inference on the FPGA, regardless of how different the input image samples are, it is considered that the maximum and minimum values of the input information are the temp_max values obtained during model training. During the training process, temp_max is continuously iterated with the input information of the model and is a variable. During the inference process of the model, temp_max is a fixed value and will no longer change with the input image information.
[0064] When deploying the LSTM neural network model to the hardware FPGA, during the inference stage of the model, the matrix-vector multiplication calculation module only performs integer multiplication operations, rather than floating-point multiplication operations.
[0065] Due to the particularity of the LSTM model, h t is continuously updated with the time dimension during the inference process of the LSTM model. Therefore, for h tThe quantization operation is implemented in real time during the inference process of the LSTM model on the FPGA, unlike the weights and input images, which can be quantized into 8-bit integers in advance before the LSTM model performs inference. The weight matrix and input data are stored in the storage module in the form of 8-bit integers.
[0066] The weight matrix is the matrix after block-cyclic compression. The weight matrix is stored in the storage module in the form of 8-bit integers after block-cyclic compression. The weight matrix after block-cyclic compression is dense and structured, which is beneficial to the implementation of FPGA hardware and reduces the storage space. After retraining the LSTM model using the block-cyclic matrix, an accuracy similar to that of the LSTM model without using the block-cyclic matrix can be achieved, and the precision is not lost due to this.
[0067] During matrix-vector multiplication calculation, the weight matrix is divided into several slice matrices of size a1*b1, and the input vector is divided into several slice input vectors of length b1. If the weight matrix cannot be divided evenly by the size of the slice matrix, 0s are padded around the weight matrix.
[0068] The calculation process between the sliced weight matrix and the input vector is as follows: After padding 0s, the size of the weight matrix is p1*q1, and the size of the input vector is q1. After performing matrix-vector multiplication operation, the resulting vector has a size of p1. The size of the slice matrix is a1*b1, the size of the slice input vector is b1, and the size of the slice result vector is a1. Then the weight matrix consists of m1*n1 block slice matrices, where m1 = p1 / a1 and n1 = q1 / b1. The input vector is also divided into n1 block slice input vectors, and the resulting vector consists of n1 block slice result vectors. The n1 block slice matrices in each row of the weight matrix perform n1 times of slice matrix-vector multiplication operations with the input vector divided into n1 block slice input vectors. After each slice matrix-vector multiplication operation, a slice result vector of size a1 is obtained. After n1 times of slice matrix-vector multiplication operations, n1 slice result vectors of size a1 are obtained. Performing element-wise addition operations on these n1 slice result vectors of size a1 results in 1 vector res of size a1. The weight matrix has m1 rows, and after each row performs matrix-vector multiplication operation, 1 vector res of size a1 is obtained. After the m1 rows of the weight matrix are calculated, m1 vectors res of size a1 are obtained. Unfolding and connecting these m1 vectors res of size a1 together gives the result vector of size m1*a1 = p1.
[0069] For the matrix-vector multiplication operation inside the slice, the pragma directive is used for code optimization. Through the logic of the algorithm inside the slice shown below, it can be seen that the matrix-vector multiplication operation inside the slice has two nested loops. For these two nested loops, the UNROLL in the pragam directive is used to perform loop unrolling operations. Matrix-vector multiplication algorithm inside the slice:
[0070]
[0071]
[0072] The size of the slice matrix is a1*b1. The inner loop runs b1 times and the outer loop runs a1 times. After the inner and outer loops are fully unrolled, the inner loop performs multiply-accumulate operations. Since the b1 multiply-accumulate operations in the inner loop are completely independent and there are no memory access conflicts, the b1 inner loops can be executed in parallel. The inner loop first performs b1 multiplication operations in parallel, and then accumulates the b1 results obtained to get one element in the final result vector.
[0073] Each unrolling of the inner loop in the slice is equivalent to an accumulation tree. Each accumulation tree performs b1 multiplication operations, and these b1 multiplication operations are parallel to each other. The outer loop runs a1 times, so the entire inside of the slice is equivalent to a1 accumulation trees that perform b1 multiplications. These accumulation trees are parallel to each other. As Figure 2 shown, it represents 2 accumulation trees, that is, two levels of outer loops. The calculations of these two accumulation trees are performed in parallel on the FPGA. There are 4 multiplication operations hanging on each accumulation tree, and each multiplication operation represents one level of the inner loop. These 4 multiplications are also completely parallel. After these 4 parallel multiplications are completed, accumulation is performed. The parallel operations inside the slice, at the cost of consuming hardware resources on the FPGA, greatly accelerate the execution speed of the LSTM model.
[0074] For convenience of processing, the block cyclic matrix and the slice matrix can have the same size.
[0075] The activation function module uses a piecewise linear function to calculate the calculation result of the matrix-vector multiplication calculation module. The sigmoid function and the tanh function are represented by 20-piece piecewise linear functions. After the activation function is piecewise linearly quantized, only the slopes and intercepts of the linear functions of each segment need to be stored. During actual operation, the input of the activation function needs to be located to calculate the corresponding index value to determine which piecewise function to use. In this way, the activation function, which is very resource-consuming on the FPGA due to transcendental function operations and floating-point division operations, is transformed into a simple linear function, only requiring the calculation of the piecewise function index value, 1 floating-point multiplication, and 1 floating-point addition.
[0076] The dot product operation module performs a dot product operation on the calculation results of the activation function module to calculate the output value and the cell state value. The dot product operation module uses a pragma instruction for loop pipelining to accelerate its operation speed. When calculating h t and c t , there is no data dependence within the loop, so the loop initiation interval can reach 1. Loop pipelining enables some operations of the loop to be overlapped, thereby accelerating the execution speed of the program.
[0077] In this embodiment, the FPGA development board Zedboard is used to construct a low-bitwidth quantization compressed LSTM accelerator to perform inference on the LSTM neural network to evaluate the performance.
[0078] In this embodiment, the pytorch deep learning framework is used to train the LSTM neural network. After multiple trainings, when the model reaches an ideal accuracy rate, the parameters such as the weights in the LSTM neural network are saved as binary files. In the pytorch deep learning framework, the network model adopted is a 2-layer LSTM network and a 1-layer fully connected layer. In the LSTM network model, the number of hidden layer nodes is 16, the number of input nodes is 28, and the time series is 28. The dataset used by the model is the MNIST handwritten dataset recognition. There are 60,000 images in the MNIST dataset as the training set and 10,000 images as the test set, and the image pixel size is 28*28. In the LSTM model, since the model input is one-dimensional, when inputting two-dimensional data such as images, one dimension (28) of the image is used as the input node, and the other dimension (28) of the image is used as the time series.
[0079] In the inference stage of the LSTM model, FPGA hardware is used for acceleration. The model of the FPGA development board used is the ZedBoard development board. The chip of the ZedBoard development board is the Zynq-7000, and the architecture of the chip is divided into two parts: the processor system and the programmable logic. The core part of the PS of the ZedBoard development board is a dual-core ARMCorex-A9 processor with a maximum main frequency of 667MHz. The programmable logic part includes hardware resources such as BRAM, DSP, FF, and LUT. The processor system core and the programmable logic communicate with each other through the AXI protocol. Generally, the input information of the model and the weight parameters of the model trained in the deep learning network framework on the GPU are stored in the DDR. The storage resources on the ZedBoard development board include 512MB DDR3, 256Mb Quad-SPI flash memory, and 4GB SD card.
[0080] Among the following experimental evaluation metrics, the three metrics of resource utilization rate, inference latency, and power consumption are all measured under the condition of inferring one image in the FPGA. In this embodiment, the slice size of the matrix-vector multiplication is 4*4, and correspondingly, the size of the block circulant matrix is also 4.
[0081] I. Resource utilization rate: Run the LSTM algorithm on the FPGA development board ZedBoard to infer one image. The usage of the on-chip resources of the FPGA, namely LUT, FF, BRAM, and DSP, is shown in Table 1.
[0082] Table 1 Resource utilization rate of the FPGA accelerator
[0083] Condition LUT FF BRAM DSP Initial 13.78% 8% 6.80% 9.09% After slicing 32.08% 16.99% 28.60% 46.82% After quantization 23.43% 11.90% 18.2% 25.45% After block circulant matrix 21.44% 10.81% 18.2% 25.45%
[0084] After matrix slicing, within a small matrix block, loop unrolling operations are performed. Therefore, the consumption of the on-chip resources of the FPGA increases. After low-bitwidth quantization of the weights and activation values, during the inference stage of the LSTM model, only 8-bit integer calculations are performed, reducing memory occupancy, decreasing the consumption of BRAM resources. Integer multiplication operations are cheaper than floating-point multiplication operations, and resources such as LUT, FF, and DSP logic calculations also decrease significantly. It can be proven that quantizing the 32-bit floating-point numbers in the LSTM model to 8-bit integers is very beneficial for the utilization rate of the on-chip resources of the FPGA. After applying the block circulant matrix to the LSTM model, the storage space required for the model parameters decreases. Since the model itself is not particularly large, after low-bitwidth quantization, the model size is greatly reduced, and the on-chip resources required to store the model change from BRAM to FF before, so the FF resources decrease. And LUT usually works closely with FF, so the resources of LUT also decrease accordingly.
[0085] II. Inference latency: Deploy the LSTM model to the ZedBoard development board, input a handwritten digit image with a pixel size of 28*28, measure the time required for the ZedBoard development board to infer one such image, and analyze the inference latency of the LSTM model under different conditions such as slicing operation, low-bitwidth quantization operation, and block circulant matrix operation, as shown in Table 2.
[0086] Table 2 Inference latency of the FPGA accelerator
[0087] Condition Inference latency (unit: ms) Initial 17.506 After slicing 1.114 After quantization 1.125 After block circulant matrix 0.917
[0088] After matrix slicing, the inner loop of the small matrix blocks is unrolled, which greatly speeds up the program execution. Therefore, when the LSTM model performs actual inference, the speed is 15.71 times that without using the slicing operation. It can be seen that the matrix slicing operation fully considers parallelism, reducing the inference latency time. It only takes 1.114 ms to recognize a 28*28 size picture. After quantization, the inference speed of the LSTM model slows down slightly, which can be ignored. After the LSTM model applies the block cyclic matrix, since the model data decreases, the time required to read the model parameters also decreases, which is beneficial to further reducing the inference latency time. Finally, it can reach a speed of 0.917 ms to recognize a 28*28 size picture, and the effect is very good.
[0089] III. Model Accuracy: The accuracy of the LSTM model was measured in the deep neural network framework pytorch and on the FPGA development board, as shown in Table 3. Since the matrix-vector multiplication slicing does not involve discarding model information and does not affect the model accuracy, it is not listed in Table 3.
[0090] Table 3 Model Accuracy
[0091] Condition Test set during training During inference Initial 97.8% 100% After quantization 97.4% 97% After block circulant matrix 97% 97%
[0092] In pytorch, after quantizing the floating-point numbers in the LSTM model to 8-bit integers and then retraining after quantization, the accuracy of the model can be basically restored and is not greatly affected. For the block cyclic matrix, the same method of retraining the LSTM model is adopted, and the model can be basically restored to the original prediction accuracy, and the accuracy loss is acceptable.
[0093] IV. Power Consumption: The power consumption was measured by inferring a single picture on the ZedBoard, and the measurement results are shown in Table 4. The static power consumption of the FPGA chip is determined by its own structure, and the dynamic power consumption can be reduced by developers designing and optimizing the FPGA accelerator.
[0094] Table 4 Power Consumption of the FPGA Accelerator
[0095] Condition Total power consumption (unit: W) Static power consumption Dynamic power consumption Initial 1.812 8% 92% After slicing 2.148 7% 93% After quantization 1.983 8% 92% After block circulant matrix 1.974 8% 92%
[0096] After performing matrix slicing calculations, the parallelism of the LSTM algorithm on the FPGA platform increases due to the loop expansion inside the slices, so that the FPGA uses multiple resources for calculations at the same time, so the power consumption of the FPGA is higher than that without using slicing. After low-bit-width quantization of floating-point numbers, the power consumption of the FPGA is reduced again because the model parameter size becomes smaller and 8-bit fixed-point multiplication consumes fewer resources than floating-point multiplication. After using block circulant matrices, the memory space required by the LSTM model is reduced, and the power consumption of the FPGA is also reduced.
[0097] In summary, in terms of resource utilization, the slicing operation requires more resources due to the internal loop expansion. The low-bit width quantization activation value and weight reduce the FPGA on-chip resources most significantly, and the DSP resources consumed by multiplication calculation are reduced by 21.37%. In terms of inference latency, the code parallelism brought by the slicing operation greatly reduces the inference latency of the LSTM model on the FPGA, and the inference speed is increased by 15.71 times. In terms of model accuracy, the slicing operation does not affect the model accuracy. The activation function in the LSTM algorithm is approximated by a piecewise linear function, and the loss of model prediction accuracy caused by this can be ignored. Therefore, the compression operation that affects the accuracy of the LSTM model is the low-bit width quantization and block cyclic matrix. The losses caused by these two compression methods can be recovered by retraining. In terms of power consumption, slicing reduces the FPGA power consumption. The application of quantization and block cyclic matrix can reduce the FPGA power consumption. The matrix slicing operation significantly reduces the inference latency at the cost of increased on-chip resource utilization.
[0098] The present invention has been described in detail above through embodiments, but the contents described are only exemplary embodiments of the present invention and cannot be considered to limit the scope of implementation of the present invention. The protection scope of the present invention is defined by the claims. Anyone who utilizes the technical solution described in the present invention, or a technician in the field, inspired by the technical solution of the present invention, designs a similar technical solution within the essence and protection scope of the present invention to achieve the above technical effects, or makes equal changes and improvements to the scope of application, etc., shall still fall within the scope of protection covered by the patent of the present invention.
Claims
1. A low-width quantization compression LSTM accelerator, characterized in that: It includes a storage module, a matrix-vector multiplication calculation module, an activation function module, and a dot product operation module. The storage module is respectively connected to the matrix-vector multiplication calculation module, the activation function module, and the dot product operation module, and the matrix-vector multiplication calculation module, the activation function module, and the dot product operation module are connected in sequence; The storage module is used to cache the weight matrix, bias vector, input data, output value, and cell state value; The matrix-vector multiplication calculation module is used to perform matrix-vector multiplication calculation of the weight matrix and the input data or the output value of the previous moment, and then add the bias vector; the weight matrix, input data, and output value of the previous moment are quantized from 32-bit floating-point numbers to 8-bit integers and calculated in the matrix-vector multiplication calculation module, and the calculation result of the matrix-vector multiplication calculation module is dequantized to 32-bit floating-point numbers; The activation function module uses a piecewise linear function to calculate the calculation result of the matrix-vector multiplication calculation module; The dot product operation module performs a dot product operation on the calculation result of the activation function module to calculate the output value and the cell state value; During matrix-vector multiplication calculation, the weight matrix is divided into several slice matrices of size a1×b1, and the input vector is divided into several slice input vectors of length b1; in the matrix-vector multiplication calculation of the slice matrix and the slice input vector, a1 accumulative calculations are completed in parallel, and the multiplication calculations of the multiplication calculation results corresponding to b1 for each accumulative calculation are completed in parallel.
2. The low-width quantization compression LSTM accelerator according to claim 1, characterized in that: The formulas used for quantization and dequantization operations are: ; Wherein, r represents the true value before quantization; scale represents the scaling factor; n represents the number of bits after quantization, n = 8; q represents the fixed-point value after quantization, which is an integer; q r represents the value after dequantization, which is a floating-point number; round represents the integer obtained by rounding; (a, b) represents the quantization range.
3. The low-width quantization compression LSTM accelerator according to claim 2, characterized in that: When backpropagating in the LSTM neural network, a straight-through estimator is used for processing, and the formula is: ; In the formula, Q represents the quantization function, and c represents the objective function.
4. The low-width quantization compression LSTM accelerator according to claim 2, characterized in that: The weight matrix and input data are stored in the storage module in the form of 8-bit integers.
5. The low-width quantization compression LSTM accelerator according to claim 2, characterized in that: The weight matrix is stored in the storage module in the form of 8-bit integers after block cyclic compression.
6. The low-width quantization compression LSTM accelerator according to claim 5, characterized in that: If the weight matrix cannot be divisible by the size of the slice matrix, 0s are supplemented around the weight matrix.
7. The low-width quantization compression LSTM accelerator according to claim 1, characterized in that: The sigmoid function and the tanh function are represented by 20-piece piecewise linear functions.
8. The low-width quantization compression LSTM accelerator according to claim 1, characterized in that: The dot product operation module adopts a loop pipelining pragma instruction.
Citation Information
Patent Citations
Machine learning sparse computation mechanism for arbitrary neural networks, and arithmetic compute microarchitecture, and sparsity for training mechanism
CN109993683A
Accelerated quantized multiply-and-add operations
CN111937010A