HE-LSTM Network Structure and Its Corresponding FPGA Hardware Accelerator
By simplifying the forget gate calculation formula of the LSTM network and replacing the nonlinear activation function with segmented linear functions, the hardware resource consumption of traditional LSTM networks on the FPGA platform is reduced, and more efficient neural network inference is achieved.
Patent Information
- Application Number
- CN202011470304.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-14
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-12-14
AI Technical Summary
When running on the FPGA platform, due to too many redundant parameters, the hardware resource consumption is too large, making it difficult to achieve efficient neural network reasoning.
A HE-LSTM network structure is proposed, which reduces redundant parameters by simplifying the calculation formula of the forgetting gate, and builds a hardware accelerator on the FPGA, and uses a segmented linear function to replace the nonlinear activation function to reduce the computational complexity.
With the same computing speed, the consumption of hardware resources is saved by about 37%, and the impact of model accuracy is almost negligible, greatly reducing the redundant parameters of traditional LSTM networks, allowing FPGAs to more efficiently complete the inference stage operations of neural networks.
Smart Images

Figure CN112561036B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural network hardware acceleration, and proposes a novel LSTM network structure (HE-LSTM) and its corresponding FPGA accelerator. Background Art
[0002] Recurrent neural networks (RNNs) are a class of neural networks for processing sequential signals and have been widely used in machine translation, movie review sentiment analysis, and other natural language-related applications. However, traditional RNNs suffer from the problems of vanishing gradients and exploding gradients. To address these issues, researchers proposed a variant of RNN called Long Short-Term Memory (LSTM), which captures long-term dependencies in sequential signals by introducing three gating units (including a forget gate, an input gate, and an output gate) and a storage unit (memory cell). Therefore, LSTM networks achieve higher accuracy and adaptability than traditional RNNs. The commonly used LSTM structure expression is as follows:
[0003] f t = σ(W fx x t + W fh h t-1 + b f )
[0004] i t = σ(W ix x t + W ih h t-1 + b i )
[0005] o t = σ(W ox x t + W oh h t-1 + b o )
[0006]
[0007]
[0008] h t = o t * tanh(c t )
[0009] i t is the input gate, which determines how much new information is added to the memory cell; f t is the forget gate, which determines how much previous information is forgotten; o t is the output gate, which determines the memory cell ct Relationship with the output of h t ; is a candidate memory unit. W is the weight matrix, and the operation between the weight matrix and the vector is matrix-vector multiplication, where * represents the dot product of vectors. b is the bias, and σ and tanh are two non-linear activation functions.
[0010] Although LSTM solves the problems of gradient vanishing and gradient explosion of traditional RNNs, due to the introduction of multiple gating units in LSTM, it has more parameters. Running the LSTM algorithm directly on a CPU does not work well. Therefore, a dedicated hardware acceleration circuit needs to be designed to accelerate the inference process of LSTM. Currently, the commonly used RNN acceleration hardware platforms are GPUs, FPGAs, and ASICs. In the LSTM network, the vast majority of operations are matrix-vector multiplications or dot products between vectors. Although the performance of GPUs in computing the LSTM algorithm is stronger than that of CPUs, the power consumption of GPUs is relatively high, so it is not suitable for scenarios with limited maximum power consumption. Although ASICs have a high energy efficiency ratio, their flexibility is relatively low, making them unsuitable for neural network structures with rapid algorithm updates. FPGAs have both high flexibility and low power consumption characteristics, so they are often used to implement RNN acceleration. However, due to limited FPGA resources, it is difficult to directly run large LSTM networks on FPGAs. Therefore, some optimization algorithms need to be introduced to reduce the large number of redundant parameters in the LSTM network.
[0011] To reduce the number of parameters in the model, many methods have been proposed by researchers, such as pruning, structured compression, replacing non-linear activation functions with piecewise linear functions, pipelined architectures, approximate computing, etc. In addition, researchers have also proposed some new LSTM variants, such as Coupled-Gate LSTM, ELSTM, GRU, etc. Although these variants of LSTM have reduced many parameters, they still have too many redundant parameters, making it difficult to run on a platform with limited hardware resources. To meet the growing demand for high-performance RNN accelerators, the research on an LSTM network structure with fewer parameters has become one of the hot research directions. Summary of the Invention
[0012] Objective of the Invention: To overcome the deficiencies in the prior art, the present invention provides a HE-LSTM network structure and an FPGA hardware accelerator. Traditional LSTM has a large number of weight matrices and contains a lot of redundant parameters. Therefore, there is still a large research space for the optimization of the LSTM unit structure. LSTM captures the long-term dependence of sequence signals by introducing three gating units and memory cells. While the input gate controls the input of information, the forget gate controls the selective forgetting of past information by the memory cells. Therefore, there is a certain functional relationship between the input gate and the forget gate. Through theoretical derivation and experimental verification, it is obtained that i t = 1 - f t . Moreover, the formula for the forget gate f t of traditional LSTM is f t = σ(W fx x t + W fh h t-1 + b f ). We believe that the size of the forget gate can be completely calculated by the output h t-1 of the previous time step in the neural network, and W fx x t is a redundant term. Therefore, the formula for calculating the forget gate of the new HE-LSTM is f t = σ(W fh h t-1 + b f ). The present invention builds a HE-LSTM network on the FPGA. Compared with the traditional LSTM network structure with the same architecture, when the operation speed is the same, the consumption of hardware resources is saved by about 37%, and the impact of this structure on the model accuracy is almost negligible. The present invention greatly reduces the redundant parameters of the traditional LSTM network, enabling the FPGA to more efficiently complete the operations in the inference stage of the neural network.
[0013] Technical Solution: To achieve the above objective, the technical solution adopted by the present invention is as follows:
[0014] A HE-LSTM network structure includes a forget gate, an output gate, and a memory cell, and the expressions are as follows:
[0015] f t = σ(W fh h t-1 + b f )
[0016] o t = σ(W ox x t + W oh h t-1 + b o )
[0017]
[0018]
[0019] h t = o t *tanh(c t )
[0020] Among them, f t is the forget gate, o t is the output gate, is the candidate memory cell, c t represents the memory cell, h t represents the current output, W is the weight matrix, W ox represents the weight matrix from the input vector x to the output gate, W cx represents the weight matrix from the input vector x to the candidate memory cell, W fh represents the weight matrix from the output of the previous time step to the forget gate, W oh represents the weight matrix from the output of the previous time step to the output gate, W ch represents the weight matrix from the output of the previous time step to the candidate memory cell, b is the bias, b f represents the bias when calculating the forget gate, b o represents the bias when calculating the output gate, b c represents the bias when calculating the candidate memory cell, σ is the first non-linear activation function, tanh is the second non-linear activation function, x t represents the input of the current network, and t represents the current time.
[0021] A method for using a HE-LSTM network structure includes the following steps:
[0022] Step 1), in the traditional LSTM network, the forget gate f t = σ(W fx x t + W fh h t-1 + b f ), is changed to f t = σ(W fh h t-1 + b f ), and 1 - f t is used to replace the input gate i t , so as to obtain a HE-LSTM network structure with fewer parameters;
[0023] Step 2), initialize the HE-LSTM network model so that all its trainable parameters (weight matrix and bias) follow a normal distribution within the range of 0 to 1, and train the network until it converges.
[0024] Step 3), replace the non-linear activation functions in the network with piecewise linear functions and train until convergence.
[0025] Use piecewise linear functions to replace the two non-linear activation functions, σ and tanh, which originally contain the exponential power form of e. In the actual FPGA implementation, such a replacement method significantly reduces the computational complexity and the consumed hardware resources.
[0026] Step 4), quantize the trained 32-bit floating-point numbers into 16-bit fixed-point numbers to obtain the finally compressed network.
[0027] Quantize the 32-bit floating-point numbers into 16-bit fixed-point numbers, including 1 sign bit, 2 integer bits, and 13 decimal bits.
[0028] An FPGA hardware accelerator based on the HE-LSTM network structure includes a weight matrix storage unit, an arithmetic processing unit, an adder, a piecewise linear function activation module, and an Element-Wise operation module, where: the weight matrix storage unit is used to store the weight matrix; the arithmetic processing unit includes a matrix-vector multiplication operation module and an accumulator. The matrix multiplication operation module performs the matrix-vector multiplication operation of the weight matrix and the input x t or the output h of the previous time step t-1 The accumulator performs a multiply-accumulate operation on the result of the matrix multiplication operation module. The adder performs an addition operation on the result obtained from the multiply-accumulate operation and the bias value. The piecewise linear function activation module uses a piecewise linear function to operate on the result of the addition operation; the Element-Wise operation unit performs a dot product operation on the value obtained by the piecewise linear function activation module to calculate the final outputs h t and the memory cell c t .
[0029] Preferably: The arithmetic processing unit includes an accumulator unit and two or more E units, and accumulates the outputs of the PE units through the accumulator unit to implement the matrix-vector multiplication operation.
[0030] Preferably: Use BRAM to store each weight matrix and intermediate values.
[0031] Preferably: The weight matrix is stored in BRAM in the form of 16-bit fixed-point numbers.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] The HE-LSTM network structure of the present invention has fewer gating units and weight matrices, and the number of parameters is greatly reduced. Therefore, it is more suitable for running on platforms with limited hardware resources such as FPGAs. In addition, by using piecewise linear functions to replace non-linear activation functions, the computational complexity is greatly reduced. In the LSTM algorithm, matrix-vector multiplication operations are the most resource-consuming and complex operations. On FPGAs, a block-parallel structure can be adopted to efficiently process the matrix multiplication operations. The hardware accelerator built using FPGAs also has characteristics such as low power consumption, fast running speed, and low power consumption. Description of the Drawings
[0034] Figure 1 is the structural diagram of the HE-LSTM network designed in this patent;
[0035] Figure 2 is the flow chart of the training strategy of the HE-LSTM network;
[0036] Figure 3 is the schematic diagram of the PE unit structure;
[0037] Figure 4 is the schematic diagram of the FPGA architecture of a single-layer HE-LSTM. Detailed Implementation Manner
[0038] The above-mentioned proposed solution will be further described below in conjunction with specific examples of machine translation (encoder-decoder structure). These examples are only used to further illustrate the present invention and do not limit the scope of application of the present invention. The experimental conditions and specific parameter settings adopted in this machine translation can be further adjusted according to the requirements of specific manufacturers, and the experimental conditions not specified are those in general experiments.
[0039] An HE-LSTM network structure applied to machine translation, as Figure 1 shown, the forget gate f t =σ(W fx x t +W fh h t-1 +b f ) in the traditional LSTM network is changed to f t =σ(W fh h t-1 +b f ), and 1 - f t is used to replace the input gate i t , thereby obtaining an HE-LSTM network structure with fewer parameters. The expression of the obtained HE-LSTM network structure is as follows:
[0040] f t =σ(W fh ht-1 +b f )
[0041] o t = σ(W ox x t +W oh h t-1 +b o )
[0042]
[0043]
[0044] h t = o t *tanh(c t )
[0045] Among them, f t is the forget gate, o t is the output gate, is the candidate memory cell, c t represents the memory cell, W is the weight matrix, W ox represents the weight matrix from the input vector x to the output gate, W cx represents the weight matrix from the input vector x to the candidate memory cell, W fh represents the weight matrix from the output of the previous time step to the forget gate, W oh represents the weight matrix from the output of the previous time step to the output gate, W ch represents the weight matrix from the output of the previous time step to the candidate memory cell, b is the bias, b f represents the bias when calculating the forget gate, b o represents the bias when calculating the output gate, b c represents the bias when calculating the candidate memory cell, σ is the first non-linear activation function, tanh is the second non-linear activation function, t represents the serial number of the current input. The encoder and decoder of the machine translation model are both built based on HE-LSTM. In the encoder, h t represents the result after encoding the sentence to be translated, x t represents the t-th word of the sentence to be translated; in the decoder, h t represents the t-th word obtained by translation, x t represents the (t - 1)-th translated word.
[0046] Such as Figure 2As shown in the figure, a machine translation model is constructed using HE-LSTM units. The training of the HE-LSTM network is divided into three stages. After the first stage of model initialization. The second stage is arranged after the learning rate decay. The third stage is arranged after replacing the non-linear activation function with a piecewise linear function. Each training trains the network until convergence, specifically including the following steps:
[0047] Step 1), construct a translation model using the HE-LSTM network and perform an initialization operation on the model. Make the initialized weight matrix and bias follow a normal distribution within the range of 0 to 1.
[0048] Step 2), train the constructed translation model until convergence.
[0049] Step 3), replace the non-linear activation function in the HE-LSTM network with a piecewise linear function and retrain until convergence.
[0050] Use a piecewise linear function to replace the original two non-linear activation functions in the form of the exponential power of e, σ and tanh. In the actual FPGA circuit implementation, such a replacement method greatly reduces the computational complexity and reduces the consumed hardware resources.
[0051] Step 4), quantize the trained 32-bit floating-point number model into a 16-bit fixed-point number to obtain the final compressed network.
[0052] Implementing fixed-point number operations requires much less hardware overhead than floating-point number operations. Therefore, it is necessary to quantize the model parameters into fixed-point numbers. In Step 4, the 32-bit floating-point numbers in the model parameters are quantized into 16-bit fixed-point numbers, which include 1 sign bit, 2 integer bits, and 13 decimal bits.
[0053] An FPGA hardware accelerator based on the HE-LSTM network structure applied to machine translation, as Figure 3 、 4 shown, includes a weight matrix storage unit, an arithmetic processing unit, an adder, a piecewise linear function activation module, and an Element-Wise operation module (EWU). The weight matrix storage unit is used to store the weight matrix. The arithmetic processing unit includes a matrix-vector multiplication operation module and an accumulator. The matrix multiplication operation module performs the weight matrix and the input x t or the output h of the previous time step t-1For the matrix-vector multiplication operation, the accumulator performs a multiply-accumulate operation on the result after the matrix multiplication operation module. The adder adds the result obtained from the multiply-accumulate operation and the bias value. The piecewise linear function activation module uses a piecewise linear function to operate on the result after the addition operation. The Element-Wise operation unit performs a dot product operation on the vector for the value obtained from the piecewise linear function activation module to calculate the final output h t and the memory cell c t . The operation processing unit includes an accumulator unit and 16 PE units. The outputs of the PE units are accumulated through the accumulator unit to implement the matrix-vector multiplication operation.
[0054] As Figure 3 shown, the operation processing unit performs a multiplication operation. Each PE unit contains 16 multipliers and one adder. The weights of each weight matrix and the values of the input vector are 16-bit fixed-point numbers. Therefore, this module can implement the multiply-accumulate operation of 32 16-bit numbers.
[0055] Figure 4 is a schematic diagram of the FPGA architecture of the single-layer HE-LSTM network structure. In the machine translation model applied in the present invention, the dimension size of the input vector of the single-layer HE-LSTM network is set to 256, and the dimension size of the output of the hidden layer is also set to 256. As Figure 4 shown, the HE-LSTM network structure requires a total of 5 operation processing units. Each operation processing unit includes an accumulator unit and 16 PE units. The outputs of the PE units are accumulated through the structure of the adder tree to implement the matrix-vector multiplication operation. The weight matrix of the HE-LSTM is stored in the BRAM in the form of 16-bit fixed-point numbers (1 sign bit, 2 integer bits, and 13 decimal bits). During the operation, the weights of each row of the weight matrix are read in sequence, and the multiply-accumulate operation is performed in the operation processing unit. The obtained value is sent to the adder, and then through the activation module, the output value of the gating unit is obtained. Among them, the activation unit uses a piecewise linear function to replace the non-linear function, so the activation module can be implemented by a shifter and an adder. Then, all the calculated data are input into the EWU unit for operation to calculate the output h t and the memory cell c t . We use a ping-pong structure to store and read the output h t and the memory cell c t . Such a storage and reading method can effectively improve the efficiency of the HE-LSTM network.
[0056] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An FPGA hardware accelerator with an HE-LSTM network structure, characterized in that: The HE-LSTM network structure includes an input gate, an output gate, and a memory cell, and the expressions are as follows: f t = σ(W fh h t-1 + b f ) o t = σ(W ox x t + W oh h t-1 + b o ) h t = o t *tanh(c t ) Among them, f t is the forget gate, o t is the output gate, is the candidate memory cell, c t represents the memory cell, h t represents the current output, W is the weight matrix, W ox represents the weight matrix from the input vector x to the output gate, W cx represents the weight matrix from the input vector x to the candidate memory cell, W fh represents the weight matrix from the output of the previous time step to the forget gate, W oh represents the weight matrix from the output of the previous time step to the output gate, W ch represents the weight matrix from the output of the previous time step to the candidate memory cell, b is the bias, b f represents the bias when calculating the forget gate, b o represents the bias when calculating the output gate, b c represents the bias when calculating the candidate memory cell, σ is the first non - linear activation function, tanh is the second non - linear activation function, x t represents the input of the current network, t represents the current time; It includes a weight matrix storage unit, an arithmetic processing unit, an adder, a piecewise linear function activation module, and an Element-Wis e arithmetic module, where: the weight matrix storage unit is used to store the weight matrix; the arithmetic processing unit includes a matrix-vector multiplication arithmetic module and an accumulator. The matrix multiplication arithmetic module performs a matrix-vector multiplication operation on the weight matrix and the input x t or the output h of the previous time step t-1 . The accumulator performs a multiply-accumulate operation on the result of the matrix multiplication arithmetic module. The adder performs an addition operation on the result obtained from the multiply-accumulate operation and the bias; the arithmetic processing unit contains an accumulator unit and two or more PE units, and accumulates the outputs of the PE units through the accumulator unit to implement the matrix-vector multiplication operation; the piecewise linear function activation module uses a piecewise linear function to operate on the result after the addition operation; the Element-Wis e arithmetic unit performs a dot product operation on the vector of the value obtained by the piecewise linear function activation module to calculate the final output h t and the memory cell c t .
2. The FPGA hardware accelerator with the HE-LSTM network structure according to claim 1, characterized in that: BRAM is used to store each weight matrix and intermediate value.
3. The FPGA hardware accelerator with the HE-LSTM network structure according to claim 2, characterized in that: The weight matrix is stored in BRAM in the form of 16-bit fixed-point numbers.
4. The FPGA hardware accelerator with the HE-LSTM network structure according to claim 3, characterized in that: The usage method of the HE-LSTM network structure includes the following steps: Step 1), initialize the HE-LSTM network structure so that all its weight matrices and biases follow a normal distribution within the range of 0 to 1, and train the network until it converges; Step 2), replace the non-linear activation function in the network with a piecewise linear function and train until convergence; Step 3), quantize the trained 32-bit floating-point numbers into 16-bit fixed-point numbers to obtain the finally compressed network.
Citation Information
Patent Citations
Computing device and method applied to long short term memory neural network
CN108510065A
Hardware Accelerator for Compressed LSTM
US20180174036A1