FPGA-based One-dimensional CNN-LSTM Acceleration Platform and Implementation Method
By constructing linear and nonlinear computing units on FPGAs, resource multiplexing of one-dimensional convolution and matrix multiplication is realized, the problem of insufficient computing resources in the prior art is solved, and parallel acceleration of CNN-LSTM neural network is realized.
Patent Information
- Application Number
- CN202210804166.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-07-07
AI Technical Summary
The existing FPGA-based acceleration solution cannot support both convolution and matrix multiplication at the same time, and the acceleration platform lacks computing resources when deploying CNN-LSTM neural networks.
A linear operation unit and a nonlinear operation unit are built on an FPGA, and the resource multiplexing of one-dimensional convolution and matrix multiplication is realized through multiplication and addition arrays, and the operation is completed in combination with ping-pong result cache.
The parallel acceleration of the CNN-LSTM neural network model is realized, which improves the computing resource utilization rate of FPGA and solves the problem of insufficient computing resources.
Smart Images

Figure CN115222028B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of heterogeneous computing, and further relates to a one-dimensional CNN-LSTM neural network acceleration platform and implementation method in the field of deep learning computing acceleration, based on a Field Programmable Gate Array (FPGA). The present invention can be used for one-dimensional CNN-LSTM neural network computing acceleration. Background Art
[0002] As the problems to be solved by deep learning become more complex and abstract, the requirements of deep learning algorithms for device computing power are getting higher and higher, and general-purpose CPUs (Central Processing Units) can no longer meet the computing power requirements of deep learning. To meet the computing power requirements of deep learning, hardware devices such as GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), or FPGAs are usually used to provide computing power support. Among them, GPUs are widely used in the training and inference processes of deep learning algorithm models due to their powerful parallel computing processing capabilities. However, the power consumption problem of GPUs has always restricted their applications in mobile terminals and portable devices. ASICs usually adopt improved computing architectures to complete computing acceleration for specific tasks and generally have a high energy efficiency ratio. However, due to reasons such as long design and development cycles and high difficulties, the applications of ASICs are restricted. FPGAs are increasingly being used in the computing acceleration of deep learning, especially in device-side computing acceleration, due to their advantages such as rich parallel computing resources, low cost, low power consumption, and programmability. However, deploying deep learning algorithms on FPGAs usually requires developers to have high software and hardware development capabilities, which greatly hinders the use of FPGAs for deep learning computing acceleration. Although there are currently many studies on the computing acceleration of convolutional neural networks or recurrent neural networks based on FPGAs, with the continuous development of artificial intelligence, many new deep learning algorithm models have been proposed. Among them, the CNN-LSTM neural network model that combines a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM) neural network extracts local features using convolutional operations and then synthesizes features using a long short-term memory neural network, showing good performance when dealing with problems related to time series. However, a single FPGA acceleration scheme for CNN or LSTM is no longer applicable.
[0003] The Suzhou Institute of the University of Science and Technology of China disclosed a deep neural network acceleration platform based on FPGA in its patent document "FPGA-based Deep Neural Network Acceleration Platform" (Patent Application No.: CN201810010938.X, Authorization Publication No.: CN 108229670 B). This computing acceleration platform includes: a general-purpose processor, an FPGA, and a DRAM (Dynamic Random Access Memory). The general-purpose processor is used to parse the neural network configuration information and weight data, and write the neural network configuration information, weight data, and image data to be processed into the DRAM. Then, the FPGA, according to the configuration information in the DRAM, completes the CNN computing acceleration through the designed convolutional layer IP (Intellectual Property) core, pooling layer IP core, fully connected layer IP core, and activation layer IP core, and writes the computing results into the DRAM. Finally, the general-purpose processor reads the computing results from the DRAM. Although the algorithm deployment process of this acceleration platform is simplified, the shortcoming of this platform is that the designed convolutional layer IP core cannot extract the matrix multiplication operation result and does not have the matrix multiplication operation function. Therefore, it can only achieve CNN computing acceleration and cannot complete the matrix operation acceleration in LSTM.
[0004] He Junhua disclosed a method for CNN and LSTM neural network computing acceleration in his published thesis "Design and Implementation of a Deep Learning Computing Platform Based on FPGA" (Beijing University of Posts and Telecommunications, Master's Thesis, 2020). This method uses the Winograd algorithm to achieve fast convolution operations, the systolic array idea to achieve fast matrix multiplication, fixed-point arithmetic instead of floating-point arithmetic, and the technique of non-linear function lookup tables to optimize the FPGA hardware resource occupancy and system latency. However, the shortcoming of this method is that there are significant differences in the operation structures of the convolution operation module implemented by the Winograd algorithm and the matrix multiplication operation module implemented by the systolic array, and it is impossible to achieve computing resource reuse. When deploying the two neural network acceleration schemes to the FPGA at the same time, there will be a shortage of computing resources, which cannot meet the computing acceleration of the CNN-LSTM neural network model. Summary of the Invention
[0005] The purpose of the present invention is to address the deficiencies of the above-mentioned existing technologies and propose a one-dimensional CNN-LSTM resource-reusable computing acceleration platform and implementation method based on FPGA, which is used to solve the problem that the existing single acceleration scheme based on FPGA cannot support both convolution and matrix multiplication operation accelerations at the same time, and the problem of insufficient computing resources when deploying the two neural network acceleration schemes to the FPGA at the same time.
[0006] To achieve the above object, the idea of the present invention is to construct an arithmetic unit module composed of a linear arithmetic unit and a non-linear arithmetic unit on an FPGA. Among them, the multiply-accumulate array in the linear arithmetic unit is used to complete one-dimensional convolution operations, and the operation results are stored in the constructed ping-pong result cache. When it is necessary to complete the multiplication operation of vectors and matrices in the LSTM neural network, the vector to be calculated is used as the input of the one-dimensional convolution operation, the matrix to be calculated is divided into blocks by rows, and multiple convolution kernels of the one-dimensional convolution operation are loaded into the multiply-accumulate array, and then the operation is completed according to the method of the one-dimensional convolution operation. Finally, the operation results of the matrix multiplication are extracted from the ping-pong result cache. The non-linear arithmetic unit completes the activation of the gating coefficient and the update of the neuron state in the LSTM neural network by looking up a non-linear function table. Through the cooperation of the linear arithmetic unit and the non-linear arithmetic unit, the calculation acceleration of the one-dimensional CNN-LSTM neural network model is completed. The present invention multiplexes the same multiply-accumulate array in the linear arithmetic unit through one-dimensional convolution operations and matrix multiplication operations, greatly improving the utilization rate of the computing resources of the FPGA, and thus solving the problems that a single acceleration scheme cannot simultaneously complete the acceleration of both convolution and matrix multiplication operations and that there will be insufficient computing resources when two acceleration schemes are deployed on the FPGA at the same time.
[0007] The platform of the present invention includes two parts: a general-purpose CPU and an FPGA. The FPGA side also includes an instruction memory, a data memory, a result memory, a controller, and an arithmetic unit, where:
[0008] The general-purpose CPU is used to write the operation instruction sequence of the one-dimensional CNN-LSTM neural network model to be accelerated provided by the user and load it into the instruction memory on the FPGA side, quantize the parameters and input data of the one-dimensional CNN-LSTM neural network model into fixed-point numbers and load them into the data memory on the FPGA side, write an operation start instruction at the highest address of the instruction memory, and read the final operation result from the result memory after all operation instructions are executed.
[0009] The instruction memory is used to store the operation instruction sequence of the one-dimensional CNN-LSTM neural network model to be accelerated provided by the user written on the general-purpose CPU side.
[0010] The data memory is used to store the parameters and input data of the one-dimensional CNN-LSTM neural network model quantized on the general-purpose CPU side.
[0011] The result memory is used to store the final operation result for the general-purpose CPU to read.
[0012] The controller is used to monitor the operation start instruction, that is, monitor the address line of the instruction memory. When the CPU writes data to the highest address of the instruction memory, it means that the calculation start instruction of the general CPU is monitored. The data written to the highest address of the instruction memory is the total number of instructions to be executed in this calculation. After monitoring the operation start instruction, read an operation instruction from the instruction memory and send it to the instruction bus, and then monitor the execution feedback bus. After monitoring the execution completion signal, send the next operation instruction to the instruction bus until all instructions are executed, then clear the highest address of the instruction memory to 0 and re-enter the start instruction monitoring state.
[0013] The arithmetic unit includes a control unit, a linear arithmetic unit, and a non-linear arithmetic unit, and is used to execute the operation instructions sent to the instruction bus by the controller; the control unit is used to generate corresponding control information according to the instruction content, control the linear arithmetic unit and the non-linear arithmetic unit to complete corresponding operations, and send the instruction execution completion signal to the execution feedback bus after the operation instruction is executed; the linear arithmetic unit consists of a multiply-accumulate array and a result cache array with a ping-pong structure, and is used to load the weight parameter w and bias parameter bias of the multiply-accumulate array from the data memory according to the control information provided by the control unit, load the input data of the multiply-accumulate array from the data memory or the row of the ping-pong result cache array or the column of the ping-pong result cache array or the result cache of the non-linear arithmetic unit, and perform max pooling P operation, rectified linear unit R operation and channel sum operation on the operation result of the multiply-accumulate array according to the control information provided by the control unit, and then store it in the ping-pong result cache array to complete and accelerate one-dimensional convolution operation and matrix multiplication operation; the non-linear arithmetic unit intercepts the parts of the independent variables of the sigmoid and tanh non-linear functions between [-4, 4), quantizes their function values into fixed-point numbers and stores them in the ROM (Read-Only Memory). When performing non-linear activation on the matrix operation result in the LSTM neural network operation according to the control information provided by the control unit, convert the matrix operation result into the address of the ROM storing the non-linear function values, read the corresponding function values, and complete the update of the LSTM neural network gate coefficient and neuron state.
[0014] The specific steps of the acceleration platform implementation method of the present invention are as follows:
[0015] Step 1, write the operation instruction sequence of the one-dimensional CNN-LSTM neural network model to be accelerated provided by the user on the general CPU side and load it into the instruction memory.
[0016] Step 2, quantize the one-dimensional CNN-LSTM neural network model parameters and input data into fixed-point numbers on the general CPU side and load them into the data memory.
[0017] Step 3: Set the controller to the startup instruction listening state and monitor the address line of the instruction memory. When the CPU writes data to the highest address of the instruction memory, it indicates that the calculation startup instruction of the general CPU is monitored, and the data written to the highest address of the instruction memory is the total number of instructions to be executed in this calculation.
[0018] Step 4: The controller reads an instruction from the instruction memory, sends it to the instruction bus, and sets the controller to the execution feedback listening state to monitor the instruction execution completion signal sent by the arithmetic unit module to the execution feedback bus.
[0019] Step 5: Execute the arithmetic instruction according to the arithmetic instruction content:
[0020] Step 5.1: The control unit of the arithmetic unit module generates corresponding control information according to the instruction content.
[0021] Step 5.2: The linear arithmetic unit of the arithmetic unit module loads the weight parameter w and bias parameter bias of the multiplication-accumulation array from the data memory according to the control information provided by the control unit.
[0022] Step 5.3: The linear arithmetic unit of the arithmetic unit module loads the input data of the multiplication-accumulation array from the data memory, or a row of the ping-pong result cache array, or a column of the ping-pong result cache array, or the result cache of the non-linear arithmetic unit according to the control information provided by the control unit.
[0023] Step 5.4: The linear arithmetic unit performs the maximum pooling P operation, the rectified linear unit R operation, and the channel summation operation on the operation result of the multiplication-accumulation array according to the control information provided by the control unit, and stores it in the ping-pong result cache array.
[0024] The maximum pooling P operation is to find the maximum value of the operation results of n multiplication-accumulation arrays, where the value of n is equal to the pooling kernel length.
[0025] The rectified linear unit R operation is to compare the operation result of the multiplication-accumulation array with 0. When the operation result of the multiplication-accumulation array is greater than 0, its own value is taken; when the operation result of the multiplication-accumulation array is less than or equal to 0, its value is taken as 0.
[0026] Step 5.5: The non-linear arithmetic unit intercepts the parts of the sigmoid and tanh non-linear function independent variables within [-4, 4), quantizes their function values into fixed-point numbers and stores them in the ROM. When performing non-linear activation on the matrix operation result in the LSTM neural network operation according to the control information provided by the control unit, the matrix operation result is converted into the address of the ROM storing the non-linear function value, and the corresponding function value is read out to complete the update of the neuron state of the LSTM neural network.
[0027] Step 6, after the arithmetic unit completes an arithmetic instruction, it sends a signal indicating the completion of the instruction execution to the execution feedback bus through the control unit.
[0028] Step 7, after the controller monitors the signal indicating the completion of the instruction execution sent by the arithmetic unit to the execution feedback bus, it determines whether all the instructions of this operation have been executed. If so, it proceeds to Step 8; otherwise, it proceeds to Step 4.
[0029] Step 8, after all the instructions are executed, the arithmetic unit writes the final operation result into the result memory.
[0030] Step 9, through the controller module, clear the data in the highest address of the instruction memory to 0 and re-enter the startup instruction monitoring state.
[0031] Step 10, after the CPU detects that the data in the highest address of the instruction memory has been cleared to 0, it indicates that this operation is completed, and then reads the final operation result from the result memory.
[0032] Compared with the prior art, the present invention has the following advantages:
[0033] First, the linear operation unit in the arithmetic unit module implemented on the FPGA of the platform of the present invention converts the multiplication operation of a vector and a matrix into a one-dimensional convolution operation, realizing the parallel acceleration of both one-dimensional convolution and matrix multiplication operations, solving the deficiency that the existing single computing acceleration platform cannot simultaneously complete the acceleration of both convolution and matrix multiplication operations, enabling the platform of the present invention to support the computing acceleration of the one-dimensional CNN-LSTM neural network model and having a wider applicability.
[0034] Second, the implementation method of the platform of the present invention is to reuse the same multiply-accumulate array for one-dimensional convolution operation and matrix multiplication operation, saving the FPGA operation resources while ensuring the acceleration performance of the platform, solving the problem of insufficient computing resources when deploying two acceleration schemes on the FPGA simultaneously in the prior art, and enabling the present invention to greatly improve the utilization rate of the FPGA computing resources. Description of the Drawings
[0035] Figure 1 is the structural diagram of the present invention;
[0036] Figure 2 is the structural diagram of the arithmetic unit of the present invention;
[0037] Figure 3 is the structural diagram of the basic processing unit of the multiply-accumulate array of the present invention;
[0038] Figure 4 is the flowchart of the method of the present invention. Detailed Embodiments
[0039] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0040] Refer to Figure 1 for a further description of the platform structure of the present invention.
[0041] The present invention consists of a general-purpose CPU (Central Processing Unit) and an FPGA that perform data interaction through a high-speed serial bus PCIE (Peripheral Component Interconnect Express).
[0042] The FPGA part consists of an instruction memory IBUF (Instruction Buffer), a data memory DBUF (Data Buffer), a result memory RBUF (Result Buffer), a controller, and an arithmetic unit. The arithmetic unit performs data interaction with the data memory and the result memory respectively through a RAM (Random Access Memory) interface. The controller also performs data interaction with the instruction memory through a RAM interface. The controller and the arithmetic unit perform data interaction through an instruction bus and an execution feedback bus.
[0043] The general-purpose CPU is used to parse the one-dimensional CNN-LSTM neural network model provided by the user, generate an operation instruction sequence, and load it into the instruction memory at the FPGA end through the PCIE bus. Then, the one-dimensional CNN-LSTM neural network model parameters and the data to be processed provided by the user are quantized into fixed-point numbers and loaded into the data memory at the FPGA end through the PCIE bus. Finally, an operation start instruction is written to the highest address of the instruction memory through the PCIE bus, and after waiting for the arithmetic unit to execute all operation instructions, the operation result is read from the result memory through the PCIE bus.
[0044] The instruction memory is used to store the operation instruction sequence generated by the CPU. This memory is composed of a dual-port RAM implemented by the Block RAM resources of the FPGA.
[0045] The data memory is used to store the user model parameters and the data to be processed quantized by the CPU. This memory is composed of a dual-port RAM implemented by the Block RAM resources of the FPGA.
[0046] The result memory is used to store the final operation result of the arithmetic unit for the CPU to read. This memory is composed of a dual-port RAM implemented by the Block RAM resources of the FPGA.
[0047] The controller is used to complete the functions of listening for operation start instructions, reading instruction sequences, and processing execution feedback information. After the system is reset, the controller module is in the operation start instruction listening state, monitoring the address lines of the instruction memory. When it detects that the CPU writes data to the highest address of the instruction memory, it indicates that the operation start instruction has been listened for, and the data written by the CPU to the highest address of the instruction memory is used as the total number of instructions to be executed by the arithmetic unit during this calculation process. After listening for the start instruction, the controller reads an instruction from address 0 of the instruction memory and sends it to the instruction bus, then enters the execution feedback listening state. When the arithmetic unit module completes the instruction content, it sends a completion signal to the execution feedback bus. After the controller listens for the feedback signal, it sends the next instruction in the instruction memory to the instruction bus. Until all instructions of this operation are executed, the controller module clears the data in the highest address of the instruction memory and re-enters the operation start instruction listening state. When the CPU detects that the data in the highest address of the instruction memory has been cleared by the controller, it indicates that this calculation is completed.
[0048] Refer to the appendix Figure 2 for a further description of the structure of the arithmetic unit.
[0049] The arithmetic unit is used to execute the arithmetic instructions sent to the instruction bus by the controller and feedback the instruction execution status to the controller through the execution feedback bus. It consists of three parts: a control unit, a linear arithmetic unit, and a non-linear arithmetic unit.
[0050] The control unit is used to listen for and parse the arithmetic instructions on the instruction bus, then generate control information to control the linear arithmetic unit and the non-linear arithmetic unit to complete the corresponding calculation tasks, and finally send the instruction execution completion signal to the execution feedback bus.
[0051] The linear arithmetic unit consists of a ping-pong result cache array CBUF (Channel Buffer) and a multiply-accumulate array. By multiplexing the multiply-accumulate array, parallel acceleration of one-dimensional convolution and matrix multiplication operations is achieved.
[0052] In the embodiment of the present invention, CBUF is a ping-pong cache composed of two 32-row and 10240-column storage arrays CBUF0 and CBUF1 implemented by the BRAM resources of the PFGA. That is, when CBUF0 is used as the data loading source, CBUF1 is used as the result memory, and vice versa, when CBUF1 is used as the data loading source, CBUF0 is used as the result memory.
[0053] Refer to the appendix Figure 3 for a further description of the basic processing unit structure of the multiply-accumulate array.
[0054] The basic processing element Pe (Processing element) of the multiply-accumulate array consists of a weight buffer unit WBUF (Weight Buffer), a data input port, a bias input port, and a result output port. Among them, WBUF consists of four independent registers, which are used to store four different weight parameters. During the operation, one of them is selected as the effective weight parameter win according to the control information provided by the control unit. The input data xin at the data input port is multiplied by the weight parameter win and then added to the input data bin at the bias input port, and the calculation result yout is output from the result output port.
[0055] In the embodiment of the present invention, the size of the multiply-accumulate array is 32 rows and 64 columns. According to the control information provided by the control unit, the input data x of the multiply-accumulate array is loaded from the row of DBUF or CBUF or the column of CBUF or the result buffer NBUF (Nonlinear results Buffer) of the nonlinear operation unit. The weight parameter w and the bias parameter bias of the multiply-accumulate array are loaded from DBUF. Each bias buffer unit bi (i = 0, 1,..., 31) of the multiply-accumulate array has four independent registers, and one of them is selected as the effective parameter for use according to the control information provided by the control unit during the operation. According to the control information provided by the control unit, the operation results of the multiply-accumulate array are subjected to maximum pooling P operation, rectified linear unit R operation, and channel sum operation, and then stored in CBUF.
[0056] The nonlinear operation unit intercepts the parts of the independent variables of the sigmoid and tanh non-linear functions between [-4, 4), quantizes their function values into fixed-point numbers and stores them in the ROM. When it is necessary to perform non-linear activation on the matrix operation results in the LSTM neural network operation, the matrix operation results are converted into ROM addresses to read the corresponding function values.
[0057] Refer to the appendix Figure 4 , and a further description is made of the implementation method of the platform of the present invention.
[0058] Step 1. Write the operation instruction sequence of the one-dimensional CNN-LSTM neural network model to be accelerated provided by the user on the CPU side, and load it into the instruction memory through the PCIE bus.
[0059] Step 2. Quantize the one-dimensional CNN-LSTM neural network model parameters and input data into fixed-point numbers, and load them into the data memory through the PCIE bus.
[0060] In the embodiment of the present invention, all the data participating in the operation are quantized into 16-bit fixed-point numbers, where 1 bit is the sign bit, 4 bits are the integer bits, and 11 bits are the fractional bits. The multiply-accumulate operation results are also processed into 16-bit fixed-point numbers by using the method of saturation truncation.
[0061] Step 3: Set the controller to the start instruction listening state, and monitor the address line of the instruction memory. When the CPU writes data to the highest address of the instruction memory, it indicates that the calculation start instruction of the general CPU is received, and the data written to the highest address of the instruction memory is the total number of instructions to be executed in this calculation.
[0062] Step 4: The controller reads an instruction from the instruction memory and sends it to the instruction bus, and sets the controller module to the execution feedback listening state to monitor the instruction execution completion signal sent by the arithmetic unit module to the execution feedback bus.
[0063] Step 5: Execute the arithmetic instruction according to the content of the arithmetic instruction.
[0064] Step 5.1: The control unit of the arithmetic unit module generates corresponding control information according to the instruction content.
[0065] Step 5.2: The linear arithmetic unit of the arithmetic unit module loads the weight parameter w and bias parameter bias of the multiplication-accumulation array from the DBUF according to the control information provided by the control unit.
[0066] Step 5.3: The linear arithmetic unit of the arithmetic unit module loads the input data x of the multiplication-accumulation array from the DBUF, the rows of the CBUF, the columns of the CBUF, or the result cache NBUF of the non-linear arithmetic unit according to the control information provided by the control unit.
[0067] Step 5.4: The linear arithmetic unit performs a maximum pooling P operation, a rectified linear R operation, and a channel summation operation on the operation result of the multiplication-accumulation array according to the control information provided by the control unit, and stores the result in the CBUF.
[0068] The maximum pooling P operation is to find the maximum value of the operation results of n multiplication-accumulation arrays, where the value of n is equal to the pooling kernel length.
[0069] The rectified linear R operation is to compare the operation result of the multiplication-accumulation array with 0. When the operation result of the multiplication-accumulation array is greater than 0, its own value is taken; when the operation result of the multiplication-accumulation array is less than or equal to 0, its value is taken as 0.
[0070] Step 5.5: The non-linear arithmetic unit intercepts the parts of the independent variables of the sigmoid and tanh non-linear functions within the range of [-4, 4), quantizes their function values into fixed-point numbers and stores them in the ROM. When performing non-linear activation on the matrix operation result in the LSTM neural network operation according to the control information provided by the control unit, the matrix operation result is converted into the address of the ROM storing the non-linear function values, and the corresponding function values are read out to complete the update of the neuron state of the LSTM neural network.
[0071] The LSTM neural network operation is defined by the following formula:
[0072] f t = sigmoid(W f * [h t-1 , x t + b f );
[0073] i t = sigmoid(W i * [h t-1 , x t + b i );
[0074] o t = sigmoid(W o * [h t-1 , x t + b o );
[0075]
[0076]
[0077] h t = o t * C t
[0078] where f t represents the forgetting gate coefficient vector of the LSTM neural network, i t represents the input gate coefficient vector of the LSTM neural network, o t represents the output gate coefficient vector of the LSTM neural network, represents the unmerged neuron state vector in the LSTM neural network, C t represents the neuron state vector at the current moment of the LSTM neural network, C t-1 represents the neuron state vector at the previous moment of the LSTM neural network, h t represents the hidden layer state vector at the current moment of the LSTM neural network, h t-1 represents the hidden layer state vector at the previous moment of the LSTM neural network, x t represents the input vector of the LSTM neural network, W f 、W i 、W o and W c are the weight matrices of f t 、i t 、o t and of the LSTM neural network model provided by the user respectively, b f, b i , b o and b c are the bias parameters of f t , i t , o t and in the LSTM neural network model provided by the user respectively.
[0079] The sigmoid function is defined by the following formula:
[0080]
[0081] where, e (·) represents the exponential operation with the natural constant e as the base, and x represents the elements in the gating coefficient vector f t , i t , o t to be activated in the LSTM operation.
[0082] The tanh function is defined by the following formula:
[0083]
[0084] where, c represents the elements in the unfused neuron state vector to be activated in the LSTM operation.
[0085] Step 6: After the arithmetic unit module completes an arithmetic instruction, it sends the execution completion signal to the execution feedback bus through the control unit.
[0086] Step 7: After detecting the instruction execution completion signal sent by the arithmetic unit module to the execution feedback bus, it judges whether all instructions of this operation have been executed. If so, it executes Step 8; otherwise, it executes Step 4.
[0087] Step 8: After all instructions are executed, the arithmetic unit writes the final operation result into the result memory
[0088] Step 9: The controller clears the data in the highest address of the instruction memory to 0 and re-enters the start instruction listening state;
[0089] Step 10: After the CPU detects that the data in the highest address of the instruction memory has been cleared to 0, it indicates that this operation is completed, and then reads the final operation result from the result memory RBUF.
Claims
1. The one-dimensional CNN-LSTM acceleration platform based on FPGA includes two parts: a general-purpose CPU and an FPGA. The FPGA side includes an instruction memory, a data memory, a result memory, a controller, and an arithmetic unit, where: The general-purpose CPU is used to write the operation instruction sequence of the one-dimensional CNN-LSTM neural network model to be accelerated provided by the user and load it into the instruction memory on the FPGA side, quantize the one-dimensional CNN-LSTM neural network model parameters and input data into fixed-point numbers and load them into the data memory on the FPGA side, write an operation start instruction to the highest address of the instruction memory, and read the final operation result from the result memory after all operation instructions are executed; The instruction memory is used to store the operation instruction sequence of the one-dimensional CNN-LSTM neural network model to be accelerated provided by the user written on the general-purpose CPU side; The data memory is used to store the parameters and input data of the one-dimensional CNN-LSTM neural network model quantized on the general-purpose CPU side; The result memory is used to store the final operation result for the general-purpose CPU to read; The controller is used to monitor the operation start instruction, that is, monitor the address line of the instruction memory. When the CPU writes data to the highest address of the instruction memory, it means that the calculation start instruction of the general-purpose CPU is monitored. The data written to the highest address of the instruction memory is the total number of instructions to be executed for this calculation. After monitoring the operation start instruction, read an operation instruction from the instruction memory and send it to the instruction bus, then monitor the execution feedback bus, and send the next operation instruction to the instruction bus after monitoring the execution completion signal until all instructions are executed, then clear the highest address of the instruction memory to 0 and re-enter the start instruction monitoring state; The arithmetic unit includes a control unit, a linear arithmetic unit, and a non-linear arithmetic unit, and is used to execute arithmetic instructions sent to the instruction bus by the controller; the control unit is used to generate corresponding control information according to the instruction content, control the linear arithmetic unit and the non-linear arithmetic unit to complete corresponding arithmetic operations, and send an instruction execution completion signal to the execution feedback bus after the arithmetic instruction is executed; the linear arithmetic unit consists of a multiply-accumulate array and a result cache array with a ping-pong structure, and is used to load the weight parameter w and bias parameter bias of the multiply-accumulate array from the data memory according to the control information provided by the control unit, load the input data of the multiply-accumulate array from the data memory or the row of the ping-pong result cache array or the column of the ping-pong result cache array or the result cache of the non-linear arithmetic unit, and perform max pooling P operation, rectified linear unit R operation and channel summation operation on the arithmetic result of the multiply-accumulate array according to the control information provided by the control unit, and then store it in the ping-pong result cache array to complete the one-dimensional convolution operation and matrix multiplication operation and accelerate them; the non-linear arithmetic unit intercepts the parts of the independent variables of the sigmoid and tanh non-linear functions within [-4, 4), quantizes their function values into fixed-point numbers and stores them in the ROM. When performing non-linear activation on the matrix operation result in the LSTM neural network operation according to the control information provided by the control unit, it converts the matrix operation result into the address of the ROM storing the non-linear function values, reads out the corresponding function values, and completes the update of the LSTM neural network gating coefficient and neuron state.
2. A method for implementing a one-dimensional CNN-LSTM acceleration platform based on FPGA for the platform according to claim 1, characterized in that, The linear arithmetic unit of the arithmetic unit accelerates the parallel operation of one-dimensional convolution and matrix multiplication by multiplexing the same multiply-accumulate array. The specific steps of this method are as follows: Step 1, write the arithmetic instruction sequence of the one-dimensional CNN-LSTM neural network model to be accelerated provided by the user on the general CPU side and load it into the instruction memory; Step 2, quantize the parameters and input data of the one-dimensional CNN-LSTM neural network model into fixed-point numbers on the general CPU side and load them into the data memory; Step 3, set the controller to the instruction listening state, monitor the address line of the instruction memory. When the CPU writes data to the highest address of the instruction memory, it means that the calculation start instruction of the general CPU is monitored, and the data written to the highest address of the instruction memory is the total number of instructions to be executed in this calculation; Step 4, the controller reads an instruction from the instruction memory and sends it to the instruction bus, and sets the controller to the execution feedback listening state, monitoring the instruction execution completion signal sent by the arithmetic unit module to the execution feedback bus; Step 5, execute the arithmetic instruction according to the arithmetic instruction content: Step 5.1, the control unit of the arithmetic unit module generates corresponding control information according to the instruction content; Step 5.2, the linear arithmetic unit of the arithmetic unit module loads the weight parameter w and bias parameter bias of the multiply-accumulate array from the data memory according to the control information provided by the control unit; Step 5.3, the linear operation unit of the arithmetic unit loads the input data of the multiplication and accumulation array from the data memory, or a row of the ping-pong result cache array, or a column of the ping-pong result cache array, or the result cache of the non-linear operation unit according to the control information provided by the control unit; Step 5.4, the linear operation unit performs max pooling P operation, rectified linear unit R operation, and channel sum operation on the operation results of each multiplication and accumulation array according to the control information provided by the control unit, and then stores the results in the ping-pong result cache array; Step 5.5, the non-linear operation unit intercepts the parts of the independent variables of the sigmoid and tanh non-linear functions within the range of [-4, 4), quantizes their function values into fixed-point numbers and stores them in the ROM. When performing non-linear activation on the matrix operation results in the LSTM neural network operation according to the control information provided by the control unit, the matrix operation results are converted into the addresses of the ROM storing the non-linear function values, and the corresponding function values are read out to complete the update of the neuron state of the LSTM neural network; Step 6, after the arithmetic unit completes an operation instruction, it sends a signal indicating the completion of the instruction execution to the execution feedback bus through the control unit; Step 7, after the controller monitors the signal indicating the completion of the instruction execution sent by the arithmetic unit to the execution feedback bus, it determines whether all the instructions of the current operation have been executed. If so, it proceeds to Step 8; otherwise, it proceeds to Step 4; Step 8, after all the instructions have been executed, the arithmetic unit writes the final operation result into the result memory; Step 9, the controller clears the data in the highest address of the instruction memory to 0 and re-enters the startup instruction monitoring state; Step 10, after the CPU detects that the data in the highest address of the instruction memory has been cleared to 0, it indicates that the current operation is completed, and then reads the final operation result from the result memory.
3. The implementation method of the one-dimensional CNN-LSTM acceleration platform based on FPGA according to claim 2, wherein, The max pooling P operation described in Step 5.4 refers to finding the maximum value among the operation results of n multiplication and accumulation arrays, where the value of n is equal to the length of the pooling kernel.
4. The method for implementing a one-dimensional CNN-LSTM acceleration platform based on FPGA according to claim 2, wherein The rectified linear unit R operation described in Step 5.4 refers to comparing the size of each operation result in the multiplication and accumulation array with 0. When the operation result of the multiplication and accumulation array is greater than 0, its own value is taken; when the operation result of the multiplication and accumulation array is less than or equal to 0, its value is taken as 0.
5. The implementation method of the one-dimensional CNN-LSTM acceleration platform based on FPGA according to claim 2, characterized in that, The LSTM neural network operation described in Step 5.5 is completed by the following formula: f t = sigmoid(W f * [h t-1 , x t + b f ); i t = sigmoid(W i * [h t-1 , x t + b i ); o t = sigmoid(W o * [h t-1 , x t + b o ); h t = o t * C t Among them, f t represents the forgetting gate coefficient of the LSTM neural network, i t represents the input gate coefficient vector of the LSTM neural network, o t represents the output gate coefficient vector of the LSTM neural network, represents the un-fused neuron state vector in the LSTM neural network, C t represents the neuron state vector at the current moment of the LSTM neural network, C t-1 represents the neuron state vector at the previous moment of the current moment of the LSTM neural network, h t represents the hidden layer state vector at the current moment of the LSTM neural network, h t-1 represents the hidden layer state vector at the previous moment of the current moment of the LSTM neural network, x t represents the input vector of the LSTM neural network, W f 、W i 、W o and W c are respectively the weight matrices of f t 、i t 、o t and of the LSTM neural network model provided by the user, b f 、b i 、b o and b c are respectively the bias parameters of f t 、i t 、o t and of the LSTM neural network model provided by the user.
6. The method for implementing a one-dimensional CNN-LSTM acceleration platform based on FPGA according to claim 2, characterized in that The sigmoid function described in Step 5.5 is defined by the following formula: Among them, e (·) represents the exponential operation with the natural constant e as the base, and x represents the gating coefficient vector f t , i t , o t in the elements of 7. The method for implementing a one-dimensional CNN-LSTM acceleration platform based on FPGA according to claim 1, wherein The tanh function described in Step 5.5 is defined by the following formula: Among them, c represents the element of the un-fused neuron state vector to be activated in the LSTM operation. in.
Citation Information
Patent Citations
Deep neural network acceleration platform based on FPGA
CN108229670A
FPGA-based deep neural network acceleration platform
CN108229670B
Design method of hardware accelerator based on LSTM recursive neural network algorithm on FPGA platform
CN108090560A
Efficient LSTM accelerator based on FPGA
CN113191494A