A General Lightweight Convolutional Neural Network Acceleration Method Based on FPGA
By dividing the convolutional neural network model into a combination of basic operators, configuring the IP core of FPGA to achieve different computing functions, solving the problems of accelerator flexibility and insufficient resource consumption in the existing technology, improving computing efficiency and reducing power consumption.
Patent Information
- Application Number
- CN202210473456.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The existing FPGA-based convolutional neural network accelerators have insufficient computing flexibility and resource consumption, and cannot meet the needs of multiple computing forms.
The convolutional neural network model is divided into a combination of basic operators, and the basic operators in the target IP core are configured through the target input data and model structure parameters, different calculation functions are realized, and data handling and calculation are carried out through the target functional operator until the overall calculation is completed.
It improves the scalability and flexibility of the convolutional neural network accelerator, reduces the overall topology complexity and hardware resource consumption, improves the computing efficiency and reduces power consumption.
Smart Images

Figure CN114936636B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer hardware acceleration, and particularly relates to a general lightweight convolutional neural network acceleration method based on FPGA. Background Art
[0002] In the field of computer vision, the method based on convolutional neural network has very high performance, but it requires more computing and storage resources compared with traditional methods. Embedded terminal devices do not have high-performance computing and storage resources due to cost, power consumption and other limiting factors, and the effect is not ideal when executing deep learning algorithms. FPGA (Field-Programmable Gate Array) has the advantages of low power consumption, good flexibility, strong parallel computing ability, etc., and has become a major customized platform for neural network inference computing.
[0003] In the prior art, most of the convolutional neural network accelerators based on FPGA are designed in the form of dedicated circuits, that is, the hardware circuit is designed for a specific neural network model. Such accelerators have good computing acceleration effects on specific neural networks, but they have poor scalability, poor flexibility, and high hardware resource consumption. Some convolutional neural network accelerators based on FPGA perform configurable design of the hardware circuit for convolutional computing, making them have a certain degree of computing flexibility, but their accelerator functions are relatively single and cannot meet the computing requirements of various computing forms such as pooling, fully connected, and feature addition. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems in the above background art, and propose a general lightweight convolutional neural network acceleration method based on FPGA.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] An embodiment of the present invention provides a general lightweight convolutional neural network acceleration method based on FPGA, and the method includes:
[0007] Obtain the current target input data and model structure parameters of the target convolutional neural network;
[0008] Configure the currently corresponding basic operator in the target IP core according to the target input data and the model structure parameters as the target functional operator; the target IP core includes all basic operators of the target convolutional neural network, and different basic operators implement different computing functions;
[0009] Read the target input data through a target functional operator, calculate the target output data, and transfer the target output data according to the model structure parameters and the capacity of the storage unit;
[0010] Repeat the above steps according to the model structure parameters until the calculation of the entire target convolutional neural network is completed.
[0011] Optionally, the input data of the target convolutional neural network includes original input data;
[0012] Before obtaining the current target input data and model structure parameters of the target convolutional neural network, the method further includes:
[0013] Obtain the original acquisition data of the target convolutional neural network; the original acquisition data includes the image to be processed and the model function parameters of the target convolutional neural network;
[0014] Perform low-bit fixed-point quantization on the original acquisition data to obtain the original input data, and store it in the input data BRAM.
[0015] Optionally, the model function parameters include at least one of convolutional kernel parameters, convolutional kernel bias parameters, fully connected layer weight parameters, and fully connected layer bias parameters.
[0016] Optionally, the input data of the target convolutional neural network further includes intermediate calculation data;
[0017] Reading the target input data through a target functional operator, calculating the target output data, and transferring the target output data according to the model structure parameters and the capacity of the storage unit includes:
[0018] Read the target input data through a target functional operator, calculate the target output data, and store it in the output data BRAM;
[0019] Determine the data type of the target output data according to the model structure parameters; the data type includes the first data that will participate in the next calculation and the second data that will not participate in the next calculation;
[0020] Store the first data as the intermediate calculation data for the next calculation in the input data BRAM, and store the second data according to the capacity of the input data BRAM.
[0021] Optionally, the input data BRAM and the output data BRAM are dual-port BRAMs, and the ping-pong operation method is adopted; while the target functional operator is performing a calculation operation, the input data BRAM and the output data BRAM perform data transfer.
[0022] Optionally, the model structure parameters include all the basic operators of the target convolutional neural network and the connection order of each basic operator.
[0023] Optionally, the target IP core includes at least one basic operator among a convolution calculation unit, a pooling padding calculation unit, a feature addition calculation unit, a fully connected calculation unit, and an activation calculation unit.
[0024] Optionally, configuring the currently corresponding basic operator in the target IP core according to the target input data and the model structure parameters as the target functional operator includes:
[0025] Determining the currently corresponding basic operator in the target IP core according to the model structure parameters as the target basic operator;
[0026] Determining the registers to be configured for the target basic operator according to the type of the target basic operator as the target registers;
[0027] Configuring the target registers according to the target input data to obtain the target functional operator.
[0028] Optionally, the target registers include at least one of a feature map start address register, a model function parameter start address register, a result output start address register, a feature map row number register, a feature map column number register, an activation enable register, a padding enable register, a feature map quantity register, a feature map size register, a fully connected layer input column vector length register, and a mode selection register.
[0029] A general lightweight convolutional neural network acceleration method based on FPGA, which obtains the current target input data and model structure parameters of the target convolutional neural network; configures the currently corresponding basic operator in the target IP core according to the target input data and the model structure parameters as the target functional operator; the target IP core includes all the basic operators of the target convolutional neural network, and different basic operators implement different calculation functions; reads the target input data through the target functional operator, calculates the target output data, and transports the target output data according to the model structure parameters and the capacity of the storage unit; sequentially repeats the above steps according to the model structure parameters until the overall calculation of the target convolutional neural network is completed. By dividing the convolutional neural network model into a combination of basic operators, configuring each basic operator according to the target input data and the type of the target basic operator, and sequentially completing the calculation of each target functional operator, the overall calculation of the target convolutional neural network is completed. The scalability and flexibility of the convolutional neural network accelerator are improved. Description of the Drawings
[0030] The present invention will be further described below with reference to the accompanying drawings.
[0031] Figure 1 Flowchart of a general lightweight convolutional neural network acceleration method based on FPGA provided by an embodiment of the present invention;
[0032] Figure 2 System block diagram of a convolutional neural network accelerator system provided by an embodiment of the present invention;
[0033] Figure 3 Schematic structural diagram of a convolutional neural network accelerator provided by an embodiment of the present invention. Specific embodiments
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0035] An embodiment of the present invention provides a general lightweight convolutional neural network acceleration method based on FPGA. Refer to Figure 1 , Figure 1 which is the flowchart of a general lightweight convolutional neural network acceleration method based on FPGA provided by an embodiment of the present invention. The method may include the following steps:
[0036] S101, Obtain the current target input data and model structure parameters of the target convolutional neural network.
[0037] S102, Configure the current corresponding basic operator in the target IP core according to the target input data and model structure parameters as the target functional operator.
[0038] S103, Read the target input data through the target functional operator, calculate the target output data, and transfer the target output data according to the model structure parameters and the capacity of the storage unit.
[0039] S104, Repeat the above steps in sequence according to the model structure parameters until the overall calculation of the target convolutional neural network is completed.
[0040] The target IP core includes all basic operators of the target convolutional neural network, and different basic operators implement different calculation functions.
[0041] Based on a general lightweight convolutional neural network acceleration method based on FPGA provided by an embodiment of the present invention, by dividing the convolutional neural network model into a combination of basic operators, configuring each basic operator according to the type of target input data and target basic operator, and sequentially completing the calculation of each target functional operator, thereby completing the overall calculation of the target convolutional neural network. The scalability and flexibility of the convolutional neural network accelerator are improved.
[0042] In one implementation, the target convolutional neural network can be composed of a combination of multiple basic operators. The multiple basic operators can be encapsulated in the same IP core, and each computing unit is scheduled through time-division multiplexing to complete the corresponding computing function, which can reduce the overall topological complexity of the system and the consumption of bus matrix resources.
[0043] In one embodiment, the model structure parameters include all the basic operators of the target convolutional neural network and the connection order of each basic operator.
[0044] In one implementation, the network structure of the target convolutional neural network can be determined through the model structure parameters.
[0045] In one embodiment, the input data of the target convolutional neural network includes the original input data.
[0046] Before step S101, the method further includes:
[0047] Step 1, obtain the original acquisition data of the target convolutional neural network.
[0048] Step 2, perform low-bit fixed-point quantization on the original acquisition data to obtain the original input data, and store it in the input data BRAM.
[0049] The original acquisition data includes the image to be processed and the model function parameters of the target convolutional neural network.
[0050] In one implementation, the model function parameters can be the deep learning model parameters trained under deep learning frameworks such as Pytorch.
[0051] In one implementation, the trained model parameters with a floating-point precision of 32 bits are quantized in fixed-point format, and the image to be processed obtained by a camera or other means is quantized in fixed-point format in the same way. Since the neural network has good robustness and is not highly sensitive to data precision, the final inference calculation result will not be greatly affected after fixed-point quantization. In addition, implementing 32-bit floating-point multiplication and addition operations on FPGA consumes more hardware computing resources and storage resources. By performing low-bit fixed-point quantization operations on the model parameters and image data, the consumption of hardware computing resources and storage resources is greatly reduced, and the computing rate is improved.
[0052] In one implementation, the original input data includes the image to be processed after fixed-point quantization (hereinafter referred to as the target image to be processed) and the model function parameters (hereinafter referred to as the target model function parameters).
[0053] In one embodiment, the model function parameters include at least one of the convolution kernel parameters, the convolution kernel bias parameters, the fully connected layer weight parameters, and the fully connected layer bias parameters.
[0054] In one implementation, by changing the convolution kernel parameters, the convolution kernel bias parameters, the fully connected layer weight parameters, and the fully connected layer bias parameters, the same convolutional neural network structure can be used to meet different application requirements.
[0055] In one embodiment, the input data of the target convolutional neural network further includes intermediate calculation data.
[0056] Step S104 includes:
[0057] Step 1: Read the target input data through the target functional operator, calculate the target output data, and store it in the output data BRAM.
[0058] Step 2: Determine the data type of the target output data according to the model structure parameters.
[0059] Step 3: Store the first data as the intermediate calculation data for the next calculation in the input data BRAM, and store the second data according to the capacity of the input data BRAM.
[0060] The data type includes the first data that will participate in the next calculation and the second data that will not participate in the next calculation.
[0061] In one implementation, the main influencing factor for data transfer is the storage capacity of the input data BRAM. If the storage capacity is sufficient, both the first data and the second data can be stored in the input data BRAM. If the storage capacity is limited, the first data needs to be stored first, and the remaining data is transferred to off-chip storage while ensuring the maximum utilization of the storage resources of the input data BRAM. Since the transfer of data between the FPGA chip and the external storage chip will cause a large delay and power consumption, through the above operations, unnecessary data interaction between the on-chip and off-chip can be reduced, the execution efficiency of the target functional operator can be improved, and the power consumption can be reduced.
[0062] In one embodiment, the input data BRAM and the output data BRAM are dual-port BRAMs, and the ping-pong operation method is adopted; while the target functional operator is performing the calculation operation, the input data BRAM and the output data BRAM perform data transfer.
[0063] In one implementation, ping-pong operation can be adopted to achieve parallel execution of computing operations and data transfer, thereby enhancing the execution efficiency of the target convolutional neural network.
[0064] In one embodiment, refer to Figure 2 , Figure 2 which is the system block diagram of the convolutional neural network accelerator system provided by the embodiment of the present invention. The system includes a general-purpose lightweight convolutional neural network accelerator, dual-port input data BRAMs (BRAM1 and BRAM2), an output data BRAM (BRAM3), and a bus interface.
[0065] In one implementation, the accelerator is directly connected to the dual-port BRAM, and the accelerator can complete the read and write operations of the data in the on-chip BRAM storage space at high speed. The target image to be processed is written into the input BRAM1; the convolutional kernel parameters, convolutional kernel bias parameters, fully connected layer weight parameters, and fully connected layer bias parameters are stored in the input data BRAM2. The data stored in BRAM1 and BRAM2 serves as the data to be calculated. After the accelerator completes one data read, the data of the target image to be processed and the input column vector data of the fully connected layer in BRAM1 will not be lost. Considering that a target image to be processed will be convolved with multiple convolutional kernels, and an input column vector data of the fully connected layer will be multiplied by multiple sets of fully connected layer weight data. For the data of the target image to be processed and the input column vector data of the fully connected layer, it can be directly read from BRAM1 during the next calculation without having to read data from the off-chip memory of the FPGA every time. Compared with reading data from the off-chip memory every time for convolutional calculation and fully connected calculation, this method effectively reduces the data transfer time of the data to be calculated and reduces the energy consumption of the system. Among them, each computing unit of the accelerator reads the data to be calculated by controlling the BRAM1_PORT and BRAM2_PORT interfaces between the accelerator and BRAM1 and BRAM2, and writes out the calculation result data by controlling the BRAM3_PORT interface between the accelerator and BRAM3. The bus interface can be a system bus such as AXI or AHB, and the relevant registers of the accelerator can be configured through the bus interface.
[0066] In one embodiment, the target IP core includes at least one basic operator among a convolutional calculation unit, a pooling padding calculation unit, a feature addition calculation unit, a fully connected calculation unit, and an activation calculation unit.
[0067] In one implementation, a convolutional neural network model is divided into a combination of basic operators (convolution, pooling, feature addition, activation, fully connected, padding). A configurable general-purpose hardware acceleration computing unit for basic operators is designed. The configurable hardware acceleration computing unit completes the corresponding operator's computing task according to the register information in step 103 above. The same hardware acceleration computing unit can perform corresponding computing processes on input data of different sizes according to different register information, that is, there is no need to design multiple computing units of the same type to adapt to different size computing requirements, greatly reducing the consumption of hardware resources.
[0068] The convolution computing unit, pooling and padding computing unit, feature addition computing unit, activation computing unit, and fully connected computing unit all adopt a pipeline architecture design, enabling the reading, computing, and output of data in each computing unit to be carried out simultaneously, improving the computing efficiency of each computing unit.
[0069] The design idea of the convolution computing unit includes: reading the convolution kernel parameters to be calculated into the convolution computing unit and temporarily storing them in the registers of the convolution computing unit; reading the input feature map data in a sliding window order and temporarily storing the read feature map data in the registers of the convolution computing unit. The convolution kernel parameters read into the convolution computing unit are convolved with the feature map data in the feature map sliding window. Considering the reading rate of the data to be calculated and the output rate of the calculation results comprehensively, structural optimization is carried out. Specifically:
[0070] If the factor restricting the overall convolution computing rate is the input data reading rate, then increase the number of convolution kernels calculated by the convolution computing unit at one time, and perform convolution calculations for different convolution kernels through time division multiplexing to improve the utilization rate of multiplier resources and increase the overall computing rate.
[0071] If the factor restricting the overall convolution computing rate is the output data writing rate, then reduce the number of convolution kernels calculated by the convolution computing unit at one time to reduce unnecessary consumption of hardware resources.
[0072] By adopting the method of sliding window data reading and computing to reduce the cache space requirement of data in the computing unit, the resource consumption is reduced. Considering factors such as the feature map data reading rate, result data output rate, and multiplier resources of the accelerator comprehensively, the convolution computing unit is structurally optimized to achieve the efficient operation of the overall convolution computing and reduce the resource consumption of hardware computing resources.
[0073] The design idea of the pooling padding calculation unit includes: Although the padding operation has a small amount of calculation, it generates a large amount of data interaction and storage requirements, increasing the overall inference calculation time of the system. Since most of the operations before the padding operation are pooling calculations, the pooling calculation and the padding calculation can be fused for inter-layer calculation, reducing the interaction of the feature map data between the FPGA chip and the off-chip memory, thereby reducing the calculation time and power consumption. According to the padding enable register information, if a padding operation is required, the pooling padding calculation unit performs corresponding padding operations before, after, and during the pooling calculation with reference to the register information in step 3. According to the padding enable register information, if no padding operation is required, the pooling padding calculation unit only performs the pooling calculation operation. The pooling calculation reads and calculates data in the form of a sliding window, reducing the consumption of hardware resources.
[0074] The design idea of the feature addition calculation unit includes: The state machine of the feature addition calculation unit sequentially reads the data at the corresponding positions of each input feature map, completes the addition calculation, and writes the calculation result into the output data BRAM3. After completing the addition calculation of the data at the same position of the feature map, the data at the subsequent positions of each input feature map is sequentially read and calculated until the addition calculation of the data at all positions is completed.
[0075] The design idea of the fully connected calculation unit includes: The fully connected calculation unit reads the input column vector data of the fully connected layer and the data stream of the model weight parameters of the fully connected layer from BRAM1 and BRAM2, and realizes the multiplication and addition operations through a pipeline architecture inside the fully connected calculation unit, reducing the consumption of hardware resources.
[0076] The design idea of the activation unit includes: The activation calculation unit is directly connected to the output data. If an activation operation needs to be performed on the output data, the accelerator can perform the activation calculation on the output data in a pipeline calculation manner, saving the transfer and calculation time for the data to be separately activated. If no activation calculation is required, the activation unit does not process the output data.
[0077] In one embodiment, refer to Figure 3 , Figure 3Schematic diagram of the structure of a convolutional neural network accelerator provided by an embodiment of the present invention. The convolutional neural network accelerator includes a convolutional computing unit, a pooling padding computing unit (pooling + padding computing unit), a feature addition computing unit, a fully connected computing unit, an activation computing unit, and three path selectors. The convolutional computing unit, the pooling padding computing unit, the feature addition computing unit, and the fully connected computing unit are connected to the interface BRAM1_PORT of BRAM1 through the path selector, and read the target image to be processed in BRAM1 through BRAM1_PORT. The convolutional computing unit and the fully connected computing unit are connected to the interface BRAM2_PORT of BRAM2 through the path selector, and read the convolutional kernel parameters, convolutional kernel bias parameters, fully connected layer weight parameters, and fully connected layer bias parameters in BRAM2 through BRAM2_PORT. The convolutional computing unit, the pooling padding computing unit, the feature addition computing unit, and the fully connected computing unit are connected to the activation computing unit through the path selector, and the activation computing unit is connected to the interface BRAM3_PORT of BRAM3, and writes the calculation result into BRAM3 through BRAM3_PORT.
[0078] In one embodiment, step S103 includes:
[0079] Step 1, determine the current corresponding basic operator in the target IP core according to the model structure parameters as the target basic operator.
[0080] Step 2, determine the registers that need to be configured for the target basic operator according to the type of the target basic operator as the target registers.
[0081] Step 3, configure the target registers according to the target input data to obtain the target functional operator.
[0082] In one embodiment, the target registers include at least one of a feature map start address register, a model function parameter start address register, a result output start address register, a feature map row number register, a feature map column number register, an activation enable register, a padding enable register, a feature map quantity register, a feature map size register, a fully connected layer input column vector length register, and a mode selection register.
[0083] In one implementation, according to the functions of the basic operators, the number and types of registers required for different basic operators are different.
[0084] In one implementation, the starting address information of the image to be processed in the input data BRAM is determined by configuring the feature map starting address register; the starting address information of the model function parameters in the input data BRAM is determined by configuring the model function parameter starting address register; the starting address information of the output result in the output data BRAM is determined by configuring the result output starting address register; the number of rows information of the feature map to be calculated is determined by configuring the feature map row number register; the number of columns information of the feature map to be calculated is determined by configuring the feature map column number register; the information on whether to perform an activation operation on the calculation result is determined by configuring the activation enable register; the information on whether to perform a padding operation while performing pooling calculation is determined by configuring the padding enable register; the number of feature maps to be added is determined by configuring the feature map quantity register; the size information of the feature maps when adding feature maps is determined by configuring the feature map size register; the length information of the input column vector of this fully connected layer is determined by configuring the fully connected layer input column vector length register; the calculation mode information to be performed is determined by configuring the mode selection register, for example, convolution calculation mode, fully connected calculation mode, etc.
[0085] The above has described in detail an embodiment of the present invention, but the content described is only a preferred embodiment of the present invention and cannot be considered as being used to limit the scope of implementation of the present invention. All equal changes and improvements made in accordance with the scope of the present invention application should still fall within the scope covered by the patent of the present invention.
Claims
1. A general lightweight convolutional neural network acceleration method based on FPGA, characterized in that, The method includes: Obtaining the original acquisition data of the target convolutional neural network; the original acquisition data includes the image to be processed and the model function parameters of the target convolutional neural network; performing low-bit fixed-point quantization on the original acquisition data to obtain the original input data, and storing it in the input data BRAM; Obtaining the current target input data and model structure parameters of the target convolutional neural network; the input data of the target convolutional neural network includes the original input data and intermediate calculation data; Determining the current corresponding basic operator in the target IP core according to the model structure parameters as the target basic operator; Determining the registers that need to be configured for the target basic operator according to the type of the target basic operator as the target registers; configuring the target registers according to the target input data to obtain the target functional operator; the target IP core includes all basic operators of the target convolutional neural network, and different basic operators implement different calculation functions; Reading the target input data through the target functional operator, calculating to obtain the target output data, and storing it in the output data BRAM; Determining the data type of the target output data according to the model structure parameters; the data type includes the first data that will participate in the next calculation and the second data that will not participate in the next calculation; Storing the first data as the intermediate calculation data for the next calculation in the input data BRAM, and storing the second data according to the capacity of the input data BRAM; Repeatedly executing the above steps in sequence according to the model structure parameters until the overall calculation of the target convolutional neural network is completed.
2. The general lightweight convolutional neural network acceleration method based on FPGA according to claim 1, characterized in that The model function parameters include at least one of convolutional kernel parameters, convolutional kernel bias parameters, fully connected layer weight parameters, and fully connected layer bias parameters.
3. A general lightweight convolutional neural network acceleration method based on FPGA according to claim 1, characterized in that, The input data BRAM and the output data BRAM are dual-port BRAMs, and the ping-pong operation method is adopted; while the target functional operator is performing the calculation operation, the input data BRAM and the output data BRAM perform data transfer.
4. A general lightweight convolutional neural network acceleration method based on FPGA according to claim 1, characterized in that The model structure parameters include all basic operators of the target convolutional neural network and the connection order of each basic operator.
5. A general lightweight convolutional neural network acceleration method based on FPGA according to claim 4, characterized in that The target IP core includes at least one of a convolutional calculation unit, a pooling padding calculation unit, a feature addition calculation unit, a fully connected calculation unit, and an activation calculation unit.
6. A general lightweight convolutional neural network acceleration method based on FPGA according to claim 1, characterized in that, The target registers include at least one of a feature map start address register, a model function parameter start address register, a result output start address register, a feature map row number register, a feature map column number register, an activation enable register, a padding enable register, a feature map quantity register, a feature map size register, a fully connected layer input column vector length register, and a mode selection register.
Citation Information
Patent Citations
Convolutional neural network IP core based on FPGA
CN109784489A
Operation method and device based on neural network
CN111767986A