A Convolutional Neural Network Accelerator Based on FPGA and Its Acceleration Method

By designing an FPGA accelerator with components such as controllers and convolution modules, and using parallel degree expansion and model quantization technology, the problem of insufficient storage and computing resources of FPGA convolutional neural network accelerator is solved, and the computing efficiency and inference speed are improved.

CN115374929BActive Publication Date: 2025-07-11XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110559767.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-21
Publication Date
2025-07-11
Estimated Expiration
2041-05-21

AI Technical Summary

Technical Problem

The existing convolutional neural network accelerators based on FPGA have storage problems and insufficient computing resources, resulting in poor network operation performance and low computing efficiency.

Method used

A convolutional neural network accelerator based on FPGA is designed, including controller, convolution module, first-in-first-out memory, adder, bias buffer, pooling module, linear rectifier function module, data transmission channel and random access memory. Through parallel degree expansion, model quantization and data alignment and parallel processing, the optimization of computing resources is achieved.

Benefits of technology

It improves the computing efficiency of convolutional neural networks on resource-constrained hardware platforms, reduces model size and memory consumption, and improves model inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115374929B_ABST
    Figure CN115374929B_ABST
Patent Text Reader

Abstract

The present invention discloses a convolutional neural network accelerator based on FPGA and its acceleration method. The accelerator is deployed on the FPGA and includes: a controller, a convolutional module, a first-in first-out memory, an adder, a bias buffer, a pooling module, a rectified linear unit module, a data transmission channel, and a random access memory. The present invention can improve the processing speed, storage capacity, and computing resources of the convolutional neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of integrated circuits, and particularly relates to a convolutional neural network accelerator based on FPGA and an acceleration method thereof. Background Art

[0002] Convolutional Neural Networks (CNN) are a class of feed-forward neural networks with deep structures and convolutional operations, and have been increasingly widely used in recent years.

[0003] The convolutional neural network is an algorithm with a very high computational density, which contains a large number of operations. In the prior art, a method of developing a convolutional neural network accelerator based on FPGA (Field Programmable Gate Array) is usually adopted to ensure the normal operation of the convolutional neural network.

[0004] However, there are still some technical problems in the accelerator developed by the prior art.

[0005] One is the storage problem: Since the internal storage of the FPGA cannot carry large-size convolutional neural network parameters and feature maps, frequent off-chip and on-chip storage accesses will occur during the operation of the convolutional neural network, which will lead to low network operation performance.

[0006] The other is the computing resource problem: When facing large-scale computations, the computing resources of its multiplication units are still insufficient. Summary of the Invention

[0007] In order to solve the above problems existing in the prior art, the present invention provides a convolutional neural network accelerator based on FPGA and an acceleration method thereof. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0008] A convolutional neural network accelerator based on FPGA, the accelerator is deployed on the FPGA, and the accelerator includes: a controller, a convolutional module, a first-in first-out memory, an adder, a bias buffer, a pooling module, a rectified linear unit module, a data transmission channel, and a random access memory; the controller is connected to the convolutional module, the first-in first-out memory, the adder, the pooling module, the rectified linear unit module, and the data transmission channel; the random access memory is connected to the convolutional module, the first-in first-out memory, the rectified linear unit module, and the data transmission channel; the adder is connected to the controller, the first-in first-out memory, the bias buffer, and the pooling module; the pooling module is connected to the adder, the controller, and the rectified linear unit module; the rectified linear unit module is connected to the pooling module, the controller, the random access memory, and the data transmission channel.

[0009] Advantages of the present invention:

[0010] 1. In the research on the hardware acceleration of the forward inference of the convolutional neural network by the present invention, the parallelism between input feature maps and output feature maps is expanded, so that the neural convolutional network can accelerate operations on a hardware platform with limited resources.

[0011] 2. The present invention adopts model quantization to reduce the model size, reduce the model memory consumption, and speed up the model inference speed.

[0012] 3. Based on data alignment parallel processing and multi-convolution kernel parallel computing, the present invention takes into account the computing flexibility and parallel efficiency, and improves the computing efficiency of the convolutional neural network.

[0013] The present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings

[0014] Figure 1 is a schematic structural diagram of a convolutional neural network accelerator based on FPGA provided by an embodiment of the present invention;

[0015] Figure 2 is a schematic flowchart of a convolutional neural network acceleration method based on FPGA provided by an embodiment of the present invention;

[0016] Figure 3 is a schematic structural diagram of a random access memory provided by an embodiment of the present invention;

[0017] Figure 4 is a schematic diagram of two working states provided by an embodiment of the present invention;

[0018] Figure 5 is a schematic connection diagram of a sliding window module provided by an embodiment of the present invention;

[0019] Figure 6 It is a schematic diagram of a convolution module structure provided by an embodiment of the present invention. Specific embodiments

[0020] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0021] Embodiment 1

[0022] Please refer to Figure 1 , Figure 1 It is a schematic diagram of a convolutional neural network accelerator structure based on FPGA provided by an embodiment of the present invention. The accelerator is deployed on the FPGA and includes: a controller, a convolution module, a first-in-first-out memory, an adder, a bias buffer, a pooling module, a rectified linear unit module, a data transmission channel, and a random access memory;

[0023] Optionally, the controller is connected to the convolution module, the first-in-first-out memory, the adder, the pooling module, the rectified linear unit module, and the data transmission channel;

[0024] Optionally, the random access memory is connected to the convolution module, the first-in-first-out memory, the rectified linear unit module, and the data transmission channel;

[0025] Optionally, the adder is connected to the controller, the first-in-first-out memory, the bias buffer, and the pooling module;

[0026] Optionally, the pooling module is connected to the adder, the controller, and the rectified linear unit module;

[0027] Optionally, the rectified linear unit module is connected to the pooling module, the controller, the random access memory, and the data transmission channel.

[0028] In summary: 1. In the research on the forward inference hardware acceleration of convolutional neural networks by the present invention, the parallelism between input feature maps and output feature maps is expanded, enabling the neural convolutional network to accelerate operations on resource-constrained hardware platforms.

[0029] 2. The present invention adopts model quantization to reduce the model size, reduce the model memory consumption, and accelerate the model inference speed.

[0030] 3. Based on data alignment parallel processing and multi-convolution kernel parallel computing, the present invention takes into account the computing flexibility and parallel efficiency, and improves the computing efficiency of the convolutional neural network.

[0031] Embodiment 2

[0032] See also Figure 2 , Figure 2 : is a flow chart of a convolutional neural network acceleration method based on FPGA provided by an embodiment of the present invention, which is applied to the accelerator, in which a convolutional neural network model is deployed, and the method includes:

[0033] Step 1: Get the image to be processed.

[0034] The purpose of the present invention is to provide a convolutional neural network accelerator to solve the problems that convolutional neural networks have a large number of multiplication and accumulation calculations, low efficiency when processing images, long running time, and difficulty in taking into account both computational flexibility and parallel efficiency. Specifically, the present invention expands the parallelism between input feature maps and output feature maps and optimizes the total amount of off-chip memory access of the convolutional layer to complete the design of the convolutional neural network accelerator and implement it on an FPGA.

[0035] Step 2: preprocessing the image to be processed to obtain a preprocessed image;

[0036] Optionally, preprocessing the image to be processed to obtain a preprocessed image includes:

[0037] Step 21, performing scale transformation processing on the image to be processed;

[0038] In order to improve data transmission speed and increase accelerator processing efficiency, the image is scaled to reduce the image pixels without seriously reducing the accuracy.

[0039] Step 22: grayscale the transformed image to be processed to obtain a preprocessed image.

[0040] In order to improve the efficiency of image transmission, the image is grayed in the accelerator. By using the graying formula, we shrink the image data while retaining its features, increasing the speed of image transmission.

[0041] Step 3: compressing the convolutional neural network model to obtain a target model;

[0042] Optionally, a distillation sub-model is deployed in the convolutional neural network model.

[0043] Optionally, compressing the convolutional neural network model to obtain a target model includes:

[0044] The convolutional neural network model is compressed according to the distillation sub-model to obtain a target model.

[0045] Since the original convolutional neural network model is relatively complex and will seriously affect the running time on the FPGA, we need to compress the parameters of the original model. The distillation model can greatly compress the model parameters without a significant decrease in accuracy. The target model is also known as the large neural convolutional network.

[0046] Step 4: Calculate the loss function according to the target model to determine multiple target sub-models.

[0047] Optionally, the random access memory is deployed with a bidirectional port, and the random access memory is also deployed with a preset threshold.

[0048] The interior of the random access memory consists of a simple dual-port RAM IP and an address control logic, and uses a dual-port for caching. Therefore, there can be two configuration processes when caching data is transferred.

[0049] Optionally, the calculating the loss function according to the target model to determine multiple target sub-models includes:

[0050] Step 41: Calculate the loss function corresponding to the sub-model according to the target model. The loss function includes a cross-entropy loss function and a distillation loss function.

[0051] After obtaining a target model with higher accuracy, start training the sub-model, which is also known as the small network. The loss function for training the small network mainly consists of two parts: one part is to calculate the distillation loss / KL divergence using the output logits of the large network and the small network, and the other part is to calculate the cross-entropy loss using the output of the small network and the data labels.

[0052] Step 42: Quantize the sub-model based on the loss function to obtain a first sub-model.

[0053] Model quantization is adopted, that is, the floating-point model weights with continuous values (or a large number of possible discrete values) or the tensor data flowing through the model are approximately fixed-pointed to a finite number (or fewer) of discrete values with a relatively low inference accuracy loss, and a data type with fewer bits is used to approximately represent the 32-bit finite-range floating-point data, so as to achieve the goals of reducing the model size, reducing the model memory consumption, and accelerating the model inference speed, etc., to obtain a first sub-model.

[0054] Step 43: Store the first sub-model in the random access memory.

[0055] The first sub-model corresponds to neuron weights. The storing the first sub-model in the random access memory includes: storing the neuron weights corresponding to the first sub-model in the random access memory.

[0056] When a weight output is completed, the internal address accumulation counter of the distribution cache is reset to zero. The starting address required for the next calculation is reconfigured, and the distribution cache is enabled. The internal address accumulation counter starts counting, the neuron weight cache increments from the initial address, and the cache control circuit outputs the weight data by controlling the simple dual-port.

[0057] The neuron thresholds are stored in the RAM in the calculation order. After the calculation in the accumulator of the arithmetic unit is completed, the neuron thresholds output data in the address order given by the control logic.

[0058] Optionally, the storing in the random access memory includes: caching neuron weights in the random access memory according to a dynamic ping-pong cache.

[0059] To meet the requirements of input data buffering during the operation process, we use a dynamic ping-pong cache. Internally, 2 true dual-port RAM IPs and corresponding control logic are used to dynamically cache data in the form of ping-pong operation. The true dual-port RAM is different from the simple dual-port RAM. The true dual-port RAM includes two independent and complete interfaces that can be read and written. The two independent interfaces can read and write the stored data simultaneously. Neuron data is dynamic data, and the data used in the calculation can be overwritten by new data.

[0060] In the dynamic ping-pong cache, 2 identical true dual-port RAMs are used, namely RAM0 and RAM1. When the neural network accelerator operates, the two dual-port RAMs with the same function are selected by the selection signal SEL. One of the RAMs participates in the calculation of the current network, and cooperates with the arithmetic unit to perform corresponding data reading and writing operations. The other RAM cooperates with the external interface to input the neuron data of the next sample input layer and output the data of the current output layer. Similar to the neural network weight and threshold cache, the dynamic ping-pong cache also has the functions of initializing the read and write addresses and resetting the address accumulation counter.

[0061] When SEL is 0, doutb of RAM0 and the neuron input layer data interface of the arithmetic unit are gated through the SEL signal, as Figure 3 dina of the left RAM0 and the result output interface of the arithmetic unit are gated through the SEL signal. At this time, the arithmetic unit can sequentially read the input layer data from the doutb interface of RAM0 from the initial address and cooperate with the operation pipeline. The output layer result of the arithmetic unit can be written into RAM0 through the dina interface of RAM0. At this time, dinb and douta of RAM0 are in a high-impedance state. See Figure 3 It is a schematic diagram of a random access memory structure provided by an embodiment of the present invention.

[0062] Step 44: Train the first sub-module based on the random access memory to obtain a target sub-model.

[0063] The douta of RAM1 and the output of the external interface are gated by the SEL signal. At this time, the external device can read the output layer result of the previous neural network from this interface. The dinb of RAM1 and the input of the external interface are gated by SEL at this time. The external device can write the sample data required for the next calculation into RAM1 through this interface. When a neural network calculation is completed, RAM0 and RAM1 are switched, and the SEL signal changes from 0 to 1. Similarly. Figure 4 These are the schematic diagrams of two working states provided by the embodiments of the present invention.

[0064] Step 5: Process the image to be processed based on the target sub-model to obtain a target image.

[0065] Optionally, a sliding window module is included in the random access memory. See Figure 5 This is a schematic diagram of the connection of a sliding window module provided by the embodiments of the present invention. The input interface of the image is connected to the first FIFO. The output of the first FIFO is connected to the input of the second FIFO. The output of the second FIFO is connected to the input of the third FIFO.

[0066] Three-parallel computing is implemented inside the convolution kernel. The image is input to each convolution operation unit in a three-row parallel input mode. The weights corresponding to each multiplier-accumulator are read by the accelerator from the SD card. The first pixel value of each row completes the multiplication calculation with the weight in the first multiplier-accumulator of each row, and then is passed to the second multiplier-accumulator as the accumulation part of the second calculation. The second pixel value corresponding to each row is multiplied by the weight of the second multiplier-accumulator and accumulated with the output result of the first multiplier-accumulator, and then passed to the third multiplier-accumulator. The third multiplier-accumulator multiplies with the third pixel value of each row and accumulates with the output result of the second multiplier-accumulator to obtain the multiplication-accumulation value of each row. Then, the output results of each row are accumulated to obtain the output value of one convolution.

[0067] The random access memory further includes a weight cache module. The weight cache module is used to receive and temporarily store the data read from the SD card.

[0068] The present invention uses 8 PE modules for parallel computing to achieve parallel operations between different convolution kernels when processing one picture at the same time. Only four are used in the first layer of convolution operations, but in the second and third layer of convolution layers, all eight convolution operation units are used because the convolution kernels are eight.

[0069] The present invention uses max pooling for downsampling, that is, a 2×2 convolution kernel processes the image with a stride of 2, and the process of retaining the largest element in the area covered by the kernel. Although pooling is destructive, since the image is scaled down proportionally and has translational, rotational, and scale invariance, the information loss of the image is small, and the main content of the image still retains most of the original spatial information, retains the main features, and can prevent overfitting.

[0070] In addition, since the intermediate data has multiple channels, the convolution results of each channel need to be temporarily stored in the RAM and sent into the addition module via the FIFO for summation and adding the bias term.

[0071] Furthermore, after the results of the convolution calculation are sent to the pooling layer and the fully connected layer, the target image is obtained and written into the RAM, and the operation process of the accelerator ends.

[0072] In summary: 1. The present invention conducts research on the hardware acceleration of the forward inference of the convolutional neural network, expands the parallelism between the input feature maps and the output feature maps, enabling the neural convolutional network to accelerate operations on resource-constrained hardware platforms.

[0073] 2. The present invention adopts model quantization to reduce the model size, reduce the model memory consumption, and accelerate the model inference speed.

[0074] 3. The present invention is based on data alignment parallel processing and multi-convolution kernel parallel computing, taking into account the computational flexibility and parallel efficiency, and improving the computational efficiency of the convolutional neural network.

[0075] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0076] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0077] In the present invention, unless otherwise clearly defined or limited, terms such as "installed", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral one; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication between two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0078] In the present invention, unless otherwise clearly defined or limited, the first feature being "on" or "under" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through additional features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes the first feature being directly above and obliquely above the second feature, or merely indicating that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "beneath" and "underneath" the second feature includes the first feature being directly under and obliquely under the second feature, or merely indicating that the horizontal height of the first feature is lower than that of the second feature.

[0079] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0080] Although the present application has been described in connection with various embodiments herein, however, in the process of implementing the claimed present application, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure content, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality of situations. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0081] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the procedures Figure 1 one or more procedures and / or blocks Figure 1 specified in the block or blocks.

[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the procedures Figure 1 one or more procedures and / or blocks Figure 1 specified in the block or blocks.

[0083] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited only to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for accelerating a convolutional neural network based on FPGA, characterized in that, Applied to an accelerator, the accelerator is deployed on an FPGA, and the accelerator includes: a controller, a convolution module, a first-in first-out memory, an adder, a bias buffer, a pooling module, a rectified linear unit module, a data transfer channel, and a random access memory; The controller is connected to the convolution module, the first-in first-out memory, the adder, the pooling module, the rectified linear unit module, and the data transfer channel; The random access memory is connected to the convolution module, the first-in first-out memory, the rectified linear unit module, and the data transfer channel; The adder is connected to the controller, the first-in first-out memory, the bias buffer, and the pooling module; The pooling module is connected to the adder, the controller, and the rectified linear unit module; The rectified linear unit module is connected to the pooling module, the controller, the random access memory, and the data transfer channel; A convolutional neural network model is deployed in the accelerator, and the method includes: Obtaining an image to be processed; preprocessing the image to be processed to obtain a preprocessed image; Performing compression processing on the convolutional neural network model to obtain a target model; Calculating a loss function according to the target model to determine a plurality of target sub-models; Processing the image to be processed based on the target sub-model to obtain a target image; wherein, The calculating a loss function according to the target model to determine a plurality of target sub-models includes: Calculating a loss function corresponding to the sub-model according to the target model, the loss function including a cross-entropy loss function and a distillation loss function; after obtaining the target model, training the sub-model, and the loss function for training the sub-model mainly consists of two parts: one part is to calculate the distillation loss from the output logits of the target model and the target sub-model, and the other part is to calculate the cross-entropy loss from the output of the target sub-model and the data label; Performing quantization processing on the sub-model based on the loss function to obtain a first sub-model; Storing the first sub-model in the random access memory; Based on the random access memory, train the first sub-model to obtain a target sub-model; the random access memory includes a sliding window module, the input interface of the image is connected to the first FIFO, the output of the first FIFO is connected to the input of the second FIFO, and the output of the second FIFO is connected to the input of the third FIFO; three-parallel computing is implemented inside the convolution kernel, and the image is input to each convolution operation unit in a three-row parallel input mode. The weights corresponding to each multiplier-accumulator are read by the accelerator from the SD card. The first pixel value of each row completes the multiplication calculation with the weight in the first multiplier-accumulator corresponding to that row, and then is passed to the second multiplier-accumulator as the accumulative part of the second calculation. The second pixel value corresponding to each row is multiplied by the weight of the second multiplier-accumulator and accumulated with the output result of the first multiplier-accumulator, and then passed to the third multiplier-accumulator. The third multiplier-accumulator multiplies with the third pixel value of each row and accumulates with the output result of the second multiplier-accumulator to obtain the multiplication-accumulation value of each row. Then, the output results of each row are accumulated to obtain the output value of one convolution.

2. The method according to claim 1, wherein Preprocessing the to-be-processed image to obtain a preprocessed image, including: Performing a scale transformation process on the to-be-processed image; Performing grayscale processing on the transformed to-be-processed image to obtain a preprocessed image.

3. The method according to claim 1, wherein A distillation sub-model is deployed in the convolutional neural network model. Compressing the convolutional neural network model to obtain a target model includes: Compressing the convolutional neural network model according to the distillation sub-model to obtain a target model.

4. The method according to claim 1, characterized in that The random access memory is deployed with a bidirectional port, and the random access memory is also deployed with a preset threshold.

Citation Information

Patent Citations

  • An FPGA parallel system of convolution neural network algorithm

    CN109032781A

  • FPGA-based neural network acceleration method and accelerator

    CN110852428A