A Method for Implementing a Convolutional Neural Network FPGA Accelerator Based on Resource Reuse

By performing a combined parallel design of multi-channel convolutional layers, fully parallel design of small-scale convolutional layers, and resource multiplexing design of full-connection layers, the problem of limited FPGA resources is solved, and the accelerator design of large-scale convolutional neural networks is realized, with a wider scope of application.

CN116542295BActive Publication Date: 2025-05-27CHONGQING UNIV OF POSTS & TELECOMM +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310414320.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-18
Publication Date
2025-05-27
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

With limited FPGA computing resources, it is difficult to complete the accelerator design of large-scale convolutional neural networks, especially highly parallel computing occupies a large amount of resources, making it difficult to meet the needs of large-scale neural networks.

Method used

By performing a combined parallel design of parameters in the multi-channel convolutional layer, the parameters in the convolutional layer with small overall parameters are designed in parallel, and the parameters in the fully connected layer are multiplexed to reduce resource usage, and the design of activation functions, intermediate data storage and pooling layers is completed.

Benefits of technology

While taking into account the data processing speed, it solves the problem that large-scale neural network accelerators cannot be designed with limited computing resources, and expands the scope of application, so that large-scale convolutional neural network accelerators can be designed on FPGAs with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116542295B_ABST
    Figure CN116542295B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for implementing a convolutional neural network FPGA accelerator based on resource reuse. The present invention is used to solve the problem that in the case of limited FPGA computing resources, the accelerator design of large-scale neural networks cannot be completed. While taking into account the data processing speed, the occupation of computing resources is greatly reduced. First, two-dimensional input data is stored in one dimension and placed in the on-chip memory of the FPGA. Secondly, according to the number of channels and the number of parameters of the convolutional layer, the convolutional layer is divided into two categories, and combined parallel design and full parallel design are carried out respectively, reducing the occupation of computing resources while ensuring the data processing speed; for the convolutional layer with combined parallel design, intermediate data storage is designed; the activation function and pooling layer are designed, and the fully connected layer is reused to reduce the generation of additional clocks, accelerating the computing speed of the network while occupying a small amount of resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of neural network applications, and particularly relates to a method for implementing a convolutional neural network FPGA accelerator based on resource reuse. Background Art

[0002] In recent years, artificial intelligence has developed rapidly. As the most representative network structure in artificial intelligence, convolutional neural networks have received extensive attention. A convolutional neural network generally includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer connected in sequence. The input layer stores the input data of the convolutional neural network. The input data first enters the convolutional layer. The convolutional layer is characterized by local connection and weight sharing, and its main function is to extract features from the data input from the previous layer. After passing through the convolutional layer, it will enter the pooling layer. The pooling layer is generally divided into max pooling and average pooling, which are mainly used to reduce the amount of data and fuse local information. An activation function is generally added between the convolutional layer and the pooling layer. The main function of the activation function is to make the data as a whole exhibit non-linearity. After the input data passes through all the above layers, it will reach the fully connected layer. The fully connected layer will classify the data, and the classified data enters the output layer to obtain the final output result. After training, the convolutional neural network has good detection results for various dimensions of data, so it is widely used in fields such as machine vision, speech recognition, image processing, and natural language processing.

[0003] Currently, FPGA is the main design platform for convolutional neural networks. On the FPGA platform, the general method for designing an accelerator for convolutional neural networks is highly parallel computing. This method can maximize the use of the parallel characteristics of FPGA and process the input data of convolutional neural networks at the fastest speed. However, highly parallel computing will occupy a large amount of computing resources. Generally, convolutional neural networks with better performance are large-scale neural networks, which have more network layers and a larger overall number of parameters. For some FPGA platforms with limited computing resources, this highly parallel computing is difficult to meet, and thus it is difficult to complete the accelerator design of large-scale neural networks. Summary of the Invention

[0004] In order to solve the problems existing in the background art, the present invention provides a method for implementing a convolutional neural network FPGA accelerator based on resource reuse. By performing combined parallel design on the parameters in the multi-channel convolutional layer, full parallel design on the parameters in the convolutional layer with a relatively small overall number of parameters, and reuse design on the parameters in the fully connected layer, a large amount of resource occupation is reduced. At the same time, the design of the activation function, intermediate data storage, and pooling layer is completed. While taking into account the data processing speed, the problem that the accelerator design of large-scale neural networks cannot be completed in the case of limited FPGA computing resources is solved, including:

[0005] The original input data is divided into multiple two-dimensional data by channels, and each two-dimensional data is expanded into a one-dimensional data by rows and stored in the on-chip memory of the FPGA;

[0006] The convolutional layers with the number of input channels greater than X or the number of parameters greater than Y are regarded as the first type of convolutional layers, and the input data is convolved by means of parallelism between input channels, parallelism within the convolutional kernel, and parallelism between convolutional kernels within the channel; after the convolution operation is completed, the data storage of all channels is completed through M register shift operations;

[0007] The convolutional layers with the number of input channels less than X or the total number of parameters less than Y are regarded as the second type of convolutional layers, and the input data is convolved in a fully parallel manner;

[0008] The weight parameters corresponding to each output neuron of the fully connected layer are stored in the on-chip memory of the FPGA, the bias parameters of the fully connected layer are stored in the FPGA in the form of fixed values, and the fully connected operation is performed on the input data by means of resource reuse method;

[0009] The weight parameters of each channel of the first type of convolutional layer are designed with the same data bit width, and the weight parameters within the convolutional kernel are stored in the on-chip memory of the FPGA channel by channel, and the storage depth is M, where M represents the number of channels of the convolutional layer; the weight parameters and bias parameters of the second type of convolutional layer are stored in the FPGA in the form of fixed values;

[0010] The Relu activation function is designed by means of combining logic to judge the highest bit to perform logical judgment on the input data. When the highest bit of the input data is 0, the corresponding output data remains unchanged; when the highest bit of the input data is 1, the corresponding output data is cleared;

[0011] The maximum pooling layer is designed by means of direct comparison, and the pooling layer directly compares the input data within each channel to obtain the output data.

[0012] Preferably, the convolution of the input data by means of parallelism between input channels, parallelism within the convolutional kernel, and parallelism between convolutional kernels within the channel includes: parallelism between input channels is expressed as extracting the input data of B channels at the same position at the same moment; parallelism within the convolutional kernel is expressed as extracting all the parameters of a single convolutional kernel at the same moment; parallelism between convolutional kernels within the channel is expressed as extracting the parameters at the same position of all convolutional kernels within a single channel at the same moment; the parameters of all convolutional kernels within a single channel of the convolutional layer are read in sequence, and the parameters of all convolutional kernels within a single channel are subjected to corresponding convolution operations with the input data of all channels, where the convolution operation is the multiplication of the input data by the weight parameters and then adding the bias parameters, and B represents the number of channels of the input data.

[0013] Preferably, the data storage of all channels completed through the memory shift operation includes: the combined parallel convolution layer outputs the result of one channel per clock, performs a shift operation on the result output per clock through M registers, after M clocks, the M registers contain the output results of all channels, inputs the output results of all channels into the Relu activation function in parallel, clears the M registers at the same time, and repeats the shift operation.

[0014] Preferably, the convolution of the input data in a fully parallel manner includes: the full parallel representation extracts all the input data of B channels at the same moment; performs corresponding convolution operations on all the input data of B channels with the weight parameters and bias parameters stored on the FPGA chip, and inputs the output results of all channels into the Relu activation function in parallel, where B represents the number of channels of the input data.

[0015] Preferably, the fully connected operation on the input data through the resource reuse method includes: successively multiplying the weight parameters corresponding to a single output neuron by the input data of the input neurons of the fully connected layer and then adding the bias parameters corresponding to the output neuron stored on the FPGA chip, and completing the fully connected operation after repeating b times, where b represents the number of output neurons.

[0016] The present invention has at least the following beneficial effects

[0017] Through the combined parallel design of the parameters in the multi-channel convolution layer, the full parallel design of the parameters in the convolution layer with a relatively small overall number of parameters, and the reuse design of the parameters in the fully connected layer, the present invention reduces the occupation of a large amount of computing resources, and at the same time completes the design of the activation function, intermediate data storage, and pooling layer. While taking into account the data processing speed, it solves the problem that the accelerator design of a large-scale neural network cannot be completed in the case of limited FPGA computing resources. Compared with the prior art, the present invention has a wider application range, and the accelerator design of a large-scale convolutional neural network can also be completed on an FPGA with limited computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a schematic diagram of the method flow framework of the present invention;

[0019] Figure 2 is a schematic diagram of the storage of the original data of the present invention;

[0020] Figure 3 is a schematic diagram of convolving the input data in a manner of parallelism between input channels, parallelism within a convolution kernel, and parallelism between convolution kernels within a channel;

[0021] Figure 4 is a schematic diagram of convolving the input data in a fully parallel manner;

[0022] Figure 5 It is a schematic diagram of a fully connected layer. Specific implementation manners

[0023] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0024] Among them, the attached drawings are only used for exemplary illustration, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.

[0025] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the attached drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the attached drawings are only used for exemplary illustration and cannot be understood as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0026] Please refer to Figure 1 , the convolutional neural network includes: an input layer connected in sequence for inputting data; a fully parallel convolutional layer for performing a convolutional operation on the input data; an activation function; a pooling layer; a combined parallel convolutional layer for performing a convolutional operation on the input data;... a fully connected layer for performing a fully connected operation on the input data; an output layer for outputting a result.

[0027] A method for implementing a convolutional neural network FPGA accelerator based on resource reuse provided by the present invention includes:

[0028] Dividing the original input data into multiple two-dimensional data according to channels, and expanding each two-dimensional data into a one-dimensional data by rows and storing it in the on-chip memory of the FPGA;

[0029] Please refer to Figure 2, Embodiment, the input data of the convolutional neural network is two-dimensional image data, which includes image data of B channels; for example, when the input image is in RGB format, B is 3. The following takes Figure 1 as an example for illustration. In the figure, 8 data are one row of the image, and the images of each channel are stored row by row in the on-chip memory of the FPGA;

[0030] Please refer to Figure 3 . The convolutional layer with the input channels greater than X or the number of parameters greater than Y is regarded as a combined parallel convolutional layer, and the input data is convolved by means of parallelism between input channels, parallelism within the convolutional kernel, and parallelism between convolutional kernels within the channel; after the convolution operation is completed, all-channel data storage is completed through M register shift operations; X and Y are randomly set by those skilled in the art, and the specific values can be set according to the implementation effect.

[0031] The convolution of the input data by means of parallelism between input channels, parallelism within the convolutional kernel, and parallelism between convolutional kernels within the channel includes: parallelism between input channels means extracting the input data of B channels at the same position at the same time; parallelism within the convolutional kernel means extracting all the parameters of a single convolutional kernel at the same time; parallelism between convolutional kernels within the channel means extracting the parameters at the same position of all convolutional kernels within a single channel at the same time; successively read the parameters of all convolutional kernels within a single channel of the convolutional layer, and perform corresponding convolution operations on the parameters of all convolutional kernels in a single channel and the input data of all channels. The convolution operation is multiplying the input data by the weight parameters and then adding the bias parameters. B represents the number of channels of the input data.

[0032] Preferably, the extraction of the input data of B channels at the same position at the same time includes:

[0033] Input each row of data of a single-channel image into the first FIFO. When the first FIFO 1 is full, the data is shifted to the second FIFO; when the second FIFO is full, the data is shifted to the third FIFO 3, and so on; when all FIFOs are full, output the data of all FIFOs at the same time. The data output of the FIFOs is carried out simultaneously with the data shift. After all C data of all FIFOs are output, there are A - 1 clock data that are invalid, and these data will be discarded through logical judgment. After repeating the operation multiple times, all the data of a single channel of the image can be read; A represents the number of FIFOs. The image data of B channels can obtain the input data of B channels by adopting the above operations. C represents the amount of data in each row of a single-channel image;

[0034] Embodiment, for a combined parallel convolutional layer with a convolutional kernel size of Q, the input data is convolved by means of parallelism between input channels, parallelism within the convolutional kernel, and parallelism between convolutional kernels within the channel;

[0035] Please refer to Figure 3 , the FIFO simultaneously obtains the input data of three channels, convolves the input data 1, input data 2, input data 3 and the parameters of all convolution kernels of the first channel to obtain the output data 1, convolves with the parameters of all convolution kernels of the second channel to obtain the output data 2, and similarly obtains the output data 3, obtaining the output data of M = 3 channels, where M represents the number of channels of the convolutional layer.

[0036] When designing on the FPGA, the specific execution method is as follows: The M×N weight parameters of the convolutional layer are designed with the same bit width and stored in the on-chip memory of the FPGA. Each bit stores the data of one channel of this convolutional layer, and the storage depth is M. The bias parameter is defined into the convolutional layer in the form of a fixed value. The multi-channel input data is input into the convolutional layer in parallel. When the total number of input data for each channel is Q, stop the data input, and at the same time read the parameters of this layer from the memory in this combined parallel manner. One bit of data is read per clock, that is, all the weight parameters of one channel, which are divided into N according to the bit width of each weight parameter and input into this multi-layer convolutional layer in parallel, and perform corresponding convolution operations with the Q input data input into this layer in parallel. Repeat this reading M times to complete all the convolutions within one channel and obtain the first output result of one channel of this layer; among them, the convolution operation formula is as follows:

[0037] h(x, y) = f(x, y) * g(m, n) = ∑ m,n f(x + m, y + n) × g(m, n)

[0038] Among them, h(x, y) represents the result after convolution of the x-th row and y-th column, f(x, y) represents the x-th row and y-th column of the input image, g represents the size of the convolution kernel, and m and n are index values in the length and width dimensions.

[0039] After obtaining the first output results of all channels, input the Q-th data of the input data into the convolutional layer and remove the first one, keeping the amount of input data in the convolutional layer as Q. Repeat the above operation of reading parameters from the memory. After M clocks, the second output results of all channels can be obtained. The number of parameters read each time and the data bit width in the above operations are the same, so the same computing unit can be reused. If a fully parallel design is carried out, this convolutional layer requires M×N multiplication computing units. In contrast, the combined parallel only requires N multiplication computing units, greatly reducing the use of computing units.

[0040] Preferably, the data storage for all channels completed through the memory shift operation includes: the combined parallel convolution layer outputs the result of one channel per clock, performs a shift operation on the result output per clock through M registers. After M clocks, the M registers contain the output results of all channels. The output results of all channels are input into the Relu activation function in parallel, and at the same time, the M registers are cleared, and the shift operation is repeated;

[0041] The convolution data output by the combined parallel convolution layer outputs the result of only one channel each time. After M clocks, the results of all channels are output. The present invention designs M registers to store these results. A shift operation is performed between the M registers. After M clocks, all registers contain the first output results of all channels, and then they are input into the Relu activation function in parallel. At the same time, the M registers are cleared and wait for the arrival of the next result. The above operation is repeated until all results are processed.

[0042] Please refer to Figure 4 , and the convolution layer with the number of input channels less than X or the total number of parameters less than Y is used as a fully parallel convolution layer, and the input data is convolved in a fully parallel manner;

[0043] Preferably, the convolution of the input data in a fully parallel manner includes: fully parallel means extracting all the input data of B channels at the same time; performing corresponding convolution operations on all the input data of B channels with the weight parameters and bias parameters stored on the FPGA chip, and inputting the output results of all channels into the Relu activation function in parallel, where B represents the number of channels of the input data; this parallel method maximizes the data processing speed of the convolution layer. The input data is input into the convolution layer in parallel, and all the convolution kernels of multiple channels of the convolution layer simultaneously perform convolution on the input data to directly obtain all the output results of this layer.

[0044] When designing on the FPGA, the specific execution method is as follows: Define all the M×N weight parameters and M bias parameters of all channels in the convolution layer in the form of fixed values, that is, store them on the FPGA chip in the form of fixed values. The input data of each channel is input into the convolution layer at the same time, and one data is input per clock. When Q data are input, all the defined parameters are multiplied correspondingly, and the first output results of multiple channels are output at the same time. The input data is input continuously in sequence, and finally all the results of all channels can be obtained. This method can complete all the convolution operations of this layer in a short time. Due to the relatively small overall number of parameters, this design can minimize the occupation of computing resources while taking into account the data processing speed.

[0045] Design the weight parameters of each channel in the combined parallel convolution layer to have the same data bit width, and sequentially store the weight parameters in the convolution kernel in the on-chip memory of the FPGA according to the channels. The storage depth is M, and M represents the number of channels in the convolution layer; store the weight parameters and bias parameters of the fully parallel convolution layer in the on-chip of the FPGA in the form of fixed values;

[0046] Embodiment. For example, the first type of convolution layer has three channels, each channel contains 3 convolution kernels, and the total number of weight parameters of the three convolution kernels is N; sequentially store the weight parameters of the three convolution layers in order in the on-chip memory, and the storage depth is 3; the bias parameters of the convolution layer are stored in the on-chip of the FPGA in the form of fixed values.

[0047] Design the Relu activation function to perform logical judgment on the input data by means of combining logic to judge the highest bit. When the highest bit of the input data is 0, the corresponding output data remains unchanged; when the highest bit of the input data is 1, the corresponding output data is cleared;

[0048] Design the Relu activation function. The characteristic of this function is that when the input value is greater than zero, the output value is equal to the input value; when the input value is less than zero, the output value is equal to zero. The convolutional neural network of the present invention is a signed number. In the FPGA, the signed number is stored in the form of a complement code, and the highest bit is the sign bit. If the highest bit is 1, it is a negative number; if the highest bit is 0, it is a positive number. Therefore, the Relu activation function is designed as follows: by combining logic to judge the highest bit, when the highest bit is 0, the output data remains unchanged; when the highest bit is 1, the output data is cleared. This processing method effectively reduces the generation of additional clocks and accelerates the calculation speed of the network while occupying a small amount of resources.

[0049] Design the max pooling layer by means of direct comparison. The pooling layer directly compares the input data in each channel to obtain the output data;

[0050] Design the K×K max pooling layer. For the combined parallel convolution layer, the input data of the pooling layer takes 4×M + 1 clocks to obtain all channel data for K×K max pooling, and directly compare the K×K data in each channel to obtain all the output data of this layer. For the fully parallel convolution layer, the input data of the pooling layer takes K×K clocks to obtain all channel data for K×K max pooling, and directly compare the K×K data in each channel to obtain all the output data of this layer. This situation enables the data to be processed most quickly.

[0051] Store the weight parameters corresponding to each output neuron in the fully connected layer in the on-chip memory of the FPGA, store the bias parameters of the fully connected layer in the on-chip of the FPGA in the form of fixed values, and perform a fully connected operation on the input data by means of resource reuse;

[0052] Preferably, the fully connected operation on the input data by the resource reuse method includes: successively multiplying the weight parameters corresponding to a single output neuron by the input data of the input neurons in the fully connected layer and then adding the bias parameters corresponding to the output neuron stored on the FPGA chip, and after repeating b times, the fully connected operation is completed, where b represents the number of output neurons.

[0053] The output neurons in the fully connected layer are connected to each input neuron. This connection structure enables the fully connected layer to map all the data in the upper layer and combines all the features extracted by the neural network. However, this connection structure has a major drawback: each neuron in the fully connected layer involves a large number of network parameters, occupying a large number of computing units, and it is difficult to design on an FPGA with limited resources. Therefore, the present invention designs a computing resource reuse method: taking Figure 5 the fully connected layer as an example, there are a output neurons and b input neurons in the figure. Each output neuron corresponds to b weight parameters and one bias parameter. In this design, each time the b weight parameters corresponding to a single output neuron are multiplied and added to all the input data in the fully connected layer. After repeating a times, the output result of the fully connected layer can be obtained.

[0054] When designing on an FPGA, the specific execution method is as follows: design all a×b weight parameters of the fully connected layer to have the same bit width and store them in the on-chip memory of the FPGA. Each bit stores the b weight parameters corresponding to an output neuron, and the storage depth is a. The bias parameters are stored on the FPGA chip in the form of fixed values. The input data corresponding to all input neurons are input in parallel into the fully connected layer at the same clock. At the same time, the b weight parameters corresponding to a single output neuron are extracted from the memory. Each clock reads the b weight parameters corresponding to an output neuron and multiplies them with the input data corresponding to all input neurons respectively. The obtained results are then added to the fixed bias value stored on the FPGA chip to obtain the output result corresponding to an output neuron. Repeating this reading b times can obtain the output of the fully connected layer, where b represents the number of output neurons and a represents the number of input neurons.

[0055] The number of parameters read each time and the data bit width in the above operations are the same, so the same computing unit can be reused. If the fully connected layer is designed using the traditional method, this fully connected layer requires a×b multiplication computing units. In contrast, the fully connected layer of this design only requires a computing units, greatly reducing the use of computing units.

[0056] Combining all the designs, the input data first passes through the first fully parallel convolutional layer, and then through the activation function. The data output from the activation function enters the first combined parallel convolutional layer. The combined parallel convolutional layer involves reading the weight parameters in the weight memory. The output data in the combined parallel convolutional layer passes through the activation function again to reach the second pooling layer. The data in the pooling layer then passes through the second combined parallel convolutional layer. The second parallel convolutional layer has the same data processing method as the first one, except for the number of parameters and the number of channels. After the data comes out of the second combined parallel convolutional layer, it passes through the activation function to reach several fully parallel convolutional layers and combined parallel convolutional layers, and finally reaches the fully connected layer. After passing through the fully connected layer of this design, the final output data is obtained.

[0057] A method for implementing a convolutional neural network FPGA accelerator based on resource reuse according to the present invention. This method converts the data from two-dimensional to one-dimensional and then back to two-dimensional for image processing. Aiming at the large consumption of computing resources in its convolutional layer, it is divided into two categories with more channels and fewer channels and designed separately, which greatly reduces the consumption of computing resources. And the Relu activation function, intermediate data storage and the corresponding pooling layer are designed to accelerate the network calculation. Finally, a reuse design of the fully connected layer is carried out, and all the designs are integrated to complete the design of the convolutional neural network accelerator that takes into account the data processing speed under the condition of limited FPGA computing resources.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A method for implementing a convolutional neural network FPGA accelerator based on resource reuse, characterized in that, it includes: Dividing the original input data into multiple two-dimensional data by channels, and expanding each two-dimensional data into one-dimensional data by rows and storing it in the on-chip memory of the FPGA; Regarding the convolutional layer with the number of input channels greater than X or the number of parameters greater than Y as a combined parallel convolutional layer, and performing convolution on the input data in a manner of parallelism between input channels, parallelism within the convolutional kernel, and parallelism between convolutional kernels within a channel; after completing the convolution operation, complete the data storage of all channels through M register shift operations; Regarding the convolutional layer with the number of input channels less than X or the total number of parameters less than Y as a fully parallel convolutional layer, and performing convolution on the input data in a fully parallel manner; Storing the weight parameters corresponding to each output neuron of the fully connected layer in the on-chip memory of the FPGA, storing the bias parameters of the fully connected layer in the on-chip in the form of fixed values, and performing a fully connected operation on the input data through a resource reuse method; Designing the weight parameters of each channel of the combined parallel convolutional layer to have the same data bit width, and storing the weight parameters within the convolutional kernel in the on-chip memory of the FPGA channel by channel, with the storage depth being M, where M represents the number of channels of the convolutional layer; Storing the weight parameters and bias parameters of the fully parallel convolutional layer in the on-chip in the form of fixed values; Designing a Relu activation function to perform a logical judgment on the input data by means of judging the highest bit through combinational logic. When the highest bit of the input data is 0, the corresponding output data remains unchanged; When the highest bit of the input data is 1, the corresponding output data is cleared; Designing a max pooling layer by means of direct comparison, and the pooling layer directly compares the input data within each channel to obtain the output data.

2. The method for implementing a convolutional neural network FPGA accelerator based on resource reuse according to claim 1, characterized in that, The performing convolution on the input data in a manner of parallelism between input channels, parallelism within the convolutional kernel, and parallelism between convolutional kernels within a channel includes: Parallelism between input channels means extracting the input data of B channels at the same position at the same time; Parallelism within the convolutional kernel means extracting all the parameters of a single convolutional kernel at the same time; Parallelism between convolutional kernels within a channel means extracting the parameters at the same position of all convolutional kernels within a single channel at the same time; Sequentially reading the parameters of all convolutional kernels within a single channel of the convolutional layer, and performing corresponding convolution operations on the parameters of all convolutional kernels in a single channel and the input data of all channels, where the convolution operation is multiplying the input data by the weight parameters and then adding the bias parameters, and B represents the number of channels of the input data.

3. The method for implementing a convolutional neural network FPGA accelerator based on resource reuse according to claim 2, characterized in that, The data storage of all channels completed through the memory shift operation includes: the combined parallel convolution layer outputs the result of one channel per clock, performs a shift operation on the result output per clock through M registers, after M clocks, the M registers contain the output results of all channels, inputs the output results of all channels into the Relu activation function in parallel, clears the M registers at the same time, and repeats the shift operation.

4. The implementation method of a convolutional neural network FPGA accelerator based on resource reuse according to claim 1, characterized in that, the convolution of the input data in a fully parallel manner includes: fully parallel means extracting all the input data of B channels at the same time; performing corresponding convolution operations on all the input data of B channels with the weight parameters and bias parameters stored on the FPGA chip, and inputting the output results of all channels into the Relu activation function in parallel, where B represents the number of channels of the input data.

5. The implementation method of a convolutional neural network FPGA accelerator based on resource reuse according to claim 1, characterized in that, the fully connected operation on the input data through the resource reuse method includes: successively multiplying the weight parameters corresponding to a single output neuron by the input data of the input neurons of the fully connected layer and then adding the bias parameters corresponding to the output neuron stored on the FPGA chip, and repeating b times to complete the fully connected operation, where b represents the number of output neurons.

Citation Information

Patent Citations

  • An FPGA parallel system of convolution neural network algorithm

    CN109032781A

  • NLP reasoning acceleration system based on FPGA

    CN111275194A