A design method for fast convolution and cache mode of convolutional neural network
By applying the Winograd algorithm and feature map cache on FPGA and optimizing the convolution sliding window operation, efficient acceleration and low-power design of convolution calculation are achieved, solving the problem of large computational complexity of deep convolutional neural networks.
Patent Information
- Application Number
- CN202210937743.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-05
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-08-05
AI Technical Summary
Existing deep convolutional neural networks are computationally intensive, resulting in enormous computational pressure. Existing hardware accelerators are unable to effectively reduce computational complexity and memory requirements.
The Winograd algorithm is adopted and FPGA design is utilized to optimize the convolution sliding window operation through feature map caching and pipeline transmission. The convolution calculation process is optimized by combining reconfigurable matrix transformation design.
Significantly reduce the number of multiplication operations, improve the efficiency of cache and conversion modules, reduce power consumption, and enhance the computing power of the convolution acceleration module.
Smart Images

Figure CN115204373B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of convolutional neural network acceleration, and in particular to a method for designing a fast convolution and cache mode of a convolutional neural network. Background Art
[0002] Currently, deep convolutional neural networks (DCNNs) have achieved remarkable performance in various computer vision tasks such as image classification, object detection and semantic segmentation.
[0003] With the emergence of larger datasets and models, the accuracy of convolutional neural networks has been greatly improved. However, this comes at the cost of enormous computational effort and processing time. To cope with this enormous computational pressure, hardware accelerators such as GPUs, FPGAs, and ASICs have been widely used to accelerate CNNs.
[0004] For the convolutional layer, which is the most computationally intensive layer in convolutional neural networks, algorithms such as Winograd and FFT have emerged to accelerate convolution operations. The Winograd algorithm can reduce the arithmetic complexity of traditional convolution by a factor of four. It also has lower memory requirements than traditional FFT convolution algorithms. This makes it possible to accelerate convolution operations using the Winograd algorithm. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for designing a fast convolution and caching mode of a convolutional neural network using the Winograd algorithm.
[0006] To achieve the above objectives, the technical solutions provided by the present invention are:
[0007] A method for designing fast convolution and caching modes for convolutional neural networks includes designing a Winograd algorithm using FPGAs, optimizing the caching of repeated row and column data required for convolutional sliding window operations through feature map caching and pipeline transmission, and implementing a reconfigurable design for matrix transformation through parameter configuration.
[0008] Furthermore, the Winograd algorithm includes:
[0009] Winograd algorithm accelerates one-dimensional convolution calculation:
[0010] For input i=[i0 i1 i2 i3] T , the convolution kernel is w=[w0 w1 w2] T , the expression of convolution is:
[0011]
[0012] in,
[0013]
[0014]
[0015] The transformation of one-dimensional convolution is extended to two-dimensional convolution, which can be expressed in matrix form as follows:
[0016] Y=A T [(GWG T )⊙(B T IB)]A
[0017] Among them, Y, W, and I represent the output feature map, weight, and input feature map, respectively; G, B, and A represent the weight conversion matrix, input feature map conversion matrix, and output feature map conversion matrix, respectively.
[0018] Furthermore, FPGA is used to complete the design of the Winograd algorithm, including the top-level control module, Winograd convolution module, and cache module;
[0019] Wherein, the top-level control module includes a pipeline scheduling unit, a control instruction parsing unit, and a matrix conversion control unit;
[0020] The pipeline scheduling unit is divided into three stages: the first stage input port includes input cache and weight cache; the second stage includes matrix conversion of input activation and weight, and Winograd convolution; the third stage output port includes intermediate value cache and output cache;
[0021] The control instruction parsing unit parses the written control instruction, the parsed control instruction including the cached data structure parameters, convolution size, and step size, and sends it to different modules or units according to requirements;
[0022] The matrix conversion control unit receives the instruction of the control instruction parsing unit and controls the input activation, weight and PE array calculation results to perform matrix conversion;
[0023] The Winograd convolution module includes a register array unit, a MUX unit, a convolution control unit, a PE array unit, an addition tree unit, a ReLU, and a POOL unit;
[0024] The register array unit caches the result of matrix conversion;
[0025] The MUX unit implements data multiplexing by selecting corresponding rows and columns to the PE array;
[0026] The convolution control unit is used for convolution control;
[0027] The PE array unit is a multiplier array implemented by DSP, which completes the dot multiplication calculation of the winograd domain;
[0028] The addition tree unit accumulates the calculation results and implements the bias addition function;
[0029] The ReLU and POOL units implement activation functions and pooling functions;
[0030] The cache module includes an input cache unit, a weight cache unit, an intermediate value cache unit, and an output cache unit;
[0031] The input cache unit caches input activations read from off-chip, caches block-based input activations according to top-level control parameters, and performs zero padding and duplicate row caching.
[0032] The weight cache unit caches weight parameters;
[0033] The intermediate value cache unit caches the calculation results of the blocks for the addition tree unit to read and complete the accumulation;
[0034] The output buffer unit buffers the final calculation result and stores it off-chip via DMA.
[0035] Furthermore, by caching feature maps and pipeline transmission, the convolution sliding window operation needs to cache repeated row and column data, including:
[0036] The H*L*N input feature map is expanded in the row direction for block operation and stored in 16 Brams in channel order for caching; the input cache is sent to the intermediate module of the conversion module, and a 4*4 block sliding window operation is completed for the 3*3 convolution kernel. It consists of a four-stage pipeline. Pipeline 0 reads the data of the same address in 16 Brams per clock cycle, and the input is passed to the subsequent pipeline. In the row direction, the sliding window data block is passed to the matrix conversion module once every two cycles, thereby optimizing the repeated data reading in the row direction; in the column direction, the repeated column data will be read into the corresponding two matrix conversion modules, that is, the 16 rows of input feature maps stored in 16 Brams are passed to the 7 input feature map matrix conversion modules, optimizing the repeated data reading in the column direction.
[0037] Furthermore, the reconfigurable design for realizing matrix transformation by configuring parameters includes:
[0038] The conversion matrix of the two-dimensional Winograd algorithm includes A, G, and B. The composition of the conversion matrix is only related to the convolution size of the Winograd algorithm and the size of the input feature map. There are repeated parameters between the conversion matrices. By configuring the parameters s and s0s1s2, a small amount of memory and resources are occupied to achieve a reconfigurable design of the conversion module. The conversion matrix B of the input feature map is derived as follows:
[0039]
[0040]
[0041] Furthermore, it also includes: in the dot product process, for the position where the weight parameter is zero, the multiplication calculation of the parameter will be turned off, and for the position where the weight parameter is not zero, the multiplication calculation of the feature map data is skipped to achieve a low-power design of the convolution calculation.
[0042] Compared with the existing technology, the principles and advantages of this solution are as follows:
[0043] 1) Fully utilize FPGA to complete the design of Winograd algorithm. Under the convolution kernel size of 3*3, the multiplication operations are reduced from 36 to 16 times, and the convolution operation can be accelerated by fully utilizing DSP resources.
[0044] 2) By caching feature maps and pipeline transmission, the convolution sliding window operation requires caching of repeated row and column data, thereby improving the efficiency of the cache and conversion modules.
[0045] 3) Realize the reconfigurable design of the matrix conversion module by configuring parameters.
[0046] 4) Make full use of the saved DSP resources, reuse the Winograd convolution calculation array, make full use of the FPGA hardware resources, further improve the computing power of the convolution acceleration module, and design a low-power design for the dot product step. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the services required for use in the embodiments or the prior art descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1 Flowchart for the hardware implementation of the Winograd algorithm;
[0049] Figure 2 Schematic diagram of input feature map cache;
[0050] Figure 3 Schematic diagram of caching input feature maps into the conversion module;
[0051] Figure 4 Schematic diagram of the conversion module. DETAILED DESCRIPTION
[0052] The present invention will be further described below in conjunction with specific embodiments:
[0053] The present embodiment describes a method for designing a fast convolution and caching mode for a convolutional neural network, including designing a Winograd algorithm using an FPGA, optimizing the caching of repeated row and column data required for convolution sliding window operations through feature map caching and pipeline transmission, and implementing a reconfigurable design for matrix conversion by configuring parameters.
[0054] Among them, the Winograd algorithm includes:
[0055] Winograd algorithm accelerates one-dimensional convolution calculation:
[0056] For input i=[i0 i1 i2 i3] T , the convolution kernel is w=[w0 w1 w2] T , the expression of convolution is:
[0057]
[0058] in,
[0059]
[0060]
[0061] The transformation of one-dimensional convolution is extended to two-dimensional convolution, which can be expressed in matrix form as follows:
[0062] y=A T [(GWG T )⊙(B T IB)]A
[0063] Among them, Y, W, and I represent the output feature map, weight, and input feature map, respectively; G, B, and A represent the weight conversion matrix, input feature map conversion matrix, and output feature map conversion matrix, respectively.
[0064] like Figure 1As shown in the figure, the implementation process of the entire algorithm involves three matrix transformations. The matrix transformation of the input activations and weights at the input end is to complete the matrix multiplication in the Winograd domain. The result of the calculation is still a data structure in the Winograd domain, so the output end needs to transform the result of the Winograd domain calculation to obtain the final output result. The overall hardware circuit implementation process also follows this process.
[0065] The design of Winograd algorithm is completed using FPGA, including top-level control module, Winograd convolution module and cache module;
[0066] Among them, the top-level control module includes a pipeline scheduling unit, a control instruction parsing unit, and a matrix conversion control unit; the pipeline scheduling unit is divided into three levels of pipelines: the first-level input port includes an input cache and a weight cache; the second level includes matrix conversion of input activations and weights, as well as winograd convolution; the third-level output port includes an intermediate value cache and an output cache; the control instruction parsing unit parses the written control instructions, and the parsed control instructions include cached data structure parameters, convolution size, and step size, and are sent to different modules or units according to requirements; the matrix conversion control unit receives instructions from the control instruction parsing unit, controls the input activations, weights, and PE array calculation results to perform matrix conversion;
[0067] The Winograd convolution module includes a register array unit, a MUX unit, a convolution control unit, a PE array unit, an addition tree unit, a ReLU and a POOL unit; the register array unit caches the results of matrix conversion; the MUX unit implements data reuse by selecting the corresponding rows and columns to the PE array unit; the convolution control unit is used for convolution control; the PE array unit is a multiplier array implemented by DSP, which completes the point multiplication calculation of the Winograd domain; the addition tree unit accumulates the calculation results and implements the bias addition function; the ReLU and POOL units implement the activation function and pooling function;
[0068] The cache module includes an input cache unit, a weight cache unit, an intermediate value cache unit, and an output cache unit; the input cache unit caches the input activations read from outside the chip, caches the input activations after blocking according to the parameters controlled by the top level, and completes zero padding and repeated row caching; the weight cache unit caches the weight parameters; the intermediate value cache unit caches the calculation results of the blocks for the addition tree unit to read and complete the accumulation; the output cache unit caches the final calculation results and stores them outside the chip through DMA.
[0069] The implementation principle is as follows:
[0070] The front end is the input activation data input buffer unit, which reads the blocked data into the on-chip buffer through DMA. The zero padding and repeated row storage of the feature map are also completed in this unit. After that, the input data matrix transformation is performed, that is, the completed H*L*N (height*width*channel) is blocked, and the blocks are divided into long strips. The reason is to make the data address continuous and improve the DMA transmission efficiency. The block size is Th*L*Tn (block height*length*block channel);
[0071] The cached input data will be read into the Winograd convolution module, and after matrix transformation, it will enter the PE array unit together with the weight for calculation. For example, for a data stream with an input size of 4*4, a convolution kernel size of 3*3, and a stride of 1, the Winograd convolution module will simultaneously read the convolution kernels of Tm output channels, map them in the PEArray, and calculate the output.
[0072] After the PE array unit calculates, the calculation result is converted into an output matrix to obtain the final block calculation result. The intermediate values are accumulated in the addition tree and cached in the intermediate value cache unit. After the accumulated output data is activated and pooled in the ReLU and POOL units, it is cached in the output cache unit and the result is output to post-processing after the DMA is idle.
[0073] Specifically, if Figure 2 As shown in , the input feature map of H*L*N is expanded in the row direction for block operation and stored in 16 Brams in channel order for cache. Figure 3 The input is cached to the middle module of the conversion module, and a 4*4 block sliding window operation is completed for the 3*3 convolution kernel. It consists of a four-stage pipeline. Pipeline 0 reads in 16 data at the same address in Bram per clock cycle, and the input is passed to the subsequent pipeline. In the row direction, the sliding window data block can be passed to the matrix conversion module once every two cycles, which can optimize the repeated data reading in the row direction; in the column direction, the repeated column data will be read into the corresponding two matrix conversion modules, that is, the 16 rows of input feature maps stored in 16 Bram are passed to the 7 input feature map matrix conversion modules, which optimizes the repeated data reading in the column direction.
[0074] Reconfigurable designs that achieve matrix transformation by configuring parameters include:
[0075] The transformation matrix of the two-dimensional Winograd algorithm includes A, G, and B. The composition of the transformation matrix is only related to the convolution size of the Winograd algorithm and the size of the input feature map, and there are repeated parameters between the transformation matrices, such as Figure 4As shown, by configuring the parameters s and s0s1s2, a small amount of memory and resources are occupied to realize the reconfigurable design of the conversion module; the input feature map conversion matrix B is derived as follows:
[0076]
[0077]
[0078] Finally, this embodiment also includes:
[0079] Since there are a large number of zero parameters in the input feature map and weights, during the dot product process, for the position where the weight parameter is zero, the multiplication calculation of the parameter will be turned off. For the position where the weight parameter is not zero, the multiplication calculation of the feature map data is skipped to achieve a low-power design for convolution calculation.
[0080] The embodiments described above are only preferred embodiments of the present invention and are not intended to limit the scope of implementation of the present invention. Therefore, any changes made based on the shape and principle of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for designing a fast convolution and caching mode for a convolutional neural network, characterized in that: This includes using FPGA to design the Winograd algorithm, optimizing the cache of repeated row and column data required for convolutional sliding window operations through feature map caching and pipeline transmission, and implementing a reconfigurable design for matrix transformation through parameter configuration. By caching feature maps and pipelining, we can optimize the convolution sliding window operation by caching repeated rows and columns of data, including: The H*L*N input feature map is expanded in the row direction for block operation and stored in 16 Brams in channel order for caching; the input cache is sent to the intermediate module of the conversion module, and a 4*4 block sliding window operation is completed for the 3*3 convolution kernel. It consists of a four-stage pipeline. Pipeline 0 reads the data of the same address in 16 Brams per clock cycle, and the input is passed to the subsequent pipeline. In the row direction, the sliding window data block is passed to the matrix conversion module once every two cycles, thereby optimizing the repeated data reading in the row direction; in the column direction, the repeated column data will be read into the corresponding two matrix conversion modules, that is, the 16 rows of input feature maps stored in 16 Brams are passed to the 7 input feature map matrix conversion modules, optimizing the repeated data reading in the column direction.
2. The method for designing a fast convolution and caching mode for a convolutional neural network according to claim 1, wherein: The Winograd algorithm includes: Winograd algorithm accelerates one-dimensional convolution calculation: For input i=[i0 i1 i2 i3] T , the convolution kernel is w=[w0 w1 w2] T , the expression of convolution is: The transformation of one-dimensional convolution is extended to two-dimensional convolution, which can be expressed in matrix form as follows: Y=A T [(GW) T )⊙(B T IB)]A Among them, Y, W, and I represent the output feature map, weight, and input feature map, respectively; G, B, and A represent the weight conversion matrix, input feature map conversion matrix, and output feature map conversion matrix, respectively.
3. The method for designing a fast convolution and caching mode for a convolutional neural network according to claim 2, wherein: The design of Winograd algorithm is completed using FPGA, including top-level control module, Winograd convolution module and cache module; Wherein, the top-level control module includes a pipeline scheduling unit, a control instruction parsing unit, and a matrix conversion control unit; The pipeline scheduling unit is divided into three stages: the first stage input port includes input cache and weight cache; the second stage includes matrix conversion of input activation and weight, and Winograd convolution; the third stage output port includes intermediate value cache and output cache; The control instruction parsing unit parses the written control instruction, the parsed control instruction including the cached data structure parameters, convolution size, and step size, and sends it to different modules or units according to requirements; The matrix conversion control unit receives the instruction of the control instruction parsing unit and controls the input activation, weight and PE array calculation results to perform matrix conversion; The Winograd convolution module includes a register array unit, a MUX unit, a convolution control unit, a PE array unit, an addition tree unit, a ReLU, and a POOL unit; The register array unit caches the result of matrix conversion; The MUX unit implements data multiplexing by selecting corresponding rows and columns to the PE array; The convolution control unit is used for convolution control; The PE array unit is a multiplier array implemented by DSP, which completes the point multiplication calculation of Winograd domain; The addition tree unit accumulates the calculation results and implements the bias addition function; The ReLU and POOL units implement activation functions and pooling functions; The cache module includes an input cache unit, a weight cache unit, an intermediate value cache unit, and an output cache unit; The input cache unit caches input activations read from off-chip, caches block-based input activations according to top-level control parameters, and performs zero padding and duplicate row caching. The weight cache unit caches weight parameters; The intermediate value cache unit caches the calculation results of the blocks for the addition tree unit to read and complete the accumulation; The output buffer unit buffers the final calculation result and stores it off-chip via DMA.
4. The method for designing a fast convolution and caching mode for a convolutional neural network according to claim 2, wherein: Reconfigurable designs that achieve matrix transformation by configuring parameters include: The conversion matrix of the two-dimensional Winograd algorithm includes A, G, and B. The composition of the conversion matrix is only related to the convolution size of the Winograd algorithm and the size of the input feature map. There are repeated parameters between the conversion matrices. By configuring the parameters s and s0s1s2, a small amount of memory and resources are occupied to achieve a reconfigurable design of the conversion module. The derivation of the input feature map conversion matrix B is as follows when the input activation size is 4*4: Matrix changes in other sizes are achieved by adding a few registers.
5. The method for designing a fast convolution and caching mode for a convolutional neural network according to claim 1, wherein: Also includes: During the dot product process, for the position where the weight parameter is zero, the multiplication calculation of the parameter will be turned off. For the position where the weight parameter is not zero, the multiplication calculation of the feature map data is skipped to achieve a low-power design for convolution calculation.
Citation Information
Patent Citations
A universal convolutional neural network accelerator based on a one-dimensional pulsation array
CN109934339A
Accelerator of deep convolutional neural network based on FPGA
CN112949845A
Convolutional neural network hardware accelerator based on Winograd algorithm and calculation method
CN113255898A