An Ultra-Low Complexity Max Pooling Backpropagation Architecture for CNN

By directly reading the maximum non-zero numerical amount from the input matrix and designing the Matrix mapping logic operation, the problem of excessive resource consumption in maximum pooling backpropagation is solved, and more efficient hardware implementation frequency and resource conservation is achieved.

CN116227576BActive Publication Date: 2025-06-17SHANDONG HAILIANG INFORMATION TECH RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211495389.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-27
Publication Date
2025-06-17
Estimated Expiration
2042-11-27

AI Technical Summary

Technical Problem

In maximum pooling backpropagation, traditional recovery matrix algorithms lead to excessive consumption of storage resources and computing resources, especially when the recovery matrix is ​​large and there are many channels.

Method used

By directly reading the maximum non-zero numerical amount inside the convolution window from the input matrix, designing the Matrix mapping logic operation, determining whether the non-zero position is within the current convolution kernel range, and selecting a non-zero value for convolution operation to avoid restoring to the original matrix.

Benefits of technology

The data storage solution is optimized, the storage resource consumption is reduced, the hardware implementation complexity is reduced, the hardware implementation frequency is improved in reverse maximum pooling, and the computing resource consumption is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116227576B_ABST
    Figure CN116227576B_ABST
Patent Text Reader

Abstract

The present invention discloses an ultra-low complexity max pooling backpropagation architecture for CNN. The main objective of the present invention is to solve the problem that in the max pooling backpropagation of convolutional neural networks, when using the traditional recovery matrix algorithm, most of the values in the recovery matrix are zero, and directly storing the recovery matrix C will cause a large waste of storage resources. According to the sparse data characteristics in the input matrix, the present invention is implemented using an improved algorithm. It designs a Matrix mapping logical operation to directly obtain the mapping address of the input matrix in the convolutional kernel window and determines whether it is located in the current window, without restoring it to the original matrix. It optimizes the data storage scheme, selects non-zero values for convolution hardware implementation, thereby improving the hardware implementation frequency in the max pooling backpropagation, reducing the consumption of storage resources, lowering the hardware implementation complexity, and reducing the consumption of computing resources by selecting non-zero values to complete the convolution operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and relates to a method for implementing ultra-low complexity of inverse max pooling based on FPGA in a convolutional neural network. Background Art

[0002] In recent years, CNN has been successfully applied to fields such as image recognition and natural language processing. The CNN architecture mainly consists of a convolutional layer, a pooling layer, and a fully connected layer. Among them, the pooling layer can be divided into average pooling and max pooling. Max pooling can reduce the error caused by the deviation of the estimated mean, thereby retaining the texture information of the picture. Its calculation method is as follows:

[0003]

[0004] Where data i represents the input matrix element N of the pooling layer pool is the window size of the pooling layer. Since max pooling has less computational complexity than average pooling, it has been applied to networks such as AlexNet, VGG, and GooglNet.

[0005] During the backpropagation process, zero values are filled by upsampling according to the positions of non-zero points in the forward direction, and the calculation of the inverse max pooling layer is completed. As shown in Table 1, since the restored matrices obtained by upsampling inside each network are large and the number of channels is large, for example, in the VGG network, the size of the restored matrix reaches 112×112 and the number of channels is 128. If data is stored according to the size of the restored matrix, it will cause excessive consumption of storage resources and computing resources.

[0006] Table 1

[0007] Summary of the Invention

[0008] The main object of the present invention is to solve the problem that when using the traditional restored matrix algorithm for backpropagation of max pooling, the restored matrix C of F×F is spliced and stored in the RAM according to the convolutional kernel size N conv ×N conv to complete the convolution operation of F×F and N conv ×N conv . At this time, the depth of the RAM memory for each channel is F×F, and most of the matrix is zero values. Directly storing the matrix C will cause a large waste of storage resources.

[0009] To solve the above problems, the technical solution adopted by the present invention is specifically as follows:

[0010] An ultra-low complexity max pooling backpropagation architecture for CNN is as follows: Let the input matrix E be R×R;

[0011] 1) Simultaneously read Q largest non-zero numerical values inside a convolution window directly from the RAM with the input matrix E of size R×R, i.e., the input data E(1) to E(Q), where the matrix E' is obtained from the matrix E through Equation (1).

[0012]

[0013] Where m and n are the row and column coordinates of the matrix E, i and j are the row and column coordinates of the matrix E' respectively, and Z is the number of elements with the same row and column coordinates in the matrix E.

[0014] 2) According to the padding size P in the transposed convolution, change the row and column coordinates r and c of the original matrix E' through Equations (2) and (3) to obtain the padded row and column coordinates r′ and c′.

[0015] r′ = r + P (2)

[0016] c′ = c + P (3)

[0017] The non-zero positions mapped to the corresponding convolution kernel are obtained by subtracting the column offset c row and c col from the internal data corresponding addresses of the matrix A'; the column offsets c row and c col are set according to the convolution stride, with the initial value being 0.

[0018] Judge whether the non-zero position mapped to the corresponding convolution kernel is less than or equal to the current convolution kernel internal coordinate. If so, within the corresponding convolution kernel range, form the logical operation module in Matrix mapping accordingly; at the same time, select the non-zero values of the input matrix at the corresponding coordinates of the convolution window through Q selectors, and obtain the address selections read_address1 to read_addressQ for the non-zero positions corresponding to the convolution kernel.

[0019] 3) Read the weight data K(1) to K(Q) corresponding to the input data E′(1) to E′(Q) from the weight RAM, multiply the weight data by the input data at the corresponding positions, and complete the operation of one convolution window through Q multipliers and Q - 1 adders.

[0020] 4) According to Equation (4), through the size of the padded matrix,

[0021]

[0022] Every time a clock cycle passes, c row is incremented once according to the convolution stride. When the size of c row is equal to the size of the padded matrix, c row is reset to zero, and c colAccumulate once according to the convolution step size, and loop the operations in steps 1) to 3) until c col The process of changing c when the size is equal to the size of the padded matrix row and c col The offset of, synchronize the step operation of the convolution window, map it to the horizontal and vertical coordinates of the convolution kernel, and calculate (F + 2P - N conv ) / S conv convolution windows and pass through the same hardware structure until the entire matrix convolution operation is completed. Among them, F is the size of the restored matrix C, P represents the padding size in the transposed convolution, N c o nv represents the size of the convolution kernel, and S c o nv represents the step size of the transposed convolution.

[0023] Advantages of the present invention:

[0024] According to the data characteristics of sparsity in the input matrix, the present invention is implemented by an improved algorithm, designs a Matrixmapping logical operation, directly obtains the mapping address of the input matrix in the convolution kernel window, and determines whether it is located in the current window, without restoring to the original matrix, optimizes the data storage scheme, selects non-zero values for convolution hardware implementation, thereby improving the hardware implementation frequency in the reverse max pooling, reducing the storage resource consumption, reducing the hardware implementation complexity, and selecting non-zero values to complete the convolution operation reduces the consumption of computing resources. Description of the drawings

[0025] Figure 1 is the hardware schematic diagram of the reverse max pooling restored matrix

[0026] Figure 2 is the algorithm schematic diagram of the present invention;

[0027] Figure 3 is the hardware structure diagram of the improved scheme. Specific implementation manners

[0028] The technical solution of the present invention will be further explained and described below in the form of specific embodiments.

[0029] Based on the principle of error backpropagation, the error function δ of layer L is obtained according to Equation 3.1 L , where δ L+1 is the error function corresponding to layer L + 1, a L-1 is the input matrix of layer L, and then the sensitivity to layer L - 1 is obtained through Equation 3.2, W L-1 is the weight of layer L - 1, z L-1 is the output of layer L - 1 without passing through the activation function, and σ'(z L-1 ) is the derivative of the activation function of layer L - 1.

[0030]

[0031]

[0032] Therefore, define the error function δ in the reverse pooling layer L+1 The output matrix of is the input matrix E, and the error function δ L The recovery matrix C of is F×F, W L-1 The size of the weight convolution kernel K of is N conv ×N conv and the stride is S conv and the padding is P. If the stride in the max pooling during forward inference is less than the size of the kernel in the pooling layer, there will be an overlapping value-taking situation. Therefore, during the reverse recovery process, the elements with the same corresponding coordinates are superimposed. If the stride in the max pooling during forward inference is greater than or equal to the size of the kernel in the pooling layer, there will be no overlapping value-taking situation, that is, the matrix E passes through Equation 3.3 to obtain the matrix E'. Here, i and j respectively represent the positions of the corresponding elements in E'. The matrix E' is upsampled through Equation 3.4 to obtain the matrix C. Then, after edge padding P, Equation 3.5 is convolved with the convolution kernel K and the reverse max pooling is completed to obtain the final output matrix D.

[0033]

[0034]

[0035] D = conv(C, K) (9)

[0036] Taking the reverse max pooling of the fifth layer of AlexNet as an example, the error matrix with an input of 6×6 and 256 channels is restored to a matrix of 13×13×256. At this time, the input error matrix 6×6×256 has completed the superposition operation of the same maximum pixel point positions. Therefore, the amount of non-zero data in a single channel is less than or equal to 36. After being restored to 13×13×256, it is padded to a matrix of 15×15×256 and convolved with a 3×3 convolution kernel. The 6×6 matrix in each channel is restored to a 15×15 error matrix through upsampling and padding operations and stored in the RAM, and the convolution operation with the 3×3 convolution kernel is completed. Therefore, there are more than 79% of the stored zero values in the memory, resulting in excessive consumption of storage and computing resources.

[0037] The upsampling operation of the backpropagation of max pooling. The number of required multipliers after convolving each restored matrix is shown in Equation 3.6

[0038] N mult = N conv 2 (10)

[0039] The number of adders is as shown in Equation 3.7

[0040] N add = N conv 2 - 1 (11)

[0041] The capacity M of the memory is

[0042] M = R 2 × W width + F 2 × W width (12)

[0043] where W width is the quantization bit width of the corresponding data.

[0044] As shown in the appendix Figure 3 In the present invention, a super - low - complexity max - pooling backpropagation architecture for CNN is as follows: Let the input matrix E be of size R×R, and the convolution kernel size be N conv × N conv , and the size of the recovery matrix C is F×F;

[0045] 1) Directly read Q maximum non - zero numerical quantities inside a convolution window simultaneously from the RAM with the size of the input matrix E' being E×E, that is, the input data E(1) to E(Q), where the matrix E' is obtained from the matrix E through Equation (1)

[0046]

[0047] where m, n are the row and column coordinates of the matrix E, i, j are the row and column coordinates of the matrix E' respectively, and Z is the number of elements with the same row and column coordinates in the matrix E.

[0048] 2) According to the padding size P in the transposed convolution, change the row and column coordinates r and c of the original matrix E' to obtain the padded row and column coordinates r′ and c′ through Equations (14) and (15)

[0049] r′ = r + P (14)

[0050] c′ = c + P (15)

[0051] Subtract the internal data corresponding address of the matrix E' from the row and column offsets c row and c col to obtain the non - zero positions mapped to the corresponding convolution kernel; the row and column offsets c row and c col are set according to the convolution stride, and the initial value is 0;

[0052] Determine whether the non-zero positions mapped to the corresponding convolution kernels are less than or equal to the current internal coordinates of the convolution kernels. If so, within the range of the corresponding convolution kernels, a logical operation module in the Matrix mapping is formed accordingly; at the same time, the non-zero values of the input matrix at the corresponding coordinates of the convolution window are selected through Q selectors, and the addresses read_address1 to read_addressQ corresponding to the non-zero positions of the convolution kernels are obtained.

[0053] 3) Read the weight data K(1) to K(Q) corresponding to the input data E′(1) to E(Q) from the weight RAM. Multiply the weight data by the input data at the corresponding positions, and complete the operation of one convolution window through Q multipliers and Q - 1 adders.

[0054] 4) According to Equation (16), through the size of the padded matrix

[0055]

[0056] Every time a clock cycle passes, c row is accumulated once according to the convolution stride. When c row is equal to the size of the padded matrix, c row is reset to zero. c col is accumulated once according to the convolution stride, and the operations in steps 1) to 3) are looped until c col is equal to the size of the padded matrix, during which the offsets of c row and c col are changed to synchronize the step operation of the convolution window, which is mapped to the horizontal and vertical coordinates of the convolution kernel, and (F + 2P - N conv ) / S conv convolution windows are calculated in sequence and passed through the same hardware structure until the entire matrix convolution operation is completed. Here, M is the size of the restored matrix C, P represents the padding size in the transposed convolution, N represents the size of the convolution kernel, and S represents the stride of the transposed convolution.

[0057]

[0058] Corollary: Let Q represent the maximum number of non-zero values corresponding to each transposed convolution window in the restored matrix C. Since the parameters of the max pooling layer and the adjacent convolution layer in different CNN networks are different, the Q value in the transposed convolution operation is different. Table 2 lists the Q values corresponding to different pooling kernels, convolution kernels, and their respective stride sizes. N pool where is the size of the max pooling kernel in the forward inference, and the stride is S pool , N conv and S conv represent the size and stride of the convolution kernel respectively.

[0059] Table 2

[0060]

[0061]

[0062] Use the selector to select the non-zero point positions and complete the convolution operation of a small window with the convolution kernel. Each clock cycle requires Q multipliers to complete a convolution window operation. There is no need to store the data after restoring the matrix. It takes (F + 2P - N conv ) / S conv + 3 clock cycles. Adopting this scheme, the restoration of the matrix and the convolution operation can be completed with a small amount of logical operations. Among them, the number of multipliers processed per clock cycle N c_mult is at most

[0063] N c_mult = Q (18)

[0064] The number of adders N c_add is

[0065] N c_add = Q - 1 (19)

[0066] The storage resources R consumed conv is

[0067] R conv = N conv × N conv × W width (20)

[0068] The number of clock cycles N consumed cycle is

[0069] N cycle = [F + 2P - N conv 2 + 4 (21)

[0070] According to the proposed algorithm, the maximum pooling algorithm in Figure 1 can be simplified to Figure 2 . Since the Q value is 4, this scheme only requires 4 multipliers. Only calculate the non-zero positions, take values according to the convolution window size, and each time take out the number of the largest non-zero values in the corresponding convolution window, judge the non-zero point positions and map them to the corresponding convolution window positions for convolution operations, as shown in the appendix Figure 2 ​As shown, taking E equal to 6, the size of F being 13, the size of N being 3, and the size of P being 1 as an example, the maximum number of non-zero values corresponding to a convolution window is 9. Then, 9 data are input in one clock cycle to complete the judgment of the 3×3 convolution kernel at the non-zero positions. The data at the non-zero positions are judged through a selector, and convolution operations are performed on the convolution kernel corresponding to the non-zero positions. In one clock cycle, the operation of one convolution window is completed. Since there is no restoration matrix in this scheme, no data storage is required, and no data storage resources are consumed. Therefore, only a small amount of logic resources are consumed to complete data selection and data multiplication and addition operations to complete the operation of one convolution window.

[0071] Under the Arria 10 hardware platform, two hardware implementation schemes of inverse max pooling in multiple architectures are tested. Table 3 shows the comparison of the consumed hardware resources.

[0072] Table 3

[0073]

[0074] Implementing inverse max pooling with low complexity has a relatively high internal operating frequency, and the original algorithm stores the restored matrix, so a large amount of storage resources for matrix C are additionally increased. Since a large amount of data selection is required in the original algorithm, affected by the size of the subsequent deconvolution kernel and the forward max pooling, when the number of non-zero values input per clock is closer to the size of the convolution kernel, the improved algorithm saves less than 10% of the ALM resources compared to the original algorithm. When the number of non-zero values input per clock is less compared to the size of the deconvolution kernel, more than 60% of the ALM resources can be saved. Therefore, the more hardware resources the improved algorithm saves, the more obvious the advantage is.

Claims

1. An ultra - low complexity max - pooling backpropagation architecture for CNN, characterized in that, The architecture is as follows: Let the input matrix E be of size R×R, and the convolutional kernel size be N conv ×N conv , and the size of the restored matrix C is F×F; 1) Simultaneously read Q = N conv × N conv maximum non-zero numerical quantities within a convolution window directly from the RAM of matrix E' with size R×R, that is, the input data E(1) to E(Q), where matrix E' is obtained from the input matrix E through Equation (1). where m and n are the row and column coordinates of matrix E, i and j are the row and column coordinates of matrix E' respectively, and Z is the number of elements with the same row and column coordinates in matrix E; 2) According to the padding size P in the deconvolution, change the row and column coordinates r and c of the original matrix E' by equations (2) and (3) to obtain the padded row and column coordinates r' and c', r′=r+P (2) c′ = c + P (3) The corresponding address of the internal data of matrix A' and the row-column offset c row and c col are subtracted to obtain the non-zero positions mapped to the corresponding convolution kernels; the row-column offsets c row and c col are set according to the convolution stride, and the initial value is 0; Determine whether the non-zero positions mapped to the corresponding convolution kernel are less than or equal to the current internal coordinates of the convolution kernel. If so, within the corresponding convolution kernel range, form the logical operation module in Matrix mapping accordingly; at the same time, select the non-zero values of the input matrix at the corresponding coordinates of the convolution window through Q selectors, and obtain the address selections read_address1 to read_addressQ of the non-zero positions corresponding to the convolution kernel; 3) Read the weight data K(1) to K(Q) corresponding to the input data E′(1) to E′(Q) from the weight RAM, multiply the weight data by the input data at the corresponding positions, and complete the operation of one convolution window through Q multipliers and Q - 1 adders; 4) According to equation (4), through the size of the padded matrix, Every time a clock cycle passes, c row is accumulated once according to the convolution stride. When c row reaches the size of the padded matrix, c row is reset to zero. c col is accumulated once again according to the convolution stride. The operations in steps 1) to 3) are looped until c col reaches the size of the padded matrix, during which the offsets of c row and c col are changed to synchronize the movement operation of the convolution window, which is mapped to the horizontal and vertical coordinates of the convolution kernel. Calculate (M + 2P - N conv ) / S conv convolution windows in sequence and pass them through the same hardware structure until the entire matrix convolution operation is completed. Here, F is the size of the restored matrix C, P represents the padding size in the transposed convolution, N conv represents the size of the convolution kernel, and S conv represents the stride of the transposed convolution.

Citation Information

Patent Citations

  • Method for accelerating convolution neutral network hardware and AXI bus IP core thereof

    CN104915322A

  • Method and device for performing deconvolution processing on feature data by using convolution hardware

    CN112686377A