Structured Fine-Grained Binary Network Double Pruning Method and Sparse Accelerator Device
Through the structured fine-grained dual pruning method and sparse binary network accelerator device, the efficient deployment problem of binary neural networks on edge devices is solved, the parameters and calculation amount are saved, and high precision is maintained, which improves the efficiency of the hardware accelerator.
Patent Information
- Application Number
- CN202211070620.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-31
AI Technical Summary
It is difficult for the existing technology to efficiently deploy large-scale deep neural networks on edge devices, and traditional pruning algorithms are not suitable for binary neural networks, resulting in high computing power consumption and large memory overhead. At the same time, fine-grained pruning brings hardware deployment challenges.
A structured fine-grained double pruning method is proposed. By generating a structured cropping index of the convolution weight channel and a plane mask for outputting feature maps, combined with Gumbel-Softmax reparameterization technology, it cuts redundant calculations, and designs a pipelined sparse binary network accelerator device based on inter-layer parallelism.
It effectively reduces the model complexity and computational amount of binary neural networks, and at the same time improves the parallelism of hardware implementation, achieves significant savings in parameters and computational amounts, maintains high accuracy, and improves the efficiency of hardware accelerators.
Smart Images

Figure CN115438774B_ABST
Abstract
Description
Technical Field
[0001] The present invention proposes a structured fine-grained pruning scheme conducive to efficient inference of a binary neural network (BNN) and the design of a corresponding sparse BNN accelerator device for high-precision and efficient inference of the binary neural network, belonging to the field of co-design of software and hardware for network compression and neural network accelerators. Background Art
[0002] In recent years, in order to achieve high precision, deep neural networks have developed to a scale with hundreds of layers and over a million parameters. Obviously, it is very difficult to deploy these large-scale networks on edge devices with power consumption limitations and low latency requirements. To solve this problem, many model compression techniques have been proposed. Among them, pruning and quantization are two widely used methods. Quantization uses low-bit data to represent the parameters and activations of the network. The activations and parameters of a binary neural network (BNN) are quantized to 1 bit, which is considered an extreme form of a quantized network. In addition, BNN has received increasing attention in recent years because after quantization to 1 bit, the original multiply-accumulate operations in the real number domain become Xnor+Popcount in the binary domain, which greatly reduces the computational power consumption, and the 1-bit data width also saves a large amount of memory overhead. Pruning is another compression method independent of quantization, which is used to remove redundant parameters and connections in the network. However, few works combine BNN and pruning techniques to explore and utilize the redundancy in BNN. There are two challenges. One is that traditional pruning algorithms applied to full-precision convolutional neural networks regard "0" as sparse, while in BNN, "0" actually represents "-1", so it is not appropriate to directly adopt traditional pruning algorithms. The other is that the accuracy of BNN is relatively low. Quantization has already caused the network to lose a lot of information, and traditional pruning algorithms are mostly coarse-grained. They ignore the differences between different neurons, and applying these algorithms in BNN may lead to unacceptable accuracy degradation. Fine-grained pruning can lead to higher sparsity and higher accuracy. However, fine-grained pruning often brings irregularity, posing a considerable challenge to hardware deployment. Summary of the Invention
[0003] Object of the Invention: In order to combine pruning with BNN, promote the realization of the limit of a compact network, consider the efficiency of hardware implementation while protecting the accuracy, a structured and fine-grained double pruning scheme is proposed. And in order to accelerate the inference of the sparse model after double pruning in actual hardware deployment, the present invention also proposes the architecture design of a corresponding sparse binary neural network accelerator, which can efficiently utilize the sparsity in the model caused by the pruning scheme.
[0004] Technical Solution:
[0005] A structured fine-grained binary network double pruning method. For a specific binary convolutional layer, according to the structured pruning index of the weight channels, the corresponding channels of the convolutional weights are pruned during inference; according to the plane mask that controls fine-grained pruning, the calculation of the output feature positions indicated by the zero values of the plane mask is masked during inference. The joint action of the two pruning methods enables the convolutional layer to only perform calculations on specific input channels at specific output positions, thereby greatly reducing the computational amount of the network. The method includes:
[0006] 1. Generate the structured pruning index for each group of convolutional weight channels. The structure is reflected in that only one channel in every N consecutive weight channels is retained for calculation;
[0007] 2. Generate the plane mask representing the importance of the output feature maps of each layer;
[0008] 3. Prune the convolutional weight channels corresponding to each binary convolutional layer according to the channel pruning index;
[0009] 4. During inference, mask the calculation of the output features at the positions indicated by the plane mask according to the plane mask of the feature map.
[0010] Assume that the size of the input feature map of a specific binary convolutional layer is [IC, H, W]; the size of the weights is [OC, IC, K, K]; the size of the output feature map is [OC, H’, W’], where H and W are the sizes of the input feature map, and H’ and W’ are the sizes of the output feature map; OC is the number of output channels; IC is the number of input channels; K is the convolutional kernel size. This convolutional layer has a total of OC groups of weights, and each group of weights has IC channels, so the total number of channels of the weights of this convolutional layer is OC×IC. The steps for generating the pruning index for each group of weight channels are as follows:
[0011] Step 1: Generate a parameter vector Pc with a length of OC×IC;
[0012] Step 2: Rearrange Pc into a matrix Mc with a size of [OC×(IC / N), N];
[0013] Step 3: Reparameterize Mc using the Gumbel-Softmax method. The method is as follows:
[0014] First, use the Softmax function σ to convert the N elements in each row of Mc into sampling probabilities P, which means the probabilities of being selected among the N convolutional kernels.
[0015] P i =σ(Mc i ) for i = 1, 2,..., N;
[0016] For any row of Mc, which has N elements [x1, x2,..., x N , use Gumbel-Softmax reparameterization to obtain a new row vector S = [S1, S2,..., S N , and the specific method refers to the sampling formula:
[0017]
[0018] Among them, the row vector S is the reparameterization result of each row of Mc during inference and training, and Gi is a Gumbel random number.
[0019] Step 4: If it is training, rearrange the reparameterized matrix into a tensor of size [OC, IC, 1, 1], and expand it by replicating it K*K times along the plane convolution kernel dimension, and multiply the expanded tensor of size [OC, IC, K, K] with the weights. In this way, the parameter vector and the weights will be updated together during the training process; if it is inference, skip Step 4 and continue with Step 5-6;
[0020] Step 5: Binarize the reparameterized matrix to obtain a binary matrix BMc, which has the same size as Mc and has OC×(IC / N) rows and N columns. The binarization method is as follows: set the element of BMc at the position of the maximum value in each row of Mc to 1, and set the remaining N-1 positions to 0;
[0021]
[0022] Step 6: Rearrange BMc into a tensor BPc of size [OC, IC, 1, 1], and this tensor is the index for cropping all OC groups of weight channels of this binary convolution layer.
[0023] Note: If IC is not divisible by N, then (IC / N) is rounded up.
[0024] The steps for generating the plane mask of the output feature map are as follows:
[0025] Step 1: For the intermediate feature map of a certain layer in the network, introduce a bypass binary convolution layer, which has only one group of weights. After this group of weights is convolved with this intermediate feature, a plane feature map Ms is obtained; by adjusting the convolution stride, the plane feature map is aligned with the target output feature map in terms of size H’ and W’.
[0026] Step 2: Perform batch normalization on this plane feature map;
[0027] Step 3: Use the Gumbel-Softmax method to reparameterize this planar feature map. Since it is a binary classification problem to determine whether the data at each point of this planar feature map is 1 or 0 at this time, the sigmoid function is used to sample Ms; the formula is as follows:
[0028] P i = sigmoid(Ms i ) for i = 1, 2,..., H'×W'
[0029]
[0030] Step 4: If it is training, perform matrix dot multiplication on the reparameterized planar feature map obtained in Step 3 and the planar features in each channel of the output feature map, so as to update the weight parameters of the bypass convolution layer during the training process; if it is inference, skip Step 4 and continue to Step 5;
[0031] Step 5: Quantize the reparameterized planar feature obtained in Step 3 Figure 2 to generate a planar mask. The specific method is: compare each element of the planar feature map with the threshold 0.5. If it is greater than 0.5, the corresponding element of the planar mask is taken as 1, otherwise it is taken as 0, so as to obtain a binary planar mask.
[0032] To protect the accuracy of the binary network after double pruning, the method for obtaining the best sparsity of the planar mask is: for the generated and already reparameterized planar feature map, introduce a trainable parameter and add it to all the data of the planar feature map, so as to change the data distribution of the planar feature map before the threshold judgment step;
[0033] Moreover, a new loss component is introduced during the training process to obtain the best sparsity of the planar mask during the training. The specific method is to add a Loss related to the sparsity of the planar mask, that is, L cls to the classification Loss of the network, that is, L prun , and the final training Loss becomes:
[0034] Loss = λ1L prun + λ2L cls ;
[0035] where, L cls is the classification loss, L prun is the pruning loss, δ is the average value of the sparsity of all planar masks, pr is the preset value of the required sparsity, the smaller L cls is, the smaller the classification loss, and L prunThe smaller it is, the closer the average sparsity of the existing planar mask is to the preset value. The hyperparameters λ1 and λ2 are used as weights to balance the two Losses and find the optimal sparsity of the planar mask.
[0036] A pipelined sparse binary network accelerator device based on inter-layer parallelism that adapts to structured pruning of weight channels and planar mask pruning within feature channels, which includes a sparse binary convolutional layer computing engine, a first-layer computing engine, a max-pooling layer computing engine, and a fully connected layer computing engine;
[0037] The pipelined accelerator architecture based on inter-layer parallelism requires a customized computing engine to be instantiated for each individual layer, and the current layer's computation starts as soon as the previous layer begins to produce output. In this pipelined manner, the overall latency of executing multiple network layers is effectively reduced;
[0038] The sparse binary convolutional layer computing engine includes a sliding window unit, a binary convolution computing unit, a channel selection unit, a planar mask generation unit, and a planar mask data channel. Among them,
[0039] The sliding window unit: includes a row buffer and a sliding window, which are used to receive and cache the input data stream, and shield redundancy according to the planar mask, and output the data required for convolution calculation in the correct order;
[0040] The binary convolution computing unit: includes an input bit-width converter, a register-based local memory, a PE array, and an output bit-width converter. The PE array is composed of PEs, and each PE consists of a certain number of XNOR gates, an adder tree for Popcount, a shift register, and a threshold comparator for performing Batchnorm and Sign functions. The binary convolution computing unit composed of the above is used to complete sparse binary convolution calculation;
[0041] The channel selection unit: includes a certain number of N-to-1 MUXes. Used in conjunction with the PEs, it is used to grab the corresponding input channels according to the N-to-1 channel pruning index and output them to each PE for calculation; since the input bit-width of the channel selection unit is configurable, multiple channels can be synchronously screened out and supplied to multiple PEs for simultaneous calculation, reflecting the advantages of N-to-1 channel structural pruning.
[0042] The planar mask generation unit: is used to generate the planar mask required for the target output feature map and coordinate the flow of input and mask data;
[0043] A planar mask data channel: This channel is connected to all computing engines sharing the same planar mask, providing mask data for the computing engines to shield redundant computations. After the sliding window unit and the PE array load the mask data from the planar mask data channel, they need to copy the data and write it back to the channel for subsequent use by the computing engines.
[0044] The first-layer computing engine is similar to the sparse binary convolution layer computing engine, except that: 1. The input is a 3-channel image and both the input pixels and weights are fixed-point numbers; 2. The XNOR gates in the binary convolution computing unit need to be replaced by multipliers; 3. Neither the sliding window unit nor the PE array needs to load mask data for judgment.
[0045] Placing the max pooling layer after the batch normalization layer is equivalent to placing it after the binarization function. Therefore, the fixed-point maximum value comparison can be replaced by a binary OR operation. The max pooling layer computing engine consists of a sliding window unit with a window size of 2×2 and an OR gate array.
[0046] Since the input and the weights of the fully connected layer have the same dimensions, the fully connected layer computing engine does not require a sliding window unit and consists only of binary convolution computing units; among them, the XNOR gates are replaced by sign converters to perform the fully connected layer convolution computation where the input is binary and the weights are fixed-point.
[0047] Beneficial effects: The structured fine-grained double pruning method proposed in the present invention not only effectively reduces the model complexity of BNN while protecting the accuracy, but also provides high parallelism that is easy to implement in hardware. On the Cifar-10 dataset, applying the double pruning method of the present invention to all convolutional layers of the binary models ResNet-18 and Vgg-Small can achieve parameter savings of 69.4% and 67.5% respectively, and computational savings of 85.4% and 80.2%, with only a 2.3% and 1.8% accuracy drop. Among them, N and pr required by the double pruning method are set to 4 and 60% respectively. On the ImageNet dataset, applying the double pruning method of the present invention to all convolutional layers of the binary models ResNet-18 and ResNet-34 realizes approximately 44% parameter savings and 70% computational savings, with only a 5% accuracy drop. Implementing the accelerator architecture on the FPGA platform, for the sparse Vgg-Small model applying the double pruning method above, the accelerator can achieve an acceleration ratio of approximately 5.4 times compared to the unpruned dense binary network model. Description of the Drawings
[0048] Figure 1 It is a schematic diagram of the process of applying the double pruning method to a specific binary convolutional layer;
[0049] Figure 2 Schematic diagram of the process for generating a structured index of channels to be pruned
[0050] Figure 3 Schematic diagram of the process for generating a planar mask of the output feature map
[0051] Figure 4 Schematic diagram of the generation and application of a planar mask in an existing classification network structure: where (a) represents the case where the network structure is a residual block; (b) represents the case where the network structure is a direct connection type and the planar mask is applied to all convolutional layers between two downsampling layers; (c) represents the case where the network structure is a direct connection type and the planar mask is applied to the network topology across downsampling layers
[0052] Figure 5 Schematic diagram of a pipeline sparse binary neural network accelerator architecture based on inter-layer parallelism
[0053] Figure 6 Schematic diagram of a sparse binary convolutional layer computing engine
[0054] Figure 7 Schematic diagram of a max pooling layer computing engine Detailed implementation manner
[0055] The present invention will be further explained below with reference to the accompanying drawings
[0056] A structured and fine-grained double pruning scheme includes generating a structured pruning index for each group of convolutional weight channels, where the structure is reflected in that only one channel in every N consecutive weight channels is retained for calculation; generating a planar mask representing the importance of the output feature maps of each layer; pruning the convolutional weight channels corresponding to each binary convolutional layer according to the channel pruning index; and during inference, masking the calculation of the output features at the positions indicated by the planar mask of the feature map
[0057] Figure 1 Shows the specific process of applying the double pruning method to a specific binary convolutional layer. In this example, the number of input channels is 4, the number of output channels is 2, and N is set to 2. First, prune the convolutional kernel corresponding to index 0 according to the channel pruning index to be pruned; then, for each output point of the output feature map, determine whether the data of the planar mask at this position is 0. If so, cancel the calculation of this output point and output 0; if it is 1, when calculating the output result at this position, each pruned convolutional filter will grab one channel from every adjacent N input channels at the corresponding position of the input feature map (i.e., the important area shown in the figure) for calculation. The combined action of the two pruning algorithms enables this convolutional layer to only complete the calculation at specific input channels for specific output positions, greatly reducing the computational amount of the network
[0058] As Figure 2 , the steps for generating the structured index of each filter's channels to be pruned are as follows: It should be noted that the figure only shows the pruning process for one filter during inference, that is, OC = 1;
[0059] Step 1: Generate a parameter vector Pc with a length of OC × IC;
[0060] Step 2: Rearrange Pc into a matrix Mc with a size of [OC × (IC / N), N];
[0061] Step 3: Reparameterize Mc using the Gumbel-Softmax method;
[0062] Step 4: If it is training, rearrange the reparameterized matrix into a tensor with a size of [OC, IC, 1, 1], and expand it by replicating it K*K times along the planar convolution kernel dimension. Multiply the expanded tensor with a size of [OC, IC, K, K] by the weights. In this way, the parameter vector and the weights will be updated together during the training process; if it is inference, skip Step 4 and continue with Steps 5-6;
[0063] Step 5: Binarize the reparameterized matrix to obtain a binary matrix BMc with the same size as Mc, having OC × (IC / N) rows and N columns. The binarization method is as follows: Set the element of BMc at the position of the maximum value in each row of Mc to 1, and set the remaining N - 1 positions to 0;
[0064]
[0065] Step 6: Rearrange BMc into a tensor BPc with a size of [OC, IC, 1, 1]. This tensor is the index for pruning all OC groups of weight channels in this binary convolutional layer.
[0066] As Figure 3 , the steps for generating the planar mask of the output feature map are as follows: The figure only shows the process during inference;
[0067] Step 1: For the intermediate feature map of a certain layer in the network, introduce a bypass binary convolutional layer with only one group of weights. After convolving this group of weights with the intermediate feature, a planar feature map Ms is obtained. By adjusting the convolution stride, the planar feature map is aligned with the target output feature map in terms of size H' and W';
[0068] Step 2: Perform batch normalization on this planar feature map;
[0069] Step 3: Reparameterize this planar feature map using the Gumbel-Softmax method. Since it is a binary classification problem to determine whether the data at each point of this planar feature map is 1 or 0, the sigmoid function is used to sample Ms.
[0070] Step 4: If it is training, perform matrix dot multiplication on the reparameterized planar feature map obtained in Step 3 and the planar features in each channel of the output feature map, so as to update the weight parameters of the bypass convolutional layer during the training process; if it is inference, skip Step 4 and continue to Step 5.
[0071] Step 5: Quantize the reparameterized planar feature obtained in Step 3 Figure 2 to generate a planar mask. The specific method is as follows: compare each element of the planar feature map with the threshold 0.5. If it is greater than 0.5, the corresponding element of the planar mask is taken as 1, otherwise it is taken as 0, so as to obtain a binary planar mask.
[0072] As Figure 4 , generating a planar mask requires an additional bypass convolutional layer. When applying it to the existing classification network structure, it is divided into the following several situations:
[0073] As Figure 4 in (a), for each residual block in the residual network (such as ResNet-like), add a bypass binary convolutional layer required to generate a planar mask. Its input is the input of this residual block, and the output is a planar binary mask shared by all convolutional layers in the residual block.
[0074] As Figure 4 in (b), for a direct-connected network (such as VggNet-like), if it is required that all convolutional layers sharing a planar binary mask are sandwiched between two downsampling layers, only add a bypass binary convolutional layer with a Stride of 1, and the input is the output of the first downsampling layer.
[0075] As Figure 4 in (c), if it is required that all convolutional layers sharing a planar binary mask span a downsampling layer, add a bypass binary convolutional layer with a Stride of 2. Its input is the input of the first convolutional layer, and the output is used by the convolutional layer after the downsampling layer. At the same time, perform nearest neighbor interpolation on the output and give the interpolation result to the convolutional layer before the downsampling layer for use.
[0076] A layer-interleaved parallel pipelined sparse binary network accelerator device adapted to structured pruning of weight channels and planar mask pruning within feature channels, as Figure 5 shown. It includes a sparse binary convolutional layer computing engine, a first-layer computing engine, a max pooling layer computing engine, and a fully connected layer computing engine;
[0077] The pipelined accelerator architecture based on inter-layer parallelism requires a customized computing engine to be instantiated for each individual layer, and the current layer's computation starts as soon as the previous layer begins to produce output. In this pipelined manner, the overall latency of executing multiple network layers is effectively reduced;
[0078] The BNN accelerator architecture roughly allocates hardware resources to the computing engines of each layer according to the workload of each layer. Since the memory overhead required for the pruned binary network parameters is extremely small, these parameters can be stored on-chip. When the accelerator is started, the image to be processed is read from the external memory (DDR) to the first-layer computing engine through DMA (Direct Memory Access). A FIFO with an appropriate depth is placed between two adjacent computing engines to store the intermediate binary feature calculation results. Once each computing engine detects the data output by the previous layer's calculation fed into its input FIFO, it will start to execute the calculation of this layer;
[0079] The accelerator architecture adopts a widely used weight stabilization scheme for the reuse of input data.
[0080] Such as Figure 6 , the sparse binary convolution layer computing engine includes a sliding window unit, a binary convolution computing unit, a channel selection unit, a planar mask generation unit, and a planar mask data channel. Sliding window unit: It includes a row buffer and a sliding window, which are used to receive and cache the input data stream, and mask the redundancy according to the planar mask and output the data required for convolution calculation in the correct order; Binary convolution computing unit: It includes an input bit-width converter, a register-based local memory, a PE array, and an output bit-width converter, which are used to complete the sparse binary convolution calculation;
[0081] The channel selection unit is used in conjunction with the PE to grab the corresponding input channels according to the indexes of the channels to be pruned and output them to each PE for calculation; The planar mask generation unit is used to generate the planar mask required for the target output feature map and coordinate the flow of input and mask data; A planar mask data channel is used to synchronize the flow of data stream and mask data in the pipelined architecture and provide the mask data to the sliding window unit and the PE array at the corresponding moment to mask the redundant calculation. This channel is connected to all computing engines sharing the same planar mask. After the sliding window unit and the PE array load a mask data from the planar mask data channel, they need to copy the data and write it back to the channel for subsequent computing engines to use.
[0082] The row buffer is composed of K shift registers connected end to end. The width and depth of each register are IC and W respectively. The data processing method of the sliding window unit is as follows: When the computing engine starts working, the sliding window unit reads the input data stream from the input FIFO with a width of IC in a row-first and column-second scanning order. The input data stream flows into one end of the row buffer and flows out from the other end. Once K rows of input data are loaded into the row buffer, a sliding window of size K×K starts to move in the row buffer. Before each movement of the sliding window, it obtains a binary mask data from the planar mask data channel to determine whether the data in the current sliding window is valid. If this mask data is 0, the K×K input features in the current sliding window will not be passed to the binary convolution calculation unit for calculation. At the same time, the sliding window will move to the next position. In this way, the input in the unimportant area will be filtered. The K×K inputs corresponding to the mask data of 1 will be sequentially output to the binary convolution calculation unit for convolution calculation. When the sliding window reaches the bottom of the row buffer, the sliding window unit loads new input row data and starts the next sliding cycle.
[0083] The binary convolution calculation unit mainly utilizes the parallelism of the input and output channels; Tm and Tn are the parallelization parameters of this unit, used to configure the computing resources it requires. Each PE in the PE array consists of Tm XNOR gates, an adder tree for Popcount, a shift register, and a comparator. The data processing method of the binary convolution calculation unit is as follows: First, the input bit-width converter sequentially receives the binary vector with a length of IC output by the sliding window and divides it into multiple segments with a length of Tn; then these data segments are broadcast to all PEs to perform binary inner product operations with the binary weights stored in the local memory. Before the calculation, the same PE will also load a mask data from the planar mask data channel to determine whether it is 1. If it is 1, the normal inner product operation will be performed. If it is 0, the operation will be cancelled and the data 0 will be directly output to the comparator. The output of the adder tree is buffered in the register for further accumulation of partial sums. After the calculation of K×K×IC / Tn input segments is completed, the accumulated value of each PE is shifted and then compared with the bias generated by the BN layer to obtain a 1-bit output. Therefore, each cycle of the PE array generates a Tm-bit vector, and then this vector is output and provided to the output bit-width converter. The register-based local memory with a size of K×K×IC is used to cache the input data to calculate the next Tm output vectors. When the output can be spliced into an OC-bit vector, the binary convolution calculation unit outputs this vector to the output FIFO with a width of OC. To minimize the storage of partial sums, the binary convolution calculation unit only starts the convolution operation at the next output feature point position after all the data in the current sliding window has been calculated.
[0084] The channel selection unit consists of Tn N-to-1 MUXes, and the output of the MUX is controlled by the index of the channel to be pruned. The channel selection unit can be used in combination with the PE. In this way, the output bit width of the input bit width converter can be set to 4×Tn, so as to increase the throughput of the PE array by 4 times. This unit only sends the processed input to the PE array in parallel within one cycle, which is much more efficient than the unstructured fine-grained pruning method.
[0085] The plane mask generation unit processes data in the following way: First, the unit loads the replicated input stream. After completing the convolution calculation, it outputs 1-bit binary mask data to the FIFO queue connected to the plane mask data channel. The length of this FIFO queue is equal to the spatial size (H’×W’) of the target output feature map. If the plane mask data channel spans a max-pooling layer, then the plane mask generation unit must upsample the output results and save them to an additional FIFO queue to support the size change of the output feature map. To keep the data stream synchronized with the mask data stream, a buffer with an appropriate depth needs to be introduced in the data stream path to wait for the generation of the mask data.
[0086] Such as Figure 7 , in the sparse BNN accelerator architecture, the max-pooling layer calculation engine consists of a sliding window unit with a window size of 2×2 and an OR gate array.
[0087] The algorithm of the present invention has been tested on the Cifar-10 and ImageNet datasets. On the Cifar-10 dataset, when applying the double pruning method of the present invention to all convolutional layers of the binary models ResNet-18 and Vgg-Small, the parameter savings can reach 69.4% and 67.5% respectively, and the computational savings can reach 85.4% and 80.2%, with only a 2.3% and 1.8% accuracy drop. Among them, N and pr required by the double pruning method are set to 4 and 60% respectively. On the ImageNet dataset, when applying the double pruning method of the present invention to all convolutional layers of the binary models ResNet-18 and ResNet-34, about 44% of the parameter savings and 70% of the computational savings can be achieved, with only a 5% accuracy drop. When implementing the accelerator architecture on the FPGA platform, for the sparse Vgg-Small model applying the double pruning method above, the accelerator can achieve an acceleration ratio of about 5.4 times compared with the unpruned dense binary network model.
[0088] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A structured fine-grained binary network double pruning method, characterized in that, The method specifically includes: generating indices for structured pruning of each group of convolutional weight channels, where the structure is such that only one channel out of every N consecutive weight channels is retained for calculation; generating planar masks representing the importance of the output feature maps of each layer; pruning the convolutional weight channels corresponding to each binary convolutional layer according to the indices of channel pruning; and during inference, masking the calculation of the output features at the positions indicated by the planar masks of the feature maps. Assume that the size of the input feature map of a specific binary convolutional layer is [IC, H, W], the size of the weights is [OC, IC, K, K], and the size of the output feature map is [OC, H', W'], where H and W are the sizes of the input feature map, H' and W' are the sizes of the output feature map, OC is the number of output channels, IC is the number of input channels, and K is the convolutional kernel size; the convolutional layer has a total of OC groups of weights, and each group of weights has IC channels, so the total number of channels of the weights of the convolutional layer is OC × IC; the steps for generating the pruning indices for each group of weight channels are as follows: Step 1: Generate a parameter vector Pc with a length of OC × IC. Step 2: Rearrange Pc into a matrix Mc with a size of [OC × (IC / N), N]. Step 3: Reparameterize Mc using the Gumbel-Softmax method. Step 4: If it is training, rearrange the reparameterized matrix into a tensor with a size of [OC, IC, 1, 1], and expand it by replicating it K*K times along the planar convolutional kernel dimension, and multiply the expanded tensor with a size of [OC, IC, K, K] by the weights; if it is inference, skip Step 4 and continue with Step 5-6. Step 5: Binarize the reparameterized matrix to obtain a binary matrix BMc, which has the same size as Mc, with OC × (IC / N) rows and N columns. The binarization method is as follows: set the element of BMc at the position of the maximum value in each row of Mc to 1, and set the remaining N - 1 positions to 0. Step 6: Rearrange BMc into a tensor BPc with a size of [OC, IC, 1, 1], and this tensor is the pruning index for all OC groups of weight channels of the binary convolutional layer. The method is implemented using a dual-pruned pipeline sparse binary network accelerator device, which includes a sparse binary convolutional layer computing engine, a first-layer computing engine, a max-pooling layer computing engine, and a fully-connected layer computing engine; when the accelerator is started, the image to be processed is read from the external memory to the first-layer computing engine through DMA.
2. The structured fine-grained binary network double pruning method according to claim 1, wherein The reparameterization of Mc using Gumbel-Softmax in Step 3 includes the following steps: Step 11: Use the Softmax function σ to convert the N elements in each row of Mc into sampling probabilities P, which means the probability of being selected among the N convolutional kernels. Step 12: Based on P obtained in Step 11 i and the Gumbel random number G i , use the following formula to transform any row of Mc into a reparameterized row vector S = [S1, S2,..., S N , 3. The structured fine-grained binary network double pruning method according to claim 1, characterized in that, The generation of the planar masks for the output feature maps includes the following steps: Step 21: For the intermediate feature map of a certain layer in the network, introduce a bypass binary convolutional layer, which has only one set of weights; after the convolution calculation of this set of weights and the intermediate feature, a planar feature map Ms is obtained; by adjusting the convolution stride, the planar feature map is aligned with the target output feature map in terms of size H' and W'. Step 22: Perform batch normalization on this planar feature map. Step 23: Use the Gumbel-Softmax method to reparameterize the planar feature map Ms into S. Step 24: If it is training, perform matrix dot multiplication on the reparameterized planar feature map S obtained in Step 23 and the planar features in each channel of the output feature map, so as to update the weight parameters of the bypass convolutional layer during the training process. If it is inference, skip Step 24 and continue with Step 25. Step 25: Binarize the reparameterized planar feature map obtained in Step 23 to generate a planar mask. The specific method is as follows: each element of the planar feature map is compared with a threshold. If it is greater than the threshold, the corresponding element of the planar mask is taken as 1, otherwise it is taken as 0, thus obtaining a binary planar mask.
4. A structured fine-grained binary network double pruning method according to claim 3, characterized in that, Introduce a new loss component during the training of the binary network to obtain the optimal sparsity of the planar mask during training. The specific method is to add a loss related to the sparsity of the planar mask, namely L cls to the classification Loss of the network, that is, L prun , and the final training Loss becomes: Among them, L cls is the classification loss, L prun is the clipping loss, δ is the average sparsity of all plane masks, pr is the preset value of the required sparsity, the smaller L cls is, the smaller the classification loss is, and the smaller L prun is, the closer the average sparsity of the existing plane masks is to the preset value. The hyperparameters λ1 and λ2 are used as weights to balance the two Losses and find the optimal sparsity of the plane masks.
5. A doubly pruned pipelined sparse binary network accelerator device based on the method according to any one of claims 1-4, characterized in that Each individual layer instantiates a customized computing engine, which starts to perform the calculation of the current layer once the previous layer starts to produce output.
6. The double-clipped pipeline sparse binary network accelerator device according to claim 5, wherein The sparse binary convolutional layer computing engine includes a sliding window unit, a binary convolution computing unit, a channel selection unit, a planar mask generation unit, and a planar mask data channel; among them, Sliding window unit: It includes a row buffer and a sliding window, which are used to receive and cache the input data stream, and shield redundancy according to the planar mask, and output the data required for convolution calculation in the correct order. Binary convolution computing unit: It includes an input bit-width converter, a register-based local memory, a PE array, and an output bit-width converter; the PE array is composed of PEs, and each PE is composed of a certain number of XNOR gates, an adder tree for Popcount, a shift register, and a threshold comparator for performing Batchnorm and Sign functions. The binary convolution computing unit composed of the above is used to complete the sparse binary convolution calculation. Channel selection unit: It includes a certain number of N-to-1 MUXs, which are used in combination with the PEs to grab the corresponding input channels according to the N-to-1 channel pruning index and output them to each PE for calculation. Planar mask generation unit: It is used to generate the planar mask required for the target output feature map and coordinate the flow of input and mask data. Planar mask data channel: This channel is connected to all computing engines sharing the same planar mask, and provides mask data for the computing engines to shield redundant calculations; after the sliding window unit and the PE array load the mask data from the planar mask data channel, they need to copy the data and write it back to the channel for subsequent use by the computing engines.
7. The double-clipped pipelined sparse binary network accelerator device according to claim 5, wherein The first-layer computing engine is similar to the sparse binary convolutional layer computing engine, except that: the input is a 3-channel image and both the input pixels and weights are fixed-point numbers; the XNOR gates in the binary convolution computing unit need to be replaced by multipliers; neither the sliding window unit nor the PE array needs to load mask data for judgment.
8. The doubly pruned pipelined sparse binary network accelerator device according to claim 5, characterized in that, The fully connected layer computing engine is composed of binary convolution computing units; among them, the XNOR gate is replaced by a sign converter to perform the convolution computing of the fully connected layer where the input is binary and the weights are fixed-point.
Citation Information
Patent Citations
An acceleration method for realizing sparse convolutional neural network inference for hardware
CN109711532A
Multi-granularity-based deep neural network structured sparse system and method
CN110276450A