A convolutional neural network compression method and edge-side FPGA accelerator
By introducing hyperparameters P and incremental dynamic pruning and retraining methods into the fine-grained pruning method, the problems of huge index amount and complex accelerator design in the existing fine-grained pruning method are solved, and efficient network compression and hardware accelerator deployment are achieved.
Patent Information
- Application Number
- CN202210515197.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-12
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-05-12
AI Technical Summary
The existing fine-grained pruning method has huge indexes during network compression, resulting in wasted storage space and transmission bandwidth, and the accelerator design is complex and difficult to deploy efficiently.
The hyperparameter P is introduced to reduce the number of indexes, so that the index is encoded into the instructions rather than stored separately, and the regularity of the effective weight position is optimized through incremental dynamic pruning and retraining methods to match the hardware accelerator structure.
It realizes that the index amount is significantly reduced without reducing network performance, and the energy efficiency and computing density of hardware accelerators are improved, so that the convolutional neural network is deployed more efficiently on the edge side.
Smart Images

Figure CN114925823B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of lightweight deep neural networks, and in particular to a compression method for convolutional neural networks, and a corresponding edge-side FPGA accelerator. Background Art
[0002] Convolutional neural networks have now been widely used in computer vision tasks such as image classification and target detection. However, in order to achieve better network performance, convolutional neural networks are constantly deepening, which leads to a rapid increase in the amount of calculation and parameters. With the rapid development of the Internet of Things in recent years, more and more edge devices have been installed in every corner of the city, making edge devices intelligent has become a future development trend. In order to realize the deployment of convolutional neural networks on low-power embedded devices, the acceleration design of convolutional neural networks based on FPGA has become a research focus in academia and industry.
[0003] One of the important methods for compressing neural network models is weight pruning. According to the characteristics of the effective weight position after pruning, the existing common pruning methods are listed from coarse to fine granularity as follows: filter pruning, kernel pruning, vector pruning, and fine-grained pruning. The finer the pruning granularity, the sparser the network is, and the smaller the network performance degradation caused by pruning. Existing studies have shown that when the network compression ratio (the ratio of the network size before and after compression) is constant, the finer the pruning granularity, the better the network performance. However, coarse-grained pruning such as filter pruning and kernel pruning are easy to accelerate on general hardware such as GPUs because of the regularity of their effective weight positions. In contrast, although the network after fine-grained pruning can achieve the best performance, it requires a dedicated hardware accelerator for deployment. The core idea of existing fine-grained pruning research is mostly to retain N of the M continuous elements in the weight matrix, but this will generate a large number of indexes because it is necessary to store the position of each effective weight in the M continuous elements; thus occupying additional storage space and transmission bandwidth, and it is necessary to design additional modules in the accelerator to select according to the index. Summary of the invention
[0004] Purpose of the invention: In order to overcome the deficiencies in the prior art, the present invention provides a model compression method including the proposed new fine-grained pruning, and a matching FPGA hardware accelerator. In view of the problems existing in the existing fine-grained pruning described in the background technology, the proposed new fine-grained pruning introduces an additional hyperparameter P, so that the number of indexes is only 1 / P of the existing fine-grained pruning method, which is convenient for encoding the index into the instruction instead of storing it separately; at the same time, the proposed new fine-grained pruning can make the effective weight position relatively regular, so as to better match the hardware accelerator structure to obtain higher performance and energy efficiency. Through the proposed incremental dynamic pruning and retraining method, the patent enables the VGG-16 network to achieve a compression ratio of 16× on the ImageNet dataset and the accuracy rate only decreases by 1.4%; the maximum compression ratio can reach 64×, and the performance degradation is within the allowable range. When the VGG-16 network compressed by the proposed compression method is deployed using the provided edge-side FPGA accelerator, compared with other designs, it can achieve 3× energy efficiency and 2.19× computing density. The convolutional neural network compression method proposed in this patent and the edge-side FPGA accelerator work together to enable neural networks to be deployed more efficiently on the edge.
[0005] In response to the deficiencies in the prior art, this patent provides a convolutional neural network compression method and a matching edge-side FPGA accelerator.
[0006] The compression method of the convolutional neural network provided by this patent is applied to all convolutional layers and fully connected layers, which includes:
[0007] The proposed fine-grained pruning method is used to dynamically prune all the selected convolutional layers and fully connected layers that need to be pruned;
[0008] Adopt incremental dynamic pruning and retraining methods to control the network performance degradation caused by pruning;
[0009] Quantize all non-zero weights in a convolutional neural network to a target accuracy.
[0010] The dynamic pruning described above includes:
[0011] If the k×k convolutional layer of layer L is currently set as the pruned layer, then its weight parameter matrix W L Flatten into k×k two-dimensional matrices, where the two dimensions of each two-dimensional matrix correspond to the input channel and output channel dimensions of the original weight four-dimensional matrix;
[0012] Group each 2D matrix according to the hyperparameter P, with each group having P elements;
[0013] According to the hyperparameters N and M, among the adjacent M groups, only the N groups with the largest L1 norm are retained;
[0014] Construct a mask matrix T L , whose shape is similar to the weight parameter matrix W L Same, T L Only the corresponding L The reserved weight position of is 1, and the rest of the positions are 0;
[0015] When the neural network is forward calculated, T is used L With W L The result of element-wise multiplication Calculated as weight parameter;
[0016] When backpropagating, update W with the gradient L Rather than
[0017] For the layers that are not set to be pruned temporarily, the forward propagation and back propagation processes remain unchanged;
[0018] After the current retraining is completed, the parameters saved by the pruned layer are W L Rather than
[0019] The incremental dynamic pruning and retraining method comprises:
[0020] Set all layers in the original trained convolutional neural network to unpruned layers;
[0021] Starting from the first layer, set p convolutional layers as pruned layers (p is a positive integer less than or equal to 3), perform dynamic pruning (the method is as described in claim 2) and retrain until the specified number of retraining rounds;
[0022] Then, p convolutional layers are incrementally set as pruned layers, and dynamic pruning and retraining are performed until the specified number of retraining rounds are reached, and this process is repeated until all layers are pruned layers.
[0023] According to the hyperparameters N, M and P, the mask matrix T of each layer is obtained l , and combine this matrix with the weight parameter matrix W of each layer l Multiply bit by bit to get the sparse weight parameter matrix after pruning
[0024] This patent provides a convolutional neural network edge-side FPGA accelerator, which is characterized by being used to accelerate the convolutional neural network compressed by the proposed compression method, including: a sparse computing acceleration module, a weight buffer, an input feature map buffer, an output feature map buffer, an activation and pooling module, a weight register group, and a control module;
[0025] Each sparse computing acceleration module includes a systolic array, a partial sum accumulator, and an input channel selector, which are used to parallelly calculate multiple input channels and generate partial sum outputs of multiple channels; a weight register group is used to assist the operation of the sparse computing acceleration module;
[0026] The weight, input feature map, and output feature map buffers are used to temporarily store data read from or about to be written into the memory;
[0027] After the activation and pooling module completes the calculation of some output channel results of the convolutional layer, it takes out the results from the partial sum accumulator of the sparse computing acceleration module for ReLu operation or pooling operation, and then writes them into the output result buffer to be written into the memory for storage;
[0028] The control module includes an instruction reading submodule, an instruction decoding submodule, and control logic connecting other modules; this edge-side accelerator comes with a proposed reduced instruction set and compiler.
[0029] The sparse computing acceleration module in the edge-side FPGA accelerator of the convolutional neural network includes:
[0030] Multiple input channel selectors, each of which is connected to the first PE unit in each row of the systolic array. The data input to the systolic array is shifted one bit to the right in each cycle.
[0031] The number of rows and columns of the systolic array in the sparse computing acceleration module is limited by the resources of the edge-side FPGA, which can be solved through a set of constraint equations. The number of columns solved will affect the selection of the hyperparameter P in the neural network compression method.
[0032] Beneficial Effects
[0033] The neural network compression method proposed in this patent can retain the characteristics of large compression ratio and small performance loss of existing fine-grained pruning, while solving the problems of large index volume and complex accelerator design. The three hyperparameters N, M and P in the proposed compression method can be solved according to the constraints of actual hardware resources, so that the proposed compression method and edge-side FPGA accelerator structure can be applied to different practical scenarios. In addition, the computing module structure of the proposed accelerator matches the compressed network, and combines multi-level buffering, pipelining and other methods to enable this accelerator to achieve high computing density and high energy efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solution of the present invention, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0035] Figure 1 A flowchart of a convolutional neural network compression method provided by an embodiment of the present invention;
[0036] Figure 2 A schematic diagram of the effective weight position distribution after pruning provided in the embodiment of this patent;
[0037] Figure 3 A schematic diagram of training the pruned layer provided for the embodiment of the present patent;
[0038] Figure 4 A schematic diagram of the structure of the edge-side FPGA accelerator provided for the embodiment of the present patent;
[0039] Figure 5 A schematic diagram of the structure of a sparse computing acceleration module in an accelerator provided in an embodiment of the present patent (N=1, M=16, P=64). DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] As described in the background technology of this application, in the prior art, fine-grained pruning with the characteristics of large compression ratio and little performance degradation has problems such as large index volume and complex accelerator design.
[0042] Therefore, in order to solve the above problems, the embodiments of the present invention partially provide a method for compressing a convolutional neural network, wherein the convolutional neural network includes a convolutional layer and a fully connected layer, see Figure 1 , Figure 1 It is a flowchart of the compression method of convolutional neural network.
[0043] Specifically, the compression method comprises the following steps:
[0044] Step S101, training to obtain an unpruned convolutional neural network model.
[0045] Step S102, setting the first p layers of the network as pruned layers.
[0046] Step S103, dynamically pruning the pruned layer and retraining the entire network.
[0047] In this step, firstly, the weight parameter matrix W of the k×k convolutional layer currently set as the pruned layer is LFlattened into k×k two-dimensional matrices, the two dimensions of each two-dimensional matrix correspond to the original four-dimensional weight parameter matrix W L The input channel and output channel dimensions of . Figure 2 As shown in the figure, each two-dimensional matrix is grouped according to the hyperparameter P, and each group has P elements. According to the hyperparameters N and M, only the N groups with the largest L1 norm are retained among the M adjacent groups. The gray squares in the figure represent the retained weights, and the white squares represent the pruned weights. Define a weight parameter matrix W of the pruned layer L The four-dimensional mask matrix T of the same shape L , this matrix corresponds to W L The positions where the weights are pruned are set to 0, and the positions where the weights are retained are set to 1, so that the pruning operation can be performed with W L With T L It can be represented by the bit-by-bit multiplication of . Figure 3 The training diagram of the pruned layer is shown in Figure 2. The weight used in the forward calculation is W L With T L The bit-by-bit multiplication result of The reverse calculation is to L Rather than Dynamic pruning is reflected in the fact that during retraining, the pruned layer T L It is based on the hyperparameters P, N and M in W L The pruned weights are not fixed at the beginning of training. After the current retraining is completed, the pruned layer saves W L Rather than
[0048] Step S104, determine whether all layers in the network are pruned layers. If yes, proceed to the next step; if not, incrementally set p layers as pruned layers and jump to step S103.
[0049] Step S105, pruning and saving the weight parameters of all layers in the network.
[0050] In this step, each pruning layer in the network obtains the mask matrix T according to the hyperparameters P, N and M and the discrimination criterion L , through W L With T L The pruned weight matrix is obtained by bit-by-bit multiplication and saved.
[0051] Step S106, quantize the non-zero effective weights and retrain.
[0052] In this step, considering the compatibility with the pruning method and the simplicity of hardware implementation, a linear symmetric quantization scheme is adopted. At the same time, in order to further reduce the performance degradation caused by quantization, this problem is solved by retraining. The most widely used int8 quantization method is used in quantization, that is, the weight parameters are quantized from 32-bit floating point numbers during training to 8-bit integers.
[0053] like Figure 4 As shown, an embodiment of the present invention partially provides an edge-side CNN accelerator, which is used to accelerate the convolutional neural network compressed by the compression method. The accelerator includes: a sparse computing acceleration module, a weight buffer, an input feature map buffer, an output feature map buffer, an activation and pooling module, a weight register group, and a control module.
[0054] The weight buffer, input feature map buffer, and output feature map buffer are used to temporarily store data read from or to be written into the memory.
[0055] The weight register group is used to assist the operation of the sparse computing acceleration module. Because it takes time to import weights from the buffer to the weight register group, there are two weight registers in the weight register group. When one is participating in the calculation, the other is importing weights. In this way, the import time overhead is covered by the ping-pong method.
[0056] After the activation and pooling module completes the calculation of the output channel results of the convolutional layer or the fully connected layer, it takes the results from the partial sum accumulator of the sparse computing acceleration module for ReLu operation or pooling operation, and then writes them into the output result buffer to be written into the memory for storage.
[0057] The control module includes an instruction reading submodule, an instruction decoding submodule, and control logic for connecting a sparse computing acceleration module, a weight buffer, an input feature map buffer, an output feature map buffer, an activation and pooling module, and a weight register group.
[0058] The structure of the sparse computing acceleration module is as follows Figure 5 As shown in the figure, it includes a systolic array, a partial sum accumulator, and an input channel selector, which are used to parallelly calculate multiple input channels and generate partial sum outputs of multiple channels. Take the case where the hyperparameters are selected as N=1, M=16, and P=64 as an example. The input channel selector selects the channel corresponding to the effective weight from the 16 input channels and sends it to the systolic array. Figure 5In the sparse computing module shown, each circle in the systolic array represents a processing element (PE). After the calculation starts, in each clock cycle, each PE unit multiplies the data received from the left data register with the weight in the weight register, and then accumulates the product result with the partial sum received from the upper partial sum register; after each clock cycle, the data moves to the right to the next data register, and the partial sum moves down to the next partial sum register. The input / output parallelism of the sparse computing acceleration module is less than the number of input / output channels of the convolution layer, and the weight loaded each time is the parameter of a certain position of the k×k convolution kernel (such as the upper left corner of the k×k matrix) and corresponds to the input / output channel of the part. Therefore, the network of each layer will be divided for calculation. After calculating each part of the same layer, its result is first stored in the SRAM of the partial sum accumulator and accumulated with the result of the next part until the calculation of the entire layer is completed.
[0059] The hyperparameters N, M and P in the proposed compression method can be solved according to the constraints of the edge-side FPGA. Take a low-cost, low-power edge-side FPGA development board Ultra96v2 as an example. The on-chip resources of this FPGA include 70560 LUTs, 141120 Flip Flops, 360 DSPs, and 216 BRAMs. In the designed accelerator, the sparse computing acceleration module is the most important part of the accelerator, which mainly occupies DSP resources. Therefore, in order to determine the structure of the systolic array in this module, constraints should be imposed according to DSP resources. Let the number of rows of the systolic array be R and the number of columns be C, where the number of columns C is numerically equal to the hyperparameter P. The following set of constraint equations can be listed:
[0060]
[0061] In the above equations, R×C is the number of DSPs used in the systolic array; R×M / N is the input parallelism of the sparse computing acceleration module. In order to ensure the regularity of the structure during data storage, the output parallelism C is equal to the input parallelism. The input / output parallelism also needs to be a power of 2, because the number of input / output channels of the convolutional layer is a power of 2. This setting facilitates the block completion of the convolution calculation.
[0062] The solutions to the above constraint equations are shown in the following table:
[0063] Solution R C(P) M / N 1 2 128 64 2 4 64 16 3 8 32 4
[0064] Among these three feasible solutions, Solution 1 has a compression ratio of 64× due to pruning. Such a huge compression ratio will lead to a sharp deterioration in network performance, so it is not considered here. Solution 2 has a compression ratio of 16×, plus a 4× compression ratio due to quantization, and a total compression ratio of 64×; experiments show that the accuracy of the VGG-16 network drops by 7.7% under a 64× compression ratio. Solution 3 has a compression ratio of 4×, plus a 4× compression ratio due to quantization, and a total compression ratio of 16×; experiments show that the accuracy of the VGG-16 network drops by only 1.4% under a 16× compression ratio. Therefore, Solution 3 can be used in scenarios with high accuracy requirements and low real-time requirements, and Solution 2 can be used in scenarios with high real-time requirements.
[0065] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.
[0066] The present invention has been described in detail above in conjunction with specific implementations and exemplary examples, but these descriptions cannot be understood as limiting the present invention. Those skilled in the art understand that, without departing from the spirit and scope of the present invention, a variety of equivalent substitutions, modifications or improvements may be made to the technical solution of the present invention and its implementation methods, all of which fall within the scope of the present invention. The scope of protection of the present invention shall be subject to the attached claims.
Claims
1. A convolutional neural network edge-side FPGA accelerator, characterized in that: The compression method applied to the convolutional neural network accelerates the compressed convolutional neural network, and the compression method of the convolutional neural network is applied to the convolutional layer and the fully connected layer of the convolutional neural network, including: Step 1: Use fine-grained pruning method to dynamically prune the selected convolutional layers and fully connected layers that need to be pruned; Step 2: Use incremental dynamic pruning and retraining to control the network performance degradation caused by pruning; Step 3: quantize all non-zero weights in the convolutional neural network to the target accuracy; The step 1 includes: Set the Lth k×k convolutional layer as the pruned layer, and its weight parameter matrix W L Flatten into k×k two-dimensional matrices, each of which has two dimensions corresponding to the weight parameter matrix W L The input channel and output channel dimensions of ; Set the hyperparameter P to group each two-dimensional matrix, with each group having P elements; Set hyperparameters N and M, and keep only the N groups with the largest L1 norm among the M adjacent groups; Construct a mask matrix T L , whose shape is the same as the weight parameter matrix W L Same, T L Corresponding to W L The retained weight position is set to 1, and the pruned weight position is set to 0; When the neural network is forward calculated, T is used L With W L The result of element-wise multiplication Calculated as weight parameter; When backpropagating, update W with the gradient L Rather than For the convolutional layers and fully connected layers that are not set to be pruned temporarily, their forward and reverse calculation processes remain unchanged; After the current retraining is completed, the parameters saved by the pruned layer are W L Rather than The edge-side FPGA accelerator of the convolutional neural network includes: a sparse computing acceleration module, a weight buffer, an input feature map buffer, an output feature map buffer, an activation and pooling module, a weight register group, and a control module; The sparse computing acceleration module includes a systolic array, a partial sum accumulator, and an input channel selector, and is used to parallelly calculate multiple input channels and output partial sum data of multiple channels; The weight buffer, input feature map buffer, and output feature map buffer are used to temporarily store data read from or to be written into the memory; After the output channel result calculation of the convolutional layer or the fully connected layer is completed, the activation and pooling module takes the result from the partial sum accumulator of the sparse computing acceleration module for ReLu operation or pooling operation, and then writes it to the output feature map buffer to be written into the memory for storage; The weight register group is used to assist the operation of the sparse computing acceleration module; The control module includes an instruction reading submodule, an instruction decoding submodule, and control logic connecting a sparse computing acceleration module, a weight buffer, an input feature map buffer, an output feature map buffer, an activation and pooling module, and a weight register group; the sparse computing acceleration module includes a plurality of input channel selectors, each of which is respectively connected to the first PE unit of each row in the systolic array; The number of rows and columns of the systolic array in the sparse computing acceleration module is limited by the resources of the edge-side FPGA accelerator, which is solved by the constraint equations. The number of columns solved will affect the selection of the hyperparameter P in the neural network compression method. The constraint relationship between the number of rows and columns of the systolic array in the sparse computing acceleration module and the DSP resources in the edge-side FPGA accelerator is: Set the number of rows of the systolic array to R and the number of columns to C, where the number of columns C is numerically equal to the hyperparameter P, and list the following set of constraint equations: Where: R×C is the number of DSPs used in the systolic array; R×M / N is the input parallelism of the sparse computing acceleration module.
2. A convolutional neural network edge-side FPGA accelerator according to claim 1, characterized in that: The step 2 includes: Set all layers in the original trained convolutional neural network to unpruned layers; Starting from the first layer, set the p-layer convolutional layer as the pruned layer, implement step 1, dynamically prune and retrain the pruned layer until the specified number of retraining rounds, Then, p convolutional layers are incrementally set as pruned layers, and step 1 is implemented to perform dynamic pruning and retraining until the specified number of retraining rounds. This iteration continues until all layers are pruned, and the mask matrix T of each layer is obtained according to the hyperparameters N, M and P. l , and combine this matrix with the weight parameter matrix W of each layer l Multiply bit by bit to get the sparse weight parameter matrix after pruning 3. A convolutional neural network edge-side FPGA accelerator according to claim 2, characterized in that: The value of p in the p-layer convolutional layer is a positive integer less than or equal to 3.
Citation Information
Patent Citations
Efficient image classification method based on structured pruning
CN110598731A
Pruning-based training method and system for acceleration hardware of a artificial neural network
KR102256288B1