Channel Global Sorting Guided Neural Network Compression Method for Joint Pruning and Quantization
Through the pruning and quantization joint method guided by channel global sorting, the problem of comparing channel importance and pruning and quantization in the prior art is solved, and a smaller classification accuracy drop value and higher model performance under the specified compression ratio are achieved.
Patent Information
- Application Number
- CN202211217914.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-09-30
AI Technical Summary
In the compression process of convolutional neural networks, the importance of the channel is only compared in the same two-dimensional convolutional layer, and pruning and quantization are essentially unrelated to each other, resulting in a larger decrease in the classification accuracy of the network compared to the uncompressed network after compression at a specified compression ratio.
The pruning and quantization joint method guided by channel global sorting is adopted. By calculating the average rank of each channel, global sorting is performed, and pruning and quantization is performed on this basis to ensure the maximum effect of pruning and quantization.
At the specified compression ratio, the classification accuracy drop in the compressed network compared to the uncompressed network is significantly reduced, and the compressed model performance is improved.
Smart Images

Figure CN115661511B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning, and relates to a method for compressing a convolutional neural network, specifically to a neural network compression method combining channel global sorting-guided pruning and quantization, which can be used to deploy an image classification convolutional neural network on edge devices with limited computing and storage resources and complete image classification tasks. Background Art
[0002] Convolutional neural networks have currently achieved great success in the processing of information such as videos, images, and voices. This is due to their increasingly deep and wide model architectures. However, this has also led to problems such as high computational complexity and large memory space occupation during the network inference process, which restricts the deployment of convolutional neural networks on edge devices with limited computing and storage resources. Especially for hardware such as mobile platforms, intelligent embedded devices, and field-programmable gate arrays that require real-time inference to complete information processing, this has further hindered the application of convolutional neural networks in many scenarios, such as forest fire rescue and face recognition.
[0003] Therefore, convolutional neural network compression methods that can both reduce computational complexity and memory occupancy while maintaining the performance of convolutional neural networks have been proposed, specifically including pruning, quantization, low-rank decomposition, knowledge distillation, etc. These methods can be used alone or in combination. Among them, pruning and quantization are more widely used. Pruning is to delete some channels of the convolutional layer in the convolutional neural network; quantization is to convert the floating-point form of the weight parameters or the output values of the activation layer into an integer form represented by bit positions, and it is an essential operation before hardware deployment. The standard for measuring an image classification convolutional neural network compression method is the magnitude of the decrease in the classification accuracy of the compressed network compared to the uncompressed network at a specified compression ratio. The smaller the decrease value, the better the performance of the compressed network.
[0004] In the paper "Joint Pruning & Quantization for Extremely Sparse Neural Networks" (arXiv preprint arXiv:2010.01892) published by Yu Po-Hsiang et al. in 2020, a method for compressing an image classification convolutional neural network model by jointly pruning and quantizing is disclosed. This method first proposes a Taylor score to evaluate the channel importance of all channels in a two-dimensional convolutional layer within the same layer, and then deletes the channels with Taylor scores lower than a preset threshold according to the Taylor scores of each channel in each two-dimensional convolutional layer. After pruning is completed, the quantization bit widths of the weights and activation layers are manually specified, and the compressed image classification convolutional neural network model compression method is obtained through fine-tuning. However, the deficiencies of this method are still that the comparison range of the channel importance of the two-dimensional convolutional layer is only within the same layer, and pruning and quantization are essentially unrelated to each other, and cannot ensure maximizing their respective advantages, resulting in a large decrease in the classification accuracy of the compressed network compared to the uncompressed network under a specified compression ratio. Summary of the Invention
[0005] The purpose of the present invention is to propose a neural network compression method guided by global channel ranking for jointly pruning and quantizing in view of the above-mentioned deficiencies of the prior art, so as to solve the problem in the prior art that the channel importance is only compared within the same two-dimensional convolutional layer and pruning and quantization are essentially unrelated to each other, resulting in a large decrease in the classification accuracy of the compressed network compared to the uncompressed network under a specified compression ratio.
[0006] To achieve the above purpose, the technical solutions adopted by the present invention include the following steps:
[0007] (1) Obtain a training sample set and a test sample set:
[0008] Obtain a data set X including M target categories and each category contains N RGB images, and label the image categories in each RGB image, then randomly select N0 images included in each category in the data set X, and form a training sample set X with the selected total of MN0 RGB images and their labels train , and form a test sample set X with the remaining M(N - N0) RGB images and their labels test , where M≥10, N≥6000, N0≥0.8N;
[0009] (2) Construct an image classification convolutional neural network model O and perform iterative training on it:
[0010] Build an image classification convolutional neural network model O including a two-dimensional convolutional layer, a batch normalization layer, a piecewise linear activation layer, multiple residual unit modules, an adaptive average pooling layer, a fully connected layer, and a softmax activation function layer connected in sequence; the first residual unit module includes a convolutional module and a piecewise linear activation layer connected in sequence, and the input of the convolutional module is jump-connected to the piecewise linear activation layer; the second residual unit module includes a convolutional module and an average pooling layer arranged in parallel, and a piecewise linear activation layer connected to the output ends of the convolutional module and the average pooling layer; the convolutional module includes multiple two-dimensional convolutional layers, multiple batch normalization layers, and a piecewise linear activation layer; where the total number of two-dimensional convolutional layers and piecewise linear activation layers is both L, L≥55, each two-dimensional convolutional layer includes I channels, I≥16;
[0011] (3) Iteratively train the image classification convolutional neural network model:
[0012] (3a) Initialize the iteration number as e, the maximum iteration number as E, E≥600, and the weight parameters of the image classification convolutional neural network model at the e-th iteration are θ e , and let e = 0;
[0013] (3b) Use the training sample set X train as the input of O, extract features for each training sample to obtain MN0 feature maps, and classify the targets in each feature map to obtain the classification results of each training sample
[0014] (3c) Adopt the cross-entropy loss function and calculate the loss value of O through the classification results of each training sample and their corresponding labels Then adopt the stochastic gradient descent method, through the partial derivative value of the weight parameter θ e to update θ for θ e to obtain the image classification convolutional neural network model O of this iteration e ;
[0015] (3d) Judge whether e≥E holds. If so, obtain the trained image classification convolutional neural network model Otherwise, let e = e + 1, O e = O, and execute step (3b);
[0016] (4) Calculate the importance scores of all channels in the trained image classification convolutional neural network model and prune and quantize the image classification convolutional neural network model:
[0017] (4a) Take from the training sample set X trainThe rank generation sample set X composed of MN1 training samples randomly selected from and their labels choose As The input of, and use the Hook function to extract The feature maps of each channel of each two-dimensional convolutional layer when the c-th image in is input Then for Perform singular value decomposition to obtain the rank of each channel when the input is input Then according to Calculate the average rank of each channel And then save it, where, N1≥0.01N0, 1≤l≤L, 1≤i≤I;
[0018] (4b) Calculate the importance score of this channel through the average rank of each channel And delete the ρ channels with the lowest importance scores in the trained image classification convolutional neural network model To obtain a pruned image classification convolutional neural network model with a pruning rate of Ω, where, a 、b l 、b l Respectively represent The scalable variable and offset variable that can be optimized in;
[0019] (4c) Calculate the sparsity S of this two-dimensional convolutional layer through the sparse mask composed of I channels of each two-dimensional convolutional layer l l =||Ψ l ||0, and calculate the weight quantization bit width of each two-dimensional convolutional layer according to S l And the quantization bit width of each piecewise linear activation layer
[0020]
[0021]
[0022] Among them, Indicates that the channel is deleted, Indicates that the channel is not deleted, ||·||0 represents the L1 norm, Indicates the ceiling operation, Is the upper bound of the weight quantization bit width of the l-th two-dimensional convolutional layer, Is the upper bound of the activation quantization bit width required by the l-th piecewise linear activation layer, p represents the penalty factor;
[0023] (4d) According to the weight quantization bit width of each two-dimensional convolutional layer And the quantization bit width of each piecewise linear activation layer For the weight vector W of each two-dimensional convolutional layer in the pruned image classification convolutional neural network model l perform quantization, and at the same time replace the activation function of each piecewise linear activation layer, and obtain the quantized weight vector as W l q and the activation function of the piecewise linear activation layer is the pruned and quantized image classification convolutional neural network model
[0024] (5) Re-prune the pruned and quantized image classification convolutional neural network model and update the weights and the quantization bit widths of the activation layers:
[0025] Through the genetic evolution algorithm, for the optimizable scaling variable a l and offset variable b l perform optimization, and through a l and b l the optimization results of a l * and b l * and the average rank of each channel recalculate the importance score of each channel, and then according to the importance scores of all channels recalculated, prune the pruned and quantized image classification convolutional neural network model re-prune and update the weights and the quantization bit widths of the activation layers, and obtain the updated pruned and quantized image classification convolutional neural network model
[0026] (6) Obtain the compression result of the image classification convolutional neural network:
[0027] Fine-tune the weight parameters of the updated pruned and quantized image classification convolutional neural network model to obtain the compressed image classification convolutional neural network model
[0028] The present invention has the following advantages compared with the existing technologies:
[0029] After calculating the importance scores of all channels of all two-dimensional convolutional layers of the trained image classification convolutional neural network model, the present invention performs a global ranking on it, and under the guidance of the global ranking of channel importance, jointly performs pruning and quantization on the image classification convolutional neural network model, solving the problem in the existing technology that only the channel importance is compared within the same two-dimensional convolutional layer and pruning and quantization are essentially unrelated to each other, resulting in a large decrease in the classification accuracy of the compressed network compared with the uncompressed network at a specified compression ratio. Brief Description of the Drawings
[0030] Figure 1 is the implementation flowchart of the present invention;
[0031] Figure 2 is the structural diagram of the residual unit module of the present invention. Detailed implementation manners
[0032] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0033] Referring to the attached Figure 1 drawings, the present invention includes the following steps.
[0034] Step 1) Obtain a training sample set and a test sample set:
[0035] Obtain a data set X including M target categories and each category contains N RGB images, and label the image categories in each RGB image, then randomly select N0 images included in each category in the data set X, and form a training sample set X with the selected total of MN0 RGB images and their labels train and form a test sample set X with the remaining M(N - N0) RGB images and their labels test In this example, M = 10, N = 6000,
[0036] Step 2) Construct an image classification convolutional neural network model O:
[0037] Construct an image classification convolutional neural network model O including a two-dimensional convolutional layer, a batch normalization layer, a piecewise linear activation layer, multiple residual unit modules, an adaptive average pooling layer, a fully connected layer, and a softmax activation function layer connected in sequence; the first residual unit module includes a convolutional module and a piecewise linear activation layer connected in sequence, and the input of the convolutional module is jump-connected to the piecewise linear activation layer; the second residual unit module includes a convolutional module and an average pooling layer arranged in parallel, and a piecewise linear activation layer connected to the output ends of the convolutional module and the average pooling layer; the convolutional module includes multiple two-dimensional convolutional layers, multiple batch normalization layers, and a piecewise linear activation layer; the total number of the two-dimensional convolutional layers and the piecewise linear activation layers is L, each two-dimensional convolutional layer includes I channels, and in this example, L = 55;
[0038] The specific arrangement of multiple residual unit modules is as follows: the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the second residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the second residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module;
[0039] Among them, the specific structure of the first residual unit module in this example is as Figure 2 (a) shown;
[0040] The specific structure of the first residual unit module is: convolutional module → piecewise linear activation layer, and the input of the convolutional module is skip-connected to the piecewise linear activation layer;
[0041] Among them, the specific structure of the second residual unit module in this example is as Figure 2 (b) shown;
[0042] The specific structure of the second residual unit module is: a convolutional module and an average pooling layer arranged in parallel → piecewise linear activation layer;
[0043] The specific structure of the convolutional module is: two-dimensional convolutional layer → batch normalization layer → piecewise linear activation layer → convolutional layer → batch normalization layer;
[0044] The parameter settings for each layer of O are as follows: The convolutional kernel size of the first convolutional layer is 3×3, the number of channels I = 16, and the stride is 1; in the convolutional modules of the first to ninth first residual unit modules, the convolutional kernel size of each two-dimensional convolutional layer is 3×3, the number of channels I = 16, and the stride is 1; in the convolutional module of the tenth second residual unit module, the convolutional kernel size of the first two-dimensional convolutional layer is 3×3, the number of channels I = 32, the stride is 2, the convolutional kernel size of the second two-dimensional convolutional layer is 3×3, the number of channels I = 32, the stride is 1, and the kernel size of the average pooling layer is 1×1, the stride is 2; in the eleventh to eighteenth first residual unit modules, the convolutional kernel size of each two-dimensional convolutional layer is 3×3, the number of channels I = 32, and the stride is 1; in the convolutional module of the nineteenth second residual unit module, the convolutional kernel size of the first two-dimensional convolutional layer is 3×3, the number of channels I = 64, the stride is 2, the convolutional kernel size of the second two-dimensional convolutional layer is 3×3, the number of channels I = 64, the stride is 1, and the kernel size of the average pooling layer is 1×1, the stride is 2; in the twentieth to twenty-seventh first residual unit modules, the convolutional kernel size of each two-dimensional convolutional layer is 3×3, the number of channels I = 64, and the stride is 1; the number of channels of each batch normalization layer is the same as that of its previous two-dimensional convolutional layer; the number of channels of the fully connected layer is 10; the kernel size of the adaptive average pooling layer is 1×1, the stride is 1; all piecewise linear activation layers are implemented by the ReLU function; the last activation layer is implemented by the softmax function;
[0045] Step 3) Iteratively train the image classification convolutional neural network model:
[0046] Step 3a) Initialize the number of iterations as e and the maximum number of iterations as E. In this example, E = 600. The weight parameters of the image classification convolutional neural network model at the e-th iteration are θ e , and let e = 0;
[0047] Step 3b) Use the training sample set X train as the input of O, extract features for each training sample to obtain MN0 feature maps, and classify the objects in each feature map to obtain the classification results of each training sample
[0048] Step 3c) Adopt the cross-entropy loss function and calculate the loss value of O through the classification results of each training sample and their corresponding labels Then adopt the stochastic gradient descent method, and through the partial derivative value of the weight parameter θ e to update θ for θ e to obtain the image classification convolutional neural network model O of this iteration e, calculate the loss value of O and for θ e perform an update. The calculation and update methods are respectively as follows:
[0049]
[0050]
[0051] where MN0 is the number of input training samples, represents the classification result output by the network when inputting the c-th image P c represents the label of the c-th input image. := is an operation that assigns the value on the right side of the formula to the left side. γ represents the learning rate. In this example, γ = 0.001;
[0052] Step 3d) Determine whether e≥E holds. If so, obtain the trained image classification convolutional neural network model Otherwise, set e = e + 1, O e = O, and execute Step 3b);
[0053] Step 4) Calculate the importance scores of all channels in the trained image classification convolutional neural network model and obtain the pruned and quantized image classification convolutional neural network model:
[0054] Step 4a) Use the MN1 training samples randomly selected from the training sample set X train and their labels to form a rank generation sample set X choose as input. In this example, N1 = 64, and use the hook Hook function to extract the feature maps of each channel of each two-dimensional convolutional layer when inputting the c-th image in Then perform singular value decomposition on to obtain the rank of each channel when inputting Then calculate the average rank of each channel according to Then save it. Among them, 1≤l≤L, 1≤i≤I, the rank of each channel and the average rank of each channel Their calculation formulas are respectively as follows:
[0055]
[0056]
[0057] where σ j , u j , v j , v jrespectively represent the left singular vectors, the first j singular values, and the right singular vectors of where the c-th input image is the rank of each channel;
[0058] In step 4b), the importance score of each channel is calculated through the average rank of each channel and the ρ channels with the lowest importance scores in the trained image classification convolutional neural network model are deleted to obtain a pruned image classification convolutional neural network model with a pruning rate of Ω, where a , b l respectively represent l the scalable variable and the offset variable that can be optimized in , and in the example, Ω = 30%;
[0059] In step 4c), the sparsity S of the two-dimensional convolutional layer is calculated through the sparse mask composed of I channels of each two-dimensional convolutional layer where S l = ||Ψ l ||0, and based on S l the weight quantization bit width of each two-dimensional convolutional layer is calculated and the quantization bit width of each piecewise linear activation layer
[0060]
[0061]
[0062] where, indicates that the channel is deleted, indicates that the channel is not deleted, ||·||0 represents the L1 norm, represents the ceiling operation, is the upper bound of the weight quantization bit width of the l-th two-dimensional convolutional layer, is the upper bound of the activation quantization bit width required for the l-th piecewise linear activation layer, p represents the penalty factor, and in this example,
[0063] In step 4d), according to the weight quantization bit width of each two-dimensional convolutional layer and the quantization bit width of each piecewise linear activation layer the weight vector W of each two-dimensional convolutional layer in the pruned image classification convolutional neural network model l is quantized, and at the same time, the activation function of each piecewise linear activation layer is replaced to obtain the quantized weight vector as W l q and the activation function of the piecewise linear activation layer is The pruned and quantized image classification convolutional neural network model The quantization and replacement formulas are respectively:
[0064]
[0065]
[0066] Among them, W l q represents the quantized weight vector of the l-th two-dimensional convolutional layer, tanh(·) represents the hyperbolic tangent function, |·| represents the absolute value operation, max(·) represents the maximum value operation, round(·) represents the operation of mapping a floating-point number to an integer, and y l represents the feature vector input to the l-th piecewise linear activation layer, represents the function used to quantize the input feature vector, and clamp(·, d1, d2) is used to map the input data to the interval [d1, d2];
[0067] Step 5) Update the pruned and quantized image classification convolutional neural network model:
[0068] Through the genetic evolution algorithm for the scalable variable a l and the offset variable b l are optimized, and through the optimization results of a l and b l and the average rank of each channel l * and b l * and the importance score of each channel is recalculated, and then the pruned and quantized image classification convolutional neural network model is pruned again according to the importance scores of all recalculated channels and updated the weights and the quantization bit widths of the activation layers to obtain the updated pruned and quantized image classification convolutional neural network model The implementation steps for the genetic evolution algorithm to optimize a l and b l are as follows:
[0069] Step 5a) Initialize a l = 1, b l = 0, calculate the standard deviation std l of the average rank of I channels in each two-dimensional convolutional layer in Initialize the population Pool as a queue of size P, containing P randomly generated samplesInitialize the mutation rate as μ, the random step size as λ, the number of gradient iteration update steps as τ, the number of iterations as k, the maximum number of iterations as K, and the highest accuracy as ACC max , P ≥ 64, 0 < μ ≤ 1, 50 ≤ τ ≤ 300, K ≥ 330, and let λ = 1, k = 0, ACC max = 0. In this example, P = 64, μ = 0.1, τ = 200, K = 336;
[0070] Step 5b) Quantize after pruning using the importance scores of all channels in S groups calculated from S samples sampled from the Pool and the average rank of all channels to obtain S pruned and quantized image classification convolutional neural network models. In this example, S = 16;
[0071] Step 5c) Initialize the weight parameters of the S pruned and quantized image classification convolutional neural network models according to the weight parameters in . Input the test sample set X test into the S pruned and quantized image classification convolutional neural network models respectively to obtain the image classification accuracy. Take the sample corresponding to the network model with the highest classification accuracy as the current optimal a l ′, b l ′;
[0072] Step 5d) Randomly select μ × L two-dimensional convolutional layers, and calculate the corresponding a l ′, b l ′ of the selected layers to obtain the updated a l ″, b l ″:
[0073]
[0074]
[0075] Among them,
[0076] Step 5e) According to the updated a l ″, b l ″ and the importance scores of each channel calculated from the average rank of all channels perform pruning and then quantization to obtain a new pruned and quantized image classification convolutional neural network model. Input the training sample set X train into the new pruned and quantized image classification convolutional neural network model, and use the cross-entropy loss function and the stochastic gradient descent method to update the gradient by τ steps to obtain a pruned and quantized image classification convolutional neural network model with good gradient iteration update;
[0077] Step 5f) Input the test sample set Xtest Input the pruned and quantized image classification convolutional neural network model updated by gradient iteration to obtain the image classification accuracy ACC k , if this accuracy is greater than ACC max , then ACC max = ACC k , and add the updated a l ″, b l ″ to Pool, otherwise, execute step 5g);
[0078] Step 5g) Judge whether k≥K holds. If so, obtain the optimal a max corresponding to ACC l * , b l * , otherwise, let k = k + 1 and execute step 5b);
[0079] Step 6) Obtain the compression result of the image classification convolutional neural network:
[0080] Fine-tune the weight parameters of the updated pruned and quantized image classification convolutional neural network model to obtain the compressed image classification convolutional neural network model The fine-tuning process is as follows:
[0081] Step 6a) Initialize the iteration number as t, the maximum iteration number as T, T≥1200, and the weight parameters of the image classification convolutional neural network model after pruning and quantization in the t-th iteration are and let t = 0. For Turn off the weight quantization operation and replace all piecewise linear activation layers with the ReLU function function;
[0082] Step 6b) Use the training sample set X train as the input of, and update to obtain the image classification convolutional neural network model of this iteration
[0083] Step 6c) Judge whether holds. If so, use the function for all piecewise linear activation layers in and execute step 6d), otherwise, let t = t + 1, and execute step 6b);
[0084] Step 6d) Judge whether holds. If so, for Restore the weight quantization operation and execute step 6e), otherwise, set t = t + 1, and execute step 6b);
[0085] In step 6e), it is judged whether t≥T holds. If so, obtain the compressed image classification convolutional neural network model Otherwise, set t = t + 1, and execute step 6b);
[0086] The technical effects of the present invention will be further described below in conjunction with simulation experiments:
[0087] 1. Experimental conditions:
[0088] The hardware platform for the simulation experiment is: NVIDIA 2080Ti GPU, and the software platform is: the operating system is Linux, the Python version is 3.7, and the Pytorch version is 1.7.1.
[0089] The simulation experiment uses the CIFAR-10 image classification dataset. The picture data type is RGB, and the size of each picture is 32 pixels × 32 pixels. This dataset contains 10 different categories of labeled pictures, and each category contains 6000 pictures; referring to the dataset division method provided by the official, 5000 pictures and their labels are randomly selected from each category contained in the dataset as the training sample set, and 1000 pictures and their labels in each of the remaining categories are used as the test sample set.
[0090] 2. Simulation content and result analysis:
[0091] The classification accuracy and compression ratio of the present invention and the existing method of jointly pruning and quantizing guided by intra-layer channel scoring are respectively compared and simulated, and the results are shown in Table 1.
[0092] The existing method for compressing neural network models by jointly pruning and quantizing refers to the method for compressing an image classification convolutional neural network model by jointly pruning and quantizing proposed by Yu Po-Hsiang et al. in their published paper "Joint Pruning&Quantization for Extremely Sparse Neural Networks" (arXiv preprint arXiv:2010.01892).
[0093] In order to evaluate the effect of the present invention, using the following evaluation index formulas, the classification accuracy ACC, the decrease value ↓ΔACC of the classification accuracy, and the compression ratio Comp of the three methods in the simulation experiment of the present invention are respectively calculated:
[0094]
[0095] ↓ΔACC = ACC 未压缩网络 -ACC 压缩后网络
[0096]
[0097] The calculation method of the number of bit operation operations BOPs is as follows:
[0098]
[0099] Among them, Ω l-1 represents the pruning rate of the (l - 1)-th layer two-dimensional convolutional layer, and Ω l represents the pruning rate of the l-th layer two-dimensional convolutional layer, and F w,l ·F h,l are respectively the sizes of the feature maps of all channels of the l-th layer two-dimensional convolutional layer, and δ w,l ·δ h,l are respectively the sizes of the corresponding convolutional kernels of the l-th layer two-dimensional convolutional layer.
[0100] Table 1. ACC 未压缩网络 、ACC 压缩后网络 、↓ΔACC, Comp comparison table
[0101] Method <![CDATA[ACC 未压缩网络 > <![CDATA[ACC 压缩后网络 > ↓ΔACC Comp Prior art 94.03% 93.28% 0.75% 54.01 times The present invention 94.03% 93.93% 0.10% 54.96 times
[0102] It can be seen from Table 1 that compared with the prior art, under a compression ratio slightly higher than the prior art, ↓ΔACC is 0.10%, while ↓ΔACC of the prior art is 0.75%, which proves that the compressed image classification convolutional neural network model of this method has a smaller classification accuracy degradation value compared with the uncompressed network model.
[0103] The above simulation experiments show that: under the guidance of the global ranking of channel importance, this invention jointly prunes and quantizes the image classification convolutional neural network model, solving the problem in the prior art that only compares the channel importance within the same two-dimensional convolutional layer and the pruning and quantization are not related to each other, resulting in a large degradation value of the classification accuracy of the compressed network compared with the uncompressed network at the specified compression ratio.
Claims
1. A convolutional neural network compression method that combines channel global sorting guidance pruning and quantization, characterized in that, It includes the following steps: (1) Obtain a training sample set and a test sample set: Obtain a dataset \(X\) that includes \(M\) target categories and each category contains \(N\) RGB images, annotate the image categories in each RGB image, then randomly select \(N_0\) images included in each category in the dataset \(X\), and form a training sample set \(X\) with the selected total of \(MN_0\) RGB images and their labels train , and form a test sample set \(X\) with the remaining \(M(N - N_0)\) RGB images and their labels test , where \(M\geq10\), \(N\geq6000\), \(N_0\geq0.8N\); (2) Construct an image classification convolutional neural network model O and perform iterative training on it: Construct an image classification convolutional neural network model O including a two-dimensional convolutional layer, a batch normalization layer, a piecewise linear activation layer, multiple residual unit modules, an adaptive average pooling layer, a fully connected layer, and a softmax activation function layer connected in sequence; the first residual unit module includes a convolutional module and a piecewise linear activation layer connected in sequence, and the input of the convolutional module is skip-connected to the piecewise linear activation layer; the second residual unit module includes a convolutional module and an average pooling layer arranged in parallel, and a piecewise linear activation layer connected to the output ends of the convolutional module and the average pooling layer; the convolutional module includes multiple two-dimensional convolutional layers, multiple batch normalization layers, and a piecewise linear activation layer; where the total number of two-dimensional convolutional layers and piecewise linear activation layers is L, L≥55, and each two-dimensional convolutional layer includes I channels, I≥16; (3) Perform iterative training on the image classification convolutional neural network model: (3a) Initialize the number of iterations as e, the maximum number of iterations as E, where E ≥ 600, and the weight parameters of the image classification convolutional neural network model at the e-th iteration are θ e , and set e = 0; (3b) Use the training sample set X train as the input of O, extract features for each training sample to obtain MN0 feature maps, and classify the targets in each feature map to obtain the classification results of each training sample (3c) The cross-entropy loss function is adopted. And the loss value of O is calculated through the classification result of each training sample and its corresponding label. Then, the stochastic gradient descent method is adopted, through the partial derivative value of e the weight parameter θ with respect to θ e is updated to obtain the image classification convolutional neural network model O for this iteration. e ; (3d) Determine whether e≥E holds. If so, obtain the trained image classification convolutional neural network model Otherwise, set e = e + 1, O e = O, and execute step (3b); (4) Calculate the importance scores of all channels in the trained image classification convolutional neural network model and perform pruning and quantization on the image classification convolutional neural network model: (4a) From the training sample set X train The rank-generated sample set X is composed of MN1 training samples and their labels randomly selected from choose As Input and use the hook function to extract Input the cth image The feature map of each channel of each 2D convolutional layer is Again Perform singular value decomposition and get the input The rank of each channel is Then according to Calculate the average rank of each channel Then save, where N1≥0.01N0, 1≤l≤L, 1≤i≤I; (4b) Average rank of each channel Calculate the importance score of this channel And for the trained image classification convolutional neural network model Delete the ρ channels with the lowest importance scores in it, and obtain a pruned image classification convolutional neural network model with a pruning rate of Ω, where a l , b l respectively represent the scalable variable and offset variable that can be optimized in it; (4c) Sparse mask composed of I channels of each two-dimensional convolutional layer Calculate the sparsity S of this two-dimensional convolutional layer l = ||Ψ l ||0, and according to S l Calculate the weight quantization bit width of each two-dimensional convolutional layer And the quantization bit width of each piecewise linear activation layer Among them, indicates that the channel is deleted, indicates that the channel is not deleted, and ||·||0 represents the L1 norm, represents the ceiling operation, is the upper bound of the weight quantization bit width of the l-th two-dimensional convolutional layer, is the upper bound of the activation quantization bit width required for the l-th piecewise linear activation layer, and p represents the penalty factor; (4d) According to the quantization bit width of the weights of each two-dimensional convolutional layer and the quantization bit width of each piecewise linear activation layer Quantize the weight vector W of each two-dimensional convolutional layer in the pruned image classification convolutional neural network model l and replace the activation function of each piecewise linear activation layer at the same time, to obtain the quantized weight vector as The activation function of the piecewise linear activation layer is The pruned and quantized image classification convolutional neural network model (5) Re-prune the pruned and quantized image classification convolutional neural network model and update the weight and the quantization bit width of the activation layer: Optimize the scalable variable a and the offset variable b l that can be optimized in l through a genetic evolution algorithm, and recalculate the importance score of each channel through the optimized results a l , b l and the average rank of each channel l * . Then, prune and update the pruned and quantized image classification convolutional neural network model l * again according to the importance scores of all recalculated channels, and update the weights and the quantization bit width of the activation layer of to obtain an updated pruned and quantized image classification convolutional neural network model . (6) Obtain the compression result of the image classification convolutional neural network: Fine-tune the weight parameters of the updated pruned and quantized image classification convolutional neural network model to obtain the compressed image classification convolutional neural network model 2. A convolutional neural network compression method for channel global sorting-guided pruning and quantization combination according to claim 1, characterized in that In the image classification convolutional neural network model O described in step (2), the numbers of the first residual unit module and the second residual unit module it contains are 25 and 2 respectively; the numbers of the two-dimensional convolutional layer and the batch normalization layer contained in the convolutional module are both 2; where: The specific arrangement of multiple residual unit modules is: the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the second residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the second residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module → the first residual unit module; The specific structure of the first residual unit module is: convolutional module → piecewise linear activation layer, and the input of the convolutional module is skip-connected to the piecewise linear activation layer; The specific structure of the second residual unit module is: convolutional module and average pooling layer arranged in parallel → piecewise linear activation layer; The specific structure of the convolutional module is: two-dimensional convolutional layer → batch normalization layer → piecewise linear activation layer → convolutional layer → batch normalization layer; The parameter settings for each layer of O are as follows: The convolutional kernel size of the first convolutional layer is 3×3, the number of channels is 16, and the stride is 1; in the convolutional modules of the first to ninth first residual unit modules, the convolutional kernel size of each two-dimensional convolutional layer is 3×3, the number of channels is 16, and the stride is 1; in the convolutional module of the tenth second residual unit module, the convolutional kernel size of the first two-dimensional convolutional layer is 3×3, the number of channels is 32, the stride is 2, the convolutional kernel size of the second two-dimensional convolutional layer is 3×3, the number of channels is 32, the stride is 1, the kernel size of the average pooling layer is 1×1, and the stride is 2; in the eleventh to eighteenth first residual unit modules, the convolutional kernel size of each two-dimensional convolutional layer is 3×3, the number of channels is 32, and the stride is 1; in the convolutional module of the nineteenth second residual unit module, the convolutional kernel size of the first two-dimensional convolutional layer is 3×3, the number of channels is 64, the stride is 2, the convolutional kernel size of the second two-dimensional convolutional layer is 3×3, the number of channels is 64, the stride is 1, the kernel size of the average pooling layer is 1×1, and the stride is 2; in the twentieth to twenty-seventh first residual unit modules, the convolutional kernel size of each two-dimensional convolutional layer is 3×3, the number of channels is 64, and the stride is 1; the number of channels of each batch normalization layer is the same as that of its previous two-dimensional convolutional layer; the number of channels of the fully connected layer is 10; the kernel size of the adaptive average pooling layer is 1×1, and the stride is 1; all piecewise linear activation layers are implemented by the ReLU function; the last activation layer is implemented by the softmax function.
3. A convolutional neural network compression method for channel global sorting-guided pruning and quantization combination according to claim 1, characterized in that Calculating the loss value of O as described in step (3c) and updating θ e The calculation and update methods are as follows: where \(MN_0\) is the number of input training samples, represents the classification result output by the network when inputting the \(c\)-th image, \(P\) c represents the label of the \(c\)-th input image, \(:=\) is an operation that assigns the value on the right side of the formula to the left side, \(\gamma\) represents the learning rate, and \(0.0001\leqslant\gamma\leqslant0.1\).
4. A convolutional neural network compression method for channel global sorting-guided pruning and quantization combination according to claim 1, characterized in that, The rank of each channel described in step (4a) and the average rank of each channel whose calculation formulas are respectively: Among them, σ j , u j , v j respectively represent the left singular vector, the first j singular values, and the right singular vector of and is the rank of each channel when the c-th image is input.
5. A convolutional neural network compression method combining channel global sorting guidance pruning and quantization according to claim 1, characterized in that The quantization bit widths of the weights of each two-dimensional convolutional layer as described in step (4d) and the quantization bit widths of each piecewise linear activation layer are used to quantize the weight vectors W of each two-dimensional convolutional layer in the pruned image classification convolutional neural network model l and replace the activation functions of each piecewise linear activation layer. The quantization and replacement formulas are as follows: Among them, W l q represents the quantized weight vector of the l-th two-dimensional convolutional layer, tanh(·) represents the hyperbolic tangent function, |·| represents the absolute value operation, max(·) represents the maximum value operation, round(·) represents the operation of mapping a floating-point number to an integer, and y l represents the feature vector input to the l-th piecewise linear activation layer. represents the function used to quantize the input feature vector, and clamp(·, d1, d2) is used to map the input data to the interval [d1, d2].
6. A convolutional neural network compression method for channel global sorting-guided pruning and quantization combination according to claim 1, characterized in that The genetic algorithm described in step (5) optimizes a l , b l , and the implementation steps are as follows: (5a) Initialize a l = 1, b l = 0, calculate the standard deviation std of the average rank of I channels in each two-dimensional convolutional layer in l , initialize the population Pool as a queue of size P, containing P randomly generated samples Initialize the mutation rate as μ, the random step size as λ, the number of gradient iteration update steps as τ, the number of iterations as k, the maximum number of iterations as K, and the highest accuracy as ACC max , P ≥ 64, 0 < μ ≤ 1, 50 ≤ τ ≤ 300, K ≥ 330, and let λ = 1, k = 0, ACC max = 0; (5b) Quantization is performed after pruning based on the importance scores of all channels in S groups calculated from S samples sampled from the Pool and the average rank of all channels, to obtain S pruned and quantized image classification convolutional neural network models, where 16 ≤ S ≤ 64; (5c) Initialize the weight parameters of S pruned and quantized image classification convolutional neural network models according to the weight parameters in, and input the test sample set X test into the S pruned and quantized image classification convolutional neural network models respectively to obtain the image classification accuracy. Take the sample corresponding to the network model with the highest classification accuracy as the current optimal a l ′, b l ′; (5d) Randomly extract μ × L two-dimensional convolutional layers, and calculate the corresponding a l ′ and b l ′ to obtain the updated a l ″ and b l ″: Among them, (5e)According to the updated a l ″, b l ″ and the importance score of each channel calculated based on the average rank of all channels After pruning, quantization is performed to obtain a new pruned and quantized image classification convolutional neural network model. The training sample set X train is input into the new pruned and quantized image classification convolutional neural network model. Using the cross-entropy loss function and the stochastic gradient descent method, the gradient is iteratively updated for τ steps to obtain a pruned and quantized image classification convolutional neural network model with the gradient iteratively updated well; (5f) Input the test sample set X test into the pruned and quantized image classification convolutional neural network model updated by gradient iteration to obtain the image classification accuracy ACC k . If this accuracy is greater than ACC max , then ACC max = ACC k , and add the updated a l ″ and b l ″ to Pool. Otherwise, execute step (5g); (5g) Determine whether k≥K holds. If so, obtain ACC max The corresponding optimal a l * and b l * Otherwise, set k = k + 1 and execute step (5b).
7. A convolutional neural network compression method for channel global sorting-guided pruning and quantization combination according to claim 1, characterized in that The weight parameters of the updated pruned and quantized image classification convolutional neural network model described in step (6) are fine-tuned, and the implementation steps are as follows: (6a) Initialize the number of iterations as t and the maximum number of iterations as T, where T ≥ 1200. After pruning and quantization in the t-th iteration, the weight parameters of the image classification convolutional neural network model are and let t = 0. For Turn off the weight quantization operation and replace all piecewise linear activation layers with the ReLU function function; (6b) Use the training sample set X train as the input to update and obtain the image classification convolutional neural network model for this iteration (6c) Determine whether it holds. If so, for all piecewise linear activation layers in , use the function and execute step (6d). Otherwise, let t = t + 1, and execute step (6b); (6d) Determine whether it holds. If so, restore the weight quantization operation and execute step (6e). Otherwise, let t = t + 1, and execute step (6b); (6e) Determine whether t≥T holds. If so, obtain the compressed image classification convolutional neural network model Otherwise, let t = t + 1, and execute step (6b).
Citation Information
Patent Citations
Joint neural network model compression method based on channel pruning and quantitative training
CN111652366A
Convolutional neural network compression method based on channel number search
CN111882040A