Convolutional neural network pruning quantization compression method based on adaptive genetic algorithm

By using adaptive genetic algorithms to evaluate the importance of filters and pruning them in the image classification network, combined with hardware-friendly quantization methods, the problem of filters being mis-pruned and slow inference speed is solved, and the effects of high accuracy and fast inference are achieved.

CN120068958APending Publication Date: 2025-05-30XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510071557.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the single indicators for evaluating the importance of the filters, resulting in the filter being missed and the accuracy of the image classification network model decreases; at the same time, the hardware-friendly bit width is limited, resulting in the inference speed of the network model being too slow.

Method used

The filter importance evaluation criteria based on adaptive genetic algorithm are used, combined with the Euclidean distance and the average rank of the feature map, and the importance of the filter is comprehensively evaluated, and the importance proportion is adjusted for pruning through the adaptive genetic algorithm. At the same time, using a hardware-friendly quantization method, the network is quantized into 1-bit, 2-bit, 4-bit, and 8-bit hardware-friendly bit width.

Benefits of technology

It greatly improves the accuracy of the image classification network model and significantly speeds up the inference speed of the model, which can meet the requirements of high inference speed while reducing storage and bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068958A_ABST
    Figure CN120068958A_ABST
Patent Text Reader

Abstract

The invention discloses a convolutional neural network pruning quantitative compression method based on an adaptive genetic algorithm. The method comprises the following implementation steps: calculating the sum of Euclidean distances between each convolution kernel of each filter in a pre-training network and other convolution kernels of a convolution layer where the filter is located; calculating the average rank of the feature map output by each filter; pruning the pre-trained convolutional neural network through an adaptive genetic algorithm, and adjusting the proportion of each filter to the convolutional layer and the network model; a hyperbolic tangent function and round operation are used to quantify the network into hardware-friendly bit widths of 2-bit, 4-bit and 8-bit, and a compressed convolutional neural network is obtained. The size of the network is reduced, it is guaranteed that the compressed network has high accuracy, the compressed network model has a hardware-friendly bit width, and the reasoning speed of the model in hardware equipment is greatly increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image data processing, and further relates to an integrated compression method of convolutional neural network pruning and quantization based on an adaptive genetic algorithm in the technical field of probability graphical models using features from images or videos. The present invention can be used in fields such as security, medical health, and intelligent driving that require image classification networks. Under the condition of less computing and storage resources of the device, compression of a larger image classification network model is achieved to improve the accuracy and inference speed of the image classification network model in the above application scenarios. Background Art

[0002] The application of image classification networks has penetrated deeply into all walks of life and has become an important technical support for promoting the intelligent development of society. It has a wide range of applications in many fields such as intelligent driving and medical image classification. In the field of intelligent driving, the image classification network model can realize functions such as traffic sign recognition, lane line detection, pedestrian and vehicle classification, etc., providing accurate data support for autonomous driving decision-making and significantly improving the driving safety and intelligent level. In the field of medical image analysis, the image classification model is widely used in disease diagnosis. For example, by analyzing X-ray films, MRI images, etc., diseases such as pneumonia and tumors can be quickly detected, improving the diagnosis efficiency and accuracy. With the wide application of image classification network models, the models have become more and more complex, resulting in continuous increase in the size and computational complexity of the models. This not only increases the storage and computational costs of the models, but also limits the application of the models in restricted environments such as embedded devices and mobile applications. The excessive size of the network model has brought significant impacts in multiple fields. In the field of intelligent driving, an overly large model will increase the computational burden, prolong the delay of vehicle perception and decision-making, threaten driving safety and reduce the reliability of the system. In the field of medical image analysis, a large model may increase the diagnosis time and computational cost, delay the disease judgment in scenarios that require rapid diagnosis, and affect the treatment effect. Therefore, compression of the image classification network model is carried out to reduce the number of parameters of the network model and accelerate the inference speed. Pruning and quantization are common means in model compression. Pruning is to remove redundant connections or neurons in the network model, significantly reducing the number of parameters and computational complexity of the model, while ensuring that the performance of the model is close to that before pruning. Quantization is to compress the weights and activation functions of the model from high-precision floating-point numbers to low-precision data types, thereby achieving the purpose of reducing storage requirements and computational complexity.

[0003] Tianjin University proposed a category-based filter pruning method in its patent document "A Category-Based Filter Pruning Method" (Patent Application No.: CN202111113265.9, Patent Publication No.: CN 113850373 B). The implementation steps of this method are as follows: Build a VGG-16 network, add an activation value generation module, and train to obtain an optimal model; Use the variance of the activation values of the filters as an evaluation index for the importance of the filters to prune the network; Finally, adjust the structure of the pruned network and retrain to restore the model accuracy. To a certain extent, this method reduces the decline in the accuracy rate of the image classification network model after pruning. However, the deficiencies of this method still lie in that only using the variance of the activation values as an evaluation index for the importance of the filters cannot comprehensively reflect the importance of the filters, easily leading to important filters being mispruned, thus causing a significant decline in the accuracy rate of the image classification network model.

[0004] Beijing University of Posts and Telecommunications proposed a neural network model compression method combining quantization and pruning in its patent document "Convolutional Neural Network Model Compression Method and Device Combining Quantization and Pruning" (Patent Application No.: CN 202310205929.7, Patent Publication No.: CN 116384470 A). The implementation steps of this method are as follows: Evaluate the importance of the filters based on importance factors such as the geometric median, determine the filters to be pruned and perform pruning operations to reduce model parameters; Divide the remaining filters into central filters for gradient backpropagation and filters to be quantized, and perform cross-layer equalization preprocessing on the pruned network; Set the quantization range using the MSE error method and perform quantization processing on the filters to be quantized in combination with the AdaRound algorithm; Fine-tune the quantized and pruned model through the distillation learning method to ensure the balance between model accuracy and compression effect. This method compresses the image classification network model to a certain extent. However, the deficiencies of this method still lie in that using the AdaRound algorithm to quantize the filters to be quantized cannot fully guarantee that the network is quantized into hardware-friendly bit widths such as 1 bit, 2 bits, 4 bits, etc., resulting in limitations in hardware acceleration of the model and unable to meet the industry's requirements for high inference speed. Summary of the Invention

[0005] The purpose of the present invention is to propose an integrated compression method for convolutional neural network pruning and quantization based on an adaptive genetic algorithm in view of the deficiencies of the above-mentioned existing technologies, so as to solve the problems of mispruning of filters caused by a single index for evaluating the importance of filters in the existing technology, and the slow inference speed of the network model due to limited hardware-friendly bit widths in the existing technology.

[0006] The technical idea for achieving the object of the present invention is as follows: Since the present invention uses an importance evaluation criterion for filters based on an adaptive genetic algorithm, this evaluation criterion brings two significant technical features: On the one hand, the sum of the Euclidean distances between each filter and other filters in the same layer is used to measure its importance in the current layer. The smaller the Euclidean distance between two convolutional kernels, the more similar these two convolutional kernels are. Calculate the sum of the Euclidean distances between each convolutional kernel and other convolutional kernels in its convolutional layer. The smaller this value is, the higher the similarity between this convolutional kernel and other convolutional kernels in this layer, then the function of this convolutional kernel can be replaced by other convolutional kernels, and the importance of this convolutional kernel is smaller. On the other hand, the average rank of the feature maps output by each filter is used to measure the importance of a single filter for the network model. The larger the average rank of the feature maps output by the filter, the more useful information this filter contains, and the greater the importance of this filter for the entire network model. The present invention reasonably adjusts the proportion of the importance of these two parts through an adaptive genetic algorithm to more objectively and comprehensively evaluate the importance of filters in the network model, thereby solving the problem in the prior art that the evaluation index for the importance of filters is too single, resulting in filters being mis-pruned and the accuracy of the pruned image classification network model decreasing too much. Since the present invention uses a hardware-friendly quantization method, through the hyperbolic tangent function and the round operation, the network is quantized into hardware-friendly bit widths such as 1 bit, 2 bits, 4 bits, and 8 bits. Hardware such as GPUs, NPUs, or AI acceleration chips used for deep learning model inference are usually optimized for bit widths that are powers of 2 in design, which greatly improves the efficiency of operations such as data alignment and cache filling. Quantizing the network model into hardware-friendly bit widths such as 1 bit, 2 bits, 4 bits, and 8 bits can better adapt to these hardware devices, making the calculation speed of matrix multiplication and convolution operations required in model inference faster. Thereby solving the defect in the prior art that weights cannot be quantized into hardware-friendly bit widths and hardware acceleration cannot be utilized, resulting in the inference speed of the image classification network model being too slow.

[0007] To achieve the above object, the technical solutions adopted by the present invention include the following steps:

[0008] Step 1, calculate the sum of the Euclidean distances between each convolutional kernel of each filter in the pre-trained convolutional neural network and other convolutional kernels in its convolutional layer, and evaluate the importance of each filter for its convolutional layer;

[0009] Step 2, calculate the average rank of the feature maps output by each filter to measure the importance of each filter for the network model;

[0010] Step 3, through an adaptive genetic algorithm, prune the pre-trained convolutional neural network and adjust the proportion of each filter for the convolutional layer and the network model;

[0011] Step 4: Quantize the pre-trained convolutional neural network. Using the hyperbolic tangent function and the round operation, quantize the pruned network into hardware-friendly bit widths of 2 bits, 4 bits, and 8 bits to obtain the compressed convolutional neural network.

[0012] Further, the sum of the Euclidean distances between each convolution kernel of each filter in the pre-trained convolutional neural network and the other convolution kernels in its convolutional layer is obtained by the following formula:

[0013]

[0014] where, represents the sum of the Euclidean distances between the i-th filter in the l-th convolutional layer and the other N l -1 filters in the l-th convolutional layer, N l represents the total number of filters in the l-th convolutional layer, C l represents the total number of channels of each filter in the l-th convolutional layer, P l , Q l respectively represent the total number of rows and columns of each filter in the l-th convolutional layer, W i represents the i-th filter, W i (c) (p,q), W j (c) (p,q) represent the values of the c-th channel of the p-th row and q-th column of the i-th and j-th filters respectively, and ∑· represents the summation operation.

[0015] Further, the average rank of the feature map output by each filter is obtained by the following formula:

[0016]

[0017] where, represents the average rank of the feature map output by the i-th filter in the l-th convolutional layer, x represents the total number of input pictures, represents the rank of the feature map matrix , I t represents the pixel matrix of the t-th picture, represents the feature map matrix generated after the i-th filter in the l-th layer of the network processes the input I t .

[0018] Further, the steps of the adaptive genetic algorithm are as follows:

[0019] The first step: Calculate the importance score of each filter according to the following formula:

[0020]

[0021] Among them, represents the importance score of the i-th filter in the l-th convolutional layer, m l represents the proportion of the importance of each filter in the l-th convolutional layer to the entire network model in the importance score of this filter, n l represents the proportion of the importance of each filter in the l-th convolutional layer to the l-th convolutional layer in the importance score of this filter;

[0022] In the second step, construct a gene pool, which contains G groups of randomly generated gene sequence pairs {m g , n g}, where {m g , n g} represents the g-th group of gene sequence pairs in the gene pool, m g ~e N(0,1) , n g ~N(0, std l ), G≥100, g≤G;

[0023] In the third step, randomly select K groups of genes (m g , n g ) pairs from the gene pool as the basic ratio, and combine them with the and of the network model to recalculate the importance scores of each filter A total of K different filter importance scores can be obtained, and all the scores are ranked from largest to smallest. Among them, represents the average rank of the feature map output by the i-th filter in the l-th layer, represents the sum of the Euclidean distances between the i-th filter in the l-th convolutional layer and the other filters in the l-th convolutional layer except the i-th filter.

[0024] Furthermore, the steps for pruning the pre-trained convolutional neural network are as follows:

[0025] In the first step, rank all K different filter importance scores from largest to smallest. The smaller the filter importance score, the lower the ranking. Obtain the positions of the q filters with lower rankings in the network model, where q represents the number of filters to be pruned;

[0026] In the second step, set the masks corresponding to the positions of the q filters with lower rankings to 0 in the way of masking with zeros to achieve pruning of the filters. A total of K pruned networks can be obtained; at the same time, the sparsity masks of each convolutional layer can be obtained Among them, l∈{1,2,3...,X}, X represents the total number of convolutional layers of the pre-trained model, Mask lDenote the sparsity mask of the l-th convolutional layer, Denote the mask corresponding to the j-th filter of the l-th convolutional layer,

[0027] Furthermore, the steps of adjusting the proportion of each filter in the convolutional layer and the network model are as follows:

[0028] First step, use the test set used by the pre-trained model as the input of the pruned model, calculate the classification accuracy of each compressed network model, and take the sample corresponding to the model with the highest accuracy as the optimal The model with the highest accuracy is denoted as W t ;

[0029] Second step, randomly select α×X convolutional layers from the model W t where α represents the mutation intensity, 0 < α ≤ 0.5, and X represents the total number of convolutional layers of the pre-trained model; according to the following formula, perform adaptive mutation on the gene pairs corresponding to each selected convolutional layer:

[0030]

[0031] where, and represent the mutated gene pairs of the l-th convolutional layer, and represent the gene pairs of the l-th convolutional layer before mutation, and represent the mutation operations on the gene pairs of the l-th convolutional layer respectively, τ represents the current iteration number, 0 < τ ≤ T, T is the maximum iteration number, T ≤ 300;

[0032] Third step, use the mutated combined with and calculated according to the network model to calculate the importance scores of all filters, prune the pre-trained model again, use the test set as the input of the model, observe the accuracy of the model, if the accuracy is higher than the accuracy before mutation, then use the mutated gene pairs as the optimal proportion, if the accuracy is lower than before mutation, execute τ = τ + 1, and perform iterative mutation in the second step again until the optimal proportion is found.

[0033] Furthermore, the quantization of the pre-trained convolutional neural network means that, based on calculating the quantization bit width of the weights of the filters and the activation quantization bits in each convolutional layer, then perform quantization operations using a hardware-friendly quantization method, and its steps are as follows:

[0034] First step, calculate the number of filters retained in each convolutional layer according to the following formula:

[0035] J l =||Mask l || 0

[0036] where J l represents the number of filters retained in the l-th convolutional layer, Mask l represents the sparsity mask of the l-th convolutional layer, and ||Mask l || 0 represents finding the number of non-zero elements in the sparsity mask Mask l of the l-th convolutional layer;

[0037] Second step, calculate the weight quantization bit width and activation quantization bit width of the filters in each convolutional layer according to the following formula:

[0038]

[0039] where b l represents the quantization bit width of the filters in the l-th convolutional layer, B l represents the available bit width budget of the l-th convolutional layer, η represents the penalty factor, which is a constant, c l represents the quantization bit width of the l-th activation layer, A l represents the available bit width budget of the l-th activation layer, B l ∈{2, 4, 8}, A l ∈{2, 4, 8}.

[0040] Furthermore, the quantization operation using a hardware-friendly quantization method is implemented by the following formula:

[0041]

[0042] where represents the value after quantization of the i-th filter in the l-th convolutional layer, b l represents the quantization bit width of the filters in the l-th convolutional layer, W l i represents the value before quantization of the i-th filter in the l-th convolutional layer, round(·) represents rounding the value to convert a floating-point number to an integer, and tanh(·) represents the hyperbolic tangent function, represents the value after quantization of the i-th activation value in the l-th activation layer, c l represents the quantization bit width of the l-th activation layer, represents restricting the value range of the number to be within 0 to 1, Represents the value before quantization of the i-th activation value input to the l-th linear activation layer.

[0043] Furthermore, the hyperbolic tangent function is as follows:

[0044]

[0045] Where W l Represents the filter of the l-th layer.

[0046] Furthermore, the round operation is as follows:

[0047]

[0048] Where γ represents the number to be rounded by round, Represents the floor operation.

[0049] The present invention has the following advantages compared with the existing technologies:

[0050] First: The present invention proposes an importance evaluation criterion for filters based on an adaptive genetic algorithm, accurately evaluating the importance of a single filter from two aspects: the importance of the filter to the convolutional layer it belongs to and the importance to the entire network model, and using the adaptive genetic algorithm to reasonably adjust the proportion of these two parts of importance during the training process, overcoming the defect in the existing technologies that only a single evaluation index for filters is considered, it is difficult to comprehensively measure the importance of a single filter, and it is easy to cause mispruning, so that the present invention greatly improves the accuracy of the image classification network model.

[0051] Second, the present invention uses a hardware-friendly quantization method. Through the hyperbolic tangent function and the round operation, the weights are quantized into 2-bit, 4-bit, and 8-bit, which can efficiently cooperate with the design of existing digital circuits and highly match the registers, instruction sets, and storage architectures of the hardware, enabling the present invention to accelerate the operation and inference speed of the network model while reducing storage and bandwidth. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 Is the flowchart of the present invention;

[0053] Figure 2 Is the structural schematic diagram of the first residual block of the present invention;

[0054] Figure 3 Is the structural schematic diagram of the second residual block of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0055] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0056] Refer toFigure 1 , a further detailed description of the implementation steps of the embodiments of the present invention will be given.

[0057] Step 1, generating a training sample set and a test sample set:

[0058] In the embodiments of the present invention, the RGB image samples used are sourced from the CIFAR10 dataset. This dataset is a classic image classification benchmark dataset, containing 60,000 RGB color pictures, each picture with a size of 32×32 pixels, and a total of 10 categories, with each category containing 6,000 pictures with category labels. Then, randomly select 5,000 images from each category, and use the obtained 50,000 RGB images as the training sample set X train , and use the remaining 10,000 RGB images as the test sample set X test .

[0059] Step 2, constructing an image classification convolutional neural network model:

[0060] Construct an image classification convolutional neural network model W that includes a two-dimensional convolutional layer, a batch normalization layer, a ReLU activation layer, 27 residual blocks, an average pooling layer, a fully connected layer, and a softmax activation function layer connected in sequence; among them, the residual blocks are composed of two structures: the first structure of the residual block includes two 3×3 convolutional layers connected in sequence, and each convolutional layer is followed by a batch normalization layer and a ReLU activation layer. Moreover, the input of this residual block will be further added to the output of the module through skip connection, thus maintaining the continuity of features. The number of input channels and output channels of the first structure of the residual block is the same. The second residual block is similar in structure to the first one, including a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer connected in sequence. Each convolutional layer is followed by a batch normalization layer and a ReLU activation layer. The difference is that the stride of the first convolutional layer of this module is 2, which is used to perform downsampling while performing convolution. This module performs skip connection through a convolutional layer with a stride of 2 to match the size change. The number of output channels doubles relative to the number of input channels, further enhancing the expressive ability of the feature map. The convolutional module includes multiple two-dimensional convolutional layers, batch normalization layers, and ReLU activation layers; the network mainly includes 3 stages, each stage includes 9 residual blocks, and the number of channels of each convolutional layer starts from 16 and doubles in each stage, finally reaching 64 channels, enabling the network to gradually extract higher-level features. Among them, the total number of convolutional layers is X.

[0061] The specific arrangement of multiple residual blocks is as follows: The first to the ninth residual blocks belong to the residual blocks of the first structure, the tenth residual block belongs to the residual blocks of the second structure, the eleventh to the eighteenth residual blocks belong to the residual blocks of the first structure, the nineteenth residual block belongs to the residual blocks of the second structure, and the twentieth to the twenty-seventh residual blocks belong to the residual blocks of the first structure.

[0062] In the embodiment of the present invention, the first type of residual block is as Figure 2 shown, and the second type of residual block is as Figure 3 shown.

[0063] Step 3: Pre-train the constructed image classification convolutional neural network.

[0064] Use the training sample set to iteratively train the image classification convolutional neural network model W. After the training is completed, save the model weight parameter W t as the basic model weight for the subsequent compression step.

[0065] Step 3.1: Set the initial number of iterations as n 0 , the maximum number of iterations as N 0 , in this embodiment, N 0 = 400. The weight parameter of the image classification convolutional neural network model at the nth iteration is W n .

[0066] Step 3.2: Randomly select 500 pictures from the training sample set X train as the input of the network model W to calculate the rank of the feature map output by each filter. For each feature map, obtain the rank through singular value decomposition, and average the rank values of different images to obtain the average rank of the feature map output by each filter and save it for later use. The calculation formula is:

[0067]

[0068] Among them, represents the rank of the feature map corresponding to the tth input image , and s is the number of input images. With the help of , the importance of a single filter to the entire network model can be accurately measured. The larger the rank of the feature map output by the filter, the more independent information in more dimensions it can capture, and the more important this filter is to the entire network model. On the contrary, if the rank of the feature map output by the filter is smaller, it means that it is difficult to capture more independent information in multiple dimensions, and the importance of this filter to the entire network model is smaller.

[0069] Step 3.3: Use the training sample set X trainAs the input of the network model W, the cross-entropy loss function L is used to compare the predicted category and the actual label of each training sample, and calculate the loss value of the model W. Then, the gradient descent algorithm is used. By calculating the partial derivatives of the weight parameters and updating them, the model W after this iteration is gradually optimized n , so as to improve the performance of the network model W in the image classification task.

[0070] The calculation formula of the cross-entropy loss function is as follows:

[0071]

[0072] Among them, 50000 is the number of pictures included in the training set, represents the classification result output by the network model when the input is the i-th picture, p i represents the label of the i-th picture.

[0073] Step 3.4, judge whether n≥N 0 , if it holds, the training is completed, and the basic network model W is obtained t , otherwise, let n=n + 1 and then execute step 3.3.

[0074] Step 4, calculate the importance scores of each filter in the trained image classification network model W t and prune and quantize the filters of the network model based on this.

[0075] Step 4.1, calculate the sum of the Euclidean distances between each filter in each convolutional layer of the network model W t and other filters in its layer With the help of the sum of Euclidean distances, the importance of a single filter for its convolutional layer can be accurately measured. If the sum of Euclidean distances is large, it means that the difference between this filter and other filters in its layer is large, and the importance of this filter for its convolutional layer is greater. On the contrary, if the sum of Euclidean distances is small, it means that the difference between this filter and other filters in its layer is small, and the importance of this filter for its convolutional layer is smaller.

[0076] Step 4.2, combine the average rank of the feature maps output by each filter calculated in step 3.2 to obtain the importance scores of each filter in each network model and for the network model O that has been trained *Delete the q filters with the lowest importance scores in the network model to obtain a pruned network model with a pruning rate of θ. The importance score combines the importance of the filter to the entire network model and the importance to its convolutional layer where it is located, and can represent the importance of the filter more comprehensively and accurately. During the pruning process, filters with relatively low importance can be accurately selected for pruning, so that the accuracy of the pruned network model decreases by a relatively small margin. Among them, m i represents the proportion of the importance of the filter to its convolutional layer in the importance score of the filter, and n i represents the proportion of the importance of the filter to the entire network model in the importance score of the filter. In the embodiment of the present invention, θ = 50%.

[0077] Step 4.3, the sparse masks of each convolutional layer in W t are Calculate the sparsity P l = ||Mask l || 0 , and calculate the weight quantization bit width b l of each layer and the quantization bit width q l of each ReLU activation layer l as follows:

[0078]

[0079] Among them, is equal to 0 or 1, 0 means the corresponding filter is deleted, and 1 means the corresponding filter is retained. ||·|| represents the L1 norm of the corresponding formula, and B l represents the maximum selectable quantization bit width of the l-th convolutional layer, and η represents the penalty factor. In the embodiment of the present invention,

[0080]

[0081] Step 4.4, according to the quantization bit width b l of each two-dimensional convolutional layer and the activation quantization bit width q l of each ReLU activation layer, perform layer-by-layer quantization on the convolutional layers of the pruned network model. The model obtained after pruning and quantization is W tc . Through this step, the network is quantized into hardware-friendly bit widths of 2 bits, 4 bits, and 8 bits. The quantized network model can be better adapted to hardware devices, further accelerating the matrix multiplication and convolution operation speeds during the model inference process and greatly improving the model inference speed.

[0082] Step 5, use the adaptive genetic evolution algorithm to optimize the parameters m and n i in the importance scorei Make further optimizations, and re-prune and quantize the network model.

[0083] Step 5.1, initialize m 1 = 1 and n 1 = 0, calculate the standard deviation std t of the average rank of the feature maps output by each filter in W l , construct a population pool pool, represented as a queue of size P, containing P groups of randomly generated gene sequences m l , n l , where m l ~ e N(0,1) , n l ~ N(0, std l ), the mutation rate is α, the step size is λ, the current iteration number is τ, the total iteration number is Τ, and the maximum accuracy is max_A. In an embodiment of the present invention, λ = 1, α = 0.1, T = 300.

[0084] Step 5.2, randomly draw K groups of gene pairs from pool, and combine and calculated according to the network model to calculate the importance scores of all filters, generate a new compression strategy, prune and quantize W t to obtain K compressed network models, use the test set as the input of the compressed models, calculate the classification accuracy of each network model, and use the sample corresponding to the model with the highest accuracy as the optimal In an embodiment of the invention, K = 10.

[0085] Step 5.3, randomly draw α × X convolutional layers from the model W t , and perform further adaptive mutation on the corresponding of each convolutional layer. The mutated value is denoted as

[0086]

[0087] where the amplitude of the mutation depends on the current generation τ and the maximum generation T. As the number of iterations increases, the adaptive adjustment of the mutation can be realized. Through adaptive mutation, the optimal importance ratio can be screened out, the most accurate filter importance score can be obtained, and pruning is performed according to this filter importance score, so that the accuracy of the pruned network model drops the least.

[0088] Step 5.4, use the latest and combine and calculated according to the network modelCalculate the importance scores of all filters, generate a new compression strategy, and perform pruning and quantization on W t Perform pruning and quantization.

[0089] Use the training set as the input of the new network model, and further train the model to obtain

[0090] Step 5.5, input the test set into the network model to obtain the classification accuracy ACC τ , if ACC τ ≥max_A, then max_A = ACC τ , and add to the pool. Otherwise, execute Step 5.6.

[0091] Step 5.6, determine whether the current iteration number τ≥T. If it holds, obtain the optimal corresponding to max_A. Otherwise, set τ = τ + 1 and then execute Step 5.2.

[0092] Step 6, detect the performance of the compressed network model.

[0093] Use the training sample set as the input, and perform the iterative training in Step 3.3 on the updated network model to further improve the accuracy of the network model so as to obtain the optimal compressed network model W o .

Claims

1. A convolutional neural network pruning quantization compression method based on an adaptive genetic algorithm, characterized in that: The importance of each filter is evaluated from two aspects: the importance of the filter to the convolutional layer where it is located and the importance of the filter to the entire network model. The pre-trained convolutional neural network is pruned and quantized by an adaptive genetic algorithm, and the network is quantized into 2-bit, 4-bit, and 8-bit hardware-friendly bit widths. The steps of the compression method include the following: Step 1, calculate the sum of the Euclidean distances between each convolution kernel of each filter in the pre-trained convolutional neural network and other convolution kernels in its convolutional layer, and evaluate the importance of each filter to its convolutional layer; Step 2: Calculate the average rank of the feature map output by each filter to measure the importance of each filter to the network model; Step 3: Prune the pre-trained convolutional neural network through an adaptive genetic algorithm to adjust the proportion of each filter to the convolutional layer and the network model; Step 4: quantize the pre-trained convolutional neural network, quantize the pruned network into hardware-friendly bit widths of 2, 4, and 8 bits, and use the hyperbolic tangent function and round operation to obtain the compressed convolutional neural network.

2. The convolutional neural network pruning quantization compression method based on adaptive genetic algorithm according to claim 1, characterized in that: The sum of the Euclidean distances between each convolution kernel of each filter in the pre-trained convolutional neural network and other convolution kernels in its convolutional layer is calculated as follows: in, represents the relationship between the i-th filter in the l-th convolutional layer and the other N l -1 The sum of the Euclidean distances between filters, N l represents the total number of filters in the lth convolutional layer, C l represents the total number of filter channels in the lth convolutional layer, P l , Q l Respectively represent the total number of rows and columns of each filter in the lth convolutional layer, W i represents the i-th filter, W i (c) (p,q),W j (c) (p,q) represents the value of the p-th row, q-th column, and c-th channel of the i-th and j-th filters respectively, and ∑· represents the summation operation.

3. The convolutional neural network pruning quantization compression method based on adaptive genetic algorithm according to claim 2, characterized in that: The average rank of the feature map output by each filter calculated in step 2 is obtained by the following formula: in, represents the average rank of the feature map output by the ith filter in the lth convolutional layer, x represents the total number of input images, Represents the feature map matrix rank, I t represents the pixel matrix of the tth image, Indicates that the i-th filter in the l-th layer of the network processes the input I t The resulting feature map matrix.

4. The convolutional neural network pruning quantization compression method based on adaptive genetic algorithm according to claim 3 is characterized in that: The steps of the adaptive genetic algorithm described in step 3 are as follows: The first step is to calculate the importance score of each filter according to the following formula: in, represents the importance score of the i-th filter in the l-th convolutional layer, m l Indicates the importance of each filter in the lth convolutional layer to the entire network model as a percentage of the filter importance score, n l Indicates the proportion of the importance of each filter in the l-th convolutional layer to the l-th convolutional layer in the importance score of the filter; The second step is to construct a gene pool, which contains G groups of randomly generated gene sequence pairs {m g ,n g }, where {m g ,n g } represents the g-th group of gene sequence pairs in the gene pool, m g ~e N(0,1) , n g ~N(0,std l ), G ≥ 100, g ≤ G; The third step is to randomly select K groups of genes (m g ,n g ) as the basic ratio, combined with the network model and Recalculate the importance scores of each filter A total of K different filter importance scores can be obtained, and all scores are ranked in descending order, where: represents the average rank of the feature map output by the i-th filter in the l-th layer, It represents the sum of the Euclidean distances between the i-th filter in the l-th convolutional layer and all filters in the l-th convolutional layer except the i-th filter.

5. The convolutional neural network pruning quantization compression method based on adaptive genetic algorithm according to claim 4 is characterized in that: The steps for pruning the pre-trained convolutional neural network described in step 3 are as follows: The first step is to rank the importance scores of all K different filters in descending order. The smaller the importance score of the filter, the lower the ranking. The positions of the q filters with the lower ranking in the network model are obtained, where q represents the number of filters to be cut off. In the second step, the masks corresponding to the positions of the last q filters are set to 0 by setting the masks to zero, so as to prune the filters. A total of K pruned networks can be obtained. At the same time, the sparsity masks of each convolutional layer can be obtained. Among them, l∈{1,2,3...,X}, X represents the total number of convolutional layers of the pre-trained model, Mask l represents the sparsity mask of the lth convolutional layer, represents the mask corresponding to the jth filter of the lth convolutional layer, 6. The convolutional neural network pruning quantization compression method based on adaptive genetic algorithm according to claim 5, characterized in that: The steps to adjust the proportion of each filter to the convolutional layer and network model described in step 3 are as follows: In the first step, the test set used by the pre-trained model is used as the input of the pruned model, the classification accuracy of each compressed network model is calculated, and the sample corresponding to the model with the highest accuracy is taken as the optimal The model with the highest accuracy is denoted as W t ; The second step is to start from the model W t α×X convolutional layers are randomly selected from the training set, where α represents the mutation intensity, 0<α≤0.5, and X represents the total number of convolutional layers in the pre-trained model. According to the following formula, the gene pairs corresponding to each extracted convolutional layer are adaptively mutated: in, and represents the gene pair after the mutation of the lth convolutional layer, and represents the gene pair before mutation in the lth convolutional layer, and They represent the mutation operation on the gene pair of the lth convolutional layer, τ represents the current number of iterations, 0<τ≤T, T is the maximum number of iterations, T≤300; The third step is to use the mutation Combined with the network model and Calculate the importance scores of all filters, prune the pre-trained model again, use the test set as the input of the model, observe the accuracy of the model, and if the accuracy is higher than the accuracy before the mutation, then the mutated gene pair As the optimal ratio, if the accuracy is lower than before the mutation, execute τ=τ+1 and perform the iterative mutation in the second step again until the optimal ratio is found.

7. The convolutional neural network pruning quantization compression method based on adaptive genetic algorithm according to claim 6, characterized in that: The quantization of the pre-trained convolutional neural network described in step 4 means that on the basis of calculating the weight quantization bit width and activation quantization bit of the filter in each convolution layer, a hardware-friendly quantization method is used for quantization operation. The steps are as follows: The first step is to calculate the number of filters retained in each convolutional layer according to the following formula: J l =||Mask l ||0 Among them, J l Indicates the number of filters retained in the lth convolutional layer, Mask l represents the sparsity mask of the lth convolutional layer, ||Mask l ||0 means to find the sparse mask Mask of the lth convolutional layer l The number of non-zero elements in; The second step is to calculate the weight quantization bit width and activation quantization bit width of the filter in each convolutional layer according to the following formula: Among them, b l represents the quantization bit width of the filter in the lth convolutional layer, B l represents the available bit width budget of the lth convolutional layer, η represents the penalty factor, which is a constant, and c l represents the quantization bit width of the lth activation layer, A l represents the available bit width budget of the lth activation layer, B l ∈{2, 4, 8}, A l ∈{2, 4, 8}.

8. The convolutional neural network pruning quantization compression method based on adaptive genetic algorithm according to claim 7, characterized in that: The quantization operation using a hardware-friendly quantization method is implemented by the following formula: in, represents the quantized value of the i-th filter in the l-th convolutional layer, b l represents the quantized bit width of the filter in the lth convolutional layer, represents the value of the i-th filter of the l-th convolutional layer before quantization, round(·) means rounding the value to convert the floating point number to an integer, tanh(·) represents the hyperbolic tangent function, represents the quantized value of the i-th activation value in the l-th activation layer, c l represents the quantization bit width of the lth activation layer, Indicates that the value is limited to the range of 0 to 1. Represents the value of the i-th activation value input to the l-th linear activation layer before quantization.

9. The convolutional neural network pruning quantization compression method based on adaptive genetic algorithm according to claim 1, characterized in that: The hyperbolic tangent function described in step 4 is as follows: Among them, W l represents the filter of the lth layer.

10. The convolutional neural network pruning quantization compression method based on adaptive genetic algorithm according to claim 1, characterized in that: The round operation described in step 4 is as follows: Among them, γ represents the number to be rounded by round, Indicates a floor operation.

Citation Information

Patent Citations

  • A Class-Based Filter Pruning Method

    CN113850373B

  • Convolutional neural network model compression method and device combining quantization and pruning

    CN116384470A