A Filter Pruning Method Based on the Distance between Feature Map Channels

By calculating the geometric distance between feature map channels and greedy strategies to remove redundant filters, the problems of complexity and redundant feature survival of existing methods are solved, and efficient model compression and recognition accuracy maintenance are achieved.

CN116757263BActive Publication Date: 2025-07-25JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310520095.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-10
Publication Date
2025-07-25
Estimated Expiration
2043-05-10

AI Technical Summary

Technical Problem

The existing filter pruning method requires modifying the loss function or embedding additional variables in the network, resulting in complexity and failure to effectively utilize inter-channel feature information, resulting in redundant feature survival.

Method used

By calculating the geometric distance between the feature map channels, a greedy strategy is used to find the optimal solution, remove redundant filters, and use greedy strategy and threshold constraints to perform filter trimming, simplifying the process and maintaining network performance.

Benefits of technology

While reducing the computational complexity, it ensures recognition accuracy and robustness. It is suitable for various trained networks and realizes efficient model compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116757263B_ABST
    Figure CN116757263B_ABST
Patent Text Reader

Abstract

The present invention discloses a filter pruning method based on the distance between feature map channels, belonging to the field of lightweight neural network image recognition. The present invention measures the correlation of feature maps based on geometric distance across channels to delete redundant feature maps and their corresponding filters. Considering filter pruning from the perspective of between channels can more stably and reliably explore redundant filters, thereby providing more accurate guidance for filter pruning. The present invention realizes the decoupling of training and pruning as well as the decoupling of network structure and pruning, defines filter pruning as an optimization problem, introduces a greedy strategy in the pruning process to seek an approximate solution to the optimal solution, and designs a highly robust and low-cost redundant filter pruning scheme; the method of the present invention eliminates additional auxiliary constraints and embedded variables, thereby simplifying the pruning process. In addition, the method of the present invention does not require modifying the loss function and does not need to know the training details, and has strong adaptability to any trained network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a filter pruning method and system based on the distance between feature map channels, and belongs to the field of lightweight image recognition neural networks. Background Art

[0002] Deep neural networks are computationally intensive and memory intensive, which makes it extremely challenging to deploy large and complex DNNs on edge devices (embedded devices, mobile platforms, etc.) with scarce computing resources. Therefore, model compression techniques must be used to reduce the number of model parameters and the amount of computation, thereby reducing the complexity of the model and enabling the DNN model to be applied to edge devices. Common model compression methods include parameter quantization, knowledge distillation, and model pruning. Among them, the model pruning technique has shown great potential among various model compression methods and can be roughly divided into weight pruning and filter pruning. Weight pruning is to delete some neurons or corresponding weights that have a low impact on the output from the network, and it is a fine-grained pruning method for individual weights. However, the sparse weight matrix obtained by weight pruning cannot achieve an ideal acceleration effect on general-purpose hardware. In contrast, filter pruning is a pruning algorithm at the filter level. The pruned network still has a good organizational structure and does not require special hardware or libraries to accelerate inference like weight pruning, and can be easily accelerated on general-purpose processors. In recent years, the research on filter pruning has received increasing attention from researchers.

[0003] The main purpose of filter pruning is to remove redundant and unimportant filters to reduce model complexity without sacrificing model performance. Among some recently proposed filter pruning methods, using the "L-norm" sparse loss and learnable mask matrices has shown good pruning effects. However, since these methods require modifying the loss function or embedding pruning-related variables in the network, they cannot directly benefit from pre-trained networks. Additionally, instead of directly using the information of the filters themselves, using feature information to prune filters has gradually become a commonly used strategy. As described by Lin et al., feature maps reflect the relationship between filters and input data and contain richer information than the filters themselves. Based on this concept, several feature-guided filter pruning algorithms have been proposed, and they perform better than filter-guided methods in pruning tasks. However, so far, some feature-guided methods only explore the importance of filters using intra-channel feature information, which may cause potential inter-channel relationships to be unnoticed during the pruning process, and similar features may receive the same scores, resulting in redundant features still surviving. In fact, if inter-channel feature information is correctly used, it can provide richer knowledge for filter pruning than intra-channel information. Specifically, there are two reasons: 1) Cross-channel pruning can better explore the potential correlations between different feature maps (and corresponding filters), thus achieving better pruning performance; 2) If the importance of the corresponding filters is evaluated only by intra-channel feature information, it may be sensitive to input data, while inter-channel feature information is more stable and reliable. All in all, many existing pruning methods are too tightly coupled with the network model itself (modifying the loss function or adding extra variables to the network), which complicates the pruning task and is not easy to use. Additionally, as mentioned above, cross-channel pruning based on feature guidance can bring more benefits than intra-channel pruning. Summary of the Invention

[0004] In order to ensure model accuracy and enhance the robustness of image recognition while reducing computational complexity, the present invention provides a filter pruning method and system based on the distance between channels of feature maps. The technical solution is as follows:

[0005] The first object of the present invention is to provide a neural network filter pruning method for image recognition, including:

[0006] Step 1: Construct a training set and a test set, train an image recognition neural network, and obtain an initial network;

[0007] Step 2: Randomly select several pictures from the test set, input them into the trained image recognition neural network, layer by layer calculate the geometric distance between channels of feature maps within the same layer, put them into the set Scores, and record the indices of the two channels;

[0008] Step 3: Use the greedy strategy to find the optimal solution, and divide the parameters \(W\) of the \(i\)-th layer of the image recognition neural network i into two subsets \(R\) i and \(K\) i , such that the geometric distance between the channels of the feature maps generated by the filters in \(R\) i approximates the filters in \(K\) i , where \(R\) i represents the set of filter index to be removed, and \(K\) i represents the set of filter index to be retained;

[0009] Step 4: Sort the set \(Scores\) in descending order according to the geometric distance, and take the top \(t\) channels;

[0010] Step 5: If the channels obtained in Step 4 are not in the set \(K\) i , and the geometric distance is greater than or equal to the preset threshold \(s\), then remove the channel and its corresponding filter, otherwise do not remove;

[0011] Step 6: After channel removal, initialize the pruned model with the filter weight parameters in the set \(K\) i ;

[0012] Step 7: Retrain the model until the accuracy is equivalent to that of the initial network in Step 1;

[0013] Step 8: If the size of the model retrained in Step 7 does not reach the preset model size, then use the model in Step 7 to go back to Step 2 to continue pruning until the preset model size is reached.

[0014] Optionally, finding the optimal solution in Step 3 is expressed as:

[0015]

[0016]

[0017]

[0018] where represents the geometric distance between the \(p\)-th and \(q\)-th channels of the output feature map of the \(i\)-th layer, represents the \(c\)-th channel of the \(m\)-th feature map in the output feature map of the \(i\)-th layer, and \(\|\cdot\|\) F represents the F-norm, \(N\) represents the number of samples input to the \(i\)-th layer, and \(L\) represents the total number of layers of the image recognition neural network.

[0019] Optionally, calculating the geometric distance between the channels of the feature maps within the same layer in Step 2 includes:

[0020] Step 21: Traverse each channel of the output feature map of each layer;

[0021] Step 22: If the channel has been removed, go to the said Step 21;

[0022] Step 23: Put the channel into the set K i ;

[0023] Step 24: Traverse each channel of the output feature map of each layer;

[0024] Step 25: Calculate the geometric distance between the two channels in the said Step 21 and Step 24;

[0025] Step 26: Put the geometric distance obtained in the said Step 25 into the set Scores, and record the indexes of the two channels, then go to the said Step 21 until the geometric distances between every two channels have been calculated.

[0026] Optionally, the image recognition neural network is a deep convolutional neural network.

[0027] Optionally, the image recognition neural network is a VGG-16 network, including a fully connected layer and thirteen convolutional layers.

[0028] Optionally, the image recognition neural network is a MobileFaceNet network, including: a plurality of bottleneck blocks, and each bottleneck block is composed of three convolutional layers: 1×1 convolution, 3×3 depth convolution, and 1×1 convolution.

[0029] Optionally, the pruning rate of the method is:

[0030]

[0031] where FLOPS pruned and FLOPS original respectively represent the floating-point operation amounts after and before pruning.

[0032] The second object of the present invention is to provide an image recognition method, which uses the image recognition neural network obtained by the filter pruning method described in any one of the above to perform image recognition.

[0033] The third object of the present invention is to provide an image recognition system, including:

[0034] An image acquisition module, which is used to acquire an image to be recognized;

[0035] A calculation and recognition module, which uses the above image recognition method to recognize the acquired image;

[0036] An output and display module, which is used to output the image recognition result.

[0037] The fourth object of the present invention is to provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method according to any one of the above.

[0038] The beneficial effects of the present invention are as follows:

[0039] The feature-oriented cross-channel filter pruning method (FOAD) of the present invention measures the correlation of feature maps across channels based on geometric distance to delete redundant feature maps and their corresponding filters, eliminating additional auxiliary constraints and embedded variables, thus simplifying the pruning process. In addition, FOAD does not require modifying the loss function and does not need to know the training details, and has strong adaptability to any trained neural network. Using the neural network pruned by the method of the present invention for image recognition can ensure the recognition accuracy on the premise of reducing the computational complexity, and the image recognition network has good robustness.

[0040] A large number of experiments have proved that the method of the present invention has good performance for various network structures. For example, on the CIFAR-10 dataset, the number of parameters and floating-point operations of VGG-16 are reduced by 87.1% and 63.7% respectively, and a high accuracy of 93.81% is achieved. The present invention also uses the lightweight network MobileFaceNet and the CASIA-WebFace dataset to evaluate the performance of the method of the present invention. The results show that after using the pruning method of the present invention, MobileFaceNet still reaches a test accuracy of 99.02% on LFW when the number of parameters and floating-point operations are reduced by 58.0% and 63.6% respectively, and the inference accuracy hardly decreases. Description of the Drawings

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 It is the architecture diagram of the filter pruning method of the present invention.

[0043] Figure 2 It is the algorithm robustness analysis diagram of the present invention.

[0044] Figure 3 It is the experimental result diagram of the present invention, where (a) is the pruning result diagram of FOAD and its variants, and (b) is the hyperparameter analysis diagram of FOAD. Detailed Embodiments

[0045] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the accompanying drawings.

[0046] First, introduce some concepts and principles:

[0047] Given a pre-trained CNN model that contains L convolutional layers, where L i represents the i-th layer. The parameters of L i can be expressed as where n i represents the number of filters in the i-th layer, represents the j-th filter in the i-th layer, and k h and k w represent the height and width of the filter, respectively.

[0048] Let the input of L i be where N represents the number of input samples, represents the p-th sample, and h i and w i represent the height and width of the input image, respectively. Let the output feature map of the i-th layer be , where

[0049] After pruning, the parameters W i of L i are divided into two subsets, namely R i and K i , represents the set of filter indices to be removed, and r i j represents the index of the j-th filter to be removed in the i-th layer. represents the set of filter indices to be retained, represents the index of the j-th filter to be retained. Among them, n ir +n ik =n i , R i ∩K i =Φ, R i ∪K i =W i . The present invention uses the number of model parameters (Params) and the number of floating-point operations (FLOPS) to evaluate the size and computational overhead of the model. The pruning rate in the present invention refers to the ratio of the reduced FLOPS to the original FLOPS. The specific mathematical formula is:

[0050]

[0051] where Pr Denotes the pruning rate, FLOPS pruned and FLOPS original respectively represent the floating-point operation amounts after and before pruning.

[0052] The goal of filter pruning is to find a subset of filters such that the impact of this subset on the CNN is negligible. The present invention aims to find subsets R i and K i , R i The geometric distance of the feature maps generated by the filters in i is approximated to the filters in K. In other words, the purpose of the present invention is to retain as many different features as possible and remove redundant features. Therefore, the filter pruning problem can be defined as the following optimization problem:

[0053]

[0054] where, is used to calculate the geometric distance between two feature maps and represents the correlation. Represents the c-th channel of the m-th feature map in the output feature map of the i-th layer. In the present invention is defined as:

[0055]

[0056] where ||·|| F represents the F-norm. In order to map the geometric distance to the interval (0, 1), the closer the geometric distance between channels is to 1, the closer the distance is and the stronger the correlation is. The above formula can be expressed as:

[0057]

[0058] The above formula maps the geometric distance to the interval (0, 1). However, the feature map is sensitive to the input image. Multiple images need to be input simultaneously, and then the average value of the geometric distances between the same channels in the feature maps generated by all input images is calculated. Therefore, the above formula can be further modified as:

[0059]

[0060] represents the geometric distance between the p-th and q-th channels of the output feature map of the i-th layer. In addition, when p = q, the value calculated according to the above formula is 1. To avoid the influence of the channel itself, the present invention constrains to return 0 when p = q. The formula is as follows:

[0061]

[0062] Finally, the filter pruning optimization problem can be expressed as:

[0063]

[0064] However, solving the above equation is an NP-Hard problem. Therefore, the present invention adopts a greedy strategy to solve the optimization problem described by the above equation, that is, FOAD selects one channel each time and removes the t channels (and their corresponding filters) with the strongest correlation with this channel in the current iteration. Obviously, the solution obtained using the greedy strategy is a sub-optimal solution, but this gap can be compensated by fine-tuning. After solving the above optimization problem, the sets R i and K i are obtained, and the filters in R i can be safely removed.

[0065] In addition, the following two problems need to be considered: First, the correlation between some feature map channels is weak, but the correlation of this channel still ranks among the top. Second, the channels that are first saved to K i may also have a relatively strong correlation with other channels. If the channels with weak correlation or those that have already been saved to K i are removed rashly, it may lead to fewer features learned by the network, thus degrading the classification performance of the network.

[0066] To solve the above problems, a threshold s and a constraint condition are introduced, that is, only the channels with a correlation greater than s and that do not belong to K i can be safely removed, and the removed channels will no longer participate in subsequent calculations. Figure 1 shows the general process of pruning one layer, where t takes the value of 1, s is 0, one channel is removed each time, the removed channel no longer participates in subsequent correlation calculations, and the retained channels are not removed.

[0067] Example 1:

[0068] This example provides a neural network filter pruning method for image recognition, including:

[0069] Step 1: Construct a training set and a test set, train an image recognition neural network, and obtain an initial network;

[0070] Step 2: Randomly select several pictures from the test set, input them into the trained image recognition neural network, calculate the geometric distance between the feature map channels in the same layer layer by layer, put them into the set Scores, and record the indexes of the two channels;

[0071] Step 3: Use the greedy strategy to seek the optimal solution, divide the parameters W i of the i-th layer of the image recognition neural network into two subsets R i and K i , so that the geometric distance between the feature map channels generated by the filters in R i approximates that of Ki The filter in which R i represents the set of filter indices to be removed, and K i is the set of filter indices to be retained;

[0072] Step 4: Sort the set Scores in descending order according to the geometric distance, and select the top t channels;

[0073] Step 5: If the channels obtained in Step 4 are not in the set K i and the geometric distance is greater than or equal to the preset threshold s, then remove the channel and its corresponding filter, otherwise do not remove;

[0074] Step 6: After removing the channels, initialize the pruned model with the filter weight parameters in the set K i ;

[0075] Step 7: Retrain the model until the accuracy is equivalent to that of the initial network in Step 1;

[0076] Step 8: If the size of the model retrained in Step 7 does not reach the preset model size, then use the model in Step 7 to go back to Step 2 to continue pruning until the preset model size is reached.

[0077] Embodiment 2:

[0078] This embodiment takes the pruning of VGG16 on CIFAR-10 as an example.

[0079] Step 1: Prepare the CIFAR-10 dataset.

[0080] The CIFAR-10 dataset contains 10 categories and a total of 60,000 images, including 50,000 training images and 10,000 test images.

[0081] Step 2: Build the VGG-16 network structure.

[0082] VGG-16 consists of thirteen convolutional layers and three fully connected layers. Since the method of the present invention mainly targets convolutional layers, the VGG-16 used in the embodiment has only one fully connected layer and thirteen convolutional layers.

[0083] Step 3: Define the loss function and the optimizer.

[0084] In this embodiment, the SGD optimizer and the cross-entropy loss function are used. The cross-entropy loss function is shown in the following formula:

[0085]

[0086] where N represents the number of samples, and y i represents the label of sample i, and pi Represents the predicted probability of sample i.

[0087] Step 4: Use the CIFAR-10 dataset prepared in Step 1, the loss function in Step 3, and the optimizer to train the VGG-16 network built in Step 2. The training parameters are as follows:

[0088] Train for 160 epochs on the CIFAR-10 dataset and select Stochastic Gradient Descent (SGD) as the optimizer. The Batch Size, initial learning rate, momentum, and weight decay coefficient are 64, 0.1, 0.9, and 1e-4 respectively. Additionally, set the learning rate to 0.01 and 0.001 at the 80th and 120th epochs respectively.

[0089] Step 5: After training the VGG-16 network, perform pruning operations on it.

[0090] Step 51: Randomly select 128 images from the test set, normalize them, and feed them into the trained VGG-16 network.

[0091] Step 52: Use formula (6) to calculate the geometric distance between the channels of the feature maps within the same layer layer by layer, put it into the Scores set, and record the indices of the two channels.

[0092] Step 53: Sort the Scores set in descending order according to the geometric distance.

[0093] Step 54: Take out the top 3 channels in the Scores set.

[0094] Step 55: If the channel is not in the set K and the geometric distance is greater than or equal to the threshold s, then remove the channel and its corresponding filter, otherwise do not remove it.

[0095] Step 6: Initialize the pruned model with the filter weight parameters in the set K.

[0096] Step 7: Retrain the pruned model until the accuracy and the error of the original model are within 1%.

[0097] Step 8: If the size of the model retrained in Step 7 does not reach the pre-set model size, then use the model in Step 7 to go back to Step 6 to continue pruning until the pre-set model size is reached.

[0098] The baseline accuracy of VGG-16 on the CIFAR-10 dataset is 93.84%. When 53.2% of the parameters and 37.3% of the floating-point operations are pruned, the accuracy after fine-tuning can reach 94.00%, which is even higher than the accuracy of the baseline model. When 70.0% of the parameters and 45.9% of the floating-point operations are pruned, the accuracy after fine-tuning can reach 93.82%, which is only 0.02% lower than the accuracy of the baseline model. When the pruning rate is above 50%, when the parameters and floating-point operations are reduced by 87.1% and 63.7% respectively, the accuracy after fine-tuning is 93.81%. For higher pruning rates, that is, when the floating-point operations are reduced to more than 70%, the accuracy of the method of the present invention after fine-tuning can reach 93.41%. These data also prove that the method of the present invention can effectively accelerate and compress convolutional neural networks, and reasonably utilizing the correlation between feature map channels is more helpful for identifying redundant features and filters.

[0099] Example 3:

[0100] In this example, MobileFaceNet pruned on the CASIA-WebFace dataset is taken as an example, and the LFW dataset is used as the evaluation dataset.

[0101] Step 1: Prepare the CASIA-WebFace dataset and the LFW dataset.

[0102] The CASIA-WebFace dataset contains 494,414 images of 10,575 individuals, all of which are used for training, and the LFW dataset is used as the dataset for evaluating the model performance after training.

[0103] Step 2: Build the MobileFaceNet network structure.

[0104] MobileFaceNet consists of multiple bottleneck blocks. A bottleneck block consists of three convolutional layers: 1×1 convolution, 3×3 depthwise convolution, and 1×1 convolution. When the input and output sizes of the bottleneck block are the same, there is a shortcut connection operation.

[0105] Step 3: Define the loss function and optimizer.

[0106] Use the Stochastic Gradient Descent (SGD) optimizer. The total number of training epochs is 60, the Batch Size is set to 128, the weight decay coefficient of PReLU is 0, the weight decay coefficient of the last fully connected layer in MoblileFaceNets is set to 4e-4, the weight decay coefficients of other parts are 4e-5, the initial learning rate and momentum are 0.1 and 0.9 respectively, and the learning rate is adjusted to 0.01, 0.001, and 0.0001 at the 25K, 36K, and 48K training iterations respectively.

[0107] Step 4: Use the CASIA-WebFace dataset prepared in Step 1, the loss function and optimizer in Step 3 to train the MobileFaceNet network built in Step 2.

[0108] Step 5: After training the MobileFaceNet network, perform pruning operations on it. Since there are short connection operations in the bottleneck block, the last convolutional layer in each bottleneck block does not perform pruning operations in the present invention. And the second convolution in the bottleneck block is a depth convolution, and the number of input channels and output channels of the depth convolution must be the same. Therefore, the first convolution layer and the second depth convolution layer in the bottleneck block share the same set K during pruning.

[0109] Step 51: Randomly extract 128 images from the test set, and after normalization, feed them into the trained VGG-16 network.

[0110] Step 52: Use formula (6) to calculate the geometric distance between the feature map channels within the same layer layer by layer, put it into the Scores set, and record the indices of the two channels.

[0111] Step 53: Sort the set Scores in descending order according to the geometric distance.

[0112] Step 54: Take out the first 3 channels from the set Scores.

[0113] Step 55: If the channel is not in the set K and the geometric distance is greater than or equal to the threshold s, then remove the channel and its corresponding filter, otherwise do not remove it.

[0114] Step 6: Initialize the pruned model with the filter weight parameters in the set K.

[0115] Step 7: Retrain the pruned model until the test accuracy on the LFW dataset and the error of the original model are within 0.5%.

[0116] Step 8: If the size of the model retrained in Step 7 does not reach the pre-set model size, then use the model in Step 7 to go to Step 6 to continue pruning until the pre-set model size is reached.

[0117] The accuracy of the trained MobileFaceNet model on LFW is 99.17%. When the number of parameters and the number of floating-point operations are decreased by 29% and 34% respectively, the accuracy on LFW only drops by 0.03% (99.17% VS 99.14%). Even when the model is compressed to 1.75MB, with the FLOPS reduced by 63.6% and the number of parameters decreased by 58%, under such extreme pruning, the test accuracy of the fine-tuned model on LFW can still reach 99.02%. Therefore, the method of the present invention also has good effects on complex data sets, and it also shows that the present invention can work well for compact networks with bottleneck blocks.

[0118] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.

[0119] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A neural network filter pruning method for image recognition, characterized in that The method includes: Step 1: Construct a training set and a test set, train an image recognition neural network, and obtain an initial network; Step 2: Randomly select several pictures from the test set, input them into the trained image recognition neural network, calculate the geometric distance between the channels of the feature maps within the same layer layer by layer, put them into the set Scores, and record the indices of the two channels; Step 3: Use the greedy strategy to find the optimal solution, and divide the parameters of the i th layer of the image recognition neural network W i into two subsets R i and K i such that the geometric distance between the feature map channels generated by the filters in R i approximates the filters in K i , where R i represents the set of filter index to be removed, and K i represents the set of filter index to be retained; Step 4: Sort the set Scores in descending order according to the geometric distance, and take the top t channels; Step 5: If the channel obtained in Step 4 is not in the set K i and the geometric distance is greater than or equal to a preset threshold s , then remove the channel and its corresponding filter; otherwise, do not remove it. Step 6: After channel removal, initialize the pruned model using the filter weight parameters in the set K i ; Step 7: Retrain the model until the accuracy is comparable to that of the initial network in Step 1; Step 8: If the size of the model retrained in Step 7 does not reach the pre-set model size, use the model in Step 7 to go back to Step 2 to continue pruning until the pre-set model size is reached.

2. The filter trimming method according to claim 1, wherein The search for the optimal solution in Step 3 is expressed as: Among them, represents the i geometric distance between the p -th and q -th channels of the output feature map of the -th layer, i represents the m -th c channel of the -th feature map in the output feature map of the N -th layer, i represents the F-norm, L represents the total number of layers of the image recognition neural network.

3. The filter trimming method according to claim 1, characterized in that The image recognition neural network is a deep convolutional neural network.

4. The filter trimming method according to claim 1, characterized in that The image recognition neural network is a VGG-16 network, which includes a fully connected layer and thirteen convolutional layers.

5. The filter trimming method according to claim 1, characterized in that The image recognition neural network is a MobileFaceNet network, which includes: a plurality of bottleneck blocks, and each bottleneck block consists of three convolutional layers: a 1×1 convolution, a 3×3 depth convolution, and a 1×1 convolution.

6. The filter trimming method according to claim 1, wherein The pruning rate of the method is: Among them, and respectively represent the floating-point operation amounts after and before pruning.

7. An image recognition method, characterized in that, The image recognition method uses the image recognition neural network obtained by the filter pruning method described in any one of claims 1-6 to perform image recognition.

8. An image recognition system, characterized in that, The image recognition system includes: An image acquisition module for acquiring an image to be recognized; A calculation and recognition module that uses the image recognition method described in claim 7 to recognize the acquired image; An output display module for outputting the image recognition result.

9. A computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed by a processor, the method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Class-based filter pruning method

    CN113850373A

  • Method of pruning convolutional neural network based on feature map variation

    US20200311549A1