Convolutional Neural Network Model Pruning Method and Device, Electronic Device, Storage Medium
By considering the relationship between filters in the convolutional neural network model, the pruning importance index of each convolutional layer is calculated and pruned, the problem of low pruning accuracy and compression accuracy in the existing technology is solved, and more efficient model compression and computing speed is achieved.
Patent Information
- Application Number
- CN202210163245.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-02-22
AI Technical Summary
The prior art fails to effectively consider the relationship between filters in the pruning of convolutional neural network models, resulting in low pruning accuracy and compression accuracy of the model.
By obtaining the convolution layer information, performing convolution calculations to obtain filter similarity values, computing the pruning importance index of each convolution layer, and pruning the model according to the preset pruning rate.
The accuracy of pruning of convolutional neural network models is improved, and the model compression accuracy and computing speed are improved.
Smart Images

Figure CN114492799B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a convolutional neural network model pruning method and device, an electronic device and a storage medium. Background Art
[0002] With the development of Internet technology and artificial intelligence, models based on convolutional neural networks have shown good performance in many tasks. For example, convolutional neural network models for target detection are widely used. However, these models require huge computational overhead and memory usage when used. Since these models usually contain a lot of redundant information, compressing the models to reduce computational overhead and memory usage during use has become an essential step. Model pruning is an important direction of model compression technology. At present, the detection model and segmentation model in the deep learning model can remove redundant parameters through pruning, ensure the model accuracy as much as possible, compress the model size, and improve the model operation speed.
[0003] However, the current model pruning method for selecting pruning filters only considers the information of a single filter, does not consider the relationship between filters, and does not obtain the redundant information of the internal filters of each convolutional layer in the model based on the relationship between filters, and then uses this redundant information for pruning, resulting in low pruning accuracy and model compression accuracy of the convolutional neural network model. Summary of the invention
[0004] The main purpose of the embodiments of the present invention is to propose a convolutional neural network model pruning method and device, electronic device and storage medium, which can improve the accuracy of convolutional neural network model pruning, and enhance the model compression accuracy and computing speed.
[0005] To achieve the above object, a first aspect of an embodiment of the present invention proposes a convolutional neural network model pruning method, comprising:
[0006] Get the convolutional layer information of the model to be pruned;
[0007] Performing convolution calculation according to the convolution layer information to obtain a filter similarity value corresponding to the filter in each convolution layer;
[0008] Calculating the pruning importance index corresponding to each convolutional layer according to the filter similarity value;
[0009] According to the preset pruning rate and the pruning importance index corresponding to each convolutional layer, the model to be pruned is pruned to obtain a pruned model.
[0010] In some embodiments, performing convolution calculation according to the convolution layer information to obtain a filter similarity value corresponding to a filter in each convolution layer includes:
[0011] Obtain at least two filters corresponding to each convolutional layer;
[0012] Perform pairwise convolution calculations on each filter in the convolutional layer to obtain multiple filter similarity values corresponding to each filter in the convolutional layer.
[0013] In some embodiments, calculating the pruning importance index corresponding to each convolutional layer according to the filter similarity values includes:
[0014] Determine the mean or sum of the multiple filter similarity values corresponding to each filter as the filter importance value corresponding to the filter;
[0015] Obtain the pruning importance index corresponding to the convolutional layer according to the filter importance values of each filter in the convolutional layer.
[0016] In some embodiments, obtaining the pruning importance index corresponding to the convolutional layer according to the filter importance values of each filter in the convolutional layer includes:
[0017] Sort the filter importance values of each filter in the convolutional layer to obtain a sorting result;
[0018] Obtain the pruning importance index corresponding to the convolutional layer according to the sorting result.
[0019] In some embodiments, pruning the model to be pruned according to the preset pruning rate and the pruning importance index corresponding to each convolutional layer to obtain a pruned model includes:
[0020] Determine the number of pruned filters in each convolutional layer according to the preset pruning rate;
[0021] Prune the model to be pruned according to the number of pruned filters and the pruning importance index corresponding to each convolutional layer to obtain a pruned model.
[0022] In some embodiments, pruning the model to be pruned according to the number of pruned filters and the pruning importance index corresponding to each convolutional layer to obtain a pruned model includes:
[0023] Determine the pruned filters from the multiple filters of each convolutional layer according to the preset pruning rate and the pruning importance index;
[0024] Prune the pruned filters to obtain the pruned model.
[0025] In some embodiments, after obtaining the pruned model, it further includes:
[0026] Select some filters of the pruned model according to a preset selection rule;
[0027] Perform model training on the remaining filters and corresponding fully connected layers in the pruning model to obtain the pruning model.
[0028] To achieve the above object, a second aspect of the present invention proposes a convolutional neural network model pruning device, including:
[0029] A convolutional layer information acquisition module, configured to acquire convolutional layer information in the model to be pruned;
[0030] A filter similarity calculation module, configured to perform convolutional calculation according to the convolutional layer information to obtain a filter similarity value corresponding to each filter in each convolutional layer;
[0031] A pruning importance index calculation module, configured to calculate a pruning importance index corresponding to each convolutional layer according to the filter similarity value;
[0032] A pruning module, configured to prune the model to be pruned according to a preset pruning rate and the pruning importance index corresponding to each convolutional layer to obtain a pruning model.
[0033] To achieve the above object, a third aspect of the present invention proposes an electronic device, including:
[0034] At least one memory;
[0035] At least one processor;
[0036] At least one program;
[0037] The program is stored in the memory, and the processor executes the at least one program to implement the method as described in the first aspect of the present invention above.
[0038] To achieve the above object, a fourth aspect of the present invention proposes a storage medium, which is a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute:
[0039] The method as described in the first aspect above.
[0040] The pruning method and device for a convolutional neural network model, electronic device, and storage medium proposed in the embodiments of the present invention obtain the convolutional layer information in the model to be pruned, then perform convolutional calculations based on the convolutional layer information to obtain the filter similarity values corresponding to the filters in each convolutional layer, and then calculate the pruning importance index corresponding to each convolutional layer according to the filter similarity values. According to the preset pruning rate and the pruning importance index corresponding to each convolutional layer, the model to be pruned is pruned to obtain the pruned model. In this embodiment, convolutional calculations are performed on the filters in the convolutional layer to obtain the filter importance values, and then the pruning importance index corresponding to each convolutional layer is obtained. By quantifying the importance of the filters in the convolutional layer through convolutional operations and obtaining the redundant information of the filters inside each convolutional layer in the model according to the filter importance values, and then using this redundant information for pruning, the accuracy of pruning the convolutional neural network model can be improved, and the model compression accuracy and operation speed can be enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 FIG. is a flowchart of the pruning method for a convolutional neural network model provided by an embodiment of the present invention.
[0042] Figure 2 FIG. is another flowchart of the pruning method for a convolutional neural network model provided by an embodiment of the present invention.
[0043] Figure 3 FIG. is a schematic diagram of a convolutional layer in a convolutional neural network model.
[0044] Figure 4 FIG. is a schematic diagram of a filter in a convolutional neural network model.
[0045] Figure 5 FIG. is another flowchart of the pruning method for a convolutional neural network model provided by an embodiment of the present invention.
[0046] Figure 6 FIG. is another flowchart of the pruning method for a convolutional neural network model provided by an embodiment of the present invention.
[0047] Figure 7 FIG. is another flowchart of the pruning method for a convolutional neural network model provided by an embodiment of the present invention.
[0048] Figure 8 FIG. is another flowchart of the pruning method for a convolutional neural network model provided by an embodiment of the present invention.
[0049] Figure 9 FIG. is a structural block diagram of the pruning device for a convolutional neural network model provided by an embodiment of the present invention.
[0050] Figure 10 FIG. is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0051] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0054] First, several nouns involved in the present invention are analyzed:
[0055] Convolutional Neural Networks (CNN): It is a type of feedforward neural network with convolutional calculations and a deep structure, and is one of the representative algorithms of deep learning. Convolutional neural networks have the ability of feature learning and can perform translation-invariant classification on input information according to their hierarchical structure. Convolutional neural networks are constructed by imitating the visual perception mechanism of organisms and can perform supervised learning and unsupervised learning. The sharing of convolutional kernel parameters and the sparsity of inter-layer connections in the hidden layer enable convolutional neural networks to process lattice features with relatively small computational amounts. A common convolutional neural network structure is the input layer - convolutional layer - pooling layer - fully connected layer - output layer.
[0056] Convolution: It is a mathematical method of integral transformation, a mathematical operator that generates a third function through two functions f and g, which represents the integral of the product of the function values of the overlapping part of functions f and g after flipping and translation over the overlapping length. If one of the functions participating in the convolution is regarded as an indicator function of an interval, the convolution can also be regarded as a "moving average".
[0057] With the development of Internet technology and artificial intelligence, models based on convolutional neural networks have shown good performance in many tasks, but these models require huge computational overhead and memory usage when used. Since these models usually contain a lot of redundant information, compressing the model to reduce the computational overhead and memory usage during use has become an essential step. Model pruning is an important direction of model compression technology. At present, the detection model and segmentation model in the deep learning model can remove redundant parameters through pruning, ensure the model accuracy as much as possible, compress the model size, and improve the model operation speed.
[0058] The operation of model pruning is mainly divided into two steps: first, select and remove the filters of relatively unimportant convolution kernels, and then fine-tune and optimize the model with the unimportant filters removed to restore the accuracy loss caused by removing the filters. Therefore, the pruning methods in related technologies are all solving how to select relatively unimportant convolution kernel filters. For example, there are three common methods: 1) directly use the size of the weight of the BN layer. This method is easy to understand and easy to implement, but the weight of the BN layer is difficult to measure the amount of information that the relevant filters actually have. There is no strong correlation between the two, so the information correlation between filters cannot be measured; 2) The size of the L1 or L2 norm value of the filter is used as the filter importance judgment indicator. This method has similar shortcomings to the first method. It only depends on the size of the value and does not consider the correlation between filters; 3) The method of using the geometric median of the space where the filter is located. This method first calculates the filter closest to the geometric median of all filters and then prunes it. However, there is no strict evidence to support whether the amount of information of the geometric median can really be replaced by the amount of information of other filters.
[0059] It can be seen that the current model pruning method for selecting pruning filters in the related technology only considers the information of a single filter, does not consider the relationship between filters, and does not obtain the redundant information of the internal filters of each convolutional layer in the model based on the relationship between filters, and then uses this redundant information for pruning, resulting in low pruning accuracy and model compression accuracy of the convolutional neural network model.
[0060] Based on this, the embodiment of the present invention provides a convolutional neural network model pruning method and device, electronic device, and storage medium, which obtains the filter importance value by performing convolution calculation on the filter in the convolution layer, and then obtains the pruning importance index corresponding to each convolution layer. The importance of the filter in the convolution layer is quantified by convolution operation, and the redundant information of the filter inside each convolution layer in the model is obtained according to the importance value of the filter, and then the redundant information is used for pruning, which can improve the accuracy of convolutional neural network model pruning, and improve the model compression accuracy and operation speed.
[0061] Embodiments of the present invention provide a method and apparatus for pruning a convolutional neural network model, an electronic device, and a storage medium, which will be specifically described through the following embodiments. First, the method for pruning a convolutional neural network model in the embodiments of the present invention will be described.
[0062] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0063] Basic artificial intelligence technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0064] The method for pruning a convolutional neural network model provided by the embodiments of the present invention relates to the field of artificial intelligence technology, and particularly to the field of data mining technology. The method for pruning a convolutional neural network model provided by the embodiments of the present invention can be applied to a terminal, or to a server, or can also be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, or a smart watch, etc.; the server can be an independent server, or can be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; the software can be an application that implements the method for pruning a convolutional neural network model, etc., but is not limited to the above forms.
[0065] The present invention can be used in numerous general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0066] Figure 1 is an optional flowchart of the convolutional neural network model pruning method provided by an embodiment of the present invention. Figure 1 The method in
[0067] Step S110, obtain the convolutional layer information in the model to be pruned.
[0068] In one embodiment, the model to be pruned is a convolutional neural network model, and the convolutional layer information can be the filters included in the convolutional layer. When pruning the model to be pruned in this embodiment, considering the relationship between filters and the redundant information of filters inside each convolutional layer in the model, the importance of filters is quantified to improve the accuracy of convolutional neural network model pruning and enhance the model compression accuracy and operation speed.
[0069] Step S120, perform convolutional calculations according to the convolutional layer information to obtain the filter similarity values corresponding to the filters in each convolutional layer.
[0070] In one embodiment, referring to Figure 2 , step S120 includes but is not limited to steps S121 to S122:
[0071] Step S121, obtain the filters corresponding to each convolutional layer, where each convolutional layer corresponds to at least two filters.
[0072] Step S122, perform pairwise convolutional calculations on the filters in each convolutional layer to obtain multiple filter similarity values corresponding to each filter in the convolutional layer.
[0073] In one embodiment, each convolutional layer contains multiple filters. Referring to Figure 3 , is a schematic diagram of a convolutional layer. Figure 3In this context, X and Y are two consecutive feature maps in a convolutional neural network model (such as the model to be pruned). There are multiple convolutional layers between X and Y. These filters are used to identify certain specific features of the image. Each filter slides over the feature map of the previous layer. For example, Conv1 shown in the figure represents one of the convolutional layers. After passing through the calculation of the convolutional layer, Feature map X can obtain Feature map Y. Among them, each convolutional layer is composed of multiple filters, and in each filter, there are several channels from front to back.
[0074] In one embodiment, a filter is a tool used to extract features in an image. For example, it can be used to extract edge features, texture features, etc. A filter is composed of multiple 2D filters. Refer to Figure 4 , which is a schematic diagram of a filter. Figure 4 The shown filter can be used to detect edges and belongs to the common LOG filter in image processing.
[0075] In this embodiment, the model to be pruned is a convolutional neural network model. The values in the filters of the convolutional neural network model are obtained by the model relying on training data for training, rather than artificially designed filters. When extracting features, the filters in the convolutional layer use the convolution operation: First, the filter filters the image, and then the filter is sequentially slid over a certain area of the original image and multiplied point by point with the original pixel values in that area. If the features in the image are very similar to the features of the filter, the value obtained after multiplying point by point and summing is very high. When there is no corresponding relationship between the image area and the filter, the value obtained by multiplying point by point and summing is very small. If the filter multiplies and sums with itself, a relatively large value will be obtained.
[0076] Since each convolutional layer is composed of multiple filters, by comparing the filters, if it is proved according to the comparison result that two filters are relatively similar, then the roles played by these two filters in this convolutional layer are also similar. Therefore, they can be regarded as redundant information in this convolutional layer, and one of the filters can be removed for model pruning. The purpose of pruning the model to be pruned is to remove a certain proportion of filters. For example, Figure 3 the filter shown by the dotted line in can be pruned.
[0077] In the above embodiments, to keep the final prediction performance of the pruned model from decreasing as much as possible, it is necessary to evaluate the importance of each filter in the convolutional layer. According to the importance, the least important partial filters are removed, so as to minimize the impact on the model accuracy. In the related art, there is a method of directly calculating the difference between the corresponding elements of the values of two filters to calculate the similarity between the filters. This method has the following defects: 1) It does not consider the role of the filters in extracting features; 2) Moreover, all elements in the filters have the same role, without considering the different positions of the elements in the filters, while the position plays a very important role in feature extraction and cannot be ignored; 3) Using the direct difference method has no physical meaning and lacks theoretical basis.
[0078] In one embodiment, the convolution calculation between two filters is used to measure the similarity of the filters. Since the more similar two filters are, the more similar the features extracted by these two filters are, and the larger the value obtained by performing the convolution operation between these two filters is, which further indicates that the importance of the corresponding filter is weaker, and it does not play a greater role in the calculation of this convolutional layer, and its information is redundant, and this filter can be deleted during the pruning process. On the contrary, if the value obtained by performing the convolution operation is large, it indicates that this filter has a significant impact on the result in the calculation of this convolutional layer, and it contains significant information and cannot be deleted. Therefore, in this embodiment, first, the filters in each convolutional layer are obtained, and pairwise convolution calculations are performed on the filters in each convolutional layer to obtain multiple filter similarity values corresponding to each filter in the convolutional layer.
[0079] In one embodiment, the convolution operation process between filters is described as: calculating the sum of the products of the corresponding positions of two filters. The process of pairwise convolution between filters is described as follows:
[0080] For example, a certain convolutional layer in the model to be pruned includes 5 filters, which are: filter 1 (denoted as F1), filter 2 (denoted as F2), filter 3 (denoted as F3), filter 4 (denoted as F4), and filter 5 (denoted as F5). Then, the pairwise convolution of each filter to obtain the filter similarity value includes: {S1 = F1 * F2, S2 = F1 * F3, S3 = F1 * F4, S4 = F1 * F5, S5 = F2 * F3, S6 = F2 * F4, S7 = F2 * F5, S8 = F3 * F4, S9 = F3 * F5, S10 = F4 * F5}, where "*" represents the convolution operation.
[0081] That is, in the above embodiments, each filter includes multiple filter similarity values. Specifically, the filter similarity values of filter 1 include: {S1, S2, S3, S4}, the filter similarity values of filter 2 include: {S1, S5, S6, S7}, the filter similarity values of filter 3 include: {S2, S5, S8, S9}, the filter similarity values of filter 4 include: {S3, S6, S8, S10}, and the filter similarity values of filter 5 include: {S4, S7, S9, S10}.
[0082] Step S130: Calculate the pruning importance index corresponding to each convolutional layer according to the filter similarity values.
[0083] In one embodiment, referring to Figure 5 , step S130 includes but is not limited to steps S131 to S132:
[0084] Step S131: Determine the mean or sum of the multiple filter similarity values corresponding to each filter as the filter importance value corresponding to the filter.
[0085] In one embodiment, that is, the filter importance value corresponding to each filter can be calculated by means of summing and averaging or summing. That is, the sum can be obtained only by summing, or the mean can be obtained by averaging after summing. The specific calculation method can be selected according to actual needs.
[0086] For example, taking the method of averaging after summing to calculate the filter importance value in the above example, it is expressed as:
[0087] The filter importance value of filter 1 is: (S1 + S2 + S3 + S4) / 4;
[0088] The filter importance value of filter 2 is: (S1 + S5 + S6 + S7) / 4;
[0089] The filter importance value of filter 3 is: (S2 + S5 + S8 + S9) / 4;
[0090] The filter importance value of filter 4 is: (S3 + S6 + S8 + S10) / 4;
[0091] The filter importance value of filter 5 is: (S4 + S7 + S9 + S10) / 4.
[0092] Step S132: Obtain the pruning importance index corresponding to the convolutional layer according to the filter importance value of each filter in the convolutional layer.
[0093] In one embodiment, referring to Figure 6 , step S132 includes but is not limited to steps S1321 to S1322:
[0094] Step S1321: Sort the filter importance values of each filter in the convolutional layer to obtain a sorting result.
[0095] Step S1322: Obtain the pruning importance index corresponding to the convolutional layer according to the sorting result.
[0096] In one embodiment, the filters included in each convolutional layer are sorted in descending order according to importance. The filters with higher rankings have greater importance. It can be understood that it is also possible to sort from large to small. Then, when selecting pruning filters for pruning, they are selected in reverse order, which is not specifically limited here. The importance information of each filter in the convolutional layer is the pruning importance index corresponding to this convolutional layer. Then, the model to be pruned is pruned according to the pruning importance index.
[0097] Step S140: Prune the model to be pruned according to the preset pruning rate and the pruning importance index corresponding to each convolutional layer to obtain a pruned model.
[0098] In one embodiment, after obtaining the pruning importance index corresponding to each convolutional layer, filters can be selected for pruning according to this pruning importance index. For example, if the filter importance value of the selected filter is larger, it means that the filter is more important in the model to be pruned. If the corresponding filter is removed, it will have a greater impact on the performance of the model to be pruned. Therefore, in this embodiment, the filters with smaller filter importance values are selected as pruning filters during pruning, that is, the filters with lower rankings in the sorting result are pruned.
[0099] In one embodiment, referring to Figure 7 , Step S140 includes but is not limited to Steps S141 to S142:
[0100] Step S141: Determine the number of pruning filters in each convolutional layer according to the preset pruning rate.
[0101] In one embodiment, the preset pruning rate needs to be set according to actual needs during the pruning operation. If the pruning rate is too high, it will lead to a decrease in model accuracy. If the pruning rate is too low, it will lead to poor improvement in model operation efficiency. Therefore, the preset pruning rate needs to be set according to actual needs, and pruning is performed according to the preset pruning rate to determine how many filters to remove, that is, the number of pruning filters in each convolutional layer can be determined according to the preset pruning rate.
[0102] Step S142: Prune the model to be pruned according to the number of pruning filters and the pruning importance index corresponding to each convolutional layer to obtain a pruned model.
[0103] In one embodiment, pruning filters are determined from multiple filters of each convolutional layer according to a preset pruning rate and pruning importance index, and the pruning model is obtained by pruning the pruning filters. For example, if the preset pruning rate is set to 75%, 3 / 4 of the filters are removed through the pruning operation, and the 3 / 4 filters with smaller filter importance values are removed as pruning filters according to the above sorting result. The filters with smaller filter importance values play a weaker role in the model to be pruned and there is redundant information. Therefore, removing them will not have a great impact on the performance of the model to be pruned, while effectively reducing the model parameters of the model to be pruned, reducing the computational amount and storage space of the model to be pruned.
[0104] In some embodiments, after obtaining the pruning model, in order to compensate for the cumulative error caused by filter pruning, it is necessary to fine-tune the pruning model to restore the accuracy of the model. Referring to Figure 8 , the steps of fine-tuning the pruning model include but are not limited to steps S810 to S820:
[0105] Step S810, select some filters of the pruning model according to a preset selection rule.
[0106] In one embodiment, the preset selection rule may be to select some filters close to the input end of the pruning model, and the number of filters can be set according to actual needs and is not limited here.
[0107] Step S820, train the remaining filters and the corresponding fully connected layers in the pruning model to obtain the pruning model.
[0108] In one embodiment, on the target data set, the remaining filters (such as the filters close to the output end) and the corresponding fully connected layers are trained to achieve fine-tuning compensation of the pruning model, and the purpose of not affecting the model operation performance under the condition of maximizing the model compression scale is achieved.
[0109] In a specific application scenario, taking the VGG16 model as an example of the model to be pruned to verify the effectiveness of the convolutional neural network model pruning method in the above embodiment. Among them, the VGG16 model is a convolutional neural network model suitable for classification and localization tasks. The model consists of 5 convolutional layers, 3 fully connected layers, and a softmax output layer. The layers are separated by max-pooling (maximization pool), and the activation units of all hidden layers use the ReLU function. And the VGG16 model uses convolutional layers with multiple smaller convolutional kernels (such as 3x3) instead of a convolutional layer with a larger convolutional kernel. On the one hand, it can reduce the parameters, and on the other hand, it is equivalent to performing more non-linear mappings, increasing the fitting / expression ability of the network.
[0110] Meanwhile, the dataset used for verification is the CIFAR-10 dataset. There are a total of 60,000 color images in the CIFAR-10 dataset. These images are 32*32 in size and are divided into 10 categories, with each category containing 6,000 images. 50,000 images in this dataset are used for the training process, which altogether constitute 5 training batches, with each batch containing 10,000 images; another 10,000 images are used for the testing process and form a separate batch. In the data of the test batch, 1,000 images are randomly selected from each of the 10 categories, and the remaining images are randomly arranged to form the training batch.
[0111] During the verification process, pruning model compression training is carried out for 100 Epochs in each experiment. The hardware used for verification is the NVIDIA V100 GPU, and the PyTorch framework is adopted. The preset pruning rate is 50%.
[0112] The compression methods (i.e., pruning methods) used for verification include:
[0113] 1) APoZ model pruning method: That is, the pruning object is determined according to the percentage of the output of the activation function being zero, and APoZ is used to predict the importance of each filter in the network.
[0114] 2) Activation value minimum model pruning method: That is, before activation, the model weights and biases are set to 0 first, and after activation, the filters that have the least impact on the activation value of the next layer are pruned, that is, the filters with the minimum average activation value (meaning the least number of uses).
[0115] 3) L1 model pruning method: That is, pruning is carried out based on the L1-norm weight parameters. Based on the L1-norm weight parameters for pruning, a certain proportion of filters are cut off using a smaller L1-norm for each convolutional layer.
[0116] 4) The convolutional neural network model pruning method in the above embodiments.
[0117] According to the verification results, the operation accuracy of the pruning model to be pruned without pruning is 93.99%. Referring to the following table, it is a comparison of the operation accuracies of the pruning models obtained by the above three different pruning methods:
[0118] Pruning method Unpruned model APoZ Minimum activation value L1 This application Operation accuracy 93.99% 92.24% 92.81% 93.05% 93.42%
[0119] As can be seen from the above table, the operation accuracy of the convolutional neural network model pruning method in the embodiments of the present application is the highest, which is 93.41%, approaching the operation accuracy of 93.99% of the to-be-pruned model without pruning. It can be seen that the convolutional neural network model pruning method in the embodiments of the present application will not have a great impact on the operation accuracy performance of the to-be-pruned model, and also has a certain regularization effect. At the same time, it can effectively reduce the model parameters of the to-be-pruned model, and reduce the operation amount and storage space of the to-be-pruned model.
[0120] The convolutional neural network model pruning method proposed in the embodiments of the present disclosure obtains the convolutional layer information in the to-be-pruned model, then performs convolutional calculations according to the convolutional layer information to obtain the filter similarity values corresponding to the filters in each convolutional layer, and then calculates the pruning importance index corresponding to each convolutional layer according to the filter similarity values. According to the preset pruning rate and the pruning importance index corresponding to each convolutional layer, the to-be-pruned model is pruned to obtain a pruned model. In this embodiment, convolutional calculations are performed on the filters in the convolutional layer to obtain the filter importance values, and then the pruning importance index corresponding to each convolutional layer is obtained. By quantifying the importance of the filters in the convolutional layer through convolutional operations, and obtaining the redundant information of the filters inside each convolutional layer in the model according to the importance values of the filters, and then using this redundant information for pruning, the accuracy of convolutional neural network model pruning can be improved, and the model compression accuracy and operation speed can be enhanced.
[0121] In addition, the embodiments of the present invention also provide a convolutional neural network model pruning device, which can implement the above-mentioned convolutional neural network model pruning method. Referring to Figure 9 , the device includes:
[0122] A convolutional layer information acquisition module 910, configured to acquire convolutional layer information in the to-be-pruned model;
[0123] A filter similarity calculation module 920, configured to perform convolutional calculations according to the convolutional layer information to obtain the filter similarity values corresponding to the filters in each convolutional layer;
[0124] A pruning importance index calculation module 930, configured to calculate the pruning importance index corresponding to each convolutional layer according to the filter similarity values;
[0125] A pruning module 940, configured to prune the to-be-pruned model according to the preset pruning rate and the pruning importance index corresponding to each convolutional layer to obtain a pruned model.
[0126] The specific implementation manner of the convolutional neural network model pruning device in this embodiment is basically the same as that of the above-mentioned convolutional neural network model pruning method, and will not be elaborated here.
[0127] The embodiments of the present invention also provide an electronic device, including:
[0128] At least one memory;
[0129] At least one processor;
[0130] At least one program;
[0131] The program is stored in the memory, and the processor executes the at least one program to implement the convolutional neural network model pruning method described above in the embodiments of the present invention. The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA for short), an in-vehicle computer, etc.
[0132] Please refer to Figure 10 , Figure 10 , which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0133] A processor 1001, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention;
[0134] A memory 1002, which can be implemented in forms such as a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory). The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the convolutional neural network model pruning method of the embodiments of the present invention;
[0135] An input / output interface 1003, which is used to implement information input and output;
[0136] A communication interface 1004, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.); and
[0137] A bus 1005, which transmits information between various components of the device (such as the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004);
[0138] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 are communicatively connected to each other inside the device through the bus 1005.
[0139] An embodiment of the present invention also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-mentioned convolutional neural network model pruning method.
[0140] The convolutional neural network model pruning method, convolutional neural network model pruning device, electronic device, and storage medium proposed in the embodiments of the present invention obtain the filter importance value by performing convolution calculations on the filters in the convolutional layer, and then obtain the pruning importance index corresponding to each convolutional layer. By quantifying the importance of the filters in the convolutional layer through convolution operations, and obtaining the redundant information of the filters inside each convolutional layer in the model based on the importance value of the filters, and then using this redundant information for pruning, it is possible to improve the accuracy of convolutional neural network model pruning, and enhance the model compression accuracy and operation speed.
[0141] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0142] The embodiments described in the embodiments of the present invention are for more clearly illustrating the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.
[0143] Those skilled in the art can understand that Figures 1-8 the technical solutions shown do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0144] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0145] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.
[0146] In the description of the present invention and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0147] It should be understood that in the present invention, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0148] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0149] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0150] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0151] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0152] The preferred embodiments of the embodiments of the present invention have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of the present invention. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present invention shall be within the scope of rights of the embodiments of the present invention.
Claims
1. A method for pruning a convolutional neural network model, characterized in that, the convolutional neural network model is used for image processing, and the method includes: Obtaining the convolutional layer information of the model to be pruned; Performing convolutional calculations according to the convolutional layer information to obtain filter similarity values corresponding to the filters in each convolutional layer; Calculating a pruning importance index corresponding to each convolutional layer according to the filter similarity values; Pruning the model to be pruned according to a preset pruning rate and the pruning importance index corresponding to each convolutional layer to obtain a pruned model; The performing convolutional calculations according to the convolutional layer information to obtain filter similarity values corresponding to the filters in each convolutional layer includes: Obtaining at least two filters corresponding to each convolutional layer; Performing a convolutional operation on each filter in the convolutional layer with an image to obtain image features, and obtaining multiple filter similarity values corresponding to each filter in the convolutional layer according to the pairwise similarity between the corresponding image features; The calculating a pruning importance index corresponding to each convolutional layer according to the filter similarity values includes: Determining the mean or sum of the multiple filter similarity values corresponding to each filter as the filter importance value corresponding to the filter; Obtaining the pruning importance index corresponding to the convolutional layer according to the filter importance values of each filter in the convolutional layer; The obtaining the pruning importance index corresponding to the convolutional layer according to the filter importance values of each filter in the convolutional layer includes: Sorting the filter importance values of each filter in the convolutional layer to obtain a sorting result; Obtaining the pruning importance index corresponding to the convolutional layer according to the sorting result.
2. The method for pruning a convolutional neural network model according to claim 1, characterized in that, the pruning the model to be pruned according to a preset pruning rate and the pruning importance index corresponding to each convolutional layer to obtain a pruned model includes: Determining the number of pruned filters in each convolutional layer according to the preset pruning rate; Pruning the model to be pruned according to the number of pruned filters and the pruning importance index corresponding to each convolutional layer to obtain a pruned model.
3. The method for pruning a convolutional neural network model according to claim 2, characterized in that, the pruning the model to be pruned according to the number of pruned filters and the pruning importance index corresponding to each convolutional layer to obtain a pruned model includes: Determining the pruned filters from the multiple filters in each convolutional layer according to the preset pruning rate and the pruning importance index; Pruning the pruned filters to obtain the pruned model.
4. The method for pruning a convolutional neural network model according to any one of claims 1 to 3, characterized in that, after obtaining the pruned model, it further includes: Selecting some filters of the pruned model according to a preset selection rule; Training the remaining filters and the corresponding fully connected layers in the pruned model to obtain the pruned model.
5. A convolutional neural network model pruning device, characterized in that, it includes: A convolutional layer information acquisition module for acquiring the convolutional layer information in the model to be pruned; A filter similarity calculation module, configured to perform convolution calculation according to the convolution layer information to obtain a filter similarity value corresponding to each filter in each convolution layer; A pruning importance index calculation module, configured to calculate a pruning importance index corresponding to each convolution layer according to the filter similarity value; A pruning module, configured to prune the model to be pruned according to a preset pruning rate and the pruning importance index corresponding to each convolution layer to obtain a pruned model; The performing convolution calculation according to the convolution layer information to obtain a filter similarity value corresponding to each filter in each convolution layer includes: Obtaining at least two filters corresponding to each convolution layer; Performing a convolution operation on each filter in the convolution layer with an image to obtain image features, and obtaining a plurality of the filter similarity values corresponding to each filter in the convolution layer according to the pairwise similarity between the corresponding image features; The calculating a pruning importance index corresponding to each convolution layer according to the filter similarity value includes: Determining the mean or sum of the multiple filter similarity values corresponding to each filter as the filter importance value corresponding to the filter; Obtaining the pruning importance index corresponding to the convolution layer according to the filter importance value of each filter in the convolution layer; The obtaining the pruning importance index corresponding to the convolution layer according to the filter importance value of each filter in the convolution layer includes: Sorting the filter importance values of each filter in the convolution layer to obtain a sorting result; Obtaining the pruning importance index corresponding to the convolution layer according to the sorting result.
6. An electronic device, characterized in that it includes: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes the at least one program to implement: the method according to any one of claims 1 to 4.
7. A storage medium, the storage medium being a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute: the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Neural network model acceleration method and platform based on filter distribution
CN112561041A
Model pruning method, device and equipment and storage medium
CN113240085A