Convolutional neural network pruning method and device based on static and dynamic combination
Through the static and dynamic joint convolutional neural network pruning method, combined with the advantages of static and dynamic pruning, the problems of irreversible pruning and high computational cost in the existing technology are solved, and the efficient compression and performance maintenance of the model are achieved.
Patent Information
- Application Number
- CN202411859766.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-17
AI Technical Summary
The existing convolutional neural network pruning methods have the problem of static pruning irreversibility and ignoring dynamic relationships. Dynamic pruning requires additional computing cost and time overhead, which affects the actual inference speed.
The convolutional neural network pruning method based on static dynamic joint is adopted to remove parameters with small contributions through static pruning, and the channel importance is dynamically evaluated according to the characteristics of the input sample during the training process, combining the two to achieve further compression and optimization of the model.
The model size reduction and computational efficiency are achieved, while keeping the model performance undamaged, avoiding the irreversibility of static pruning and the additional computational cost of dynamic pruning.
Smart Images

Figure CN119990229A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of compressing and pruning convolutional neural networks, and in particular to a convolutional neural network pruning method and device based on static and dynamic combination. Background Art
[0002] At present, convolutional neural networks have achieved remarkable performance in many computer vision tasks, such as image classification, anomaly detection, and target detection. Convolutional neural networks require huge memory overhead and high-performance computing units, which makes it difficult to deploy and run convolutional neural networks on resource-constrained embedded devices. There are still many obstacles to the real application of convolutional neural networks in real life.
[0003] Existing network model compression generally reduces the size of convolutional neural networks and reduces computational complexity through methods such as pruning, quantization, low-rank decomposition, and knowledge distillation, thereby reducing the resource requirements of network models. Among them, pruning reduces the size of network models by identifying and removing unimportant weights in neural networks, and is currently a popular method in the field of network model compression.
[0004] Pruning can be divided into static pruning and dynamic pruning. Static pruning prunes the parameters of the network model and designs a fixed sparse structure for all samples. This method has been widely used in most weight pruning, channel pruning and filter pruning, and has shown significant effects in network model compression and acceleration. However, static pruning is usually irreversible. Once some parts are pruned, these parts cannot be restored in the subsequent process, which may cause the network model to permanently lose some important parameters, which will have an adverse effect on the accuracy and generalization ability of the network model. At the same time, static pruning ignores the dynamic relationship between channel saliency and input samples, and does not take into account the different importance of channels in different input samples. Compared with static pruning, dynamic pruning takes a more flexible approach. Dynamic pruning does not permanently remove any parameters in the network, but dynamically determines which part of the convolutional neural network should participate in the calculation based on the characteristics of each input image, thereby effectively reducing unnecessary calculations and redundancy. However, dynamic pruning requires certain additional operations to be repeated for each input sample, such as index replication or weight replication, which brings additional computational costs and time overhead. Therefore, although dynamic pruning can significantly improve inference efficiency in theory, in practical applications, the actual inference speed may be much slower than the theoretical speed due to these additional operations.
[0005] In summary, the current pruning methods have the following problems: First, static pruning generates a fixed pruning structure and cannot prune flexibly, which affects the accuracy and generalization of the network model and ignores the dynamic relationship between channel saliency and input samples. Second, dynamic pruning requires additional operations such as index replication or weight replication to be repeated for each input sample, which reduces the actual inference speed.
[0006] For example: the invention application with application number 202410904785.9 discloses a model pruning method and a computer-readable storage medium. The application prunes the first preset number of filter channels in the sorting based on a preset global pruning rate; obtains the model to be pruned after pruning each set of convolutional layers, and obtains the target model. Another example: the invention application with application number 202410601678.9 discloses a fast pruning method, system, device and storage medium for convolutional neural networks. The application scheme uses the gradient descent method to generate the pruning rate, which can search for the optimal pruning rate in fewer rounds, and only needs one fine-tuning during the pruning process, which greatly improves the pruning efficiency.
[0007] Although the above scheme improves the pruning effect of the model, it also has the following problems: static pruning and dynamic pruning are not used together, and the complementary advantages of the two pruning strategies are not realized, which can easily lead to damage to the model performance and increase additional computing costs and time expenses. Summary of the invention
[0008] In view of the above problems, the purpose of the present invention is to provide a convolutional neural network pruning method and device based on static and dynamic combination, which utilizes the complementary advantages of static pruning and dynamic pruning strategies to achieve model size reduction and improve computational efficiency, while keeping the performance of the model intact as much as possible.
[0009] The embodiment of the present invention provides a convolutional neural network pruning method and device based on static and dynamic combination.
[0010] A first aspect: A convolutional neural network pruning method based on static and dynamic combination, comprising the steps of:
[0011] S1. Static pruning of convolutional neural network parameters, screening out parameters with low weights for pruning, and optimizing convolutional neural network parameters;
[0012] S2. During the training of the convolutional neural network, the importance of the remaining channels is dynamically evaluated according to the specific characteristics of the input samples, and the parameters of the convolutional neural network are dynamically pruned and optimized;
[0013] S3. Repeat steps S1 to S2 until the convolutional neural network pruning is completed.
[0014] Furthermore, in the static pruning stage in S1, parameters with low importance are screened out based on a trainable threshold T.
[0015] Furthermore, in the dynamic pruning stage in S2, the channel attention mechanism is optimized to dynamically evaluate the importance of the remaining channels.
[0016] Furthermore, the method of screening out parameters with low importance based on the trainable threshold T comprises the steps of:
[0017] S11. Use the trainable threshold T to evaluate the importance of the convolution weight parameters, and pass each convolution weight parameter through the threshold function to obtain the corresponding mask M i , the formula is:
[0018]
[0019] Among them, M∈R K*K is the mask set for each convolution kernel, T represents the threshold, σ represents the sigmoid function, M ij ∈M i , i represents the number of convolution layers, and j represents the number of convolution kernels.
[0020] S12, if the convolution weight parameter exceeds the threshold, the corresponding mask is set to 1, the weight is retained and participates in subsequent training, otherwise, the convolution weight is clipped;
[0021] S13, the original weight parameters Perform dot multiplication with the obtained mask parameter, the formula is:
[0022]
[0023] in, is the weight parameter of the i-th convolutional layer, M i is the convolution kernel mask of the i-th convolution layer.
[0024] Furthermore, the channel attention mechanism is optimized to dynamically evaluate the importance of the remaining channels, including the steps of:
[0025] S21. For input feature X∈R C*H*W , extract the global feature F through global average pooling and global maximum pooling a and texture feature F m , the formula is:
[0026]
[0027] Among them, H is the height of the input feature, W is the width of the input feature, and F a ∈R C*1*1 is the global information of the input features, Fm ∈R C*1*1 is the texture information in the input feature;
[0028] S22, F a and F m One-dimensional convolution is used to realize local interaction between channels, and the Sigmoid activation function is used to calculate the attention parameter. After splicing, the attention weight F of the i-th convolution layer is generated. i :
[0029] F i =σ(conv(F a ))+σ(conv(F m ))
[0030] Among them, conv represents one-dimensional convolution, σ represents Sigmoid activation function;
[0031] Furthermore, the S2 dynamically prunes and optimizes the parameters of the convolutional neural network, including: i The weight parameters obtained by static pruning Multiply point by point to get the weight parameter
[0032]
[0033] Among them, i represents the number of convolutional layers, It is the convolution weight parameter after static pruning.
[0034] Furthermore, in the static-dynamic joint pruning process, the trainable threshold T is optimized using the loss function, as follows:
[0035]
[0036] Among them, L c represents the cross entropy loss, λ represents the coefficient factor, N represents the number of convolutional layers, and C represents the number of filters in the convolutional layer.
[0037] A second aspect: A device for a convolutional neural network pruning method based on static and dynamic combination, comprising:
[0038] Static pruning module, which performs static pruning on the parameters of the convolutional neural network, selects parameters with low weights for pruning, and optimizes the parameters of the convolutional neural network;
[0039] Dynamic pruning module: During the training process of the convolutional neural network, the importance of the remaining channels is dynamically evaluated according to the specific characteristics of the input samples, and the parameters of the convolutional neural network are dynamically pruned and optimized;
[0040] The static pruning module is combined with the dynamic pruning module to jointly complete the pruning of the convolutional neural network.
[0041] A third aspect: An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method provided in the first aspect when executing the program.
[0042] A fourth aspect: A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method provided in the first aspect.
[0043] Beneficial effects of the present invention:
[0044] 1. The present invention removes parameters and structures with smaller contributions by applying static pruning. In the static pruning stage, a sparse method based on a trainable threshold is used to identify convolution weights with smaller contributions. By setting an appropriate threshold, parameters with lower importance can be effectively screened out and pruned, thereby completing the static pruning of the model without significantly damaging the model performance.
[0045] 2. The present invention optimizes the traditional channel attention mechanism in the dynamic pruning stage by applying dynamic pruning, so that it can more accurately generate attention weights for the input feature map, thereby more accurately evaluating the importance of the remaining channels after static pruning, and dynamically adjusting the pruning process of the model according to the specific features of each input sample.
[0046] 3. The present invention combines static pruning and dynamic pruning, making full use of the complementary advantages of the two pruning strategies to achieve model size reduction and improved computational efficiency, while trying to keep the performance of the model intact. It is not only more flexible than static pruning, but also reduces additional computational cost and time overhead compared to dynamic pruning.
[0047] 4. The present invention improves the channel attention mechanism in the dynamic pruning stage, generates more accurate attention weights to guide further pruning work, and dynamically adjusts the pruning structure according to different input samples; significantly reduces the computational complexity and number of parameters of the network model while maintaining the accuracy of the convolutional neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Schematic diagram of the flow of the network pruning method of the present invention;
[0049] Figure 2 It is a structural schematic diagram of the network pruning device of the present invention;
[0050] Figure 3 It is a principle flow chart of the network pruning method of the present invention;
[0051] Figure 4 It is a schematic diagram of the structure of the electronic device of the present invention. DETAILED DESCRIPTION
[0052] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar symbols throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0053] When pruning convolutional neural networks, static pruning generates a fixed pruning structure and cannot prune flexibly, which affects the accuracy and generalization of the network model; while dynamic pruning requires repeated execution of additional operations such as index copying or weight copying for each input sample, which reduces the actual inference speed.
[0054] In view of the above problems, the present invention provides a convolutional neural network pruning method based on static and dynamic combination. Figure 1 A schematic diagram of a convolutional neural network pruning method based on static and dynamic combination provided in an embodiment of the present invention, the method comprising:
[0055] S1. Static pruning of convolutional neural network parameters is performed to filter out parameters with low weights for pruning and optimize the convolutional neural network parameters.
[0056] like Figure 2 As shown in the static pruning process of convolutional neural network parameters, in the static pruning stage, the parameters with low importance are screened out based on the trainable threshold T, and each convolution layer weight is passed through a smooth threshold function to obtain a corresponding mask M i , the formula is defined as:
[0057]
[0058] Among them, M∈R K*K is the mask set for each convolution kernel, T represents the threshold, σ represents the sigmoid function, M ij ∈M i , i represents the number of convolution layers, j represents the number of convolution kernels, that is, the number of output channels.
[0059] The mask M is generated based on the comparison result between the weight parameter of the convolution kernel and the threshold T. If the convolution weight parameter exceeds the threshold T, the corresponding mask is set to 1, and the weight is retained and participates in subsequent training. Otherwise, the convolution weight is clipped.
[0060] The original weight parameter With the obtained mask M i We do this by performing a dot product:
[0061]
[0062] in, is the weight parameter of the i-th convolutional layer, M i is the convolution kernel mask of the i-th convolution layer.
[0063] S2. During the training process of the convolutional neural network, the importance of the remaining channels is dynamically evaluated according to the specific characteristics of the input samples, and the convolutional neural network parameters are dynamically pruned and optimized.
[0064] The dynamic pruning process is as follows Figure 3 As shown. First, for the input feature X∈R C*H*W , extract important information of input features from different angles, and extract the global features F of the image through global average pooling and global maximum pooling operations. a and texture feature F m , the formula is:
[0065]
[0066] Among them, H represents the height of the input feature, W represents the width of the input feature, and F a ∈R C*1*1 reflects the global information of the input features, F m ∈R C*1*1 Highlights the texture information in the input features.
[0067] Traditional channel attention usually uses two fully connected layers to complete the operation of dimensionality reduction and then dimensionality increase to generate the attention weight of each channel. However, this design adds extra parameters and increases the burden on the network model. In addition, the dimensionality reduction operation of the first fully connected layer may cause important information loss, which will also affect the accuracy of the network model.
[0068] The present invention uses one-dimensional convolution to replace two fully connected layers to process the channel dimension of the input feature map. Through one-dimensional convolution, the attention weight of each channel is affected by the current channel and adjacent channels, and the network model can more effectively capture the dynamic relationship between channels.
[0069] Therefore, we can a and F m One-dimensional convolution is used to realize local interaction between channels, and the Sigmoid activation function is used to calculate the attention parameter. After splicing, the attention weight F of the i-th convolution layer is generated. i , represents the importance of each channel, and the formula is expressed as:
[0070]
[0071] Among them, conv represents one-dimensional convolution and σ represents the Sigmoid activation function.
[0072] S3. Repeat steps S1 to S2 until the convolutional neural network pruning is completed.
[0073] Finally, the weight parameters obtained by static pruning are And the attention weight F generated by dynamic pruning i Multiply point by point to get the weight parameter
[0074]
[0075] Through dynamic combined with static pruning, the network model parameters can be further compressed while ensuring that the performance of the network model will not be greatly affected until the convolutional neural network pruning is completed.
[0076] In the static joint dynamic pruning process, the trainable threshold T is optimized using the loss function, and the loss function formula is:
[0077]
[0078] Among them, L c represents the cross entropy loss, λ represents the coefficient factor, N represents the number of convolutional layers, and C represents the number of filters in the convolutional layer.
[0079] During the training process, the threshold T will be updated during the back propagation of the loss function, and will self-adjust according to the performance of the network model on the training data in order to find the optimal threshold setting. Ultimately, the threshold T in each filter will reach an optimal state, which not only reduces the computational complexity of the network model, but also maintains the performance of the network model.
[0080] like Figure 2 As shown, the present invention also discloses a convolutional neural network pruning device for the above method, the device comprising:
[0081] Static pruning module, which performs static pruning on the parameters of the convolutional neural network, selects parameters with low weights for pruning, and optimizes the parameters of the convolutional neural network;
[0082] Dynamic pruning module: During the training process of the convolutional neural network, the importance of the remaining channels is dynamically evaluated according to the specific characteristics of the input samples, and the parameters of the convolutional neural network are dynamically pruned and optimized;
[0083] The static pruning module is combined with the dynamic pruning module to jointly complete the pruning of the convolutional neural network.
[0084] In the present invention, a static pruning module is first applied to remove parameters and structures with smaller contributions, and then the model is further sparsed and adjusted through a dynamic pruning module during the training process.
[0085] In the static pruning stage, a sparse method based on a trainable threshold is used to identify convolution weights with smaller contributions. By setting an appropriate threshold, parameters with lower importance can be effectively screened out and pruned, and the static pruning of the model is completed without significantly damaging the performance of the model. In the dynamic pruning stage, the traditional channel attention mechanism is optimized. Through one-dimensional convolution, the attention weight of each channel is affected by the channel and the adjacent channels. The model can more effectively capture the dynamic relationship between channels and can more accurately generate attention weights for the input feature map, thereby more accurately evaluating the importance of the remaining channels after static pruning, and dynamically adjusting the pruning process of the model according to the specific features of each input sample. In the present invention, the complementary advantages of the two pruning strategies are fully utilized to achieve the reduction of model size and the improvement of computational efficiency, while keeping the performance of the model as unimpaired as possible.
[0086] The present invention also provides an electronic device, Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, such as Figure 4 As shown, the electronic device may include: a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory communicate with each other via the communications bus. The processor may call the logic instructions in the memory, for example, to execute the following method:
[0087] S1. Static pruning of convolutional neural network parameters, screening out parameters with low weights for pruning, and optimizing convolutional neural network parameters;
[0088] S2. During the training of the convolutional neural network, the importance of the remaining channels is dynamically evaluated according to the specific characteristics of the input samples, and the parameters of the convolutional neural network are dynamically pruned and optimized;
[0089] S3. Repeat steps S1 to S2 until the convolutional neural network pruning is completed.
[0090] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0091] An embodiment of the present invention further provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in each of the above embodiments is implemented, for example, including:
[0092] S1. Static pruning of convolutional neural network parameters, screening out parameters with low weights for pruning, and optimizing convolutional neural network parameters;
[0093] S2. During the training of the convolutional neural network, the importance of the remaining channels is dynamically evaluated according to the specific characteristics of the input samples, and the parameters of the convolutional neural network are dynamically pruned and optimized;
[0094] S3. Repeat steps S1 to S2 until the convolutional neural network pruning is completed.
[0095] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative effort.
[0096] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A convolutional neural network pruning method based on static and dynamic combination, characterized in that: include: S1. Static pruning of convolutional neural network parameters, screening out parameters with low weights for pruning, and optimizing convolutional neural network parameters; S2. During the training of the convolutional neural network, the importance of the remaining channels is dynamically evaluated according to the specific characteristics of the input samples, and the parameters of the convolutional neural network are dynamically pruned and optimized; S3. Repeat steps S1 to S2 until the convolutional neural network pruning is completed.
2. The method according to claim 1, characterized in that In the static pruning stage in S1, parameters with low importance are screened out based on a trainable threshold T.
3. The method according to claim 1, characterized in that In the dynamic pruning stage in S2, the channel attention mechanism is optimized to dynamically evaluate the importance of the remaining channels.
4. The method according to claim 2, characterized in that: The method of screening out parameters with low importance based on a trainable threshold T comprises the following steps: S11. Use the trainable threshold T to evaluate the importance of the convolution weight parameters, and pass each convolution weight parameter through the threshold function to obtain the corresponding mask M i , the formula is: Where M∈R K*K is the mask set for each convolution kernel, T represents the threshold, σ represents the sigmoid function, M ij ∈M i , i represents the number of convolution layers, j represents the number of convolution kernels; S12, if the convolution weight parameter exceeds the threshold, the corresponding mask is set to 1, the weight is retained and participates in subsequent training, otherwise, the convolution weight is clipped; S13, the original weight parameters Perform dot multiplication with the obtained mask parameter, the formula is: in, is the weight parameter of the i-th convolutional layer, M i is the convolution kernel mask of the i-th convolution layer.
5. The method according to claim 3, characterized in that: The optimization of the channel attention mechanism and the dynamic evaluation of the importance of the remaining channels include the following steps: S21. For input feature X∈R C*H*W , extract the global feature F through global average pooling and global maximum pooling a and texture feature F m , the formula is: Among them, H is the height of the input feature, W is the width of the input feature, and F a ∈R C*1*1 is the global information of the input features, F m ∈R C*1*1 is the texture information in the input feature; S22, F a and F m One-dimensional convolution is used to realize local interaction between channels, and the Sigmoid activation function is used to calculate the attention parameter. After splicing, the attention weight F of the i-th convolution layer is generated. i : F i =σ(conv(F a ))+σ(conv(F m )) Among them, conv represents one-dimensional convolution and σ represents the Sigmoid activation function.
6. The method according to claim 4 or 5, characterized in that: In S2, the parameters of the convolutional neural network are dynamically pruned and optimized, including The attention weights F generated by dynamic pruning i The weight parameters obtained by static pruning Multiply point by point to get the weight parameter Among them, i represents the number of convolutional layers, It is the convolution weight parameter after static pruning.
7. The method according to claim 1, characterized in that In the static joint dynamic pruning process, the trainable threshold T is optimized using the loss function, as follows: Among them, L c represents the cross entropy loss, λ represents the coefficient factor, N represents the number of convolutional layers, and C represents the number of filters in the convolutional layer.
8. A convolutional neural network pruning device based on the method according to any one of claims 1 to 7, characterized in that: The device comprises: Static pruning module, which performs static pruning on the parameters of the convolutional neural network, selects parameters with low weights for pruning, and optimizes the parameters of the convolutional neural network; Dynamic pruning module: During the training process of the convolutional neural network, the importance of the remaining channels is dynamically evaluated according to the specific characteristics of the input samples, and the parameters of the convolutional neural network are dynamically pruned and optimized; The static pruning module is combined with the dynamic pruning module to jointly complete the pruning of the convolutional neural network.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Rapid pruning method, system and device for convolutional neural network, and storage medium
CN118313432A
Model pruning method and computer readable storage medium
CN118886468A
Deep neural network compression method based on joint dynamic pruning
CN112613610A
Channel attention guided convolutional neural network dynamic channel pruning method and device
CN112949840A
Convolutional neural network compression method and device combining dynamic pruning and conditional convolution
CN116306808A