Pruning method for deep convolutional neural network filter based on multi-layer channel joint measurement

By pruning deep convolutional neural network filters using multi-channel joint metric and knowledge distillation, the hardware resource bottleneck problem of edge devices is solved, enabling the deployment of efficient and lightweight models and improving the intelligent application of edge devices.

CN121480591APending Publication Date: 2026-02-06SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511559677.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Deep convolutional neural networks face hardware resource bottlenecks when deployed on edge devices, making it difficult to achieve efficient deployment while maintaining model performance.

Method used

A deep convolutional neural network filter pruning method using multi-channel joint metric is proposed. The filter importance metric is calculated through pre-training and multi-channel joint metric, the pruning rate is set by combining BN layer parameters, and the model is fine-tuned by knowledge distillation.

Benefits of technology

While maintaining model performance, the number of parameters and computational complexity are significantly reduced, resulting in a high-performance, lightweight model suitable for resource-constrained edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121480591A_ABST
    Figure CN121480591A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, and discloses a multilayer channel joint measurement deep convolutional neural network filter pruning method, and the method comprises the steps: carrying out the pre-training of a deep convolutional neural network model, and obtaining a to-be-pruned network model; performing multi-layer channel joint measurement by using the filter information of the previous layer, the current layer and the next layer, and calculating the importance measurement value of the filter of the current layer; setting a pruning rate by using BN layer parameters; performing filter pruning according to the importance measurement value and the pruning rate of the filter; according to the method, multi-layer channel joint measurement is utilized, an importance measurement value of a filter is calculated for the network model to be pruned, the network model is effectively compressed, a proper pruning rate is set for each layer by utilizing BN layer parameters, and the importance measurement value and the pruning rate of the filter are synthesized, so that the pruning efficiency of the network model to be pruned is improved. Filter pruning is achieved, model fine adjustment is conducted through knowledge distillation, and a high-performance lightweight model is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and more specifically to a pruning method for deep convolutional neural network filters with multi-channel joint metric. Background Technology

[0002] In recent years, Deep Convolutional Neural Networks (DCNNs) and their derivative technologies have continued to generate a research boom in the field of artificial intelligence, constantly pushing the performance limits in core vision tasks such as image classification, object detection, and image segmentation. With the rapid development of technologies such as the Internet of Things (IoT), autonomous driving, and smart terminals, the intelligent upgrade of edge devices has created an urgent need for the lightweight deployment of deep learning models.

[0003] However, high-performance network models often come with a large number of parameters and computational complexity, making them difficult to run efficiently on edge devices (such as embedded chips and mobile terminals) with limited computing power and storage capacity. Therefore, how to overcome hardware resource bottlenecks and achieve efficient model deployment while maintaining model performance has become a key challenge for the practical application of artificial intelligence. Summary of the Invention

[0004] This invention provides a pruning method for deep convolutional neural network filters with multi-channel joint metric, in order to solve the problem of how to achieve efficient model deployment by overcoming hardware resource bottlenecks while maintaining model performance.

[0005] In a first aspect, the present invention provides a pruning method for a deep convolutional neural network filter with a multi-channel joint metric, the method comprising: Pre-train the deep convolutional neural network model to obtain the network model to be pruned; For the filters in the network model to be pruned, multi-channel joint measurement is performed using filter information from the previous layer, the current layer, and the next layer, and the importance metric value of the current layer filter is calculated. Use BN layer parameters to set the pruning rate; Filter pruning is performed based on the filter's importance metric and pruning rate; The pruned network model is fine-tuned by incorporating knowledge distillation.

[0006] This invention utilizes multi-channel joint metric to calculate the importance metric of filters in the network model to be pruned, effectively compresses the network model, sets appropriate pruning rates for each layer using BN layer parameters, integrates the importance metric of filters and pruning rates to achieve filter pruning, and uses knowledge distillation to fine-tune the model, ultimately obtaining a high-performance lightweight model.

[0007] In one optional implementation, for the filters of the network model to be pruned, multi-channel joint measurement is performed using filter information from the previous layer, the current layer, and the next layer to calculate the importance metric value of the current layer filter, including: Calculate a metric for the information reception capability of the current layer filter itself; Calculate a metric for the receiving capability of the main information of the current layer filter; Calculate a metric for the information output capability of the current layer filter itself; A metric for calculating the ability of the current layer filter to output main information; The importance metric of the current layer filter is calculated based on the metric values ​​of its own information reception capability, its main information reception capability, its own information output capability, and its ability to output main information.

[0008] This invention is based on the information reception capability of the current layer filter itself, the information reception capability of the current layer filter, the information output capability of the current layer filter itself, and the ability of the current layer filter to output the main information, so as to realize a comprehensive evaluation of the importance of the current layer filter from both information reception and information output aspects.

[0009] In one optional implementation, the importance metric of the current layer filter is calculated based on the metric of the current layer filter's own information reception capability, the metric of the current layer filter's main information reception capability, the metric of the current layer filter's own information output capability, and the metric of the current layer filter's ability to output main information, including: For the first layer filter, the importance metric of the filter is determined by multiplying the metric of the current layer filter's own information receiving capability, the metric of the current layer filter's own information output capability, and the metric of the current layer filter's ability to output main information. For filters that are neither the first nor the last layer, the importance metric of the filter is determined by multiplying the metric of the current layer filter's own information receiving capability, the metric of the current layer filter's main information receiving capability, the metric of the current layer filter's own information output capability, and the metric of the current layer filter's main information output capability. For the last layer filter, the importance metric of the filter is determined by multiplying the information reception capability metric of the current layer filter itself by the main information reception capability metric of the current layer filter.

[0010] This invention takes into account the different information transmission tasks undertaken by filters at different positions in a deep convolutional neural network. It employs different calculation methods to comprehensively calculate the importance metric value of filters for the first layer, the last layer, and non-first and non-last layers, thereby improving the accuracy of filter importance measurement.

[0011] In one alternative implementation, the metric for the information reception capability of the current layer filter is calculated according to the following formula:

[0012] in, For the first Layer A measure of the information reception capability of each filter itself. For the first Layer Each filter No. The L1 norm of each sub-channel For the first The number of input information for each convolutional layer; The metric for the receiving capability of the current layer filter's main information is calculated using the following formula:

[0013] in, For the first Layer The receiver capability metric of the main information of each filter. For the first Layer Each filter No. The L1 norm of each sub-channel For the first The number of input information for each convolutional layer; The information output capability metric of the current layer filter is calculated using the following formula:

[0014] in, For the first Layer A measure of the information output capability of each filter itself. For the first Layer Each filter No. The L1 norm of each sub-channel For the first The number of output information of each convolutional layer; The measure of the current layer filter's ability to output main information is calculated using the following formula:

[0015] in, For the first Layer A measure of the ability of a filter to output main information. For the first Layer The first information receiving channel The L1 norm of each sub-channel For the first The number of input information for each convolutional layer.

[0016] In one alternative implementation, filter pruning is performed based on the filter's importance metric and pruning rate, including: Calculate the number of filters to be pruned in each convolutional layer based on the pruning rate; The importance metric of the filter is compared with the importance metric of the filter at a preset position in the layer to obtain the pruning mask for each channel. The pruning mask is used to characterize whether the filter needs to be pruned. Prune the filters that need to be removed.

[0017] This invention determines the number of filters to be pruned based on the pruning rate, ensuring that the pruning target is quantified and controllable, avoiding over-pruning or under-pruning. It determines whether a filter needs to be pruned based on the pruning mask, accurately selecting the filters that need to be pruned, and avoiding the accidental pruning of critical filters.

[0018] In one alternative implementation, the number of filters to be pruned from each convolutional layer is calculated according to the following formula:

[0019] in, For the first The number of filters that need to be pruned per convolutional layer For the first The number of filters in each convolutional layer For the first The pruning rate of each convolutional layer.

[0020] In one alternative implementation, the filter importance metrics for each layer are sorted from lowest to highest importance according to the following formula:

[0021] in, It is in the sorting position of Filter importance metrics at each location; Obtain the pruned mask for each filter layer. M :

[0022] in, Used to represent the Layer Does a filter need pruning? When the value is 1, the corresponding filter needs to be retained; when... When the value is 0, the corresponding filter needs to be clipped.

[0023] In a second aspect, the present invention provides a pruning device for a deep convolutional neural network filter with a multi-layer channel joint metric, the device comprising: The pre-training module is used to pre-train the deep convolutional neural network model to obtain the network model to be pruned; The calculation module is used for the filters of the network model to be pruned. It uses the filter information of the previous layer, the current layer and the next layer to perform multi-layer channel joint measurement and calculate the importance metric value of the current layer filter. The configuration module is used to set the pruning rate using BN layer parameters; The pruning module is used to prune filters based on their importance metrics and pruning rate. The fine-tuning module is used to fine-tune the pruned network model by combining knowledge distillation.

[0024] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the pruning method of the deep convolutional neural network filter with multi-channel joint metric described in the first aspect or any corresponding embodiment thereof.

[0025] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the pruning method of a deep convolutional neural network filter with multi-channel joint metric as described in the first aspect or any of its corresponding embodiments. Attached Figure Description

[0026] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0027] Figure 1This is a flowchart illustrating a pruning method for a deep convolutional neural network filter with multi-channel joint metric according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the information flow of three adjacent convolutional layers according to an embodiment of the present invention; Figure 3 This is a structural block diagram of a pruning device for a deep convolutional neural network filter with multi-channel joint metric according to an embodiment of the present invention. Figure 4 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0030] In intelligent application scenarios, deep network models are key to enhancing the intelligence of terminal devices. However, due to their massive number of parameters, these models pose deployment challenges for devices with limited resources. They not only require significant storage and computing resources but also increase energy consumption, impacting the long-term operation of the devices.

[0031] Model compression technology is a core solution that emerged to address this contradiction. This technology can significantly reduce the number of parameters and computational complexity of a model while maintaining its predictive accuracy, ultimately resulting in a lightweight model suitable for edge computing. By reducing the number and size of parameters, model compression effectively reduces computational complexity without compromising performance. This makes models easier to deploy on terminal devices, improves efficiency, and promotes the expansion of intelligent technology applications. Advances in this technology are of great significance in promoting the widespread application of artificial intelligence.

[0032] Currently, mainstream methods for model compression include parameter quantization, knowledge distillation, low-rank decomposition, and network pruning.

[0033] Network pruning, a commonly used method in model compression, removes redundant parameters from the network by constructing specific evaluation criteria, thereby reducing the model size. Based on the pruning granularity, network pruning can be broadly classified into four categories: weight pruning, convolutional kernel pruning, filter pruning, and layer pruning.

[0034] Weight pruning is called unstructured pruning, while kernel pruning, filter pruning, and layer pruning all fall under the category of structured pruning. Weight pruning first evaluates the importance of the parameters in the model, then sets the unimportant parameters to zero, performing pruning operations at a very fine level to generate a highly sparse parameter matrix.

[0035] Convolutional kernel pruning uses two-dimensional convolutional kernels as the basic unit and assigns zero values ​​to unimportant kernel parameters according to predefined evaluation criteria; however, these parameters are still retained in the model. Because layer pruning has a significant impact on network architecture, it is rarely used directly in practical applications.

[0036] Filter pruning achieves model compression by directly removing redundant filters from convolutional layers. Pruned networks are easier to optimize using existing computing architectures, making filter pruning an important direction in the field of network pruning.

[0037] To address the problem that traditional methods treat convolutional layers as independent structures, neglecting inter-layer dependencies and information transmission methods between multiple layers, this invention provides a pruning method for deep convolutional neural network filters with multi-layer channel joint metric. While maintaining or even improving model performance, this method can significantly reduce the number of parameters and floating-point operations in deep convolutional neural network models.

[0038] According to an embodiment of the present invention, a pruning method for a deep convolutional neural network filter with multi-layer channel joint metric is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0039] This embodiment provides a pruning method for deep convolutional neural network filters with multi-channel joint metric. Figure 1 This is a flowchart of a pruning method for a deep convolutional neural network filter with multi-channel joint metric according to an embodiment of the present invention, as shown below. Figure 1 As shown, the process includes the following steps: Step S101: Pre-train the deep convolutional neural network model to obtain the network model to be pruned.

[0040] In this embodiment of the invention, a deep convolutional neural network model is pre-trained to obtain a network model to be pruned. The network model includes single-branch networks such as VGGNet, multi-branch residual networks such as ResNet, and multi-branch densely connected networks such as DenseNet. These network models can be used for computer vision tasks such as image classification and object detection. Existing deep learning neural network training frameworks, such as PyTorch, are used to train the selected network model to obtain the original model, which also serves as the teacher model for knowledge distillation.

[0041] Step S102: For the filters of the network model to be pruned, perform multi-channel joint measurement using the filter information of the previous layer, the current layer, and the next layer, and calculate the importance metric value of the current layer filter.

[0042] In this embodiment of the invention, in DCNNs, the filters of the current layer are responsible for both receiving and integrating information from the previous layer and outputting information to the next layer. In fact, filters that are not in the last layer correspond to two information channels: one is the information receiving channel (i.e.,...) IC R The other is the information output channel (i.e.) IC O ).

[0043] Among them, such as Figure 2 As shown, the information receiving channel is responsible for receiving and integrating the feature information from the previous layer and generating output information. It consists of a set of sub-channels in the current layer, each implemented by a single convolutional kernel, and each sub-channel is responsible for receiving different features from the previous layer. Due to the characteristics of the information receiving channel in generating output information, the information receiving channel of the previous layer filter will affect the information receiving of the current layer filter.

[0044] The information output channel is responsible for transmitting the output information of the current layer filter to the next layer and distributing it to the corresponding information receiving channel. IC O It consists of multiple sub-channels in the next layer, each responsible for transmitting the output information of the same filter. Because the next layer... IC R Responsible for receiving IC O The information distributed will affect the information output of the filter in the next layer's information receiving channel.

[0045] A multi-layer channel joint metric model is established using the information receiving and output channels of the upper, current, and lower layers to calculate the importance metric value of the current layer filter.

[0046] Step S103: Set the pruning rate using BN layer parameters.

[0047] In this embodiment of the invention, since the sensitivity of each convolutional layer in DCNNs to pruning is different, it is necessary to set an appropriate pruning rate for each convolutional layer. Several methods for setting the pruning rate already exist. These methods utilize the scaling factor and offset factor of the Batch Normalization (BN) layer to construct a layer pruning sensitivity evaluation, and finally, by setting an overall pruning rate, the pruning rate of each convolutional layer is allocated to confirm the pruning rate of each layer.

[0048] Step S104: Prune the filter according to its importance metric and pruning rate.

[0049] In this embodiment of the invention, after obtaining the filter importance metric value of each convolutional layer and the pruning rate of each layer, the number of filters that need to be pruned in each convolutional layer is calculated, and the filters of each convolutional layer are pruned.

[0050] Step S105: Fine-tune the pruned network model using knowledge distillation.

[0051] In this embodiment of the invention, during the fine-tuning stage of the network model after pruning, a strategy combining knowledge distillation and fine-tuning is adopted to organically combine the advantages of the two compression methods, ultimately resulting in better model performance and compression effect.

[0052] Knowledge distillation, a model compression technique in deep learning, aims to transfer knowledge from large, complex models (i.e., teacher models) to smaller, lightweight models (i.e., student models). The key to knowledge distillation is minimizing the difference in probability distributions between the outputs of the teacher and student models.

[0053] By setting the original model as the teacher model and the pruned model as the student model, and by fine-tuning the pruned network model, a lightweight network with excellent performance can be obtained.

[0054] The deep convolutional neural network filter pruning method with multi-channel joint metric provided in this embodiment uses multi-channel joint metric to calculate the importance metric value of the filter in the network model to be pruned, effectively compresses the network model, sets appropriate pruning rates for each layer using BN layer parameters, integrates the importance metric value of the filter and the pruning rate to achieve filter pruning, and uses knowledge distillation to fine-tune the model, finally obtaining a high-performance lightweight model.

[0055] This embodiment provides a pruning method for deep convolutional neural network filters with multi-layer channel joint metric, the process of which includes the following steps: Step S201: Pre-train the deep convolutional neural network model to obtain the network model to be pruned.

[0056] Please see details Figure 1 Step S101 of the illustrated embodiment will not be described again here.

[0057] Step S202: For the filters of the network model to be pruned, perform multi-channel joint measurement using the filter information of the previous layer, the current layer, and the next layer, and calculate the importance metric value of the current layer filter.

[0058] Specifically, step S202 includes: Step S2021: Calculate the metric value of the information reception capability of the current layer filter itself.

[0059] Step S2022: Calculate the metric value of the receiving capability of the main information of the current layer filter.

[0060] Step S2023: Calculate the metric value of the information output capability of the current layer filter itself.

[0061] Step S2024: Calculate a metric for the ability of the current layer filter to output main information.

[0062] Step S2025: Calculate the importance metric of the current layer filter based on the metric of the current layer filter's own information receiving capability, the metric of the current layer filter's main information receiving capability, the metric of the current layer filter's own information output capability, and the metric of the current layer filter's ability to output main information.

[0063] In embodiments of the present invention, such as Figure 2 As shown, the information reception capability of the current layer filter is not only related to the filter's... IC R The intensity of the feature map is related to the intensity of the feature map, and also depends on the intensity of the filter in the previous layer. IC R The current layer filter's IC R Composed of multiple sub-channels, the information reception capability of the current layer filter itself is the accumulation of the feature mapping intensities of these sub-channels, reflecting the filter's maximum potential for information reception and feature generation. In the current layer... IC R In this process, the sub-channel with the strongest feature mapping strength is the main channel for receiving information from the previous layer, and its information reception capability depends on the previous layer. IC R The intensity of the feature mapping.

[0064] like Figure 2 As shown, the information output capability of the current layer filter is not only related to the current layer filter's... IC O The feature mapping capability is related to the feature mapping capability, and also depends on the next layer filter. IC RThe current layer filter's IC O Composed of multiple sub-channels in the next layer, the current layer filter's output capability is the accumulation of the feature mapping intensities of these sub-channels, reflecting the filter's maximum potential to output information to the next layer. The current layer filter's... IC O The feature information of the filter is mapped and distributed to each element in the next layer. IC R In the current layer filter. IC O In this model, the sub-channel with the strongest feature mapping strength is the main channel for transmitting information to the lower layer, and the information output capability of this sub-channel depends on the next layer. IC R The feature mapping intensity and the sub-channel in the next layer IC R The degree of importance.

[0065] Taking into account the previous layer, the current layer, and the next layer. IC R and IC O This approach overcomes the limitations of traditional single-layer evaluation systems, enabling a more reasonable evaluation of filters. By using multi-channel joint metrics, the role of filters in information transmission can be evaluated more accurately, ensuring that important filters not only generate crucial feature information but also effectively transmit that information to lower layers.

[0066] First, construct a metric for the information reception capability of the current layer filter itself.

[0067] The information reception capability metric of the current layer filter itself is... IC R Perform measurement. IC R The feature mapping intensity of a channel is the sum of the feature mapping intensities of all its sub-channels. The larger the norm of a sub-channel, the more significant the extracted features, and the higher the feature mapping intensity of that sub-channel. In this invention, the L1 norm is used to measure the feature mapping intensity of a sub-channel. Therefore, the feature mapping intensity of the first channel... l Layer n The information reception capability of each filter is measured as follows: (1) in, For the first Layer A measure of the information reception capability of each filter itself. For the first Layer Each filter No. The L1 norm of each sub-channel For the first The number of input information for each convolutional layer This reflects the maximum potential of the filter's information reception capability.

[0068] Secondly, a metric is constructed to measure the receiving capability of the main information of the current layer filter.

[0069] The main information reception capability metric of the current layer filter is... IC R The strongest sub-channel is used for measurement. The information reception capability of this sub-channel depends on the corresponding upper-layer filter. IC R .Should IC R The stronger the feature mapping strength, the more information the previous layer filter outputs to the strongest sub-channel, and the stronger the information reception capability of the strongest sub-channel. Conversely, the stronger the feature mapping strength, the more information the previous layer filter outputs to the strongest sub-channel. IC R The weaker the feature mapping strength, the less information the previous layer filter outputs to the strongest sub-channel, and the weaker the information reception capability of the strongest sub-channel. Therefore, the... l Layer n The receiver capability metrics for the main information of each filter are as follows: (2) in, It is the first l Layer n Each filter IC R The strongest subchannel index: (3) in, For the first Layer The receiver capability metric of the main information of each filter. For the first Layer Each filter No. The L1 norm of each sub-channel For the first The number of input information for each convolutional layer.

[0070] Then, a metric is constructed to measure the information output capability of the current layer filter itself.

[0071] The information output capability metric of the current layer filter itself is... IC O The intensity of the feature mapping is measured. ICO The feature mapping intensity of a channel is the accumulation of the feature mapping intensities of all its sub-channels. The larger the norm of a sub-channel, the more significant the extracted features, and the higher the feature mapping intensity of that sub-channel. In this embodiment of the invention, the L1 norm is used to measure the feature mapping intensity of a sub-channel. Therefore, the feature mapping intensity of the first channel... l Layer n The information output capability of each filter is measured as follows: (4) in, For the first Layer A measure of the information output capability of each filter itself. For the first Layer Each filter No. The L1 norm of each sub-channel For the first The number of output information of each convolutional layer This reflects the maximum potential of the filter's information output capability.

[0072] Finally, a measure of the ability to construct the main information output by the current layer filter.

[0073] The ability of the current layer filter to output main information is a metric for the current layer filter. IC O The strongest sub-channel is used for measurement. The information output capability of this sub-channel depends on the corresponding next-layer filter. IC O .Should IC O The stronger the feature mapping strength, the better the information features generated by the next layer filter, and the better the information of that sub-channel is passed on in subsequent layers. When the strongest sub-channel is its... IC R If a sub-channel is considered important, then the information transmitted in that strongest sub-channel is important. Therefore, when that sub-channel... IC O The feature mapping strength is the strongest, and this strongest sub-channel is also the strongest. IC O When the strongest sub-channel is reached, the current layer filter has the strongest ability to output main information. Therefore, the first... l Layer n The ability of each filter to output main information is quantified as follows: (5) in, It is the first l Layer n Each filterIC O The strongest sub-channel: (6) yes Index: (7) yes l +1 floor Each filter IC R The strongest sub-channel: (8) in, For the first Layer A measure of the ability of a filter to output main information. For the first Layer The first information receiving channel The L1 norm of each sub-channel For the first The number of input information for each convolutional layer.

[0074] Based on the information reception capability, main information reception capability, information output capability, and main information output capability of the current layer filter, the importance metric of the current layer filter is comprehensively evaluated from both information reception and information output perspectives.

[0075] Specifically, step S2025 above includes: Step S20251: For the first layer filter, the product of the information receiving capability metric of the current layer filter itself, the information output capability metric of the current layer filter itself, and the main information output capability metric of the current layer filter is determined as the importance metric of the filter.

[0076] Step S20252: For filters that are neither the first nor the last layer, the product of the information reception capability metric of the current layer filter, the main information reception capability metric of the current layer filter, the information output capability metric of the current layer filter, and the main information output capability metric of the current layer filter is determined as the importance metric of the filter.

[0077] Step S20253: For the last layer filter, the product of the information reception capability metric of the current layer filter itself and the main information reception capability metric of the current layer filter is determined as the importance metric of the filter.

[0078] In this embodiment of the invention, the importance of a filter is defined as a comprehensive reflection of its role in receiving information from the preceding layer and transmitting information to subsequent layers. For the filter in the first layer, since its input information is the input image, and all input information is equally important, it does not have a metric for its ability to receive main information. Therefore, the importance metric for the filter in the first layer is: (9) For intermediate layer filters that are neither the first nor the last layer, there are measures of the current layer filter's own information receiving capability, its main information receiving capability, its own information output capability, and its ability to output main information. Therefore, when 1 < l < L , No. l Layer n The overall importance score of the filters is : (10) For the filter in layer L, there is no metric for its ability to provide output information. Therefore, the importance metric for the filter in the last layer is: (11) The current layer filter's own information receiving capability is its maximum potential ability to receive information from the previous layer, considering only the filter's information receiving channel. The current layer filter's main information receiving capability is the information receiving capability of its largest information receiving sub-channel, considering the sub-channel's actual ability to receive information from the previous layer filter. The current layer filter's own information output capability is its maximum potential ability to output information to the next layer, considering only the filter's information output channel. The current layer filter's main information output capability is the information output capability of its largest information output sub-channel, considering the sub-channel's actual ability to output information to the previous layer filter. Multiplying these four factors reflects the filter's importance in the information transmission process. Information generated by the previous layer is received by the current layer filter and integrated into new characteristic information, ultimately achieving effective information output to the next layer.

[0079] Only when a filter has both high information reception and information output capabilities can it truly play a key role in the information transmission of the model. This comprehensive evaluation method can better measure filters that have strong information reception capabilities but poor information output capabilities, or good information output capabilities but weak information reception capabilities. Thus, while ensuring model performance, it can more effectively simplify the model structure and improve the model's operating efficiency.

[0080] Considering that the information transmission tasks undertaken by filters at different positions in a deep convolutional neural network are different, different calculation methods are used to comprehensively calculate the importance metric of filters for the first layer, the last layer, and non-first and non-last layers, thereby improving the accuracy of filter importance measurement.

[0081] Step S203: Set the pruning rate using BN layer parameters.

[0082] In this embodiment of the invention, the pruning sensitivity score of each set of BN parameters corresponding to each convolutional layer is calculated, the pruning sensitivity score is subjected to maximum-minimum normalization, the global pruning sensitivity scores are sorted by size, a pruning sensitivity flag is set for each layer, and the pruning rate of each layer is calculated based on the pruning sensitivity flag of each layer.

[0083] Step S204: Prune the filter according to the importance metric and pruning rate.

[0084] Specifically, step S204 includes: Step S2041: Calculate the number of filters to be pruned in each convolutional layer based on the pruning rate.

[0085] Step S2042: Compare the importance metric of the filter with the importance metric of the filter at the preset position of the layer to obtain the pruning mask of each layer channel. The pruning mask is used to characterize whether the filter needs to be pruned.

[0086] Step S2043: Prune the filters that need to be removed.

[0087] In this embodiment of the invention, the pruning rate of each layer is obtained. The number of filters that need to be pruned in each layer can be determined based on the pruning rate. (12) in, For the first The number of filters that need to be pruned per convolutional layer For the first The number of filters in each convolutional layer For the first The pruning rate of each convolutional layer.

[0088] After obtaining the number of pruning operations, the importance metric of the filter in each layer is compared with the number of its corresponding layers. p By comparing the importance metrics of large filters, the filters that need to be retained in each layer can be determined.

[0089] Specifically, the measure of filter importance for each convolutional layer can be obtained according to the above formulas (9), (10), and (11). Based on the metric values, the importance metrics of the filters in each layer are sorted from lowest to highest importance: (13) in, It is the first in the above sorting p Filter importance metrics at each location: (14) Finally, the pruned mask for each filter layer is obtained. M : (15) in, Used to represent the Layer Does a filter need pruning? When the value is 1, the corresponding filter needs to be retained; when... When the value is 0, the corresponding filter needs to be clipped.

[0090] By determining the number of filters to be pruned based on the pruning rate, the pruning target can be quantified and controlled, avoiding over-pruning or under-pruning. The pruning mask determines whether a filter needs to be pruned, accurately selecting the filters to be pruned and avoiding the accidental pruning of critical filters.

[0091] Step S205: Fine-tune the pruned network model using knowledge distillation.

[0092] Please see details Figure 1 Step S105 of the illustrated embodiment will not be described again here.

[0093] The multi-channel joint metric deep convolutional neural network filter pruning method provided in this embodiment effectively compresses and accelerates the network model, significantly reducing the number of parameters and floating-point operations while maintaining almost no loss in accuracy. Furthermore, this method is applicable to various common convolutional neural networks, such as VGG16, ResNet56, and DenseNet40, and is also suitable for datasets of different sizes, such as the small and simple CIFAR-10 dataset, the small and complex CIFAR-100 dataset, and the large and complex ImageNet dataset.

[0094] This embodiment also provides a pruning device for a deep convolutional neural network filter with multi-layer channel joint metric. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0095] This embodiment provides a pruning device for a deep convolutional neural network filter with multi-channel joint metric, such as... Figure 3 As shown, it includes: The pre-training module 301 is used to pre-train the deep convolutional neural network model to obtain the network model to be pruned.

[0096] The calculation module 302 is used for the filters of the network model to be pruned. It uses the filter information of the previous layer, the current layer and the next layer to perform multi-layer channel joint measurement and calculates the importance metric value of the current layer filter.

[0097] Module 303 is used to set the pruning rate using BN layer parameters.

[0098] The pruning module 304 is used to prune the filter based on the filter's importance metric and pruning rate.

[0099] The fine-tuning module 305 is used to fine-tune the pruned network model by combining knowledge distillation.

[0100] In some alternative implementations, the computing module 302 includes: The first calculation unit is used to calculate a metric of the information reception capability of the current layer filter itself.

[0101] The second calculation unit is used to calculate a metric value for the receiving capability of the main information of the current layer filter.

[0102] The third calculation unit is used to calculate a metric of the information output capability of the current layer filter itself.

[0103] The fourth calculation unit is used to calculate a metric of the current layer filter's ability to output main information.

[0104] The fifth calculation unit is used to calculate the importance metric of the current layer filter based on the metric of the current layer filter's own information receiving capability, the metric of the current layer filter's main information receiving capability, the metric of the current layer filter's own information output capability, and the metric of the current layer filter's ability to output main information.

[0105] In some alternative implementations, the fifth computing unit includes: The first determining subunit is used to determine the importance metric of the filter by multiplying the metric of the information receiving capability of the current layer filter itself, the metric of the information output capability of the current layer filter itself, and the metric of the capability of the current layer filter to output main information.

[0106] The second determining subunit is used to determine the importance metric of a filter for a filter that is neither the first nor the last layer by multiplying the metric of the filter's own information receiving capability, the metric of the filter's main information receiving capability, the metric of the filter's own information output capability, and the metric of the filter's ability to output main information.

[0107] The third determining subunit is used to determine the importance metric of the filter by multiplying the metric of the current layer filter's own information receiving capability and the metric of the current layer filter's main information receiving capability for the last layer filter.

[0108] In some alternative implementations, the pruning module 304 includes: The sixth calculation unit is used to calculate the number of filters to be pruned from each convolutional layer based on the pruning rate.

[0109] The comparison unit is used to compare the importance metric of the filter with the importance metric of the filter at a preset position in the layer to obtain the pruning mask for each channel. The pruning mask is used to characterize whether the filter needs to be pruned.

[0110] The pruning unit is used to prune filters that need to be removed.

[0111] The pruning device for deep convolutional neural network filters with multi-channel joint metric provided in this embodiment of the invention can execute the pruning method for deep convolutional neural network filters with multi-channel joint metric provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method. Further functional descriptions of the above modules and units are the same as in the corresponding embodiments described above, and will not be repeated here.

[0112] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0113] The following is a detailed reference. Figure 4 This diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from memory 408 into random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the electronic device. The processor 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0114] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0115] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 409, or installed from a memory 408, or installed from a ROM 402. When the computer program is executed by the processor 401, it performs the functions defined in the pruning method of the deep convolutional neural network filter with multi-layer channel joint metric according to embodiments of the present invention.

[0116] Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0117] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the pruning method for deep convolutional neural network filters with multi-layer channel joint metric shown in the above embodiments is implemented.

[0118] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0119] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the present invention.

Claims

1. A pruning method for deep convolutional neural network filters with multi-channel joint metric, characterized in that, The method includes: Pre-train the deep convolutional neural network model to obtain the network model to be pruned; For the filters in the network model to be pruned, multi-channel joint measurement is performed using filter information from the previous layer, the current layer, and the next layer, and the importance metric value of the current layer filter is calculated. Use BN layer parameters to set the pruning rate; Filter pruning is performed based on the importance metric of the filter and the pruning rate; The pruned network model is fine-tuned by incorporating knowledge distillation.

2. The method according to claim 1, characterized in that, The filters in the network model to be pruned are used to perform multi-channel joint measurement using filter information from the previous, current, and next layers, and the importance metric value of the current layer filter is calculated, including: Calculate a metric for the information reception capability of the current layer filter itself; Calculate a metric for the receiving capability of the main information of the current layer filter; Calculate a metric for the information output capability of the current layer filter itself; A metric for calculating the ability of the current layer filter to output main information; The importance metric of the current layer filter is calculated based on the metric values ​​of its own information reception capability, its main information reception capability, its own information output capability, and its main information output capability.

3. The method according to claim 2, characterized in that, The step of calculating the importance metric of the current layer filter based on the metric of the current layer filter's own information reception capability, the metric of the current layer filter's main information reception capability, the metric of the current layer filter's own information output capability, and the metric of the current layer filter's ability to output main information includes: For the first layer filter, the product of the information receiving capability metric of the current layer filter itself, the information output capability metric of the current layer filter itself, and the capability of the current layer filter to output main information is determined as the importance metric of the filter. For filters that are neither the first nor the last layer, the importance metric of the filter is determined by multiplying the information reception capability of the current layer filter itself, the main information reception capability of the current layer filter, the information output capability of the current layer filter itself, and the main information output capability of the current layer filter. For the last layer filter, the product of the information reception capability metric of the current layer filter itself and the main information reception capability metric of the current layer filter is determined as the importance metric of the filter.

4. The method according to claim 2, characterized in that, The information reception capability metric of the current layer filter is calculated using the following formula: in, For the first Layer A measure of the information reception capability of each filter itself. For the first Layer Each filter No. The L1 norm of each sub-channel For the first The number of input information for each convolutional layer; The metric for the receiving capability of the current layer filter's main information is calculated using the following formula: in, For the first Layer The receiver capability metric of the main information of each filter. For the first Layer Each filter No. The L1 norm of each sub-channel For the first The number of input information for each convolutional layer; The information output capability metric of the current layer filter is calculated using the following formula: in, For the first Layer A measure of the information output capability of each filter itself. For the first Layer Each filter No. The L1 norm of each sub-channel For the first The number of output information of each convolutional layer; The measure of the current layer filter's ability to output main information is calculated using the following formula: in, For the first Layer A measure of the ability of a filter to output main information. For the first Layer The first information receiving channel The L1 norm of each sub-channel For the first The number of input information for each convolutional layer.

5. The method according to claim 1, characterized in that, The step of pruning the filter based on the importance metric and the pruning rate includes: Based on the pruning rate, calculate the number of filters that need to be pruned in each convolutional layer; The importance metric of the filter is compared with the importance metric of the filter at a preset position in the layer to obtain the pruning mask for each channel. The pruning mask is used to characterize whether the filter needs to be pruned. Prune the filters that need to be removed.

6. The method according to claim 5, characterized in that, Calculate the number of filters to be pruned from each convolutional layer using the following formula: in, For the first The number of filters that need to be pruned per convolutional layer For the first The number of filters in each convolutional layer For the first The pruning rate of each convolutional layer.

7. The method according to claim 5, characterized in that, The importance metrics of the filters in each layer are sorted from lowest to highest according to the following formula: in, It is in the order of Filter importance metrics at each location; Obtain the pruned mask for each filter layer. M : in, Used to represent the Layer Does a filter need pruning? When the value is 1, the corresponding filter needs to be retained; when... When the value is 0, the corresponding filter needs to be clipped.

8. A pruning device for a deep convolutional neural network filter with multi-channel joint metric, characterized in that, The device includes: The pre-training module is used to pre-train the deep convolutional neural network model to obtain the network model to be pruned; The calculation module is used for the filters of the network model to be pruned. It uses the filter information of the previous layer, the current layer and the next layer to perform multi-layer channel joint measurement and calculate the importance metric value of the current layer filter. The configuration module is used to set the pruning rate using BN layer parameters; A pruning module is used to prune the filter based on the importance metric of the filter and the pruning rate; The fine-tuning module is used to fine-tune the pruned network model by combining knowledge distillation.

9. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the pruning method of the deep convolutional neural network filter with multi-layer channel joint metric as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform a pruning method for a deep convolutional neural network filter with a multi-layer channel joint metric as described in any one of claims 1 to 7.