Convolutional neural network filter pruning method based on mixed importance score guidance
Through the convolutional neural network filter pruning method based on mixed importance score guidance, the problems of inaccurate evaluation and difficulty in reducing computational complexity in the prior art are solved, and more accurate pruning and network performance improvement are achieved, which is suitable for resource-constrained devices.
Patent Information
- Application Number
- CN202510316366.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
AI Technical Summary
The existing convolutional neural network filter pruning method is inaccurate in evaluation, making it difficult to effectively reduce the computational complexity or meet the memory requirements of specific hardware.
Using a method based on mixed importance score guidance, the filter activation value and weight information are extracted by training the benchmark model, the mixed importance score of each filter is calculated, and the pruning threshold is dynamically set according to the requirements of computing resources and storage resources, filters with importance scores below the threshold are filtered and pruned, and finally the pruning model is fine-tuned and trained to restore model accuracy.
It significantly improves the accuracy of pruning and network performance, effectively reduces the overall computing complexity and storage requirements of the model, and improves the inference speed and resource utilization of the model, which is especially suitable for deployment on resource-constrained devices.
Smart Images

Figure CN120146133A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and artificial intelligence, and specifically to a convolutional neural network filter pruning method guided by hybrid importance scoring. Background Art
[0002] In recent years, convolutional neural networks (CNNs) have achieved remarkable results in fields such as image recognition, object detection, and semantic segmentation. However, they also bring huge computational costs and storage requirements. This poses a severe challenge when deploying CNNs on resource-constrained embedded or edge devices. To address this issue, existing technologies have proposed model compression methods such as weight pruning and filter pruning. Among them, filter pruning has received extensive attention because it can directly reduce the network layer structure, lower the model computational complexity (FLOPs), and storage requirements.
[0003] However, most current filter pruning methods only perform importance evaluation based on the activation values or weight information of filters, resulting in inaccurate evaluation and being unable to fully reflect the true contribution of filters to the model prediction performance. In addition, existing pruning strategies often ignore the non-linear relationship between computational cost and the number of parameters. As a result, in practical applications, pruning a large number of filters may not necessarily effectively reduce the computational complexity or meet the memory requirements of specific hardware. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a convolutional neural network filter pruning method guided by hybrid importance scoring, which solves the problems of inaccurate evaluation of current filter pruning methods and difficulty in effectively reducing computational complexity or meeting the memory requirements of specific hardware.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A convolutional neural network filter pruning method guided by hybrid importance scoring, comprising the following steps:
[0006] Train a baseline model and extract the filter activation values and weight information of the target convolutional layer;
[0007] Based on the activation values and weight information, calculate the hybrid importance score of each filter;
[0008] According to the computational resources and storage resource requirements of each layer, dynamically set the pruning threshold, and filter and prune the filters with importance scores lower than the threshold;
[0009] Fine-tune the pruned model to restore the model accuracy.
[0010] Preferably, the baseline model is any one of the VGG, ResNet, or DenseNet architectures.
[0011] Preferably, the reference model adopts a convolutional neural network including residual modules, and synchronously extracts the activation values and weight information of the filters during the training process. The activation value is the output matrix of the input image passing through the filter, and the weight information is the result of the eigenvalue decomposition of the convolution kernel in the complex domain.
[0012] Preferably, the calculation of the hybrid importance score includes:
[0013] Score based on activation value: By counting the rank of the activation matrix output by the filter, calculate the normalized activation value score;
[0014] Score based on weight: By decomposing the convolution kernel matrix, analyze the scaling factor and rotation angle of the eigenvalues, and calculate the normalized weight score;
[0015] Take the minimum value of the activation value score and the weight score as the final hybrid importance score.
[0016] Preferably, the calculation of the activation value score specifically includes:
[0017] Input multiple images into the network, and extract the activation matrix output by the target filter in the convolutional layer;
[0018] Calculate the rank of each activation matrix, where the rank represents the number of linearly independent features in the activation matrix;
[0019] Normalize the rank and calculate the average value of multiple images to obtain the activation value score.
[0020] Preferably, the calculation of the weight score specifically includes:
[0021] Decompose the convolution kernel matrix of the target filter to extract the complex domain eigenvalues;
[0022] Calculate the scaling factor and rotation angle according to the eigenvalues. The scaling factor represents the modulus length of the eigenvalue, and the rotation angle represents the direction of the eigenvalue in the complex plane;
[0023] Normalize the scaling factor and rotation angle respectively and aggregate them into the weight score.
[0024] Preferably, the step of dynamically setting the pruning threshold according to the calculation resources and storage resource requirements of each layer includes:
[0025] Calculate the parameter ratio P of each layer i and the floating-point operation amount ratio F i , where the parameter ratio P i represents the ratio of the number of parameters in the i-th layer to the total number of parameters in the model, and the floating-point operation amount ratio F i represents the ratio of the floating-point operation amount in the i-th layer to the total floating-point operation amount in the model;
[0026] According to P i and F i Calculate the resource ratio index and select a pruning strategy based on the value of φ i :
[0027] When φ i ≥ 1, adopt a computation-intensive pruning strategy;
[0028] When φ i < 1, adopt a storage-optimized pruning strategy;
[0029] Dynamically adjust the pruning threshold according to the selected pruning strategy.
[0030] Preferably, the step of dynamically adjusting the pruning threshold according to the selected pruning strategy includes:
[0031] Determine the quantile value Q corresponding to the pruning ratio ε according to the sorting of the mixed importance scores of the current layer filters i ;
[0032] Based on the resource ratio index φ i , combined with the pruning ratio ε and the adjustment parameter m, dynamically calculate the pruning threshold T i :
[0033] When φ i ≥ 1: T i = Q i + m·Q i ·ln(1 + ∈), increase the threshold to increase the number of pruned filters;
[0034] When φ i < 1: T i = Q i - m·Q i ·ln(1 + ∈), decrease the threshold to reduce the number of pruned filters.
[0035] Preferably, in the step of screening and pruning the filters with importance scores lower than the threshold, the operation of pruning the filters includes:
[0036] Retain the first and last convolutional layers and the corresponding batch normalization layers of the residual module to ensure the coherence of the network structure.
[0037] Preferably, the fine-tuning training adopts the same optimizer and learning rate strategy as the baseline model.
[0038] The present invention provides a convolutional neural network filter pruning method guided by mixed importance scores.
[0039] It has the following beneficial effects:
[0040] 1. The present invention introduces a resource ratio index and dynamically selects a suitable pruning strategy according to the proportion of computing resources and storage resources in each layer. This adaptive pruning method can more accurately identify and prune redundant filters, avoid mis-pruning of key feature extraction layers, and significantly improve the pruning accuracy and network performance.
[0041] 2. The present invention optimizes and prunes the computing resources or storage resources respectively according to the resource bottlenecks of different layers. For computationally intensive layers, FLOPs are preferentially reduced; for storage-intensive layers, the number of parameters is preferentially reduced. This balanced optimization strategy can effectively reduce the overall computational complexity and storage requirements of the model, significantly improve the inference speed and resource utilization rate of the model, and is especially suitable for deployment on resource-constrained devices.
[0042] 3. By comprehensively evaluating the importance of filters based on the Hybrid Importance Score (HIS), the present invention ensures that only redundant filters with little impact on the model accuracy are pruned. In multiple experiments, even under high pruning ratios, the model can still maintain a high accuracy rate and is superior to various existing pruning methods, effectively solving the problem of "pruning - accuracy loss" in traditional pruning methods.
[0043] 4. The HP2 pruning method of the present invention does not rely on a specific network structure and is applicable to various classic convolutional neural networks such as VGG, ResNet, and DenseNet. At the same time, excellent pruning effects have been achieved on various datasets such as CIFAR-10, CIFAR-100, and ImageNet. This method has good versatility and scalability and is convenient for wide application in different deep learning scenarios.
[0044] 5. Traditional pruning methods usually rely on manually setting thresholds or adjusting through repeated experiments, while the present invention can automatically determine the optimal pruning strategy according to the characteristics of network layers by adaptively adjusting the pruning threshold. This automated pruning process not only improves the pruning efficiency, reduces the manual intervention cost, but also enhances the consistency and reliability of pruning, facilitating large-scale application and rapid deployment. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a schematic flow chart of the method of the present invention;
[0046] Figure 2 is a schematic diagram of a misjudgment situation that may be caused by relying solely on activation value evaluation;
[0047] Figure 3 is a schematic diagram of the decomposition process of the convolutional kernel matrix in the filter of the present invention;
[0048] Figure 4 is a schematic diagram of the distribution of FLOPs and parameters in different layers;
[0049] Figure 5 Schematic diagram of the pruning strategy reserved for the key layer in the residual module of the present invention;
[0050] Figure 6 Schematic diagram of the characteristic that HIS has stability in the experimental example of the present invention;
[0051] Figure 7 Schematic diagram of the accuracy decline under different pruning ratios in the experimental example of the present invention. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] Please refer to the attached Figure 1 - attached Figure 5 , the present invention proposes a convolutional neural network (CNN) filter pruning method guided by Hybrid Importance Score (HIS). Its core principle is to dynamically evaluate the importance of each filter by fusing the activation value features and weight features of the filter, and adaptively adjust the pruning threshold in combination with the differentiated requirements of computing and storage resources to achieve efficient and accurate model compression.
[0054] The specific technical principle is as follows:
[0055] 1. Hybrid Importance Score (HIS):
[0056] Activation value scoring: By analyzing the rank of the activation matrix output by the filter, measure its ability to express input features. The higher the rank, the more diverse the features extracted by the filter and the higher the importance.
[0057] Weight scoring: By decomposing the complex domain eigenvalues of the convolution kernel matrix, calculate its Scaling Factor and Rotation Angle to quantify the ability of the filter to transform spatial features.
[0058] Scoring fusion: Take the minimum value of the activation value scoring and the weight scoring to avoid the deviation of a single index and ensure the robustness of the importance evaluation.
[0059] Such as Figure 2As shown, it demonstrates the bias problem that may occur when scoring filters solely relying on activation values or weights. In this example, K1, K2, and K3 represent convolutional kernels respectively. If only based on activation values, filters with relatively weak feature extraction capabilities may be misjudged as important filters, and vice versa. Especially in residual modules, the performance of filters directly affects the accuracy of feature maps, and single scoring may mislead the pruning strategy. Therefore, the present invention adopts the fusion of activation value scoring and weight scoring to ensure comprehensive and accurate identification of redundant filters.
[0060] 2. Adaptive pruning threshold:
[0061] Calculate the resource ratio index according to the parameter ratio and floating-point operation amount ratio of each layer, and dynamically adjust the pruning threshold:
[0062] Resource ratio index ≥ 1: For compute-intensive layers, increase the pruning threshold to preferentially reduce compute resource consumption (FLOPs).
[0063] Resource ratio index < 1: For storage-optimized layers, lower the pruning threshold to preferentially reduce the number of parameters (Storage).
[0064] 3. Residual structure protection:
[0065] Retain the first and last convolutional layers and their batch normalization (BatchNorm) layers for residual modules to ensure the coherence of the gradient propagation path and avoid network structure breakage caused by pruning.
[0066] The method of the present invention will be elaborated in detail below in combination with specific steps.
[0067] As Figure 1 shown, the convolutional neural network filter pruning method guided by hybrid importance scoring may include the following steps:
[0068] S1. Train a baseline model and extract the filter activation values and weight information of the target convolutional layer;
[0069] S2. Calculate the hybrid importance score of each filter based on the activation value and weight information;
[0070] S3. Dynamically set the pruning threshold according to the compute resources and storage resource requirements of each layer, and filter out and prune the filters with importance scores lower than the threshold;
[0071] S4. Fine-tune and train the pruned model to restore the model accuracy.
[0072] For step S1, the process of training the baseline model includes constructing the network architecture, data preprocessing, training strategy setting, and extracting the filter activation values and weight information of the target convolutional layer.
[0073] First, establish a standard deep learning model architecture. The benchmark model can adopt typical convolutional neural network models such as VGG, ResNet, or DenseNet, where ResNet contains residual modules and DenseNet has cross-layer connection characteristics. The model structure needs to support the extraction of activation values and weight information of convolutional layers to meet the requirements of subsequent pruning steps. The training of the benchmark model can be carried out based on public datasets, including but not limited to CIFAR-10, CIFAR-100, or ImageNet.
[0074] During the training process, the input data needs to be normalized so that the data mean is close to zero and the variance is normalized to optimize the stability of model training. The normalization method uses the mean-standard deviation normalization method, which can improve the stability of data during training and ensure the consistency of data distribution in different batches.
[0075] To enhance the generalization ability of the model, data augmentation strategies can be adopted for training data, including but not limited to random cropping, horizontal flipping, random rotation, etc. Random cropping can extract sub-images from the original image based on a fixed-size window to ensure the consistency of data distribution. The probability of horizontal flipping can be set to a predetermined value to enhance the model's adaptability to inputs in different directions.
[0076] The training process adopts a gradient-optimization-based method. The optimizer is preferably the momentum optimization method, combined with a learning rate decay strategy to improve training stability. The learning rate decay can adopt the cosine annealing strategy, so that the initial learning rate is gradually adjusted during training to reduce convergence oscillation. This strategy can optimize the parameter adjustment of the model in different training stages and improve the final training effect.
[0077] After training is completed, extract the filter activation values and weight information of all target convolutional layers and store them for subsequent pruning analysis.
[0078] For step S2, calculate the mixed importance score of each filter based on the activation values and weight information. This score is used to quantify the importance of the filter in a specific convolutional layer and provide a basis for subsequent pruning decisions. The calculation process includes the score calculation based on activation values, the score calculation based on weights, and the comprehensive calculation of both to ensure more accurate evaluation results.
[0079] Calculate the score based on activation values. Select K randomly sampled images as input data, pass them through the trained convolutional neural network, and obtain the output activation matrix of the filter in the τ-th layer. Let the output activation matrix of the filter for the j-th image be Calculate the representational ability of the filter for the input data through matrix rank calculation. The activation value score calculation is as follows:
[0080]
[0081] Among them, f(·) is a normalization function that maps the matrix rank value to the standardized interval [0, 1] to ensure the comparability of scores between different filters. The level of the matrix rank reflects the complexity of the features extracted by the filter. The higher the rank value, the greater the contribution of the filter to feature extraction.
[0082] As Figure 3 shown, the score is calculated based on the weight information. The filter consists of multiple convolutional kernels, each of size s×s, and its weight matrix can be expressed as:
[0083]
[0084] Among them, C in is the number of input channels of the filter. Perform eigenvalue decomposition on , and let its eigenvalues be:
[0085] λ m = a m + b m i, m = 1, 2, …, k
[0086] Among them, a m and b m are the real part and the imaginary part of the eigenvalue respectively, and i is the imaginary unit. Calculate the modulus length (scaling factor) and rotation angle of each eigenvalue to measure the weight distribution of the filter. The scaling factor is calculated as follows:
[0087]
[0088] Among them, r m reflects the absolute amplitude of the eigenvalue and represents the amplification degree of the filter to the input signal. The rotation angle is calculated as follows:
[0089]
[0090] Among them, θ m reflects the direction characteristic of the weight matrix. Take the mean values of the scaling factors and rotation angles of all convolutional kernels respectively to obtain the feature scores of each convolutional kernel:
[0091]
[0092] Perform normalization processing on the convolutional kernel scores and sum them to obtain the final weight score of the filter:
[0093]
[0094] Among them, f(·) is a normalization function used to map the calculation result to a standardized interval.
[0095] Finally, calculate the hybrid importance score of the filters to ensure that the evaluation result comprehensively considers the activation information and weight information and reduces the bias that may be caused by a single metric. The formula for the hybrid score is as follows:
[0096]
[0097] Among them, is the final importance score of the i-th filter in the τ-th layer, is the weight score, is the activation value score. Adopting the minimum value criterion can ensure that even if a certain part of the information is weak, the scoring result can still accurately reflect the importance of the filter, thereby improving the reliability of the pruning decision.
[0098] After completing the above calculations, normalize the hybrid importance scores of all filters to form the input data for the pruning strategy.
[0099] For step S3, based on the hybrid importance scores of the filters in each layer and the computational resource distribution of each layer, adopt a dynamic pruning strategy to optimize the computational efficiency and storage resources of the network. This strategy aims to optimize the computational amount (FLOPs) or the number of parameters by adaptively calculating the pruning threshold.
[0100] Perform in-layer resource statistics and metric calculations. For the i-th layer, calculate the proportion P i of the number of parameters in this layer to the total number of parameters in the entire model, and the proportion F i of the floating-point operation amount (FLOPs) in this layer to the total FLOPs of the entire model.
[0101] Define the resource ratio metric φ i to reflect the relative relationship between the computational resources and storage resources of the i-th layer:
[0102]
[0103] Among them, φ i ≥1 indicates that the computational amount is dominant, that is, the computational resources of this layer are tight; while φ i <1 indicates that the memory resources are dominant, that is, the memory resources of this layer are tight. Through this metric, the resource bottleneck of each layer can be accurately judged, so as to adopt different pruning strategies for different layers.
[0104] As Figure 4 shown, it shows the imbalance of FLOPs and parameter distribution in different layers.
[0105] Dynamically set the pruning threshold. In this process, first set a preset pruning ratio ε and sort the combined importance scores of all filters in the i-th layer from low to high, and take the score Q at the ε-th percentile as the pruning reference benchmark. At this time, the pruning threshold T is calculated according to the resource ratio φ and dynamically adjusted. The specific formula is as follows: i As the pruning reference benchmark. At this time, the pruning threshold T i is calculated according to the resource ratio φ i and is dynamically adjusted. The specific formula is as follows:
[0106] When φ i ≥1 (i.e., when computing resources are tight), the pruning threshold is calculated as:
[0107] T i = Q i + m·Q i ·ln(1 + ∈)
[0108] When φ i <1 (i.e., when memory resources are tight), the pruning threshold is calculated as:
[0109] T i = Q i - m·Q i ·ln(1 + ∈)
[0110] Among them, m is a parameter that controls the pruning amplitude and can adjust the pruning intensity according to actual needs. Through this dynamic adjustment method, the pruning threshold can be flexibly set according to the resource usage of the layer, so as to achieve a more precise pruning strategy.
[0111] As Figure 5 shown, it shows a schematic diagram of the residual pruning strategy.
[0112] For filter pruning, for each filter in each layer if its combined importance score is less than the pruning threshold T i of this layer, then this filter is removed from the network. For modules containing residual structures, such as convolutional layers in residual networks, special processing is required: retain the key convolutional layers (such as the first layer and the last layer) in the residual module, as well as the corresponding batch normalization layers, to ensure the coherence of the network structure and the normal propagation of gradients.
[0113] Through this strategy, it is possible to dynamically adjust the pruning strength according to the resource requirements of different layers, which can not only effectively reduce the pressure on computing resources and memory, but also maximize the retention of network performance, ensuring the efficiency and accuracy of the network after pruning.
[0114] For step S4, a network fine-tuning process is carried out after pruning to restore or optimize the performance of the model. The core objective of this step is to adjust the network weights through retraining, compensate for the possible performance loss caused by the pruning operation, and ensure that the network maintains high prediction ability while optimizing computational resources and storage resources.
[0115] After the pruning operation is completed, the integrity of the structure is retained, and only unimportant filters are removed. The pruned network structure becomes sparser, and the number of filters decreases. Therefore, the original training weight distribution may no longer be applicable to the new network structure. Therefore, it is necessary to fine-tune the pruned network.
[0116] During the fine-tuning process, the same training methods as the original training are adopted, including but not limited to optimization algorithms based on gradient descent. Common optimization algorithms such as Adam and SGD can be used in the fine-tuning process. The learning rate should be adjusted appropriately, and a warm-up strategy can be adopted, that is, a smaller learning rate is used in the initial stage, and the learning rate is gradually increased as the training process progresses.
[0117] To ensure the fine-tuning effect, first use the weights of the pruned network as the initial weights to start the fine-tuning training process. At this time, the model training does not start from scratch but continues to optimize the pruned network. The training objective of fine-tuning is to restore or further improve the accuracy of the model by minimizing the loss function.
[0118] The fine-tuned network can further restore the accuracy on the basis of pruning. And because pruning optimizes the computational complexity of the network, the execution efficiency of the model in actual applications is improved.
[0119] Through the above fine-tuning steps, it is ensured that the pruned network is not only optimized in computational efficiency but also can still maintain a high accuracy performance. This process can significantly improve the running efficiency of the network, especially on devices with limited computational resources, enabling deep learning models to achieve more efficient deployment while ensuring performance.
[0120] Generally speaking, the present invention evaluates the contribution of each filter in a specific layer by calculating the mixed importance score of each filter, combining activation value and weight information. Then, according to the ratio of computational resources and storage resources of each layer, the pruning threshold is dynamically set to accurately screen the filters to be pruned. Finally, the performance of the pruned network is restored through fine-tuning training to ensure that the network maintains high prediction accuracy while reducing computational overhead. This method can effectively balance the accuracy and efficiency of the model and is particularly suitable for resource-constrained application scenarios.
[0121] Experimental example:
[0122] This experiment aims to verify the performance of the pruning method of the present invention on different neural network architectures (VGG16, ResNet56, DenseNet169) and different datasets (CIFAR-10, CIFAR-100, ImageNet). The experimental results show that the proposed method can maintain a high model accuracy while significantly reducing the model's computational complexity and storage requirements, and is superior to various existing pruning methods. The following table shows the experimental effects of this method at different pruning ratios.
[0123]
[0124]
[0125]
[0126] As Figure 6 shown, the stability characteristics of HIS (Hybrid Importance Score) are presented. The HIS values calculated for different input images still maintain similar numerical values, indicating the robustness and consistency of this method on different images.
[0127] In addition, as Figure 7 shown, the decline in accuracy at different pruning ratios is presented, where HP 2 is the method of the present invention, HP 2 -F and HP 2 -P correspond to pruning strategies based on FLOPs (computational volume) and the number of parameters (Params), respectively. It can be clearly seen that the method proposed by the present invention can effectively reduce the decline in accuracy compared with other pruning methods, and while reducing the computational and storage overheads, maintain a high performance.
[0128] Through these experimental data, the advantages of the method of the present invention in computational efficiency and model accuracy are verified. Especially in the case of limited parameters and computational resources, it can effectively optimize the performance of the neural network.
[0129] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A convolutional neural network filter pruning method based on hybrid importance score guidance, characterized in that: The following steps are involved: Train the baseline model to extract the filter activation values and weight information of the target convolutional layer; Calculating a hybrid importance score for each filter based on the activation value and weight information; According to the computing and storage resource requirements of each layer, the pruning threshold is dynamically set to filter and prune filters with importance scores below the threshold; Fine-tune the pruned model to restore model accuracy.
2. The convolutional neural network filter pruning method based on hybrid importance score guidance according to claim 1, characterized in that: The benchmark model is any one of the VGG, ResNet or DenseNet architectures.
3. The convolutional neural network filter pruning method based on hybrid importance score guidance according to claim 1, characterized in that: The benchmark model adopts a convolutional neural network including a residual module, and synchronously extracts the activation value and weight information of the filter during the training process, wherein the activation value is the output matrix after the input image passes through the filter, and the weight information is the complex domain eigenvalue decomposition result of the convolution kernel.
4. The convolutional neural network filter pruning method based on hybrid importance score guidance according to claim 1, characterized in that: The calculation of the hybrid importance score includes: Activation-based scoring: The normalized activation score is calculated by taking the activation matrix rank output by the statistical filter. Weight-based scoring: By decomposing the convolution kernel matrix, analyzing the scaling factor and rotation angle of the eigenvalues, and calculating the normalized weight score; The minimum value of the activation score and the weight score is taken as the final mixed importance score.
5. The convolutional neural network filter pruning method based on hybrid importance score guidance according to claim 4, characterized in that: The calculation of the activation value score specifically includes: Input multiple images to the network and extract the activation matrix of the target filter output at the convolutional layer; Calculate the rank of each activation matrix, the rank representing the number of linearly independent features in the activation matrix; The rank is normalized, and the average value of multiple images is calculated to obtain the activation value score.
6. The convolutional neural network filter pruning method based on hybrid importance score guidance according to claim 4, characterized in that: The calculation of the weight score specifically includes: Decompose the convolution kernel matrix of the target filter and extract the complex domain eigenvalues; Calculate a scaling factor and a rotation angle according to the eigenvalue, wherein the scaling factor represents the modulus of the eigenvalue, and the rotation angle represents the direction of the eigenvalue on the complex plane; The scaling factor and rotation angle are normalized separately and aggregated into a weighted score.
7. The convolutional neural network filter pruning method based on hybrid importance score guidance according to claim 1, characterized in that: The step of dynamically setting the pruning threshold according to the computing resource and storage resource requirements of each layer includes: Calculate the parameter ratio P of each layer i And the floating point operation ratio F i , where the parameter ratio P i Indicates the ratio of the number of parameters in the i-th layer to the total number of model parameters and the floating-point operation ratio F i Indicates the ratio of the floating-point operations of the i-th layer to the total floating-point operations of the model; According to P i and F i Computing resource ratio indicator Based on φ i The value of selects the pruning strategy: When φ i When ≥1, a computationally intensive pruning strategy is adopted; When φ i When <1, a storage-optimized pruning strategy is adopted; The pruning threshold is dynamically adjusted according to the selected pruning strategy.
8. The convolutional neural network filter pruning method based on hybrid importance score guidance according to claim 7, characterized in that: The step of dynamically adjusting the pruning threshold according to the selected pruning strategy includes: According to the mixed importance score sorting of the current layer filter, determine the quantile value Q corresponding to the pruning ratio ε i ; Based on the resource ratio indicator φ i , combined with the pruning ratio ε and the adjustment parameter m, dynamically calculate the pruning threshold T i : When φ i ≥1: T i =Q i +m·Q i ln1+∈), increase the threshold to increase the number of pruning; When φ i <1: T i =Q i -m·Q i ln(1+∈), lower the threshold to reduce the amount of pruning.
9. The convolutional neural network filter pruning method based on hybrid importance score guidance according to claim 1, characterized in that: In the step of screening and pruning filters with importance scores lower than a threshold, the operation of pruning filters includes: The first and last convolutional layers and the corresponding batch normalization layers are retained for the residual module to ensure the consistency of the network structure.
10. The convolutional neural network filter pruning method based on hybrid importance score guidance according to claim 1, characterized in that: The fine-tuning training uses the same optimizer and learning rate strategy as the baseline model.
Citation Information
Cited By
Intelligent inhaul cable monitoring method, edge calculation unit and monitoring system
CN122242614A