Target detection and segmentation pruning method and system based on multistage feature decorrelation
By adding feature decorrelation constraints in neural network model training and pruning according to the mode value of the convolution kernel, the problems of unstable model performance and incomplete redundant feature removal in the existing pruning methods are solved, and more efficient model compression and performance improvement are achieved.
Patent Information
- Application Number
- CN202510124317.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-23
AI Technical Summary
The existing neural network pruning methods have problems such as unstable model performance, incomplete redundant feature removal, and difficult to control the pruning effect.
During the network model training process, the modulus values of each convolution kernel of each convolution layer are calculated, and the convolution kernel with smaller modulus values and their parameters and connections are cut according to the preset pruning rate to reduce redundant parameters.
It improves the accuracy and efficiency of pruning, reduces the number of parameters and calculations of the model, and ensures the performance and generalization capabilities of the model.
Smart Images

Figure CN120032108A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a target detection and segmentation pruning method and system based on multi-level feature decorrelation. Background Art
[0002] As one of the core technologies of deep learning, convolutional neural network (CNN) has made remarkable achievements in many fields such as image recognition, semantic segmentation, object detection, natural language processing and speech analysis in recent years. With the continuous development of CNN architecture, its performance continues to improve, but it is also accompanied by the increase of model complexity, which is reflected in the expansion of parameter scale, increased memory consumption, and the increase of energy demand and computing cost in the inference process. Therefore, high-performance CNN models usually need to be deployed on servers or devices with powerful computing power and abundant resources. However, for edge computing platforms, such as embedded systems and small mobile terminals, these devices are limited by low power budget, limited computing resources and storage space, and it is difficult to directly support complex CNN models. With the rapid development of the edge application market and the growing demand for compact and efficient models, researchers have begun to pay attention to neural network optimization technology, among which pruning methods have become an important means to reduce model size and computational burden and improve model adaptability.
[0003] The motivation for neural network pruning is to solve the inherent inefficiencies of large-scale models, which are usually manifested as redundancy of parameters and connections. By removing these redundancies in a targeted manner, pruning aims to simplify the model architecture, reduce computational overhead, and increase the speed of inference. This method not only reduces the computational burden in deep learning tasks, but also helps to deploy models on resource-constrained devices and accelerate the training and use of models. Existing neural network pruning methods have certain shortcomings. Most researchers' ideas are to impose penalties on model parameters, such as the weight parameters of the convolutional layer and the weights and biases of the BN layer, to achieve a sparse effect, thereby removing smaller parameters and connections. However, this method is relatively crude and the network performance during training is highly sensitive to the penalty coefficient, and the pruning effect is often difficult to predict. In addition, existing technologies usually ignore the optimization from the perspective of reducing the correlation of responses between layers of the network, especially the correlation of output features of the convolutional layer, so that important features cannot be effectively retained during the pruning process, thereby affecting the overall performance. Summary of the invention
[0004] In view of the defects in the prior art, the purpose of the present invention is to propose a target detection and segmentation pruning method and system based on multi-level feature decorrelation, which effectively solves the problems of unstable model performance, incomplete removal of redundant features and difficult to control pruning effects in existing pruning methods, thereby improving the accuracy and efficiency of pruning.
[0005] To achieve the above object, the present invention is implemented by the following technologies:
[0006] The present invention provides a target detection and segmentation pruning method based on multi-level feature decorrelation, comprising:
[0007] Step S1, during the network model training process, constraints are added to the correlation between the hidden features at each level of the network to perform constraint training, so as to optimize the network and reduce feature redundancy;
[0008] Step S2, based on the constrained training network model, calculating the modulus of each convolution kernel in the convolution layer of the network model, which is used to measure the importance of the corresponding channel of the convolution layer;
[0009] Step S3, sorting the modulus values of the convolution kernels according to a preset pruning rate, and pruning the convolution kernels with smaller modulus values and their corresponding parameters and connections according to the pruning ratio to reduce redundant parameters;
[0010] Step S4, calculating evaluation results of the pruned network model on the test image set, wherein the evaluation results include any evaluation type including parameter and computational compression ratio, model performance loss, and model running speed, which are used to measure the pruning effect;
[0011] Step S5: When the evaluation result does not reach the preset index, steps S1 to S4 are executed repeatedly until the preset index is met.
[0012] Furthermore, in step S1, during the network model training process, constraint conditions are added to the correlation between the hidden features at each level of the network to perform constraint training, so as to optimize the network and reduce redundant features, including:
[0013] Step S11, extract the output features of each convolutional layer in the network model after the activation function, and calculate the correlation value F of the hidden features of each layer IJ , used to set the hyperparameters of constrained training to adjust the constraint strength of feature decorrelation;
[0014] Step S12, according to the correlation value F IJ The multi-stage feature decorrelation loss function is calculated and combined with the traditional task loss to form a joint loss function. The loss value based on the joint loss function reaches the preset threshold, which is used to optimize task performance and reduce feature redundancy during the constraint training process to complete the constraint training.
[0015] Furthermore, in step S11, the correlation value F IJ The calculation formula is:
[0016]
[0017] Among them, F IJis the correlation value between the i-th channel feature and the j-th channel feature, and the correlation value ranges from -1 to 1, b is the number of samples in the batch, I k , J k The observed value of the kth sample representing the Ith and Jth dimension features is a two-dimensional feature map. Represents the average value of the I-th and J-th dimension features in this batch.
[0018] Furthermore, in step S12, the multi-stage feature decorrelation loss function is:
[0019]
[0020] Among them, d is the number of channels, that is, the number of features, i and j are channels; F IJ is the correlation value between the i-th channel feature and the j-th channel feature.
[0021] Furthermore, in step S12, the calculation formula of the joint loss function of constraint training is:
[0022] L=λ 1 L loc +λ 2 L obj +λ 3 L cls +λL MFD ;
[0023] Among them, L loc is the positioning loss, L obj is the confidence loss, L cls is the classification loss, s indicates that the feature decorrelation constraints need to be performed on s convolutional layers. λ is the balancing factor of each loss term, which is used to control the contribution of the feature decorrelation loss function and the traditional task loss function.
[0024] Furthermore, in step S2, based on the constrained trained network model, the modulus of each convolution kernel in the convolution layer of the network model is calculated to measure the importance of the corresponding channel of the convolution layer, including:
[0025] Step S21, extracting the weight parameters of each convolution kernel in the convolution layer, wherein the shape of each convolution kernel is mxn;
[0026] Step S22, based on the height and width of each convolution kernel K and the corresponding weight parameter, calculate the modulus value of the convolution kernel, the formula is:
[0027]
[0028] Among them, m, n are the height and width of the convolution kernel respectively, and V is the modulus value;
[0029] Step S23, based on the number of convolution kernels, that is, the number of output channels, the modulus of all convolution kernels is calculated to be [V 0 , V 1 …V c ].
[0030] Furthermore, in step S3, the modulus values of the convolution kernels are sorted according to a preset pruning rate, and the convolution kernels with smaller modulus values and their corresponding parameters and connections are pruned according to the pruning ratio to reduce redundant parameters, including:
[0031] According to the set pruning rate, that is, the pruning ratio R, the modulus value [V 0 , V 1 …V c ] are sorted by size, and according to the pruning ratio R, the convolution kernels of the last C*R channels and the corresponding parameters and connections are selected.
[0032] Furthermore, step S4 also includes, after pruning the network model, evaluating the performance of the pruned model through a data validation set and performing fine-tuning training; the fine-tuning training includes removing feature decorrelation loss, reducing the pruning rate and / or reducing the balance factor λ of the feature decorrelation loss.
[0033] Based on the same inventive concept, the present invention provides a target detection and segmentation pruning system based on multi-level feature decorrelation, which adopts the target detection and segmentation pruning method of multi-level feature decorrelation as described above, including:
[0034] The constraint training module is used to add constraints to the correlation between the hidden features of each layer of the network during the network model training process to perform constraint training, which is used to optimize the network and reduce feature redundancy, including extracting the output features of each convolutional layer in the network model after the activation function, and calculating the correlation value F of the hidden features of each layer. IJ , used to set the hyperparameters of constraint training to adjust the constraint strength of feature decorrelation; according to the correlation value F IJ Calculate the multi-stage feature decorrelation loss function and combine it with the traditional task loss to form a joint loss function. The loss value based on the joint loss function reaches the preset threshold, which is used to optimize task performance and reduce feature redundancy during constraint training to complete constraint training;
[0035] The feature evaluation module is used for the network model based on constraint training to calculate the modulus of each convolution kernel in the convolution layer of the network model, which is used to measure the importance of the corresponding channel of the convolution layer;
[0036] The model pruning module is used to sort the modulus values of the convolution kernels according to the preset pruning rate, and prune the convolution kernels with smaller modulus values and their corresponding parameters and connections according to the pruning ratio to reduce redundant parameters;
[0037] The model evaluation module is used to calculate the evaluation results of the pruned network model on the test image set. The evaluation results include any evaluation type including parameter and computational compression ratio, model performance loss and model running speed, which are used to measure the pruning effect. When the evaluation result does not meet the preset indicators, the constraint training module is cycled until the preset indicators are met.
[0038] Furthermore, a fine-tuning module is included for evaluating the performance of the pruned model through a data validation set and performing fine-tuning training; the fine-tuning training includes removing feature decorrelation loss, reducing the pruning rate and / or reducing the balancing factor λ of the feature decorrelation loss.
[0039] Compared with the prior art, the present invention has at least one of the following technical effects:
[0040] The technical solution of the present invention effectively solves the problems of unstable model performance, incomplete removal of redundant features and difficult control of pruning effects in existing pruning methods, thereby improving the accuracy and efficiency of pruning. First, by adding constraints on the correlation between hidden features at all levels of the network in the training of the network model for target detection and segmentation, a network model with more refined features is obtained, and then the importance of the corresponding channel is judged by calculating the modulus of the convolution kernel of each convolution layer of the network model. According to the expected pruning rate, the unimportant convolution layer channels of the corresponding proportion are cut off, and a lighter neural network model is obtained after pruning. Then, the pruning effect of the pruned model is evaluated on the test image set, and finally, the performance of the model can be restored by fine-tuning training. The above process can be repeated until the best effect is achieved. The present invention is applicable to most commonly used target detection and segmentation models, such as SSD and yolo series networks. On the VOC data set, the present invention can reduce the number of parameters and the amount of calculation of yolov5 by half while sacrificing the 1.3% mAP of the model. While ensuring the accuracy and reliability of the neural network model, the present invention can effectively compress the size of the model, reduce the number of model parameters and the amount of calculation, so that the network model does not need to consume a lot of resources and is easier to implement and deploy.
[0041] (1) The present invention uses the loss function MFD Loss to enable the network to autonomously select more effective and refined features for learning during the learning process without destroying the trainability of the network;
[0042] (2) The pruning rate of the present invention can be set by the user, and the user can choose to prune any proportion of the convolution kernel. The user can also choose to constrain and prune a specific convolution layer, and can screen the effect of global redundant parameters by constraining a single layer, which can meet different user needs.
[0043] (3) The present invention can achieve different effects by adjusting the constraint strength, that is, the size of the balance factor. At a smaller compression rate, the performance of the model can be improved while compressing the network model;
[0044] (4) The present invention comprehensively evaluates the pruned neural network model on the test image set to objectively evaluate the model performance and ensure the effectiveness of the pruning operation. This evaluation method helps to verify the generalization ability of the pruned model, reduce the risk of overfitting, and achieve an effective closed loop of the neural network model pruning operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solution of the embodiment of the present invention, the following is a brief introduction to the drawings required for describing the embodiment:
[0046] Figure 1 The present invention is a flowchart of the method implementation steps in the embodiment of the present invention.
[0047] Figure 2 Schematic diagram of the use of loss function in constrained training in an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the chimney detection method based on YOLO-RSOD provided by the present invention is described in detail below in combination with the drawings in the embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and specific operation process are given.
[0049] First embodiment
[0050] like Figure 1 As shown, the present invention provides a target detection and segmentation pruning method based on multi-level feature decorrelation, comprising:
[0051] Step S1, during the network model training process, constraints are added to the correlation between the hidden features of each level of the network for constraint training, which is used to optimize the network and reduce feature redundancy; this step is constraint training, which adds constraints on the correlation between the hidden features of each level of the network during the network model training to obtain a neural network model with more refined features
[0052] Step S2, based on the constrained training network model, calculate the modulus of each convolution kernel in the convolution layer of the network model to measure the importance of the corresponding channel of the convolution layer; this step is feature evaluation, and the importance of the corresponding channel is measured by calculating the modulus of each convolution kernel in the convolution layer of the network model;
[0053] Step S3, sorting the modulus values of the convolution kernels according to the preset pruning rate, and pruning the convolution kernels with smaller modulus values and their corresponding parameters and connections according to the pruning ratio to reduce redundant parameters; this step is channel pruning: according to the preset pruning rate, pruning the unimportant parameters and connections of the convolution layer in a corresponding proportion, and obtaining a lighter neural network model after pruning;
[0054] Step S4, by calculating the evaluation results of the pruned network model on the test image set, wherein the evaluation results include any evaluation type including parameter and computational compression ratio, model performance loss and model running speed, which are used to measure the pruning effect; this step is to judge the effectiveness, and the pruning effect of the pruned model is evaluated on the test image set, specifically including parameter and computational compression ratio, model performance loss and model running speed.
[0055] Step S5: When the evaluation result does not reach the preset index, steps S1 to S4 are executed repeatedly until the preset index is met.
[0056] Furthermore, in step S1, during the network model training process, constraint conditions are added to the correlation between the hidden features at each level of the network to perform constraint training, so as to optimize the network and reduce redundant features, including:
[0057] Step S11, extract the output features of each convolutional layer in the network model after the activation function, and calculate the correlation value F of the hidden features of each layer IJ , used to set the hyperparameters of constrained training to adjust the constraint strength of feature decorrelation;
[0058] Step S12, according to the correlation value F IJ The multi-stage feature decorrelation loss function is calculated and combined with the traditional task loss to form a joint loss function. The loss value based on the joint loss function reaches the preset threshold, which is used to optimize task performance and reduce feature redundancy during the constraint training process to complete the constraint training.
[0059] Furthermore, in step S11, the correlation value F IJ The calculation formula is:
[0060]
[0061] Among them, F IJ is the correlation value between the i-th channel feature and the j-th channel feature, and the correlation value ranges from -1 to 1, b is the number of samples in the batch, I k , J k The observed value of the kth sample representing the Ith and Jth dimension features is a two-dimensional feature map. Represents the average value of the I-th and J-th dimension features in this batch.
[0062] Furthermore, in step S12, the multi-stage feature decorrelation loss function is:
[0063]
[0064] Among them, d is the number of channels, that is, the number of features, i and j are channels; F IJ is the correlation value between the i-th channel feature and the j-th channel feature.
[0065] Furthermore, in step S12, the calculation formula of the joint loss function of constraint training is:
[0066] L=λ 1 L loc +λ 2 L obj +λ 3 L cls +λL MFD ;
[0067] Among them, L loc is the positioning loss, L obj is the confidence loss, L cls is the classification loss, s indicates that the feature decorrelation constraints need to be performed on s convolutional layers. λ is the balancing factor of each loss term, which is used to control the contribution of the feature decorrelation loss function and the traditional task loss function.
[0068] Specifically, the constraint training is mainly to increase the constraints on the correlation between the hidden features of each layer of the neural network model. First, extract the output features of each layer, which are the output features after the activation function in the convolutional neural network. The correlation value F between the hidden features of each layer of the neural network is IJ The calculation formula of the multi-stage feature decorrelation loss function is:
[0069]
[0070] When describing the relationship between features, the Pearson correlation coefficient matrix F is constructed. The matrix is symmetrical and the element values on its diagonal are 1, indicating that each feature is completely positively correlated with itself. The dimension of the matrix is determined by the number of features (i.e., the number of channels d), where the element F in the i-th row and j-th column is IJ represents the Pearson correlation coefficient between the i-th channel feature and the j-th channel feature. Since the value range of the Pearson correlation coefficient is limited to between -1 and 1, when calculating the loss function, these coefficients are usually squared to strengthen the impact of the correlation. b is the number of samples in each batch, I k , J k The observation value of the kth sample representing the Ith and Jth dimension features is a two-dimensional feature map with different resolutions at each level. Represents the average value of the I-th and J-th dimension features in this batch.<A,B> It is used to represent the inner product operation between matrices A and B, that is, the sum of the multiplication of the elements at corresponding positions of the two matrices, and the final result is a single value or scalar. This operation is used to quantify the similarity or correlation strength between two feature vectors. In this way, we can evaluate the degree of linear dependence between different features and incorporate this dependency into the loss calculation during model training, thereby guiding the optimization algorithm to adjust the model parameters more effectively. By calculating the correlation value between the implicit features of each layer of the original network model, the size of the correlation between the features within the layer is determined, which serves as the basis for setting hyperparameters in subsequent constraint training;
[0071] The joint loss function expression for constrained training is:
[0072] L=λ 1 L loc +λ 2 L obj +λ 3 L cls +λL MFD
[0073] λ 1 L loc +λ 2 L obj +λ 3 L cls is a loss function commonly used in target detection and segmentation network models, where L loc is the positioning loss, L obj is the confidence loss, L cls is the classification loss. s means that the feature decorrelation constraints need to be performed on s convolutional layers. λ is a constant, which is a balancing factor. According to the calculated L MFD The value can set the number of layers, that is, s different λ values. Increasing λ can enhance the degree of constraint on feature correlation. For different data sets and neural network models, choosing an appropriate λ value can improve model performance while pruning the model. Too large a λ will cause the model to lose its normal learning ability, and too small a λ will weaken or lose the constraint effect on the correlation between features. In the process of network model training, by adding a penalty term for feature correlation to the loss function, the model automatically reduces the redundancy between features, autonomously learns more refined features, and reduces the modulus of redundant convolutional layer weights, which is conducive to screening out redundant parameters and connections.
[0074] The key to constrained training is the use of MFD Loss. Specifically, taking yolov5 as an example, see Figure 2:CBS module, Resx module, CSP_x module and SSPF module. Each module has a CBS module part. In constraint training, we choose to extract the output of the CBS module, input it into the MFD calculation module, add the output of all MFD modules to the original loss of the model, and back propagate in training to achieve the constraint effect.
[0075] Furthermore, in step S2, based on the constrained trained network model, the modulus of each convolution kernel in the convolution layer of the network model is calculated to measure the importance of the corresponding channel of the convolution layer, including:
[0076] Step S21, extracting the weight parameters of each convolution kernel in the convolution layer, wherein the shape of each convolution kernel is mxn;
[0077] Step S22, based on the height and width of each convolution kernel K and the corresponding weight parameter, calculate the modulus value of the convolution kernel, the formula is:
[0078]
[0079] Among them, m, n are the height and width of the convolution kernel respectively, and V is the modulus value;
[0080] Specifically, the weight parameters of the convolution layer are calculated and evaluated. For a convolution kernel K with a shape of m*n, its L1 norm, that is, the sum of the absolute values of all elements in the matrix V, can be expressed as:
[0081]
[0082] For a specific layer, C is the number of convolution kernels, that is, the number of output channels. Calculate the modulus of all convolution kernels [V 0 , V 1 …V c ], the smaller the modulus value, the less important the channel corresponding to the convolution kernel is to the convolution layer;
[0083] Taking yo lov5 as an example, for the network basic module CBS, we directly extract the weight parameters of its convolution layer, and use step S21. to calculate the modulus of the convolution kernel of the CBS module convolution layer. For the network structure with residual connection, we do not simply use the parameters of the single-level layer as the basis for parameter evaluation, but take into account the convolution layers connected by the skip connection of the residual structure. Taking yo lov5 as an example, see Figure 2In the CSP_x module, the Conv layer parameters of the CBS module connected to the residual module Resx and the Conv layer parameters of the second CBS module in the x residual modules Resx, and the x+1 convolution kernel parameters are used as the basis for pruning the residual structure of the entire CSP_x module. In the residual module Resx, for the first CBS module, it is not directly connected to the residual, so we do not need to consider the residual connection and only need to treat it as a normal CBS module. In the SSPF module, there is no residual structure, so we also treat the CBS module as a normal CBS module.
[0084] Step S23, based on the number of convolution kernels, that is, the number of output channels, the modulus of all convolution kernels is calculated to be [V 0 , V 1 …V c ].
[0085] Furthermore, in step S3, the modulus values of the convolution kernels are sorted according to a preset pruning rate, and the convolution kernels with smaller modulus values and their corresponding parameters and connections are pruned according to the pruning ratio to reduce redundant parameters, including:
[0086] According to the set pruning rate, that is, the pruning ratio R, the modulus value [V 0 , V 1 …V c ] are sorted by size, and according to the pruning ratio R, the convolution kernels of the last C*R channels and the corresponding parameters and connections are selected. Specifically, channel pruning mainly selects the pruning ratio R for [V 0 , V 1 …V c ] are sorted by size, and the smaller C*R channels are selected, and these channels and connections are pruned. When the output channel of a layer is pruned, the input channel of the next layer is pruned accordingly. Pruning is flexible, and the pruning rate can be set by the user, and any proportion of convolution kernels can be pruned. Users can also choose to constrain and prune specific convolution layers, and can adjust the size of the constraints by controlling the size of the balance factor to meet different needs.
[0087] Taking yo lov5 as an example, for a simple CBS module, we cut off the corresponding output channels of the Conv layer of the CBS module, and at the same time cut off the corresponding channels of the BN layer. To maintain channel consistency, the corresponding input channels of the convolution layer connected to the CBS module should also be cut off. For the CSP module, channel consistency is guaranteed. At this time, it is necessary to pay attention to the special connection of Concat. The input channel index to be pruned of the CBS module after Concat is equal to the left CBS pruned channel index plus the Resx output channel number and the right Resx module output channel index. Similarly, for the SSPF module, after three Maxpool modules, the input channel index to be pruned of the CBS module after Concat is equal to the left CBS pruned channel index plus the index plus one times the CBS output channel number, the index plus two times the CBS output channel number, and the index plus three times the CBS output channel number, a total of four indexes are spliced to maintain channel consistency.
[0088] Furthermore, step S4 also includes, after pruning the network model, evaluating the performance of the pruned model through a data validation set, and performing fine-tuning training; the fine-tuning training includes removing feature decorrelation loss, reducing the pruning rate, and / or reducing the balance factor λ of the feature decorrelation loss. This step is mainly to restore the performance of the model through fine-tuning training.
[0089] The embodiment of the present invention enables the neural network to learn more refined and effective features by adding feature decorrelation constraints during the neural network training process, then determines the importance of the corresponding channel according to the module value of each convolution kernel of the convolution layer, prunes the unimportant channels, and finally uses a test image set to test the performance of the pruned neural network model to determine whether the pruning of the neural network model is effective, thereby ensuring that the pruned model has better generalization ability in practical applications and reducing the risk of overfitting, thereby achieving an effective closed loop of the neural network model pruning operation.
[0090] Second embodiment
[0091] Based on the same inventive concept, the present invention provides a target detection and segmentation pruning system based on multi-level feature decorrelation, which adopts the target detection and segmentation pruning method of multi-level feature decorrelation as described above, including:
[0092] The constraint training module is used to add constraints to the correlation between the hidden features of each layer of the network during the network model training process to perform constraint training, which is used to optimize the network and reduce feature redundancy, including extracting the output features of each convolutional layer in the network model after the activation function, and calculating the correlation value F of the hidden features of each layer. IJ , used to set the hyperparameters of constraint training to adjust the constraint strength of feature decorrelation; according to the correlation value F IJCalculate the multi-stage feature decorrelation loss function and combine it with the traditional task loss to form a joint loss function. The loss value based on the joint loss function reaches the preset threshold, which is used to optimize task performance and reduce feature redundancy during the constraint training process to complete the constraint training;
[0093] The feature evaluation module is used for the network model based on constraint training to calculate the modulus of each convolution kernel in the convolution layer of the network model, which is used to measure the importance of the corresponding channel of the convolution layer;
[0094] The model pruning module is used to sort the modulus values of the convolution kernels according to the preset pruning rate, and prune the convolution kernels with smaller modulus values and their corresponding parameters and connections according to the pruning ratio to reduce redundant parameters;
[0095] The model evaluation module is used to calculate the evaluation results of the pruned network model on the test image set. The evaluation results include any evaluation type including parameter and computational compression ratio, model performance loss and model running speed, which are used to measure the pruning effect. When the evaluation result does not meet the preset indicators, the constraint training module is cycled until the preset indicators are met.
[0096] Furthermore, a fine-tuning module is included for evaluating the performance of the pruned model through a data validation set and performing fine-tuning training; the fine-tuning training includes removing feature decorrelation loss, reducing the pruning rate and / or reducing the balancing factor λ of the feature decorrelation loss.
[0097] Although the present invention has been disclosed as above in the form of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the protection scope of the technical solution of the present invention.
Claims
1. A target detection and segmentation pruning method based on multi-level feature decorrelation, characterized in that: include: Step S1, during the network model training process, constraints are added to the correlation between the hidden features at each level of the network to perform constraint training, so as to optimize the network and reduce feature redundancy; Step S2, based on the network model trained by the constraint, calculating the modulus of each convolution kernel of the convolution layer in the network model, so as to measure the importance of the corresponding channel of the convolution layer; Step S3, sorting the modulus values of the convolution kernels according to a preset pruning rate, and pruning the convolution kernels with smaller modulus values and their corresponding parameters and connections according to the pruning ratio to reduce redundant parameters; Step S4, calculating an evaluation result of the pruned network model on a test image set, wherein the evaluation result includes any evaluation type including parameter and computational compression ratio, model performance loss, and model running speed, which is used to measure the pruning effect; Step S5, when the evaluation result does not reach the preset index, steps S1 to S4 are executed in a loop until the preset index is met.
2. The target detection and segmentation pruning method with multi-level feature decorrelation according to claim 1, characterized in that: In step S1, during the network model training process, constraints are added to the correlation between the hidden features at each level of the network to perform constraint training, which is used to optimize the network and reduce redundant features, including: Step S11, extracting the output features of each convolutional layer in the network model after the activation function, and calculating the correlation value F of the hidden features of each level IJ , used to set the hyperparameters of the constraint training to adjust the constraint strength of feature decorrelation; Step S12, according to the correlation value F IJ A multi-stage feature decorrelation loss function is calculated, and the multi-stage feature decorrelation loss function is combined with the traditional task loss to form a joint loss function; based on the loss value of the joint loss function reaching a preset threshold, it is used to simultaneously optimize the task performance and reduce the feature redundancy during the constraint training process to complete the constraint training.
3. The target detection and segmentation pruning method with multi-level feature decorrelation according to claim 2, characterized in that: In step S11, the correlation value F IJ The calculation formula is: Among them, F IJ is the correlation value between the i-th channel feature and the j-th channel feature, and the correlation value ranges from -1 to 1, b is the number of samples in the batch, I k , J k The observed value of the kth sample representing the Ith and Jth dimension features is a two-dimensional feature map. Represents the average value of the I-th and J-th dimension features in this batch.
4. The target detection and segmentation pruning method with multi-level feature decorrelation according to claim 3, characterized in that: In step S12, the multi-stage feature decorrelation loss function is: Among them, d is the number of channels, that is, the number of features, i and j are channels; F IJ is the correlation value between the i-th channel feature and the j-th channel feature.
5. The target detection and segmentation pruning method with multi-level feature decorrelation according to claim 4, characterized in that: In step S12, the calculation formula of the joint loss function of the constraint training is: L=λ1L loc +λ2L obj +λ3L cls +λL MFD ; Among them, L loc is the positioning loss, L obj is the confidence loss, L cls is the classification loss, s indicates that the s convolutional layers need to be subjected to feature decorrelation constraints; λ is the balancing factor of each loss term, which is used to control the contribution of the feature decorrelation loss function and the traditional task loss function.
6. The target detection and segmentation pruning method with multi-level feature decorrelation according to claim 1, characterized in that: In step S2, based on the network model trained by the constraint, the modulus of each convolution kernel in the convolution layer of the network model is calculated to measure the importance of the channel corresponding to the convolution layer, including: Step S21, extracting the weight parameter of each convolution kernel in the convolution layer, wherein the shape of each convolution kernel is mxn; Step S22, based on the height and width of each convolution kernel K and the corresponding weight parameter, calculate the modulus value of the convolution kernel, the formula is: Wherein, m and n are the height and width of the convolution kernel respectively, and V is the modulus value; Step S23, based on the number of the convolution kernels, that is, the number of output channels, the modulus of all the convolution kernels is calculated to be [V0, V1…V c ].
7. The target detection and segmentation pruning method with multi-level feature decorrelation according to claim 6, characterized in that: In step S3, the modulus values of the convolution kernels are sorted according to a preset pruning rate, and the convolution kernels with smaller modulus values and their corresponding parameters and connections are pruned according to the pruning ratio to reduce redundant parameters, including: According to the set pruning rate, i.e., the pruning ratio R, for all the modulus values of the convolution kernels [V0, V1…V c ] are sorted by size, and according to the pruning ratio R, the convolution kernels of the last C*R channels and the corresponding parameters and connections are selected.
8. The target detection and segmentation pruning method with multi-level feature decorrelation according to claim 4, characterized in that: The step S4 also includes, after the network model is pruned, evaluating the performance of the pruned model through a data validation set and performing fine-tuning training; the fine-tuning training includes removing the feature decorrelation loss, reducing the pruning rate and / or reducing the balancing factor λ of the feature decorrelation loss.
9. A target detection and segmentation pruning system based on multi-level feature decorrelation, using any one of the multi-level feature decorrelation target detection and segmentation pruning methods according to claims 1 to 8, characterized in that: include: The constraint training module is used to add constraints to the correlation between the hidden features of each layer of the network during the network model training process to perform constraint training, which is used to optimize the network and reduce feature redundancy. Including, extracting the output features of each convolutional layer in the network model after the activation function, and calculating the correlation value F of the hidden features of each level IJ , used to set the hyperparameters of the constraint training to adjust the constraint strength of feature decorrelation; according to the correlation value F IJ Calculating a multi-stage feature decorrelation loss function, and combining the multi-stage feature decorrelation loss function with a traditional task loss to form a joint loss function; based on the loss value of the joint loss function reaching a preset threshold, the task performance and the feature redundancy are simultaneously optimized during the constraint training process to complete the constraint training; A feature evaluation module, used to calculate the modulus of each convolution kernel of the convolution layer in the network model based on the network model trained by the constraint, so as to measure the importance of the corresponding channel of the convolution layer; A model pruning module is used to sort the modulus values of the convolution kernels according to a preset pruning rate, and prune the convolution kernels with smaller modulus values and their corresponding parameters and connections according to the pruning ratio to reduce redundant parameters; The model evaluation module is used to calculate the evaluation result of the pruned network model on the test image set, wherein the evaluation result includes any evaluation type including parameter and computational compression ratio, model performance loss and model running speed, which is used to measure the pruning effect; when the evaluation result does not meet the preset indicator, the constraint training module operation is looped until the preset indicator is met.
10. The multi-level feature decorrelation target detection and segmentation pruning system according to claim 9, characterized in that: It also includes a fine-tuning module for evaluating the performance of the pruned model through a data validation set and performing fine-tuning training; the fine-tuning training includes removing the feature decorrelation loss, reducing the pruning rate and / or reducing the balancing factor λ of the feature decorrelation loss.
Citation Information
Cited By
Model pruning method for heterogeneous cloud edge-end cooperative system
CN120806024A
Convolutional neural network pruning method and device based on Grubrum matrix orthogonality
CN121525769A
Convolutional neural network pruning method and device based on gram matrix orthogonality
CN121525769B