Remote sensing image classification model compression method based on incremental information guidance
By combining incremental information and static information, using gradient information and rank information to comprehensively evaluate the importance of the filter, and using integrated pruning and quantization design, the problems of insufficient characterization of the dynamic characteristics of the filter and the lack of collaborative optimization of pruning and quantization in the existing technology are solved, and efficient model compression and performance balance are achieved.
Patent Information
- Application Number
- CN202510205103.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-27
AI Technical Summary
The existing model compression technology is limited by the static information of model parameters, and the insufficient characterization of the dynamic characteristics of the filter leads to evaluation deviations, as well as the lack of collaborative optimization of pruning and quantization techniques.
By combining incremental information and static information, the model compression is guided by using gradient information, and the importance of the filter is comprehensively evaluated in combination with rank information. The integrated design of pruning and quantization is adopted, so that the model can achieve coordinated optimization of filter cropping and weight quantization during the compression process.
Effectively balance computing performance and model accuracy, significantly reduce model size and computational complexity, while maintaining good accuracy performance and improving compression effect.
Smart Images

Figure CN120219798A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and further relates to model compression technology. Specifically, it is a method for compressing a remote sensing image classification model guided by incremental information, which can be used for optimizing deep learning models, compressing models in edge computing devices, and other fields that require efficient storage and computing resources. Background Art
[0002] Remote sensing scene classification (RSCC) is a key task in the field of optical remote sensing image processing and analysis. In recent years, with the development of remote sensing technology and the increasing complexity of remote sensing data, traditional artificial feature extraction methods have shown limitations in the face of these challenging factors. Therefore, deep learning methods based on convolutional neural networks (CNNs) have become the mainstream choice due to their excellent feature learning and classification capabilities. However, CNN models usually contain a large number of parameters, resulting in excessive consumption of computing and storage resources, especially in the processing of high-resolution, multi-band remote sensing images. This resource requirement limits its deployment on edge devices, and model compression has thus become the key to solving this problem.
[0003] In existing model compression algorithms, network pruning precisely compresses the model scale through parameter pruning to improve the deployment efficiency; quantization techniques convert floating-point parameters into fixed-point formats to reduce memory and computational resource consumption. One of the most classic methods, Lin, Ji, and others found in "Hrank: Filter pruning using high-rank feature map[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2020:1529-1538" that the average rank of multiple feature maps generated by a single filter is always the same, and effective compression of the network is achieved by pruning filters with low-rank feature maps. Different from directly calculating importance scores using trained weights, He, Liu, and others proposed in "Filter pruning via geometric median for deep convolutional neural networks acceleration[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019:4340-4349." to prune redundant filters by calculating the geometric median of filters within the same layer, rather than compressing those with relatively less importance. In addition, Dong, et al. proposed in "Hawq: Hessian aware quantization of neural networks with mixed-precision[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019:293-302." that the relative quantization levels of each layer can be determined according to the Hessian spectrum of each layer. However, traditional deep learning model optimization methods have some inherent limitations. Such as the possible decline in model performance during pruning and the possible accuracy loss introduced by parameter quantization. These problems are particularly prominent in resource-constrained environments.
[0004] In recent years, hybrid strategies that combine multiple compression methods have received extensive attention. Compared with applying them separately, this combined strategy is more conducive to achieving an ideal balance between performance and efficiency, thus providing broader support for practical applications in multiple fields. Therefore, many scholars have conducted research on it. Han et al. proposed a three-stage deep compression method that combines pruning, post-training quantization, and Huffman coding in "Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding[J]. Fiber, 2015, 56(4): 3--7. DOI: 10.48550 / arXiv.1510.00149." Through the organic integration of multiple compression techniques, significant model compression effects were achieved while minimizing the loss of model accuracy. This method systematically verified the combined effects of multiple compression techniques for the first time and provided an efficient solution for neural network compression. However, this method adopted a globally unified fixed bit-width allocation strategy (such as 8-bit for convolutional layers and 5-bit for fully connected layers) in the quantization stage, did not establish a dynamic mapping relationship between filter importance and quantization bit-width, and lacked an adaptive quantization strategy. This limitation makes it impossible to dynamically adjust the quantization accuracy according to the importance of different network layers, task requirements, and the characteristics of the target hardware platform, which may affect the computational efficiency and inference performance of the model in specific scenarios. Li et al. proposed an efficient knowledge distillation method in "Few Sample Knowledge Distillation for Efficient Network Compression. 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE." This method regards the original network as the teacher network and the lightweight network after pruning as the student network, and adds 1×1 convolutional layers at the ends of each layer of the student network. Through least squares regression, the outputs of these additional layers are aligned with the teacher network to compensate for the information loss caused by pruning, while ensuring that no additional parameters and computational costs are introduced, thus achieving efficient model compression and fast distillation. However, this method has a strong dependence on the network structure, assuming that the student network and the teacher network are strictly aligned at the block level. This limitation restricts its applicability to heterogeneous compression architectures (such as networks generated by NAS) and makes it difficult to directly migrate to model compression tasks with significantly different structures. Although the above methods have made certain progress in terms of compression effects and efficiency, most of them are still based on the simple superposition of multiple compression techniques and fail to achieve collaborative optimization among them.Due to the lack of collaborative learning mechanism, these technologies fail to fully cooperate and exert their advantages under the same compression framework, resulting in limited improvement in the storage and computing efficiency of the model after compression, and often a large loss in accuracy. Especially when dealing with complex remote sensing image classification tasks, the compressed model may still face high computing requirements and cannot fully utilize hardware resources.
[0005] In addition, existing criteria are usually limited to static information about model parameters when evaluating filter importance, and do not fully consider the impact of incremental information on model performance. Static information mainly focuses on the current weights and structure of the model, but it ignores the dynamic update and adjustment of model parameters during the training process; while incremental information is the change in the model's response to input data in each training step, reflecting the model's adaptability and learning ability to data during the optimization process; therefore, incremental information plays a vital role in model training, and it can capture subtle changes and trends in model performance during training, providing a more comprehensive and accurate basis for filter importance evaluation. Summary of the invention
[0006] The purpose of the present invention is to propose a remote sensing image classification model compression method based on incremental information guidance in view of the deficiencies of the above-mentioned prior art, which is used to solve the technical problems that the existing model compression technology is limited by the static information of the model parameters, the insufficient characterization of the dynamic characteristics of the filter leads to evaluation deviation, and the lack of coordinated optimization of pruning and quantization technology. The present invention takes into account both incremental information and static information, uses gradient information to guide model compression, improves the adaptability of the model, and combines rank information to make up for the problem of incomplete information that may occur when only incremental information is used, thereby ensuring comprehensive coverage of static information. In addition, the present invention integrates pruning and quantization design, so that the model can achieve coordinated optimization of filter clipping and weight quantization during the compression process, effectively strikes a balance between computing performance and model accuracy, and can significantly reduce the model scale and computational complexity without affecting accuracy, thereby improving the compression effect.
[0007] To achieve the purpose of the invention, the technical solution adopted by the present invention includes the following steps:
[0008] (1) Obtain a remote sensing image dataset and preprocess it to obtain a preprocessed image x as the model input;
[0009] (2) The convolutional neural network (CNN) model is used as the remote sensing image classification model. The model is pre-trained using the pre-processed image x, that is, the model weights are optimized through forward propagation and back propagation algorithms to obtain the filter weights W of each layer. l ′, get the pre-trained model;
[0010] (3) Evaluate the importance of filters in the pre-trained model:
[0011] (3.1) Propagate the preprocessed image data through the pre-trained model forward, and then perform low-rank decomposition on the feature map output by the model to obtain its rank information r:
[0012]
[0013] Among them, is the feature map output by the i-th filter in the l-th convolutional layer, where l = 1, 2,..., D, and D is the number of convolutional layers included in the pre-trained model; s j represents the j-th singular value, u j represents the j-th left singular vector, v j T represents the j-th right singular vector;
[0014] (3.2) Use some preprocessed images in the validation set as the input data for calculating gradients, calculate the loss value through forward propagation, and accumulate the gradient information of each image using backpropagation to obtain the gradient statistics result of each filter where y is the label;
[0015] (3.3) According to the rank information obtained in step (3.1) and the gradient information obtained in step (3.2), calculate the importance score of each layer of filters:
[0016]
[0017] Among them, represents the importance score of the i-th filter in the l-th layer;
[0018] (4) Perform a global importance ranking on the filters to generate the final global importance ranking result:
[0019] (4.1) Initialize the pair of affine transformation parameters, which are the parameters to be learned for the l-th convolutional layer of each pre-trained model, including the scaling factor γ l and the constant term β l ;
[0020] (4.2) Calculate the global importance score of the i-th filter in the l-th convolutional layer according to the following formula and perform a global importance score ranking on the filters to obtain the ranking result {S↓} with the importance scores from high to low:
[0021]
[0022] (4.3) Design a candidate compression scheme according to the ranking result:
[0023] Define the pruning operation as 0-bit quantization, perform discrete bit-width quantization on the remaining filters, and set the allowable bit-width set as {2, 4, 8}; remove redundant filters according to {S↓}, for the remaining filters, that is, non-0 bits, introduce two threshold parameters, a low threshold ρ1 and a high threshold ρ2, and dynamically allocate bit-widths for them according to the importance scores of the current filters;
[0024] (4.4) Compress the preprocessed model according to the compression scheme in step (4.3), then calculate the accuracy of the compressed model on the validation set, and use the regular evolutionary algorithm to optimize the parameters γ l and β l , with the goal of maximizing the validation set accuracy to obtain the optimized affine transformation parameters;
[0025] (4.5) Calculate the final global importance scores of each filter based on the optimized affine transformation parameters, and sort these scores from high to low to generate the final global importance ranking result;
[0026] (5) Obtain a lightweight model through staged compression:
[0027] (5.1) Determine the filters that need to be 0-bit quantized, that is, the filters that need to be pruned, according to the final global importance ranking result and the preset pruning rate; after completing the pruning operation, fine-tune the model;
[0028] (5.2) For the remaining filters, that is, non-0 bits, introduce two threshold parameters, a low threshold ρ1 and a high threshold ρ2, and dynamically allocate bit-widths for them according to the importance scores of the current filters, and then perform mixed-precision quantization operations on the weight parameters and activation values of the convolutional layer respectively;
[0029] (5.3) After the mixed-precision quantization operation is completed, fine-tune the model again in the same way as in step (5.1) to obtain a compressed lightweight remote sensing image classification model.
[0030] Compared with other existing model compression methods, the present invention can make full use of incremental information and introduce key elements of real-time performance and dynamic adaptability during the compression process; through the incremental form based on gradients and combined with rank information, it ensures comprehensive coverage of dynamic information and static information; this method effectively generates an optimized compression strategy, while minimizing the model size and computational complexity, maintaining good accuracy performance.
[0031] The present invention has the following advantages compared with the existing technologies:
[0032] First, since the present invention proposes a joint compression framework guided by incremental information, it can make full use of the dynamic performance and learning ability of filters during model training to perform precise and efficient compression optimization on the model. Compared with traditional methods, this framework introduces real-time performance and dynamic adaptability in model compression, enabling the model to more flexibly adapt to the changes in remote sensing data distribution and the diversity of task requirements; through the deep combination of incremental information and pruning quantization technology, not only efficient compression of the model is achieved, but also the performance stability of the model in resource-constrained scenarios is significantly improved, ensuring efficient operation on edge devices; this framework can dynamically evaluate the importance of filters and optimize the model structure in real time according to task requirements, thus achieving a better balance between compression rate and performance and providing a more accurate and intelligent solution for remote sensing scene classification.
[0033] Second, the present invention introduces an incremental form of a gradient-based method to guide model compression by utilizing gradient information, enabling it to better adapt to dynamic changes and data distribution in remote sensing tasks; at the same time, in order to enhance the accuracy of gradient information, the present invention combines rank information for comprehensive evaluation to ensure the balance between the feature expression ability and learning ability of filters, thereby comprehensively evaluating the importance of each filter and avoiding the limitations that may occur when using gradient information alone; in addition, the randomly introduced strategy adopted by the present invention can effectively reduce the computational complexity and provide a more feasible and efficient solution for practical remote sensing applications.
[0034] Third, the present invention proposes a unified framework for pruning and quantization, regarding pruning as a special 0-bit quantization to achieve collaborative learning optimization during the pruning and quantization processes. Through mixed-precision design, higher bit widths are assigned to important filters, improving the balance between model compression efficiency and performance. This collaborative learning method effectively reduces the implementation complexity and achieves a better balance between compression rate and performance.
[0035] Fourth, since the present invention introduces a strategy based on learning cross-layer importance ranking, the sorting of filters is not limited to single-layer evaluation but comprehensively considers global-level information. Through the regular evolutionary algorithm, the compression strategy is further optimized, making the compressed model more stable and adaptable globally in terms of performance.
[0036] Fifth, the present invention simplifies the compression process by regarding pruning as 0-bit quantization and setting a unified quantization bit width for each layer; by directly removing unimportant filters through pruning operations, the computational and storage burdens are effectively reduced; the unified quantization bit width design avoids complex parameter adjustments, making the entire compression process more efficient and easy to implement. Brief Description of the Drawings
[0037] Figure 1It is the overall implementation flowchart of the method of the present invention;
[0038] Figure 2 It is the simulation result diagram of inputting the NWPU-45 dataset and compressing it on three mainstream networks respectively by the present invention. Among them, (a)-(c) are the result comparison diagrams of compressing the NWPU-45 dataset on the VGG-16, VGG-19, and ResNet-56 networks respectively by the present invention; (d) is the change situation of the compression accuracy of the three mainstream networks before and after combining incremental information on the NWPU-45 dataset;
[0039] Figure 3 It is the simulation result diagram of inputting the UCML-21 dataset and compressing it on three mainstream networks respectively by the present invention. Among them, (a)-(c) are the result comparison diagrams of compressing the UCML-21 dataset on the VGG-16, VGG-19, and ResNet-56 networks respectively by the present invention; (d) is the change situation of the compression accuracy of the three mainstream networks before and after combining incremental information on the UCML-21 dataset;
[0040] Figure 4 It is the schematic diagram of the application process of image classification compression using the method of the present invention. Detailed implementation manners
[0041] The present invention will be described in detail below with reference to the accompanying drawings.
[0042] The technical idea of the present invention is to calculate the gradient information obtained by the backpropagation of model training, combine the rank information of the feature maps extracted during the forward propagation of the model, and use it as the key basis for evaluating the importance of filters. In addition, by introducing a regularization evolutionary algorithm, the affine parameters required for calculating the global score are learned and optimized, so as to generate a global importance score, and use this to guide the design of the optimal compression strategy, and finally realize the integrated compression of model pruning and quantization.
[0043] Example 1: Refer to Figure 1 , the present invention proposes a method for compressing a remote sensing image classification model guided by incremental information, which specifically includes the following steps:
[0044] Step 1. Obtain a remote sensing image dataset and preprocess it to obtain the preprocessed image x as the model input, which is realized as follows:
[0045] (1.1) Obtain the original remote sensing image W from the publicly available remote sensing image dataset W×H×C Construct a sample set and divide it into a training set and a validation set; where W×H is the spatial resolution of the image and C is the number of channels;
[0046] (1.2) Preprocess the images in the sample set to obtain the preprocessed image x, which is used as the input image of the model; specifically, the preprocessing is to randomly crop the training set images to 256 pixels and perform random horizontal flipping; then convert them into tensors and standardize them, where the standardization uses the preset mean and standard deviation; at the same time, use central cropping to adjust the validation set images to 256 pixels, then convert them into tensors and standardize them.
[0047] Step 2. Use a convolutional neural network CNN model as the remote sensing image classification model. In this embodiment, the preferred CNN model is any one of VGG-16, VGG-19, and ResNet-56; use the preprocessed image x to pre-train the model, that is, optimize the model weights through forward propagation and backpropagation algorithms to obtain the filter weights W l ′ of each layer to obtain the pre-trained model.
[0048] The above optimization of the model weights through forward propagation and backpropagation algorithms to obtain the filter weights W l ′ is specifically implemented as follows:
[0049] O l (x) = f(W l ·O l-1 (x) + b l ),
[0050]
[0051] where O l represents the feature map of the l-th layer, f(·) is the activation function, W l and b l are the original weight matrix and bias of the l-th layer respectively, W l ′ and b l ′ represent the optimized weight matrix and bias of the l-th layer, δ l is the error term, L is the loss function, and η is the learning rate.
[0052] Step 3. Evaluate the importance of the filters in the pre-trained model:
[0053] (3.1) Perform forward propagation of the preprocessed image data through the pre-trained model to extract the feature maps of each layer. This feature map can be regarded as the activation matrix generated after the convolution operation, where each channel corresponds to a two-dimensional matrix representing the feature information extracted by a specific convolution kernel; to further evaluate the expression ability of these features, then perform low-rank decomposition on the feature map matrix of each channel, and approximate the rank information by decomposing the principal components of the feature map, that is, perform low-rank decomposition on the feature map output by the model to obtain its rank information r:
[0054]
[0055] Among them, is the feature map output by the i-th filter in the l-th convolutional layer, where l = 1, 2,..., D, and D is the number of convolutional layers included in the pre-trained model; s j represents the j-th singular value, and u j represents the j-th left singular vector, and v j T represents the j-th right singular vector;
[0056] (3.2) Use some preprocessed images in the validation set as the input data for calculating gradients. Specifically, in this embodiment, this data is obtained by randomly selecting 90 - 100 preprocessed images from the validation set; calculate the loss value through forward propagation, and accumulate the gradient information of each image using backpropagation to obtain the gradient statistical results of each filter where y is the label;
[0057] (3.3) According to the rank information obtained in step (3.1) and the gradient information obtained in step (3.2), calculate the importance score of each layer of filters:
[0058]
[0059] Among them, represents the importance score of the i-th filter in the l-th layer;
[0060] Step Four. Perform a global importance ranking on the filters to generate the final global importance ranking result:
[0061] (4.1) Initialize the pair of affine transformation parameters, which are the parameters to be learned for the l-th convolutional layer of each pre-trained model, including the scaling factor γ l and the constant term β l ;
[0062] (4.2) Calculate the global importance score of the i-th filter in the l-th convolutional layer according to the following formula and perform a global importance score ranking on the filters to obtain the ranking result {S↓} with the importance scores from high to low:
[0063]
[0064] (4.3) Design a candidate compression scheme according to the ranking result:
[0065] Define the pruning operation as 0-bit quantization, perform discrete bit-width quantization on the retained filters, and set the allowable bit-width set as {2, 4, 8}; remove redundant filters according to {S↓}, and for the retained filters, that is, non-0 bits, introduce two threshold parameters, a low threshold ρ1 and a high threshold ρ2, and dynamically allocate bit-widths for them according to the importance scores of the current filters.
[0066] The dynamic bit-width allocation performed in this embodiment in this step is carried out in the following manner: When holds, allocate a quantization bit-width of 2 to the current filter; when holds, allocate a quantization bit-width of 4 to the current filter; when holds, allocate a quantization bit-width of 8 to the current filter; for the same convolutional layer l of the model, use the maximum value of the filter bit-widths within this layer as the unified bit-width B l :
[0067]
[0068] where c l represents the number of filters contained in the l-th convolutional layer of the model.
[0069] (4.4) Compress the preprocessed model according to the compression scheme in step (4.3), and then calculate the accuracy of the compressed model on the validation set, and use the regular evolutionary algorithm to optimize the parameters γ l and β l , with the goal of maximizing the validation set accuracy to obtain the optimized affine transformation parameters. In this embodiment, the accuracy of the network is calculated on the validation set, and preferably the regular evolutionary algorithm is used for parameter optimization. Of course, here it is also possible to achieve the optimization goal through genetic algorithms, differential evolution algorithms, or Bayesian optimization algorithms, etc. Preferably, the number of iterations is set to 400 for optimization training.
[0070] (4.5) Calculate the final global importance score of each filter based on the optimized affine transformation parameters, and sort the scores from high to low to generate the final global importance ranking result;
[0071] Step Five. Obtain a lightweight model through staged compression:
[0072] (5.1) In order to alleviate the performance loss that may be brought about by pruning and quantization, the present invention adopts a staged compression and fine-tuning strategy for the model in this step.
[0073] According to the final global importance ranking result and the preset pruning rate, determine the filters that need to be quantized to 0 bits, that is, the filters that need to be pruned; after completing the pruning operation, in order to recover the performance loss caused by pruning, fine-tune the model. In this embodiment, during fine-tuning, the preprocessed image is used to perform the same pre-training process on the compressed model as in step (2), and the weights of the compressed model are optimized through forward propagation and backpropagation algorithms, and its training process is not less than 300 rounds.
[0074] (5.2) For the retained filters, that is, non-0 bits, introduce two threshold parameters, a low threshold ρ1 and a high threshold ρ2, and dynamically allocate bit widths according to the importance scores of the current filters, and then perform mixed-precision quantization operations on the weight parameters and activation values of the convolutional layer respectively. The specific method of dynamically allocating bit widths in this step of this embodiment is the same as the allocation method in step (4.3); the mixed-precision quantization operations performed on the weight parameters and activation values of the convolutional layer are specifically expressed as follows:
[0075]
[0076] Among them, Q(·) is the quantization function, A l is the activation value of the l-th layer.
[0077] (5.3) After the mixed-precision quantization operation is completed, fine-tune the model again in the same way as in step (5.1) to obtain a compressed and lightweight remote sensing image classification model.
[0078] The present invention combines gradient information and rank information, overcomes the problem of insufficient attention to incremental information in the prior art when evaluating the importance of filters, and uses static information to make up for the problem of incomplete information that may be caused by only using incremental information. Using global ranking can more comprehensively and reliably evaluate the importance of filters. In addition, using an integrated compression strategy can give full play to the respective advantages of pruning and quantization under the condition of collaborative learning, effectively improving the compression efficiency and performance of the model.
[0079] Embodiment 2: The overall implementation steps of the model compression method proposed in this embodiment are the same as those in Embodiment 1. Now, the calculation process of the model gradient information involved in step (3.2) will be further described in detail:
[0080] (3.2.1) Data preprocessing and preparation: In order to improve the randomness and coverage of gradient calculation, we preprocess the validation set data. First, randomly adjust the data order through the shuffle operation to ensure the diversity of input samples and avoid the calculation results being affected by a specific sample distribution. Subsequently, select the samples of the first 6 batches from the randomly shuffled validation set as the input data for gradient calculation to ensure the efficiency and representativeness of the calculation. Let the selected sample set be where x i is the input sample, y i is the corresponding label, and N is the total number of selected samples;
[0081] (3.2.2) Forward propagation to calculate the loss value: For each sample (x i , y i ) ∈ D, the input data x i is propagated forward through the neural network to obtain the predicted output
[0082]
[0083] where f(·; W) represents the neural network model with weights W, and W is the set of all parameters of the model. Subsequently, the loss value L i of each sample is calculated. The cross-entropy loss function is selected in the present invention:
[0084]
[0085] C represents the number of classes, y i,k and represent the true label and predicted probability of the k-th class of the i-th sample, respectively.
[0086] (2.2.3) Backward propagation to calculate the gradient and sum:
[0087]
[0088] where W l i represents the weight parameter of the i-th filter in the l-th layer, is the activation output of the i-th filter in the l-th layer, represents the gradient of the b-th sample with respect to W l i . By summing the gradient values, the incremental information of the filter during the training process can be quantified, thereby evaluating its importance.
[0089] Example 3: The overall implementation steps of the model compression method proposed in this example are the same as those in Example 1. Now, the candidate compression scheme designed in step (4.3) is further described in detail:
[0090] (4.3.1) Calculate the target FLOPs: According to the set compression rate p, calculate the target FLOPs value:
[0091]
[0092] FLOPs target = (1 - p) * FLOPs
[0093] Among them, p is the compression ratio, c i-1 is the input channel, c i is the output channel, w i *h i is the spatial size of the output feature map, K i is the convolution kernel size of the i-th convolutional layer.
[0094] (4.3.2) Pruning is performed according to FLOPs target : Calculate the FLOPs of each layer l , and allocate the target FLOPs to each layer:
[0095]
[0096] Among them, FLOPs l is the computational complexity of the l-th layer, and FLOPs l,target is the target computational complexity of each layer.
[0097] According to the target FLOPs of each layer, calculate the number of filters N that need to be subtracted from each layer prune,l :
[0098]
[0099] Among them, is the computational amount of the i-th filter in the l-th layer. On this basis, according to the importance score of the filter and the computational requirements of each layer, select the least important filter and set its bit width to 0 bit.
[0100] (4.3.3) Allocate quantization bit widths: For other filters that have not been pruned, we introduce two threshold parameters ρ1 and ρ2, and allocate quantization bit widths according to the global importance score of each filter. Specifically, according to the sorted importance score set {S↓}, allocate quantization bit widths to each filter so that filters with higher importance get larger bit widths:
[0101]
[0102] Among them, represents the quantization bit width allocated to the i-th filter in the l-th layer. To simplify the calculation and reduce the complexity in the quantization operation, for the same convolutional layer l, select the maximum quantization bit width of all filters in this layer as the unified quantization bit width of this layer. The specific calculation method is:
[0103]
[0104] Among them, c l is the number of filters in the l-th layer, B lis the layer width of the l-th layer, representing the maximum quantization bit width among all filters in this layer. This approach ensures that filters in each layer use the same bit width during quantization, thus simplifying the network structure.
[0105] Example 4: The overall implementation steps of the model compression method proposed in this example are the same as those in Example 1. Now, the optimization process of the affine transformation parameter pair in step (4.4) will be further described in detail:
[0106] (4.4.1) Define the optimization objective:
[0107]
[0108] where γ and β are the affine transformation parameter pair, is the compression scheme, and Accuarcy val is the accuracy of the lightweight model obtained by adopting this compression strategy on the validation set.
[0109] (4.4.2) Initialize the parameters: During the compression process, we first create a "sample pool" to store the candidate compression strategies (i.e., different γ_β parameter pairs) of the model and their corresponding evaluation results. Its size is initialized to P×P and filled with several initial samples randomly.
[0110] (4.4.3) Iterative learning: Assume the total number of iterations is E. Each iteration attempts to find a better compression strategy by updating γ and β. At the beginning of each iteration, set γ = 1 and β = 0 as the initial hyperparameters.
[0111] In each iteration, first randomly select S samples from the sample pool and select the optimal γ_β from these samples. Set the mutation rate as u, and randomly select u% of the layers in the network for mutation (i.e., adjust the parameters of the layers) to change the compression effect of the model. Specifically, for each layer l, calculate the standard deviation of the filter weights of this layer as an index to measure the complexity of this layer:
[0112]
[0113] The larger the standard deviation, the greater the change in the weights of this layer, and the more significant the effect after mutation. According to the calculated standard deviation, adjust γ l and β l as follows:
[0114]
[0115] Calculate the global importance based on the updated γ and β parameters and sort them. According to the sorting results, compress and fine-tune the model. Calculate the accuracy Acc of the compressed model using the validation set. The higher the accuracy, the more effective the current compression strategy is. According to the current evaluation results, add the new (γ, β) parameters and Acc to the sample pool and replace the earliest sample in the pool.
[0116] Example 5: The overall implementation steps of the model compression method proposed in this example are the same as those in Example 1. Now refer to Figure 4 , and based on Examples 1-4, give a more detailed example to further illustrate the implementation process of the present invention:
[0117] Step 1. Preprocess the dataset and initialize the model. The specific steps are as follows:
[0118] The first step is to obtain the original dataset from the public dataset where x i is the input data and y i is the corresponding label. After the original data is divided, the ratio of the training set to the validation set is set to 80%:20%, and the data consistency is ensured through standardization processing. Specifically, the training set images will undergo a series of random enhancement operations, including randomly cropping to a size of 256×256, randomly flipping horizontally, and normalizing the pixel values. The standardization formula is
[0119]
[0120] where μ and σ are the mean and standard deviation of the training dataset. At the same time, to ensure the consistency of the validation set processing, the validation set images are center-cropped to maintain a size of 256×256 and are subjected to the same normalization and tensorization processing. The above operations ensure the consistency of the data distributions of the training set and the validation set, providing high-quality data input for subsequent model training.
[0121] The second step is to construct a deep convolutional neural network (CNN) model and design a network structure including multiple convolutional layers, pooling layers, and fully connected layers to meet the requirements of a specific task.
[0122] The third step is to use the preprocessed training set data to initialize the training of the model, providing a reasonable starting point for subsequent compression operations. The initial learning rate set in the present invention is 0.001, the weight decay is 5e-4, and the batch size is 128.
[0123] Step 2. Rank information of feature maps: At each layer of the neural network, after passing the input data to the network through forward propagation, extract the feature maps of each layer These feature maps can be regarded as activation matrices after convolutional operations, where each channel corresponds to a two-dimensional matrix representing the feature information extracted by a specific convolutional kernel. To further evaluate the expressive power of these features, we perform low-rank decomposition on the feature map matrices of each channel to approximately estimate the rank information of the feature maps.
[0124] Low-rank decomposition decomposes each feature map matrix into multiple principal components and represents the original feature map as a weighted sum of the principal components. Given a feature map matrix it is decomposed into the following form using Singular Value Decomposition (SVD):
[0125]
[0126] where U and V are orthogonal matrices containing the left and right singular vectors respectively, and ∑ is a diagonal matrix containing the singular values σ1, σ2, …, σ r . A threshold σ is preset, and all singular values greater than this threshold are regarded as valid principal components. The rank of the feature map is the number of non-zero singular values greater than this threshold.
[0127] Step 3. Incremental information of the model:
[0128] First step, calculate the loss value through forward propagation. Specifically, for each sample (x i , y i ) ∈ D, the input data x i is propagated forward through the neural network to obtain the predicted output
[0129]
[0130] where f(·; W) represents the neural network model with weights W, and W is the set of all parameters of the model. Subsequently, calculate the loss value L i of each sample. The cross-entropy loss function is selected in the present invention:
[0131]
[0132] where C represents the number of classes, and y i,k and represent the true label and predicted probability of the k-th class of the i-th sample respectively.
[0133] Second step, calculate the gradient through backpropagation and sum them up:
[0134]
[0135] where W l iDenote the weight parameter of the \(i\)-th filter in the \(l\)-th layer. is the activation output of the \(i\)-th filter in the \(l\)-th layer. Denote the gradient of the \(b\)-th sample with respect to \(W\) l i . By summing up the gradient values, the incremental information of the filter during the training process can be quantified, thereby evaluating its importance.
[0136] To ensure the representativeness of the results, we randomly shuffle the validation set and select the samples from the first 6 batches of the shuffled dataset for gradient calculation. This method can ensure that the calculation results have good diversity and reflect the overall performance of the model.
[0137] Step 4. Initialize the affine transformation parameter pair: Before the optimization starts, first create a sample pool. The sample pool is used to store the affine parameter pairs \((\gamma,\beta)\) generated in each iteration and the corresponding accuracy value Acc. The size of the sample pool is set to \(S\), and part of the samples are randomly initialized as the starting samples. Assume the total number of iterations is \(E\), and in each iteration, an attempt is made to find a better compression strategy by updating \(\alpha\) and \(k\). At the beginning of each iteration, set \(\gamma = 1\) and \(\beta = 0\) as the initial hyperparameters.
[0138] Step 5. Global importance score of the filter:
[0139] First step, according to the rank information and incremental information obtained in Step 2 and Step 3, evaluate the importance of the filters in each layer. Specifically, by performing a weighted sum of the feature map rank information and incremental information, the importance score of the filter is obtained:
[0140]
[0141] Second step, in order to unify the importance measurement criteria for filters between different layers, the present invention introduces a global importance score. For the filter \(i\) in the convolutional layer \(l\), its global importance score can be expressed as:
[0142]
[0143] Step 6. Candidate compression scheme: Consider pruning as per-channel quantization with 0 bits, and at the same time, limit the quantization bitwidth to \(\{0, 2, 4, 8\}\). First, perform 0-bit quantization on the model according to the preset compression rate. For other filters, two threshold parameters \(\rho_1\) and \(\rho_2\) are introduced. According to the sorted importance score set \(\{S\downarrow\}\), assign the corresponding quantization bitwidth to each filter, so that the more important the filter, the larger the bitwidth obtained. Specifically, when , the quantization bitwidth of the corresponding filter is 2. When , the quantization bitwidth of the corresponding filter is 4. When When the quantization bit width of the corresponding filter is 8. To simplify the calculation, for the same convolutional layer, we select the maximum bit width obtained by the filters in this layer as the layer bit width:
[0144]
[0145] Step 7. Evaluate the number of iteration rounds: The number of iteration rounds of the present invention is fixedly set to 400 rounds, and an optimization operation is performed once in each round. By gradually adjusting the affine hyperparameters (γ, β) within a fixed number of iterations, it is ensured that the efficient compression of the model is achieved under the limited computing resources. This fixed-round design can provide a stable optimization process, avoid the uncertainty that may be introduced by dynamically adjusting the number of rounds, and at the same time ensure the controllability and consistency of the optimization process.
[0146] Step 8. Optimization by evolutionary algorithm: The present invention adopts a parameter optimization method based on evolutionary algorithm, and the specific steps are as follows:
[0147] The first step is to randomly select S samples from the sample pool as the candidate set for the current optimization. Based on the performance evaluation of these samples, the optimal pair of affine parameters (γ, β) is selected as the starting point for this round of optimization, providing a basis for subsequent mutation and compression.
[0148] The second step is to set the mutation rate as u, that is, randomly select u% of the layers in the network for mutation operations to adjust the layer parameters so as to change the compression effect of the model. Specifically, for each layer l, calculate the standard deviation of the filter weights of this layer as an index to measure the complexity of this layer:
[0149]
[0150] The larger the standard deviation, the more significant the change in the weight distribution of this layer, so it may have a more obvious impact on the model performance after the mutation operation. Based on the calculated standard deviation std l , adjust the affine parameters γ l and β l of each layer:
[0151]
[0152] The third step is to recalculate the global importance score according to the updated affine parameters (γ, β) and sort all the filters. The sorted importance scores will directly guide the adjustment of pruning and quantization strategies to achieve the compression and fine-tuning of the model. Subsequently, use the validation set to evaluate the compressed model and calculate the accuracy index Acc. The higher the model accuracy, the better the current compression strategy.
[0153] Step 4: According to the current evaluation results, add the new (γ,β) parameters and their corresponding accuracy Acc to the sample pool, and at the same time replace the earliest sample in the pool to ensure the update and diversity of the sample pool.
[0154] Step 9. Compression and fine-tuning: To alleviate the possible performance loss caused by pruning and quantization, a phased compression and fine-tuning strategy is adopted for the model.
[0155] Step 1: According to the cross-layer importance ranking and the preset pruning rate, determine the filters that need to be 0-bit quantized, that is, the filters that need to be pruned. After completing the pruning operation, to recover the performance loss caused by pruning, the model is fine-tuned, and the fine-tuning process lasts for 300 epochs;
[0156] Step 2: For the remaining filters, introduce two threshold parameters, and assign bit widths to them according to the importance scores. Larger bit widths are preferentially assigned to more important filters, while smaller bit widths are assigned to less important filters. To further reduce the computational complexity, select the filter with the largest bit width in each convolutional layer as the overall bit width of the layer. Subsequently, perform mixed-precision quantization operations on the weight parameters of the convolutional layer and the activation layer respectively:
[0157]
[0158] Among them, And after that, fine-tune the model for another 300 epochs.
[0159] The present invention comprehensively evaluates the importance of filters by combining the gradient information and rank information of the network. In terms of incremental information, the present invention uses the gradient of the model to capture the dynamic changes of each filter, reflecting the sensitivity of the filter to the loss function and the learning ability. At the same time, the rank information analyzes each layer of feature maps through low-rank decomposition, evaluates the role of filters in information expression, and provides supplements from the perspective of static features. In addition, the present invention introduces an evolutionary algorithm, and through the optimization learning of hyperparameters, calculates the affine parameters that can measure the global importance. This evolutionary algorithm adjusts the parameters in multiple iterative processes, gradually improving the model's ability to judge the importance of filters, thereby realizing the global ranking of filters. Through the comprehensive ranking of the importance of each filter, the present invention can effectively perform pruning and quantization while ensuring the model accuracy, optimize the computational complexity, and improve the model performance. This multi-dimensional comprehensive evaluation method not only improves the efficiency of model compression but also makes the compression process more flexible and intelligent.
[0160] The following further illustrates the effect of the present invention through simulation experiments:
[0161] 1. Simulation conditions:
[0162] The first dataset used in the experiment is the NWPU-RESISC45 (NWPU-45) dataset, which is an open benchmark dataset created by Northwestern Polytechnical University and belongs to the REmote Sensing Image Scene Classification (RESISC) series. This dataset contains 31,500 images, each with a pixel size of 256*256, covering 45 different scene categories, and each category contains 700 images.
[0163] The second dataset used in the experiment is the UC Merced Land-Use (UCML-21) dataset, created by the University of California, Merced, and is an open benchmark dataset widely used in remote sensing image classification tasks. This dataset contains high-resolution remote sensing images of 21 land use categories, with 100 images in each category. These images were manually extracted from the US Geological Survey's National Map Urban Area Image Collection, with a pixel resolution of up to 0.3 meters and a pixel size of 256*256.
[0164] 2. Simulation content:
[0165] The simulation experiment of the present invention uses the method proposed by the present invention and various existing model compression methods. Under the same simulation conditions, experiments are carried out on three network models respectively. By using a variety of advanced compression comparison algorithms to compress the network, and comparing the accuracy performance of the method of the present invention and the comparison algorithms at the same compression rate, and calculating relevant performance indicators at the same time. The experimental results show that compared with the existing compression algorithms, at the same compression rate, the method of the present invention can achieve higher accuracy.
[0166] 3. Analysis of simulation results:
[0167] The first experiment was carried out on the NWPU-45 dataset. Refer to Figure 2 , and a detailed description of Simulation Experiment 1 of the method of the present invention and existing compression algorithms is given. Among them, Figure 2 (a) shows the results after compression using various methods on the VGG-16 network, Figure 2 (b) shows the performance of VGG-19 under multiple compression algorithms, Figure 2 (c) shows the compression results of the ResNet-56 network, Figure 2 (d) compares the changes in compression accuracy before and after combining incremental information for three different networks on the NWPU-45 dataset. Through these results, it can be intuitively seen the impact of different compression algorithms on model accuracy and computational performance.
[0168] See Figure 2 (a), Figure 2(a) shows the experimental results of using the method of the present invention and existing compression algorithms NPIC, LeGR, HRank, and FSKD on the VGG-16 network. It can be seen that when the FLOPs are significantly reduced by 50.13%, the method of the present invention can still effectively retain the key features in the network, ensuring that important information is retained. Although the computational complexity of the model is significantly reduced, compared with traditional compression methods, the method of the present invention can still maintain the highest accuracy. Figure 2 (b) shows the experimental results of the VGG-19 network under different compression methods. In this experiment, the method of the present invention improves the accuracy by 0.19%, exceeding the traditional method, and also shows significant advantages in terms of computational resource consumption. Compared with the Network Slimming based on channel scaling factors and the NPIC method based on interpretable CNN, the FLOPs of the method of the present invention decrease from 51.200G to 25.585G, but it can still maintain the best performance. Figure 2 (c) shows the performance of the ResNet-56 network under different compression methods. The experimental results show that after compression, the method of the present invention still maintains a high classification accuracy. Compared with other compression methods, while the FLOPs of the method of the present invention are reduced by 50.18%, the accuracy is improved by 0.27%. Figure 2 (d) shows the comparison of the accuracy before and after combining incremental information on three different networks (VGG-16, VGG-19, and ResNet-56). The experimental results clearly show that the introduction of incremental information significantly improves the compression effect of the model. After combining incremental information, the model can more accurately evaluate the importance of filters, thereby optimizing the pruning and quantization strategies, effectively alleviating the problem of accuracy loss in traditional methods.
[0169] The second experiment was conducted on the UCML-21 dataset. Refer to Figure 3 , and a detailed description of the simulation experiment 2 of the present method and existing compression algorithms is given. Among them, Figure 3 (a) shows the results after compressing the VGG-16 network using multiple methods, Figure 3 (b) shows the performance of VGG-19 under multiple compression algorithms, Figure 3 (c) shows the compression results of the ResNet-56 network, Figure 3 (d) compares the changes in accuracy before and after combining incremental information on three different networks.
[0170] See Figure 3 (a), Figure 3(a) shows the experimental results of using the method of the present invention and existing compression algorithms LeGR, HRank, and FSKD on the VGG-16 network. After compression, the method of the present invention successfully reduces the FLOPs from 40.088G to 20.000G, a decrease of 50.11%, while the Top-1 accuracy increases by 2.62%. In contrast, for other methods such as HRank and LeGR, the reduction in FLOPs is similar, but the increase in accuracy is smaller, at 1.43% and 1.94% respectively. Figure 3 (b) shows the experimental results of the VGG-19 network under different compression methods. After compression, the FLOPs of the method of the present invention decrease by 50.06% and the Top-1 accuracy increases by 1.06%. Compared with the NPIC and Network Slimming methods, the method of the present invention demonstrates higher compression efficiency and accuracy retention ability. Figure 3 (c) shows the performance of the ResNet-56 network under different compression methods. The experimental results show that the FLOPs of the method of the present invention decrease by 50.09% while the Top-1 accuracy increases by 1.91%. In contrast, for the LeGR and HRank methods, with a similar reduction in FLOPs, the accuracy increases by 0.53% and 0.46% respectively. Figure 3 (d) respectively shows the comparison of the compression accuracies before and after combining incremental information for the VGG-16, VGG-19, and ResNet-56 networks on the UCML-21 dataset. The experimental results show that the introduction of incremental information significantly improves the compression effect of the model. By utilizing incremental information, the model can more accurately measure the importance of filters, thereby optimizing the strategy design of pruning and quantization, effectively reducing the loss of accuracy in traditional compression methods.
[0171] To further comprehensively demonstrate the superiority of the method of the present invention, multiple indicators such as Params, Params reduction rate, BOPs (Bitwise Operations Per Second), and BOPs compression rate are calculated respectively, and the performance of each method is compared and evaluated from different perspectives. These indicators can more comprehensively reflect the performance of the model in terms of resource occupancy and computational efficiency, providing more powerful support for verifying the comprehensive performance of the method of the present invention. Table 1 shows the detailed performance indicator comparison of the experimental results of the method of the present invention and various existing compression methods on the NWPU-45 dataset.
[0172]
[0173]
[0174] By analyzing the performance metrics of the compression results of the NWPU dataset in Table 1, it can be found that compared with other compression methods, the present invention shows significant advantages in terms of accuracy retention, computational complexity reduction, and resource optimization. This method achieved a Top-1 accuracy of 94.41% in the VGG-16 network, with only a 0.16% decrease, while the BOPs and the number of parameters decreased by 40.05% and 77.23% respectively. In the VGG-19 network, it achieved a Top-1 accuracy of 94.36%, an increase of 0.19%, while the BOPs and the number of parameters decreased by 40.00% and 13.74% respectively. In the ResNet-56 network, it achieved a Top-1 accuracy of 92.11%, an increase of 0.27%, while the BOPs and the number of parameters decreased by 40.00% and 53.67% respectively. These results fully demonstrate that this method can not only optimize the model in terms of computational complexity but also achieve greater advantages in terms of storage and hardware requirements, further verifying the comprehensive performance and practical application value of the method of the present invention.
[0175] Table 2 shows the experimental results on the UCML-21 dataset, listing the detailed performance metric comparisons between the present invention and various existing compression methods on the UCML-21 dataset.
[0176]
[0177] By analyzing the performance metrics of the compression results of the UCML-21 dataset in Table 2, it can be found that compared with other compression methods, the present invention shows the best accuracy performance on multiple networks, and the TOP-1 accuracy is much greater than that of other compression algorithms, ensuring the efficient retention of the performance of the model after compression. At the same time, in terms of the reduction rate of FLOPs and the compression rate of BOPs, the method of the present invention also shows a greater compression effect. Compared with traditional compression methods, the method of the present invention can achieve the greatest reduction in FLOPs and BOPs, greatly reducing the computational complexity and the occupation of hardware resources, and is particularly suitable for devices with limited computing resources and edge computing scenarios. In addition, on VGG-16 and ResNet-56, the compression rate of the number of parameters of the present invention far exceeds that of other methods, significantly reducing the memory occupation and providing better support for the deployment of the model on embedded and mobile devices.
[0178] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the users or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0179] The above simulation analysis proves the correctness and effectiveness of the method proposed by the present invention.
[0180] The parts not described in detail in the present invention belong to the common general knowledge of those skilled in the art.
[0181] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Obviously, for those skilled in the art, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these modifications and changes based on the idea of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A remote sensing image classification model compression method based on incremental information guidance, characterized in that: The steps include: (1) Obtain a remote sensing image dataset and preprocess it to obtain a preprocessed image x as the model input; (2) The convolutional neural network (CNN) model is used as the remote sensing image classification model. The model is pre-trained using the pre-processed image x, that is, the model weights are optimized through forward propagation and back propagation algorithms to obtain the filter weights W of each layer. l ′, get the pre-trained model; (3) Evaluate the importance of filters in the pre-trained model: (3.1) The preprocessed image data is forward propagated through the pre-trained model, and then the feature map output by the model is low-rank decomposed to obtain its rank information r: in, is the feature map output by the ith filter in the lth convolutional layer, l = 1, 2, ..., D, D is the number of convolutional layers contained in the pre-trained model; s j represents the jth singular value, u j represents the jth left singular vector, v j T represents the jth right singular vector; (3.2) Use some preprocessed images in the validation set as input data for calculating the gradient, calculate the loss value through forward propagation, and use back propagation to accumulate the gradient information of each image to obtain the gradient statistics of each filter Where y is the label; (3.3) According to the rank information obtained in step (3.1) and the gradient information obtained in step (3.2), the importance score of each layer filter is calculated: in, represents the importance score of the i-th filter in the l-th layer; (4) Rank the filters by global importance and generate the final global importance ranking result: (4.1) Initialize the affine transformation parameter pair, which is the parameter to be learned for the lth convolutional layer of each pre-trained model, including the scaling factor γ l and the constant term β l ; (4.2) The global importance score of the i-th filter in the l-th convolutional layer is calculated according to the following formula: And sort the filters by their global importance scores to get the ranking results {S↓} from high to low importance scores: (4.3) Design candidate compression schemes based on the sorting results: The pruning operation is defined as 0-bit quantization, and the retained filters are discretely quantized, and the allowed bit width set is set to {2, 4, 8}; redundant filters are removed according to {S↓}, and two threshold parameters, low threshold ρ1 and high threshold ρ2, are introduced for the retained filters, i.e., the non-0 bits, and the bit width is dynamically allocated to them according to the importance score of the current filter; (4.4) Compress the preprocessed model according to the compression scheme in step (4.3), then calculate the accuracy of the compressed model on the validation set, and use the regularized evolution algorithm to optimize the parameter γ l and β l , the goal is to maximize the accuracy of the validation set and obtain the optimized affine transformation parameters; (4.5) Calculate the final global importance score of each filter based on the optimized affine transformation parameters, and sort the scores from high to low to generate the final global importance sorting result; (5) Obtaining a lightweight model through staged compression: (5.1) According to the final global importance ranking result and the preset pruning rate, determine the filter that needs to be 0-bit quantized, that is, the filter that needs to be pruned; after completing the pruning operation, fine-tune the model; (5.2) For the retained filters, i.e., the non-zero bits, two threshold parameters, low threshold ρ1 and high threshold ρ2, are introduced, and the bit width is dynamically allocated to them according to the importance score of the current filter. Then, the weight parameters and activation values of the convolutional layer are subjected to mixed precision quantization operations respectively; (5.3) After the mixed precision quantization operation is completed, the model is fine-tuned again to obtain a compressed and lightweight remote sensing image classification model.
2. The method according to claim 1, characterized in that: The preprocessed image x in step (1) is obtained according to the following steps: (1.1) Obtain the original remote sensing image W from the public remote sensing image dataset W×H×C Construct a sample set and divide it into training set and validation set; Where W×H is the spatial resolution of the image, and C is the number of channels; (1.2) Preprocess the images in the sample set to obtain the preprocessed image x, which is used as the model input image; the preprocessing is specifically to randomly crop the training set images to 256 pixels and randomly flip them horizontally; then convert them into tensors and standardize them, and the standardization uses a preset mean and standard deviation; at the same time, the validation set images are adjusted to 256 pixels by center cropping, and then converted into tensors and standardized.
3. The method according to claim 1, characterized in that: The CNN models in step (2) include VGG-16, VGG-19 and ResNet-56.
4. The method according to claim 1, characterized in that: Step (2) optimizes the model weights through forward propagation and back propagation algorithms to obtain the filter weights W of each layer l ′, and is implemented as follows: O l (x)=f(W l ·O l-1 (x)+b l ), Among them, O l represents the feature map of the lth layer, f(·) is the activation function, W l and b l are the original weight matrix and bias of the lth layer, W l ′ and b l ′ represents the weight matrix and bias after optimization of the lth layer, δ l is the error term, L is the loss function, and η is the learning rate.
5. The method according to claim 1, characterized in that: The input data for calculating the gradient in step (3.2) is image data obtained by randomly selecting 90 to 100 preprocessed images in the validation set.
6. The method according to claim 1, characterized in that: The dynamic allocation of bit width described in steps (4.3) and (5.2) is performed as follows: When , the quantization bit width of the current filter is assigned to 2; when When , the quantization bit width assigned to the current filter is 4; when , the quantization bit width of the current filter is assigned to 8; for the same convolution layer l of the model, the maximum bit width of the filter in the layer is used as the uniform bit width B of the layer l : Among them, c l Indicates the number of filters contained in the lth convolutional layer of the model.
7. The method according to claim 1, characterized in that: The regular evolution algorithm described in step (4.4) is replaced by a genetic algorithm, a differential evolution algorithm or a Bayesian optimization algorithm.
8. The method according to claim 1, characterized in that: The fine-tuning described in steps (5.1) and (5.3) is to use the preprocessed images to perform the same pre-training process as step (2) on the compressed model, and optimize the weights of the compressed model through forward propagation and back propagation algorithms, and the training process is no less than 300 rounds.
9. The method according to claim 1, characterized in that: In step (5.2), mixed precision quantization operations are performed on the weight parameters and activation values of the convolutional layer, respectively, as shown below: Where Q(·) is the quantization function, A l is the activation value of the lth layer.
Citation Information
Cited By
Target detection deep learning network quantification method
CN120562490A
Building feature extraction and mapping method based on BIM technology
CN120951452A
Behavior adaptive continuous identity authentication method and system
CN121389091A
Risk perception neural network quantification method and device for aero-engine health management
CN121920453A