Method for Fine-Tuning Basic Model Parameters for Remote Sensing Scene Classification Task

By self-supervised pre-training of the ConvNeXt backbone network and introducing an efficient quantization adapter module and a context-aware prompt module, the basic model parameters efficient fine-tuning method for remote sensing scene classification tasks is solved, and the high computing cost and memory requirements of large-scale pre-trained basic models in remote sensing scene classification tasks are achieved, and the efficient and highly adaptable model fine-tuning effect is achieved.

CN118212547BActive Publication Date: 2025-06-10HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410319866.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-20
Publication Date
2025-06-10
Estimated Expiration
2044-03-20

AI Technical Summary

Technical Problem

The basic model of large-scale pre-training has high computational cost and memory requirements in the fine-tuning process of remote sensing scene classification tasks, poor adaptability, and ignores the prior information and spatiotemporal correlation in remote sensing data.

Method used

A basic model parameter efficient fine-tuning method for remote sensing scenario classification tasks is proposed, and the ConvNeXt backbone network is pre-trained by self-supervising, and an efficient quantization adapter module and context-aware prompt module are introduced during the fine-tuning process, which only updates less than 1% of the model parameters.

Benefits of technology

It significantly reduces computing and storage costs, maintains or even improves model performance, and can effectively adapt to the prior information and spatiotemporal correlation of remote sensing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118212547B_ABST
    Figure CN118212547B_ABST
Patent Text Reader

Abstract

Method for fine-tuning basic model parameters for remote sensing scene classification tasks. The present invention relates to a method for fine-tuning basic model parameters. The purpose of the present invention is to solve the problems of high computational cost and memory requirements, poor adaptability during fine-tuning of large-scale pre-trained basic models, and the fact that existing fine-tuning methods ignore the prior information and spatio-temporal correlation in remote sensing data. The process is as follows: 1: Obtain a pre-trained ConvNeXt backbone network; 2: Construct a scene classification network model; the scene classification network model includes: a ConvNeXt backbone network, an adapter module, a context-aware prompt module, and a scene classification task head; 3: Fine-tune the parameters of the scene classification network model using labeled training samples. When fine-tuning, freeze all parameters except those of the adapter module, the context-aware prompt module, and the scene classification task head to obtain a fine-tuned scene classification network model. The present invention is used in the field of remote sensing image processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of remote sensing image processing, relates to remote sensing scene classification, and specifically relates to an efficient fine-tuning method for the basic model parameters for remote sensing scene classification tasks. Background Art

[0002] With the rapid development of earth observation and deep learning technologies, large-scale vision models have gradually dominated the remote sensing field in the past few years. Following the principle of transfer learning, downstream tasks can benefit from the knowledge learned by pre-trained models. Early pre-trained models such as CNNs and Transformers can be pre-trained through supervised learning or self-supervised learning. Recently, in order to unify visual understanding tasks, advanced pre-training paradigms such as MoCo and MAE have been proposed, and these models are collectively referred to as basic models.

[0003] Basic models generally follow the "pre-train & fine-tune" paradigm. In this paradigm, the basic model is initially pre-trained on a large amount of data through self-supervised learning to learn general representations. Subsequently, fine-tuning is used to transfer the understanding ability of the pre-trained model to achieve satisfactory performance on downstream tasks.

[0004] As a mainstream fine-tuning strategy, full fine-tuning requires maintaining different parameter sets for each task and deployment scenario, and the size of these parameter sets is equivalent to that of the pre-trained basic model. These parameters operate independently and cannot be shared, making them resource-intensive in real-world applications. In addition, as downstream tasks and deployment environments continue to expand, the storage cost increases linearly, especially when deploying large-scale vision basic models such as ViT-Huge (632M parameters) on mobile systems. In contrast, although only fine-tuning the downstream task head can avoid updating the entire backbone model, it usually results in unsatisfactory performance.

[0005] It is worth noting that inspired by natural language processing (NLP), an emerging solution for pre-trained vision models is to replace full fine-tuning with parameter-efficient fine-tuning (PEFT). This method only adjusts a small number of trainable parameters while freezing most of the parameters shared by multiple downstream tasks. By fine-tuning a more limited set of parameters, not only is the complexity of optimization reduced, but overfitting is effectively avoided when generalizing large-scale basic models to specific target datasets, ultimately achieving performance equivalent to or even better than full fine-tuning. However, the challenges of considering prior information in remote sensing data samples and spatio-temporal correlations in multi-scale features during the fine-tuning process are still significant. Summary of the Invention

[0006] The object of the present invention is to solve the problems of high computational cost, large memory requirements, poor adaptability during the fine-tuning of large-scale pre-trained basic models, and the neglect of prior information and spatio-temporal correlations in remote sensing data by existing fine-tuning methods, and to propose a basic model parameter fine-tuning method for remote sensing scene classification tasks.

[0007] The specific process of the basic model parameter fine-tuning method for remote sensing scene classification tasks is as follows:

[0008] Step 1: Unlabeled remote sensing images are used for self-supervised pre-training of the ConvNeXt backbone network to obtain a pre-trained ConvNeXt backbone network;

[0009] Step 2: Construct a scene classification network model; the specific process is as follows:

[0010] The scene classification network model includes: a ConvNeXt backbone network, an adapter module, a context-aware prompt module, and a scene classification task head;

[0011] The adapter module is inserted into each convolutional block in the pre-trained ConvNeXt backbone network;

[0012] The context-aware prompt module is inserted after each convolutional block in the pre-trained ConvNeXt backbone network;

[0013] The scene classification task head is inserted after the pre-trained ConvNeXt backbone network;

[0014] Step 3: Use labeled training samples to fine-tune the parameters of the scene classification network model. During fine-tuning, all parameters except those of the adapter module, the context-aware prompt module, and the scene classification task head are frozen to obtain a fine-tuned scene classification network model.

[0015] Beneficial effects

[0016] The present invention proposes an efficient basic model parameter fine-tuning method for remote sensing scene classification tasks, which reduces the number of parameters that need to be updated during the fine-tuning of the basic model, reduces computational cost and memory requirements, and at the same time maintains or even improves the performance of the model.

[0017] Compared with the prior art, the present invention has obvious beneficial effects. By means of the above technical solutions, the method provided by the present invention can achieve considerable technological progressiveness and practicality, and has broad industrial utilization value. It has at least the following beneficial effects:

[0018] (1) The present invention can effectively reduce the computational and storage costs generated during the fine-tuning of the basic model in remote sensing scene classification tasks.

[0019] (2) By updating only less than 1% of the model parameters, the present invention can make the performance of the base model after fine-tuning close to or even exceed that of the full fine-tuning method. The method proposed by the present invention is simple and effective, and has great practical application value;

[0020] (3) The present invention designs an efficient fine-tuning method for the parameters of the base model for remote sensing scene classification tasks, which significantly reduces the computational and storage requirements during the fine-tuning of the base model used for remote sensing scene classification tasks. Specifically, the parameter-efficient fine-tuning method proposed by the present invention integrates two complementary key modules: an efficient quantization adapter module and a context-aware prompt module. The efficient quantization adapter module is specifically designed to learn task-specific feature representations during the fine-tuning process, and effectively reduces the number of model parameters without sacrificing performance through its innovative quantization mechanism. At the same time, the context-aware prompt module endows the base model with the ability to dynamically generate trainable and context-related prompts, significantly enhancing the model's perception and processing ability of context information. Through condition-related prompts, a new way of context understanding is provided for the model, which is particularly important when dealing with complex remote sensing scenes, as it enables the model to better adapt to different tasks and environmental changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is the flowchart of the implementation of the present invention;

[0022] Figure 2 is the framework diagram of the efficient parameter fine-tuning of the base model designed by the present invention;

[0023] Figure 3 is the convolutional block diagram of the inserted adapter module EQAM designed by the present invention;

[0024] Figure 4 is the schematic diagram of the fine-tuning of the scene classification model on the scene classification task. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] DETAILED DESCRIPTION OF THE EMBODIMENT 1: The specific process of the method for fine-tuning the parameters of the base model for remote sensing scene classification tasks in this embodiment is as follows:

[0026] The following further describes the method of the present invention in detail with reference to the drawings and embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0027] Step 1: Self-supervised pre-training of the ConvNeXt backbone network with a large number of unlabeled remote sensing images to obtain a pre-trained ConvNeXt backbone network;

[0028] Step 2: Construct a scene classification network model; the specific process is as follows:

[0029] The scene classification network model includes: a ConvNeXt backbone network, an efficient quantization adapter module, a context-aware prompt module, and a scene classification task head;

[0030] The efficient quantization adapter module is inserted into each convolutional block in the pre-trained ConvNeXt backbone network;

[0031] The context-aware prompt module is inserted after each convolutional block in the pre-trained ConvNeXt backbone network to process the feature maps output by each convolutional block in the ConvNeXt backbone network;

[0032] The scene classification task head is inserted after the pre-trained ConvNeXt backbone network;

[0033] Step 3: Use the labeled training samples to fine-tune the parameters of the scene classification network model. When fine-tuning, freeze all parameters except those of the efficient quantization adapter module, the context-aware prompt module, and the scene classification task head to obtain a fine-tuned scene classification network model;

[0034] Step 4: The fine-tuned scene classification network model is used for subsequent remote sensing scene classification tasks.

[0035] Specific Embodiment 2: Different from Specific Embodiment 1, in Step 1, the ConvNeXt backbone network sequentially includes: a chunk head and N convolutional blocks;

[0036] The connection relationship of the ConvNeXt backbone network is: the image is sequentially input into the chunk head, the first convolutional block, the second convolutional block, until the Nth convolutional block, and the output feature of the Nth convolutional block is used as the output feature of the ConvNeXt backbone network;

[0037] Each convolutional block sequentially includes a two-dimensional depth convolutional layer with a kernel size of 7×7, a layer normalization layer LayerNorm, a two-dimensional convolutional layer with a kernel size of 1×1, an activation function GELU, a two-dimensional convolutional layer with a kernel size of 1×1, a layer scale LayerScale, and a path dropout Droppath;

[0038] The specific processing process of each convolutional block is: the input feature α is sequentially input into the two-dimensional depth convolutional layer with a kernel size of 7×7, the layer normalization layer LayerNorm, the two-dimensional convolutional layer with a kernel size of 1×1, GELU, the two-dimensional convolutional layer with a kernel size of 1×1, the layer scale LayerScale, and the path dropout Droppath, and the path dropout Droppath outputs the feature β;

[0039] The feature α and the feature β are element-wise added to obtain the feature γ, and the feature γ is used as the output feature of each convolutional block.

[0040] In step 1, the ConvNeXt backbone network is pre-trained in a self-supervised manner using a large number of unlabeled remote sensing images to obtain a pre-trained ConvNeXt backbone network. The specific process is as follows:

[0041] A sample set containing millions of unlabeled remote sensing images is constructed for the self-supervised pre-training of the ConvNeXt backbone network to learn a general and generalized feature representation;

[0042] Based on the unlabeled remote sensing image sample set, the ConvNeXt backbone network is pre-trained for masked image reconstruction. The optimization objective is to minimize the normalized pixel error between the reconstruction target and the original image at the masked positions, resulting in a pre-trained ConvNeXt backbone network.

[0043] Other steps and parameters are the same as those in the first specific implementation manner.

[0044] The third specific implementation manner: The difference between this implementation manner and the first or second specific implementation manner is that the efficient quantization adapter module EQAM in step 2 sequentially includes: a flattening layer Flatten, a first quantization linear layer Q-Linear, a ReLU activation function, a second quantization linear layer Q-Linear, a scaling scalar Scaling, and a reshape Reshape;

[0045] The specific processing process of the adapter module EQAM is as follows: Features are sequentially input into the flattening layer Flatten, the first quantization linear layer Q-Linear, the ReLU activation function, the second quantization linear layer Q-Linear, and the scaling scalar Scaling. The scaling scalar Scaling outputs feature B′. The element-wise addition of the output feature B′ of Scaling and the output feature of the flattening layer Flatten is performed to obtain feature B″. Feature B″ is subjected to a reshape Reshape process to obtain feature B, which is the output feature of the adapter module EQAM.

[0046] Other steps and parameters are the same as those in the first or second specific implementation manner.

[0047] The fourth specific implementation manner: The difference between this implementation manner and one of the first to third specific implementation manners is that the efficient quantization adapter module in step 2 is inserted into each convolutional block of the pre-trained ConvNeXt backbone network. The specific process is as follows:

[0048] The adapter module EQAM is inserted after the Droppath layer of each convolutional block in the pre-trained ConvNeXt backbone network to obtain a convolutional block with the adapter module EQAM inserted;

[0049] Each convolutional block of the inserted adapter module EQAM sequentially includes a two-dimensional depth convolutional layer with a convolutional kernel size of 7×7, a layer normalization layer LayerNorm, a two-dimensional convolutional layer with a convolutional kernel size of 1×1, an activation function GELU, a two-dimensional convolutional layer with a convolutional kernel size of 1×1, a layer scale LayerScale, a path dropout Droppath, and an adapter module EQAM;

[0050] The specific processing process of each convolutional block of the inserted adapter module EQAM is as follows: The input feature A is sequentially input into a two-dimensional depth convolutional layer with a convolutional kernel size of 7×7, a layer normalization layer LayerNorm, a two-dimensional convolutional layer with a convolutional kernel size of 1×1, GELU, a two-dimensional convolutional layer with a convolutional kernel size of 1×1, a layer scale LayerScale, a path dropout Droppath, and an adapter module EQAM, and the adapter module EQAM outputs the feature B;

[0051] The element-wise addition of the feature A and the feature B is performed to obtain the feature C, and the feature C is used as the output feature of each convolutional block.

[0052] Other steps and parameters are the same as those in one of the specific embodiments 1 to 3.

[0053] Specific embodiment 5: The difference between this embodiment and one of the specific embodiments 1 to 4 is that the specific processing process of each adapter module EQAM is as follows:

[0054] 1), Assume that the input feature map of the efficient quantization adapter module is represented as

[0055] where H, W, and C respectively represent the 1 height, width, and channel dimensions of F;

[0056] The feature map F 1 is flattened along the channel dimension to form a two-dimensional feature map

[0057]

[0058] where D = H×W, and Flatten() represents being flattened along the channel dimension;

[0059] 2), For simplicity, ignoring the bias parameter, the two-dimensional feature map is projected to obtain the projected feature map The expression is:

[0060]

[0061] where s represents the scaling scalar Scaling, and s is set to 1;

[0062] · is matrix dot product;

[0063] ReLU is an activation function;

[0064] is the learnable weight of the first quantization linear layer, is the learnable weight of the second quantization linear layer; c' is the hidden feature dimension, and c' is set to 8;

[0065] The learnable weights of the two quantization linear layers Q-Linear are set to and aiming to adjust the spatial size of the low-dimensional representation through downward projection and upward projection;

[0066] 3), Reshape into a three-dimensional feature map with the same spatial size as F 1

[0067]

[0068] Among them, the feature map generated by the efficient quantization adapter module is represented as Reshape means reshaping.

[0069] Other steps and parameters are the same as those in any one of the specific embodiments one to four.

[0070] Specific embodiment six: The difference between this embodiment and any one of the specific embodiments one to five is that the processing process of each quantization linear layer in the first quantization linear layer Q-Linear and the second quantization linear layer Q-Linear is as follows:

[0071] For the quantization linear layer Q-Linear in the EQAM module, the present invention applies clustering-based quantization to reduce the bit width of the linear layer weights;

[0072] 1), Introduce the bit width b to divide the weights of the quantization linear layer (continuous weight real number space) into B = 2 b non-overlapping discrete value sets j = 1, 2,..., B;

[0073] The bit width b is set to 8;

[0074] 2), Set the codebook, and the codebook is mapped to a set of numerical values {c 1 ,..., c j ,..., c B}, and there are B numbers in {c 1 ,..., c j ,..., c B}, and there are B numbers in {c 1 ​,...,c j ,...,c B} is preset;

[0075] j = 1, 2, …, B;

[0076] Each corresponds to a specific code c in the codebook j ;

[0077] During the quantization process, the present invention introduces a pre-set codebook. The set is associated with a codebook containing B codes {c 1 ,...,c j ,...,c B};

[0078] 3), introduce 1D clustering to minimize the quantization error:

[0079]

[0080] where, w i represents the i-th element of the weight matrix W of the quantization linear layer Q-Linear, and m represents the number of elements in the weight matrix W of the quantization linear layer Q-Linear; represents the quantized value of w i after being processed by the quantization function ;

[0081] It is assumed that the parameters in the weight matrix W follow a Gaussian distribution;

[0082] The processing process of the quantization function is as follows:

[0083] The quantization function quantizes all the values in into the code c j :

[0084]

[0085] where, j = 1, 2, …, B; represents the weight from the set ;

[0086] The present invention standardizes the weights according to the mean and variance;

[0087] 4), based on the set and the codes {c 1 ,...,c j ,...,c B}, for the i-th element w of the weight matrix W of the quantized linear layer Q-Linear i Perform standardization, quantize each standardized element, and restore the element to its original mean and variance through de-standardization;

[0088] The expression is:

[0089]

[0090]

[0091]

[0092] Among them, w i Represents the i-th element of the weight matrix W of the quantized linear layer Q-Linear;

[0093] μ represents the mean, MEAN() represents the mean operation, and m represents the number of elements in the weight matrix W of the quantized linear layer Q-Linear;

[0094] σ represents the variance, STD() represents the variance operation;

[0095] w′ i Indicates w i Normalized weights;

[0096] Indicates w i The values ​​after standardization and quantization;

[0097] Indicates w i The value after quantization;

[0098] In the entire quantization process, the only non-differentiable is the quantization operation To solve this problem, the present invention adopts a straight-through estimator (STE) to approximate the gradient;

[0099] 5) For w k ∈W, calculate the approximate gradient:

[0100]

[0101] Among them, w′ k Indicates w k After the normalization, the weight w k Represents the kth element of the weight matrix W of the quantized linear layer Q-Linear.

[0102] Other steps and parameters are the same as those in any one of the first to fifth specific embodiments.

[0103] Specific embodiment seven: The difference between this embodiment and any one of the first to sixth specific embodiments is that in step 2, the context-aware prompt module includes: a global average pooling layer, a linear layer, a Softmax layer, a prompt weight, a bilinear upsampling layer, and a convolutional layer with a convolution kernel size of 1×1;

[0104] The specific processing process of the context-aware prompt module is: feature map F 2 It is sequentially input into the global average pooling layer, the linear layer, and the Softmax layer, and the Softmax layer generates a compressed prompt weight ε;

[0105] A prompt P composed of a set of learnable parameters is introduced;

[0106] The prompt P is multiplied by the weight ε for output, and the condition input-related prompt P w ;

[0107] For the condition input-related prompt P w After bilinear upsampling, it is input into a convolutional layer with a convolution kernel size of 1×1, and the output feature is used as the output feature of the context-aware prompt module.

[0108] Other steps and parameters are the same as those in any one of the first to sixth specific embodiments.

[0109] Specific embodiment eight: The difference between this embodiment and any one of the first to seventh specific embodiments is that the context-aware prompt module is inserted after each convolutional block in the pre-trained ConvNeXt backbone network, and the output feature of each convolutional block is element-wise added to the output feature of the corresponding context-aware prompt module (the context-aware prompt module following the output feature of each convolutional block is the corresponding context-aware prompt module) as the input feature of the next convolutional block;

[0110] The specific processing process is as follows:

[0111] Given the feature map generated by a certain convolutional block

[0112] First, the feature map F is processed by using global average pooling (GAP) in the spatial dimension 2 to extract a feature vector

[0113] Subsequently, v is processed through a linear layer responsible for channel scaling and a Softmax operation to generate a compressed prompt weight L represents the length of the prompt weight;

[0114] The expression is:

[0115] ε = [ε 1 , …, ε l , …, ε L = Softmax(Linear(GAP(F 2 )))

[0116] Where ε 1 represents the first hint weight component, ε l represents the l-th hint weight component, ε L represents the L-th hint weight component, l = 1, 2, …, L;

[0117] Then, a hint P = [P 1 , …, P l , …, P L composed of a set of learnable parameters is introduced to interact with the hint weight ε to embed context information. P 1 represents the first hint component, P l represents the l-th hint component, P L represents the L-th hint component;

[0118] To maintain the low computational cost of the context-aware hint module, L is set to 5, is set to 8, representing the height, width, and number of channels of the hint composed of learnable parameters respectively;

[0119] In this way, the context-aware hint module dynamically predicts weights based on the context information of the input feature map and applies them to the hint components to generate hints P related to the conditional input w . In addition, the context-aware hint module establishes a shared space to promote the exchange and sharing of relevant knowledge between hint components;

[0120] The expression is:

[0121]

[0122] Where P w represents the hint related to the conditional input, · represents dot product;

[0123] Since the spatial dimensions between the feature map F 2 and the hint P w do not match, bilinear upsampling and Conv1×1 operations are used to upsample the hint to the same spatial size as the input feature;

[0124] The prompts generated by the context-aware prompting module are added to the input features through skip connections as the input features for the next convolutional block, promoting the interaction between the features and the prompt information; this enables efficient learning of task-specific representations during the fine-tuning process;

[0125] The expression is:

[0126] F′ 2 = F 2 + Conv 1×1 (Upsample(P w ))

[0127] where F′ 2 represents the feature map integrated with the prompts; Upsample represents bilinear upsampling.

[0128] Other steps and parameters are the same as those in any one of the first to seventh specific embodiments.

[0129] Specific embodiment nine: The difference between this embodiment and any one of the first to eighth specific embodiments is that in step 2, the scene classification task head sequentially includes a fully connected layer and a Softmax activation function;

[0130] The specific processing process of the scene classification task head is: the features are sequentially input into the fully connected layer and the Softmax activation function, and the output of the Softmax activation function is used as the output feature of the scene classification task head;

[0131] The scene classification task head is inserted after the pre-trained ConvNeXt backbone network.

[0132] Other steps and parameters are the same as those in any one of the first to eighth specific embodiments.

[0133] Specific embodiment ten: The difference between this embodiment and any one of the first to ninth specific embodiments is that in step 3, the parameters of the scene classification network model are fine-tuned using the labeled training samples. During fine-tuning, all parameters except those of the efficient quantization adapter module, the context-aware prompting module, and the scene classification task head are frozen to obtain the fine-tuned scene classification network model;

[0134] The specific process is as follows:

[0135] The scene classification network model includes: a chunk head, N convolutional blocks inserted with the equalization adapter module EQAM, N context-aware prompting modules CAPM, and a scene classification task head;

[0136] The specific processing process of the scene classification network model is:

[0137] The image input block header, the block header output feature map is input into the convolutional block of the first insertion adapter module EQAM. The output feature map of the convolutional block of the first insertion adapter module EQAM is input into the first context-aware prompt module CAPM. The output feature map of the first context-aware prompt module CAPM and the output feature map of the convolutional block of the first insertion adapter module EQAM are element-wise added to obtain a feature. Figure 1 ;

[0138] Feature Figure 1 is input into the convolutional block of the second insertion adapter module EQAM. The output feature map of the convolutional block of the second insertion adapter module EQAM is input into the second context-aware prompt module CAPM. The output feature map of the second context-aware prompt module CAPM and the output feature map of the convolutional block of the second insertion adapter module EQAM are element-wise added to obtain a feature. Figure 2 ;

[0139] Until,

[0140] The feature map N-1 is input into the convolutional block of the Nth insertion adapter module EQAM. The output feature map of the convolutional block of the Nth insertion adapter module EQAM is input into the Nth context-aware prompt module CAPM. The output feature map of the Nth context-aware prompt module CAPM and the output feature map of the convolutional block of the Nth insertion adapter module EQAM are element-wise added to obtain the feature map N;

[0141] The feature map N is input into the scene classification task header, and the scene classification task header outputs the scene classification task;

[0142] Based on the pre-trained ConvNeXt backbone network, efficient quantization adapter modules are inserted into each convolutional block in the backbone network; these adapter modules are designed to effectively learn task-specific feature representations during the fine-tuning process while reducing the number of model parameters;

[0143] At the same time, in order to further enhance the model's ability to perceive context information in remote sensing images, context-aware prompt modules are connected after each convolutional block; these modules can dynamically adjust the model's behavior according to the context of the input features, improving the classification performance;

[0144] A scene classification task header is connected after the ConvNeXt backbone network; the task header is responsible for converting the features extracted from the ConvNeXt backbone network and its associated modules into specific scene classification decisions, and finally outputting the prediction probabilities for each class.

[0145] The parameters of the scene classification network model are fine-tuned using the labeled training samples. During fine-tuning, all parameters except for the efficient quantization adapter modules, context-aware prompt modules, and scene classification task header are frozen to obtain the fine-tuned scene classification network model.

[0146] This fine-tuning strategy can significantly reduce the number of parameters to be optimized, thereby accelerating the training speed and reducing the risk of overfitting. The goal of fine-tuning is to minimize the difference between the actual scene labels and the model predictions, which is achieved by the cross-entropy loss function in the present invention.

[0147] After fine-tuning, the present invention evaluates the performance of the scene classification network on an independent test set, including metrics such as accuracy, recall, and F1-score. The design of the efficient quantization adapter module and the context-aware prompt module ensures that the model can still maintain high performance even when the number of parameters is significantly reduced. Finally, this fine-tuned scene classification network can be deployed into actual remote sensing image analysis applications for automatically identifying and classifying surface scenes.

[0148] Other steps and parameters are the same as those in any one of the specific embodiments one to nine.

[0149] Based on the pre-trained ConvNeXt backbone network, an efficient quantization adapter module is inserted into each ConvNeXt Block in the backbone network. These adapter modules are designed to effectively learn task-specific feature representations during the fine-tuning process while reducing the number of model parameters.

[0150] Meanwhile, in order to further enhance the model's ability to perceive context information in remote sensing images, a context-aware prompt module is connected after each ConvNeXt Block. These modules can dynamically adjust the behavior of the model according to the context of the input features, improving the classification performance.

[0151] After adding the above two enhancement modules, the present invention connects a scene classification task head after the ConvNeXt backbone network. Specifically, the scene classification task head consists of a fully connected layer and a Softmax activation function. The task head is responsible for converting the features extracted from the ConvNeXt backbone network and its associated modules into specific scene classification decisions, and finally outputting the predicted probabilities for each class.

[0152] During the training process, the scene classification network is fine-tuned using an annotated remote sensing image dataset. During this process, most of the parameters of the backbone network are frozen, and only the parameters of the efficient quantization adapter module, the context-aware prompt module, and the scene classification task head are updated. This fine-tuning strategy can significantly reduce the number of parameters to be optimized, thereby accelerating the training speed and reducing the risk of overfitting. The goal of fine-tuning is to minimize the difference between the actual scene labels and the model predictions, which is achieved by the cross-entropy loss function in the present invention.

[0153] After fine-tuning, the present invention evaluates the performance of the scene classification network on an independent test set, including metrics such as accuracy, recall, and F1-score. The design of the efficient quantization adapter module and the context-aware prompt module ensures that the model can maintain high performance even when the number of parameters is significantly reduced. Finally, this fine-tuned scene classification network can be deployed into practical remote sensing image analysis applications for automatically identifying and classifying surface scenes.

[0154] It should be understood that the above description of the preferred embodiment is relatively detailed and should not be considered as a limitation on the protection scope of the present invention. Under the inspiration of the present invention, those of ordinary skill in the art can also make substitutions or deformations without departing from the protection scope defined by the claims of the present invention, and all fall within the protection scope of the present invention. The scope of the present invention claimed shall be subject to the appended claims.

Claims

1. A basic model parameter fine-tuning method for remote sensing scene classification tasks, characterized by: The specific process of the method is: Step 1: Use unlabeled remote sensing images to perform self-supervised pre-training on the ConvNeXt backbone network to obtain a pre-trained ConvNeXt backbone network; Step 2: Build a scene classification network model; The specific process is: The scene classification network model includes: ConvNeXt backbone network, adapter module, context-aware prompt module and scene classification task head; The adapter module is inserted into each convolutional block in the pre-trained ConvNeXt backbone network; The context-aware cue module is inserted after each convolutional block in the pre-trained ConvNeXt backbone network; The scene classification task head is inserted into the pre-trained ConvNeXt backbone network; Step 3: Use the labeled training samples to fine-tune the parameters of the scene classification network model. During fine-tuning, freeze all parameters except the adapter module, context-aware prompt module, and scene classification task head to obtain a fine-tuned scene classification network model. The adapter module EQAM in step 2 includes in sequence: a flattening layer Flatten, a first quantized linear layer Q-Linear, an activation function ReLU, a second quantized linear layer Q-Linear, a scaling scalar Scaling, and a deformation Reshape; The specific processing process of the adapter module EQAM is as follows: the features are sequentially input into the flatten layer Flatten, the first quantized linear layer Q-Linear, the activation function ReLU, the second quantized linear layer Q-Linear, and the scaling scalar Scaling. The scaling scalar Scaling outputs the feature B′. The Scaling output feature B′ is element-wise added to the output feature of the flatten layer Flatten to obtain the feature B″. The feature B″ is reshaped to obtain the feature B. Feature B is the output feature of the adapter module EQAM. The context-aware prompt module in step 2 includes: a global average pooling layer, a linear layer, a Softmax layer, a prompt weight, a bilinear upsampling layer, and a convolution layer with a convolution kernel size of 1×1; The specific processing process of the context-aware prompt module is as follows: the feature map F2 is sequentially input into the global average pooling layer, the linear layer, and the Softmax layer, and the Softmax layer generates a compressed prompt weight ε; A hint P consisting of a set of learnable parameters is introduced; The prompt P is multiplied by the weight ε to output the conditional input related prompt P w ; Input related prompts for conditions w After bilinear upsampling, the convolution layer with a kernel size of 1×1 is input, and the output features are used as the output features of the context-aware cue module.

2. The basic model parameter fine-tuning method for remote sensing scene classification tasks according to claim 1, characterized in that: The ConvNeXt backbone network in step 1 includes: a block header and N convolution blocks in sequence; The connection relationship of the ConvNeXt backbone network is: the image is sequentially input into the block header, the first convolution block, the second convolution block, and so on until the Nth convolution block. The output features of the Nth convolution block are used as the output features of the ConvNeXt backbone network. Each convolution block includes a two-dimensional depth convolution layer with a convolution kernel size of 7×7, a layer normalization layer LayerNorm, a two-dimensional convolution layer with a convolution kernel size of 1×1, an activation function GELU, a two-dimensional convolution layer with a convolution kernel size of 1×1, a layer scaling LayerScale, and a path discarding Droppath. The specific processing process of each convolution block is as follows: the input feature α is sequentially input into the two-dimensional deep convolution layer with a convolution kernel size of 7×7, the layer normalization layer LayerNorm, the two-dimensional convolution layer with a convolution kernel size of 1×1, GELU, the two-dimensional convolution layer with a convolution kernel size of 1×1, the layer scaling LayerScale, the path discarding Droppath, and the path discarding Droppath outputs the feature β; Feature α and feature β are element-wise added to obtain feature γ, which is used as the output feature of each convolution block; In the step 1, the unlabeled remote sensing image performs self-supervisory pre-training on the ConvNeXt backbone network to obtain a pre-trained ConvNeXt backbone network; the specific process is: Construct an unlabeled remote sensing image sample set; The ConvNeXt backbone network is pre-trained for mask image reconstruction based on an unlabeled remote sensing image sample set. The optimization goal is to minimize the normalized pixel error between the reconstructed target and the original image at the mask position, and a pre-trained ConvNeXt backbone network is obtained.

3. The basic model parameter fine-tuning method for remote sensing scene classification tasks according to claim 2 is characterized in that: The adapter module in step 2 is inserted into each convolution block in the pre-trained ConvNeXt backbone network; the specific process is: After the adapter module EQAM is inserted into the Droppath layer of each convolutional block in the pre-trained ConvNeXt backbone network, the convolutional block with the adapter module EQAM inserted is obtained; Each convolution block inserted into the adapter module EQAM includes a two-dimensional depth convolution layer with a convolution kernel size of 7×7, a layer normalization layer LayerNorm, a two-dimensional convolution layer with a convolution kernel size of 1×1, an activation function GELU, a two-dimensional convolution layer with a convolution kernel size of 1×1, a layer scaling LayerScale, a path drop Droppath, and an adapter module EQAM; The specific processing process of each convolution block inserted into the adapter module EQAM is as follows: the input feature A is sequentially input into the two-dimensional deep convolution layer with a convolution kernel size of 7×7, the layer normalization layer LayerNorm, the two-dimensional convolution layer with a convolution kernel size of 1×1, GELU, the two-dimensional convolution layer with a convolution kernel size of 1×1, the layer scaling LayerScale, the path drop Droppath, the adapter module EQAM, and the adapter module EQAM outputs the feature B; Feature A and feature B are element-wise added to obtain feature C, which is used as the output feature of each convolution block.

4. The basic model parameter fine-tuning method for remote sensing scene classification tasks according to claim 3 is characterized by: The specific processing process of each adapter module EQAM is: 1) The input feature map of the adapter module is expressed as Where H, W and C represent the height, width and channel dimension of F1 respectively; The feature map F1 is flattened along the channel dimension to form a two-dimensional feature map Where D = H × W, Flatten() means flattening along the channel dimension; 2) For the two-dimensional feature map Projection is performed to obtain the projected feature map The expression is: Among them, s represents the scaling scalar Scaling, and s is set to 1; · is the matrix dot product; ReLU is the activation function; are the learnable weights of the first quantized linear layer, is the learnable weight of the second quantized linear layer; c′ is the hidden feature dimension, c′ is set to 8; 3) Reshape into a 3D feature map of the same spatial size as F1 Among them, the feature map generated by the adapter module is represented as Reshape means reshaping.

5. The basic model parameter fine-tuning method for remote sensing scene classification tasks according to claim 4, characterized in that: The processing process of each quantized linear layer in the first quantized linear layer Q-Linear and the second quantized linear layer Q-Linear is: 1) Introduce bit width b to divide the weight of the quantized linear layer into B = 2 b non-overlapping sets of discrete values The bit width b is set to 8; 2) Set the codebook, which maps to a set of values ​​{c1,...,c j ,...,c B }, {c1,...,c j ,...,c B }, there are B numbers in total, {c1,...,c j ,...,c B } is preset; j=1,2,…,B; Each Corresponding to a specific code c in the code book j ; 3) Introduce 1-dimensional clustering to minimize quantization error: Among them, w i represents the i-th element of the weight matrix W of the quantized linear layer Q-Linear, and m represents the number of elements in the weight matrix W of the quantized linear layer Q-Linear; Indicates w i After quantization function The processed quantized value is the quantization function The processing process is: Quantization function Will All values ​​in are quantized to code c j : Where, j = 1, 2, ..., B; Indicates that it comes from a collection The weight of 4) Collection-based and the code {c1,...,c j ,...,c B }, for the i-th element w of the weight matrix W of the quantized linear layer Q-Linear i Perform standardization, quantize each standardized element, and restore the element to its original mean and variance through de-standardization; The expression is: Among them, w i Represents the i-th element of the weight matrix W of the quantized linear layer Q-Linear; μ represents the mean, MEAN() represents the mean operation, and m represents the number of elements in the weight matrix W of the quantized linear layer Q-Linear; σ represents the variance, STD() represents the variance operation; w′ i Indicates w i Normalized weights; Indicates w i The values ​​after standardization and quantization; Indicates w i The value after quantization; 5) For Compute the approximate gradient: Among them, w′ k Indicates w k After the normalization, the weight w k Represents the kth element of the weight matrix W of the quantized linear layer Q-Linear.

6. The basic model parameter fine-tuning method for remote sensing scene classification tasks according to claim 5, characterized in that: After the context-aware prompt module is inserted into each convolution block in the pre-trained ConvNeXt backbone network, the output features of each convolution block are element-wise summed with the output features of the corresponding context-aware prompt module as the input features of the next convolution block; the specific processing process is as follows: Given a feature map generated by a convolutional block First, the feature map F2 is processed using global average pooling to extract a feature vector Subsequently, v is processed through a linear layer and a Softmax operation to produce a compressed cue weight L represents the length of the prompt weight; The expression is: ε=[ε1,…,ε l ,…,eh L ]=Softmax(Linear(GAP(F2))) Among them, ε1 represents the first prompt weight component, ε l represents the lth prompt weight component, ε L represents the Lth prompt weight component, l = 1, 2, ..., L; Then, we introduce a hint consisting of a set of learnable parameters P=[P1,…,P l ,…,P L ] is used to interact with the prompt weight ε; P1 represents the first prompt component, P l represents the lth prompt component, P L represents the Lth cue component; Set L to 5, Set to 8, Respectively represent the height, width and number of channels of the prompt composed of learnable parameters; The expression is: Among them, P w Indicates prompts related to conditional input, · indicates dot product; The hints are upsampled to the same spatial size as the input features using bilinear upsampling and Conv1×1 operations; the hints generated by the context-aware hint module are added to the input features through skip connections as the input features of the next convolutional block; The expression is: F′2=F2+Conv 1×1 (Upsample(P w )) Among them, F′2 represents the feature map integrated with the prompt; Upsample represents bilinear upsampling.

7. The basic model parameter fine-tuning method for remote sensing scene classification tasks according to claim 6, characterized in that: The scene classification task head in step 2 includes a fully connected layer and a Softmax activation function in sequence; The specific processing process of the scene classification task head is as follows: the features are sequentially input into the fully connected layer and the Softmax activation function, and the output features of the Softmax activation function are used as the output features of the scene classification task head; The scene classification task head is inserted into the pre-trained ConvNeXt backbone network.

8. The basic model parameter fine-tuning method for remote sensing scene classification tasks according to claim 7, characterized in that: In step 3, the labeled training samples are used to fine-tune the parameters of the scene classification network model, and all parameters except the adapter module, the context-aware prompt module and the scene classification task head are frozen during the fine-tuning to obtain a fine-tuned scene classification network model; The specific process is: The scene classification network model includes: a block head, N convolution blocks inserted into adapter modules EQAM, N context-aware prompt modules CAPM, and a scene classification task head; The specific processing process of the scene classification network model is: The image is input into a block header, the feature map output by the block header is input into the convolution block of the first inserted adapter module EQAM, the feature map output by the convolution block of the first inserted adapter module EQAM is input into the first context-aware prompt module CAPM, the feature map output by the first context-aware prompt module CAPM and the feature map output by the convolution block of the first inserted adapter module EQAM are element-wise added to obtain a feature map 1; The feature map 1 is input into the convolution block of the second insertion adapter module EQAM, the feature map output by the convolution block of the second insertion adapter module EQAM is input into the second context-aware prompt module CAPM, the feature map output by the second context-aware prompt module CAPM and the feature map output by the convolution block of the second insertion adapter module EQAM are element-wise added to obtain the feature map 2; Until, The feature map N-1 is input into the convolution block of the Nth inserted adapter module EQAM, the output feature map of the convolution block of the Nth inserted adapter module EQAM is input into the Nth context-aware prompt module CAPM, the output feature map of the Nth context-aware prompt module CAPM is element-wise added with the output feature map of the convolution block of the Nth inserted adapter module EQAM to obtain the feature map N; The feature map N is input into the scene classification task head, and the scene classification task head outputs the scene classification task; The labeled training samples are used to fine-tune the parameters of the scene classification network model. During fine-tuning, all parameters except the adapter module, context-aware prompt module and scene classification task head are frozen to obtain a fine-tuned scene classification network model.