Visual language remote sensing scene classification method for domain discrimination attention enhancement

By introducing domain discriminant attention enhancement module and visual language pretrained model, integrating remote sensing images and text semantic features, the problems of data limitations and poor cross-domain adaptability in remote sensing scene classification are solved, and classification accuracy and generalization capabilities are improved.

CN120259801AActive Publication Date: 2025-07-04ANHUI UNIV

Patent Information

Application Number
CN202510755943.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-07
Publication Date
2025-07-04
Estimated Expiration
2045-06-07

AI Technical Summary

Technical Problem

The existing remote sensing scenario classification methods have problems such as data limitations, insufficient use of semantic information and poor cross-domain adaptability, especially in complex scenarios and cross-domain data processing.

Method used

The domain discriminative attention enhancement module is introduced, combined with the visual language pre-trained model, the remote sensing image visual features and text semantic features are fused through the image encoder and text encoder, and the model adaptability is improved using the cross-domain attention and gradient inversion layers, and the domain discriminative attention enhancement spatial state embedding network is constructed.

Benefits of technology

The accuracy and generalization ability of remote sensing scenario classification are improved, and the model's adaptability to remote sensing data in different domains is enhanced, especially in complex scenarios and cross-domain data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259801A_ABST
    Figure CN120259801A_ABST
Patent Text Reader

Abstract

The invention relates to a visual language remote sensing scene classification method for domain discrimination attention enhancement. The method comprises the following steps: selecting and processing a remote sensing scene classification data set, constructing a domain discriminant attention enhancement spatial state embedded network, training the domain discriminant attention enhancement spatial state embedded network, and testing the domain discriminant attention enhancement spatial state embedded network. Compared with the prior art, the domain discrimination attention enhancement module can dynamically adjust the weight of the source domain feature according to the target domain feature, so that the target domain feature can focus on the most valuable information in the source domain; an encoder of a CLIP model is pre-trained by using a contrast language image, and judgment can be carried out by using related knowledge learned in large-scale data through association with natural language description, so that the generalization ability of a network in different scenes is enhanced, and meanwhile, a spatial state embedding module constructs richer feature representation, so that the robustness of the network is improved. And the model captures an interaction relationship between the two modes, and potential information is mined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image analysis and computer vision, and particularly relates to a visual language remote sensing scene classification method with domain discriminative attention enhancement. Background Art

[0002] Remote sensing scene classification aims to accurately classify different scenes (such as farmland, city, forest, etc.) in remote sensing images, which is an important basic task in remote sensing image analysis and is applicable to fields such as land use monitoring, urban planning, and disaster warning.

[0003] Traditional remote sensing scene classification methods mostly rely on manually designed features, such as texture features, spectral features, etc. Such methods have high requirements for feature engineering and limited classification effects in complex scenes. With the development of deep learning, methods based on convolutional neural networks (CNNs) have significantly improved the classification accuracy by automatically extracting image features. However, the existing deep learning methods still have the following problems: ① Data limitations. The cost of annotating high-quality remote sensing image data is high, and the limited training data leads to insufficient generalization ability of the model when facing unseen scenes or data distribution changes. ② Insufficient utilization of semantic information. Existing deep learning methods mainly focus on image visual features and are difficult to effectively fuse text semantic information related to the scene, which limits the understanding of complex scene semantics. ③ Poor cross-domain adaptability. Remote sensing images obtained by different sensors, at different shooting times, and in different regions have data distribution differences (i.e., domain differences), and existing models perform poorly in cross-domain scene classification tasks.

[0004] In recent years, contrastive learning and pre-training techniques have achieved remarkable results in the fields of natural language processing and computer vision. For example, the CLIP (Contrastive Language-Image Pretraining) model maps images and texts to a shared feature space through contrastive learning and achieves zero-shot learning ability. However, directly applying it to remote sensing scene classification still cannot fully solve the specific problems of remote sensing data, such as the effective processing of cross-domain data and the accurate modeling of complex semantics of remote sensing scenes. Therefore, there is an urgent need for a new method that can combine the characteristics of remote sensing data, effectively utilize multi-modal information and cross-domain knowledge, and improve the performance of remote sensing scene classification. Summary of the Invention

[0005] The purpose of the present invention is to provide a visual language remote sensing scene classification method with domain discriminative attention enhancement. By introducing a domain discriminative attention enhancement module, the adaptability of the model to remote sensing data in different domains is enhanced; combined with visual language pre-training, the visual features of remote sensing images and text semantic features are effectively fused, thereby improving the accuracy and generalization ability of remote sensing scene classification and solving the problems of existing methods in aspects such as limited data, insufficient semantic utilization, and poor cross-domain adaptability.

[0006] To achieve the above object, the technical solution of the present invention is as follows:

[0007] A visual language remote sensing scene classification method with domain discriminative attention enhancement, comprising the following steps:

[0008] 11) Select and process the remote sensing scene classification dataset, including the following steps:

[0009] 111) Select three datasets, namely AID, NWPU-RESISC45, and UC Merced LandUse, from the publicly available remote sensing scene classification datasets, extract 8 common scene categories among them as the training dataset, and respectively select two of the datasets, one as the source domain data and the other as the target domain data;

[0010] 112) The source domain data includes source domain images and source domain labels, where the source domain labels exist in the form of text information, and the target domain data only contains target domain images without corresponding label information;

[0011] 113) Crop all images to a size of 224×224 pixels, and perform preprocessing operations of rotation, Gaussian blur, and CutMix according to a ratio of 20%. The CutMix operation specifically randomly selects a certain area in the image and removes the pixel values of that area, fills the removed area with the pixel values of the corresponding area of other images randomly selected from the training set, and simultaneously distributes the classification result labels proportionally according to the ratio of the filled area;

[0012] 12) Construct a domain discriminative attention enhanced spatial state embedding network, including an image encoder, a text encoder, a domain discriminative attention enhancement module, and a spatial state embedding module;

[0013] 13) Train the domain discriminative attention enhanced spatial state embedding network, including the following steps:

[0014] 131) First, use the pre-trained weights of the Contrastive Language-Image Pretraining (CLIP) model to initialize the image encoder and the text encoder respectively;

[0015] 132) Input the source domain images and the target domain images into the initialized image encoder to obtain source domain image features and target domain image features respectively, and input the source domain labels into the initialized text encoder to obtain text features;

[0016] 133) Then input the source domain image features and the target domain image features into the domain discriminative attention enhancement module, and after processing, obtain image features that fuse domain information;

[0017] 134) Then, input the image features and text features fused with domain information into the spatial state embedding module to obtain fused modality features, and calculate the similarity between the fused modality features and the text features. The category with the highest similarity is the classification result predicted for the target domain image;

[0018] 135) Finally, update the network parameters according to the loss through backpropagation and save the learned optimal weight parameters;

[0019] 14) Test the domain discriminative attention enhanced spatial state embedding network: Input the target domain image and the source domain label into the trained domain discriminative attention enhanced spatial state embedding network, and obtain the classification result of the target domain image through the proposed image encoder, text encoder, domain discriminative attention enhanced module, and spatial state embedding module.

[0020] The construction of the domain discriminative attention enhanced spatial state embedding network includes the following steps:

[0021] 21) Construct an image encoder to extract image features. Use the ViT-B / 16 model in CLIP as the encoder, which includes a 12-layer transformer structure with 12 attention heads and three fully connected Linear layers;

[0022] 211) The 16 in ViT-B / 16 indicates that the patch size input to the Patch Embedding layer is 16×16, and B represents Base;

[0023] 212) The transformer includes two layer normalizations LN, a 12-head self-attention, a multi-layer perceptron MLP, and two residual connections;

[0024] 213) The three Linear layers are respectively used to generate the target domain query Q, source domain key K, and source domain value V matrices;

[0025] 22) Construct a text encoder to extract text features. The text encoder uses a 12-layer transformer structure, which is the same as the image encoder structure and has 8 self-attention heads;

[0026] 23) Construct a domain discriminative attention enhanced module, which includes a cross-domain attention layer, a feature enhancement module, and a gradient reversal layer GRL;

[0027] 231) The feature enhancement module specifically includes: a global average pooling GAP, two fully connected FC layers, a ReLU layer, a Sigmoid activation layer, a 1×3 convolution, a 3×1 convolution, a 1×1 convolution, and an element-wise multiplication operation;

[0028] 24) Construct a spatial state embedding module, which includes the following operations: pointwise convolution PWC, 3×3 convolution, Conv1d convolution, ReLU activation function, SiLU activation function, Linear layer, state space model SSM, element-wise addition operation, and concatenation operation.

[0029] The training domain discriminative attention enhanced spatial state embedding network includes the following steps:

[0030] 31) Initialize the domain discriminative attention enhanced spatial state embedding network using the pre-trained ViT-B / 16 encoder of the contrastive language-image pre-training CLIP model;

[0031] 32) Input the source domain image and the target domain image into the image encoder to obtain the source domain image feature S feature and the target domain image feature T feature ;

[0032] 321) Perform the Patch Embedding layer on the source domain image and the target domain image respectively to obtain the feature embeddings S embed and T embed with position information;

[0033] 322) Input the two feature embeddings into the transformer respectively. First, input the feature embeddings S embed and T embed into layer normalization LN respectively to obtain the source domain query Q s , the source domain key K s , the source domain value V s , the target domain query Q t , the target domain key K t , and the target domain value V t six matrices, and each matrix calculates the attention scores attn1 and attn2 through multi-head self-attention;

[0034] ,

[0035] ,

[0036] where is the scaling factor, whose value is 64 in ViT-B / 16, represents the transpose of matrix K, and softmax is the activation function;

[0037] 323) Residually add the attention scores to the feature embeddings, then the result of the addition is further passed through layer normalization LN, and then passed through a multi-layer perceptron MLP for output. The output result is residually added to the attention scores to obtain the two outputs output sand output t ;

[0038] ,

[0039] ,

[0040] Among them, is the scaling factor, with a value of 64 in ViT-B / 16, represents the transpose of matrix K;

[0041] 324) Pass the two outputs of the first transformer through 11 transformers with the same structure respectively to obtain higher-level semantic features S feature and T feature ;

[0042] 325) Input S feature into two fully connected Linear layers to obtain the source domain key K and source domain value V matrices respectively, and input T feature into the Linear layer to obtain the target domain query Q matrix;

[0043] 33) Input the source domain label into the text encoder to obtain the text feature z t ;

[0044] 331) First, perform the Tokenize operation on the source domain label, encode it into tokens for encapsulation, add a start word and an end word to each token, map the tokens to the indices corresponding to the CLIP-defined vocabulary, and use zero-padding to make the lengths of all text sequences the same;

[0045] 332) Then, through the trainable word embedding matrix, convert the word indices into low-dimensional word vectors with a dimension of 512;

[0046] 333) Then input the low-dimensional word vectors into the transformer model for processing. Its calculation process is the same as that of the image operation process. After passing through 12 layers of transformers in sequence, the text feature z t ;

[0047] ,

[0048] Among them, TextEncoder is the text encoder, Tokenize is the tokenization operation, and SourceLabel represents the source domain label;

[0049] 34) Input the target domain query Q, source domain key K, and source domain value V matrices output by the image encoder into the domain discriminative attention enhancement module. After being processed by this module, the image feature z with fused domain information is obtainedi , during the processing, the Gradient Reversal Layer (GRL) is used to promote the model to learn domain-invariant features;

[0050] 341) First, input the target domain query Q, source domain key K, and source domain value V matrices into the cross-domain attention to calculate the cross-domain attention score matrix CrossAttn between the target domain query Q and the source domain key-value;

[0051] ,

[0052] 342) Split the target domain query Q, source domain key K, and source domain value V along the feature dimension into 8 heads, and each head calculates the attention independently and concatenates the results;

[0053] ,

[0054] Among them, , i represents the number of the 1...8 attention heads, is the output projection matrix, and concat is the concatenation operation by channels;

[0055] 343) Perform a residual connection on the cross-domain attention score CrossAttn and the target domain query Q matrix, and stabilize the training through layer normalization LN to obtain the output image feature z i ;

[0056] ,

[0057] Among them, is the optional dimension alignment matrix;

[0058] 344) Pass the output image feature z i through the feature enhancement module to obtain the image feature z j ;

[0059] 3441) First, perform a global average pooling (GAP) operation on the image feature z i to compress it in the spatial dimension, so as to obtain the statistical information of each channel;

[0060] 3442) Subsequently, input the result of the GAP operation into two fully connected layers (FC) in sequence, and a ReLU activation is used between these two fully connected layers (FC) to introduce non-linear transformation;

[0061] 3443) Activate the output of the second FC layer through Sigmoid to obtain the channel weight vector w;

[0062] ,

[0063] 3444) Next, pass the image feature zi Perform a 1×3 convolution, a 3×1 convolution, and a 1×1 convolution in sequence to obtain the asymmetric convolution output feature z a ;

[0064] ,

[0065] 3445) Finally, multiply the channel weight vector w by the asymmetric convolution output feature z a to obtain the image feature z of the fused domain information j ;

[0066] ,

[0067] where denotes element-wise multiplication;

[0068] 345) Use the gradient reversal layer for implicit adversarial learning to promote the model to learn domain-invariant features. GRL acts as an identity mapping during forward propagation and multiplies the gradient by a negative learning rate scaling factor during backpropagation ;

[0069] ,

[0070] ,

[0071] where L is the loss function, x is the input feature, y is the output feature of GRL(x), is an adjustable hyperparameter with a range between (0, 1);

[0072] 35) Input the image feature z of the fused domain information j and the text feature z t into the spatial state embedding module to obtain the fused modality feature z f ;

[0073] 351) First, pass the image feature z of the fused domain information j and the text feature z t through a 3×3 convolution, a pointwise convolution PWC, and ReLU respectively to obtain out j , out t ;

[0074] ,

[0075] ,

[0076] 352) Combine out j and out tAfter element-wise addition, it goes through a 3×3 convolution and ReLU activation, a linear layer, a Conv1d convolution and SiLU activation to obtain out1;

[0077] ,

[0078] Among them, denotes element-wise addition, and Linear denotes a linear layer;

[0079] 353) Then, out j and out t After concatenation by channel, it goes through a 3×3 convolution and ReLU activation, a linear layer, a Conv1d convolution and SiLU activation, and a state space model SSM to obtain out2;

[0080] ,

[0081] 354) Finally, out1 and out2 are concatenated and passed through a linear layer to obtain the fused modality feature z f ;

[0082] ,

[0083] 36) Calculate the similarity between the fused modality feature z f and the text feature z t In the feature space, determine the text with the highest similarity as the predicted label of the image;

[0084] ,

[0085] Among them, is the L2 norm, and Similarity denotes similarity;

[0086] 37) Calculate the cross-entropy loss Loss between the predicted label and the source domain label;

[0087] ,

[0088] Among them, cross_entorpy_loss is the cross-entropy loss function encapsulated in the pytorch framework, and logit denotes the predicted label;

[0089] 38) Use Loss for backpropagation to update the network parameters, and determine whether the set number of rounds is reached, or stop training if Loss no longer decreases, save the best weights of the training, otherwise return to step 32) to continue loading data for training.

[0090] The specific steps for testing the domain discriminative attention enhanced spatial state embedding network are as follows:

[0091] 41) Input a single target domain image to be tested and source domain labels into the domain discriminative attention enhanced spatial state embedding network;

[0092] 411) The target domain image is input into the image encoder to obtain an output, which then passes through three Linear layers to generate the target domain query Q, target domain key K, and target domain value V matrices, and is input into the domain discriminative attention enhanced module to obtain image features fused with domain information;

[0093] 412) The source domain labels are input into the text encoder to obtain text features;

[0094] 413) Input the image features fused with domain information and the text features into the spatial state embedding module to obtain fused modality features;

[0095] 42) Calculate the similarity between the fused modality features and the text features, and the text label with the highest similarity is the category of the test image.

[0096] Beneficial effects

[0097] The present invention relates to a domain discriminative attention enhanced visual - language remote - sensing scene classification method. Compared with the prior art, by loading the CLIP pre - trained weights, it reduces the model's dependence on large - scale labeled data, improves the generalization ability of the model, and combines visual - language pre - training to deeply fuse the visual features of remote - sensing images and text semantic features, enabling the model to better understand the semantic information of remote - sensing scenes. Especially for complex scenes and scenes that are difficult to directly distinguish by visual features, the classification accuracy is significantly improved; through the domain discriminative attention enhanced module, it effectively captures the feature differences and commonalities of remote - sensing data in different domains, enhances the model's adaptability to cross - domain data, and improves the classification performance on remote - sensing images from different sensors and different regions; through the spatial state embedding module, it effectively captures the global view and better fuses the image and text modality information. This method can be applied to various remote - sensing scene classification tasks and has high practical value and promotional significance. Brief description of the drawings

[0098] Figure 1 is a flowchart of a domain discriminative attention enhanced visual - language remote - sensing scene classification method, showing each step from data pre - processing to the final scene classification;

[0099] Figure 2 is a model structure diagram for constructing a domain discriminative attention enhanced spatial state embedding network, presenting in detail the structures of the image encoder, text encoder, domain discriminative attention enhanced module, and spatial state embedding module and their connection relationships;

[0100] Figure 3 is the structural details of the feature enhancement module;

[0101] Figure 4 It is the structural details of the spatial state embedding module;

[0102] Figure 5 It is the test result graph of a visual language remote sensing scene classification method with domain discriminative attention enhancement. The "class" shown above the image is the predicted label, and "probability" is the possibility of belonging to the class; Specific implementation manners

[0103] To have a further understanding and recognition of the structural features and achieved effects of the present invention, the following is a detailed description with preferred embodiments and accompanying drawings:

[0104] As Figure 1 shown, a visual language remote sensing scene classification method with domain discriminative attention enhancement according to the present invention includes the following steps:

[0105] The first step is to select and process the remote sensing scene classification dataset, and the specific steps are as follows:

[0106] (1) Select three datasets, namely AID, NWPU-RESISC45, and UC Merced LandUse, from the publicly available remote sensing scene classification datasets, extract 8 common scene categories of the three as the training dataset, and respectively select two of the datasets, one as the source domain data and the other as the target domain data;

[0107] (2) The source domain data includes the source domain image SourceImage and the source domain label SourceLabel, where the source domain label SourceLabel exists in the form of text information, and the target domain data only contains the target domain image TargetImage without corresponding label information;

[0108] (3) Crop all images to a size of 224×224 pixels, and perform preprocessing operations of rotation, Gaussian blur, and CutMix according to a ratio of 20%. The CutMix operation is specifically to randomly select a certain area in the image and remove the pixel values of this area, fill the removed area with the pixel values of the corresponding area of other images randomly selected from the training set, and at the same time allocate the classification result labels proportionally according to the proportion of the filled area.

[0109] The second step is to construct a domain discriminative attention enhanced spatial state embedding network, and the specific steps are as follows:

[0110] (1) Construct an image encoder ImageEncoder to extract image features, using the ViT-B / 16 model in CLIP, including a Patch Embedding layer, 12 layers of transformer structures, and three fully connected Linear layers;

[0111] In (1-1), the 16 in ViT-B / 16 indicates that the size of the patch input to the Patch Embedding layer is 16×16, and B represents Base;

[0112] In (1-2), the transformer includes two layer normalizations LN, a 12-head self-attention, a multi-layer perceptron MLP, and two residual connections;

[0113] In (1-3), three Linear layers are respectively used to generate the target domain query Q, the source domain key K, and the source domain value V matrices;

[0114] In (2), a TextEncoder is constructed to extract text features. The TextEncoder uses a 12-layer transformer structure, which is the same as the ImageEncoder of the image encoder and has 8 self-attention heads;

[0115] In (3), a domain discriminative attention enhancement module is constructed, which includes a cross-domain attention layer, a feature enhancement module, and a gradient reversal layer GRL;

[0116] In (3-1), the feature enhancement module specifically includes: a global average pooling (Global Average Pooling, GAP), two fully connected FC layers, a ReLU layer, a Sigmoid activation layer, a 1×3 convolution, a 3×1 convolution, a 1×1 convolution, and an element-wise multiplication operation;

[0117] In (4), a spatial state embedding module is constructed, which includes the following operations: pointwise convolution PWC, 3×3 convolution, Conv1d convolution, ReLU activation function, SiLU activation function, Linear linear layer, state space model SSM, element-wise addition operation, and concatenation operation.

[0118] First, the preprocessed source domain image SourceImage and target domain image TargetImage are input into the ImageEncoder of the image encoder, and then the image features fused with domain information are obtained through the domain discriminative attention enhancement module. Then, the text label SourceLabel is input into the TextEncoder to obtain text features. The image features and text features are fused through the spatial state embedding module. Finally, the similarity between the text features and the fused modality features is calculated to complete the construction of the entire model.

[0119] In the third step, the domain discriminative attention enhancement spatial state embedding network is trained, and the specific steps are as follows:

[0120] (1) Initialize the domain discriminative attention enhanced spatial state embedding network using the pre-trained ViT-B / 16 encoder of the contrastive language-image pre-training CLIP model;

[0121] (2) Input the source domain image SourceImage and the target domain image TargetImage into the image encoder ImageEncoder respectively to obtain the source domain image feature S feature and the target domain image feature T feature ;

[0122] (2-1) Perform the PatchEmbedding layer on the source domain image SourceImage and the target domain image TargetImage respectively to obtain the feature embeddings S embed and T embed ;

[0123] (2-2) Input the two feature embeddings into the transformer respectively. First, input the feature embeddings S embed and T embed into the layer normalization LN respectively to obtain the source domain query Q s , the source domain key K s , the source domain value V s , the target domain query Q t , the target domain key K t , the target domain value V t six matrices, and the matrices are respectively calculated through the multi-head self-attention to obtain the attention scores attn1 and attn2;

[0124] ,

[0125] ,

[0126] Among them, is the scaling factor, and its value in ViT-B / 16 is 64, represents the transpose of matrix K, and softmax is the activation function;

[0127] (2-3) Residually add the attention scores to the feature embeddings, then the added result is further passed through the layer normalization LN, and then passed through the MLP layer for output. The output result is residually added to the attention scores to obtain the two outputs output s and output t ;

[0128] ,

[0129] ,

[0130] Among them, is the scaling factor, with a value of 64 in ViT-B / 16, represents the transpose of matrix K;

[0131] (2-4) Pass the two outputs of the first transformer through 11 transformers with the same structure respectively to obtain higher-level semantic features S feature and T feature ;

[0132] (2-5) Input S feature into two Linear layers to obtain the source domain key K and source domain value V matrices respectively, and input T feature into the Linear layer to obtain the target domain query Q matrix;

[0133] (3) Input the source domain label SourceLabel into the text encoder TextEncoder to obtain the text feature z t ;

[0134] (3-1) First, perform the Tokenize operation on the source domain label SourceLabel, encode it into tokens for encapsulation, add a start word and an end word to each token, map the tokens to the indices corresponding to the CLIP-defined vocabulary, and use the method of filling with 0 to make the lengths of all text sequences consistent;

[0135] (3-2) Then, through the trainable word embedding matrix, convert the word indices into low-dimensional word vectors with a dimension of 512;

[0136] (3-3) Then input the low-dimensional word vectors into the transformer model for processing. Its calculation process is the same as that of the image operation process. After the low-dimensional word vectors pass through the calculations of 12 layers of transformers in sequence, the text feature z t ;

[0137] ,

[0138] where Tokenize is the word segmentation operation;

[0139] (4) Input the target domain query Q, source domain key K, and source domain value V matrices output by the image encoder into the domain discriminative attention enhancement module. After being processed by this module, the image feature z i is obtained. During the processing, the gradient reversal layer GRL is used to promote the model to learn domain-invariant features;

[0140] (4-1) First, input the target domain query Q, source domain key K, and source domain value V matrices into the cross-domain attention CrossAttention to calculate the cross-domain attention score matrix CrossAttn between the target domain query Q and the source domain key-value;

[0141] ,

[0142] (4-2) Split the target domain query Q, source domain key K, and source domain value V along the feature dimension into 8 heads Head, and calculate the attention independently for each head and concatenate the results;

[0143] ,

[0144] Among them, , i represents the number of the 1...8 attention heads, is the output projection matrix, and concat is the concatenation operation by channel;

[0145] (4-3) Perform a residual connection on the cross-domain attention score CrossAttn and the target domain query Q matrix, and stabilize the training through layer normalization LN to obtain the output image feature z i ;

[0146] ,

[0147] Among them, is the optional dimension alignment matrix;

[0148] (4-4) Pass the output image feature z i through the feature enhancement module to obtain the image feature z j ;

[0149] (4-4-1) First, perform a global average pooling GAP operation on the image feature z i to compress it in the spatial dimension, so as to obtain the statistical information of each channel;

[0150] (4-4-2) Subsequently, input the result of the GAP operation into two fully connected layers FC in sequence, and a ReLU activation is adopted between these two fully connected layers FC to introduce a non-linear transformation;

[0151] (4-4-3) Obtain the channel weight vector w by activating the output of the second FC layer through Sigmoid;

[0152] ,

[0153] (4-4-4) Next, perform 1×3 convolution, 3×1 convolution, and 1×1 convolution on the image feature z i in sequence to obtain the asymmetric convolution output feature za ;

[0154] ,

[0155] (4-4-5) Finally, multiply the channel weight vector w by the asymmetric convolution output feature z a to obtain the image feature z of the fused domain information j ;

[0156] ,

[0157] where denotes element-wise multiplication;

[0158] (4-5) Use the gradient reversal layer for implicit adversarial learning to promote the model to learn domain-invariant features. The GRL acts as an identity mapping during forward propagation and multiplies the gradient by a negative learning rate scaling factor during backpropagation ;

[0159] ,

[0160] ,

[0161] where L is the loss function, x is the input feature, y is the output feature of GRL(x), is an adjustable hyperparameter with a range between (0,1);

[0162] (5) Input the image feature z j of the fused domain information and the text feature z t into the spatial state embedding module to obtain the fused modality feature z f ;

[0163] (5-1) First, pass the image feature z j of the fused domain information and the text feature z t through a 3×3 convolution, a pointwise convolution PWC, and ReLU respectively to obtain out j , out t ;

[0164] ,

[0165] ,

[0166] (5-2) After element-wise adding out j and out t , pass through a 3×3 convolution and ReLU activation, a linear layer, a Conv1d convolution, and SiLU activation to obtain out1;

[0167] ,

[0168] Among them, denotes element-wise addition, and Linear denotes a linear layer;

[0169] (5-3) Then, after concatenating out j and out t by channel, passing through a 3×3 convolution and ReLU activation, a linear layer, a Conv1d convolution and SiLU activation, and a state space model SSM, out2 is obtained;

[0170] ,

[0171] (5-4) Finally, after concatenating out1 and out2 and passing through a linear layer, the fused modal feature z f is obtained;

[0172] ,

[0173] (6) Calculate the similarity between the fused modal feature z f and the text feature z t , and in the feature space, determine the text with the highest similarity as the predicted label logit of the image;

[0174] ,

[0175] Among them, is the L2 norm, and Similarity denotes similarity;

[0176] (7) Calculate the cross-entropy loss Loss between the predicted label logit and the source domain label SourceLabel;

[0177] ,

[0178] Among them, cross_entorpy_loss is the cross-entropy loss function encapsulated in the pytorch framework;

[0179] (8) Use Loss for backpropagation to update the network parameters, and determine whether the set number of epochs is reached, or stop training if Loss no longer decreases, save the best weights during training, otherwise return to step (2) to continue loading data for training.

[0180] Step 4: Test the domain discriminative attention enhanced spatial state embedding network, and the specific steps are as follows:

[0181] (1) Input a single target domain image TargetImage to be tested and eight common class labels SourceLabel into the network;

[0182] (1-1) The TargetImage is input into the ImageEncoder to obtain an output, which then passes through three Linear layers to generate the target domain query Q, target domain key K, and target domain value V matrices. These are input into the domain discriminative attention enhancement module to obtain image features fused with domain information.

[0183] (1-2) The SourceLabel is input into the TextEncoder to obtain text features.

[0184] (1-3) The image features fused with domain information and the text features are input into the spatial state embedding module to obtain fused modality features.

[0185] (2) Calculate the similarity between the fused modality features and the text features. The text label with the highest similarity is the category of the test image TargetImage.

[0186] The method proposed by the present invention will be described below by taking NWPU-RESISC45 as the source domain and AID as the target domain dataset as an example:

[0187] Select 8 common categories as the label set. Select an image from AID, and obtain an image with a size of 224×224 through the preprocessing operation described in the present invention. Using the method described in the present invention, the predicted image categories are as Figure 5 shown. Above the image are the predicted labels and the probabilities belonging to that category, and 6 images are listed for display.

[0188] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A visual language remote sensing scene classification method with domain discriminative attention enhancement, characterized in that, Including the following steps: 11) Select and process the remote sensing scene classification dataset: Select 8 categories common to the three datasets of AID, NWPU-RESISC45, and UC Merced Land Use from the publicly available remote sensing scene classification datasets as training data, and select two training data as the source domain data and the target domain data; the source domain data includes source domain images and source domain labels, and the target domain data only contains target domain images; 12) Construct a domain discriminative attention-enhanced spatial state embedding network: including an image encoder, a text encoder, a domain discriminative attention-enhanced module, and a spatial state embedding module; 13) Train the domain discriminative attention-enhanced spatial state embedding network: First, initialize the image and text encoders with the pre-trained weights of the CLIP model, then input the source domain data and the target domain data into the domain discriminative attention-enhanced spatial state embedding network respectively, update the network parameters according to the loss by backpropagation, and save the best weight parameters; 14) Test the domain discriminative attention-enhanced spatial state embedding network: Input the target domain images and the source domain labels into the trained domain discriminative attention-enhanced spatial state embedding network to obtain the classification results of the target domain images.

2. The visual language remote sensing scene classification method with domain discriminative attention enhancement according to claim 1, wherein, The construction of the domain discriminative attention-enhanced spatial state embedding network includes the following steps: 21) Construct an image encoder to extract image features, use the ViT-B / 16 model in CLIP as the encoder, including a 12-layer transformer structure with 12 attention heads and three fully connected Linear layers; 211) The 16 in ViT-B / 16 indicates that the patch size input to the Patch Embedding layer is 16×16, and B represents Base; 212) The transformer includes two layer normalizations LN, a 12-head self-attention, a multi-layer perceptron MLP, and two residual connections; 213) The three Linear layers are respectively used to generate the target domain query Q, the source domain key K, and the source domain value V matrices; 22) Construct a text encoder to extract text features, the text encoder uses a 12-layer transformer structure, and the transformer structure is the same as the image encoder, with 8 self-attention heads; 23) Construct a domain discriminative attention-enhanced module, which includes a cross-domain attention layer, a feature enhancement module, and a gradient reversal layer GRL; 231) The feature enhancement module specifically includes: a global average pooling GAP, two fully connected FC layers, a ReLU layer, a Sigmoid activation layer, a 1×3 convolution, a 3×1 convolution, a 1×1 convolution, and an element-wise multiplication operation; 24) Construct a spatial state embedding module, which includes the following operations: pointwise convolution PWC, 3×3 convolution, Conv1d convolution, ReLU activation function, SiLU activation function, Linear linear layer, state space model SSM, element-wise addition operation, and concatenation operation.

3. A visual language remote sensing scene classification method with domain discriminative attention enhancement according to claim 1, characterized in that, The training of the domain discriminative attention-enhanced spatial state embedding network includes the following steps: 31) Initialize the domain discriminative attention enhanced spatial state embedding network using the pre-trained ViT-B / 16 encoder of the contrastive language-image pre-training CLIP model; 32) Input the source domain image and the target domain image into the image encoder to obtain the source domain image feature S feature and the target domain image feature T feature ; 321) Perform the Patch Embedding layer on the source domain image and the target domain image respectively to obtain the feature embeddings S embed and T embed ; 322) Input the two feature embeddings into the transformer respectively. First, input the feature embedding S embed and T embed into layer normalization LN respectively to obtain the source domain query Q s , the source domain key K s , the source domain value V s , the target domain query Q t , the target domain key K t , and the target domain value V t These six matrices are each obtained by calculating the attention scores attn1 and attn2 through multi-head self-attention; , , Among them, is the scaling factor, with a value of 64 in ViT-B / 16, represents the transpose of matrix K, and softmax is the activation function; 323) Note that the fraction and the feature embedding are added residually, and then the result of the addition is further normalized by layer normalization LN, and then output through a multi-layer perceptron MLP. The output result is added residually to the attention score to obtain the two outputs output s and output t ; , , Among them, is the scaling factor, with a value of 64 in ViT-B / 16, represents the transpose of matrix K; 324) Pass the two outputs of the first transformer through 11 transformers with the same structure respectively to obtain higher-level semantic features S feature and T feature ; 325) Input S feature Input two fully-connected Linear layers to obtain the source domain key K and source domain value V matrices respectively, T feature Input a Linear layer to obtain the target domain query Q matrix; 33) Input the source domain label into the text encoder to obtain the text feature z t ; 331) First, perform a Tokenize operation on the source domain labels, encode them into tokens for encapsulation, add a start word and an end word to each token, map the tokens to the indices corresponding to the vocabulary defined by CLIP, and use zero-padding to make the lengths of all text sequences consistent; 332) Then, through a trainable word embedding matrix, convert the word indices into low-dimensional word vectors with a dimension of 512; 333) Then, the low-dimensional word vectors are input into the transformer model for processing. The calculation process is the same as that of the image. After passing through the calculations of 12 layers of transformers in sequence, the text feature z is obtained t ; , Among them, TextEncoder is the text encoder, Tokenize is the tokenization operation, and SourceLabel represents the source domain label; 34) Input the target domain query Q, source domain key K, and source domain value V matrices output by the image encoder into the domain discriminative attention enhancement module. After being processed by this module, the image feature z fused with domain information is obtained. i , during the processing, the gradient reversal layer GRL is used to promote the model to learn domain-invariant features; 341) First, input the target domain query Q, source domain key K, and source domain value V matrices into the cross-domain attention to calculate the cross-domain attention score matrix CrossAttn between the target domain query Q and the source domain key-value; , 342) Split the target domain query Q, source domain key K, and source domain value V along the feature dimension into 8 heads, and calculate the attention independently for each head and concatenate the results; , Among them, , i represents the numbers of 1...8 attention heads, is the output projection matrix, and concat is the channel-wise concatenation operation; 343) Perform a residual connection between the cross-domain attention score CrossAttn and the target domain query Q matrix, and stabilize the training through layer normalization LN to obtain the output image feature z i ; , Among them, is an optional dimension alignment matrix; 344) The output image feature z i passes through the feature enhancement module to obtain the image feature z with fused domain information j ; 3441) First, perform a global average pooling (GAP) operation on the image feature z i to compress it in the spatial dimension, thereby obtaining the statistical information of each channel; 3442) Subsequently, input the results after the GAP operation into two fully connected layers FC in sequence, and introduce a non-linear transformation using the ReLU activation between these two fully connected layers FC; 3443) Obtain the channel weight vector w by activating the output of the second FC layer through Sigmoid; , 3444) Next, the image feature z i will be successively subjected to 1×3 convolution, 3×1 convolution, and 1×1 convolution to obtain the asymmetric convolution output feature z a ; , Finally, multiply the channel weight vector w by the output feature z of the asymmetric convolution to obtain the image feature z of the fusion domain information a j ;​ , Among them, represents element-wise multiplication; 345) Use a gradient reversal layer for implicit adversarial learning to promote the model to learn domain-invariant features. The GRL acts as an identity mapping during the forward propagation process and multiplies the gradient by a negative learning rate scaling factor during the backward propagation ; , , where L is the loss function, x is the input feature, and y is the output feature of GRL(x), is an adjustable hyperparameter with a range between (0, 1); 35) Input the image feature z j fused with domain information and the text feature z t into the spatial state embedding module to obtain the fused modality feature z f ; 351) First, the image feature z j of the fusion domain information and the text feature z t are respectively passed through a 3×3 convolution, a pointwise convolution PWC, and ReLU to obtain out j and out t ; , , 352) Add out j and out t element-wise, then pass through a 3×3 convolution and ReLU activation, a linear layer, a Conv1d convolution and SiLU activation to obtain out1; , Among them, represents element-wise addition, and Linear represents a linear layer; 353) Then take out j and out t After concatenating them by channel, pass them through a 3×3 convolution and ReLU activation, a linear layer, a Conv1d convolution and SiLU activation, and a state space model SSM to obtain out2; , Finally, out1 and out2 are concatenated and passed through a linear layer to obtain the fused modality feature z f ; , 36) Calculate the similarity between the fused modal feature z f and the text feature z t and determine the text with the highest similarity in the feature space as the predicted label of the image; , Among them, is the L2 norm, and Similarity represents the similarity; 37) Calculate the cross-entropy loss Loss between the predicted label and the source domain label; , Among them, cross_entorpy_loss is the cross-entropy loss function encapsulated in the pytorch framework, and logit represents the predicted label; 38) Use Loss for backpropagation to update the network parameters, and determine whether the set number of rounds is reached, or stop training if Loss no longer decreases, save the best weights of the training, otherwise return to step 32) to continue loading data for training.

4. A visual language remote sensing scene classification method with domain discriminative attention enhancement according to claim 1, characterized in that, The specific steps for testing the domain discriminative attention enhanced spatial state embedding network are as follows: 41) Input a single target domain image to be tested and the source domain label into the domain discriminative attention enhanced spatial state embedding network; 411) Input the target domain image into the image encoder to obtain an output, and then generate the target domain query Q, target domain key K, and target domain value V matrices through three fully connected Linear layers and input them into the domain discriminative attention enhanced module to obtain the image features fused with domain information; 412) Input the source domain label into the text encoder to obtain text features; 413) Input the image features fused with domain information and the text features into the spatial state embedding module to obtain the fused modality features; 42) Calculate the similarity between the fused modality features and the text features, and the text label with the highest similarity is the category of the test image.

Citation Information

Patent Citations

  • Remote sensing image scene classification method based on multi-mode airspace transformation network

    CN116503753A

  • Multimode self-supervised mixed Mangbar hyperspectral image classification method

    CN119007024A

  • Remote sensing image change detection method based on language guidance

    CN119169449A

  • Clustering-based autism diagnosis method for cross-domain facial expression recognition

    CN119580998A

  • Cross-domain small sample learning hyperspectral image classification method and system based on text perception

    CN119649213A

Cited By

  • General model pre-training and adaptive optimization system and method for remote sensing image

    CN121415186A