A Remote Sensing Image Scene Classification Method Based on a Multimodal Airspace Transformation Network
By using a multimodal airspace transformation network in remote sensing image scene classification to fuse images and semantic information, the problem of insufficient fusion of cross-modal information in the existing technology is solved, and effective classification and feature identification of complex scenarios are achieved.
Patent Information
- Application Number
- CN202310476470.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-04-28
AI Technical Summary
The prior art lacks effective fusion of cross-modal information in remote sensing image scene classification, resulting in insufficient feature identification ability of extracted features to complex scenes.
Using a method based on multimodal airspace transformation network, ResNet50 pre-trained network module, cyclic airspace transformation module and class name embedding module, multi-layer features of the image are acquired and fused with the semantic information of the image category, classification loss and similarity loss of images and text are established, and the model is optimized to achieve effective semantic alignment.
Effectively utilize multimodal information, explore the intrinsic correlation between modes, realize the effective classification of remote sensing images, and improve the feature identification ability of complex scenes.
Smart Images

Figure CN116503753B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing image classification and recognition, and particularly relates to a method for remote sensing image scene classification based on a multi-modal spatial domain transformation network. Background Art
[0002] Multi-modal data, that is, data containing multiple data types, such as text, images, videos, audio, etc., has currently been widely used in many practical application scenarios, such as image classification, autonomous driving, and saliency detection. The research on multi-modal data has broad development prospects and can provide richer and more accurate information for artificial intelligence applications. Combining the internal information of multi-modal data can effectively fuse complementary features and avoid omission of certain information in a single modality. However, most of the research work based on multi-modal only uses the images captured by different sensors as different modalities, without achieving true cross-modal, and the extracted features still have certain limitations.
[0003] Remote sensing image scene classification mainly maps the input image to discrete labels, but the features extracted by the network from the image are limited, and other forms of information related to each image are completely ignored during the training process. Most of the existing research content is carried out for a single modality of images, lacking relevant cross-modal work. Due to the lack of complementary information between different modalities, the features extracted by the network have insufficient feature discrimination ability for complex scenes. The types of data are diverse, and other forms of information can be learned from these multi-modal data to help identify image categories. Currently, many multi-modal frameworks have been proposed in the field of natural images to explore the potential dependencies between different modalities. However, due to the diversity and complexity of remote sensing images, the methods proposed for natural images cannot be used to establish relationships between remote sensing modalities well. Therefore, how to effectively utilize multi-modal information and explore the internal correlation between modalities to achieve effective semantic alignment remains a difficult problem. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a method for remote sensing image scene classification based on a multi-modal spatial domain transformation network, including the following steps:
[0005] S1: Obtain a training data set composed of remote sensing images with scene category labels;
[0006] S2: Establish a remote sensing image classification model; the model includes a ResNet50 pre-training network module, a cyclic spatial domain transformation module, and a class name embedding module;
[0007] The ResNet50 pre-trained network module includes Conv-1, Res-2, Res-3, Res-4, Res-5, a dilated spatial pyramid pooling layer, a global average pooling layer, and a Softmax layer;
[0008] S3: Input the remote sensing images in the training dataset into the remote sensing image classification model for model training;
[0009] S31: Input the remote sensing image into the ResNet50 pre-trained network module to obtain multi-layer features. The multi-layer features undergo feature interaction through the dilated spatial pyramid and output the overall feature f through global average pooling 1 , the feature f 1 passes through the Softmax layer to obtain the predicted classification result of the image;
[0010] S32: The cyclic spatial transformation module performs cyclic adaptive spatial transformation on features at different levels;
[0011] S33: Input the class label of the image into the class name embedding module, extract the semantic information of the remote sensing image class through the GloVe model and the multi-head self-attention mechanism, and obtain the predicted classification result of the text through the Softmax layer;
[0012] S34: Perform pixel-by-pixel weighted fusion on the semantic information of the class name and the features after cyclic adaptive spatial transformation to obtain the discriminative feature f 2 ;
[0013] S35: Establish the classification losses of the image and text respectively according to the predicted classification results of the image and text, and establish the similarity loss according to the overall feature f 1 and the discriminative feature f 2 ;
[0014] S36: Use the classification losses of the image and text and the similarity loss as the final loss function of the remote sensing image classification model. When the value of the loss function is the smallest, the model training is completed;
[0015] S4: Input the remote sensing image to be classified into the trained remote sensing image classification model for classification to obtain the classification result.
[0016] Advantages of the present invention:
[0017] The present invention effectively utilizes multi-modal information and explores the internal correlation between modalities to achieve effective semantic alignment by fusing the multi-layer features of the image with the semantic information of the image class; at the same time, the remote sensing image classification model obtained by jointly optimizing the classification losses of the image and text and the similarity loss can achieve the classification of remote sensing images. Description of the Drawings
[0018] Figure 1 This is a framework diagram of a remote sensing image scene classification method based on a multi-modal spatio-temporal transformation network of the present invention. Specific embodiments
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] A remote sensing image scene classification method based on a multi-modal spatio-temporal transformation network, as Figure 1 shown, includes:
[0021] S1: Obtain remote sensing images with scene category labels to form a training data set;
[0022] S2: Establish a remote sensing image classification model; the model includes a ResNet50 pre-trained network module, a cyclic spatio-temporal transformation module, and a class name embedding module;
[0023] The ResNet50 pre-trained network module includes Conv-1, Res-2, Res-3, Res-4, Res-5, a dilated spatial pyramid pooling layer, a global average pooling layer, and a Softmax layer;
[0024] S3: Input the remote sensing images in the training data set into the remote sensing image classification model for model training;
[0025] S31: Input the remote sensing image into the ResNet50 pre-trained network module to obtain multi-level features. The multi-level features undergo feature interaction through the dilated spatial pyramid and output the overall feature f 1 , the feature f 1 passes through the Softmax layer to obtain the predicted classification result of the image;
[0026] S32: The cyclic spatio-temporal transformation module performs cyclic adaptive spatial transformation on features at different levels;
[0027] S33: Input the category label of the image into the class name embedding module, extract the semantic information of the remote sensing image category through the GloVe model and the multi-head self-attention mechanism, and obtain the predicted classification result of the text through the Softmax layer;
[0028] S34: Perform pixel-by-pixel weighted fusion on the semantic information of the class name and the features after cyclic adaptive spatial transformation to obtain the discriminative feature f 2 ;
[0029] S35: Establish the classification losses of the image and text respectively according to the predicted classification results of the image and text, and establish the similarity loss according to the overall feature f 1 and the discriminative feature f 2 ;
[0030] S36: Take the classification losses of the image and text and the similarity loss as the final loss function of the remote sensing image classification model, and complete the training of the model when the value of the loss function is the smallest;
[0031] S4: Input the remote sensing image to be classified into the trained remote sensing image classification model for classification to obtain the classification result.
[0032] Input the remote sensing image into the ResNet50 pre-trained network module to obtain multi-level features, including:
[0033] Input the remote sensing image into the Conv-1 layer for image enhancement, and the enhanced image undergoes layer-by-layer feature extraction through Res-2, Res-3, Res-4, and Res-5 to obtain multi-level features.
[0034] Perform cyclic adaptive spatial transformation on features at different levels, including:
[0035] Perform cyclic adaptive spatial transformation on features at different levels of Res-2, Res-3, and Res-4 in the ResNet50 pre-trained network: First, input the feature map into the localization network to generate the transformation parameter θ. In the localization network, features are extracted successively through convolutional kernels of different sizes of 5×5 and 3×3, and then 1×1 convolution is used to achieve cross-channel information fusion. The final transformation parameter θ is obtained through the MLP regression layer. The grid generator performs corresponding spatial transformation on the positions in the image using the transformation parameter θ regressed by the localization network, and the sampler uses bilinear interpolation to obtain the output feature map.
[0036] The corresponding spatial transformation includes operations such as image scaling, rotation, and translation.
[0037] Extract the semantic information of the remote sensing image category through the GloVe model and the multi-head self-attention mechanism, including:
[0038] Embed the labels of K scene categories into an m-dimensional vector space R through the GloVe model, thereby generating K semantic feature vectors S 1 , S 2 , …, S k , select the word vector S of the category label corresponding to the input image category label i , split it into n segments and perform replication and extension to obtain the word vector X iPerform the multi-head self-attention mechanism operation as the input of the multi-head self-attention model. Among them, the scaled dot-product attention mechanism obtains the attention scores by performing dot-product operations on the query vector Q, the key vector K, and the value vector V, normalizes the attention scores, and calculates the output result by weighted summation of V. Then, the semantic information vector of the class name is obtained, and the semantic information vector of the class name is converted to the specified dimension using a fully connected layer. Finally, the Sigmoid activation function is used for processing to obtain the deep semantic information of the class name.
[0039] Establish the classification losses of the image and text respectively according to the classification results and the image category information. The establishment methods of the two classification losses are the same. Calculate the probability of the class label separated by the classification results and the true class label, and calculate the classification loss function according to the probability.
[0040] The calculation method of the sample classification label belonging to the true label:
[0041]
[0042] Among them, represents the probability that the i-th sample classification label belongs to the true label y i , z i represents the image category and text category values output by the model, and K represents the number of sample class labels.
[0043] The classification loss of the image includes:
[0044]
[0045] Among them, L img represents the classification loss of the image, N represents the number of samples, represents the probability that the i-th sample classification label output by the pre-trained network belongs to the true label y i .
[0046] The classification loss of the text includes:
[0047]
[0048] Among them, L txt represents the classification loss of the text, N represents the number of samples, represents the probability that the i-th sample classification label output by the class name embedding module belongs to the true label y i .
[0049] The similarity loss includes:
[0050]
[0051] Among them, Lsim Denote the similarity loss as f 1 and f 2 respectively represent the overall feature obtained by passing the image through the pre-trained network and the discriminative feature obtained by fusing the semantic information of the class name.
[0052] The loss function of the model includes:
[0053]
[0054] Among them, L img and L txt respectively represent the classification losses of the image and the text, N represents the number of samples, and respectively represent the probabilities that the classification label of the i-th sample output by the pre-trained network and the class name embedding module belongs to the true label y i The L sim denotes the similarity loss, f 1 and f 2 respectively represent the overall feature obtained by passing the image through the pre-trained network and the discriminative feature obtained by fusing the semantic information of the class name.
[0055] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A remote sensing image scene classification method based on multimodal spatial transformation network, It is characterized in that include: S1: Obtain remote sensing images with scene category labels to form a training data set; S2: Establish a remote sensing image classification model; the model includes a ResNet50 pre-trained network module, a cyclic spatial domain transformation module, and a class name embedding module; The ResNet50 pre-trained network module includes Conv-1, Res-2, Res-3, Res-4, Res-5, a dilated space pyramid pooling layer, a global average pooling layer, and a Softmax layer; S3: Input the remote sensing images in the training data set into the remote sensing image classification model for model training; S31: Input the remote sensing image into the ResNet50 pre-trained network module to obtain multi-layer features. The multi-layer features are subjected to feature interaction through the atrous spatial pyramid and the overall feature f is output through global average pooling 1 , feature f 1 passes through the Softmax layer to obtain the predicted classification result of the image; S32: The cyclic spatial transformation module performs cyclic adaptive spatial transformation on features at different levels; S33: Input the category label of the image into the class name embedding module, extract the semantic information of the remote sensing image category through the GloVe model and the multi-head self-attention mechanism, and obtain the predicted classification result of the text through the Softmax layer; S34: Perform pixel-by-pixel weighted fusion of the semantic information of the class name and the features after cyclic adaptive spatial transformation to obtain discriminative feature f 2 ; S35: Establish the classification losses of the image and text respectively according to the predicted classification results of the image and text, and establish the similarity loss according to the overall feature f 1 and the discriminative feature f 2 ; S36: The classification loss of the image and text and the similarity loss are used as the final loss function of the remote sensing image classification model. When the loss function value is the minimum, the model training is completed. S4: Input the remote sensing image to be classified into the trained remote sensing image classification model for classification to obtain the classification result.
2. According to claim 1, a remote sensing image scene classification method based on a multimodal spatial transformation network, It is characterized in that Input the remote sensing image into the ResNet50 pre-trained network module to obtain multi-layer features, including: The remote sensing image is input into the Conv-1 layer for image enhancement. The enhanced image is extracted layer by layer through Res-2, Res-3, Res-4, and Res-5 to obtain multi-layer features.
3. The remote sensing image scene classification method based on a multimodal spatial transformation network according to claim 1, It is characterized in that The features at different levels are subjected to cyclic adaptive spatial transformation, including: The features of different levels of Res-2, Res-3, and Res-4 in the ResNet50 pre-trained network are subjected to cyclic adaptive spatial transformation: the feature map is input into the positioning network to generate the transformation parameter θ. In the positioning network, features are extracted by convolution kernels of different sizes, such as 5×5 and 3×3, and 1×1 convolution is used to achieve cross-channel information fusion. The final transformation parameter θ is obtained through the MLP regression layer. The grid generator uses the final transformation parameter θ of the positioning network to perform corresponding spatial transformation on the position in the image, and the sampler uses bilinear interpolation to obtain the output feature map.
4. The remote sensing image scene classification method based on a multimodal spatial transformation network according to claim 1, It is characterized in that The semantic information of remote sensing image categories is extracted through the GloVe model and the multi-head self-attention mechanism, including: Embed the labels of K scene categories into an m-dimensional vector space R through the GloVe model to obtain K semantic feature vectors S 1 , S 2 , …, S k , select the word vector S of the category label corresponding to the input image category label i , segment it into n segments and perform copy expansion to obtain the word vector X i Perform the multi-head self-attention mechanism operation as the input of the multi-head self-attention model. Among them, the scaled dot-product attention mechanism calculates the attention score by performing the dot-product operation on the query vector Q, the key vector K, and the value vector V, normalizes the attention score, and calculates the output result by weighted summation of V to obtain the semantic information vector of the class name. Then use the fully connected layer to convert the semantic information vector of the class name into the specified dimension, and finally use the Sigmoid activation function for processing to obtain the deep semantic information of the class name.
5. The remote sensing image scene classification method based on a multimodal spatial transformation network according to claim 1, It is characterized in that The classification loss of the image includes: Among them, L img represents the classification loss of the image, N represents the number of samples, represents the probability that the classification label of the i-th sample output by the pre-trained network belongs to the true label y i of.
6. A remote sensing image scene classification method based on a multi-modal spatio-temporal transformation network according to claim 1, characterized in that, the classification loss of the text includes: Among them, L txt represents the classification loss of the text, N represents the number of samples, represents the probability that the classification label of the i-th sample output by the class name embedding module belongs to the true label y i of.
7. A remote sensing image scene classification method based on a multi-modal spatio-temporal transformation network according to claim 1, characterized in that, the similarity loss includes: Among them, L sim represents the similarity loss, and f 1 and f 2 respectively represent the overall feature obtained by the image passing through the pre-trained network and the discriminative feature obtained by fusing the semantic information of the class name.
8. A remote sensing image scene classification method based on a multi-modal spatio-temporal transformation network according to claim 1, characterized in that, the loss function of the model includes: Among them, L img and L txt respectively represent the classification losses of images and texts, N represents the number of samples, and respectively represent the probabilities that the classification label of the i-th sample output by the pre-trained network and the class name embedding module belongs to the true label y i , L sim represents the similarity loss, f 1 and f 2 respectively represent the overall feature obtained by the image through the pre-trained network and the discriminative feature obtained by fusing the semantic information of the class name.
Citation Information
Patent Citations
Remote sensing image scene classification method based on multi-similarity measurement deep learning
CN111723675A
Cross-modal retrieval method based on modal relation learning
CN114817673A