Remote sensing image cross-modal retrieval method based on deep learning
By extracting the image texture features and local features of remote sensing images, combined with the enhanced self-attention and cross-attention mechanism, the problems of insufficient utilization of local features and semantic alignment in cross-modal retrieval of remote sensing images are solved, and high-precision cross-modal retrieval is achieved.
Patent Information
- Application Number
- CN202510481048.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-15
AI Technical Summary
Existing cross-modal retrieval methods for remote sensing images cannot effectively utilize the local features of images and text, and it is difficult to achieve deep semantic alignment, resulting in low retrieval accuracy.
By extracting the image texture features and local features of the remote sensing image, combining the graph attention network to enhance feature interaction, and using the enhanced self-attention and cross-attention mechanism to achieve the fusion of global and local features, and obtain the fine-grained alignment of the remote sensing image with the text description.
It significantly improves the accuracy and robustness of cross-modal retrieval, enhances the model's perceptual ability in complex contexts, and can generate accurate cross-modal retrieval results.
Smart Images

Figure CN120492660A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image retrieval, and in particular to a cross-modal retrieval method for remote sensing images based on deep learning. Background Art
[0002] The cross-modal retrieval of remote sensing images aims to more efficiently locate relevant remote sensing images from massive image collections by taking input text and extracting the corresponding remote sensing images from a database. This technology plays a crucial role in address surveying, disaster relief, and other fields. However, due to the complex backgrounds of remote sensing images and the significant modality differences between images and text, existing methods still struggle to achieve accurate cross-modal retrieval of remote sensing images. Previous approaches typically employ convolutional neural networks to extract remote sensing image features, then employ the BERT model to encode these features. A contrastive loss function is then used to calculate and optimize the differences between image and text features. However, these methods also have limitations. Existing cross-modal image retrieval methods only leverage the global features of images and text for learning, neglecting to leverage local features within images and text, which are crucial for remote sensing image retrieval. Furthermore, existing retrieval methods struggle to achieve deep interaction between image and text features, hindering deep semantic alignment and resulting in low retrieval accuracy. Summary of the Invention
[0003] The main purpose of this application is to provide a cross-modal retrieval method for remote sensing images based on deep learning, aiming to solve the problem of low accuracy of existing model retrieval.
[0004] To achieve the above-mentioned objectives, the present application provides a cross-modal retrieval method for remote sensing images based on deep learning, comprising: obtaining image texture features of remote sensing images; performing feature extraction on the remote sensing images to obtain global features and multiple local feature representations of the image; performing feature fusion using all local feature representations and image texture features to obtain a first fused feature; performing feature fusion on the global features of the image and the first fused feature to obtain a visual embedding; performing feature extraction on text information to obtain global features and multiple local features of the text; performing feature fusion on the global features of the text and the local features of the text to obtain a text embedding; determining a similarity score between the visual embedding and the text embedding, and generating a ranking result for cross-modal retrieval based on the similarity score.
[0005] Optionally, feature fusion is performed using multiple local feature representations and image texture features to obtain a first fused feature, including: multiplying all local feature representations with the image texture features respectively to obtain multiple channel features; determining the similarity between the channel features, and based on the similarity, using a preset number of channel features as local features of the image; using the local feature representation to enhance the local features of the image to obtain an enhanced feature; and fusing the enhanced feature with the image texture feature to obtain the first fused feature.
[0006] Optionally, local features of the image are enhanced using local feature representations to obtain enhanced features, including: constructing a graph structure using all local feature representations and local features of the image; in the graph structure, enhancing the features of the local features of the image using neighbor nodes of the local features to obtain enhanced features; wherein the neighbor nodes are local feature representations.
[0007] Optionally, the enhanced feature and the image texture feature are fused, including: performing an enhanced self-attention operation on the enhanced feature to obtain a first self-attention feature; and performing an enhanced cross-attention operation on the first self-attention feature and the image texture feature to obtain a first fused feature.
[0008] Optionally, the global features of the text and the local features of the text are fused, including: performing an enhanced self-attention operation on the local features of the text to obtain a second self-attention feature; An enhanced cross-attention operation is performed on the second self-attention feature and the global feature of the text to obtain a second fusion feature; and the global feature of the text is fused with the second fusion feature.
[0009] Optionally, the enhanced self-attention operation adopts a multi-head attention mechanism, in which the results of the multi-head attention operation are weighted.
[0010] Optionally, the enhanced cross-attention operation adopts a multi-head attention mechanism, in which the results of the multi-head attention operation are weighted.
[0011] Optionally, obtaining image texture features of the remote sensing image includes: performing feature extraction on the remote sensing image using a convolution encoder to obtain image texture features.
[0012] Optionally, feature extraction is performed on the remote sensing image to obtain global features and multiple local feature representations of the image, including: A window attention encoder is used to extract features from remote sensing images to obtain multiple encoded features. The encoded feature at the first position is used as the global feature of the image, and the encoded features at all positions are represented as local features.
[0013] To achieve the above-mentioned objectives, the present application also provides a cross-modal retrieval device for remote sensing images based on deep learning, including: a local feature extraction module for obtaining image texture features of remote sensing images; a global feature extraction module for performing feature extraction on remote sensing images to obtain global features and multiple local feature representations of the image; a feature enhancement module for performing feature fusion using all local feature representations and image texture features to obtain a first fused feature; a visual feature fusion module for performing feature fusion on the global features of the image with the first fused feature to obtain a visual embedding; a feature extraction module for performing feature extraction on text information to obtain global features and multiple local features of the text; a text feature fusion module for performing feature fusion on the global features of the text with the local features of the text to obtain a text embedding; a retrieval module for determining a similarity score between the visual embedding and the text embedding, and generating a ranking result for cross-modal retrieval based on the similarity score.
[0014] Compared with the prior art, the present invention has the following advantages: The deep learning-based cross-modal retrieval method for remote sensing images of the present invention extracts image texture features, local feature representations and global features of remote sensing images respectively, and uses a graph attention network to enhance the interaction between image texture features and local features of the image. The model can effectively capture the relationship between multi-scale targets and targets in remote sensing images; at the same time, the global features and the enhanced local features are fused to achieve accurate image cross-modal retrieval in the global and local directions; by extracting the global and local features of text information and combining enhanced self-attention and enhanced cross-attention mechanisms, fine-grained alignment between remote sensing images and text descriptions is achieved, significantly improving the accuracy and robustness of cross-modal retrieval; through global-local feature fusion and enhanced attention mechanism, the model's perception of complex backgrounds is enhanced, and accurate cross-modal retrieval results can be obtained in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a flowchart of a cross-modal retrieval method for remote sensing images based on deep learning in this application; Figure 2 This is a flowchart of a cross-modal retrieval method for remote sensing images based on deep learning for this application; Figure 3 This is a structural diagram of the enhanced self-attention module in a cross-modal retrieval method for remote sensing images based on deep learning in this application; Figure 4 This is the image retrieval result of a cross-modal retrieval method for remote sensing images based on deep learning in this application; Figure 5 This application presents the text retrieval results of a cross-modal retrieval method for remote sensing images based on deep learning.
[0016] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0017] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0018] The first embodiment of the present invention provides a cross-modal retrieval method for remote sensing images based on deep learning, such as Figure 1 As shown, the specific steps include: Step S1, obtaining image texture features of remote sensing images; Specifically, a convolutional encoder is used to extract features from remote sensing images to obtain image texture features. For example, the convolutional encoder may be a ResNet50 network. For example, the image texture feature may be a texture or a geometric structure.
[0019] Step S2, performing feature extraction on the remote sensing image to obtain global features and multiple local feature representations of the image; Specifically, a window attention encoder is used to extract features from remote sensing images to obtain multiple encoded features. The encoded features at the first position are used as the global features of the image, and the encoded features at all positions are represented as local features. For example, the window attention encoder can be a Swin Transformer network, which extracts global visual features of remote sensing images through a sliding window mechanism. , the global visual features The feature of the first position in is regarded as the global feature of the image, and the features of the first 50 positions are regarded as local features. .
[0020] Step S3, using all local feature representations and image texture features to perform feature fusion to obtain a first fused feature; the details are as follows.
[0021] Step S31: represent all local features Image texture features Perform multiplication operations to obtain multiple channel features; Step S32, determine the similarity between each channel feature, and according to the similarity, use a preset number of channel features as local features of the image; specifically, sort all channel features according to the similarity, and use the first 40 channel features as the local features. Most relevant local feature representation ; Step S33, using the local feature representation to enhance the local features of the image to obtain enhanced features; the details are as follows.
[0022] Step S331, constructing a graph structure using all local feature representations and local features of the image; Step S332: In the graph structure, the neighboring nodes of the local features of the image are used to enhance the features to obtain enhanced features; wherein the neighboring nodes are local feature representations.
[0023] Exemplarily, the local features and local feature representations of all images can be input into a graph neural network, and a graph structure can be constructed through the graph neural network to enhance the contextual interaction of regional features. In the graph structure, the local features of the image collect the feature information of their neighboring nodes, enhance themselves, and obtain enhanced features.
[0024] Step S34: Fusing the enhanced feature and the image texture feature to obtain a first fused feature. The details are as follows.
[0025] Step S341, performing an enhanced self-attention operation on the enhanced feature to obtain a first self-attention feature; Step S342: perform an enhanced cross-attention operation on the first self-attention feature and the image texture feature to obtain a first fusion feature.
[0026] The enhanced self-attention operation utilizes a multi-head attention mechanism, in which the results of the multi-head attention operation are weighted. Furthermore, the enhanced self-attention module performs an enhanced self-attention operation on the enhanced features. A weighted operation is added to the traditional self-attention module to multiply the attention score of each head by its corresponding weight before inputting it into the activation function. The weights used in this weighted operation are trainable multi-head weights.
[0027] This means that in the self-attention and cross-attention mechanisms, each attention head can focus on different attention weights, that is, information from different spaces in the input sequence. However, some of this information is important, while some is irrelevant. To obtain more accurate features, we design trainable multi-head weights to weight the attention heads. Trainable multi-head weights are automatically learned variables that are automatically updated during model training.
[0028] The working process of the enhanced self-attention module is to perform multi-head attention calculation on the input features to obtain the attention score of each head, multiply the attention score of each head by the corresponding trainable weight, merge the results, pass the merged result through the activation function to obtain the output, multiply the output by the merged result again, input the FFN (feedforward neural network), input the result of the FFN output into the activation function, and obtain the first self-attention feature by summing the output and the result of the FFN output.
[0029] It is worth noting that the enhanced cross-attention operation is implemented through the enhanced cross-attention module. Similar to the enhanced self-attention operation, it also adds a weighted operation on the results of the multi-head attention operation to the traditional cross-attention module. The rest of the process is the same as the traditional cross-attention module and will not be repeated here.
[0030] Step S4, performing feature fusion on the global features of the image and the first fusion features to obtain a visual embedding; Specifically, the global features of the image are added to the first fusion features, and then projected and normalized to obtain the visual embedding.
[0031] In this embodiment, the graph neural enhancement network can effectively establish the relationship between the local features extracted by the convolutional network and the global features extracted by the SwinTransformer, enhance the overall expression ability of the visual features, and effectively capture the multi-scale targets and complex relationships in the remote sensing images; further, by enhancing the self-attention mechanism and the enhanced cross-attention mechanism, the complete global visual features can be effectively extracted from the remote sensing images, the perception ability of complex backgrounds is enhanced, and accurate cross-modal retrieval results can be generated in complex scenes.
[0032] Step S5, extracting features from the text information to obtain global features and local features of the text; Specifically, a text feature encoder is used to encode the text information to obtain multiple encoding features, and the encoding feature at the first position is used as the global feature of the text. , the encoding features of the first 40 positions are used as local features of the text . Exemplarily, the text feature encoder can be a BERT model.
[0033] Step S6: perform feature fusion on the global features of the text and the local features of the text to obtain text embedding; the details are as follows.
[0034] Step S61, performing an enhanced self-attention operation on the local features of the text to obtain a second self-attention feature; Step S62, performing an enhanced cross-attention operation on the second self-attention feature and the global feature of the text to obtain a second fusion feature; Step S63: Fusing the global features of the text with the second fusion features.
[0035] In this embodiment, global features and local features are extracted from remote sensing images and text information respectively, and these global features and local features are fused. By enhancing the self-attention and enhanced cross-attention mechanisms, high-quality representations of remote sensing images and text description features are obtained. By designing a graph neural network model to model the association between the local features obtained by the ResNet network and the local features obtained by the Swin Transformer network, the most critical visual features in the remote sensing image are dynamically captured, enabling the model to generate accurate cross-modal retrieval results in complex scenarios.
[0036] Step S7: Determine the similarity score between the visual embedding and the text embedding, and generate a ranking result for cross-modal retrieval based on the similarity score.
[0037] Specifically, the cosine similarity between visual embedding and text embedding is used as the similarity score, and the retrieval objects (remote sensing images or text information) are sorted according to the similarity score to obtain the retrieval results.
[0038] It is worth noting that the cross-modal retrieval method of remote sensing images based on deep learning in this embodiment is implemented through a retrieval network. The retrieval network includes an image feature extraction network for obtaining visual embedding and a text feature extraction network for obtaining text embedding. The image feature extraction network includes a convolutional encoder, a window attention encoder, a graph neural network, an enhanced self-attention module and an enhanced cross-attention module. The text feature extraction network includes a text feature encoder, an enhanced self-attention module and an enhanced cross-attention module. Before retrieval, the retrieval network needs to be trained in advance, and the contrast loss function and the triple loss function are used to optimize it during the training process. Among them, the contrast loss function optimizes the model's ability to distinguish between positive and negative sample pairs by calculating the similarity between visual embedding and text embedding; the triple loss function further optimizes the model's cross-modal retrieval performance by introducing boundary constraints on positive and negative samples.
[0039] Example 1 The network was built using the PyTorch framework on an RTX3090 GPU server and the CentOS operating system. The dataset used in the experiment is a dataset of RSITMD remote sensing images and their corresponding text information. The remote sensing images were uniformly resized to 224×224 pixels. The text information was preprocessed and word segmented. The dataset was split into training, validation, and test sets in a ratio of approximately 8:1:1 to ensure a balanced distribution of images and text descriptions across the training, validation, and test sets.
[0040] A cross-modal retrieval model for remote sensing images was trained on the training set, and the trained cross-modal retrieval model was tested on the cross-modal retrieval test set for remote sensing images. The results were evaluated using seven indicators: mR and R@K (where K = 1, 5, 10).
[0041] Table 1 Performance of the model on the RSITMD dataset
[0042] After many experiments, the experimental results show that this method can search for corresponding text descriptions or remote sensing images from the database based on remote sensing images or texts. The mR index reaches 39.8, and it is significantly ahead of the comparison methods in all comparison indicators.
[0043] A second embodiment of the present invention provides a cross-modal remote sensing image retrieval device based on deep learning, comprising: Local feature extraction module, used to obtain image texture features of remote sensing images; The global feature extraction module is used to extract features from remote sensing images and obtain global features and multiple local feature representations of the image; A feature enhancement module is used to perform feature fusion using multiple local feature representations and image texture features to obtain a first fused feature; A visual feature fusion module is used to fuse the global features of the image with the first fusion features to obtain a visual embedding; Feature extraction module, used to extract features from text information to obtain global features and multiple local features of the text; The text feature fusion module is used to fuse the global features of the text with the local features of the text to obtain text embedding; The retrieval module is used to determine the similarity score between the visual embedding and the text embedding, and generate the ranking results of the cross-modal retrieval based on the similarity score.
[0044] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A cross-modal retrieval method for remote sensing images based on deep learning, characterized in that: include: Acquiring image texture features of the remote sensing image; Performing feature extraction on the remote sensing image to obtain global features and multiple local feature representations of the image; Performing feature fusion using all the local feature representations and image texture features to obtain a first fused feature; Performing feature fusion on the global image feature and the first fusion feature to obtain a visual embedding; Perform feature extraction on text information to obtain global features and multiple local features of the text; Performing feature fusion on the global features of the text and the local features of the text to obtain text embedding; A similarity score between the visual embedding and the text embedding is determined, and a ranking result of the cross-modal retrieval is generated based on the similarity score.
2. The cross-modal retrieval method for remote sensing images based on deep learning according to claim 1, characterized in that: The step of performing feature fusion using all the local feature representations and the image texture features to obtain a first fused feature includes: Multiplying all the local feature representations with the image texture features to obtain multiple channel features; Determining the similarity between the channel features, and based on the similarity, using a preset number of channel features as local features of the image; Performing feature enhancement on local features of the image using the local feature representation to obtain enhanced features; The enhancement feature and the image texture feature are fused to obtain a first fused feature.
3. The cross-modal retrieval method for remote sensing images based on deep learning according to claim 2, characterized in that: The utilizing the local feature representation to perform feature enhancement on the local features of the image to obtain enhanced features includes: Utilize all local feature representations and local features of the image to build a graph structure; In the graph structure, the neighboring nodes of the local features of the image are used to enhance the features to obtain enhanced features; The neighbor nodes are represented by local features.
4. The cross-modal retrieval method for remote sensing images based on deep learning according to claim 2, characterized in that: The fusing of the enhanced features and the image texture features comprises: Performing an enhanced self-attention operation on the enhanced feature to obtain a first self-attention feature; An enhanced cross-attention operation is performed on the first self-attention feature and the image texture feature to obtain a first fusion feature.
5. The cross-modal retrieval method for remote sensing images based on deep learning according to claim 1, characterized in that: The feature fusion of the global features of the text and the local features of the text includes: Performing an enhanced self-attention operation on the local features of the text to obtain a second self-attention feature; Performing an enhanced cross-attention operation on the second self-attention feature and the global feature of the text to obtain a second fusion feature; The global feature of the text is fused with the second fusion feature.
6. The cross-modal retrieval method for remote sensing images based on deep learning according to claim 4 or 5, characterized in that: The enhanced self-attention operation adopts a multi-head attention mechanism, in which the results of the multi-head attention operation are weighted.
7. The cross-modal retrieval method for remote sensing images based on deep learning according to claim 4 or 5, characterized in that: The enhanced cross-attention operation adopts a multi-head attention mechanism, in which the results of the multi-head attention operation are weighted.
8. The cross-modal retrieval method for remote sensing images based on deep learning according to claim 1, characterized in that: The acquiring of the image texture features of the remote sensing image comprises: A convolutional encoder is used to extract features from the remote sensing image to obtain image texture features.
9. The cross-modal retrieval method for remote sensing images based on deep learning according to claim 1, characterized in that: The step of extracting features from the remote sensing image to obtain global features and multiple local feature representations of the image includes: A window attention encoder is used to extract features from the remote sensing image to obtain multiple coding features. The coding feature at the first position is used as the global feature of the image, and the coding features at all positions are represented as local features.
10. A remote sensing image cross-modal retrieval device based on deep learning, characterized in that: include: A local feature extraction module, used for obtaining image texture features of the remote sensing image; A global feature extraction module is used to extract features from the remote sensing image to obtain global features and multiple local feature representations of the image; A feature enhancement module, configured to perform feature fusion using the plurality of local feature representations and image texture features to obtain a first fused feature; A visual feature fusion module, configured to fuse the global features of the image with the first fusion feature to obtain a visual embedding; Feature extraction module, used to extract features from text information to obtain global features and multiple local features of the text; A text feature fusion module is used to fuse the global features of the text with the local features of the text to obtain text embedding; A retrieval module is configured to determine a similarity score between the visual embedding and the text embedding, and generate a ranking result of a cross-modal retrieval based on the similarity score.