Remote sensing image-text retrieval method based on multi-scale feature modeling and feature consistency enhancement

By employing multi-scale feature modeling and feature consistency enhancement methods, the problems of multi-scale feature fusion and cross-modal feature interaction in remote sensing images were solved, thereby improving the efficiency and accuracy of remote sensing image and text retrieval.

CN120929628APending Publication Date: 2025-11-11XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511031848.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

There is a lack of research on multi-scale feature fusion and regional feature extraction methods for remote sensing images. The interaction between image and text features is complex, and cross-modal feature alignment is difficult, resulting in low accuracy and efficiency of remote sensing image and text retrieval.

Method used

We employ a multi-scale feature modeling and feature consistency enhancement approach, utilizing ResNet-50 to extract image features and combining a multi-scale dilated fusion attention module and a triplet loss function to optimize the alignment process of image and text features.

Benefits of technology

It improves the accuracy and efficiency of remote sensing image retrieval, can better capture the details and global information of image content, reduces the amount of computation and storage requirements, and enhances retrieval performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929628A_ABST
    Figure CN120929628A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image-text retrieval algorithm based on multi-scale feature modeling and feature consistency enhancement, and aims to solve the problems of non-uniform distribution of remote sensing image objects, inconsistent target scales and the like and improve the remote sensing image-text retrieval performance. Firstly, multi-scale features of an image are extracted, and shallow details and deep global semantic information are covered. Afterwards, through a multi-scale cavity fusion attention module, in combination with a channel and space attention mechanism, feature expression is enhanced, and intra-modal differences are reduced; in the text aspect, a Glove model is used for obtaining word vectors, and context information is captured through a Bi-GRU encoder. And finally, performing feature alignment by adopting a triple loss function, and optimizing model parameters. Experiments show that the retrieval performance of the algorithm on a remote sensing image-text data set is remarkably superior to that of an existing method, and the algorithm is suitable for the fields of geographic information analysis, disaster monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information retrieval technology, and relates to remote sensing image and text retrieval, particularly to a remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement. Background Technology

[0002] With the development of satellite technology, remote sensing technology has been increasingly applied to fields such as oil resource exploration, geological structure analysis, and disaster detection. Especially in resource exploration, remote sensing technology, through platforms such as satellites, aircraft, and drones equipped with sensors, enables non-contact observation of coalfield areas, collecting vast amounts of geographic information data. The utilization rate of remote sensing survey images and technical documents in the oil exploration and development field is growing exponentially, and researchers' demand for knowledge retrieval is increasing daily. How to efficiently retrieve valuable information from massive amounts of remote sensing data has become a major challenge for research and production management personnel.

[0003] First, unlike natural images, remote sensing images typically exhibit uneven object distribution, inconsistent target scales, and complex backgrounds, posing a significant challenge to acquiring robust visual representations. Current research on multi-scale feature fusion and region feature extraction methods for remote sensing images is insufficient. Second, previous methods primarily focused on extracting textual features directly from the entire input query text, ignoring inherent redundancy in textual features and interactions between words in the query text. Remote sensing images contain many small-scale semantic objects, and their semantic representations are easily influenced by non-semantic components (such as background and irrelevant objects). Finally, the differences between image and text modal data pose a challenge to achieving better cross-modal feature alignment. High similarity between images in different scenes can easily lead to inaccurate visual semantic representations. To address the complex interactions between multi-scale image and text features, the improvement and optimization of fine-grained feature alignment methods are particularly urgent. Summary of the Invention

[0004] To overcome the shortcomings of the prior art, the present invention aims to provide a remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement. This method optimizes the fine-grained interaction between multi-scale remote sensing image and text features, solves problems such as uneven distribution of objects and inconsistent target scales in remote sensing images, improves the retrieval performance of the remote sensing image and text retrieval model, and meets the needs of scientific research and production management personnel to efficiently retrieve valuable information from massive amounts of remote sensing data.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A remote sensing image retrieval method based on multi-scale feature modeling and feature consistency enhancement includes the following steps:

[0007] Step 1: Read the remote sensing dataset, which consists of remote sensing images and their corresponding text descriptions. Extract local and global features from the remote sensing images, and extract initial text features from the text descriptions.

[0008] Step 2: Using a multi-scale dilated fusion attention module, multi-scale feature modeling and feature consistency enhancement are performed on the local and global features respectively, to obtain the enhanced local feature V. l and the enhanced global feature V g The multi-scale dilated fusion attention module includes a local feature enhancement module and a global feature enhancement module. The local feature enhancement module includes an attention fusion module and several parallel convolutional branches with different dilation rates. The global feature enhancement module includes an attention fusion module and a depthwise separable convolutional module.

[0009] Step 3: Use the triplet loss function to perform... V l and V g Feature alignment;

[0010] Step 4: Use the remote sensing dataset for training. After training, use remote sensing images or text descriptions as input to perform cross-modal remote sensing image and text retrieval.

[0011] In one embodiment, step 1, extracting local and global features from the remote sensing image, is implemented as follows:

[0012] ResNet-50, pre-trained on the remote sensing dataset AID, was used as an image feature extractor to extract features from remote sensing images. The outputs of ResNet-50's layer 0 and layer 1 were used. As local features, the outputs of layer2, layer3, and layer4 As global features, i = 2, 3, 4, j = 0, 1.

[0013] In one embodiment, initial text features are extracted from the text description. The implementation method is as follows:

[0014] Use a pre-trained GloVe model to obtain the word vector corresponding to each word;

[0015] Using Bi-GRU as the text encoder, the contextual relationships between words are learned to obtain initial text features.

[0016] In one embodiment, the local feature enhancement module performs multi-scale feature modeling and feature consistency enhancement on the local features to obtain enhanced local features V.l The implementation method is as follows:

[0017] To each Extract features to obtain feature V j j = 0, 1;

[0018] Inputting V0 into several parallel convolutional branches enables feature extraction at different scales, resulting in several output features. Feature T1 is obtained by concatenating the data along the channel dimension, with each branch configured with a different hole rate.

[0019] V1 and T1 are fed into the attention fusion module for feature enhancement to obtain feature F. fusion1 and F fusion2 The attention fusion module includes a channel attention module and a spatial attention module.

[0020] F fusion1 and F fusion2 Feature F is obtained by concatenating along the channel dimension. f Then F f F is obtained by concatenating V0 along the channel dimension. f ′, will F f After passing through an FC layer, the enhanced local feature V is obtained. l .

[0021] In one embodiment, the pair Extract features to obtain feature V j The implementation method is as follows:

[0022] First to Perform 1×1 convolution, then batch normalization layer followed by ReLU activation, and finally average pooling to obtain feature V. j .

[0023] In one embodiment, the number of convolutional branches is 4. The first branch uses a 1×1 convolutional kernel, the second branch uses a 3×3 convolutional kernel with a dilatation rate of 6, the third branch uses a 3×3 convolutional kernel with a dilatation rate of 12, and the fourth branch uses a 3×3 convolutional kernel with a dilatation rate of 18. Each convolutional branch is a convolutional block, which includes convolution, batch normalization, ReLU activation, and adaptive average pooling operations.

[0024] In one embodiment, V1 and T1 are respectively fed into the attention fusion module for feature enhancement to obtain feature F. fusion1 and F fusion2 , represented as:

[0025] F fusion1 =Concat(F ca1 ,F spa1 )

[0026] F fusion2 =Concat(F ca2 ,F spa2 )

[0027] Among them, F ca1 For the result of V1 passing through the channel attention module, F ca2 For T1, the result of the channel attention module, F spa1 For the result of V1 passing through the spatial attention module, F spa2 This is the result of T1 passing through the spatial attention module.

[0028] In one embodiment, the local feature enhancement module further includes a WSA model, wherein the F f After passing through an FC layer, the input is fed into the WSA model to establish the relationship between features of different granularities, represented as:

[0029] V l =Avg(WSA(F f 'W f +b f ))

[0030] Among them, W f and b f This represents the weights and biases of the FC layer.

[0031] In one embodiment, the global feature enhancement module performs multi-scale feature modeling and feature consistency enhancement on the global features to obtain the enhanced global features V. g The implementation method is as follows:

[0032] Each global feature is processed using three 1×1 convolutional layers. Obtain features

[0033] right Downsampling, for Upsampling, followed by Multi-scale visual features V are obtained by splicing. d ;

[0034]

[0035] F fusion3 =Concat(F ca3 ,F spa3 )

[0036] V d The enhanced global feature V is obtained by feeding it into the attention fusion module for feature enhancement. g .

[0037] In one embodiment, V is first... d The intermediate feature V3 is obtained by processing the data in a depthwise separable convolution module. Then, the intermediate feature V3 is fed into an attention fusion module for feature enhancement to obtain the enhanced global feature V. g .

[0038] In one embodiment, the intermediate feature V3 is fed into the attention fusion module for feature enhancement to obtain the enhanced global feature V. g The formula is as follows:

[0039] V g =Concat(F ca3 ,F spa3 )

[0040] Among them, F ca3 For the result of V3 through the channel attention module, F spa3 The results are from the spatial attention module.

[0041] In one embodiment, the triplet loss function is expressed as:

[0042]

[0043] Where m represents the margin parameter, [x] + =max(x,0), where Representing paired RS image features V and text features similarity, Representing unpaired RS image features V and text features The similarity. These are text features in a single batch that are not paired with RS image feature V. It is a single batch that does not match the text features Paired RS image features, RS image features V are composed of V g and V l It is obtained through addition.

[0044] Compared with the prior art, the beneficial effects of the present invention are:

[0045] Existing technologies often struggle to balance detail and global information when processing multi-scale remote sensing images, leading to inaccurate feature representations and impacting retrieval accuracy. This invention, however, employs a multi-scale feature modeling method, utilizing ResNet-50 to extract features at different levels. It combines rich shallow spatial information with deep global semantic information, further enhanced by a multi-scale dilated fusion attention module (MDFA). This approach comprehensively and accurately describes image content, effectively capturing both small target details and the overall scene, significantly improving the accuracy of feature representation.

[0046] In cross-modal feature interaction, existing technologies suffer from significant semantic differences and difficulties in interaction. This invention introduces a feature consistency enhancement algorithm, which strengthens important feature channels and spatial regions through channel attention and spatial attention mechanisms, making features more consistent and robust. Simultaneously, a vision-guided multimodal dynamic fusion module utilizes multi-scale features and region-of-interest features to promote fine-grained alignment of text and image features, effectively solving the cross-modal feature interaction challenge and improving retrieval performance.

[0047] Furthermore, existing retrieval algorithms suffer from limited efficiency and accuracy when dealing with complex remote sensing scenarios. This invention reduces computational and storage requirements and improves retrieval efficiency by optimizing feature extraction, enhancement, and fusion processes. Simultaneously, by employing triplet loss as the objective function and optimizing model parameters, retrieval accuracy is further enhanced. Experimental results demonstrate that this invention significantly outperforms existing methods on datasets such as RSICD and RSITMD, exhibiting greater application value.

[0048] In summary, this invention has significant advantages in multi-scale feature modeling, cross-modal feature interaction, and improved retrieval performance, providing strong support for the development of remote sensing image and text retrieval. Attached Figure Description

[0049] Figure 1 The flowchart of the remote sensing image retrieval network based on multi-scale feature modeling and feature consistency enhancement provided by this invention is shown below.

[0050] Figure 2 This is a diagram of the remote sensing image retrieval network structure based on multi-scale feature modeling and feature consistency enhancement provided by this invention.

[0051] Figure 3 This is a schematic diagram of the Scale Hole Fusion Attention Module (MDFA) structure of the present invention.

[0052] Figure 4 This is a heat map of the significant area of ​​the present invention.

[0053] Figure 5 The Rank-5 search results of this invention. Detailed Implementation

[0054] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and examples.

[0055] like Figure 1As shown, this invention is a remote sensing image-text retrieval method based on multi-scale feature modeling and feature consistency enhancement. The method utilizes main modules including an image feature extractor, a text feature extraction module, a multi-scale dilated fusion attention module (MDFA), and a feature alignment submodule. First, the invention uses an image feature extractor and a text feature encoding module to extract features from remote sensing images and text, respectively. Then, the multi-scale dilated fusion attention module (MDFA) enhances image feature representation and extracts multi-scale remote sensing image targets. Simultaneously, Bi-GRU is employed to capture bidirectional contextual information in sequential data, improving the model's feature representation capability and retrieval performance. During feature alignment, features from different modalities (images and text) are aligned in the feature space, meaning that image and text features with the same semantics are placed close together in the feature space. In the feature alignment submodule, by continuously optimizing the triplet loss, the model can learn more discriminative feature representations. These features can better capture the semantic relationships between images and text, improving the accuracy of cross-modal retrieval.

[0056] Following the above technical solution, refer to Figure 1 and Figure 2 The present invention specifically includes the following steps:

[0057] Step 1: Read the remote sensing dataset, loading the remote sensing images and their corresponding text description datasets, such as the RSITMD and RSICD datasets.

[0058] This step primarily involves loading remote sensing image data and corresponding textual descriptions from the storage medium. The two remote sensing datasets contain a large number of remote sensing images that record various geographical information about the Earth's surface; the textual descriptions are specific textual explanations of the scenes, features, and other content within the remote sensing images. For the images, this invention first scales them to 278×278 pixels, then randomly crops and rotates the images to enhance the training samples. The final size of the input image is 256×256 pixels. To ensure experimental consistency and prevent the model from being affected by overfitting caused by deep networks, this invention uses ResNet-50 as the visual feature extractor, with a visual embedding dimension of 512. The word embedding dimension is set to 300, and the hidden layers of the bidirectional GRU are set to 512. Reading this data provides the foundation for subsequent feature extraction and processing.

[0059] Step 2 involves cropping and rotating the input remote sensing image, followed by the extraction of visual features.

[0060] like Figure 2As shown, this invention uses ResNet-50 as the image feature extractor. Slightly different, this invention uses a ResNet-50 pre-trained on the AID dataset, which, compared to a model pre-trained on ImageNet, has a more refined ability to perceive edge features in remote sensing images. The shallow features of CNNs with smaller receptive fields contain fine-grained semantic information (texture, color, etc.). Utilizing this information can improve the stability of image-text retrieval. Given an input image... Where (H, W) is the resolution of the original image. This invention first utilizes the output features of ResNet-50's layer 0 and layer 1 as local features. Recorded as and That is, j = 0, 1. Similarly, this invention uses the semantic-level features output by layer 2, layer 3, and layer 4 as global features. Recorded as

[0061] Step 3: Input the preliminarily extracted visual feature vectors into MDFA for feature enhancement.

[0062] like Figure 3 As shown, a Multi-Scale Hollow Fusion Attention Module (MDFA) is presented. Addressing the issues of uneven object distribution and inconsistent target scales in remote sensing images, this invention designs a Multi-Scale Hollow Fusion Attention Module (MDFA). The MDFA aims to enhance feature representation by utilizing multiple hole rates and integrating channel and spatial attention mechanisms. This design meets the need to simultaneously capture detailed information and extensive contextual information in the image, which is crucial for complex remote sensing image retrieval tasks. The module consists of two parts: the first part is a local feature enhancement module, which enhances local features... Multi-scale feature modeling and feature consistency enhancement are performed to obtain the enhanced local features V. l The second part is the global feature enhancement module, which enhances global features. Multi-scale feature modeling and feature consistency enhancement are performed to obtain the enhanced global feature V. g .

[0063] In convolutional neural networks, different channels often represent different features or patterns. However, when processing complex images, some channels may contain more useful information than others. Traditional convolutional operations cannot dynamically adjust the importance of channels. Furthermore, the information contained in different spatial locations may have different importance. For example, the region of a target object may be more important than the background. Traditional convolutional operations treat all spatial locations equally, failing to highlight key regions.

[0064] To enhance visual feature representation, this invention introduces two approaches: firstly, channel attention: by evaluating the importance of each channel and making corresponding adjustments, useful feature channels can be strengthened while unimportant channels are suppressed, thereby improving the quality of feature representation; secondly, spatial attention: by focusing on the importance of specific regions in the image, it helps the model concentrate on processing those local regions that have the greatest impact on the final task. These two approaches will be described separately below.

[0065] A. Channel Attention Module

[0066] The channel attention module aims to identify which channels in the input feature map contain important information. It does this by performing global average pooling on the feature map across its spatial dimensions. And Global Max Pooling Then, after processing by a shared multilayer perceptron (MLP), the two pooling results are finally combined. and The sums are then passed through a sigmoid activation function to obtain the channel attention weights M. c Let the input features be F∈R. C ×H×W Where C is the number of channels, H is the height, and W is the width. The formula for the channel attention module is expressed as follows:

[0067]

[0068] B. Spatial Attention Module

[0069] The spatial attention module aims to identify which spatial locations in the input feature map contain important information. It does this by performing average pooling along the channel dimension. and max pooling The two pooling results are then concatenated along the channel dimension, passed through a convolutional layer and a sigmoid activation function, to obtain the spatial attention weights M. s The formula for the spatial attention module is expressed as follows:

[0070]

[0071] M s =σ(M s )

[0072]

[0073] Finally, the output feature maps from the channel attention and spatial attention mechanisms are fused using an element-wise addition operation to obtain the enhanced feature map. At this point, the feature maps that have undergone channel and spatial calibration are element-wise added to the original merged feature map to integrate and enhance related features. The enhanced feature map is then passed through a 1x1 convolutional layer for dimensionality reduction and integration, generating the final output feature map, as shown below:

[0074] F fusion =F ca +F spa

[0075] Where F ca F is the output of the channel attention module. spa This is the output of the spatial attention module.

[0076] Therefore, the feature enhancement in this step can be described as follows:

[0077] Step 31, Local Feature Enhancement.

[0078] like Figure 3 As shown, the present invention firstly addresses... Perform 1×1 convolution, then pass through batch normalization layers, activate using the ReLU activation function, and finally obtain the visual feature V through average pooling. j .

[0079]

[0080] V j =AvgPool(Z)

[0081] The local feature enhancement module uses four parallel convolutional branches to extract features at different scales. Each branch is configured with a different dilation rate to expand the receptive field and capture spatial information of different ranges. For the input feature V... j The first branch uses a 1×1 convolutional kernel, extracting features directly without changing the spatial scale. The second branch uses a 3×3 convolutional kernel with a dilation rate of 6, moderately expanding the receptive field. The third branch uses a 3×3 convolutional kernel with a dilation rate of 12, further expanding the receptive field to capture broader contextual information. The fourth branch uses a 3×3 convolutional kernel with a dilation rate of 18, providing the widest receptive field. After processing through four convolutional blocks, four output feature maps are obtained.

[0082]

[0083] Here, `conv2d_block` represents the k-th convolutional block, which includes convolution, batch normalization, ReLU activation, and adaptive average pooling operations. The output feature maps of the four convolutional blocks are concatenated along the channel dimension (dim=1) to obtain the features:

[0084] V1 and T1 are respectively added to the AttentionFusion module for feature enhancement to obtain F. fusion1 and F fusion2 :

[0085] F fusion1 =Concat(F ca1 ,F spa1 )

[0086] F fusion2 =Concat(F ca2 ,F spa2 )

[0087] The attention fusion module here includes a channel attention module and a spatial attention module, where F ca1 and F ca2 The results of V1 and T1 passing through the channel attention module are respectively, F spa1 and F spa2 These are the results for V1 and T1 respectively, processed through the spatial attention module. Then, F... fusion1 and F fusion2 Feature F is obtained by concatenating features along the channel dimension. f To enhance the feature representation, feature F f F is obtained by concatenating features of V0 along the channel dimension. f ′.

[0088] F f =Concat(F fusion1 ,F fusion2 )

[0089] F f ′=F f +V0

[0090] Ultimately, F f After passing through an FC layer, the enhanced local feature V is obtained. l , This indicates local features at a fine-grained level.

[0091] To enhance the semantic interaction between visual and textual features, this invention can further apply the WSA model to establish connections between features of different granularities, as follows:

[0092] V l =Avg(WSA(Ff 'W f +b f ))

[0093] In the formula, and This represents the weights and biases of the FC layer.

[0094] Step 32, global feature enhancement.

[0095] This invention posits that deeper features in a network contain more visual global information, while shallower features contain more visually salient spatial information. For example... Figure 2 As shown, this invention uses three 1×1 convolutional layers to process different global features. Obtain features

[0096]

[0097] in Let i represent the two-dimensional convolution of the i-th global feature, where i = 2, 3, 4. For global features, this invention... Downsampling, for Upsampling, followed by Multi-scale visual features V are obtained by splicing. d , represented as:

[0098]

[0099] F fusion3 =Concat(F ca3 ,F spa3 )

[0100] Downsample(·) and Upsample(·) represent downsampling and upsampling operations, respectively.

[0101] Then, in order to better grasp the overall information, this invention will V d The intermediate feature V3 is obtained by processing the data in a depthwise separable convolutional module. Then, the intermediate feature V3 is fed into an attention fusion module for feature enhancement to obtain the enhanced global feature V. g , means as follows:

[0102] V g =Concat(F ca3 ,F spa3 )

[0103] Among them, F ca3 For the result of V3 through the channel attention module, F spa3 The results are from the spatial attention module.

[0104] Finally, the enhanced global feature V g and the enhanced local feature V l are used to obtain the RS image feature V through an addition operation.

[0105] Step 4: Preprocess the input text description statement and extract the initial text features from the text description.

[0106] This step can be carried out in parallel with Steps 1 - Step 3.

[0107] The present invention uses a pre-trained GloVe model to obtain the word vectors corresponding to each word. GloVe is a pre-trained model for obtaining word vectors. It maps each word into a low-dimensional vector space, such that words with similar semantics are close in the vector space. In this step, each word in the text description sentence is input into the pre-trained GloVe model to obtain the word vector corresponding to each word. These word vectors can serve as the initial representation of the text features and provide a basis for subsequent text feature extraction.

[0108] For example, the original text is “There is a double-tower bridge over the river” (Chinese interpretation: There is a double-tower bridge over the river). After preprocessing, it is converted to “there is a double tower bridge overthe river.” The preprocessed text is mapped using the GloVe model word vector table, where each word is replaced with the corresponding 300-dimensional GloVe word vector in the table. See Table 1. Due to space limitations, only the first and last five dimensions of each word vector are selected.

[0109] Table 1 GloVe text vector table

[0110]

[0111]

[0112] Step 5: In order to achieve text-image alignment, the present invention extracts text features from the sentence.

[0113] First, the present invention obtains word vectors through pre-training with GloVe, and then uses a Bi-GRU as a text encoder to learn the context relationships between words. Given the text input T, the word vectors {τ1, τ2,..., τ L} are obtained. Then the present invention embeds them into vectors using GloVe, that is, e i =W e τ iNext, this invention uses Bi-GRU to obtain the outputs of different hidden layers, as shown below:

[0114]

[0115] and This represents the output of the i-th hidden layer. Ultimately, this invention yields the following text features:

[0116]

[0117] Where h i This represents the average value of the bidirectional output of the i-th layer. From this, the initial text features can be obtained.

[0118] Step 6: Use the triplet loss function to perform... V l and V g Feature alignment.

[0119] In the field of cross-modal RS retrieval, triplet loss, as one of the mainstream loss functions for multimodal feature matching, is often used for feature alignment. The purpose of triplet loss is to make the distance between a sample and a positive sample as close as possible, while also making the distance between a sample and a negative sample as far as possible. Throughout the training process, this invention uses triplet loss as the objective function, expressed as follows:

[0120]

[0121] Where m represents the margin parameter, [x] + =max(x,0), where Representing paired RS image features V and text features similarity, Representing unpaired RS image features V and text features The similarity. These are text features in a single batch that are not paired with RS image feature V. It is a single batch that does not match the text features Paired RS image features, RS image features V are composed of V g and V l It is obtained through addition.

[0122] Step 7: Use the remote sensing dataset for training. After training, use remote sensing images or text descriptions as input to perform cross-mode remote sensing image and text retrieval.

[0123] This invention effectively reduces modal differences between remote sensing images and corresponding text descriptions, significantly improving the accuracy and robustness of remote sensing image-text retrieval. It is applicable to fields such as geographic information analysis, disaster monitoring, and environmental assessment. To further demonstrate the advantages of the proposed remote sensing image-text retrieval algorithm based on multi-scale feature modeling and feature consistency enhancement, extensive comparative experiments were conducted. The model was trained and evaluated on the publicly available remote sensing image-text datasets RSICD and RSITMD, and a comprehensive comparison was made with existing remote sensing image-text retrieval methods. This invention also provides a visual demonstration of salient region masking, proving through experiments that the MDFA module can effectively focus on salient objects in RS images. Qualitative analysis of image-text retrieval shows that the retrieval performance and stability of this invention have achieved the expected results.

[0124] 1. Dataset

[0125] RSICD: The RSICD dataset is a large dataset specifically designed for remote sensing imagery tasks. It encompasses 10,921 remote sensing images collected from multiple platforms, including Google Earth and Baidu Maps. Each image has been resized to 224×224 pixels while preserving the diversity of the original resolutions. Each image is accompanied by five detailed descriptive sentences, totaling 54,605 ​​annotations and using 3,325 unique words. Furthermore, the dataset is subdivided into 30 scene categories, providing rich resources for the understanding and analysis of remote sensing images, and has wide applications in fields such as geological structure analysis, environmental monitoring, and disaster response.

[0126] RSITMD: The RSITMD dataset focuses on matching and retrieving remote sensing images with text information. It contains 4,743 remote sensing images collected from authoritative sources, each accompanied by detailed text descriptions and 1 to 5 keywords, totaling 23,715 annotations (21,829 of which are unique). The descriptions are more granular than those in RSICD. The dataset is divided into 32 scene categories, covering urban and natural landscapes, demonstrating high diversity and complexity.

[0127] 2. Evaluation Criteria

[0128] R@K (K = 1, 5, 10): R@K represents the percentage of correctly retrieved results within the top K positions in all retrieval attempts. This invention performs image-to-text and text-to-image retrieval modes and obtains three R@K values ​​for each mode: R@1, R@5, and R@10.

[0129] mR: mR represents the average of all six R@K values ​​for the two retrieval modes. This metric provides an overall evaluation indicator to measure the model's average performance across the entire retrieval task.

[0130] 3. Quantitative comparison

[0131] This invention compares the model presented in this paper with traditional image and text retrieval methods VSE++, CAMP-triplet, CAMP-bce, and current mainstream remote sensing image and text retrieval methods MTFN, AMFMN, HVSA, and SWAN on the RSICD and RSIMD datasets. Since the original papers on traditional image and text retrieval methods did not conduct experiments on RSICD and RSIMD, this invention, for a fair comparison, used the same image encoder and text encoder as the model presented in this invention, and conducted three experiments, averaging the results of the traditional image and text retrieval methods. Furthermore, this invention directly references the best results of the remote sensing image and text retrieval methods from the original papers. The algorithm was trained on the RSICD and RSIMD datasets respectively, and the experimental results are shown in Tables 1 and 2.

[0132] Table 2 Comparison of the algorithm with existing algorithms on the RSICD dataset.

[0133]

[0134]

[0135] Results on the RSICD dataset: Table 2 shows the experimental results on the RSICD dataset. From these results, it can be seen that the model presented in this paper exhibits significant improvements compared to the baseline model method SWAN. For example, R@1 is improved by 0.8% (7.41 vs 78.20) and 0.9% (5.56 vs 6.41) in sentence and image retrieval, respectively. Overall, the mR metric is improved by 0.6% (20.61 vs 21.24). Experimental results show that a medium-sized model using ResNet-50 as the backbone outperforms HVSA and SWAN. Therefore, it can be said that the method of this invention outperforms traditional remote sensing retrieval methods on the RSICD dataset, reflecting its superior performance in addressing visual-semantic imbalance.

[0136] Table 3 compares the algorithm with existing algorithms on the RSITMD dataset.

[0137]

[0138] Results on the RSITMD dataset: RSITMD has more fine-grained sentence descriptions than RSICD; therefore, its overall retrieval performance can be improved. Table 3 shows the performance on the RSITMD test set; the method of the present invention improves on almost all metrics. Thanks to the superiority of visual representation, the R@1 value for image retrieval reaches 17.84 on the RSITMD test set, an improvement of 4.5% (13.35 vs. 17.84), and overall, the mR metric improves by 0.8% (34.11 vs. 34.97). The results show that the model significantly improves R@1 and can outperform baseline model methods on mR. In summary, it can be seen that the model of the present invention has a high-performance advantage because it has substantial improvements compared to state-of-the-art methods in most metrics.

[0139] 4. Visualization of salient region masks

[0140] like Figure 4 As shown, this invention visualizes salient region masks to analyze the functionality of the MDFA module. The salient mask enables the network to extract salient features of an image and adaptively analyze queries that focus on which regions of the image are relevant.

[0141] The salient masks of five typical images are as follows Figure 4 As shown. In Figure 4 In (a), it is most likely to be described as "two white domed buildings in the center of the ground," therefore the model focuses more on the two white domed buildings in the corresponding prominent mask. Figure 4 In (b), the model emphasizes the white running track portion of the playground, representing the algorithm's conclusion from the dataset that people prefer to describe the relationship between the running track and the playground. Figure 4 In (c), as can be seen in this invention, the model focuses on the twin-tower bridge, and this sample fully demonstrates that the model believes people are more likely to start with the adjacent relationship between two intersecting bridge surfaces when describing an image. Figure 4 In example (d), strangely, the model focuses on the outer edge of the white rectangles around the playground, rather than the entire playground. This might be because the algorithm considers the shape of the white rectangles more representative. Figure 4 In (e), the model is highlighted on the six oil storage tanks, consistent with the labeled title. These experiments are sufficient to demonstrate that the MDFA module can effectively focus salient objects in RS images.

[0142] 5. Qualitative Analysis of Image and Text Retrieval

[0143] This invention selects and visualizes the top-5 search results, such as... Figure 5As shown, the five example sentences to be searched in sentence retrieval are: "the irregular shaped terminal building sits alongside runaways", "this is asphalt, roads, grass, buildings and many planes", "there is a bridge on the river with grass on both sides", "stadium consists of five baseball fields". Their corresponding Chinese semantics are: "This irregularly shaped terminal building stands side by side with those planes that deviate from the runway", "This is asphalt roads, roads, grass, buildings and many planes", "There is a bridge on the river, and there is grass on both sides of the bridge", "A stadium has five baseball fields". In image retrieval, the sentences corresponding to the first image retrieval result are: "a row of tennis courts are near a baseball field", "There are six tennis courts next to the baseball field", "Six tennis courts are next to the baseball field", "A stadium consists of five baseball fields", "Next to the orange baseball field is a parking lot". Their corresponding Chinese semantics are "A row of tennis courts is next to a baseball field", "There are six tennis courts next to the baseball field", "There are six tennis courts near the baseball field", "A stadium is composed of five baseball fields", "There is a parking lot next to the orange baseball field". Among them, the two sentences marked in red do not match the image description.The statements corresponding to the second image retrieval results are: "The roof of the church is cardboard and colored with brown and orange", "The roof of the church is wavy and is coloured with brown and orange", "The church roof is corrugated and coloured with brown and orange", "The roof of the church is brown and orange waves", "The roof of the church is corrugated, brown and orange". The corresponding Chinese semantics are respectively: "The roof of the church is made of cardboard and painted with brown and orange pigments", "The roof of the church is wavy and painted with brown and orange pigments", "The roof of the church is corrugated and painted with brown and orange pigments", "The roof of the church is a brown and orange wavy structure", "The roof of the church is corrugated and painted with brown and orange pigments". All five retrieved texts match the description of the image. The statements corresponding to the third image retrieval results are: "one side of the four table tennis courts is a bare ground", "one side of the four table tennis courts is a green playground", "The four table tennis courts are bare on one side and green on the other", "There is a bare place on one side of the four table tennis courts", "a path passes through the lawn, with several ponds and trees nearby". The corresponding Chinese semantics are respectively: "One side of the four table tennis tables is an empty ground", "One side of the four table tennis tables is a green playground", "One side of the four table tennis tables is an empty ground and the other side is a green lawn", "One side of the four table tennis tables is an empty place", "A path passes through the lawn and there are several ponds and trees nearby". Sentence retrieval refers to using an image as a query to search for matching texts. Similar to image retrieval, it uses text as a query to retrieve matching images. From. Figure 5 As can be seen in (a) and (b), the top 5 search results are highly similar and difficult to distinguish. In contrast, the method of this invention can obtain more accurate search results by enhancing visual and textual representations. In sentence retrieval, matching sentences to images without significant objects is a challenge. In image retrieval, the model of this invention exhibits good retrieval performance regardless of whether the retrieved image contains a salient target. In summary, the model of this invention can better achieve bidirectional retrieval of images and text.

Claims

1. A remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement, characterized in that, Includes the following steps: Step 1: Read the remote sensing dataset, which consists of remote sensing images and their corresponding text descriptions. Extract local and global features from the remote sensing images, and extract initial text features from the text descriptions. Step 2: Using a multi-scale dilated fusion attention module, multi-scale feature modeling and feature consistency enhancement are performed on the local and global features respectively, to obtain the enhanced local feature V. l and the enhanced global feature V g The multi-scale dilated fusion attention module includes a local feature enhancement module and a global feature enhancement module. The local feature enhancement module includes an attention fusion module and several parallel convolutional branches with different dilation rates. The global feature enhancement module includes an attention fusion module and a depthwise separable convolutional module. Step 3: Use the triplet loss function to perform... V l and V g Feature alignment; Step 4: Use the remote sensing dataset for training. After training, use remote sensing images or text descriptions as input to perform cross-modal remote sensing image and text retrieval.

2. The remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement according to claim 1, characterized in that, Step 1, which extracts local and global features from the remote sensing image, is implemented as follows: ResNet-50, pre-trained on the remote sensing dataset AID, was used as an image feature extractor to extract features from remote sensing images. The outputs of ResNet-50's layer 0 and layer 1 were used. As local features, the outputs of layer2, layer3, and layer4 As global features, i = 2, 3, 4, j = 0, 1.

3. The remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement according to claim 2, characterized in that, The local feature enhancement module performs multi-scale feature modeling and feature consistency enhancement on the local features to obtain the enhanced local features V. l The implementation method is as follows: To each Extract features to obtain feature V j j = 0, 1; Inputting V0 into several parallel convolutional branches enables feature extraction at different scales, resulting in several output features. Feature T1 is obtained by concatenating the data along the channel dimension, with each branch configured with a different hole rate. V1 and T1 are fed into the attention fusion module for feature enhancement to obtain feature F. fusion1 and F fusion2 The attention fusion module includes a channel attention module and a spatial attention module. F fusion1 and F fusion2 Feature F is obtained by concatenating along the channel dimension. f Then F f F is obtained by concatenating V0 along the channel dimension. f ′, will F f After passing through an FC layer, the enhanced local feature V is obtained. l .

4. The remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement according to claim 3, characterized in that, The pair Extract features to obtain feature V j The implementation method is as follows: First to Perform 1×1 convolution, then batch normalization layer followed by ReLU activation, and finally average pooling to obtain feature V. j .

5. The remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement according to claim 3, characterized in that, The number of convolutional branches is 4. The first branch uses a 1×1 convolutional kernel, the second branch uses a 3×3 convolutional kernel with a dilatation rate of 6, the third branch uses a 3×3 convolutional kernel with a dilatation rate of 12, and the fourth branch uses a 3×3 convolutional kernel with a dilatation rate of 18. Each convolutional branch is a convolutional block, which includes convolution, batch normalization, ReLU activation, and adaptive average pooling operations.

6. The remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement according to claim 3, characterized in that, The local feature enhancement module also includes a WSA model, the F f After passing through an FC layer, the input is fed into the WSA model to establish the relationship between features of different granularities, represented as: V l =Avg(WSA(F f ′W f +b f )) Among them, W f and b f This represents the weights and biases of the FC layer.

7. The remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement according to claim 2, characterized in that, The global feature enhancement module performs multi-scale feature modeling and feature consistency enhancement on the global features to obtain the enhanced global features V. g The implementation method is as follows: Each global feature is processed using three 1×1 convolutional layers. Obtain features right Downsampling, for Upsampling, followed by Multi-scale visual features V are obtained by splicing. d ; V d The enhanced global feature V is obtained by feeding it into the attention fusion module for feature enhancement. g .

8. The remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement according to claim 7, characterized in that, First, V d The intermediate feature V3 is obtained by processing the data in a depthwise separable convolution module. Then, the intermediate feature V3 is fed into an attention fusion module for feature enhancement to obtain the enhanced global feature V. g The formula is as follows: V g =Concat(F ca3 ,F spa3 ) Among them, F ca3 For the result of V3 through the channel attention module, F spa3 The results are from the spatial attention module.

9. The remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement according to claim 8, characterized in that, Step 1 involves extracting initial text features from the text description. The implementation method is as follows: Use a pre-trained GloVe model to obtain the word vector corresponding to each word; Using Bi-GRU as the text encoder, the contextual relationships between words are learned to obtain initial text features.

10. The remote sensing image and text retrieval method based on multi-scale feature modeling and feature consistency enhancement according to claim 8, characterized in that, The triplet loss function is expressed as follows: Where m represents the margin parameter, [x] + Let max(x,0) represent the expression where Representing paired RS image features V and text features similarity, Representing unpaired RS image features V and text features similarity, These are text features in a single batch that are not paired with RS image feature V, where V is a text feature in a single batch that is not paired with RS image feature V. Paired RS image features, RS image features V are composed of V g and V l It is obtained through addition.