Remote sensing image text retrieval method and device based on Mamba adapter and sharing prompt enhancement

By introducing Mamba adapter and shared prompt mechanism in remote sensing image text retrieval, the problems of high computing resource requirements, complex background influence and cross-modal interaction in the prior art are solved, and efficient and accurate remote sensing image text retrieval is achieved.

CN120144816AActive Publication Date: 2025-06-13CHINA UNIV OF MINING & TECH
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510307454.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Existing remote sensing image text retrieval methods face the problems of high computing resource requirements, susceptibility to complex backgrounds and lack of cross-modal interaction mechanisms, resulting in low retrieval accuracy and efficiency in remote sensing scenarios.

Method used

The CLIP enhancement network based on Mamba adapter and shared prompts is adopted to reduce training parameters through the adapter, enhance key feature extraction, and make up for the missing cross-modal interaction of the model through the shared prompt mechanism to achieve efficient correlation between images and text.

Benefits of technology

While reducing trainable parameters, the accuracy and efficiency of remote sensing image text retrieval is improved, and the content features of the image subject can be extracted more effectively and sensitive to the association between the image and the text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144816A_ABST
    Figure CN120144816A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image text retrieval method and device based on Mama adapter and sharing prompt enhancement. The remote sensing image text retrieval device comprises a CLIP enhancement network based on a multi-mode adapter and sharing prompt. A modal exclusive adapter based on a Mama structure is introduced into a CLIP trunk model, and the modeling capability of remote sensing image long-distance spatial features and text semantic association is enhanced, so that the focusing capability of a remote sensing image main body is enhanced; meanwhile, a dynamic prompt fusion module is provided, initial prompt feature vectors of an image and a text are extracted based on a pre-trained CLIP, modal exclusive dynamic prompts are generated through projection, a learnable shared prompt matrix is combined, composite prompt vectors fusing intra-modal features and cross-modal interaction information are constructed, and the composite prompt vectors are injected into a 12-layer Mama adapter of the CLIP, so that the dynamic prompt fusion is realized. And guiding the model to optimize cross-modal alignment. According to the method, the remote sensing feature extraction efficiency is improved through the Mama adapter, modal internal and external information collaboration is realized in combination with dynamic prompt, and the precision and generalization of remote sensing image-text retrieval are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a remote sensing image text retrieval method and device enhanced based on a Mamba adapter and shared prompts, and belongs to the technical field of remote sensing image text retrieval. Background Art

[0002] With the booming development of space exploration in China and globally, the application of remote sensing technology and data has become increasingly widespread and has become a key tool in multiple fields such as agricultural management, environmental monitoring, and urban planning. Remote sensing images, as the main means of obtaining surface information, are growing at an unprecedented rate in terms of data volume. However, in the face of a vast amount of remote sensing image data, how to quickly and accurately retrieve this data has become an urgent problem to be solved. Currently, image text retrieval technologies have shown great potential and effectiveness in addressing this challenge. They can understand the content of images and texts, perform matching between images and texts, and greatly promote the analysis and application of image data.

[0003] With the rapid development of vision-language foundation models, transfer learning based on large-scale pre-trained models has become the mainstream paradigm for remote sensing image text retrieval. Such models first extract cross-modal semantic features through pre-training on large-scale natural image-text pairs, and then fine-tune the models using image-text paired data in the remote sensing field, such as satellite images and geographic tags, disaster monitoring reports and radar images, etc. Since the remote sensing field has rich image-text paired data, this method has achieved quite good success in the field of remote sensing image text retrieval.

[0004] However, the existing methods still face three key challenges: 1. Vision-language foundation models usually contain billions of parameters (for example, the number of parameters of CLIP-ViT-L / 14 reaches 350 million), and full-parameter fine-tuning of them requires extremely high computing resources and time costs, which limits their lightweight deployment in remote sensing scenarios; 2. Remote sensing images have significant regional focusing characteristics (such as monitoring farmland requires attention to specific crop areas, and detecting geological disasters requires locating landslide areas), but existing models are vulnerable to complex backgrounds (such as cloud cover and terrain interference), resulting in difficulty in effectively extracting key feature information of ground objects; 3. Traditional vision-language models use independent encoders to process images and texts (such as the two-tower structure of CLIP), lacking a cross-modal interaction mechanism, while in remote sensing tasks, there is often a strong spatial correlation between images and texts (such as "river pollution in the northeast direction" requires text to guide the attention of the image area), and independent encoding may lose fine-grained semantic alignment information. Summary of the Invention

[0005] Objective of the Invention: To overcome the deficiencies in the prior art, the present invention provides a remote sensing image text retrieval method and device based on Mamba adapter and shared prompt enhancement, which comprehensively utilizes adapter and prompt technologies to reduce the required training parameters, and at the same time uses the Mamba adapter structure to enhance the extraction of key object features, and uses prompts with shared content to make up for the lack of cross-modal interaction of the model, so as to enhance the accuracy of remote sensing image text retrieval while using as few parameters as possible.

[0006] Technical Solution: To achieve the above objective, the technical solution adopted by the present invention is as follows:

[0007] A remote sensing image text retrieval method based on Mamba adapter and shared prompt enhancement, comprising the following steps:

[0008] Step1: Obtain the original remote sensing image and the corresponding text description (RS-Text), and perform cropping and standardization processing on the original remote sensing image to obtain a remote sensing image (RS-Image) of a set size;

[0009] Step2: Divide the remote sensing image and the corresponding text description into a training set and a test set according to a ratio, where: the training set includes remote sensing image I train and the corresponding text description T train , and the test set includes remote sensing image I test and the corresponding text description T test ;

[0010] Step3: Construct a CLIP enhanced network based on multi-modal adapter and shared prompt, including a prompt generation network and a feature extraction network; the prompt generation network includes a prompt vector extraction module and a prompt generation module, and the prompt vector extraction module includes an image prompt vector extractor and a text prompt vector extractor; the feature extraction network includes a feature extraction module and an adapter module, and the adapter module includes a group of image adapters and a group of text adapters. The feature extraction module is designed based on the CLIP model. The CLIP model includes 12 layers of TransformerBlock, and an image feature extractor is embedded at the image end of each Transformer Block layer, and a text adapter is embedded at the text end of each Transformer Block layer; replace the CLIP model parameters in the prompt vector extraction module and the feature extraction module with the pre-trained CLIP-B / 32 parameters;

[0011] Input I train and T train into the prompt generation network, and I train generates an image pre-prompt vector feature through the image prompt vector extractorp_img , T train Generate the text pre-prompt vector feature through the text prompt vector extractor p_text , and use feature p_img and feature p_text as the input of the prompt generation module;

[0012] Use feature p_img and feature p_text as the input of the prompt generation module to generate the final image prompt vector and the final text prompt vector j = 1, 2,..., 12;

[0013] Use I train and T train respectively as and the input of the feature extraction network. The feature extraction network includes 12 layers of Transformer Block; the image-side input of the j-th layer of Transformer Block is denoted as The image-side output is denoted as Use and as the input of the image adapter to generate the optimized image feature The text-side input of the j-th layer of TransformerBlock is denoted as The text-side output is denoted as Use and as the input of the text adapter to generate the optimized text feature

[0014] Use 's cls_token marker as the input of the linear projection layer, and output the image feature vector F after dimensionality reduction img , use 's [EOS] marker as the input of the linear projection layer, and output the text feature vector F after dimensionality reduction text ;

[0015] Step4: Calculate the loss value of F img and F text using the weighted combination of symmetric cross-modal contrast loss and membership loss;

[0016] Step5: Freeze the parameters of the prompt vector extraction module and the feature extraction module, and repeat Step3~Step4 until the loss value converges to complete the training of the prompt generation module and the adapter module;

[0017] Step6: Use I test and T testInput prompt generation network, output the final prompt vector of the image and the final prompt vector of the text Take I test 、T test 、 and Input feature extraction network, output F img and F text ;

[0018] Step7: Based on F img and F text Perform bidirectional cross-modal retrieval, calculate the retrieval performance metrics and output them.

[0019] Specifically, in the prompt vector extraction module, the image prompt vector extractor includes a PatchEmbedding layer, a Position Embedding layer, a learnable embedding vector P class ′, 12 sequentially connected Transformer Encoder layers and a linear projection layer; Input I train into the image prompt vector extractor, and it is split and flattened into a vector sequence P img ′ by the Patch Embedding layer. Concatenate P img ′ and P class ′, and then output through the PositionEmbedding layer The input of the j-th Transformer Encoder layer is The output of the j-th Transformer Encoder layer is Put into the linear projection layer to transform the dimension, and the output is feature p_i ;

[0020] In the prompt vector extraction module, the text prompt vector extractor includes a tokenizer, a TokenEmbedding layer, a Position Embedding layer, 12 sequentially connected Transformer Encoder layers and a linear projection layer; Input T train into the text prompt vector extractor, and it is tokenized by the tokenizer to output T token ′, output P text ′ through the TokenEmbedding layer, and output through the Position Embedding layer The input of the j-th TransformerEncoder layer is The output of the j-th Transformer Encoder layer is Put The dimension is transformed in the input linear projection layer, and the output is the feature p_t .

[0021] Specifically, the prompt generation module includes an image linear layer, a text linear layer, and a shared prompt matrix; the feature p_i and the feature p_i are input into the prompt generation module. The feature p_i generates the feature p_i ' through the image linear layer, and the feature p_t generates the feature p_t ' through the text linear layer. The feature p_i ' is concatenated with the shared prompt matrix to output the final image prompt vector The feature p_t ' is concatenated with the shared prompt matrix to output the final text prompt vector

[0022] Specifically, in the prompt generation module, the dimension of the image linear layer is [512, 128], the dimension of the text linear layer is [512, 128], and the dimension of the shared prompt matrix is [12, 5, 128].

[0023] Specifically, in the feature extraction module, the image end of the Transformer Block layer is called the image feature extractor, and the text end of the Transformer Block layer is called the text feature extractor;

[0024] The image feature extractor includes a Patch Embedding layer, a Position Embedding layer, a learnable embedding vector P class , 12 sequentially connected Transformer Encoder layers, and a linear projection layer; the T train is input into the image feature extractor and is segmented and flattened into a vector sequence P img through the Patch Embedding layer. P img and P class are concatenated and then output through the Position Embedding layer The input of the j-th Transformer Encoder layer is The output of the, -th Transformer Encoder layer is The and are input into the image adapter to generate the image optimized feature The The cls_token is input into the linear projection layer to transform the dimension, and the output is the image feature vector F img ;

[0025] The text feature extractor includes a tokenizer, a Token Embedding layer, a Position Embedding layer, 12 Transformer Encoder layers and a linear projection layer; input T train into the text feature extractor, and after tokenization by the tokenizer, output T token , and after passing through the Token Embedding layer, output P text , and after passing through the Position Embedding layer, output The input of the j-th Transformer Encoder layer is The output of the j-th Transformer Encoder layer is Input and into the text adapter to generate the text optimized feature Input 's [EOS] token into the linear projection layer to transform the dimension, and the output is the text feature vector F text .

[0026] Specifically, in the adapter module, the image adapter includes a dimensionality reduction linear layer, a data preprocessing layer, two TreeMamba deep feature extraction layers, two residual connections, a gated mechanism connection and a dimensionality increase linear layer. The data preprocessing layer converts the input from the length dimension to the length and width dimensions;

[0027] Input and into the image adapter, and output after passing through the dimensionality reduction linear layer Concatenated with and output after passing through the GELU activation function After dimensionality conversion by the data preprocessing layer, output Output successively through two TreeMamba deep feature extraction layers Perform gated mechanism connection on and and output Perform residual connection on and and output Remove belonging to The Token obtained Output through the dimensionality-increasing linear layer For and Perform residual connection and output

[0028] Specifically, in the adapter module, the text adapter includes a dimensionality-reducing linear layer, two Hydra deep feature extraction layers, a data preprocessing layer, two residual connections, a gated mechanism connection, and a dimensionality-increasing linear layer;

[0029] Put and into the text adapter, Output through the dimensionality-reducing linear layer Concatenate with and output through the GELU activation function Output successively through two Hydra deep feature extraction layers For and Perform gated mechanism connection and output For and Perform residual connection and output Remove The Tokens belonging to to obtain Output through the dimensionality-increasing linear layer For, and Perform residual connection and output

[0030] Specifically, in the image adapter, the dimension of the dimensionality-reducing linear layer is [768, 128], the receptive dimension of the TreeMamba deep feature extraction layer is 128, the SSM spatial feature dimension is 1, and the dimension of the dimensionality-increasing linear layer is [128, 768]; in the text adapter, the dimension of the dimensionality-reducing linear layer is [512, 128], the receptive dimension of the Hydra deep feature extraction layer is 128, the SSM spatial feature dimension is 16, and the dimension of the dimensionality-increasing linear layer is [128, 512].

[0031] A remote sensing image text retrieval device based on Mamba adapter and shared prompt enhancement, including a remote sensing image and text acquisition unit, a CLIP enhancement network based on multi-modal adapter and shared prompt, a training unit, and a testing unit;

[0032] The remote sensing image and text acquisition unit is used to acquire the original remote sensing image and the corresponding text description, and perform cropping and standardization processing on the original remote sensing image to obtain a remote sensing image of a set size;

[0033] The CLIP enhanced network based on multi-modal adapter and shared prompt includes a prompt generation network and a feature extraction network; the prompt generation network includes a prompt vector extraction module and a prompt generation module. The image pre-prompt vector and text pre-prompt vector are respectively extracted from the remote sensing image and the corresponding text description through the prompt vector extraction module, and then through the prompt generation module, the image final prompt vector and text final prompt vector are generated; the feature extraction network includes a feature extraction module and an adapter module. The feature extraction module and the adapter module work together. The feature extraction module extracts the shallow features of the remote sensing image and the corresponding text description, and the adapter module extracts the deep features of the remote sensing image and the corresponding text description, and finally generates the feature vectors of the remote sensing image and the corresponding text description;

[0034] The training unit uses a loss function to supervise and train the CLIP enhanced network based on multi-modal adapter and shared prompt;

[0035] The testing unit uses the trained CLIP enhanced network based on multi-modal adapter and shared prompt to perform image-to-text retrieval and text-to-image retrieval on paired remote sensing images and corresponding text descriptions.

[0036] Specifically, in the CLIP enhanced network based on multi-modal adapter and shared prompt, it includes a prompt generation network and a feature extraction network; the prompt generation network includes a prompt vector extraction module and a prompt generation module, and the prompt vector extraction module includes an image prompt vector extractor and a text prompt vector extractor; the feature extraction network includes a feature extraction module and an adapter module, and the adapter module includes a group of image adapters and a group of text adapters. The feature extraction module is designed based on the CLIP model. The CLIP model includes 12 Transformer Block layers, and an image feature extractor is embedded at the image end of each Transformer Block layer, and a text adapter is embedded at the text end of each Transformer Block layer; the CLIP model parameters in the prompt vector extraction module and the feature extraction module are replaced with the pre-trained CLIP-B / 32 parameters.

[0037] Specifically, in the CLIP enhanced network based on multi-modal adapter and shared prompt:

[0038] Input I train and T train into the prompt generation network, Itrain Generate the image pre-prompt vector feature through the image prompt vector extractor p_img , T train Generate the text pre-prompt vector feature through the text prompt vector extractor p_text , and use feature p_img and feature p_text as the input of the prompt generation module;

[0039] Use feature p_img and feature p_text as the input of the prompt generation module to generate the final image prompt vector and the final text prompt vector j = 1, 2,..., 12;

[0040] Use I train and T train as the inputs of and respectively into the feature extraction network. The feature extraction network includes 12 layers of Transformer Block; the image-side input of the j-th layer of Transformer Block is denoted as The image-side output is denoted as Use and as the input of the image adapter to generate the image optimized feature The text-side input of the j-th layer of Transformer Block is denoted as The text-side output is denoted as Use and as the input of the text adapter to generate the text optimized feature

[0041] Use 's cls_token marker as the input of the linear projection layer for dimensionality reduction and output the image feature vector F img , use 's [EOS] marker as the input of the linear projection layer for dimensionality reduction and output the text feature vector F text .

[0042] Beneficial effects: The remote sensing image text retrieval method and device based on Mamba adapter and shared prompt enhancement provided by the present invention establish a joint retrieval model by introducing an adapter based on Mamba and a prompt with a sharing mechanism into CLIP, perform image-text retrieval on remote sensing image text, greatly reduce the number of trainable parameters, enable the model to extract more effective features of the main content of the image and be more sensitive to the association between the image and the text, and have a higher accuracy for remote sensing image text retrieval. Description of the Drawings

[0043] Figure 1 Schematic diagram of the frame structure of the device of the present invention;

[0044] Figure 2 Flow chart of the implementation of the method of the present invention;

[0045] Figure 3 Pair of remote sensing images and corresponding text descriptions;

[0046] Figure 4 Prediction effect of the trained CLIP enhanced network based on multi-modal adapter and shared prompt, where the text in the rightmost box is the correctly retrieved text. Detailed implementation manners

[0047] The present invention will be specifically introduced below in conjunction with the accompanying drawings and specific embodiments.

[0048] As Figure 1 shown, the remote sensing image text retrieval device enhanced based on Mamba adapter and shared prompt includes a remote sensing image and text acquisition unit, a CLIP enhanced network based on multi-modal adapter and shared prompt, a training unit and a testing unit.

[0049] The remote sensing image and text acquisition unit is used to acquire the original remote sensing image and the corresponding text description, and perform cropping and standardization processing on the original remote sensing image to obtain a remote sensing image of a set size.

[0050] The CLIP enhanced network based on multi-modal adapter and shared prompt includes a prompt generation network and a feature extraction network; the prompt generation network includes a prompt vector extraction module and a prompt generation module, which respectively extract the image pre-prompt vector and text pre-prompt vector from the remote sensing image and the corresponding text description through the prompt vector extraction module, and then generate the image final prompt vector and text final prompt vector through the prompt generation module; the feature extraction network includes a feature extraction module and an adapter module, and the feature extraction module and the adapter module work together. The feature extraction module extracts the shallow features of the remote sensing image and the corresponding text description, and the adapter module extracts the deep features of the remote sensing image and the corresponding text description, and finally generates the feature vectors of the remote sensing image and the corresponding text description.

[0051] The training unit uses a loss function to supervise and train the CLIP enhanced network based on multi-modal adapter and shared prompt.

[0052] The testing unit uses the trained CLIP enhanced network based on multi-modal adapter and shared prompt to perform image-to-text retrieval and text-to-image retrieval on pairs of remote sensing images and corresponding text descriptions.

[0053] Use Figure 1 the remote sensing image text retrieval device enhanced by Mamba adapter and shared prompt shown in the figure to perform multi-source remote sensing image text retrieval. The process is as follows Figure 2 shown. The following will be described in detail in combination with each step

[0054] Step 1: Obtain remote sensing images and corresponding text descriptions

[0055] Obtain the original remote sensing images and corresponding text descriptions (RS-Text), and perform cropping and standardization processing on the original remote sensing images to obtain remote sensing images (RS-Image) with the same size for subsequent processing. As Figure 3 shown, the paired remote sensing images and corresponding text descriptions obtained are shown

[0056] Step2: Divide the training set and the test set

[0057] Divide the remote sensing images and corresponding text descriptions into a training set and a test set according to a ratio of 7:3. Among them, the training set includes remote sensing images I train and corresponding text descriptions T train , and the test set includes remote sensing images I test and corresponding text descriptions T test .

[0058] Step3: Construct a CLIP enhancement network based on multi-modal adapter and shared prompt

[0059] The CLIP enhancement network based on multi-modal adapter and shared prompt includes a prompt generation network and a feature extraction network

[0060] 3.1 Structure of the prompt generation network

[0061] The said prompt generation network contains a prompt vector extraction module and a prompt generation module

[0062] 3.1.1 Structure of the prompt vector extraction module

[0063] The said prompt vector extraction module contains an image prompt vector extractor and a text prompt vector extractor

[0064] The said image prompt vector extractor contains a Patch Embedding layer, a PositionEmbedding layer, a learnable embedding vector P class ′, 12 consecutive Transformer Encoder layers in series, and a linear projection layer; Input I train into the image prompt vector extractor, and it is segmented and flattened into a vector sequence P by the Patch Embedding layerimg ′, concatenate P img ′ and P class ′, and then output through the Position Embedding layer The input of the j-th Transformer Encoder layer is The output of the j-th Transformer Encoder layer is Put into the linear projection layer to transform the dimension, and the output is feature p_i .

[0065] The text prompt vector extractor includes a tokenizer, a Token Embedding layer, a Position Embedding layer, 12 sequentially concatenated Transformer Encoder layers, and a linear projection layer; input T train into the text prompt vector extractor, tokenize it through the tokenizer to output T token ′, output P text ′ through the Token Embedding layer, and output through the Position Embedding layer The input of the j-th Transformer Encoder layer is The output of the j-th Transformer Encoder layer is Put into the linear projection layer to transform the dimension, and the output is feature p_t .

[0066] 3.1.2 Structure of the prompt generation module

[0067] The prompt generation module includes an image linear layer, a text linear layer, and a shared prompt matrix; input feature p_i and feature p_t into the prompt generation module, generate feature p_i ′ through the image linear layer, generate feature p_i ′ through the text linear layer, concatenate feature p_t ′ with the shared prompt matrix to output the final image prompt vector p_t ′, concatenate feature p_i ′ with the shared prompt matrix to output the final text prompt vector feature p_t ′, concatenate with the shared prompt matrix to output the final text prompt vector

[0068] In the prompt generation module, the dimension of the image linear layer is [512, 128], the dimension of the text linear layer is [512, 128], and the dimension of the shared prompt matrix is [12, 5, 128].

[0069] 3.2 Structure of the Feature Extraction Network

[0070] The feature extraction network includes a feature extraction module and an adapter module. The adapter module includes a group of image adapters and a group of text adapters. The feature extraction module is designed based on the CLIP model. The CLIP model includes 12 Transformer Block layers. An image feature extractor is embedded at the image end of each Transformer Block layer, and a text adapter is embedded at the text end of each Transformer Block layer.

[0071] 3.2.1 Structure of the Feature Extraction Module

[0072] In the feature extraction module, the image end of the Transformer Block layer is called the image feature extractor, and the text end of the Transformer Block layer is called the text feature extractor.

[0073] The image feature extractor includes a Patch Embedding layer, a Position Embedding layer, a learnable embedding vector P class , and 12 sequentially connected Transformer Encoder layers and a linear projection layer; input T train into the image feature extractor, and after being segmented and flattened by the Patch Embedding layer into a vector sequence P img , splice P img and P class , and then output through the Position Embedding layer The input of the j-th Transformer Encoder layer is The output of the, -th Transformer Encoder layer is Input and into the image adapter to generate the optimized image feature Input 's cls_token tag into the linear projection layer to transform the dimension, and the output is the image feature vector F img .

[0074] The text feature extractor includes a word segmenter, a Token Embedding layer, a Position Embedding layer, 12 Transformer Encoder layers, and a linear projection layer; input T train into the text feature extractor, and after word segmentation by the word segmenter, output T token , and after passing through the Token Embedding layer, output P text , and after passing through the Position Embedding layer, output The input of the j-th Transformer Encoder layer is The output of the j-th Transformer Encoder layer is Input and into the text adapter to generate optimized text features Input 's [EOS] token into the linear projection layer to transform the dimension, and the output is the text feature vector F text .

[0075] 3.2.2 Structure of the Adapter Module

[0076] The adapter module includes a group of image adapters and a group of text adapters.

[0077] The image adapter includes a dimensionality reduction linear layer, a data preprocessing layer, two TreeMamba deep feature extraction layers, two residual connections, a gated mechanism connection, and a dimensionality increase linear layer. The data preprocessing layer converts the input from the length dimension to the length and width dimensions; input and into the image adapter and output after passing through the dimensionality reduction linear layer Concatenated with and output after activation by the GELU activation function Output after dimensionality conversion by the data preprocessing layer Output sequentially through two TreeMamba deep feature extraction layers Perform gated mechanism connection on and and output Perform residual connection on and and output Remove the Tokens belonging to from to get Output of the upsampling linear layer For And Perform residual connection and output

[0078] In the image adapter, the dimension of the downsampling linear layer is [768, 128], the receptive dimension of the TreeMamba deep feature extraction layer is 128, the SSM spatial feature dimension is 1, and the dimension of the upsampling linear layer is [128, 768].

[0079] The text adapter contains a downsampling linear layer, two Hydra deep feature extraction layers, a data preprocessing layer, two residual connections, a gated mechanism connection, and an upsampling linear layer; input And Into the text adapter, Output after passing through the downsampling linear layer Concatenate with And output after passing through the GELU activation function Output successively after passing through two Hydra deep feature extraction layers For And Perform gated mechanism connection and output For And Perform residual connection and output Remove The Tokens in Belonging to Output after passing through the upsampling linear layer For, And Perform residual connection and output

[0080] In the text adapter, the dimension of the downsampling linear layer is [512, 128], the receptive dimension of the Hydra deep feature extraction layer is 128, the SSM spatial feature dimension is 16, and the dimension of the upsampling linear layer is [128, 512].

[0081] 3.3 Execution process of the CLIP enhanced network based on the multimodal adapter and shared prompt

[0082] Replace the CLIP model parameters in the prompt vector extraction module and the feature extraction module with the pre-trained CLIP-B / 32 parameters.

[0083] Input I train And T train Into the prompt generation network, I trainGenerate the image pre-prompt vector feature through the image prompt vector extractor p_img , T train Generate the text pre-prompt vector feature through the text prompt vector extractor p_text , and use feature p_img and feature p_text as the input of the prompt generation module.

[0084] Use feature p_img and feature p_text as the input of the prompt generation module to generate the final image prompt vector and the final text prompt vector j = 1, 2, …, 12.

[0085] Use I train and T train as and input the feature extraction network respectively. The feature extraction network includes 12 layers of Transformer Block; the image-side input of the j-th layer of Transformer Block is denoted as the image-side output is denoted as Use and input the image adapter to generate the image optimized feature The text-side input of the j-th layer of TransformerBlock is denoted as the text-side output is denoted as Use and input the text adapter to generate the text optimized feature

[0086] Use the cls_token of to input the linear projection layer to reduce the dimension and output the image feature vector F img , use the [EOS] of to input the linear projection layer to reduce the dimension and output the text feature vector F text .

[0087] Step4: Use the loss function for supervised training

[0088] The training unit uses a weighted combination of symmetric cross-modal contrast loss and membership loss to calculate and perform supervised training on the CLIP enhanced network based on the multi-modal adapter and shared prompt.

[0089] During the supervised training of the CLIP enhancement network based on multimodal adapters and shared prompts, the image final prompt vector and the text final prompt vector are respectively generated by the prompt generation network. Then, the remote sensing image and the corresponding text description, as well as their corresponding image final prompt vector and text final prompt vector, are input into the feature extraction network to output the image feature vector and the text feature vector. The similarity matrix logits are calculated using the image feature vector and the text feature vector. The weighted combination of the symmetric cross-modal contrast loss and the membership loss is calculated using the logits and the labels, and the final loss value is obtained:

[0090] L = L c + λ cs L a

[0091]

[0092] where: L represents the final loss value, L a represents the membership loss value, L c represents the contrast loss value, λ cs represents the central scale; V i represents the image feature vector of the i-th sample pair, t i represents the text feature vector of the i-th sample pair, and N represents the number of sample pairs; represents the image center of the class to which the i-th sample belongs, represents the text center of the class to which the i-th sample belongs.

[0093] If the final loss value L converges, the training of the CLIP enhancement network based on multimodal adapters and shared prompts is completed; otherwise, new sample pairs I train and T train are selected, and the image final prompt vector and the text final prompt vector are generated again through the prompt generation network. Then, the image feature vector and the text feature vector are extracted through the feature extraction network, and the final loss value is calculated for the results until the final loss value converges.

[0094] Step6: Use the trained CLIP enhancement network based on multimodal adapters and shared prompts for testing

[0095] As Figure 4 shown, input I test and T test into the prompt generation network, and output the image final prompt vector and the text final prompt vector Input I test 、T test 、 and into the feature extraction network, and output the image feature vector Fimg and the text feature vector F text , calculate F img and F text the similarity between them, and output the predicted label. From Figure 4 the comparison between the predicted results and labels in it, it can be seen that the present invention can accurately lock the core ground object main body in the remote sensing image, achieve a high-accuracy retrieval effect in a complex multi-target scene, and can still effectively focus on the main body even in the face of dense distribution or partial occlusion of ground objects of the same category, effectively avoiding background noise interference.

[0096] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form, and all technical solutions obtained by using equivalent replacements or equivalent transformations fall within the protection scope of the present invention.

Claims

1. A remote sensing image text retrieval method based on Mamba adapter and shared hint enhancement, characterized by: The steps include: Step 1: Get the original remote sensing image and the corresponding text description, crop and standardize the original remote sensing image to obtain a remote sensing image of a set size; Step 2: Divide the remote sensing images and the corresponding text descriptions into training sets and test sets according to the proportion, where the training set includes remote sensing images I train and the corresponding text description T train , the test set includes remote sensing images I test and the corresponding text description T test ; Step 3: Construct a CLIP enhanced network based on multimodal adapters and shared prompts, including a prompt generation network and a feature extraction network; the prompt generation network includes a prompt vector extraction module and a prompt generation module, the prompt vector extraction module includes an image prompt vector extractor and a text prompt vector extractor; the feature extraction network includes a feature extraction module and an adapter module, the adapter module includes a group of image adapters and a group of text adapters, the feature extraction module is designed based on the CLIP model model, the CLIP model includes 12 layers of TransformerBlock, an image feature extractor is embedded in the image end of each Transformer Block layer, and a text adapter is embedded in the text end of each Transformer Block layer; replace the CLIP model parameters in the prompt vector extraction module and the feature extraction module with the pre-trained CLIP-B / 32 parameters; Will I train and T train Input to the prompt generation network, I train Generate image pre-hint vector feature through image hint vector extractor p_img , T train Generate text pre-hint vector feature through text hint vector extractor p_text , feature p_img and feature p_text As input to the prompt generation module; The feature p_img and feature p_text Input prompt generation module to generate the final prompt vector of the image and text final prompt vector j = 1, 2, ..., 12; Will I train and T train As and Input feature extraction network, the feature extraction network includes 12 layers of Transformer Block; the image input of the jth layer Transformer Block is recorded as The image output is recorded as Will and Input image adapter, generate image optimization features The text input of the j-th layer TransformerBlock is recorded as The text output is recorded as Will and Input text adapter, generate text optimization features Will The cls_token tag input linear projection layer outputs the image feature vector F after dimensionality reduction img ,Will The [EOS] tag input linear projection layer outputs the text feature vector F after dimensionality reduction text ; Step 4: Calculate F using a weighted combination of symmetric cross-modal contrast loss and membership loss img and F text The loss value of Step 5: Freeze the parameters of the prompt vector extraction module and the feature extraction module, repeat Step 3 to Step 4 until the loss value converges, and the training of the prompt generation module and the adapter module is completed; Step 6: I test and T test Input prompt generation network, output image final prompt vector and text final prompt vector Will I test 、T test , and Input feature extraction network, output F img and Ft. test ; Step 7: Based on F img and F text Perform bidirectional cross-modal retrieval, calculate retrieval performance indicators and output them.

2. The remote sensing image text retrieval method based on Mamba adapter and sharing prompt enhancement according to claim 1 is characterized in that: In the hint vector extraction module, the image hint vector extractor includes a Patch Embedding layer, a Position Embedding layer, a learnable embedding vector P class ′, 12 TransformerEncoder layers connected in series and a linear projection layer; I train Input image hint vector extractor, segmented and flattened into vector sequence P by Patch Embedding layer img ′, for P img ′ and P class ' is spliced ​​and then output through the Position Embedding layer The input of the jth Transformer Encoder layer is The output of the jth Transformer Encoder layer is Will The input linear projection layer transforms the dimension and the output is feature p_i ; In the prompt vector extraction module, the text prompt vector extractor includes a word segmenter, a TokenEmbedding layer, a Position Embedding layer, 12 Transformer Encoder layers connected in series, and a linear projection layer; train Input text prompt vector extractor, segmented by word segmenter output T token ′, output P through the TokenEmbedding layer text ′, output by the Position Embedding layer The input of the jth Transformer Encoder layer is The output of the jth Transformer Encoder layer is Will The input linear projection layer transforms the dimension and the output is feature p_t .

3. The remote sensing image text retrieval method based on Mamba adapter and sharing prompt enhancement according to claim 1 is characterized in that: The prompt generation module includes an image linear layer, a text linear layer and a shared prompt matrix; p_i and feature p_t Input prompt generation module, feature p_i Generate features through the image linear layer p_i ′, feature p_i Generate features through text linear layer p_t ′, feature p_i ′ is concatenated with the shared hint matrix to output the final hint vector of the image feature p_t ′ is concatenated with the shared prompt matrix to output the final prompt vector of the text 4. The remote sensing image text retrieval method based on Mamba adapter and sharing prompt enhancement according to claim 3 is characterized in that: In the hint generation module, the dimension of the image linear layer is [512, 128], the dimension of the text linear layer is [512, 128], and the dimension of the shared hint matrix is ​​[12, 5, 128].

5. The remote sensing image text retrieval method based on Mamba adapter and sharing prompt enhancement according to claim 1 is characterized in that: In the feature extraction module, the image end of the Transformer Block layer is called an image feature extractor, and the text end of the Transformer Block layer is called a text feature extractor; The image feature extractor includes a Patch Embedding layer, a Position Embedding layer, and a learnable embedding vector P class , 12 Transformer Encoder layers connected in series and a linear projection layer; train Input image feature extractor, segmented and flattened into vector sequence P by Patch Embedding layer img , for P img and P class Splicing, then output through the Position Embedding layer The input of the jth Transformer Encoder layer is The output of the ,th Transformer Encoder layer is Will and Input image adapter, generate image optimization features Will The cls_token tag is input into the linear projection layer to transform the dimension, and the output is the image feature vector F img ; The text feature extractor includes a word segmenter, a Token Embedding layer, a PositionEmbedding layer, 12 Transformer Encoder layers and a linear projection layer; train Input into the text feature extractor, and output T after word segmentation by the word segmenter token , output P through the Token Embedding layer text , output by the Position Embedding layer The input of the jth Transformer Encoder layer is The output of the jth Transformer Encoder layer is Will and Input text adapter, generate text optimization features Will The [EOS] tag is input into the linear projection layer to transform the dimension, and the output is the text feature vector F test .

6. The remote sensing image text retrieval method based on Mamba adapter and sharing hint enhancement according to claim 3 is characterized in that: In the adapter module, the image adapter includes a dimensionality reduction linear layer, a data preprocessing layer, two TreeMamba deep feature extraction layers, two residual connections, a gating mechanism connection and a dimensionality increase linear layer, and the data preprocessing layer converts the input from the length dimension to the two dimensions of length and width; Will and Input image adapter, Output of the dimensionality reduction linear layer and After splicing, the output is activated by GELU function Output after dimension conversion by data preprocessing layer Output from two TreeMamba deep feature extraction layers in turn right and Make gate control mechanism connection and output right and Perform residual connection and output Removal belongs to Get the token Output of the linear layer after dimensionality increase right and Perform residual connection and output In the adapter module, the text adapter includes a dimensionality reduction linear layer, two Hydra deep feature extraction layers, a data preprocessing layer, two residual connections, a gating mechanism connection and a dimensionality increase linear layer; Will and Input text adapter, Output of the dimensionality reduction linear layer and After splicing, the output is activated by GELU function Output from two Hydra deep feature extraction layers in sequence right and Make gate control mechanism connection and output right and Perform residual connection and output Removal belongs to Get the token Output of the linear layer after dimensionality increase right, and Perform residual connection and output 7. The remote sensing image text retrieval method based on Mamba adapter and sharing hint enhancement according to claim 6 is characterized in that: In the image adapter, the dimension of the dimensionality reduction linear layer is [768, 128], the acceptance dimension of the TreeMamba deep feature extraction layer is 128, the SSM spatial feature dimension is 1, and the dimension of the dimensionality increase linear layer is [128, 768]; in the text adapter, the dimension of the dimensionality reduction linear layer is [512, 128], the acceptance dimension of the Hydra deep feature extraction layer is 128, the SSM spatial feature dimension is 16, and the dimension of the dimensionality increase linear layer is [128, 512].

8. A remote sensing image text retrieval device based on Mamba adapter and sharing prompt enhancement, characterized in that: It includes remote sensing image and text acquisition unit, CLIP enhancement network based on multimodal adapter and shared prompt, training unit and testing unit; The remote sensing image and text acquisition unit is used to acquire the original remote sensing image and the corresponding text description, and to crop and standardize the original remote sensing image to obtain a remote sensing image of a set size; The CLIP enhancement network based on multimodal adapter and shared prompts includes a prompt generation network and a feature extraction network; The prompt generation network includes a prompt vector extraction module and a prompt generation module, through which the prompt vector extraction module respectively extracts image pre-prompt vectors and text pre-prompt vectors from the remote sensing image and the corresponding text description, and then generates the image final prompt vector and the text final prompt vector through the prompt generation module; the feature extraction network includes a feature extraction module and an adapter module, the feature extraction module and the adapter module work together, the feature extraction module extracts shallow features of the remote sensing image and the corresponding text description, and the adapter module extracts deep features of the remote sensing image and the corresponding text description, and finally generates feature vectors of the remote sensing image and the corresponding text description; The training unit uses a loss function to perform supervised training on a CLIP enhanced network based on a multimodal adapter and shared prompts; The test unit uses a trained CLIP enhanced network based on a multimodal adapter and shared prompts to perform image-to-text retrieval and text-to-image retrieval on paired remote sensing images and corresponding text descriptions.

9. The remote sensing image text retrieval device based on Mamba adapter and sharing prompt enhancement according to claim 8, characterized in that: The CLIP enhancement network based on multimodal adapter and shared prompts includes a prompt generation network and a feature extraction network; The prompt generation network includes a prompt vector extraction module and a prompt generation module, the prompt vector extraction module includes an image prompt vector extractor and a text prompt vector extractor; the feature extraction network includes a feature extraction module and an adapter module, the adapter module includes a group of image adapters and a group of text adapters, the feature extraction module is designed based on the CLIP model model, the CLIP model includes 12 layers of TransformerBlock, an image feature extractor is embedded in the image end of each Transformer Block layer, and a text adapter is embedded in the text end of each Transformer Block layer; the CLIP model parameters in the prompt vector extraction module and the feature extraction module are replaced with pre-trained CLIP-B / 32 parameters.

10. The remote sensing image text retrieval device based on Mamba adapter and sharing prompt enhancement according to claim 9, characterized in that: In the CLIP-enhanced network based on multimodal adapters and shared cues: Will I train and T train Input to the prompt generation network, I train Generate image pre-hint vector feature through image hint vector extractor p_img , T train Generate text pre-hint vector feature through text hint vector extractor p_test , feature p_img and feature p_text As input to the prompt generation module; The feature p_img and feature p_text Input prompt generation module to generate the final prompt vector of the image and text final prompt vector Will I train and T train As and Input feature extraction network, the feature extraction network includes 12 layers of Transformer Block; the image input of the jth layer Transformer Block is recorded as The image output is recorded as Will and Input image adapter, generate image optimization features The text input of the j-th layer TransformerBlock is recorded as The text output is recorded as Will and Input text adapter, generate text optimization features Will The cls_token tag input linear projection layer outputs the image feature vector F after dimensionality reduction img ,Will The [EOS] tag input linear projection layer outputs the text feature vector F after dimensionality reduction text .

Citation Information

Patent Citations

  • Cross-modal remote sensing image-text matching network based on collaborative learning and matching method thereof

    CN116578737A

  • Refined category image generation method based on retrieval enhancement

    CN118170937A

  • Building image retrieval method, device and equipment based on multiple modes

    CN118277603A

  • Image-text retrieval method and device in remote sensing field based on prompt learning

    CN118690034A

  • Cross-modal hash learning method based on Mama and covariance interactive attention

    CN118747843A