Remote sensing cross-modal retrieval method and system based on large model fine-tuning
Through cross-modal asymmetric adapter and dual-task consistency loss function, the problem of modal features difference in remote sensing image-text retrieval is solved, and the search performance and adaptability are improved.
Patent Information
- Application Number
- CN202510412423.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing remote sensing image-text search method fails to fully consider the difference in the modal features of images and text, resulting in modal imbalance, affecting the performance and generalization capabilities of the model.
A cross-modal asymmetric adapter is adopted, combined with a visual enhancement adapter and a text semantic adapter, and image and text features are enhanced by differential attention mechanism and hierarchical attention mechanism, and optimized with a dual-task consistency loss function.
It realizes better alignment and optimization of the features of images and text modalities, and improves the performance and adaptability of remote sensing image-text retrieval.
Smart Images

Figure CN119917691B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of remote sensing cross-modal retrieval, and in particular to a remote sensing cross-modal retrieval method and system based on large model fine-tuning. Background Art
[0002] Remote sensing image-text retrieval is of great importance and profound significance in remote sensing data processing and applications. Its core lies in achieving efficient matching and retrieval between images and text. By integrating multimodal data, remote sensing image-text retrieval not only improves the accuracy and comprehensiveness of remote sensing data analysis but also provides strong support for diverse application scenarios.
[0003] Currently, remote sensing image-text retrieval methods are primarily divided into traditional deep learning approaches and those based on large-scale vision-language pre-trained models. Traditional deep learning methods typically rely on deep learning operators such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and transformers to extract features from remote sensing images and text data, respectively. These features are then mapped into a shared feature space and cross-modal matching is achieved through similarity metrics. However, these methods often struggle to achieve accurate feature alignment when dealing with the complex and abstract cross-modal relationships between remote sensing images and text, limiting retrieval performance. To address this, some studies have introduced methods based on large-scale vision-language pre-trained models. By freezing most model parameters and fine-tuning only a few key parameters, they retain the original learning information while incorporating domain knowledge. This approach not only significantly reduces training costs but also improves the model's cross-modal matching accuracy in remote sensing scenarios. Among efficient parameter fine-tuning methods, adapters are the most common fine-tuning solution. However, in remote sensing image-text retrieval tasks, this presents the following problem: due to the text modality's greater discriminability and ease of optimization, it often becomes the dominant modality, thereby inhibiting the optimization of the image modality. More importantly, the image and text modalities often use the same adapter structure, which ignores the potential differences in representation between the modalities and leads to modality imbalance, thus affecting the overall model performance and generalization ability. Summary of the Invention
[0004] The embodiments of the present application provide a remote sensing cross-modal retrieval method and system based on large model fine-tuning, which is used to solve the problem that the fine-tuning method in the above-mentioned prior art fails to fully consider the differences in image and text modal features.
[0005] On the one hand, an embodiment of the present application provides a remote sensing cross-modal retrieval method based on large model fine-tuning, the method comprising:
[0006] S1. Using the text describing the remote sensing image as the description text, the remote sensing image and / or the description text are input into a trained remote sensing image text retrieval network for processing to obtain cross-modal fused image features and text features;
[0007] S2. Obtaining text and / or images that match the remote sensing image and / or description text based on image features and / or text features for output;
[0008] Wherein, step S1 includes:
[0009] S11, extracting initial image features and / or initial text features from the remote sensing image and / or description text through an image-text encoder;
[0010] S12. Performing cross-modal fusion processing on the initial image features and / or initial text features through a cross-modal asymmetric adapter to obtain cross-modal fused image features and / or text features;
[0011] S13. Optimize image features and / or text features through a dual-task consistency loss function.
[0012] Preferably, the image-text encoder includes an image encoder and a text encoder; step S11 includes:
[0013] Extracting initial image features from remote sensing images through an image encoder;
[0014] Initial text features are extracted from the description text through a text encoder.
[0015] Preferably, step S12 includes:
[0016] S121, performing dimensionality reduction processing on the initial image features and the initial text features;
[0017] S122, using the nonlinear activation function GELU to process the initial image features and initial text features after dimensionality reduction;
[0018] S123, performing feature expression enhancement processing on the initial image features and the initial text features processed by the nonlinear activation function respectively;
[0019] S124: Inputting the initial image features and initial text features processed based on feature expression enhancement into a shared layer for cross-modal fusion processing to obtain cross-modal fused image features and / or text features.
[0020] Preferably, the cross-modal asymmetric adapter includes: a visual enhancement adapter and a text semantic adapter; the visual enhancement adapter and the text semantic adapter respectively include a linear transformation and a nonlinear activation function.
[0021] Preferably, in step S123, the visual enhancement adapter performs feature expression enhancement processing on the initial image features through a differential attention mechanism, including:
[0022] Map the query, key vector, and value vector to the output;
[0023] Calculate the attention score based on the query and key vectors, and then sum it with the value vector;
[0024] Eliminate attention score noise based on the softmax function.
[0025] Preferably, in step S123, the text semantic adapter performs feature expression enhancement processing on the initial text features through a hierarchical attention mechanism, including:
[0026] Obtain word-level hidden representations from word-level annotations through a multi-layer perceptron;
[0027] Calculate word importance weights based on word-level hidden representations and word-level context vectors;
[0028] The sentence vector is obtained by weighted summing of word-level annotations based on importance weights.
[0029] Preferably, the dual-task consistency loss function includes the following loss terms:
[0030] Consistency constraints, based on the exponential moving average mechanism, use the text modality as a teacher model to guide image modality learning;
[0031] Cross-modal constraints, measuring the global alignment of images and texts via cosine similarity;
[0032] Classification constraints, which calculate the difference between the predicted probability distribution and the true category distribution;
[0033] Based on the adaptive weight adjustment strategy, the consistency constraint loss term, the cross-modal constraint loss term and the classification constraint loss term are dynamically weighted to optimize the image features and / or text features.
[0034] On the other hand, this application also provides a remote sensing cross-modal retrieval system based on large model fine-tuning, which includes:
[0035] An image text input module is used to use the text describing the remote sensing image as the description text and output the remote sensing image and / or the description text;
[0036] A network composition and processing module is used to process the remote sensing image and / or description text input by the image and text input module to obtain image features and / or text features after cross-modal fusion;
[0037] An image text output module, which outputs text and / or images that match the remote sensing image and / or the description text based on the image features and / or the text features;
[0038] Among them, the network composition and processing modules include:
[0039] An image-text encoder that extracts initial image features and / or initial text features from remote sensing images and / or description texts;
[0040] Performing cross-modal fusion processing on the initial image features and / or initial text features to obtain a cross-modal asymmetric adapter of cross-modal fused image features and / or text features;
[0041] A dual-task consistency loss function that optimizes image features and / or text features.
[0042] On the other hand, the present application also provides a terminal device, including: a memory and a processor, the memory is used to store a remote sensing cross-modal retrieval program based on large model fine-tuning, and when the remote sensing cross-modal retrieval program based on large model fine-tuning is executed by the processor, the steps of the remote sensing cross-modal retrieval method based on large model fine-tuning as above are implemented.
[0043] On the other hand, the present application also provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed, it implements the above-mentioned remote sensing cross-modal retrieval method based on large model fine-tuning.
[0044] The remote sensing cross-modal retrieval method and system based on large model fine-tuning in this application has the following advantages:
[0045] The solution proposed in this application effectively addresses the differences in image and text modal features through the collaborative work of a cross-modal asymmetric adapter, a visual enhancement adapter, and a text semantic adapter. This allows for better alignment and optimization of features between image and text modalities during fine-tuning. Furthermore, by expanding the single-task model to a multi-task model and introducing a dual-task consistency loss function, inter-modal consistency learning is strengthened, thereby improving retrieval performance. This approach has demonstrated excellent adaptability and significant performance improvements in remote sensing image-text retrieval tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0047] Figure 1 A flowchart of a remote sensing cross-modal retrieval method based on large model fine-tuning provided in an embodiment of the present application;
[0048] Figure 2 A schematic diagram of a process for processing remote sensing images and / or description texts by a remote sensing image text retrieval network provided in an embodiment of the present application;
[0049] Figure 3 A schematic diagram of a process for processing initial image features and / or initial text features by a cross-modal asymmetric adapter provided in an embodiment of the present application;
[0050] Figure 4 A schematic diagram of the structure of a remote sensing cross-modal retrieval system based on large model fine-tuning provided in an embodiment of the present application;
[0051] Figure 5 A schematic diagram of the training process of a remote sensing image text retrieval network provided in an embodiment of the present application;
[0052] Figure 6 The image-text retrieval analysis comparison table provided in the embodiment of this application;
[0053] Figure 7 This is a text-image retrieval analysis comparison table provided in the embodiments of this application. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0055] refer to Figure 1 As shown, a flow chart of a remote sensing cross-modal retrieval method based on large model fine-tuning provided by an embodiment of the present application is provided. An embodiment of the present application provides a remote sensing cross-modal retrieval method based on large model fine-tuning, the method comprising:
[0056] S1. Using the text describing the remote sensing image as the description text, the remote sensing image and / or the description text are input into a trained remote sensing image text retrieval network for processing to obtain cross-modal fused image features and / or text features;
[0057] S2. Based on the cross-modal fusion features, text and / or image matching the remote sensing image and / or description text is obtained and outputted.
[0058] In an embodiment of the present application, during the execution of a remote sensing image-text retrieval task, a remote sensing image or descriptive text is obtained and input into a remote sensing image text retrieval network previously trained through deep learning of big data for processing, thereby obtaining image features and text features after cross-modal fusion. Furthermore, text information matching the image features or image information matching the text features is obtained through the image features and text features, and the matched text information or image information is output as text or image matching the remote sensing image or descriptive text.
[0059] In one embodiment, when performing a remote sensing image-text retrieval task, if only a remote sensing image is available, it is processed through a remote sensing image text retrieval network to obtain image features that are cross-modally fused with auxiliary text features. For example, in these fused features, the regular arrangement of factory roofs (visual) is semantically combined with "dense industrial facilities" (text), eliminating interference from "residential areas" during retrieval; and the outline of a lake (visual) is associated with "a lake surrounded by mountains" (text), avoiding mismatching with "a plain reservoir." During retrieval, the text description with the highest matching degree is obtained based on these image features and output as the text matching the remote sensing image.
[0060] In one embodiment, when performing a remote sensing image-text retrieval task and only obtaining descriptive text, the remote sensing image text retrieval network processes this descriptive text to obtain text features that are cross-modally fused with auxiliary image features. For example, the fused features associate visual patterns such as "regularly arranged factory buildings" and "highly reflective roofs," excluding images of "scattered residential buildings." The semantic meaning of "oasis" in the features is strongly associated with visual features such as "small patches of vegetation" and "contrasting colors on the sand," improving fine-grained retrieval accuracy. During retrieval, the image with the highest matching degree based on this text feature is obtained and output as the image that matches the descriptive text.
[0061] In one embodiment, when performing a remote sensing image-text retrieval task, when a remote sensing image and text describing the image are simultaneously acquired, a remote sensing image-text retrieval network performs cross-modal fusion processing on the remote sensing image and the text description to obtain image features and text features corresponding to the remote sensing image and the text description, respectively. A search is then performed based on the image features and text features to obtain text and images that match the remote sensing image and text description, and these are output.
[0062] refer to Figure 2 As shown, a schematic diagram of a process for processing remote sensing images and / or description texts by a remote sensing image text retrieval network provided in an embodiment of the present application is shown. The steps for processing remote sensing images and / or description texts by a remote sensing image text retrieval network provided in an embodiment of the present application include:
[0063] S11, extracting initial image features and / or initial text features from the remote sensing image and / or description text through an image-text encoder;
[0064] In the embodiment of the present application, after receiving the remote sensing image and / or the description text, the image text encoder extracts the initial image features from the received remote sensing image through the image encoder. , extracting initial text features from the received description text through the text encoder In one embodiment, a vision transformer (ViT) is selected as an image encoder to extract the initial image features in the remote sensing image. , select Bidirectional Encoder Representations from Transformers (BERT) as the text encoder to extract the initial text features in the description text In this embodiment, both encoders perform feature extraction through Transformer blocks, which contain 12 Transformer blocks in total.
[0065] S12. Performing cross-modal fusion processing on the initial image features and / or initial text features through a cross-modal asymmetric adapter to obtain cross-modal fused image features and / or text features;
[0066] In an embodiment of the present application, the cross-modal asymmetric adapter performs cross-modal fusion on the extracted image features and / or text features, as follows Figure 3 Follow the steps shown below to process:
[0067] S121, performing dimensionality reduction processing on the initial image features and the initial text features;
[0068] S122, using the nonlinear activation function GELU (Gaussian error linear unit) to process the initial image features and initial text features after dimensionality reduction;
[0069] S123, performing feature expression enhancement processing on the initial image features and the initial text features processed by the nonlinear activation function respectively;
[0070] S124: Inputting the initial image features and initial text features processed based on feature expression enhancement into a shared layer for cross-modal fusion processing to obtain cross-modal fused image features and / or text features.
[0071] In an embodiment of the present application, the cross-modal asymmetric adapter first performs a multi-step process on the initial image features extracted by the Transformer block. and initial text features Perform dimensionality reduction and project it into a low-dimensional representation; then use the nonlinear activation function GELU to transform the initial image features and initial text features Then, the initial image features of the image modality are processed by the differential attention mechanism, and the initial text features of the text modality are processed by the hierarchical attention mechanism to further enhance the feature expression of each. Then, these processed features are input into the shared layer for interaction between the image modality and the text modality, and finally restored to the original input size through the dimensionality increase layer. In this embodiment, the cross-modal asymmetric adapter includes a visual enhancement adapter and a text semantic adapter, wherein a typical adapter includes a series of linear transformations and nonlinear activation functions, and its calculation formula is as follows:
[0072] .
[0073] Among them, Adapter means adapter, is the extracted features; and are the weights of the down-projection and up-projection respectively; is a nonlinear activation function, is a GELU activation function, and the superscript T represents transposition.
[0074] In an embodiment of the present application, the visual enhancement adapter of the cross-modal asymmetric adapter, when performing feature expression enhancement processing on the initial image features, introduces a differential attention mechanism to extract the initial image features of the image modality in a more fine-grained manner. The differential attention mechanism maps the query, key vector, and value vector to the output, and uses the query and key vectors to calculate the attention score, and then weighted sums the value vector. Its core lies in introducing a pair of softmax functions to eliminate noise in the attention score. Its specific calculation formula is as follows:
[0075]
[0076] in, Q 1 and Q 2 is the query vector, K 1 and K 2 is the key vector, V is the value vector; They are Parameters, where Q for Q 1 and Q 2, K for K 1 andK 2; DiffAttn represents the differential attention mechanism; d The dimension of the key vector is used to scale the dot product result to prevent the gradient from exploding or disappearing; is a learnable scalar. In order to coordinate the dynamic changes of the learning process, Reparameterize to:
[0077]
[0078] in, is a learnable vector; Is used for initialization constant.
[0079] In this embodiment, features extracted by the differential attention mechanism are passed to the shared layer, sharing information with the text modality, thereby achieving complementarity between the two modalities. For example, when certain key information in the text is omitted or unclear, the visual information of the image can serve as a supplement, helping the model better understand and reason. Finally, before the features are upgraded to the input dimension, a gating mechanism is introduced to regulate the information interaction between different modalities, preventing the features of a certain modality from being overly dominant or being ignored, thereby achieving a more balanced cross-modal fusion.
[0080] In an embodiment of the present application, the text semantic adapter of the cross-modal asymmetric adapter, when performing feature expression enhancement processing on the initial text features, extracts key information words by introducing a hierarchical attention mechanism and aggregates their representations to construct a sentence vector, and performs the following steps, wherein the calculation formula for each step is as follows:
[0081]
[0082]
[0083]
[0084] in, For word-level annotations, for The hidden representation of W w represents the learnable weight matrix, b w represents the learnable bias term, is the normalized importance weight, is the word-level context vector, is the sentence vector.
[0085] First, the word-level annotation Pass a single-layer multi-layer perceptron (MLP) to obtain its hidden representation .
[0086] Then, through With word-level context vector The importance of the word is measured by the similarity of the word, and the normalized importance weight is obtained through the softmax function .
[0087] Finally, the word-level annotations are weighted and summed according to these weights to obtain the sentence vector . Among them, the word-level context vector It can be seen as a high-level representation of a fixed query "Which word is the most informative word", similar to the query mechanism used in Memory Networks.
[0088] In this embodiment, during the training process, It is randomly initialized and learned jointly with other parameters of the model. After the hierarchical attention mechanism, the subsequent operations of the visual enhancement adapter and the text semantic adapter are consistent.
[0089] S13. Optimize image features and / or text features through a dual-task consistency loss function.
[0090] In an embodiment of the present application, after the cross-modal asymmetric adapter obtains the fused image features and text features, the two final features are subjected to similarity calculation, cross entropy loss calculation, and mean square error calculation. These three items are used as components of the loss function and are dynamically weighted and optimized.
[0091] The loss term corresponding to the mean squared error calculation is a consistency constraint loss term. By introducing the moving average (EMA) mechanism, the text modality is used as the teacher model and the image modality as the student model, using the guidance of the text modality to improve the classification accuracy of the image modality. By measuring the gap between the predicted outputs of the image modality and the text modality, the learning process of the image modality is optimized. The formula for calculating the mean squared error is as follows:
[0092]
[0093] in, L consist is the mean square error, is the sample size; It is The true value of the samples; It is The predicted value of the sample.
[0094] Among them, the loss term corresponding to the similarity calculation is the cross-modal loss term. The cross-modal alignment of remote sensing images and texts is achieved by measuring their global similarity. The formula for similarity calculation is as follows:
[0095]
[0096] in, L cross For similarity, represents a matched pair of samples; For remote sensing images Unmatched text; For text Mismatched remote sensing images; is the cosine similarity; is the boundary value of the cross-modal constraint; .
[0097] The calculation formula of cosine similarity is as follows:
[0098]
[0099] in, I i For images I No. i Dimension values, T i For text T No. i The loss term corresponding to the cross entropy loss calculation is the classification constraint loss term, which measures the classification performance of the model by calculating the difference between the probability distribution predicted by the model and the true category distribution. The formula for calculating the cross entropy loss is as follows:
[0100]
[0101] in, L cls is the cross entropy loss, is the number of categories; Is the one-hot encoding of the true label, if the sample belongs to the category ,but ,otherwise ; Is the model for the category The predicted probability of .
[0102] In the embodiments of this application, in view of the different effects of different loss terms on model training, this application introduces an adaptive weight adjustment strategy to avoid the weight imbalance problem caused by directly accumulating each loss term in previous studies, and dynamically weights and optimizes the above three loss terms. This strategy can dynamically adjust the weight ratio according to the contribution of each loss term during training, thereby more effectively guiding model learning and improving the effect of cross-modal feature alignment. The specific form of the dual-task consistency loss function is as follows:
[0103]
[0104] in, L total is the dual-task consistency loss, It is The learnable parameters corresponding to the loss terms; are respectively , representing the A specific loss item; For dynamic weighting; It is a regularization term used to limit the growth of weights and avoid excessive bias towards a certain loss term.
[0105] refer to Figure 4 The present application also provides a remote sensing cross-modal retrieval system based on large model fine-tuning, which includes: an image and text input module, a network composition and processing module, and an image and text output module.
[0106] In an embodiment of the present application, after acquiring a remote sensing image or descriptive text, the remote sensing cross-modal retrieval system based on large model fine-tuning inputs the remote sensing image or descriptive text into the network composition and processing module through the image-text input module for cross-modal fusion processing to obtain the image features or text features after cross-modal fusion, and the image-text output module obtains text information or image information that matches the remote sensing image or descriptive text based on the image features or text features.
[0107] In an embodiment of the present application, the network composition and processing module is composed of an image-text encoder, a cross-modal asymmetric adapter and a dual-task consistency loss function, wherein the image-text encoder adopts a vision transformer (ViT) large model and a bidirectional encoder representation method Bidirectional Encoder Representations from Transformers (BERT) large model to extract initial image features or initial text features respectively; the cross-modal asymmetric adapter processes the image modality and text modality respectively by adopting a differential attention mechanism and a hierarchical attention mechanism, and performs inter-modal interaction; the dual-task consistency loss function uses an adaptive weight adjustment strategy to dynamically adjust the weight ratio according to the contribution of each loss item during the training process, thereby more effectively guiding model learning.
[0108] In this embodiment, the cross-modal asymmetric adapter first performs a multi-step process on the initial image features extracted by the Transformer block. and initial text features Perform dimensionality reduction and project it into a low-dimensional representation; then use the nonlinear activation function GELU to transform the initial image features and initial text features Then, the initial image features of the image modality are processed by the differential attention mechanism, and the initial text features of the text modality are processed by the hierarchical attention mechanism to further enhance the feature expression of each. Then, these processed features are input into the shared layer for interaction between the image modality and the text modality, and finally restored to the original input size through the dimensionality increase layer. In this embodiment, the cross-modal asymmetric adapter includes a visual enhancement adapter and a text semantic adapter, wherein a typical adapter includes a series of linear transformations and nonlinear activation functions, and its calculation formula is as follows:
[0109]
[0110] in, is the extracted features; and are the weights of the down-projection and up-projection respectively; is a nonlinear activation function and is a GELU activation function.
[0111] In an embodiment of the present application, the visual enhancement adapter of the cross-modal asymmetric adapter, when performing feature expression enhancement processing on the initial image features, introduces a differential attention mechanism to extract the initial image features of the image modality in a more fine-grained manner. The differential attention mechanism maps the query, key vector, and value vector to the output, and uses the query and key vectors to calculate the attention score, and then weighted sums the value vector. Its core lies in introducing a pair of softmax functions to eliminate noise in the attention score. Its specific calculation formula is as follows:
[0112]
[0113] in, They are query, key vector and value vector respectively; They are Parameters; is a learnable scalar. In order to coordinate the dynamic changes of the learning process, Reparameterize to:
[0114]
[0115] in, is a learnable vector; Is used for initialization constant.
[0116] In this embodiment, features extracted by the differential attention mechanism are passed to the shared layer, sharing information with the text modality, thereby achieving complementarity between the two modalities. For example, when certain key information in the text is omitted or unclear, the visual information of the image can serve as a supplement, helping the model better understand and reason. Finally, before the features are upgraded to the input dimension, a gating mechanism is introduced to regulate the information interaction between different modalities, preventing the features of a certain modality from being overly dominant or being ignored, thereby achieving a more balanced cross-modal fusion.
[0117] In an embodiment of the present application, the text semantic adapter of the cross-modal asymmetric adapter, when performing feature expression enhancement processing on the initial text features, extracts key information words by introducing a hierarchical attention mechanism and aggregates their representations to construct a sentence vector, and performs the following steps, wherein the calculation formula for each step is as follows:
[0118]
[0119]
[0120]
[0121] First, the word-level annotation Pass a single-layer multi-layer perceptron (MLP) to obtain its hidden representation .
[0122] Then, through With word-level context vector The importance of the word is measured by the similarity of the word, and the normalized importance weight is obtained through the softmax function .
[0123] Finally, the word-level annotations are weighted and summed according to these weights to obtain the sentence vector . Among them, the word-level context vector It can be seen as a high-level representation of a fixed query "Which word is the most informative word", similar to the query mechanism used in Memory Networks.
[0124] In the embodiment of the present application, the dual-task consistency loss function avoids the weight imbalance problem caused by directly accumulating each loss term. It dynamically adjusts the weight ratio according to the contribution of each loss term during training, thereby more effectively guiding model learning and improving the effect of cross-modal feature alignment. Among them, each loss term includes:
[0125] Corresponding to the consistency constraint loss term in the mean squared error calculation, we introduce the moving average (EMA) mechanism, using the text modality as the teacher model and the image modality as the student model. We leverage the guidance of the text modality to improve the classification accuracy of the image modality. By measuring the gap between the predicted outputs of the image modality and the text modality, we optimize the learning process of the image modality. The formula for calculating the mean squared error is as follows:
[0126]
[0127] in, is the sample size; It is The true value of the samples; It is The predicted value of the sample.
[0128] Corresponding to the cross-modal loss term of similarity calculation, the cross-modal alignment of remote sensing images and texts is achieved by measuring their global similarity. The formula for similarity calculation is as follows:
[0129]
[0130] in, represents a matched pair of samples; For remote sensing images Unmatched text; For text Mismatched remote sensing images; is the boundary value of the cross-modal constraint; ;
[0131] The calculation formula of cosine similarity is as follows:
[0132] .
[0133] The classification constraint loss term is calculated corresponding to the cross entropy loss. The classification performance of the model is measured by calculating the difference between the probability distribution predicted by the model and the true category distribution. The formula for calculating the cross entropy loss is as follows:
[0134]
[0135] in, is the number of categories; Is the one-hot encoding of the true label, if the sample belongs to the category ,but ,otherwise ; Is the model for the category The predicted probability of .
[0136] In the embodiments of this application, in view of the different effects of different loss terms on model training, this application introduces an adaptive weight adjustment strategy to avoid the weight imbalance problem caused by directly accumulating each loss term in previous studies, and dynamically weights and optimizes the above three loss terms. This strategy can dynamically adjust the weight ratio according to the contribution of each loss term during training, thereby more effectively guiding model learning and improving the effect of cross-modal feature alignment. The specific form of the dual-task consistency loss function is as follows:
[0137]
[0138] in, It is The learnable parameters corresponding to the loss terms; are respectively , representing the A specific loss item; For dynamic weighting; It is a regularization term used to limit the growth of weights and avoid excessive bias towards a certain loss term.
[0139] In an embodiment of the present application, the image-text output module obtains matching text information from a text library based on image features after cross-modal fusion, and obtains matching image information from an image library based on text features after cross-modal fusion.
[0140] refer to Figure 5 As shown, this application also provides a training method for a remote sensing image text retrieval network, including:
[0141] To construct training samples for a remote sensing image text retrieval network, we selected the RSICD and RSITMD datasets and divided them into training, validation, and test sets, accounting for 80%, 10%, and 10%, respectively. The training set was used for model training, optimizing model parameters using the data in the training set; the validation set was used for fine-tuning the model during training; and the test set was used to evaluate the final model performance. The remote sensing images in the training, validation, and test sets were resized to a fixed size of 224×224. The resulting features were linearly projected onto a 512-dimensional common space.
[0142] Construct a remote sensing cross-modal retrieval network based on large model fine-tuning, including an image-text encoder, a cross-modal asymmetric adapter, and a dual-task consistency loss function.
[0143] The Vision Transformer (ViT) was selected as the image encoder, and the Bidirectional Encoder Representations from Transformers (BERT) was selected as the text encoder. At the start of training, the parameters of the ViT image encoder and BERT text encoder were frozen and not updated during training. All images were resized to a fixed size of 224×224 for training. The dropout value and the cross-modal loss function were both set to 0.2. The network's initial learning rate was set to 2e-4, with a weight decay of 0.7 after every 20 training epochs. During training, the Adam optimizer was used for network parameter optimization, with a batch size of 32 and a total training epoch of 30. To ensure the stability and reliability of the experimental results, k-fold cross-validation was used for evaluation, and the average of the 5-fold cross-validation (k=5) was used for performance evaluation and presentation.
[0144] In remote sensing image-text retrieval tasks, existing methods typically design adapters with the same architecture for both image and text modalities. However, due to significant representational differences between image and text data, using the same adapter architecture to process both often fails to accurately extract key information from each modality, thus affecting model performance. In the embodiments of this application, a cross-modal asymmetric adapter, including a visual enhancement adapter and a text semantic adapter, is designed to fine-tune the remote sensing image text retrieval network based on the characteristics of the two modalities.
[0145] In remote sensing image-text retrieval tasks, bidirectional triplet loss has become the mainstream optimization strategy. However, due to the differences in image and text modal features, there are significant differences in their classification performance, making it impossible to accurately align and match features. In addition, this task usually involves multiple loss terms. If the losses of all samples are directly accumulated, it may cause difficulties in the training process to converge, making it difficult for the model to effectively capture subtle feature differences. To this end, this application expands the single-task model to a multi-task model and proposes a dual-task consistency loss function with adaptive weight adjustment capabilities. This function can dynamically mine and optimize difficult samples, strengthen feature alignment and consistency between modalities, and thus significantly improve retrieval performance.
[0146] During training, the remote sensing image text retrieval network was run on an NVIDIA RTX 4070 12GB GPU. For feature representation, the final features of the CLIP (Contrastive Language-Image Pre-training) or GeoRSCLIP (Geographic Remote Sensing Contrastive Language-Image Pre-training) models were linearly projected into a 512-dimensional common space. For remote sensing images, all images were resized to a fixed size of 224×224 for training. Regarding parameter settings, the dropout value and the cross-modal loss function were both set to 0.2. The network's initial learning rate was set to 2e-4, with a weight decay of 0.7 applied after every 20 training epochs. During training, the Adam optimizer was used for network parameter optimization, with a batch size of 32 and a total training epoch of 30. To ensure the stability and reliability of the experimental results, k-fold cross-validation was used for evaluation, and the average of the 5-fold cross-validation (k=5) was used for performance evaluation and presentation.
[0147] To more intuitively illustrate the effectiveness of the remote sensing cross-modal retrieval method based on large model fine-tuning provided in this application, this application uses two indicators, R@K (K=1, 5, and 10) and mR (mean recall) to evaluate model performance. R@K is used to measure the proportion of true matches in the query results that appear in the first K retrieval results, where K is 1, 5, and 10, to evaluate the accuracy of the model at different retrieval depths; mR represents the average of each R@K indicator in text retrieval and image retrieval tasks, and is used to comprehensively evaluate the overall performance of the model in cross-modal retrieval tasks. The specific calculation formula is as follows:
[0148]
[0149] As shown in Tables 1 and 2, the performance comparison of the remote sensing cross-modal retrieval method based on large model fine-tuning in this application and other retrieval methods on the RSICD (Remote Sensing Image Captioning) and RSITMD (Remote Sensing Image Cross-Modal Retrieval Fine-Grained Multi-Scale) datasets is given. The bold black represents the method proposed in this application and the fully fine-tuned GeoRSCLIP, respectively.
[0150] Table 1 Performance comparison of the retrieval method of this application and other retrieval methods on the RSICD dataset
[0151]
[0152] Table 2 Performance comparison of the retrieval method of this application and other retrieval methods on the RSITMD dataset
[0153]
[0154] In Tables 1 and 2, AMFMN-soft is an asymmetric multimodal feature matching network-soft method, AMFMN-fusion is an asymmetric multimodal feature matching network-fusion method, AMFMN-sim is an asymmetric multimodal feature matching network-similarity method, GaLR w / o MR is a geographic attention-guided cross-modal retrieval method without edge reconstruction, GaLR with MR is a geographic attention-guided cross-modal retrieval method with edge reconstruction, PIR is a private information retrieval method, Full-FT CLIP is a full-scale fine-tuned contrast language-image pre-training model, Full-FT GeoRSCLIP is a full-scale fine-tuned geographic remote sensing contrast language-image pre-training model, and PE-RSITR is a remote sensing image target recognition technology based on piezoelectric effect.
[0155] As shown in Tables 1 and 2, the remote sensing cross-modal retrieval method based on large-scale model fine-tuning in this application improves performance by 13%-23% over traditional methods and surpasses advanced traditional retrieval methods in recent years. The CLIP-based method fully utilizes its strong generalization ability and rich visual-language prior knowledge obtained from pre-training on large-scale natural scene datasets, thereby achieving significant performance improvements in remote sensing image text retrieval tasks. Specifically, the remote sensing cross-modal retrieval method based on large-scale model fine-tuning provided in this application improves MR (edge reconstruction) on RSICD by approximately 16.44% compared to CLIP-Adapter and 6.97% compared to PE-RSITR; on RSITMD, the retrieval method provided in this application improves by 13.98% compared to Cross-ModalAdapter and 7.57% compared to PE-RSITR. It is worth noting that with the help of the GeoRSCLIP pre-trained model, the retrieval method provided in this application outperforms all current parameter-efficient fine-tuning methods on the RSICD and RSITMD datasets, and even exceeds the retrieval performance of the fully fine-tuned GeoRSCLIP model, verifying the effectiveness of the retrieval method provided in this application.
[0156] In order to further verify the superiority of the remote sensing cross-modal retrieval method based on large model fine-tuning in this application, the following Figure 6 and Figure 7 The visual analysis shown. Figure 6 It can be seen from the image-text retrieval results that the remote sensing cross-modal retrieval method based on large model fine-tuning provided by this application is superior to the fully fine-tuned GeoRSCLIP in retrieval accuracy. For example, in the second image, the erroneous text retrieved by the remote sensing cross-modal retrieval method based on large model fine-tuning provided by this application can correctly identify the number of "baseballfield", while the Text-1 and Text-3 retrieved by the fully fine-tuned GeoRSCLIP only focus on key targets and fail to capture the quantity information of the targets. The above phenomenon shows that when performing image feature extraction in complex remote sensing scenes, the remote sensing cross-modal retrieval method based on large model fine-tuning provided by this application shows stronger capabilities in the fine-grained understanding and expression of image modal features. From Figure 7It can be seen from the text-image retrieval results that the remote sensing cross-modal retrieval method based on large model fine-tuning provided by this application is significantly better than the fully fine-tuned GeoRSCLIP in retrieval performance. For example, for the first text, the remote sensing cross-modal retrieval method based on large model fine-tuning provided by this application can accurately identify the image corresponding to the text in Top-1, while the fully fine-tuned GeoRSCLIP only retrieves the correct result in Top-3. These results fully demonstrate the advantages of the remote sensing cross-modal retrieval method based on large model fine-tuning provided by this application in fine-grained feature extraction and cross-modal information alignment, and effectively bridges the gap between image and text modalities, significantly improving the retrieval performance and generalization ability of the model.
[0157] Based on the above analysis results, the experimental results on multiple benchmark datasets show that the remote sensing cross-modal retrieval method based on large model fine-tuning in this application outperforms other state-of-the-art retrieval methods in multiple indicators, and surpasses the effect of full fine-tuning, verifying its effectiveness and reliability in remote sensing image text retrieval tasks.
[0158] The present application also provides a terminal device, including: a memory and a processor, the memory is used to store a remote sensing cross-modal retrieval program based on large model fine-tuning, and when the remote sensing cross-modal retrieval program based on large model fine-tuning is executed by the processor, the steps of the remote sensing cross-modal retrieval method based on large model fine-tuning as above are implemented.
[0159] The present application also provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed, it implements the above-mentioned remote sensing cross-modal retrieval method based on large model fine-tuning.
[0160] In summary, the solution provided in this application effectively addresses the differences in image and text modal features through the collaborative work of a cross-modal asymmetric adapter, a visual enhancement adapter, and a text semantic adapter, achieving better alignment and optimization of image and text modal features during fine-tuning. At the same time, by expanding the single-task model to a multi-task model and introducing a dual-task consistency loss function, inter-modal consistency learning is strengthened, thereby improving retrieval performance. This approach has demonstrated excellent adaptability and significant performance improvements in remote sensing image-text retrieval tasks.
[0161] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0162] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A remote sensing cross-modal retrieval method based on large model fine-tuning, characterized by: The method comprises: S1. Using text describing a remote sensing image as description text, the remote sensing image and the description text are input into a trained remote sensing image text retrieval network for processing to obtain image features and text features after cross-modal fusion; S2. Based on the image features and the text features, obtain and output text and images that match the remote sensing image and the description text; Wherein, the step S1 includes: S11, extracting initial image features and initial text features from the remote sensing image and the description text through an image-text encoder; S12, performing cross-modal fusion processing on the initial image features and the initial text features through a cross-modal asymmetric adapter to obtain the cross-modal fused image features and text features; S13, optimizing the image features and the text features by using a dual-task consistency loss function; The step S12 includes: S121, performing dimensionality reduction processing on the initial image features and the initial text features; S122, using a nonlinear activation function GELU to process the initial image features and the initial text features after the dimensionality reduction process; S123, performing feature expression enhancement processing on the initial image features and the initial text features processed by the nonlinear activation function respectively; S124: Inputting the initial image features and the initial text features processed based on feature expression enhancement into a shared layer for cross-modal fusion processing to obtain cross-modal fused image features and text features; The cross-modal asymmetric adapter includes: a visual enhancement adapter and a text semantic adapter; wherein the visual enhancement adapter and the text semantic adapter respectively contain a linear transformation and a nonlinear activation function; the visual enhancement adapter performs feature expression enhancement processing on the initial image features through a differential attention mechanism, and the text semantic adapter performs feature expression enhancement processing on the initial text features through a hierarchical attention mechanism; The dual-task consistency loss function includes the following loss terms: The consistency constraint loss term uses the text modality as a teacher model to guide image modality learning based on the exponential moving average mechanism; Cross-modal constraint loss term, which measures the global alignment of image and text through cosine similarity; The classification constraint loss term calculates the difference between the predicted probability distribution and the true category distribution; The consistency constraint loss item, the cross-modal constraint loss item, and the classification constraint loss item are dynamically weighted based on an adaptive weight adjustment strategy to optimize the image features and the text features.
2. The remote sensing cross-modal retrieval method based on large model fine-tuning according to claim 1 is characterized in that: The image-text encoder includes an image encoder and a text encoder; the step S11 includes: extracting initial image features from the remote sensing image by the image encoder; Initial text features are extracted from the description text by the text encoder.
3. The remote sensing cross-modal retrieval method based on large model fine-tuning according to claim 1 is characterized in that: In step S123, the feature expression enhancement processing performed by the visual enhancement adapter on the initial image features includes: Map the query, key vector, and value vector to the output; Calculate an attention score based on the query and the key vector, and then perform a weighted sum of the attention score and the value vector; Eliminate attention score noise based on the softmax function.
4. The remote sensing cross-modal retrieval method based on large model fine-tuning according to claim 1 is characterized in that: In step S123, the feature expression enhancement processing performed by the text semantic adapter on the initial text features includes: Obtain word-level hidden representations from word-level annotations through a multi-layer perceptron; Calculating the importance weight of a word based on the word-level hidden representation and the word-level context vector; The word-level annotations are weighted and summed based on the importance weights to obtain a sentence vector.
5. A system using the remote sensing cross-modal retrieval method based on large model fine-tuning according to any one of claims 1 to 4, characterized in that: The system comprises: An image text input module, configured to use text describing a remote sensing image as a description text, and output the remote sensing image and the description text; A network composition and processing module, configured to process the remote sensing image and the description text input by the image and text input module to obtain cross-modal fused image features and text features; An image-text output module, which obtains text and images that match the remote sensing image and the description text based on the image features and text features after the cross-modal fusion and outputs them; The network composition and processing module includes: An image-text encoder that extracts initial image features and initial text features from the remote sensing image and the description text; Performing cross-modal fusion processing on the initial image features and the initial text features to obtain a cross-modal asymmetric adapter of the cross-modal fused image features and text features; A dual-task consistency loss function is used to optimize the image features and the text features.
6. A terminal device, characterized in that: The terminal device includes: a memory and a processor, the memory is used to store a remote sensing cross-modal retrieval program based on large model fine-tuning, and when the remote sensing cross-modal retrieval program based on large model fine-tuning is executed by the processor, the steps of the remote sensing cross-modal retrieval method based on large model fine-tuning as described in any one of claims 1 to 4 are implemented.
7. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed, the computer program implements the remote sensing cross-modal retrieval method based on large model fine-tuning according to any one of claims 1 to 4.
Citation Information
Patent Citations
Multi-scale information dynamic fusion remote sensing cross-modal image-text retrieval method
CN118939821A