A Remote Sensing Image-Text Retrieval Method Based on a Priori Indication Representation Framework

By using a method based on a priori indication representation framework in remote sensing text retrieval, visual and text modules are constructed and characterization alignment is performed, the problem of semantic noise affecting performance in the prior art is solved, and the accuracy of remote sensing text retrieval and the amplification capability of the model are improved.

CN117171373BActive Publication Date: 2025-06-10ZHEJIANG UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311179519.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-13
Publication Date
2025-06-10
Estimated Expiration
2043-09-13

AI Technical Summary

Technical Problem

The existing remote sensing graphic search methods have deteriorated performance when dealing with semantic noise, making it difficult to effectively model the remote dependence relationship between modes, affecting the amplification capability of the model.

Method used

Using a method based on a priori indication representation framework, the visual indication representation module and language cyclic attention module are constructed through graphic preprocessing, and feature extraction of visual and text modalities is realized, and the performance alignment of visual and text modes is performed by comparing the loss function and attribution loss function to reduce the impact of semantic noise.

Benefits of technology

The accuracy of remote sensing image and text retrieval is improved, and image redundancy is filtered through prior instructions, and the modeling of remote dependencies is realized, which improves the model's performance and amplification capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117171373B_ABST
    Figure CN117171373B_ABST
Patent Text Reader

Abstract

A remote sensing image-text retrieval method based on a prior indication representation framework, including image-text preprocessing, construction of an image-text retrieval model, and representation alignment. The construction of the image-text retrieval model includes three parts: constructing an image-text pre-encoding, constructing a visual indication representation module, and a language recurrent attention module. The visual indication representation module sorts and filters image features through a belief matrix to achieve redundancy filtering and filter redundant information in remote sensing images. The language recurrent attention module enhances text representation by recursively using the features of the previous time step to activate the features of the current time step. Among them, an attribution loss function is designed in the representation alignment to constrain the inter-class relationship and reduce the semantic confusion area. The present invention solves the problem of performance decline caused by semantic noise in existing remote sensing image-text retrieval methods and improves the performance of remote sensing image-text retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of remote sensing, and particularly relates to a method for remote sensing image-text retrieval based on a prior indication representation framework. Background Art

[0002] Remote sensing image-text retrieval is an important information retrieval technology, aiming to associate remote sensing images with relevant text information and retrieve matching image patches and text content. The main goal of this technology is to accurately retrieve results that match a given query text or image patch from a large-scale remote sensing image database, such as data collected by satellites or aerial drones, and it plays a significant role in fields such as resource investigation, disaster detection, and agricultural production. Small-scale targets in remote sensing images are easily interfered by semantic noise, such as backgrounds and irrelevant objects. Excessive attention to semantic noise will affect visual and text representations, increase semantic confusion areas, and thus seriously affect retrieval performance.

[0003] In the field of remote sensing image-text retrieval, most existing methods are based on convolutional neural network-based visual representations and recurrent neural network-based text representations, and are optimized using pairwise triplet losses, but this is not conducive to modeling long-range dependencies between modalities and improving the expandability of remote sensing image-text retrieval models. Summary of the Invention

[0004] To solve the performance degradation caused by semantic noise in existing remote sensing image-text retrieval methods, the present invention proposes a remote sensing image-text retrieval method based on a prior indication representation framework to improve remote sensing image-text retrieval performance.

[0005] To achieve the above objective, the technical solution of the present application is as follows:

[0006] A remote sensing image-text retrieval method based on a prior indication representation framework, the method comprising the following steps:

[0007] Step 1, image-text preprocessing, performing preprocessing operations on the image and text inputs of a remote sensing image-text dataset;

[0008] Step 2, building an image-text retrieval model, including constructing image-text pre-encoding, constructing a visual indication representation module and a language recurrent attention module, to achieve feature extraction of visual and text modalities and obtain final visual and text embedding features;

[0009] Step 3, representation alignment, including similarity measurement and subspace representation; calculating the cosine similarity matrix of visual and text modality features, and designing loss functions, including a contrast loss function and an attribution loss function, to achieve alignment of images and texts by minimizing the loss functions.

[0010] Further, the first step includes the following sub-steps:

[0011] Step 1.1: Preprocessing of remote sensing images; dividing the image data into a training set, a validation set, and a test set; performing data augmentation on the training and validation data, including scaling, random cropping, random flipping, and normalization processing;

[0012] Step 1.2: Preprocessing of text data; each remote sensing image corresponds to five text descriptions; first, add a special Token [CLS] at the beginning of the first sentence to mark the start of the sentence, and use [SEP] to mark the end of the sentence; then establish a word vector table to convert each word into a one-dimensional vector.

[0013] Furthermore, the second step includes the following sub-steps:

[0014] Step 2.1: Constructing a text-image pre-encoder; the text-image pre-encoder includes a visual encoder, an indicator encoder, and a text encoder, and the process is as follows:

[0015] Step 2.1.1: Using the Swin Transformer network as the visual encoder to extract the global and local relevant features of the image;

[0016] Step 2.1.2: Using the ResNet network pre-trained on the AID dataset as the indicator encoder to obtain the indicator embedding features, so as to help obtain an unbiased visual representation;

[0017] Step 2.1.3: Using a pre-trained Bert as the text encoder to extract the global and local relevant features of the text;

[0018] Step 2.2: Constructing a progressive attention encoder; the Transformer encoding layer consists of a self-attention layer and a cross-attention layer, and the transmission method is divided into a spatial progressive attention encoder and a time-slot progressive attention encoder according to the different information transmission methods between the Transformer encoding layers;

[0019] Step 2.3: Constructing a visual indicator representation module; first, calculate the belief matrix using the indicator embedding features obtained in Step 2.1.1 and the image features obtained in Step 2.1.2, then sort and filter the image features through the belief matrix to achieve redundant filtering, remove the redundant information in the remote sensing image, activate the filtered image features through the spatial progressive attention encoder to obtain the unbiased local relevant embedding features of the image, and finally map the unbiased local relevant embedding features of the image and add them to the global relevant features of the image to obtain the final visual embedding features;

[0020] Step 2.4: Construct a language cyclic attention module; put the text features obtained in Step 2.1.3 into a slot progressive attention encoder for activation to obtain unbiased local correlation embedding features of the text. Finally, map the unbiased local correlation embedding features of the text and add them to the text global correlation features to obtain the final text embedding features.

[0021] Furthermore, the third step includes the following sub-steps:

[0022] Step 3.1: Design a contrast loss function; calculate the cosine similarity of the corresponding final visual embedding features and final text embedding features to obtain the visual-to-text contrast loss and the text-to-visual contrast loss, thereby obtaining the overall contrast loss;

[0023] Step 3.2: Design an attribution loss function; for each image, calculate the text clustering center corresponding to the same category according to the scene category information to form a set of positive sample pairs from image to text; similarly, for each text, calculate the clustering center of the images corresponding to the same category according to the scene category information to form a set of positive sample pairs from text to image; then calculate the cosine similarity of the features of these positive sample pairs to obtain the visual-to-text and text-to-visual attribution loss representations, and finally obtain the overall attribution loss;

[0024] Step 3.3: Combine the overall contrast loss function and the overall attribution loss function as the overall loss function for model training.

[0025] The beneficial effects of the present invention are as follows: using prior indication to filter image redundancy, designing a progressive attention encoder to perform long-range dependency modeling, and using loss functions to constrain the relationships between classes, thereby improving the accuracy of remote sensing image-text retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flowchart of the image-text retrieval method of the present invention.

[0027] Figure 2 is a schematic diagram of the image-text retrieval network framework of the present invention.

[0028] Figure 3 is a schematic diagram of the Transformer encoding layer of the present invention.

[0029] Figure 4 is a schematic diagram of the progressive attention encoder of the present invention, where (a) is Spatial-PAE and (b) is Temporal-PAE. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0031] Referring to Figures 1 to 4 , a remote sensing image-text retrieval method based on a prior indication representation framework, includes the following steps:

[0032] Step 1. Image-text preprocessing: Perform preprocessing operations on the image and text inputs of the remote sensing image-text dataset;

[0033] The said Step 1 includes the following sub-steps:

[0034] Step 1.1: Preprocessing of remote sensing images; Divide the image data into a training set, a validation set, and a test set; Perform data augmentation processing on the training and validation data, including scaling, random cropping, random flipping, and normalization processing;

[0035] Step 1.2: Preprocessing of text data; Each remote sensing image corresponds to five text descriptions; First, add a special Token [CLS] at the beginning of the first sentence to mark the start of the sentence, and use [SEP] to mark the end of the sentence; Then establish a word vector table to convert each word into a one-dimensional vector;

[0036] Step 2. Construction of the image-text retrieval model, including constructing image-text pre-encoding, constructing a visual indication representation module and a language recurrent attention module, to achieve feature extraction of visual and text modalities and obtain the final visual and text embedding features;

[0037] The said Step 2 includes the following sub-steps:

[0038] Step 2.1: Construct image-text pre-encoding; The said image-text pre-encoding includes a visual encoder, an indication encoder, and a text encoder;

[0039] The process of the said Step 2.1 is as follows:

[0040] Step 2.1.1: First, divide the input image into image patches of a fixed size, and then use the Swin Transformer network to encode these image patches to obtain global correlation features and local correlation features which can be expressed as:

[0041]

[0042] Where represents the visual encoder, is the fine-tuned weight, [·,·] represents stacking and concatenating in the sequence length dimension, and m is the number of locally relevant features;

[0043] Step 2.1.2: Use the ResNet network pre-trained on the AID dataset as the indication encoder where is the pre-trained weight to obtain the indication embedding features

[0044] Step 2.1.3: Use a pre-trained Bert as the text encoder to encode the text T to obtain the globally relevant features and the locally relevant features can be expressed as:

[0045]

[0046] where represents the text encoder, is the fine-tuned weight, and n is the number of locally relevant features;

[0047] Step 2.2: As Figure 3 shown in the Transformer encoding layer, denoted as TEL; it includes a self-attention layer and a cross-attention layer. Given the query vector key vector and the value vector the scaled dot-product attention MHA(Q,K,V) can be calculated, and the formula is expressed as:

[0048] MHA(Q,K,V) = [head 1 ,head 2 ,...,head h T , (3)

[0049] where and are the projection matrices, and Softmax(·) is the Softmax function. Given two different sequences and the output TEL(S l-1 ,C l-1 ) of the Transformer encoding layer can be obtained, and the formula is expressed as:

[0050] S l = S l-1 +LN(MHA(S l-1 ,S l-1 ,S l-1 ))), (4)​

[0051] S l+1 = S l + LN(MLP(S l ))), (5)

[0052] C l = S l+1 + LN(MHA(C l-1 , S l+1 , S l+1 ))), (6)

[0053] C l+1 = C l + LN(MLP(C l ))), (7)

[0054] where LN(·) represents layer normalization, and MLP(·) represents a multi-layer perceptron, which is a feed-forward artificial neural network model.

[0055] As Figure 4 shown, according to the different information transfer methods between Transformer encoding layers, the transfer methods are divided into a spatial progressive attention encoder and a temporal attention encoder, denoted as Spatial-PAE and Temporal-PAE respectively. Among them, Spatial-PAE uses linear projection to make a spatial connection with the input sequence of the external source, and uses external knowledge containing global information to assist in remote dependence modeling. Temporal-PAE uses linear projection to make a temporal connection with the input sequence at the last moment, and calculates the attention map using the sequence outputs of the previous and current time steps.

[0056] Step 2.3: Construct a visual indication representation module (Spatial-PAE); First, use the indication embedding feature v ins obtained in Step 2.1.2 to calculate the belief matrix of the local correlation feature [v cls , E v . The formula is expressed as: The formula is expressed as:

[0057]

[0058] Then, sort and filter the features to achieve redundancy filtering, filtering the redundant information in the remote sensing image. The formula is expressed as:

[0059]

[0060] where represents sorting and filtering the sequence A according to to obtain the sequence B. Then, use Spatial-PAE for the external information Model the remote dependencies of the activated filtering features, which is expressed by the formula:

[0061]

[0062] where is the output of the i-th TEL, represents the i-th weight of the linear mapping. After that, an unbiased local correlation embedding feature is obtained and expressed by the formula:

[0063]

[0064] where Head(·) represents mapping the head embedding feature of the last layer to the unbiased embedding feature. Finally, the final visual embedding feature v emb = v cls + v loc ;

[0065] Step 2.4: Construct the Temporal-PAE (Temporal-Point Attention Embedding) module; first, recursively activate the text features using Temporal-PAE, which is expressed by the formula:

[0066]

[0067] where is the output of the i-th TEL, is the i-th weight of the linear mapping, and after that, an unbiased local correlation embedding feature is obtained and expressed by the formula:

[0068]

[0069] Finally, the final text embedding feature t emb = t cls + t loc .

[0070] Step 3: Feature alignment, including similarity measurement and subspace representation; calculate the cosine similarity matrix of the visual and text modalities, design loss functions, including the contrastive loss function and the attribution loss function, and achieve the alignment of images and texts by minimizing the loss functions;

[0071] The third step includes the following sub-steps:

[0072] Step 3.1: Construct the contrastive loss function; given a mini-batch of samples, randomly sample N positive pairs{(I 1 , T 1 ), (I 2 , T 2 ),..., (I i , Ti ),...,(I N ,T N )}, and its corresponding final visual embedding features and final text embedding features are {(v 1 , t 1 ), (v 2 , t 2 ),...,(v i , t i ),...,(v N , t N ). Then calculate the cosine similarity Thus, the visual-to-text contrast loss and text-to-visual contrast loss can be obtained, which are expressed by the formula as:

[0073]

[0074]

[0075] where τ is the temperature parameter. Finally, the contrast loss can be obtained, which is expressed by the formula as:

[0076]

[0077] Step 3.2: Construct the attribution loss function; Given a small sample, for each image, calculate the clustering center of the text features corresponding to the same category according to the scene category information, and form a set of positive sample pairs from image to text where represents the clustering center of the i-th text. Calculate the mean value of each dimension of the text features in the same category. Similarly, for each text, calculate the clustering center of the image features corresponding to the same category according to the scene category information, and form a set of positive sample pairs from text to image. Then calculate the cosine similarity from vision to text and the cosine similarity from text to vision Finally, the visual-to-text and text-to-visual attribution loss representations can be obtained, which can be expressed by the formula as:

[0078]

[0079]

[0080] Thus, the overall attribution loss is obtained, which is expressed by the formula as:

[0081]

[0082] Step 3.3: Combine the contrast loss function and the attribution loss function as the overall loss function for model training, which is expressed by the formula as:

[0083]

[0084] where λ cs is the central factor, representing the aggregation degree of the class center.

[0085] The present invention also provides a remote sensing image text retrieval system implementing the described prior indication representation framework, including the following modules:

[0086] Preprocessing before graphic and text encoding, preprocessing the images and texts of the remote sensing image-text dataset to obtain standard model input sample data;

[0087] Construction of the graphic and text retrieval model, including constructing graphic and text encoding, constructing a visual indication representation module and a language cyclic attention module, to implement the extraction of visual and text modality features and obtain the final visual and text embedding features;

[0088] Representation alignment, including similarity measurement and subspace representation, to obtain the cosine similarity matrix of visual and text modalities, designing loss functions, including a contrast loss function and an attribution loss function, and realizing the alignment of images and texts by minimizing the loss functions.

[0089] The above-mentioned modules correspond to steps one to three of the described method.

[0090] The above-described embodiments only represent two implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A remote sensing image-text retrieval method, characterized in that, the method comprises the following steps: Step 1, remote sensing image-text preprocessing, performing preprocessing operations on the image and text inputs of the remote sensing image-text dataset; Step 2, building a remote sensing image-text retrieval model, including constructing text-image pre-encoding, constructing a visual indication representation module and a language recurrent attention module, to achieve extraction of visual and text modality features and obtain final visual and text embedding features; Step 3, representation alignment, including similarity measurement and subspace representation; calculating the cosine similarity matrix of visual and text modality features, designing loss functions, including a contrastive loss function and an attribution loss function, and achieving alignment of images and texts by minimizing the loss functions; The said Step 2 comprises the following sub-steps: Step 2.1: Construct text-image pre-encoding; the said text-image pre-encoding includes a visual encoder, an indication encoder and a text encoder, and the process is as follows: Step 2.1.1: Use the Swin Transformer network as the visual encoder to extract the global relevant features and local relevant features of the image; Step 2.1.2: Use the ResNet network pre-trained on the AID dataset as the indication encoder to obtain indication embedding features, so as to guide the remote sensing image-text retrieval model to perform unbiased visual representation; Step 2.1.3: Use a pre-trained Bert as the text encoder to extract the global relevant features and local relevant features of the text; Step 2.2: Construct a progressive attention encoder; the Transformer encoding layer consists of a self-attention layer and a cross-attention layer. According to the different information transfer methods between Transformer encoders, the transfer methods are divided into a spatial progressive attention encoder and a time-slot progressive attention encoder. The spatial progressive attention encoder is denoted as Spatial-PAE, and the time-slot progressive attention encoder is denoted as Temporal-PAE. Among them, Spatial-PAE uses linear projection to perform spatial connection with the input sequence of the external source, and uses the external knowledge containing global information to assist in remote dependence modeling. Temporal-PAE uses linear projection to perform time connection with the input sequence of the last moment, and calculates the attention map using the sequence outputs of the previous and current time steps; Step 2.3: Construct a visual indication representation module; first calculate and obtain a belief matrix using the global relevant features and local relevant features of the image obtained in Step 2.1.1 and the indication embedding features obtained in Step 2.1.2, then sort and filter the local relevant features through the belief matrix to achieve redundancy filtering, remove the redundant information in the remote sensing image, activate the filtered local relevant features through the spatial progressive attention encoder to obtain the unbiased local relevant embedding features of the image, and finally map the unbiased local relevant embedding features of the image and add them to the global relevant features of the image to obtain the final visual embedding features; Step 2.4: Construct a language cyclic attention module; put the text features obtained in Step 2.1.3, i.e., the global relevant features and local relevant features of the text, into a time-slot progressive attention encoder for activation to obtain unbiased local relevant embedding features of the text, and finally map the unbiased local relevant embedding features of the text and add them to the global relevant features of the text to obtain the final text embedding features.

2. A remote sensing image-text retrieval method as claimed in claim 1, wherein, the first step includes the following sub-steps: Step 1.1: Preprocessing of remote sensing images; dividing the image data into a training set, a validation set, and a test set; performing data augmentation processing on the training and validation data, including scaling, random cropping, random flipping, and normalization processing; Step 1.2: Preprocessing of text data; each remote sensing image corresponds to five text descriptions; first, add a special Token [CLS] at the beginning of the first sentence to mark the start of the sentence, and use [SEP] to mark the end of the sentence; Then establish a word vector table to convert each word into a one-dimensional vector.

3. A remote sensing image-text retrieval method as claimed in claim 1, wherein, In the said step 2.1.1, first, the input image is divided into image patches of a fixed size, and then these image patches are encoded by the Swin Transformer network to obtain global correlation features and local correlation features which are expressed as: Among them represents the visual encoder, are the fine-tuning weights, [·,·] represents stacking and concatenation in the sequence length dimension, and m is the number of locally relevant features of the image; In the step 2.1.2, the ResNet network pre-trained on the AID dataset is used as the indication encoder wherein are the pre-trained weights, so as to obtain the indication embedding features In the step 2.1.3, a pre-trained Bert is used as a text encoder to encode the text T, so as to obtain global correlation features and local correlation features which is expressed as: Among them represents the text encoder, is the fine-tuning weight, and n is the number of text locally relevant features; In step 2.2, the Transformer encoding layer is denoted as TEL; it includes a self-attention layer and a cross-attention layer. Given a query vector key vector and value vector the scaled dot-product attention MHA(Q, K, V) is calculated, and the formula is expressed as: MHA(Q, K, V) = [head 1 , head 2 ,..., head h T , (3)​ where and are projection matrices, Softmax(·) is the Softmax function. Given two different sequences and the output TEL(S l-1 , C l-1 ) of the Transformer encoding layer is obtained, which is expressed by the formula as: S l = S l-1 + LN(MHA(S l-1 , S l-1 , S l-1 ))), (4) S l+1 = S l + LN(MLP(S l ), (5) C l = S l+1 + LN(MHA(C l-1 , S l+1 , S l+1 ), (6) C l+1 = C l + LN(MLP(C l )), (7) where LN(·) represents layer normalization, and MLP(·) represents a multi-layer perceptron, which is a feed-forward artificial neural network model; In step 2.3, first, use the indication embedding feature v obtained in step 2.1.2 ins to calculate the global correlation feature and the local correlation feature [v cls , E v belief matrix The formula is expressed as: Then sort and filter the local relevant features to achieve redundancy filtering and filter redundant information in the remote sensing image, which is expressed by the formula: Among them indicates that sequence A is sorted and filtered according to to obtain sequence B. Then, Spatial-PAE is used to model the long-range dependencies of the filtered features activated by external information The formula is expressed as: Among them is the output of the i-th TEL, represents the i-th weight of the linear mapping. After that, an unbiased local correlation embedding feature is expressed by the formula: where Head(·) represents mapping the head embedding features of the last layer to unbiased embedding features, and finally obtaining the final visual embedding feature v emb = v cls + v loc ; In the said Step 2.4, first use Temporal-PAE to recursively activate the text features, which is expressed by the formula: where is the output of the i-th TEL, is the i-th weight of the linear mapping, After that, an unbiased local correlation embedding feature is obtained Expressed by the formula: Finally, the final text embedding feature t is obtained. emb = t cls + t loc .

4. A remote sensing image-text retrieval method as claimed in any one of claims 1 to 3, wherein, the third step includes the following sub-steps: Step 3.1: Design a contrast loss function; calculate the cosine similarity of the corresponding final visual embedding features and final text embedding features to obtain the visual-to-text contrast loss and the text-to-visual contrast loss, so as to obtain the contrast loss; Step 3.2: Design an attribution loss function; for each image, calculate the text clustering center corresponding to the same category according to the scene category information to form a set of positive sample pairs from image to text, and for each text, calculate the clustering center of the images corresponding to the same category according to the scene category information to form a set of positive sample pairs from text to image, and then calculate the cosine similarity of the features of these positive sample pairs to obtain the visual-to-text and text-to-visual attribution loss representations, and finally obtain the overall attribution loss; Step 3.3: Combine the contrast loss function and the attribution loss function as the overall loss function for training the remote sensing image-text retrieval model.

5. A remote sensing image-text retrieval method as claimed in claim 4, wherein, In step 3.1, randomly sample N positive pairs {(I 1 ,T 1 ),(I 2 ,T 2 ),...,(I i ,T i ),...,(I N ,T N )} from all the training sets, and their corresponding final visual embedding features and final text embedding features are {(v 1 ,t 1 ),(v 2 ,t 2 ),...,(v i ,t i ),...,(v N ,t N ). Then calculate the cosine similarity Thus, the visual-to-text contrastive loss and the text-to-visual contrastive loss are obtained, which are expressed by the formula as follows: where τ is the temperature parameter, and finally obtain the contrast loss, which is expressed by the formula: In step 3.2, a belonging loss function is constructed; given a small sample, for each image, the clustering center of text features corresponding to the same category is calculated according to the scene category information, forming a set of positive sample pairs from image to text. Where represents the clustering center of the i-th text, and the mean value of each dimension of the text features of the same category is calculated. Similarly, for each text, the clustering center of the image features corresponding to the same category is calculated according to the scene category information, forming a set of positive sample pairs from text to image, and then the cosine similarity from vision to text is calculated. And the cosine similarity from text to vision. Finally, the belonging loss representations from vision to text and from text to vision are obtained, which are expressed by the formula: Thus obtain the overall attribution loss, which is expressed by the formula: In the said Step 3.3, combine the contrast loss function and the attribution loss function as the overall loss function for training the remote sensing image-text retrieval model, which is expressed by the formula: where λ cs is the central factor, representing the aggregation degree of the class center.

Citation Information

Patent Citations

  • Image-text retrieval system and method based on multi-angle self-attention mechanism

    CN109992686A

  • Weak annotation Hash image retrieval architecture of knowledge graph embedded attention mechanism

    CN115329120A

  • Image-text cross-modal retrieval network training method, application method and electronic equipment

    CN116304307A

  • Remote sensing image cross-modal retrieval method based on layout semantic joint significant representation

    CN116561365A