Training Method of Cross-Modal Retrieval Model and Remote Sensing Image Text Retrieval Method
By using the training method of cross-modal search model in remote sensing image text retrieval, the global and local features of remote sensing images and text are extracted and aligned, the problem of insufficient local fine-grained information capture in the prior art is solved, and higher matching accuracy and retrieval accuracy are achieved.
Patent Information
- Application Number
- CN202510174770.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Existing remote sensing image text retrieval methods have shortcomings in capturing local fine-grained information between images and text, especially when there are subtle differences between images and text content, the retrieval performance is limited.
The training method of cross-modal retrieval model is adopted, and the global and local features of remote sensing images and text are extracted through visual encoder and text encoder, and the global and local features are aligned based on the contrast learning method, and mask modeling is performed to enhance the correlation learning of single-modal features.
It improves the matching accuracy between remote sensing images and text and the accuracy of cross-modal retrieval, can effectively handle subtle differences between images and text, and significantly improves the performance of remote sensing image and text retrieval.
Smart Images

Figure CN119646255B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing cross-modal retrieval, and specifically relates to a training method for a cross-modal retrieval model and a remote sensing image-text retrieval method. Background Art
[0002] With the continuous development of remote sensing technology, the acquisition of remote sensing images has become easier, and a large amount of remote sensing image data has been accumulated. However, how to quickly and accurately retrieve the information required by users from the vast amount of remote sensing images is still an important research issue. Therefore, the problem of remote sensing image-text retrieval (RSITR) has been proposed. Remote sensing image-text retrieval refers to querying through natural language descriptions to retrieve remote sensing images that match them, or finding relevant text descriptions through image queries. The diversity and similarity of remote sensing images make the retrieval between images and texts more complex.
[0003] In the existing remote sensing image-text retrieval methods, the remote sensing image-text retrieval method based on a convolutional neural network has deficiencies in capturing local fine-grained information between images and texts. Especially when there are subtle differences between the image and text contents, the retrieval performance is often limited; while the remote sensing image-text retrieval method based on Transformer still has technical problems such as insufficient alignment of fine-grained features. Summary of the Invention
[0004] In view of the above problems, the present invention provides a training method for a cross-modal retrieval model and a remote sensing image-text retrieval method that improve the matching accuracy and cross-modal retrieval efficiency.
[0005] According to the first aspect of the present invention, a training method for a cross-modal retrieval model is provided, including:
[0006] Using the visual encoder of the cross-modal retrieval model to extract features from the preprocessed remote sensing image samples to obtain the global features of the image samples and the local features of the image samples;
[0007] Using the text encoder of the cross-modal retrieval model to extract features from the preprocessed text samples to obtain the global features of the text samples and the global features of the phrase samples;
[0008] Based on the contrastive learning method, globally align the global features of the image sample global features and the text sample global features by calculating the global similarity between the constructed positive sample pairs and negative sample pairs, and calculate the global alignment loss value;
[0009] Perform local feature alignment on the local features of the image samples and the global features of the phrase samples using the local similarity between the local features of the image samples and the global features of the phrase samples, and calculate the local alignment loss value;
[0010] Perform masked modeling on the preprocessed remote sensing image samples and the preprocessed text samples respectively, and calculate the masked loss values of the remote sensing image samples and the text samples during the masked modeling process respectively;
[0011] Perform weighted calculation on the global alignment loss value, the local alignment loss value and the masked loss value to obtain a weighted loss value, and use the weighted loss value to optimize the parameters of the cross-modal retrieval model.
[0012] According to an embodiment of the present invention, the above method for training a cross-modal retrieval model further includes:
[0013] Iteratively perform the feature extraction operation of the remote sensing image samples and the text samples, the weighted loss value calculation operation and the model parameter optimization operation until the preset training conditions are met, and obtain a trained cross-modal retrieval model;
[0014] Among them, the above-mentioned use of the visual encoder of the cross-modal retrieval model to extract features from the preprocessed remote sensing image samples to obtain the global features and local features of the image samples includes:
[0015] Based on the dimension information of the remote sensing image, perform image segmentation on the remote sensing image samples to obtain a plurality of non-overlapping rectangular region image samples;
[0016] Use the learnable linear mapping module of the cross-modal retrieval model to project each rectangular region image sample into a one-dimensional space to obtain a plurality of initial local features of the image samples;
[0017] Add a special encoding for aggregating global information to each initial local feature of the image sample to obtain an initial local feature of the image sample with special encoding;
[0018] Add image position encoding to each initial local feature of the image sample with special encoding to obtain a plurality of encoded initial local features of the image sample;
[0019] Use the visual encoder of the cross-modal retrieval model to process all the encoded initial local features of the image samples to obtain the local features and global features of the image samples.
[0020] According to an embodiment of the present invention, the above-mentioned use of the text encoder of the cross-modal retrieval model to extract features from the preprocessed text samples to obtain the global features of the text samples and the global features of the phrase samples includes:
[0021] Segment the text sample using the byte pair encoding method to obtain multiple phrases from the text sample, and insert a start identifier and an end identifier at the start position and the end position of the text sample respectively;
[0022] Map each phrase to a word vector using the encoding matrix of the cross-modal retrieval model, and add text position encoding to each word vector;
[0023] Use the text encoder of the cross-modal retrieval model to extract features from the start identifier, the end identifier, and all word vectors with text position encoding to obtain the text global features;
[0024] Use a predefined natural language processing tool to extract multiple noun phrases from the text sample to complete the enhancement operation of the text sample, and use the text encoder to extract global features from each noun phrase to obtain the global features of the phrase sample.
[0025] According to an embodiment of the present invention, the above visual encoder is constructed based on the Vision Transformer model;
[0026] Among them, the text encoder is constructed based on the Bert model;
[0027] Among them, the predefined natural language processing tool includes the Spacy library.
[0028] According to an embodiment of the present invention, based on the contrast learning method, the global features of the image sample global features and the text sample global features are globally feature-aligned by calculating the global similarity between the constructed positive sample pairs and negative sample pairs, and calculating the global alignment loss value includes:
[0029] Based on the pairing information between the remote sensing image sample and the text sample, construct positive sample pairs and negative sample pairs;
[0030] Based on the contrast learning method, calculate the global similarity of the positive sample pairs and negative sample pairs respectively to obtain the positive global similarity and the negative global similarity;
[0031] Use the positive global similarity and the negative global similarity to globally feature-align the image sample global features and the text sample global features;
[0032] Use a predefined global alignment loss function to calculate the global alignment loss value during the global feature alignment process.
[0033] According to an embodiment of the present invention, the local features of the image sample and the global features of the phrase sample are locally feature-aligned using the local similarity between the local features of the image sample and the global features of the phrase sample, and calculating the local alignment loss value includes:
[0034] Calculate the cosine similarity between the global features of the phrase samples and the local features of the image samples;
[0035] Perform an averaging operation on all cosine similarities higher than a preset similarity threshold to obtain the local similarity;
[0036] Align the local features of the image samples and the global features of the phrase samples using the local similarity;
[0037] Calculate the local alignment loss value during the local feature alignment process using a predefined local alignment loss function.
[0038] According to an embodiment of the present invention, the above-mentioned respectively perform masked modeling on the preprocessed remote sensing image samples and the preprocessed text samples, and calculate the masked loss values of the remote sensing image samples and the text samples during the masked modeling process, including:
[0039] Based on a preset image masking ratio, randomly mask the preprocessed remote sensing image samples to obtain masked remote sensing image samples;
[0040] Extract features from the masked remote sensing image samples using the visual encoder of the cross-modal retrieval model to obtain multiple masked feature samples;
[0041] Process all masked feature samples using the visual mask regression head of the cross-modal retrieval model to obtain the reconstructed remote sensing image samples;
[0042] Process the reconstructed remote sensing image samples and the remote sensing image samples using a predefined image masking loss function to obtain the image masking loss value.
[0043] According to an embodiment of the present invention, the above-mentioned respectively perform masked modeling on the preprocessed remote sensing image samples and the preprocessed text samples, and calculate the masked loss values of the remote sensing image samples and the text samples during the masked modeling process further includes:
[0044] Randomly replace some words in the text samples with masked specific words based on a preset image masking ratio to obtain masked text samples;
[0045] Extract features from the masked text samples using the text encoder of the cross-modal retrieval model to obtain multiple word features;
[0046] Predict each word feature using the text mask classification head of the cross-modal retrieval model to obtain the predicted text;
[0047] Process the text samples and the predicted text using a predefined text masking loss function to obtain the text masking loss value.
[0048] According to an embodiment of the present invention, the above-mentioned predefined global alignment loss function and predefined local alignment loss function are constructed based on a contrastive learning loss function;
[0049] Among them, the predefined image mask loss function is constructed based on the mean squared error function;
[0050] Among them, the predefined text mask loss function is constructed based on the cross-entropy loss function.
[0051] According to a second aspect of the present invention, a remote sensing image text retrieval method is provided, including:
[0052] Using the trained cross-modal retrieval model to preprocess the remote sensing image and the target text respectively, obtaining the preprocessed remote sensing image and the preprocessed target text, where the trained cross-modal retrieval model is trained based on the above-mentioned training method for the cross-modal retrieval model for remote sensing image text retrieval;
[0053] Using the trained cross-modal retrieval model to extract features from the preprocessed remote sensing image, obtaining the global image feature and the local image feature, and extracting features from the preprocessed target text, obtaining the global text feature and the global phrase feature;
[0054] Using the trained cross-modal retrieval model to calculate the global similarity between the global image feature and the global text feature, and calculating the local similarity between the local image feature and the global phrase feature;
[0055] Using the trained cross-modal retrieval model to perform weighted calculation on the global similarity and the local similarity, obtaining the weighted similarity, and screening the matching result between the remote sensing image and the target text based on the weighted similarity, obtaining the retrieval result between the remote sensing image and the target text.
[0056] The training method of the cross-modal retrieval model provided by the present invention can be used in the technical field of remote sensing image text retrieval. By extracting the global and local features of remote sensing image samples and text samples through the constructed cross-modal retrieval model, and aligning the global features and local features based on the contrastive learning method, it not only realizes the feature alignment of image-text samples on a large scale, but also can realize the feature alignment of image-text samples on a fine-grained scale; through the image-text sample mask modeling, it enhances the correlation learning of single-modal features and further improves the matching ability of images and texts. In addition, by combining the global similarity and local similarity scores, the accuracy of the trained cross-modal retrieval model in remote sensing image-text retrieval is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0058] Figure 1 is an application scenario diagram of a training method for a cross-modal retrieval model and a remote sensing image text retrieval method according to an embodiment of the present invention;
[0059] Figure 2 is a flowchart of a training method for a cross-modal retrieval model according to an embodiment of the present invention;
[0060] Figure 3 is a flowchart of a remote sensing image text retrieval method according to an embodiment of the present invention;
[0061] Figure 4 is a structural block diagram of a remote sensing image text retrieval device according to an embodiment of the present invention;
[0062] Figure 5 is a block diagram of an electronic device suitable for implementing a training method for a cross-modal retrieval model and a remote sensing image text retrieval method according to an embodiment of the present invention. Detailed Embodiments
[0063] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0064] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0065] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0066] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0067] In existing remote sensing image-text retrieval methods, most rely on convolutional neural networks (CNNs) to extract image features and recurrent neural networks (RNNs) to extract text features, and achieve image and text retrieval through the alignment of global features. However, these methods are insufficient in capturing local fine-grained information between images and texts. Especially when there are subtle differences between image and text contents, the retrieval performance of existing methods is often limited. Recent Transformer-based image-text retrieval methods utilize the powerful global information fusion ability of Transformer, which can improve the matching accuracy between images and texts. Among them, the prominent one is the CLIP (Contrastive Language-Image Pre-training) model, which has made significant progress in image-text retrieval performance. However, although Transformer can effectively process global features, there is still a problem of insufficient alignment of fine-grained features. To make up for this defect, it is necessary to combine the alignment of global features and local features to improve the accuracy and robustness of remote sensing image-text retrieval. How to effectively combine global alignment with fine-grained local alignment has become a key issue in improving the accuracy of remote sensing image-text retrieval.
[0068] To at least solve one of the problems in the prior art, the present invention provides a remote sensing image and text retrieval method of Global and Local Alignment with Phrase Argumentation (GLAPA) that combines global alignment with phrase enhancement. The above method provided by the present invention can effectively improve the matching accuracy between remote sensing images and texts and the cross-modal retrieval accuracy rate, and can be widely applied to fields such as disaster monitoring, agricultural production, and environmental monitoring in the remote sensing field.
[0069] An embodiment of the present invention provides a remote sensing image text retrieval method, which relates to the technical field of remote sensing cross-modal retrieval and can be applied to fields such as disaster monitoring, agricultural production, and environmental monitoring in the remote sensing field.
[0070] Figure 1 It is an application scenario diagram of a training method of a cross-modal retrieval model and a remote sensing image text retrieval method according to an embodiment of the present invention.
[0071] Such as Figure 1As shown, the application scenario 100 according to this embodiment may include fields such as disaster monitoring, agricultural production, and environmental monitoring in the field of remote sensing cross-modal retrieval technology. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0072] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for examples).
[0073] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0074] The server 105 may be a server that provides various services, such as a background management server that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only for examples). The background management server may analyze and process data such as received user requests, etc., and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0075] It should be noted that the training method of the cross-modal retrieval model and the remote sensing image text retrieval method provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the remote sensing image text retrieval device provided by the embodiments of the present invention can generally be set in the server 105. The training method of the cross-modal retrieval model and the remote sensing image text retrieval method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the remote sensing image text retrieval device provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0076] It should be understood, Figure 1The numbers of the terminal devices, networks, and servers in [it] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0077] Based on the Figure 1 described scenario, through Figure 2 and Figure 3 the training method of the cross-modal retrieval model and the remote sensing image text retrieval method of the disclosed embodiments will be described in detail.
[0078] Figure 2 is a flowchart of the cross-modal retrieval model according to an embodiment of the present invention.
[0079] As Figure 2 shown, the above-mentioned training of the cross-modal retrieval model based on the multi-modal pre-trained neural network obtains the trained retrieval model including operations S210 to S270.
[0080] In operation S210, the visual encoder of the cross-modal retrieval model is used to extract features from the pre-processed remote sensing image samples to obtain the global features and local features of the image samples.
[0081] The above cross-modal retrieval model is constructed based on a multi-modal pre-trained neural network. The multi-modal pre-trained neural network includes CLIP (Contrastive Language-Image Pre-Training), and CLIP can perform image classification, image retrieval, text generation, multi-modal search, etc.
[0082] In operation S220, the text encoder of the cross-modal retrieval model is used to extract features from the pre-processed text samples to obtain the global features of the text samples and the global features of the phrase samples.
[0083] Since the text samples are composed of multiple words or phrases, the global features of the obtained phrase samples can be used to represent the local features of the text samples.
[0084] In operation S230, based on the contrastive learning method, the global features of the image sample global features and the text sample global features are globally feature-aligned by calculating the global similarity between the constructed positive sample pairs and negative sample pairs, and the global alignment loss value is calculated.
[0085] The contrastive learning method can improve the generalization ability of the trained model, reduce the dependence on labeled data, enhance the model performance, and can adapt to different data types.
[0086] In operation S240, local feature alignment is performed on the local features of the image samples and the global features of the phrase samples using the local similarity between the local features of the image samples and the global features of the phrase samples, and the local alignment loss value is calculated.
[0087] The local feature alignment realizes the alignment of images and texts at a fine-grained level, and the local alignment process is also based on the contrast learning method.
[0088] In operation S250, masked modeling is respectively performed on the preprocessed remote sensing image samples and the preprocessed text samples, and the masked loss values of the remote sensing image samples and the text samples during the masked modeling process are respectively calculated.
[0089] In computer vision tasks, masks can be used to mark specific regions in images, such as object detection or image segmentation. A mask is a matrix with the same size as the image, where each element represents the state of the corresponding pixel, which can better process remote sensing images.
[0090] In operation S260, weighted calculation is performed on the global alignment loss value, the local alignment loss value, and the masked loss value to obtain the weighted loss value, and the cross-modal retrieval model is optimized with the weighted loss value.
[0091] In operation S270, the feature extraction operation of the remote sensing image samples and the text samples, the weighted loss value calculation operation, and the model parameter optimization operation are iteratively performed until the preset training conditions are met, and the trained cross-modal retrieval model is obtained.
[0092] The training method of the cross-modal retrieval model provided by the present invention can be used in the technical field of remote sensing image text retrieval. By extracting the global and local features of remote sensing image samples and text samples, and aligning the global features and local features based on the contrast learning method, it not only realizes the feature alignment of image and text samples at a large scale, but also can realize the feature alignment of image and text samples at a fine-grained level; through the masked modeling of image and text samples, the correlation learning of single-modal features is enhanced, and the matching ability of images and texts is further improved. In addition, by combining the global similarity and local similarity scores, the accuracy of remote sensing image and text retrieval is significantly improved.
[0093] According to the embodiments of the present invention, the above-mentioned predefined global alignment loss function and predefined local alignment loss function are constructed based on the contrast learning loss function; wherein, the predefined image masked loss function is constructed based on the least square error function; wherein, the predefined text masked loss function is constructed based on the cross-entropy loss function.
[0094] The following further elaborates on the training process of the cross-modal retrieval model involved in the above embodiments through specific examples.
[0095] The training process of the cross-modal retrieval model of the present invention mainly involves feature extraction of remote sensing images and texts, global feature alignment, local feature alignment, masked modeling, as well as training and inference.
[0096] Among them, global feature alignment: Use the contrastive learning method to align the global features of images and texts to ensure the consistency of global semantics.
[0097] Among them, local feature alignment: By extracting concept phrases in the text and aligning them with the local region features in the image, capture the relationship between images and texts at the fine-grained level.
[0098] Among them, masked modeling: Introduce the masked modeling strategy to enhance the correlation learning of unimodal features and further improve the matching ability of images and texts.
[0099] Among them, training and inference: In the training stage, train by weighting the global and local alignment losses and additionally using the masked modeling losses of images and texts. In the inference stage, combine the global similarity and local similarity, and obtain the final retrieval similarity score through weighted average, and balance the influence of the two through weight parameters.
[0100] The following will detail each operation involved in the training process of the cross-modal retrieval model through specific embodiments.
[0101] According to an embodiment of the present invention, the above-mentioned preprocessing of remote sensing image samples and feature extraction of the preprocessed remote sensing image samples using the visual encoder of the cross-modal retrieval model to obtain the global features and local features of the image samples include: Based on the dimension information of the remote sensing image, perform image segmentation on the remote sensing image samples to obtain multiple non-overlapping rectangular region image samples; Use the learnable linear mapping module of the cross-modal retrieval model to project each rectangular region image sample into a one-dimensional space to obtain multiple initial local features of the image samples; Add special encoding for aggregating global information to each initial local feature of the image sample to obtain the initial local features of the image sample with special encoding; Add image position encoding to each initial local feature of the image sample with special encoding to obtain multiple encoded initial local features of the image sample; Use the visual encoder of the cross-modal retrieval model to process all the encoded initial local features of the image sample to obtain the local features and global features of the image sample.
[0102] According to an embodiment of the present invention, the above-mentioned visual encoder is constructed based on the Vision Transformer model; among them, the text encoder is constructed based on the Bert model; among them, the predefined natural language processing tool includes the Spacy library.
[0103] In the above embodiments, for the extraction process of the global features and local features of the remote sensing image samples, a vision encoder based on Vision Transformer is utilized.
[0104] Specifically, given an input image , denotes -dimensional Euclidean space (the same hereinafter). First, the image is divided into non-overlapping rectangular regions, where , and represent the height, width, and number of channels of the image respectively, and is the side length of the rectangular region. The features of each rectangular region are projected into a one-dimensional vector space through a learnable linear mapping to obtain the feature representation of each rectangular region, where is the feature dimension, represents the image (the same hereinafter), represents the number or quantity (the same hereinafter), that is, the th rectangular region. Then, position encoding and a special feature [CLS] are added to the features of all rectangular regions. The [CLS] feature is used to converge the global information representing the entire picture. Then, all the feature vectors are input into the -layer Transformer model to obtain the global feature of the image and the local features of each rectangular region, where, i.e., global, means global (the same hereinafter), and is expressed by formula (1) as follows:
[0105] (1),
[0106] where, represents the CLIP vision encoder, is the global feature obtained by converging the [CLS] feature, are the local features of different rectangular regions.
[0107] According to an embodiment of the present invention, the above preprocessing of the text sample and the feature extraction of the preprocessed text sample using the text encoder of the cross-modal retrieval model to obtain the global text feature and the global phrase sample feature include: segmenting the text sample using the byte pair encoding method to obtain multiple phrases from the text sample, and inserting a start identifier and an end identifier at the start position and the end position of the text sample respectively; mapping each phrase to a word vector using the encoding matrix of the cross-modal retrieval model, and adding text position encoding to each word vector; using the text encoder of the cross-modal retrieval model to perform feature extraction on the start identifier, the end identifier, and all word vectors with text position encoding to obtain the global text feature; using a predefined natural language processing tool to extract multiple noun phrases from the text sample to complete the enhancement operation of the text sample, and using the text encoder to perform global feature extraction on each noun phrase to obtain the global phrase sample feature.
[0108] In the above feature extraction process of the text sample, the text encoder of the cross-modal retrieval model is used, and the text encoder is constructed based on the Bert model.
[0109] For a given input text , first, the text is segmented using Byte Pair Encoding (BPE), and special words [SOS] and [EOS] are inserted at the start and end of the text respectively. Then, each word in the text is mapped to a word vector through the encoding matrix, and position encoding is added. Next, all word vectors pass through layers of Transformer to obtain the feature representation of each word , where , that is, start of sentence, represents the start of the text (the same below), , that is, end of sentence, represents the end of the text (the same below), , that is, text, represents the text (the same below). Among them is used as the global feature of the text , is the local feature of the text. This process is represented by formula (2) as follows:
[0110] (2)
[0111] Among them, represents the CLIP text encoder, is the global feature of the text, is the local feature of each word in the text.
[0112] In addition to performing feature extraction on the entire text, the present invention also performs phrase argumentation on each independent phrase in the text. The present invention uses the well-known NLP processing tool, the spacy library, to extract from the text noun phrases , and inputs each noun phrase separately into the text encoder to obtain the global and local features of each noun phrase. Here, only the global features of each noun phrase are used. For example, for the sentence "an empty port is near some buildings and many green trees.", three noun phrases {"an empty port", "some buildings", "many green trees"} are extracted. Then, the three noun phrases are separately fed into the text encoder to obtain the global features of the three noun phrases, , i.e., text phrase, which represents the text phrase (the same hereinafter), represents the th global feature of the text phrase. The above process can be expressed by formula (3) as follows:
[0113] (3),
[0114] wherein, represents the noun phrase extraction tool in the Spacy library.
[0115] According to an embodiment of the present invention, based on the contrast learning method, the global feature alignment of the global features of the image samples and the global features of the text samples is performed by calculating the global similarity between the constructed positive sample pairs and negative sample pairs, and the global alignment loss value is calculated during the global feature alignment process, including: constructing positive sample pairs and negative sample pairs based on the pairing information between the remote sensing image samples and the text samples; calculating the global similarity of the positive sample pairs and the negative sample pairs respectively based on the contrast learning method to obtain the positive global similarity and the negative global similarity; performing global feature alignment on the global features of the image samples and the global features of the text samples by using the positive global similarity and the negative global similarity; calculating the global alignment loss value during the global feature alignment process by using a predefined global alignment loss function.
[0116] The above-mentioned embodiments involve the global feature alignment operation of the cross-modal retrieval model. The present invention aligns the global features of images and texts through contrastive learning. Contrastive learning methods have been widely used in tasks such as image recognition and natural language processing. Especially in cross-modal learning, they can effectively align feature representations of different modalities. The core idea is to bring similar samples (positive samples) closer and push dissimilar samples (negative samples) farther away. In this process, the similarity between positive samples and negative samples will be pulled apart to ensure that the model can learn more critical features. Specifically in the image-text retrieval task, paired image-text pairs are treated as positive samples, and unpaired images and texts are treated as negative samples. Use cosine similarity to measure the similarity between image and text feature vectors, record is the global similarity between the image and the text, For the Pictures and The global similarity between texts, Indicates Pictures and The global similarity between the texts, then the loss of global alignment As shown in formula (4):
[0117] (4),
[0118] in, Represents the global similarity between positive pairs of images and texts, is the batch size, A hyperparameter for controlling temperature in contrastive learning.
[0119] According to an embodiment of the present invention, the above-mentioned method of using the local similarity between the local features of the image samples and the global features of the phrase samples to perform local feature alignment on the local features of the image samples and the global features of the phrase samples, and calculating the local alignment loss value during the local feature alignment process includes: calculating the cosine similarity between the global features of the phrase samples and the local features of the image samples; averaging all cosine similarities above a preset similarity threshold to obtain the local similarity; using the local similarity to perform local feature alignment on the local features of the image samples and the global features of the phrase samples; and using a predefined local alignment loss function to calculate the local-to-partial alignment loss value during the local feature alignment process.
[0120] The above embodiments involve the local feature alignment operation of cross-modal retrieval model training. In addition to global alignment, the present invention also aligns fine-grained local features to supplement the details ignored in global alignment. The global features of the noun phrases are aligned with the local features of the image, that is, for each noun phrase, find the rectangular area of the image whose similarity is greater than a certain threshold for alignment. Specifically, For the picture The local similarity between a rectangular region and the -th noun phrase in the text, where i.e., local, means local (the same below). Then for the -th noun phrase, its similarity with the picture is the average of the similarities of the rectangular regions with higher similarities, which is expressed by formula (5) as follows:
[0121] (5),
[0122] where is a threshold hyperparameter, represents the set of local similarities with the -th noun phrase greater than , and represents an element in the set . Then, the average of the similarities between all the name phrases in a text and the picture is used as the local similarity between the text and the picture , as shown in formula (6):
[0123] (6),
[0124] where is the number of noun phrases in a text. Denote as the local similarity between the -th picture and the -th text, and as the local similarity between the -th picture and the -th text. Then the local alignment loss is as shown in formula (7):
[0125] (7),
[0126] where represents the local similarity between the picture and the positive text pair, is the batch size, and is the temperature hyperparameter.
[0127] According to an embodiment of the present invention, the above-mentioned masked modeling of the preprocessed remote sensing image samples and the preprocessed text samples is respectively performed, and the masked loss values of the remote sensing image samples and the text samples in the masked modeling process are respectively calculated, including: randomly masking the preprocessed remote sensing image samples based on a preset image masking ratio to obtain masked remote sensing image samples; extracting features of the masked remote sensing image samples by using the visual encoder of the cross-modal retrieval model to obtain a plurality of masked feature samples; processing all the masked feature samples by using the visual masked regression head of the cross-modal retrieval model to obtain the reconstructed remote sensing image samples; and processing the reconstructed remote sensing image samples and the remote sensing image samples by using a predefined image masking loss function to obtain the image masking loss value.
[0128] The above embodiment relates to the image masked modeling operation in the training process of the cross-modal retrieval model. In order to further capture the correlation of features within a single modality, the present invention introduces a masked modeling strategy. The idea of masked modeling is to randomly mask a part of the input, and then let the model predict the masked part through the unmasked part, so as to learn the correlation inside the input data.
[0129] In order to better understand the local information in the picture and capture the relationship between rectangular regions, the present invention adopts visual masked modeling. Given a picture , the present invention randomly masks 75% of the picture area to obtain the masked picture . is input into the visual encoder to obtain the features of each rectangular region. Then these features pass through a visual masked regression head for reconstructing the original picture, and the reconstructed value is denoted as . The least square error is used to optimize the masked part of the reconstructed picture and the original picture. MVM represents Masked Visual Modeling, represents the masked visual modeling loss, as shown in formula (8):
[0130] (8),
[0131] where, represents the set of subscripts masked in the picture rectangular region.
[0132] According to an embodiment of the present invention, the above-mentioned operations of performing masked modeling on the preprocessed remote sensing image samples and the preprocessed text samples respectively, and calculating the masked loss values of the remote sensing image samples and the text samples during the masked modeling process further include: randomly replacing some words in the text samples with masked specific words based on a preset image masking ratio to obtain masked text samples; using the text encoder of the cross-modal retrieval model to extract features from the masked text samples to obtain multiple word features; using the text masked classification head of the cross-modal retrieval model to predict each word feature to obtain the predicted text; and using a predefined text masked loss function to process the text samples and the predicted text to obtain the text masked loss value.
[0133] The above embodiment involves the text masked modeling operation in the training process of the cross-modal retrieval model. Text masked modeling is very common in language models and multi-modal models and is used to enhance the model's context understanding ability. Given a text , the present invention randomly replaces 15% of the words in it with the [MASK] special word to obtain the masked text . Then, is input into the text encoder to obtain the features of each word. Then, the features pass through a text masked classification head for predicting the probability of the word. Denote the probability as , where is the number of words in , and is the dictionary size of the words in CLIP. Use the cross-entropy loss
[0134] (9) to optimize the prediction probability, as shown in formula (9):
[0135] where, represents the set of subscripts of the masked words, represents the index of the -th word in the vocabulary. represents the probability that the -th word is predicted as .
[0136] In the training and inference stages of the cross-modal retrieval model, the present invention performs weighted averaging on the global alignment loss, the local alignment loss, and the masked modeling loss to train the model, as shown in formula (10):
[0137] (10),
[0138] where, is a hyperparameter used to balance the importance of different loss functions. Specifically in the experiment, the present invention sets .
[0139] In the inference stage, the present invention combines the global similarity and the local similarity to calculate the final retrieval similarity score , as shown in formula (11):
[0140] (11),
[0141] wherein, is the weighting coefficient, which controls the importance of the global similarity and the local similarity. By adjusting the value of , the contributions of the global information and the local information can be balanced, and the retrieval performance can be optimized. Usually, the value of can be determined through experiments to maximize the retrieval accuracy.
[0142] The training method of the above cross-modal retrieval model of the present invention captures the relationship between images and texts at a fine-grained level by extracting the conceptual phrases in the text and aligning them with the local region features in the images, complementing the neglect of fine-grained features in the global alignment. At the same time, a masked modeling strategy is introduced to enhance the correlation learning of unimodal features, further improving the matching ability of images and texts. Finally, by combining the global similarity and the local similarity scores, the accuracy of remote sensing image-text retrieval is significantly improved.
[0143] Figure 3 is a flowchart of the remote sensing image-text retrieval method according to an embodiment of the present invention.
[0144] As Figure 3 shown, the remote sensing image-text retrieval method of this embodiment includes operations S310 to S340.
[0145] In operation S310, the trained cross-modal retrieval model is used to preprocess the remote sensing image and the target text respectively, to obtain the preprocessed remote sensing image and the preprocessed target text, wherein the trained cross-modal retrieval model is trained based on the above training method of the cross-modal retrieval model for remote sensing image-text retrieval.
[0146] The present invention first constructs a cross-modal retrieval model based on a pre-trained neural network, trains the model, and uses the trained model to implement the retrieval between the remote sensing image and the target text.
[0147] Before retrieving the remote sensing image and the target text, it is necessary to preprocess the remote sensing image and the target text, so as to improve the matching accuracy and the retrieval efficiency.
[0148] In operation S320, the trained cross-modal retrieval model is used to extract features from the preprocessed remote sensing image, obtaining the global image features and local image features, and to extract features from the preprocessed target text, obtaining the global text features and global phrase features.
[0149] In the process of image feature extraction, a vision encoder based on VisionTransformer is utilized; in the process of target text feature extraction, a text encoder based on the Bert model is utilized.
[0150] In operation S330, the trained cross-modal retrieval model is used to calculate the global similarity between the global image features and the global text features, and to calculate the local similarity between the local image features and the global phrase features.
[0151] The local similarity can achieve fine-grained cross-modal retrieval and can complement the lack of matching of global similarity in detailed features. The present invention realizes fine-grained cross-modal matching, thus greatly improving the matching accuracy and retrieval accuracy.
[0152] In operation S340, the trained cross-modal retrieval model is used to perform weighted calculation on the global similarity and the local similarity to obtain the weighted similarity, and based on the weighted similarity, the matching result between the remote sensing image and the target text is screened to obtain the retrieval result between the remote sensing image and the target text.
[0153] In the process of calculating the weighted similarity, those skilled in the art can adjust the weighting coefficient according to actual needs, thereby balancing the contributions of global information and local information, and thus optimizing the retrieval performance. The weighting coefficient can also be determined through experiments.
[0154] The remote sensing image text retrieval method provided by the present invention realizes high-precision matching between the remote sensing image and the target text by using the trained cross-modal retrieval model; the remote sensing image text retrieval method provided by the present invention extracts the features of the remote sensing image and the target text, and locally aligns the features of the remote sensing image and the target text, capturing the relationship between the image and the text at the fine-grained level, and complementing the neglect of fine-grained features in the global alignment; at the same time, due to the combination of global similarity and local similarity, the cross-modal retrieval accuracy between the remote sensing image and the target text is significantly improved.
[0155] Figure 4 It is a structural block diagram of a remote sensing image text retrieval device according to an embodiment of the present invention.
[0156] As Figure 4As shown, the remote sensing image text retrieval device 400 of this embodiment includes a graphic and text preprocessing module 410, a feature extraction module 420, a similarity calculation module 430, and a weighting and screening module 440.
[0157] The graphic and text preprocessing module 410 is used to train a cross-modal retrieval model based on a multi-modal pre-trained neural network to obtain a trained cross-modal retrieval model, and use the trained cross-modal retrieval model to preprocess the remote sensing image and the target text respectively to obtain a preprocessed remote sensing image and a preprocessed target text, wherein the trained cross-modal retrieval model is trained based on the above training method for the cross-modal retrieval model for remote sensing image text retrieval; in one embodiment, the graphic and text preprocessing module 410 can be used to perform the operation S310 described above, which will not be elaborated here.
[0158] The feature extraction module 420 is used to extract features from the preprocessed remote sensing image by using the trained cross-modal retrieval model to obtain an image global feature and an image local feature, and extract features from the preprocessed target text to obtain a text global feature and a phrase global feature; in one embodiment, the feature extraction module 420 can be used to perform the operation S320 described above, which will not be elaborated here.
[0159] The similarity calculation module 430 is used to calculate the global similarity between the image global feature and the text global feature by using the trained cross-modal retrieval model, and calculate the local similarity between the image local feature and the phrase global feature; in one embodiment, the similarity calculation module 430 can be used to perform the operation S330 described above, which will not be elaborated here.
[0160] The weighting and screening module 440 is used to perform weighted calculation on the global similarity and the local similarity by using the trained cross-modal retrieval model to obtain a weighted similarity, and screen the matching result between the remote sensing image and the target text based on the weighted similarity to obtain the retrieval result between the remote sensing image and the target text; in one embodiment, the weighting and screening module 440 can be used to perform the operation S340 described above, which will not be elaborated here.
[0161] According to an embodiment of the present invention, any plurality of modules among the graphic and text preprocessing module 410, the feature extraction module 420, the similarity calculation module 430, and the weighting and screening module 440 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the graphic and text preprocessing module 410, the feature extraction module 420, the similarity calculation module 430, and the weighting and screening module 440 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the graphic and text preprocessing module 410, the feature extraction module 420, the similarity calculation module 430, and the weighting and screening module 440 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.
[0162] Figure 5 It is a block diagram of an electronic device suitable for implementing the training method of the cross-modal retrieval model and the remote sensing image text retrieval method according to an embodiment of the present invention.
[0163] As Figure 5 shown, the electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage part 508 into the random access memory (RAM) 503. The processor 501 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 501 may also include on-board memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0164] In the RAM 503, various programs and data required for the operation of the electronic device 500 are stored. The processor 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. The processor 501 performs various operations of the method flow according to an embodiment of the present invention by executing programs in the ROM 502 and / or the RAM 503. It should be noted that the programs may also be stored in one or more memories other than the ROM 502 and the RAM 503. The processor 501 may also perform various operations of the method flow according to an embodiment of the present invention by executing programs stored in the one or more memories.
[0165] According to an embodiment of the present invention, the electronic device 500 may further include an input / output (I / O) interface 505, and the input / output (I / O) interface 505 is also connected to the bus 504. The electronic device 500 may further include one or more of the following components connected to the input / output (I / O) interface 505: an input part 506 including a keyboard, a mouse, etc.; an output part 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part 508 including a hard disk, etc.; and a communication part 509 including a network interface card such as a LAN card, a modem, etc. The communication part 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 510 as needed so that a computer program read from it can be installed into the storage part 508 as needed.
[0166] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to an embodiment of the present invention is implemented.
[0167] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include one or more memories other than the ROM 502 and / or RAM 503 and / or ROM 502 and RAM 503 described above.
[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0169] Those skilled in the art can understand that the features described in various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0170] The above describes the embodiments of the present invention. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all these substitutions and modifications should fall within the scope of the present invention.
Claims
1. A method for training a cross-modal retrieval model, characterized in that: The method comprises: The visual encoder of the cross-modal retrieval model is used to extract features from the preprocessed remote sensing image samples to obtain the global features and local features of the image samples. Using the text encoder of the cross-modal retrieval model to extract features from the preprocessed text samples, and obtaining global features of the text samples and global features of the phrase samples; Based on the pairing information between the remote sensing image sample and the text sample, construct a positive sample pair and a negative sample pair; Based on the contrastive learning method, the global similarities of the positive sample pair and the negative sample pair are calculated respectively to obtain the positive global similarity and the negative global similarity; Performing global feature alignment on the image sample global feature and the text sample global feature using the positive global similarity and the negative global similarity; The global alignment loss value in the global feature alignment process is calculated using a predefined global alignment loss function; Performing local feature alignment on the local features of the image samples and the global features of the phrase samples by using the local similarity between the local features of the image samples and the global features of the phrase samples, and calculating a local alignment loss value; Performing mask modeling on the preprocessed remote sensing image samples and the preprocessed text samples respectively, and calculating mask loss values of the remote sensing image samples and the text samples in the mask modeling process respectively; The global alignment loss value, the local alignment loss value and the mask loss value are weightedly calculated to obtain a weighted loss value, and the weighted loss value is used to optimize the parameters of the cross-modal retrieval model.
2. The method according to claim 1, characterized in that Also includes: Iteratively performing a feature extraction operation, a weighted loss value calculation operation, and a model parameter optimization operation on the remote sensing image sample and the text sample until a preset training condition is met, thereby obtaining a trained cross-modal retrieval model; The visual encoder of the cross-modal retrieval model is used to extract features from the preprocessed remote sensing image samples to obtain global features and local features of the image samples, including: Based on the dimensional information of the remote sensing image, the remote sensing image sample is segmented to obtain a plurality of non-overlapping rectangular area image samples; Using the learnable linear mapping module of the cross-modal retrieval model, each rectangular area image sample is projected into a one-dimensional space to obtain a plurality of local features of the initial image samples; Adding a special code for aggregating global information to each of the local features of the initial image samples to obtain the local features of the initial image samples having the special code; Adding an image position code to each local feature of the initial image sample having the special code to obtain a plurality of encoded local features of the initial image sample; The visual encoder of the cross-modal retrieval model is used to process all the encoded local features of the initial image samples to obtain local features and global features of the image samples.
3. The method according to claim 1, characterized in that The text encoder of the cross-modal retrieval model is used to extract features from the preprocessed text samples, and the global features of the text samples and the global features of the phrase samples are obtained, including: Segmenting the text sample using a byte pair encoding method to obtain a plurality of phrases from the text sample, and inserting a start identification word and an end identification word at the start position and the end position of the text sample respectively; Mapping each of the phrases into a word vector using the encoding matrix of the cross-modal retrieval model, and adding a text position code to each of the word vectors; Using the text encoder of the cross-modal retrieval model to perform feature extraction on the start identification word, the end identification word, and all word vectors with text position encoding, to obtain the global features of the text sample; A predefined natural language processing tool is used to extract multiple noun phrases from the text sample to complete the enhancement operation of the text sample, and the text encoder is used to extract global features of each noun phrase to obtain the global features of the phrase sample.
4. The method according to claim 2 or 3, characterized in that: The visual encoder is constructed based on the VisionTransformer model; Wherein, the text encoder is constructed based on the Bert model; Wherein, the predefined natural language processing tool includes the Spacy library.
5. The method according to claim 1, characterized in that Performing local feature alignment on the local features of the image sample and the global features of the phrase sample by using the local similarity between the local features of the image sample and the global features of the phrase sample, and calculating the local alignment loss value comprises: Calculating the cosine similarity between the global features of the phrase sample and the local features of the image sample; Averaging all the cosine similarities above a preset similarity threshold to obtain the local similarity; Performing local feature alignment on the local features of the image sample and the global features of the phrase sample by using the local similarity; The predefined local alignment loss function is used to calculate the local-to-partial alignment loss value in the local feature alignment process.
6. The method according to claim 1, characterized in that Performing mask modeling on the preprocessed remote sensing image sample and the preprocessed text sample respectively, and calculating the mask loss values of the remote sensing image sample and the text sample in the mask modeling process respectively includes: Based on a preset image mask ratio, randomly masking the preprocessed remote sensing image samples to obtain masked remote sensing image samples; Using the visual encoder of the cross-modal retrieval model to extract features from the masked remote sensing image samples, to obtain a plurality of mask feature samples; Processing all the mask feature samples using the visual mask regression head of the cross-modal retrieval model to obtain reconstructed remote sensing image samples; The reconstructed remote sensing image sample and the remote sensing image sample are processed using a predefined image mask loss function to obtain an image mask loss value.
7. The method according to claim 6, characterized in that Also includes: Based on a preset image mask ratio, randomly replacing some words in the text sample with masked specific words to obtain a masked text sample; Using the text encoder of the cross-modal retrieval model to perform feature extraction on the masked text sample to obtain multiple word features; Using the text mask classification head of the cross-modal retrieval model to predict each of the word features, to obtain predicted text; The text sample and the predicted text are processed using a predefined text mask loss function to obtain a text mask loss value.
8. The method according to any one of claims 5 to 7, characterized in that: The predefined global alignment loss function and the predefined local alignment loss function are constructed based on the contrastive learning loss function; Wherein, the predefined image mask loss function is constructed based on a minimum square error function; Wherein, the predefined text mask loss function is constructed based on the cross entropy loss function.
9. A remote sensing image text retrieval method, characterized in that: The method comprises: Preprocessing the remote sensing image and the target text respectively using the trained cross-modal retrieval model to obtain the preprocessed remote sensing image and the preprocessed target text, wherein the trained cross-modal retrieval model is trained based on the training method described in any one of claims 1 to 8; Using the trained cross-modal retrieval model, feature extraction is performed on the pre-processed remote sensing image to obtain image global features and image local features, and feature extraction is performed on the pre-processed target text to obtain text global features and phrase global features; Calculating the global similarity between the image global features and the text global features, and calculating the local similarity between the image local features and the phrase global features using the trained cross-modal retrieval model; The global similarity and the local similarity are weightedly calculated using the trained cross-modal retrieval model to obtain a weighted similarity, and the matching results between the remote sensing image and the target text are screened based on the weighted similarity to obtain a retrieval result between the remote sensing image and the target text.
Citation Information
Patent Citations
Single stream multi-level alignment for vision-language pretraining
US20230281963A1
Model training method and apparatus, and device, storage medium and product
WO2024174583A1