Marine remote sensing image and text retrieval method and system based on adaptive view matching
By adopting adaptive perspective matching methods in marine remote sensing graphic and text retrieval, and using technologies such as compensation network and graph migration, the problem of discomfort qualitativeness is solved and the retrieval accuracy is improved.
Patent Information
- Application Number
- CN202411638733.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-18
AI Technical Summary
The prior art is difficult to perform effective information mining on unfair ocean remote sensing data, resulting in limited performance of ocean remote sensing graphic and text retrieval models.
Adaptive viewing angle matching is adopted to generate full-view text features through the compensation network, supervise image feature extraction, combine graph migration and cascading Transformer feature alignment to achieve robust matching across modal features.
The matching accuracy of marine remote sensing graphic and text search is improved, the problem of discomfort qualitativeness is solved, and the accuracy of the search results is enhanced.
Smart Images

Figure CN119149769B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image and text retrieval, and in particular relates to an ocean remote sensing image and text retrieval method and system based on adaptive view matching. Background Art
[0002] Remote sensing image and text retrieval uses cross-modal retrieval algorithms to automatically retrieve text data that accurately describes the image based on satellite remote sensing images, or to automatically retrieve remote sensing images in the database that match the given text data. Remote sensing image and text retrieval includes two key processes. First, feature engineering is performed on text data and image data to extract corresponding text features and image features; second, the similarity between the two features is calculated, and the image features and text features with the highest similarity are used as the best retrieval matching pair.
[0003] The main problem faced by traditional methods is the difficulty in extracting effective features. This is because there is a lot of redundant / background information in marine remote sensing data, and the spatial distribution of targets is relatively scattered. Effective targets will be disturbed by background noise, thus affecting the performance of the retrieval model. Therefore, the cutting-edge method is an asymmetric multimodal feature matching network. The advantage of this method is that after feature extraction is realized in text data and image data, different strategies are used to extract key information.
[0004] Cross-modal retrieval of marine remote sensing data aims to establish matching relationships between different modal data, improve the representation of marine remote sensing objects through multi-modal data fusion, and provide important technical support for its application. As one of the important means of cross-modal retrieval of marine remote sensing data, image-text retrieval uses comprehensive and rich image space and concise text information to more accurately and detailedly characterize marine objects. It has attracted the attention of more and more researchers in recent years. At present, cutting-edge image-text retrieval methods are committed to using saliency mining methods such as attention mechanisms and graph convolutional networks to filter out redundant background information in images and texts, and improve the matching accuracy of image-text retrieval. However, the above methods still have the following problems when applied to the ocean: There is a big difference between marine remote sensing images and ordinary images. In addition to low resolution, small targets and cross-scale problems, there is another problem that has been ignored by existing research, namely, the ill-posed problem. Ordinary images are generally shot from the main view, and are generally captured by optical cameras, close to the target, and have a focus point. Remote sensing images are usually presented in a bird's-eye view, usually received by remote sensing satellites and other capturers that are far away from the target, without a focus point, and do not meet the uniqueness principle in the well-posedness. Existing methods cannot carry out effective information mining on ill-posed ocean remote sensing data, which greatly limits the performance of remote sensing image and text retrieval models. Therefore, in response to the above problems, the present invention proposes an ocean remote sensing image and text retrieval method and system based on adaptive view matching. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides an ocean remote sensing image and text retrieval method and system based on adaptive view matching.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0007] Firstly, the present invention provides a marine remote sensing image and text retrieval method based on adaptive view matching, which includes two parts: feature extraction and feature processing. In the feature extraction part, image features X and text features F are extracted from input image data and text data respectively;
[0008] In the feature processing part, it includes:
[0009] Steps of full-view feature modeling based on compensation network: Generate full-view text feature F based on compensation network l , in the full-view text feature F l Under supervision, based on the image feature X, the neural network is trained so that the neural network extracts the comprehensive and complete perspective information in the image, and finally obtains the full-view image feature X l ;
[0010] Steps for discriminable view modeling based on graph transfer: In the full view image feature X l In , the perspective described by the text feature F to be queried is mined;
[0011] Steps based on cascaded Transformer feature alignment: In the full-view image feature X l Based on the text feature F, the Transformer decoder is used to align the sample information in two steps, namely text-guided image feature extraction and image-guided text feature extraction.
[0012] Furthermore, in the feature extraction part, the Transformer encoder is used for the image data to obtain the image feature X, and the GRU is used for the text data to extract the text feature F.
[0013] Furthermore, in the feature processing part, in the step of full-view feature modeling based on the compensation network, the full-view text feature F l The generation of is through the compensation network. On the one hand, the compensation network uses the text of multiple perspectives corresponding to an image, gathers the effective information in all texts with the help of the attention mechanism, and generates significant features; on the other hand, it uses the pre-trained semantic classifier , generate semantic features; finally, the saliency features and semantic features are cascaded to obtain the full-view text feature F l .
[0014] Furthermore, the full-view image feature X l The representation is as follows:
[0015] (1);
[0016] Conv 1×1 Represents a convolution operation with a convolution kernel of 1×1, and then X l Dependency triplet loss Supervised training; the loss function is expressed as follows:
[0017] (2);
[0018] Full-view text feature F l The generation steps are as follows: On the one hand, for the text T k First, GRU is used to generate text features F k ,
[0019] (3);
[0020] Where k represents the number of texts; then the attention mechanism is used Extract the salient information from each text and generate the salient feature Z k ,
[0021] (4);
[0022] in, Represents the parameters in the attention mechanism;
[0023] On the other hand, using pre-trained semantic classifiers , for text feature F k Perform semantic classification and generate semantic features M k ,
[0024] (5);
[0025] in, Represents the parameters in the semantic classifier;
[0026] Finally, the significant feature Z k With semantic feature M k Fusion, generate the final full-view text feature F l ,
[0027] (6);
[0028] in, Indicates a cascade operation.
[0029] Furthermore, the steps of discriminable view modeling based on graph transfer include two units: full view explicit representation based on graph transfer and discriminable view feature extraction based on distribution. The full view explicit representation unit based on graph transfer uses point construction and graph generation mechanism to extract full view image features X l The text feature F is converted into an explicit graph representation R, and the viewpoint described by the text feature F to be queried in the image is located by analyzing the connection density and quantity in the graph representation R;
[0030] According to the judged viewing angle, the distribution-based discriminable viewing angle feature extraction unit extracts the full viewing angle image feature X l The image is transmitted to the neural network that focuses on processing the view to extract the correct view information and generate the discriminable view image feature X. new ;
[0031] In addition, the text feature F is also processed by the attention mechanism to generate the effective text feature F new .
[0032] Furthermore, in the full-view explicit representation unit based on graph migration, the order of point construction is: text features, full-view image features, written as (F, X l ); The mechanism of graph generation is to use affinity matrix to calculate the correlation between features and generate graph representation R,
[0033] (7);
[0034] in, and is a trainable parameter, T is a transposition operation; the connection density and quantity in R are represented by the graph to locate the perspective described by the text feature F to be queried in the image.
[0035] In the distribution-based discriminable view feature extraction unit, a graph pattern selection mechanism is designed, which includes a graph pattern. The generated graph representation R and the graph representation G in the graph pattern are i Compare and locate the perspective described by the text feature F to be queried in the image;
[0036] Among them, the comparison adopts KL-divergence, which is expressed as follows:
[0037] (8);
[0038] Among them, i is the index of the graph representation in the graph mode, j is the dimension of the feature, R(j) is the feature corresponding to the j dimension in the graph representation R, G i (j) is a graph representing G i The feature corresponding to the j dimension, J is the maximum dimension of the feature; select the graph representation that maximizes the KL-divergence in the graph mode, and record the index.
[0039] (9);
[0040] Among them, o represents the index corresponding to the graph representation that maximizes the KL-divergence in the graph mode; according to the index, the corresponding network is selected to transform the full-view image feature X l The image is transmitted to the neural network that focuses on processing the view to extract the correct view information and generate the discriminable view image feature X. new ,
[0041] (10);
[0042] in Represents the parameters in the attention mechanism;
[0043] In turn, the graph representation in the graph model also relies on the newly generated discriminable view image features X new To update,
[0044] (11);
[0045] is the updated graph representation, where + represents element addition;
[0046] In addition, the text feature F is also processed by the attention mechanism to generate the effective text feature F new ,
[0047] (12);
[0048] in Represents the parameters in the attention mechanism.
[0049] Furthermore, in the step of cascaded Transformer feature alignment, the text-guided image feature extraction is specifically as follows: the full-view image features are regarded as the query sentence Q, the text features are regarded as K and V, and the query sentence samples are subjected to re-mining and information alignment of the effective image information to generate a robust image feature X upd ;
[0050] The specific method of image-guided text feature extraction is: consider the text feature as Q, the robust image feature as K and V, use the robust image feature to filter the noise of the text feature, and generate the robust text feature F upd .
[0051] Furthermore, the marine remote sensing image and text retrieval method based on adaptive view matching also includes a similarity matching step, and the loss calculation includes four parts, namely: the triple loss of the original image feature X and the text feature F , full view image feature Xl and the full-view text feature F l The triplet loss , discriminable view image feature X new and effective text features F new The triplet loss And the robust image feature X upd and robust text features F upd The triplet loss .
[0052] Secondly, the present invention provides an ocean remote sensing image and text retrieval system based on adaptive perspective matching, which is used to implement the ocean remote sensing image and text retrieval method based on adaptive perspective matching as described above, including an image feature extraction module, a text feature extraction module, a full-view feature modeling module based on a compensation network, a discriminable perspective modeling module based on graph migration, a cascaded Transformer feature alignment module, and a loss calculation module.
[0053] The image feature extraction module is used to extract image features X from input image data;
[0054] The text feature extraction module is used to extract text features F from input text data;
[0055] The full-view feature modeling module based on the compensation network generates a full-view text feature F based on the compensation network. l , in the full-view text feature F l Under supervision, based on the image feature X, the neural network is trained so that the neural network extracts the comprehensive and complete perspective information in the image, and finally obtains the full-view image feature X l ;
[0056] The discriminable view modeling module based on graph migration is used to l In the above method, the perspective described by the text feature F to be queried is mined, which specifically includes two parts: extraction of discriminable perspective features based on distribution and full perspective display representation based on graph migration;
[0057] The cascaded Transformer-based feature alignment module is used to align the full-view image features X l Based on the text feature F, the Transformer decoder is used to align the sample information in two steps, namely, text-guided image feature extraction and image-guided text feature extraction;
[0058] The loss calculation module is used to calculate the triplet loss.
[0059] Compared with the prior art, the present invention has the following advantages:
[0060] (1) In order to extract robust information from images and solve the problem of ill-posedness, this method designs a full-view feature modeling module based on a compensation network. This module uses full-view text features to supervise the feature extraction process of the image. The generation of full-view text features combines multiple text descriptions and semantic classification information. Since multiple text descriptions can comprehensively and reliably describe images from multiple perspectives, and attention and semantic classification can mine effective information in the text, the generated full-view text features are robust. Under the supervision of this feature, the deep neural network is trained to mine the robust areas in the image and filter out invalid noise to generate full-view image features. This is the first stage of image feature extraction, which is equivalent to extracting robust information from the image without considering the query text, laying the foundation for the subsequent targeted modeling of feature representation.
[0061] (2) Based on the generation of full-view image features, the discriminable view modeling module based on graph transfer can realize adaptive positioning of the view represented by the query text in the image. The full-view explicit representation unit based on graph transfer uses point construction and graph generation mechanisms to convert full-view image features and text features into explicit graph representations. By analyzing the density and number of connections in the graph representation, the view described by the query text features in the image can be located. For example, sparse graphs represent local view, dense graphs represent global view, etc. According to the judged view, the distribution-based discriminable view feature extraction unit transfers the full-view image features to the network dedicated to processing the view, and mines the correct view information. By converting image and text features into graph representations, the view expressed by the text in the image is displayed, which guides the subsequent image transmission to the network dedicated to processing the view, thereby accurately locating and extracting effective view information in the image, solving the ill-posed problem and improving the accuracy of the retrieval results.
[0062] (3) Based on the full-view image features and text features, the Transformer decoder is used to align sample information in two steps based on the cascaded Transformer feature alignment module. The text-guided image feature extraction unit regards the full-view image features as the query sentence Q and the text features as K and V, and re-mines and aligns the effective image information of the query sentence samples. The image-guided text feature extraction unit regards the text features as Q and the robust image features as K and V, and uses the robust image features to filter the noise of the text features, ultimately obtaining a more robust cross-modal feature representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0064] Figure 1 It is a schematic diagram of the system structure of the present invention;
[0065] Figure 2 A schematic diagram of a full-view feature modeling module based on a compensation network of the present invention;
[0066] Figure 3 It is a schematic diagram of a discriminable view modeling module based on graph migration of the present invention;
[0067] Figure 4 It is a schematic diagram of the feature alignment module based on cascaded Transformer of the present invention. DETAILED DESCRIPTION
[0068] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0069] Example 1
[0070] like Figure 1 As shown, this embodiment designs a marine remote sensing image and text retrieval system based on adaptive view matching, following the main architecture of AMFMN. Unlike AMFMN, in the feature extraction part, the Transformer encoder is used to obtain the image feature X for the image data, and the GRU is still used to extract the text feature F for the text data. Afterwards, in the feature processing part, the present invention does not use the visual self-attention module and the visual guided text attention module to globally mine the significant information in the data, but innovatively designs a cross-modal adaptive view matching strategy, including (1) a full view feature modeling module based on a compensation network (2) a discriminable view modeling module based on graph migration (3) a feature alignment module based on a cascaded Transformer.
[0071] Among them, (1) the full-view feature modeling module based on the compensation network designs a compensation network with the purpose of exploring comprehensive and complete view information in ill-posed remote sensing images. The idea is: l Under supervision, based on the image feature X, the neural network is trained so that the neural network extracts comprehensive and complete perspective information from the image, and finally obtains the full-perspective image feature representation X l . Full-view text feature F lThe generation relies on the compensation network. On the one hand, the compensation network makes full use of multiple texts corresponding to an image, and uses the attention mechanism to gather effective information in all texts to improve the integrity and comprehensiveness of the signal. On the other hand, the pre-trained classifier is used to generate semantic features, which fully guarantees the reliability of the signal. Through multiple text descriptions in the image, the model can mine supervisory features from multiple perspectives, and finally achieves the preliminary filtering of invalid information. (The full-view feature modeling module based on the compensation network is the first innovation of the present invention).
[0072] (2) The purpose of the discriminable view modeling module based on graph transfer is to l In the above method, the perspective described by the text feature F to be queried is mined, which includes two units: full-perspective explicit representation based on graph migration and discriminable perspective feature extraction based on distribution. The former uses point construction and graph generation mechanism to transform full-perspective image features and text features into explicit graph representation R. By analyzing the connection density and number in the graph representation, the perspective described by the text feature F to be queried can be located. For example, sparse graphs represent local perspectives, dense graphs represent global perspectives, etc. According to the judged perspective, the latter transforms the full-perspective image feature X into a discriminable perspective feature. l Transmitted to the network that focuses on processing this perspective, mining the correct perspective information, and generating discriminable perspective image features X new In addition, in this module, the text feature F is also processed by the attention mechanism to generate the effective text feature F new By converting image and text features into graph representations, the perspective expressed by the text in the image is displayed, which guides the subsequent image transmission to the network dedicated to processing the perspective, thereby accurately locating and extracting effective perspective information in the image, solving the ill-posed problem and improving the accuracy of the retrieval results. (The discriminable perspective modeling module based on graph migration is the second innovation of the present invention).
[0073] (3) Based on the cascaded Transformer feature alignment module, the full-view image feature X l Based on the text feature F, the Transformer decoder is used in two steps to align sample information. Two units, text-guided image feature extraction and image-guided text feature extraction, are proposed. The former regards the full-view image feature as the query sentence Q and the text feature as K and V. It re-mines and aligns the effective image information of the query sentence sample to generate a robust image feature X. upd The latter regards text features as Q, robust image features as K and V, and uses robust image features to filter the noise of text features to generate robust text features F upd . (The feature alignment module based on the cascaded Transformer is the third innovation of the present invention).
[0074] In the similarity matching part, in addition to calculating the triple loss between the original image feature X and the text feature F, the feature X l and F l , X new and F new , X upd and F upd The triplet loss between .
[0075] like Figure 1 As shown, the system includes an image feature extraction module, a text feature extraction module, a full-view feature modeling module based on a compensation network, a discriminable view modeling module based on graph migration, a cascaded Transformer feature alignment module, and a loss calculation module.
[0076] In the feature extraction part, the image feature extraction module is used to extract image features X from the input image data. The text feature extraction module is used to extract text features F from the input text data.
[0077] In the feature processing part, such as Figure 2 As shown in the figure, the full-view feature modeling module based on the compensation network is designed to explore the comprehensive and complete view information in the ill-posed remote sensing image. The idea is to generate the full-view text feature F based on the compensation network. l , in the full-view text feature F l Under supervision, based on the image feature X, the neural network is trained so that the neural network extracts comprehensive and complete perspective information from the image, and finally obtains the full-perspective image feature representation X l The function and implementation of this module can be found in the description of the steps of full-view feature modeling based on the compensation network in Example 2, and will not be repeated here.
[0078] like Figure 3 As shown, the discriminable view modeling module based on graph migration is used to l In the example, the perspective described by the text feature F to be queried is mined, which specifically includes two parts: discriminable perspective feature extraction based on distribution and full perspective display representation based on graph migration. The function and implementation of this module can be found in the steps of discriminable perspective modeling based on graph migration in Example 2, which will not be repeated here.
[0079] like Figure 4 As shown, the cascaded Transformer feature alignment module is used to align the full-view image features X lBased on the text feature F, the Transformer decoder is used to align the sample information in two steps, namely, text-guided image feature extraction and image-guided text feature extraction. The function and implementation of this module can be found in the description of the steps of cascaded Transformer feature alignment in Example 2, which will not be repeated here.
[0080] Finally, the loss calculation module is used to calculate the triplet loss. The specific loss calculation can be found in the loss calculation section of Example 2, which will not be repeated here.
[0081] Example 2
[0082] This embodiment provides a method for marine remote sensing image and text retrieval based on adaptive view matching, which is implemented by the marine remote sensing image and text retrieval system based on adaptive view matching described in Example 1. The method includes two parts: feature extraction and feature processing.
[0083] In the feature extraction part, image feature X and text feature F are extracted from the input image data and text data respectively; specifically, the image feature X is obtained by using the Transformer encoder for the image data, and the text feature F is extracted by using the GRU for the text data.
[0084] The feature processing part includes three steps: full-view feature modeling based on compensation network, discriminable view modeling based on graph migration, and feature alignment based on cascaded Transformer. Each step is introduced in detail below.
[0085] 1. Steps of full-view feature modeling based on compensation network:
[0086] like Figure 2 As shown, Figure 2 The specific examples of multi-view text in the , such as "a ship is docked at the port", are only examples and have no specific meaning. Generating full-view text features F based on compensation network l , in the full-view text feature F l Under supervision, based on the image feature X, the neural network is trained so that the neural network extracts comprehensive and complete perspective information from the image, and finally obtains the full-perspective image feature representation X l .
[0087] In the step of full-view feature modeling based on the compensation network, the full-view text feature F l The generation of is based on the compensation network. On the one hand, the compensation network uses the texts of multiple perspectives corresponding to an image, gathers the effective information in all the texts with the help of the attention mechanism, and generates significant features to improve the integrity and comprehensiveness of the signal; on the other hand, it uses the pre-trained semantic classifier , generate semantic features, and fully guarantee the reliability of the signal; finally, the saliency features and semantic features are cascaded to obtain the full-view text feature F l .
[0088] Full view image features X l The representation is as follows:
[0089] (1);
[0090] Conv 1×1 Represents a convolution operation with a convolution kernel of 1×1, and then X l Dependency triplet loss Supervised training; the loss function is expressed as follows:
[0091] (2);
[0092] Full-view text feature F l The generation steps are as follows: On the one hand, for the text T k First, GRU is used to generate text features F k ,
[0093] (3);
[0094] Where k represents the number of texts; then the attention mechanism is used Extract the salient information from each text and generate the salient feature Z k ,
[0095] (4);
[0096] in, Represents the parameters in the attention mechanism.
[0097] On the other hand, using pre-trained semantic classifiers , for text feature F k Perform semantic classification and generate semantic features M k ,
[0098] (5);
[0099] in, Represents the parameters in the semantic classifier;
[0100] Finally, the significant feature Z k With semantic feature M k Fusion, generate the final full-view text feature F l ,
[0101] (6);
[0102] in, Indicates a cascade operation.
[0103] 2. Steps for discriminable perspective modeling based on graph migration:
[0104] like Figure 3 As shown, in the full-view image feature X l In , the perspective described by the text feature F to be queried is mined.
[0105] Among them, the step of discriminable perspective modeling based on graph migration includes two units: full perspective explicit representation based on graph migration and discriminable perspective feature extraction based on distribution.
[0106] 1. The full-view explicit representation unit based on graph migration uses point construction and graph generation mechanism to transform the full-view image feature X l The text feature F is converted into an explicit graph representation R, and the perspective described by the text feature F to be queried is located by analyzing the connection density and number in the graph representation. For example, a sparse graph represents a local perspective, a dense graph represents a global perspective, and so on.
[0107] According to the judged viewing angle, the distribution-based discriminable viewing angle feature extraction unit extracts the full viewing angle image feature X l Transmitted to the network that focuses on processing this perspective, mining the correct perspective information, and generating discriminable perspective image features X new .
[0108] In addition, in this step, the text feature F also passes through the attention mechanism to generate the effective text feature F new By converting image and text features into graph representations, the perspective expressed by the text in the image is displayed, which guides the subsequent transmission of images to the network that focuses on processing the perspective, thereby accurately locating and extracting effective perspective information in the image, solving the ill-posed problem and improving the accuracy of retrieval results.
[0109] In the full-view explicit representation unit based on graph migration, the order of point construction is: text features, full-view image features, written as (F, X l ); The mechanism of graph generation is to use affinity matrix to calculate the correlation between features and generate graph representation R,
[0110] (7);
[0111] in, and is a trainable parameter, T is a transposition operation; the connection density and quantity in R are represented by the graph to locate the perspective described by the text feature F to be queried in the image.
[0112] 2. In the distribution-based discriminable view feature extraction unit, a graph pattern selection mechanism is designed, which includes a graph pattern. The generated graph representation R and the graph representation G in the graph pattern are selected by i Compare and locate the perspective described by the text feature F to be queried in the image.
[0113] The comparison uses KL-divergence, which is expressed as follows:
[0114] (8);
[0115] Among them, i is the index of the graph representation in the graph mode, j is the dimension of the feature, R(j) is the feature corresponding to the j dimension in the graph representation R, G i (j) is a graph representing G i The feature corresponding to the j dimension, J is the maximum dimension of the feature; select the graph representation that maximizes the KL-divergence in the graph mode, and record the index.
[0116] (9);
[0117] Where o represents the index corresponding to the graph representation that maximizes the KL-divergence in the graph mode. According to the index, the corresponding network is selected to transform the full-view image feature X l The image is transmitted to the neural network that focuses on processing the view to extract the correct view information and generate the discriminable view image feature X. new ,
[0118] (10);
[0119] in Represents the parameters in the attention mechanism. For networks that process different perspectives, In Varies, for example, Dedicated to handling features from a global perspective.
[0120] In turn, the graph representation in the graph model also relies on the newly generated discriminable view image features X new To update,
[0121] (11);
[0122] is the updated graph representation, + represents element addition. In addition, in this module, the text feature F is also processed by the attention mechanism to generate the effective text feature F new ,
[0123] (12);
[0124] in Represents the parameters in the attention mechanism.
[0125] 3. Steps of feature alignment based on cascaded Transformer:
[0126] like Figure 4 As shown, in the full-view image feature X l Based on the text feature F, the Transformer decoder is used to align the sample information in two steps, namely text-guided image feature extraction and image-guided text feature extraction.
[0127] In the step of cascaded Transformer feature alignment, the text-guided image feature extraction is as follows: the full-view image features are regarded as query sentences Q, the text features are regarded as K and V, the effective image information of the query sentence samples is re-mined and aligned, and the robust image features X are generated. upd The Transformer decoder consists of two steps, namely multi-head attention MultiHead and feed-forward neural network FFN.
[0128] (13);
[0129] (14);
[0130] in, represents the output of the middle layer of the Transformer decoder, t and t-1 are the t-th stack and t-1-th stack in the Transformer decoder. The output of the last stack is used as the robust image feature X upd .
[0131] The specific method of image-guided text feature extraction is: consider the text feature as Q, the robust image feature as K and V, use the robust image feature to filter the noise of the text feature, and generate the robust text feature F upd .
[0132] (15);
[0133] (16);
[0134] in, Represents the output of the middle layer of the Transformer decoder, and the output of the last stack is used as the robust text feature F upd .
[0135] In addition, the present invention also includes a similarity matching step, in addition to calculating the triplet loss between the original image feature X and the text feature F, it also calculates the feature Xl and F l , X new and F new , X upd and F upd The triplet loss between , specifically:
[0136] Use AMFMN to calculate triplet loss,
[0137] (17)
[0138] in, Represents the similarity between the image feature X and the corresponding text feature F, Represents image feature X and corresponding negative text feature The similarity of Represents the text feature F and the corresponding negative image feature The similarity.
[0139] The loss calculation consists of four parts, namely: the triple loss of the original image feature X and the text feature F , full view image feature X l and the full-view text feature F l The triplet loss , discriminable view image feature X new and effective text features F new The triplet loss And the robust image feature X upd and robust text features F upd The triplet loss .
[0140] Total loss (18).
[0141] In summary, the present invention 1) proposes a full-view feature modeling module based on a compensation network, the purpose of which is to explore comprehensive and complete view information in ill-posed remote sensing images. l Under supervision, based on the image feature X, the neural network is trained so that the neural network extracts the comprehensive and complete perspective information in the image, and finally obtains the full-view image feature X l . Full-view text feature F l The generation of is based on the compensation network. On the one hand, the compensation network makes full use of multiple texts corresponding to an image and uses the attention mechanism to gather effective information from all texts to improve the integrity and comprehensiveness of the signal. On the other hand, the pre-trained classifier is used to generate semantic features, which fully guarantees the reliability of the signal.
[0142] 2) A discriminable view modeling module based on graph transfer is proposed, the purpose is to l In the method, the perspective described by the text feature F to be queried in the image is mined. It includes two units: full-view explicit representation based on graph migration and discriminable perspective feature extraction based on distribution. The former uses point construction and graph generation mechanism to transform full-view image features and text features into explicit graph representation R. By analyzing the connection density and number in the graph representation, the perspective described by the text feature F to be queried in the image can be located. For example, sparse graphs represent local perspectives, dense graphs represent global perspectives, etc. According to the judged perspective, the latter converts the full-view image feature X into a discriminable perspective feature. l Transmitted to the network that focuses on processing this perspective, mining the correct perspective information, and generating discriminable perspective image features X new In addition, in this module, the text feature F is also processed by the attention mechanism to generate the effective text feature F new .
[0143] 3) A cascaded Transformer feature alignment module is proposed to achieve the re-mining and information alignment of effective information in cross-modal features. l and text features Based on the above, the Transformer decoder is used in two steps to realize the alignment of sample information, including two units: text-guided image feature extraction and image-guided text feature extraction. The former regards the full-view image features as the query sentence Q and the text features as K and V, and re-mines and aligns the effective image information of the query sentence sample to generate a robust image feature X. upd The latter regards text features as Q, robust image features as K and V, and uses robust image features to filter the noise of text features to generate robust text features F upd .
[0144] The present invention improves the accuracy of image and text retrieval through innovative improvements in three aspects.
[0145] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Any changes, modifications, additions or substitutions made by ordinary technicians in this technical field within the essential scope of the present invention should fall within the protection scope of the present invention.
Claims
1. A marine remote sensing image and text retrieval method based on adaptive view matching includes two parts: feature extraction and feature processing, and also includes a similarity matching step, characterized in that: In the feature extraction part, it includes extracting image feature X and text feature F from the input image data and text data respectively; In the feature processing part, it includes: Steps of full-view feature modeling based on compensation network: Generate full-view text feature F based on compensation network l , in the full-view text feature F l Under supervision, based on the image feature X, the neural network is trained so that the neural network extracts the comprehensive and complete perspective information in the image, and finally obtains the full-view image feature X l ; In the step of full-view feature modeling based on the compensation network, the full-view text feature F l The generation of is through the compensation network. On the one hand, the compensation network uses the text of multiple perspectives corresponding to an image, gathers the effective information in all texts with the help of the attention mechanism, and generates significant features; on the other hand, it uses the pre-trained semantic classifier Generate semantic features; finally, the saliency features and semantic features are cascaded to obtain the full-view text feature F l ; Steps for discriminable view modeling based on graph transfer: In the full view image feature X l In , the perspective described by the text feature F to be queried is mined; Steps based on cascaded Transformer feature alignment: In the full-view image feature X l Based on the text feature F, the Transformer decoder is used to align the sample information in two steps, namely text-guided image feature extraction and image-guided text feature extraction.
2. The method for marine remote sensing image and text retrieval based on adaptive perspective matching according to claim 1, characterized in that: In the feature extraction part, the Transformer encoder is used for image data to obtain image features X, and the GRU is used for text data to extract text features F.
3. The method for marine remote sensing image and text retrieval based on adaptive perspective matching according to claim 1, characterized in that: Full view image features X l The representation is as follows: X l =Conv 1×1 (X) (1); Conv 1×1 Represents a convolution operation with a convolution kernel of 1×1, after which Xl depends on the triplet loss Supervised training; the loss function is expressed as follows: Full-view text feature F l The generation steps are as follows: On the one hand, for the text T k First, GRU is used to generate text features F k , F k =GRU(T k ) (3); Where k represents the number of texts; then the attention mechanism is used Extract the salient information from each text and generate the salient feature Z k , Among them, ω a Represents the parameters in the attention mechanism; On the other hand, using pre-trained semantic classifiers For text feature F k Perform semantic classification and generate semantic features M k , Among them, ω g Represents the parameters in the semantic classifier; Finally, the significant feature Z k With semantic feature M k Fusion, generate the final full-view text feature F l , Among them, [,] represents a cascade operation.
4. The method for marine remote sensing image and text retrieval based on adaptive perspective matching according to claim 1, characterized in that: The steps of discriminable view modeling based on graph migration include two units: full view explicit representation based on graph migration and discriminable view feature extraction based on distribution. The full view explicit representation unit based on graph migration is to use point construction and graph generation mechanism to extract the full view image feature X l The text feature F is converted into an explicit graph representation R, and the viewpoint described by the text feature F to be queried in the image is located by analyzing the connection density and quantity in the graph representation R; According to the judged viewing angle, the distribution-based discriminable viewing angle feature extraction unit extracts the full viewing angle image feature X l The image is transmitted to the neural network that focuses on processing the view to extract the correct view information and generate the discriminable view image feature X. new ; In addition, the text feature F is also processed by the attention mechanism to generate the effective text feature F new .
5. The method for marine remote sensing image and text retrieval based on adaptive view matching according to claim 4, characterized in that: In the full-view explicit representation unit based on graph migration, the order of point construction is: text features, full-view image features, written as (F, X l ); The mechanism of graph generation is to use affinity matrix to calculate the correlation between features and generate graph representation R, in, and φ are trainable parameters, and T is the transposition operation; the connection density and number in R are represented by the graph to locate the perspective described by the text feature F to be queried in the image; In the distribution-based discriminable view feature extraction unit, a graph pattern selection mechanism is designed, which includes a graph pattern. The generated graph representation R and the graph representation G in the graph pattern are i Compare and locate the perspective described by the text feature F to be queried in the image; Among them, the comparison adopts KL-divergence, which is expressed as follows: Among them, i is the index of the graph representation in the graph mode, j is the dimension of the feature, R(j) is the feature corresponding to the j dimension in the graph representation R, G i (j) is a graph representing G i The feature corresponding to the j dimension, J is the maximum dimension of the feature; select the graph representation that maximizes the KL-divergence in the graph mode, and record the index. Among them, o represents the index corresponding to the graph representation that maximizes the KL-divergence in the graph mode; according to the index, the corresponding network is selected to transform the full-view image feature X l The image is transmitted to the neural network that focuses on processing the view to extract the correct view information and generate the discriminable view image feature X. new , where ω o Represents the parameters in the attention mechanism; In turn, the graph representation in the graph model also relies on the newly generated discriminable view image features X new To update, is the updated graph representation, where + represents element addition; In addition, the text feature F is also processed by the attention mechanism to generate the effective text feature F new , Among them, ω f Represents the parameters in the attention mechanism.
6. The method for marine remote sensing image and text retrieval based on adaptive view matching according to claim 4, characterized in that: In the step of cascaded Transformer feature alignment, the text-guided image feature extraction is as follows: the full-view image features are regarded as query sentences Q, the text features are regarded as K and V, the effective image information of the query sentence samples is re-mined and aligned, and the robust image features X are generated. upd ; The specific method of image-guided text feature extraction is: consider the text feature as Q, the robust image feature as K and V, use the robust image feature to filter the noise of the text feature, and generate the robust text feature F upd .
7. The method for marine remote sensing image and text retrieval based on adaptive view matching according to claim 6, characterized in that: The loss calculation consists of four parts, namely: the triple loss of the original image feature X and the text feature F Full view image features X l and the full-view text feature F l The triplet loss Discriminable view image feature X new and effective text features F new The triplet loss And the robust image feature X upd and robust text features F upd The triplet loss 8. The marine remote sensing image and text retrieval system based on adaptive perspective matching is characterized by: The method for realizing the marine remote sensing image and text retrieval method based on adaptive view matching as described in any one of claims 1 to 7 comprises an image feature extraction module, a text feature extraction module, a full view feature modeling module based on a compensation network, a discriminable view modeling module based on graph migration, a cascaded Transformer feature alignment module, and a loss calculation module. The image feature extraction module is used to extract image features X from input image data; The text feature extraction module is used to extract text features F from input text data; The full-view feature modeling module based on the compensation network generates a full-view text feature F based on the compensation network. l , in the full-view text feature F l Under supervision, based on the image feature X, the neural network is trained so that the neural network extracts comprehensive and complete perspective information from the image, and finally obtains the full-perspective image feature representation X l ; The discriminable view modeling module based on graph migration is used to l In the above method, the perspective described by the text feature F to be queried is mined, which specifically includes two parts: extraction of discriminable perspective features based on distribution and full perspective display representation based on graph migration; The cascaded Transformer-based feature alignment module is used to align the full-view image features X l Based on the text feature F, the Transformer decoder is used to align the sample information in two steps, namely, text-guided image feature extraction and image-guided text feature extraction; The loss calculation module is used to calculate the triplet loss.
Citation Information
Patent Citations
Marine remote sensing image-text retrieval method and device, electronic equipment and storage medium
CN117407558A
Semantic knowledge guided vehicle re-identification method
CN118230321A