A remote sensing cross-modal retrieval method based on a deep ternary fusion perception network

By using a deep ternary fusion sensing network, the problem of multimodal representation in remote sensing cross-modal retrieval is solved. It realizes unified feature representation of image, text and audio data, improves the utilization rate and retrieval accuracy of remote sensing data, overcomes the problem of scarce labeled data, and enhances the adaptability and robustness of the model.

CN119336968BActive Publication Date: 2025-11-28CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411371357.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-11-28
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

In remote sensing cross-modal retrieval, there is a challenge in constructing a unified representation space across multiple modalities, and high-quality labeled data is scarce, which affects the learning effect and retrieval performance of the model.

Method used

A deep ternary fusion sensing network is adopted, which combines remote sensing images, text and audio data through a ternary feature representation module, a fusion sensing module and a self-supervised modality enhancement module. The model parameters are optimized by self-supervised learning and feature alignment loss function to generate a unified multimodal feature representation.

Benefits of technology

It improves the accuracy and efficiency of cross-modal retrieval of remote sensing data, enhances the robustness and scalability of the model, reduces the dependence on high-quality labeled data, and improves the performance of the retrieval model in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119336968B_ABST
    Figure CN119336968B_ABST
Patent Text Reader

Abstract

The application discloses a remote sensing cross-modal retrieval method based on a deep ternary fusion perception network, which uses a trained remote sensing cross-modal retrieval model to perform cross-modal retrieval on multi-modal remote sensing data; the remote sensing cross-modal retrieval model takes the multi-modal remote sensing data as the input of a self-supervised modal enhancement module for preprocessing, and inputs the multi-modal remote sensing data into a ternary feature expression module to independently capture and extract single-modal image features, text features and audio features; the features are taken as the input of a fusion perception module to generate multi-modal feature embeddings by fusion, and remote sensing cross-modal retrieval results of the multi-modal remote sensing data to be processed are obtained according to the multi-modal feature embeddings. The application uses a ternary feature expression strategy, a fusion perception mechanism and a self-supervised modal enhancement technology to solve key problems such as model modal scalability and the scarcity and high cost of remote sensing labeled data, and significantly enhances the precision and efficiency of the remote sensing data cross-modal retrieval task.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing image processing and modal fusion, and particularly relates to a remote sensing cross-modal retrieval method based on a deep ternary fusion perception network. BACKGROUND

[0002] Remote sensing image retrieval, as the basis for important decision-making and data utilization in various industries, has received extensive attention and research in recent years. However, with the advent of the era of remote sensing big data, different types of data sources are becoming increasingly diverse, and single-modal retrieval has been unable to meet the needs of practical applications. How to break through the limitations of information, overcome the neglect of context and semantic association, and handle the differences between sensors have become the focus and hotspot of research in the field of remote sensing cross-modal retrieval.

[0003] Current research on cross-modal retrieval in the field of remote sensing mainly focuses on image-text retrieval tasks. However, due to the limitations of manual input in practical applications, text retrieval is inefficient in some emergency scenarios. In order to improve the comprehensive utilization of remote sensing data and enhance the modal scalability of remote sensing image-text cross-modal retrieval, some research has integrated audio data into image-text retrieval, promoting the extraction of scalable image information from heterogeneous generalized remote sensing data. With the increasing number of modalities in cross-modal retrieval, the field of remote sensing image cross-modal retrieval faces more and more problems and challenges, including feature fusion and representation learning, domain adaptation and transfer learning, semantic consistency and context understanding, system construction and optimization, and application field expansion. In terms of feature fusion and representation learning, it is necessary to effectively fuse the features of remote sensing images obtained by different sensors and learn more representative and robust representations. In terms of domain adaptation and transfer learning, it is necessary to address the differences in data distribution between different domains or sensors to achieve good generalization ability of the model. In addition, achieving semantic consistency matching between cross-modal images and accurately understanding and analyzing images in different contexts are also challenges. In terms of system construction and optimization, designing an efficient cross-modal retrieval system, including effective feature extraction, similarity measurement, and retrieval algorithm, is one of the current research focuses.

[0004] Although some research achievements have been made in the fields of remote sensing cross-modal image-text retrieval and remote sensing cross-modal audio-image retrieval, there are still many problems and challenges to be solved in this field. When introducing the audio modality into remote sensing cross-modal image-text retrieval, the difficulty of constructing a unified representation space among multiple modalities is faced. Due to the natural semantic differences between different modalities and the limited paired data, the model is difficult to capture rich modal features, thereby affecting the learning effect of modal representation. In addition, due to factors such as the coverage range and acquisition cost of remote sensing data, high-quality labeled data are scarce, which is particularly prominent in multi-modal retrieval, because a small amount of labeled data is difficult to support effective spatial representation learning.

[0005] Therefore, how to optimize the sparse uniform space representation of multi-modal data under the condition of limited labeled data to improve the retrieval performance is a key challenge in the current remote sensing cross-modal retrieval field. SUMMARY

[0006] In view of the above problems of the prior art, the present application provides a remote sensing cross-modal retrieval method based on a deep ternary fusion perception network, which uses a ternary feature expression strategy, a fusion perception mechanism and a self-supervised modal enhancement technology to solve the key problems of model modal scalability and the scarcity and high cost of remote sensing labeled data, and significantly improves the accuracy and efficiency of the remote sensing data cross-modal retrieval task.

[0007] To solve the above technical problems, the present application adopts the following technical solutions:

[0008] A remote sensing cross-modal retrieval method based on a deep ternary fusion perception network comprises the following steps:

[0009] S1, obtaining multi-modal remote sensing data to be processed, wherein the multi-modal remote sensing data comprises remote sensing images, text data and audio data;

[0010] S2, inputting the multi-modal remote sensing data into a trained remote sensing cross-modal retrieval model to output a remote sensing cross-modal retrieval result of the multi-modal remote sensing data to be processed; the remote sensing cross-modal retrieval model comprises a ternary feature expression module, a fusion perception module and a self-supervised modal enhancement module;

[0011] The training steps of the remote sensing cross-modal retrieval model are as follows:

[0012] S201, the remote sensing cross-modal retrieval model pre-processes the input multi-modal remote sensing data as the input of the self-supervised modal enhancement module, and outputs the pre-processed multi-modal remote sensing data as the training sample of the remote sensing cross-modal retrieval model and inputs it into the ternary feature expression module;

[0013] S202, the ternary feature expression module comprises a visual Transformer network, a BERT extraction unit and a convolution extraction network, which respectively extract single-modal image features, text features and audio features from the remote sensing images, text data and audio data of the pre-processed multi-modal remote sensing data, and input them into the fusion perception module;

[0014] S203, the fusion perception module dynamically adjusts the input single-modal image features, text features and audio features through a cross-modal self-attention module and a memory unit to generate multi-modal feature embeddings, and obtains a remote sensing cross-modal retrieval result of the multi-modal remote sensing data to be processed according to the multi-modal feature embeddings;

[0015] S204, the self-supervised modal enhancement module optimizes and updates the model parameters of the remote sensing cross-modal retrieval model aiming at minimizing the total loss function constructed by the triple loss function and the triple discriminative loss function;

[0016] S205, repeating steps S201 to S204, iterative training is performed until the remote sensing cross-modal retrieval model converges or reaches a preset iteration number.

[0017] As a preferred solution, in step S201, the pre-processing of the input to-be-processed multi-modal remote sensing data as the input of the self-supervised modal enhancement module specifically includes:

[0018] The remote sensing image is processed by selecting cropping, adjusting size, color jittering, random rotation and Gaussian blur to obtain data-enhanced remote sensing image; the text data is data-enhanced by using rule-based method and back-translation method to obtain fine-grained text data; the audio data is processed by using time shift, speed change cropping and noise mixing method to obtain data-enhanced audio data; then, the multi-modal remote sensing data before pre-processing and the data after pre-processing are aligned by using alignment loss function; the calculation formula of the alignment loss function is:

[0019]

[0020] In the formula, represents the alignment loss function, M represents the total number of multi-modal remote sensing data samples, f j represents the feature vector of the multi-modal remote sensing data before pre-processing, f' k represents the data feature vector after pre-processing, ||.|| represents the norm calculation of the data feature vector.

[0021] As a preferred solution, in step S202, the visual Transformer network comprises two image feature extraction units which are connected in sequence and have the same structure; each image feature extraction unit comprises a first layer normalization layer, a window multi-head self-attention module, a second layer normalization layer and a multi-layer perception in cascade; wherein the input of the visual Transformer network is taken as the input of the first layer normalization layer in the first image feature extraction unit, the output of the first layer normalization layer is taken as the input of the window multi-head self-attention module, the output of the window multi-head self-attention module is added to the input of the first layer normalization layer to serve as the input of the second layer normalization layer, the output of the second layer normalization layer is taken as the input of the multi-layer perception, the output of the multi-layer perception is added to the input of the second layer normalization layer to serve as the input of the second image feature extraction unit; the second image feature extraction unit has the same connection mode as the first image feature extraction unit, and the output of the second image feature extraction unit is taken as the output of the visual Transformer network.

[0022] In the visual Transformer network, the input remote sensing image is first subjected to image cutting processing to obtain a plurality of image blocks; then, each image block is subjected to linear mapping by using a full connection layer to obtain a feature vector, and the average value e p,CLS of the feature vectors of the image blocks is calculated. The feature vectors of the image blocks are respectively represented as e Finally, the image feature I trans of the remote sensing image is extracted by using a window multi-head self-attention module and a multi-layer perception regularization.

[0023] The expression of the image feature I trans of the remote sensing image is as follows:

[0024]

[0025] In the expression, I trans represents the image feature extracted by the visual Transformer network, LN(.) represents a layer normalization layer, WinAtt(.) represents a window multi-head self-attention module, MLP(.) represents a multi-layer perception, and e represents the embedding vector of the remote sensing image.

[0026] As a preferred solution, in step S202, in the BERT extraction unit, the sentences in the input text are first subjected to word segmentation processing to obtain a plurality of words, and the plurality of words are subjected to linear mapping conversion into embedding vectors by using an embedding layer; then, the average value e w,CLS of the embedding vectors of the plurality of words is calculated. respectively represent the embedding vectors of a plurality of words w1, w2, …, w N ; finally, the word sequence features of each sentence are input into the BERT model to extract the text features T bert of the text data.

[0027] The expression of the text features T bert of the text data is as follows:

[0028]

[0029] In the formula, T bert represents the text features extracted by the BERT extraction unit, Att(.) represents a multi-head attention module, represents the word sequence features of each sentence.

[0030] As a preferred solution, in step S202, the convolution extraction network comprises a pre-emphasis unit, a time frame unit, a Hamming window function, a Mel spectrogram conversion unit and a convolution neural network connected in sequence; wherein the input of the convolution extraction network is taken as the input of the pre-emphasis unit, the output of the pre-emphasis unit is taken as the input of the time frame unit, the output of the time frame unit is taken as the input of the Hamming window function, the output of the Hamming window function is taken as the input of the Mel spectrogram conversion unit, the output of the Mel spectrogram conversion unit is taken as the input of the convolution neural network, and the output of the convolution neural network is taken as the output of the convolution extraction network.

[0031] In the convolution extraction network, the input audio data is first pre-emphasized to obtain an enhanced continuous audio signal; secondly, the continuous audio signal is used to obtain a series of discrete time frames by using a time frame unit, and a Hamming window is used for spectral analysis on each time frame to convert the time-domain signal in the time frame into a frequency-domain signal by using a fast Fourier transform; then, the frequency-domain signal is converted into a Mel frequency scale and a Mel spectrogram is generated by using a Mel spectrogram conversion unit; finally, the Mel spectrogram is taken as the input of the convolution neural network, so as to extract the audio features A cnn .

[0032] As a preferred solution, in step S203, the fusion perception module comprises a cross-modal self-attention module, a memory unit, a layer normalization self-attention unit and a layer normalization multi-layer perception unit connected in sequence; the layer normalization self-attention unit comprises a layer normalization layer and a cross-modal self-attention module connected in cascade, and the layer normalization multi-layer perception unit comprises a layer normalization layer and a multi-layer perception machine connected in cascade; wherein the input of the fusion perception module is taken as the input of the cross-modal self-attention module and the memory unit respectively, the output of the cross-modal self-attention module and the output of the memory unit are subjected to multiplication operation and then addition operation, the output of the layer normalization self-attention unit is taken as the input of the layer normalization multi-layer perception unit, and the output of the layer normalization multi-layer perception unit is taken as the output of the fusion perception module;

[0033] In the fusion perception unit, the input single-modal image features, text features and audio features are first processed by the cross-modal self-attention module to reconstruct the context semantic information of each modal feature, and the long-distance dependency relationship information in each single-modal feature is processed by the memory unit, the outputs of the cross-modal self-attention module and the memory unit are subjected to multiplication operation to generate a triple-modal feature vector; then, the triple-modal feature vector is unified to the same dimension by using a full connection layer and is mapped into an embedding feature vector; finally, cross-modal attention and regularization operations are performed on the dimension-aligned feature vector to obtain a multi-modal feature embedding.

[0034] The expression of the multi-modal feature embedding is as follows:

[0035]

[0036] In the formula, CrossAtt(.) represents a cross-modal attention module, S represents an embedding feature vector, represents an embedding expression of the multi-modal feature.

[0037] As a preferred solution, after the multi-modal feature embedding is generated in step S203, the method further comprises: enhancing the similarity of features between samples with the same semantics of the multi-modal feature embedding by using a contrastive loss function in a self-supervised modal enhancement module; and the calculation formula of the contrastive loss function is as follows:

[0038]

[0039] In the formula, represents a contrastive loss function, represents a multi-modal feature embedding vector of sample j, represents a multi-modal feature embedding of sample j with the same semantics but subjected to data enhancement processing, and sin(.,.) represents a similarity measurement function.

[0040] As a preferred solution, in step S204, the triple loss function is:

[0041]

[0042] wherein x represents data of a specific modality, x - represents negative sample data, and sim(.) represents a cosine similarity function.

[0043] As a preferred solution, in step S204, the triple discriminative loss function is:

[0044]

[0045] wherein and respectively represent an image-text discriminative loss function, an image-audio discriminative loss function, and an audio-text discriminative loss function.

[0046] Compared with the prior art, the present application has the following technical effects:

[0047] (1) The present application integrates the audio modality into the remote sensing image-text retrieval task by constructing a remote sensing cross-modality retrieval model, providing a more comprehensive and richer data perspective, complementing the image and text data, enhancing the diversity of data, effectively improving the scalability of the modality model, so that the retrieval model can process and understand multi-source information of images, texts and audios, at the same time, improving the robustness of the model under different environmental conditions, improving the accuracy of retrieval.

[0048] (2) The triple feature expression module constructed by the present application extracts features from remote sensing images, text data and audio data respectively. Such feature extraction operation independently captures the features of image, text and audio data, wherein the introduction of BERT model enhances the understanding ability of deep semantic of text data, and visual Transformer and ResNet-18 provide high-level feature representation of image and audio respectively. Compared with image-text retrieval, the diversity and utilization of remote sensing data are enhanced, which can effectively improve the accuracy and efficiency of cross-modality retrieval of the model

[0049] (3) The present application effectively aggregates the independently captured triple features through the fusion perception module based on modality memory to generate a unified feature representation, solving the problem of semantic alignment complexity change caused by the increase of modalities. The self-attention mechanism is used to reconstruct the context semantic information on each modality, and the storage unit is used to efficiently dynamically supplement and capture long-term memory, which is used to overcome the semantic alignment complexity change caused by the increase of modalities, realize the semantic alignment of cross-modality features, and ensure the accuracy of cross-modality retrieval task.

[0050] (4) The self-supervised modal enhancement module is designed, a large amount of unlabeled multi-modal remote sensing data is pre-trained, the scalability and generalization ability of the model are enhanced, the dependence on expensive and limited manual annotation data is reduced, the data preparation cost is reduced, meanwhile, the self-supervised learning optimizes the feature representation of the model through the feature alignment loss and contrast learning, so that the data with the same semantics are closer in the feature space, and the accuracy of retrieval is improved. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the present application with reference to the accompanying drawings, in which:

[0052] Figure 1 A remote sensing cross-modal retrieval method flow chart based on a deep ternary fusion perception network is disclosed in the present application.

[0053] Figure 2 An architecture schematic diagram of a remote sensing cross-modal retrieval model of an embodiment of the present application is shown in the figure.

[0054] Figure 3 An image, text and audio data enhancement example diagram of an embodiment of the present application is shown in the figure.

[0055] Figure 4 A comparison result diagram of R@10 retrieval results of an embodiment of the present application on an M-RSITMD dataset with respect to alpha value is shown in the figure.

[0056] Figure 5 A top 5 visualization result diagram of an embodiment of the present application in M-RSITMD dataset image-text and image-audio retrieval is shown in the figure. DETAILED DESCRIPTION

[0057] In order to make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the present application with reference to the accompanying drawings, in which:

[0058] The present application will be further described in detail below with reference to the accompanying drawings.

[0059] Remote sensing image retrieval is the basis for important decision-making and data utilization in various industries. Currently, the research on cross-modal retrieval in the field of remote sensing mainly focuses on the image-text retrieval task. However, when introducing the audio modality into remote sensing cross-modal image-text retrieval, it faces the difficulty of constructing a unified representation space among multiple modalities. Due to the natural semantic difference among different modalities and the limited paired data, the model is difficult to capture rich modality features, thereby affecting the learning effect of modality representation. In addition, due to the coverage range and acquisition cost of remote sensing data, high-quality labeled data is scarce, and a small amount of labeled data is difficult to support effective spatial representation learning, which is particularly prominent in multi-modal retrieval.

[0060] To solve this problem, the present application provides a remote sensing cross-modal retrieval method based on a deep ternary fusion perception network, and the method flow is as follows Figure 1 Specifically, it includes:

[0061] S1, obtaining the multi-modal remote sensing data to be processed, the multi-modal remote sensing data including remote sensing images, text data and audio data;

[0062] S2, inputting the multi-modal remote sensing data into the trained remote sensing cross-modal retrieval model to output the remote sensing cross-modal retrieval result of the multi-modal remote sensing data to be processed; the remote sensing cross-modal retrieval model includes a ternary feature expression module, a fusion perception module and a self-supervised modality enhancement module;

[0063] The training steps of the remote sensing cross-modal retrieval model are as follows:

[0064] S201, the remote sensing cross-modal retrieval model pre-processes the input multi-modal remote sensing data as the input of the self-supervised modality enhancement module, and outputs the pre-processed multi-modal remote sensing data as the training sample of the remote sensing cross-modal retrieval model and inputs it into the ternary feature expression module;

[0065] S202, the ternary feature expression module includes a visual Transformer network, a BERT extraction unit and a convolution extraction network, which respectively extract single-modal image features, text features and audio features from the remote sensing images, text data and audio data of the pre-processed multi-modal remote sensing data, and input them into the fusion perception module;

[0066] S203, the fusion perception module dynamically adjusts the input single-modal image features, text features and audio features through the cross-modal self-attention module and the memory unit to generate multi-modal feature embedding, and obtains the remote sensing cross-modal retrieval result of the multi-modal remote sensing data to be processed according to the multi-modal feature embedding;

[0067] S204, the self-supervised modal enhancement module optimizes and updates the model parameters of the remote sensing cross-modal retrieval model with the objective of minimizing the total loss function constructed by the triplet loss function and the triadic discriminative loss function;

[0068] S205, repeating steps S201 to S204, iterative training is performed until the remote sensing cross-modal retrieval model converges or reaches a preset number of iterations.

[0069] Through the above process of the present application, it can be seen that the cross-modal retrieval constructed by the present application inherits the audio modality to the remote sensing image-text retrieval task, providing a more comprehensive and rich data perspective, which is complementary to image and text data, effectively improving the model modality scalability, and at the same time, enhancing the utilization rate of diverse remote sensing data. The fusion of audio data overcomes the influence of environmental factors on the performance of the model, improves the robustness and reliability of the retrieval model in various environments. At the same time, the model of the present application aggregates the obtained triadic features using the fusion perception module to generate a unified feature representation, solving the problem of semantic alignment complexity change due to the increase of modalities. In view of the problem that high-quality labeled data is scarce, which makes it difficult to effectively learn spatial representation, the present application uses a large amount of unlabeled multi-modal data for pre-training through a self-supervised modal enhancement module, solving the problem of difficulty in obtaining high-quality paired data and high labeling cost. The model can be fully learned and generalized under limited paired data.

[0070] In order to better introduce the technical scheme of the present application, the following parts will be explained in more detail.

[0071] 1、Multi-modal remote sensing data preprocessing

[0072] In specific application implementation, the multi-modal remote sensing data for training the remote sensing cross-modal retrieval model can be downloaded from the public data sets M-RSICD and M-RSITMD. Among them, the M-RSICD data set M-RSICD data set contains 10921 high-resolution remote sensing images, the images are from multiple map services and are equipped with five descriptive sentences and audio provided by remote sensing experts, aiming to improve the diversity and accuracy of remote sensing image description task. In addition, the M-RSITMD data set contains remote sensing images of diversified ground features and accurate text descriptions, and each sample is equipped with an audio, which is recorded by two speakers and adjusted in volume and speed to enhance the robustness of the audio annotation data.

[0073] 2、Remote sensing cross-modal retrieval model

[0074] Figure 2 The architecture diagram of the remote sensing cross-modal retrieval model of the present application is as follows, Figure 2It can be known that the remote sensing cross-modal retrieval model of the application comprises a ternary feature expression module, a fusion perception module and a self-supervised modal enhancement module. The ternary feature expression module, the fusion perception module and the self-supervised modal enhancement module will be introduced in detail below.

[0075] 2.1, ternary feature expression module

[0076] The ternary feature expression module provides powerful feature extraction and embedding capabilities, which uses a visual Transformer network, a BERT extraction unit and a convolution extraction network to independently capture single-modal image features, text features and audio features from the remote sensing image text data and audio data of the pre-processed multi-modal remote sensing data. Compared with graphic retrieval, the diversity and utilization of remote sensing data are enhanced, which can effectively improve the accuracy and efficiency of cross-modal retrieval of the model. The visual Transformer network, the BERT extraction unit and the convolution extraction network will be introduced in detail below.

[0077] 2.1.1, visual Transformer network

[0078] The visual Transformer network comprises two image feature extraction units which are connected in sequence and have the same structure; each image feature extraction unit comprises a cascaded first layer normalization layer, a window multi-head self-attention module, a second layer normalization layer and a multi-layer perception machine; wherein the input of the visual Transformer network is taken as the input of the first layer normalization layer in the first image feature extraction unit, the output of the first layer normalization layer is taken as the input of the window multi-head self-attention module, the output of the window multi-head self-attention module is added with the input of the first layer normalization layer as the input of the second layer normalization layer, the output of the second layer normalization layer is taken as the input of the multi-layer perception machine, and the output of the multi-layer perception machine is added with the input of the second layer normalization layer as the input of the second image feature extraction unit; the second image feature extraction unit has the same connection mode as the first image feature extraction unit, and the output of the second image feature extraction unit is taken as the output of the visual Transformer network.

[0079] In the visual Transformer network, the input remote sensing image with 256*256 pixels is first subjected to image cutting processing, which is divided into 32*32 pixel image blocks, a total of 64 image blocks are obtained, and these image blocks are formulaized as The image blocks can simultaneously retain semantic features and positional information; then, a fully connected layer is used to linearly map each image block into a 512-dimensional feature vector, and the average value e of the feature vectors of each image block is calculated. CLS This yields the embedding vector of the remote sensing image, represented as... Representing image blocks The feature vectors are obtained; finally, the image features I of the remote sensing image are extracted through a window multi-head self-attention module and a multilayer perceptron regularization. trans .

[0080] Image features I of the remote sensing image trans The expression is as follows:

[0081]

[0082] In the formula, I trans This represents the image features extracted through the Visual Transformer network, LN(.) represents the layer normalization layer, WinAtt(.) represents the window multi-head self-attention module, and MLP(.) represents the multilayer perceptron. This represents the embedding vector of a remote sensing image.

[0083] 2.1.2 BERT Extraction Unit

[0084] In the BERT extraction unit, each sentence in the input text is first segmented into several words, denoted as T = w1, w2, ..., w N The process involves using an embedding layer to linearly map several words into embedding vectors; then, the average value e of the embedding vectors of these words is calculated. w,CLS This allows us to obtain the word sequence features of each sentence in the text data, represented as... Each represents a number of words w1, w2, ..., w N The embedding vectors are then processed; finally, the word sequence features of each sentence are input into the BERT model. Pooling is applied to the output of the last encoder layer of the BERT model to obtain a vector integrating the full-text information. This vector is then transformed into a 512-dimensional feature representation through a fully connected layer to extract the text features T from the text data. bert .

[0085] The text features T of the text data bert The expression is as follows:

[0086]

[0087] In the formula, T bert This represents the text features extracted by the BERT extraction unit, and Att(.) represents the multi-head attention module. represent the word sequence features of each sentence.

[0088] 2.1.3, convolutional extraction network

[0089] The convolutional extraction network comprises a pre-emphasis unit, a time frame unit, a Hamming window function, a Mel spectrogram conversion unit and a convolutional neural network connected in sequence; wherein the input of the convolutional extraction network is taken as the input of the pre-emphasis unit, the output of the pre-emphasis unit is taken as the input of the time frame unit, the output of the time frame unit is taken as the input of the Hamming window function, the output of the Hamming window function is taken as the input of the Mel spectrogram conversion unit, the output of the Mel spectrogram conversion unit is taken as the input of the convolutional neural network, and the output of the convolutional neural network is taken as the output of the convolutional extraction network;

[0090] In the convolutional extraction network, the input audio data is first pre-emphasized to obtain an enhanced continuous audio signal; secondly, the continuous audio signal is used to obtain a series of discrete time frames by using a time frame unit, and a Hamming window is used for spectral analysis on each time frame to convert the time-domain signal in the time frame into a frequency-domain signal by using fast Fourier transform; then, the frequency-domain signal is converted into a Mel frequency scale and a Mel spectrogram is generated by using a Mel spectrogram conversion unit; finally, RessNet-18 is selected as the convolutional neural network, and the Mel spectrogram is taken as the input of RessNet-18, so that the audio feature A of the audio data is extracted. cnn .

[0091] 2.2, fusion perception module

[0092] Although the image, text and audio features are expressed in the same embedding space, due to the independent extraction and expression of the multi-modal features, and the existence of semantic differences, in order to realize the semantic alignment between the cross-modal retrieval features, fuse the embedding features of these independent modalities, optimize the clustering center distribution between different modalities, and promote the distribution of the feature space to be both discrete and close to each other, a fusion perception module based on modal memory is proposed.

[0093] The fusion perception module comprises a cross-modal self-attention module, a memory unit, a layer normalization self-attention unit and a layer normalization multi-layer perception unit connected in sequence; the layer normalization self-attention unit comprises a layer normalization layer and a cross-modal self-attention module connected in cascade, and the layer normalization multi-layer perception unit comprises a layer normalization layer and a multi-layer perception machine connected in cascade; wherein the input of the fusion perception module is taken as the input of the cross-modal self-attention module and the memory unit respectively, the output of the cross-modal self-attention module and the output of the memory unit are subjected to multiplication operation and then addition operation, and the output of the layer normalization self-attention unit is taken as the input of the layer normalization multi-layer perception unit, and the output of the layer normalization multi-layer perception unit is taken as the output of the fusion perception module;

[0094] In the fusion perception unit, the input single-modal image feature I trans , the text feature T bert and the audio feature A cnn are first processed by the cross-modal self-attention module to reconstruct the context semantic information of each modal feature, and the long-distance dependency relationship information in each single-modal feature is processed by the memory unit, and the outputs of the cross-modal self-attention module and the memory unit are subjected to multiplication operation to generate a ternary modal feature vector and

[0095] The process of correcting the ternary modal feature vector based on the cross-modal self-attention module and the memory unit is shown in the following formula:

[0096]

[0097] In the formula, SelfAtt(.) represents a self-attention mechanism, represents vector multiplication operation, M I,T,A represents a trainable multi-modal memory matrix;

[0098] Then, the ternary modal feature vector is unified to the same dimension by using a full connection layer, and is mapped into an embedding feature vector, denoted as Finally, cross-modal attention and regularization operations are performed on the dimension-aligned feature vector S to obtain a multi-modal feature embedding;

[0099] The expression of the multi-modal feature embedding is as follows:

[0100]

[0101] In the formula, CrossAtt(.) represents a cross-modal attention module, S represents an embedding feature vector, represents an embedding expression of the multi-modal feature.

[0102] The fusion perception module effectively aggregates the independently captured ternary features, generates a unified feature representation, realizes unified coding and semantic alignment of multi-modal data, so as to subsequently perform cross-modal retrieval and sorting, and solves the problem of semantic alignment complexity change due to the increase of modalities. The self-attention mechanism is used to reconstruct the context semantic information on each modality, and the storage unit is used to efficiently dynamically supplement and capture long-term memory, so as to overcome the semantic alignment complexity change due to the increase of modalities, realize cross-modal feature semantic alignment, and ensure the accuracy of the cross-modal retrieval task.

[0103] 2.3, self-supervised modality enhancement module

[0104] The self-supervised modality enhancement module improves the performance of the model under limited labeled data by preprocessing the multi-modal remote sensing data to be processed, and enables the model to better learn useful representations from unlabeled data through a self-supervised learning mechanism, thereby exhibiting higher accuracy and stronger adaptability in cross-modal retrieval and sorting tasks in the field of remote sensing.

[0105] As shown in Figure 3 The present application pre-processes the above-mentioned obtained multi-modal remote sensing data as the input of the self-supervised modality enhancement module, and uses the self-supervised module to pre-train on a large amount of unlabeled multi-modal remote sensing data to improve the scalability and generalization ability of the model. In specific implementation, first, the enhancement and expansion strategy for the multi-modal remote sensing data to be processed is constructed, wherein the remote sensing image is processed by selecting cropping, adjusting size, color jittering, random rotation and Gaussian blur to obtain data-enhanced remote sensing images; the text data is data-enhanced by using a rule-based method and a back-translation method to obtain fine-grained text data; the audio data is processed by using time shift, speed change cropping and noise mixing methods to obtain data-enhanced audio data; then, the multi-modal remote sensing data before pre-processing and the data after pre-processing are aligned in feature by using an alignment loss function; the calculation formula of the alignment loss function is as follows:

[0106]

[0107] In the formula, represents the alignment loss function, M represents the total number of multi-modal remote sensing data samples, f represents the feature vector of the multi-modal remote sensing data before pre-processing, f' represents the data feature vector after pre-processing, and ||.|| represents the norm calculation of the data feature vector.

[0108] At the same time, the embedding of the semantic-aligned multi-modal features output by the fusion perception module is used as the input of the cross-modal retrieval and sorting module The similarity of features between the samples with the same semantics of the multi-modal feature embedding is enhanced through a contrast loss function in the self-supervised modal enhancement module, and a calculation formula of the contrast loss function is:

[0109]

[0110] In the formula, denotes the contrast loss function, denotes the multi-modal feature embedding vector of the sample j, denotes the multi-modal feature embedding of the sample j with the same semantics but after data enhancement processing, and sin(.,.) denotes a similarity measurement function.

[0111] 3. Training of the remote sensing cross-modal retrieval model

[0112] In a specific implementation, the training of the remote sensing cross-modal retrieval model is performed in the following manner: the preprocessed multi-modal remote sensing data is taken as a training sample of the remote sensing cross-modal retrieval model, input into the remote sensing cross-modal retrieval model, a triplet loss function and a triplet discriminant loss function for optimizing the distance in the feature space to improve the recognition and classification ability of the model are constructed, and the model parameters of the remote sensing cross-modal retrieval model are optimized and updated to minimize the total loss function constructed by the triplet loss function and the triplet discriminant loss function, and then the remote sensing cross-modal retrieval model is trained.

[0113] The total loss function is denoted as:

[0114]

[0115] In the formula, denotes the total loss function, denotes the triplet loss function, denotes the triplet discriminant loss function.

[0116] In the application of the present application, for a given set of image, text and audio samples (I, T, A), according to whether these samples constitute a relevant match in semantics, they are divided into positive samples and negative samples, and the triplet loss is preferably calculated by the following formula:

[0117]

[0118] In the formula, x denotes data of a specific modality, x - denotes negative sample data, and sim(.) denotes a cosine similarity function.

[0119] The triplet loss function is minimized to reduce the distance between positive samples and expand the gap between negative samples, so as to optimize the learning process of the model and constrain it.

[0120] In the application of the present application, a ternary discriminative loss is introduced to combine three modalities to realize the generalization of common features between similar samples and the optimization of cross-modal retrieval accuracy, preferably by calculating the ternary discriminative loss according to the following formula:

[0121]

[0122] In the formula, and respectively represent the image-text discriminative loss function, the image-audio discriminative loss function and the audio-text discriminative loss function.

[0123] Among them, the calculation methods of the three loss functions are similar. Taking the image-text retrieval task as an example, the calculation process of the image-text discriminative loss function is as follows: first, the similarities s I→T and s T→I between the image samples and the text samples in the batch are calculated, and the calculation formula is as follows:

[0124]

[0125] In the formula, I trans and T bert represent the embedding feature representations of the image and text modalities respectively, and ω represents the trainable learning parameter; then, the image-text discriminative loss function L is calculated, and the expression is as follows:

[0126]

[0127] In the formula, g represents the similarity between the real matching image and text in the sample, and H(.) represents the cross-entropy function, which is defined as 1 for positive samples and 0 for negative samples.

[0128] 4. Embodiment

[0129] In order to better illustrate the advantages of the technical scheme of the present application, the following experiments are disclosed in this embodiment.

[0130] The multi-modal remote sensing data applied in this embodiment mainly consists of remote sensing images with high spatial resolution, descriptive texts corresponding to these images and audios of the texts. In this embodiment, the training set, test set and validation set are selected from the M-RSICD and M-RSITMD two public data sets, and the cross-modal retrieval model is experimented after the division according to the ratio of 7:2:1.

[0131] 4.1. Evaluation index

[0132] R@k: is a recall-based evaluation metric that measures the proportion of relevant documents retrieved by the search system within the top k positions of the search results. It is a special case of recall, which focuses on the top k search results. Recall is an important performance metric in information retrieval, defined as the proportion of positive samples successfully retrieved and returned by the search system among all relevant data for a user query. R@k focuses on whether the system can effectively retrieve the information that users are truly interested in within a limited result display. It is commonly used to evaluate the quality of a specific number of search results, such as R@1, R@5, or R@10, which represent the proportion of relevant documents in the top 1, 5, or 10 search results, respectively. The higher the value of R@k, the better the performance of the search system. The calculation formula is as follows:

[0133]

[0134] where |R| represents the total number of query-related data, r i represents the ith search result, k represents the top k positions of the search results, represents the indicator function, when r i is a positive sample, otherwise 0.

[0135] mR: is the arithmetic average operation of R@k indicators for a series of queries, used to evaluate the overall performance and consistency of the search system when processing multiple queries. The calculation formula is as follows:

[0136] mR = (R@1 + R@5 + R@10) / 3 # (4.13)

[0137] 4.2, Comparison of experimental results

[0138] 4.2.1, Comparative analysis of search performance

[0139] The cross-membrane retrieval model of this embodiment and six latest remote sensing multi-modal retrieval methods VSE++, SCAN, CAMP, LW-MCR, AMFMN and MCRN are compared in performance analysis in M-RSICD and M-RSITMD two data sets.

[0140] Table 1 shows the comparison of the retrieval results of the R@k index when performing the image-text retrieval and image-audio retrieval tasks on the M-RSICD dataset with the current latest method. Specifically, the retrieval performance of the remote sensing cross-modal retrieval model shows significant performance advantages in the four tasks, with mR values of 24.09%, 22.39%, 19.81%, and 16.82% in the four tasks, with a maximum increase of 2.55 percentage points and an average increase of about 11.39 percentage points compared to the suboptimal method. This indicates the effectiveness of the method in the ternary feature expression fusion perception stage, which can accurately extract key features and exhibit good performance. However, due to the limitations of the M-RSICD dataset in terms of sample diversity, label completeness, and data representativeness, the model has limitations in learning and generalizing to more extensive application fields, which further affects the further improvement of performance.

[0141] Table 1 Comparison of R@k retrieval results on M-RSICD dataset with the latest method

[0142]

[0143]

[0144] The M-RSITMD dataset contains remote sensing images with diverse ground features and accurate text descriptions, and innovatively adds adjusted audio annotations to each sample. Table 2 shows the comparison of the retrieval results of the R@k index when performing the image-text retrieval and image-audio retrieval tasks on the M-RSICD dataset with the current latest method. Specifically, the method of the embodiment shows excellent performance in the image-text retrieval and image-audio retrieval tasks. In the image-text retrieval task, the method of the embodiment has a Top-1 accuracy of 15.02%, with a maximum improvement of 6.45%. In terms of Top-5 and Top-10 accuracy, the method of the embodiment leads with 12.45% and 64.24%, respectively, highlighting its excellent ability in the image-to-text retrieval task. In the image-audio retrieval task, the method of the embodiment ranks second with a Top-1 accuracy of 8.05%, but leads in other evaluation indicators. This result shows that although the lack of labels in remote sensing cross-modal data poses challenges to the robustness of model training, the pre-training method effectively alleviates this problem by utilizing unpaired samples.

[0145] Table 2 Comparison of R@k retrieval results on M-RSITMD dataset with the latest method

[0146]

[0147]

[0148] It can be seen that the method of the embodiment performs better in most tasks, and makes significant progress in the R@10 index, while the improvement in the R@1 index is relatively more challenging. This result reveals that it can effectively face the change in semantic alignment complexity due to the increase in modalities, and by pre-training a large amount of unlabeled multi-modal remote sensing data, not only enhances the generalization and expansibility of the model, but also effectively alleviates the problems of scarcity and high cost of remote sensing data labeling.

[0149] 4.2.2, comparison and analysis of retrieval efficiency

[0150] The transmembrane state retrieval model of the embodiment and the three methods of VSE++, AMFMN and MCRN are compared and evaluated in detail in terms of test and inference time on the M-RSICD and M-RSITMD datasets. To ensure the accuracy and fairness of the time difference analysis, the experiment is carried out in a CPU environment without other loads, and is performed multiple times to obtain reliable data.

[0151] Table 3 details the test and inference time of each method on the M-RSICD and M-RSITMD datasets. It is worth noting that the fusion method such as CAMP is not included in the experiment, because these methods use a serial execution mode in feature extraction and similarity calculation, resulting in relatively low retrieval efficiency. The experimental results show that VSE++ performs best in test and inference time, which is largely due to its relatively simple network structure. Although the test and inference time of the DTFP method is relatively long, the difference compared with other comparison methods is not significant, and still remains within the same order of magnitude. In summary, DTFP achieves a good balance between retrieval performance and efficiency.

[0152] Table 3 Test and inference time of different methods on M-RSICD and M-RSITMD datasets

[0153]

[0154] 4.3, self-comparison experimental results

[0155] 4.3.1, ablation experiment analysis

[0156] As shown in Table 4, the R@k retrieval results of different ablation methods for the image-text retrieval and image-audio retrieval tasks on the M-RSITMD dataset are compared. Among them, TFE, FP and SME represent ternary feature expression, fusion perception module and self-supervised modal enhancement, and represent ternary triplet loss and ternary discriminative loss, respectively. Specifically, the following ablation experiments are set up:

[0157] (1) m1: No ablation module, representing the complete state of the model;

[0158] (2) m2: Ablate the triple feature expression (TFE) module. In image feature expression, remove the local window multi-head self-attention mechanism and multi-layer perception operation, and only express the image embedding as image features. In text feature expression, remove the multi-head attention mechanism and text encoder operation, and only express the token sequence of the sentence as text features.

[0159] (3) m3: Fusion perception (FP) module, remove the memory unit correction process, and only use the original features as the aggregated embedding features S.

[0160] (4) m4: Ablate the self-supervised modal enhancement (SME) module, remove the pre-training process on a large amount of unlabeled multi-modal remote sensing data.

[0161] (5) m 5-6 : Ablate the triple value triple loss and the triple discriminative loss respectively, i.e. remove the corresponding loss function in the training process.

[0162] It can be seen that after removing the triple feature expression module, the retrieval accuracy of the m2 model decreases significantly, reaching the lowest value in the table. This shows that the triple feature expression module plays a crucial role in capturing the detailed features of images and the context semantics of text in the cross-modal retrieval task. In addition, the experimental results of the m3 model also show a significant downward trend, which shows that when introducing the audio modality in remote sensing cross-modal image-text retrieval, it is difficult to unify the representation space among multi-modalities. Due to the inherent semantic differences between different modalities and the limited paired data, the model has difficulty in capturing rich cross-modal features, which in turn affects the learning effect of modality representation. In the absence of the self-supervised modal enhancement module, the retrieval accuracy of the m4 model is also slightly affected, which shows that a small amount of labeled data is difficult to support effective spatial representation learning. Finally, by comparing the effects of the triple value triple loss and the triple discriminative loss in the m 5-6 model, it can be found that the precision of the m5 model decreases the most, and the influence of the m6 model is relatively small. This shows that the triple value triple loss can effectively generalize the common features between similar samples, thereby optimizing the accuracy of cross-modal retrieval.

[0163] Table 4 R@k retrieval results of different ablation methods on M-RSITMD dataset

[0164]

[0165]

[0166] 4.3.2, Hyperparameter experiment analysis

[0167] In the cross-modal retrieval task, the tri-value triplet loss function aims to learn the similarity measure between samples by minimizing the distance between anchor samples and positive samples, while maximizing the distance between anchor samples and negative samples. The margin parameter a has a significant impact on the performance of the model. To this end, as shown in Figure 4 , a series of self-comparison experiments were designed on the M-RSITMD dataset to analyze the impact of a value on retrieval results. Finally, based on the results of multiple experiments, it is shown that the best retrieval results can be obtained when a = 0.4.

[0168] 4.4, Visualization analysis of experimental results

[0169] In order to more intuitively evaluate the effectiveness of the method proposed in the present application, as shown in Figure 5 , the visualization samples of the top 5 results of image-text retrieval and image-audio retrieval on the M-RSITMD dataset are shown respectively. Through this intuitive display, the performance of the method in terms of retrieval accuracy and relevance can be clearly observed, thereby deeply understanding the capabilities and limitations of various methods in processing cross-modal data.

[0170] As can be seen, Figure 5 , the three columns represent the retrieval task category, the image, text and audio information of the query input, and the corresponding top five retrieval results. In this embodiment, the positive samples are highlighted in green for easy identification. Specifically, in the image-text retrieval task, "park" and "resort" are selected as the retrieval theme, and the results show that the correct information is successfully returned in the top five retrieval results, which confirms that even in the image-text retrieval task, the introduction of additional audio modalities does not affect the original retrieval performance of the system. In the image-audio retrieval task, "school" and "highway" are used as the retrieval theme, although the model fails to accurately match the retrieval results in the first position, but the positive samples are included in the top five results. This indicates that the model has effective expression capability for three-modal features in high-dimensional feature space, and can further realize the unification of heterogeneous feature spaces.

[0171] 5, Summary

[0172] In summary, the present method uses tri-modal feature expression strategy, fusion perception mechanism and self-supervised modal enhancement technology to solve the key problems of "model modality scalability" and "remote sensing labeled data scarcity and high cost", significantly enhancing the accuracy and efficiency of remote sensing data cross-modal retrieval task. Specifically, the present method has the following technical advantages:

[0173] (1) The present application integrates the audio modality into the remote sensing image-text retrieval task by constructing a remote sensing cross-modal retrieval model, providing a more comprehensive and richer data perspective, complementing image and text data, enhancing data diversity, effectively improving modality model scalability, so that the retrieval model can process and understand multi-source information of images, texts and audios, at the same time, improve the robustness of the model in different environmental conditions, and improve the accuracy of retrieval.

[0174] (2) The ternary feature expression module constructed in the present application extracts features from remote sensing images, text data and audio data respectively. Such feature extraction operation independently captures the features of image, text and audio data, wherein the introduction of BERT model enhances the understanding ability of deep semantic of text data, and visual Transformer and ResNet-18 provide high-level feature representation of image and audio respectively. Compared with image-text retrieval, the diversity and utilization of remote sensing data are enhanced, which can effectively improve the accuracy and efficiency of cross-modal retrieval of the model

[0175] (3) The present application generates a unified feature representation by effectively aggregating the independently captured ternary features through the fusion perception module based on modal memory, solving the problem of semantic alignment complexity change caused by the increase of modal. The self-attention mechanism is used to reconstruct the context semantic information on each modal, and the storage unit is used to efficiently dynamically supplement and capture long-term memory, which is used to overcome the semantic alignment complexity change caused by the increase of modal, realize the semantic alignment of cross-modal features, and ensure the accuracy of cross-modal retrieval task.

[0176] (4) The present application designs a self-supervised modal enhancement module, which enhances the scalability and generalization ability of the model by pre-training a large amount of unlabeled multi-modal remote sensing data, reduces the dependence on expensive and limited manual annotation data, reduces the data preparation cost, at the same time, the self-supervised learning optimizes the feature representation of the model through feature alignment loss and contrast learning, so that the data with the same semantics are closer in the feature space, thereby improving the accuracy of retrieval.

[0177] Finally, it should be pointed out that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described by referring to the preferred embodiments of the present application, those skilled in the art should understand that various changes can be made in form and detail without departing from the spirit and scope of the present application defined in the appended claims.

Claims

1. A remote sensing cross-modal retrieval method based on a deep ternary fusion sensing network, characterized in that, Includes the following steps: S1. Acquire multimodal remote sensing data to be processed, wherein the multimodal remote sensing data includes remote sensing images, text data and audio data; S2. Input the multimodal remote sensing data into the trained remote sensing cross-modal retrieval model, and output the remote sensing cross-modal retrieval results of the multimodal remote sensing data to be processed; the remote sensing cross-modal retrieval model includes a ternary feature representation module, a fusion perception module, and a self-supervised modality enhancement module; The training steps for the remote sensing cross-modal retrieval model are as follows: S201. The remote sensing cross-modal retrieval model uses the input multimodal remote sensing data as the input of the self-supervised modality enhancement module for preprocessing, and outputs the preprocessed multimodal remote sensing data as the training samples of the remote sensing cross-modal retrieval model, and inputs it into the ternary feature expression module. S202. The ternary feature expression module includes a visual Transformer network, a BERT extraction unit, and a convolutional extraction network, which extract single-modal image features, text features, and audio features from the remote sensing images, text data, and audio data of the preprocessed multimodal remote sensing data, respectively, and use them as input to the fusion perception module. S203. The fusion sensing module dynamically adjusts the input single-modal image features, text features, and audio features through a cross-modal self-attention module and a memory unit to fuse and generate a multimodal feature embedding, and obtains the remote sensing cross-modal retrieval result of the multimodal remote sensing data to be processed based on the multimodal feature embedding; in step S203, the fusion sensing module includes a cross-modal self-attention module, a memory unit, a layer normalization self-attention unit, and a layer normalization multilayer sensing unit connected in sequence; the layer normalization self-attention unit includes a cascaded layer normalization layer and a cross-modal self-attention module, and the layer normalization multilayer sensing unit includes a cascaded layer normalization layer and a multilayer perceptron; wherein, the input of the fusion sensing module is used as the input of the cross-modal self-attention module and the memory unit respectively, the output of the cross-modal self-attention module and the output of the memory unit are multiplied and then added as the input of the layer normalization self-attention unit, the output of the layer normalization self-attention unit is used as the input of the layer normalization multilayer sensing unit, and the output of the layer normalization multilayer sensing unit is used as the output of the fusion sensing module; In the fusion perception unit, the input single-modal image features, text features, and audio features are first processed through a cross-modal self-attention module to reconstruct the contextual semantic information of each modality feature. Simultaneously, a memory unit processes the long-distance dependency information in each single-modal feature. The outputs of the cross-modal self-attention module and the memory unit are multiplied to fuse and generate a ternary modality feature vector. Then, a fully connected layer is used to unify the ternary modality feature vector to the same dimension and map it into an embedded feature vector. Finally, cross-modal attention and regularization operations are performed on the dimension-aligned feature vector to fuse and obtain a multimodal feature embedding. The expression for the multimodal feature embedding is as follows: ; In the formula, This indicates a cross-modal attention module. Represents the embedded feature vector. This represents an embedded representation that incorporates multimodal features; S204. The self-supervised modality enhancement module optimizes and updates the model parameters of the remote sensing cross-modal retrieval model with the goal of minimizing the total loss function constructed by the triplet loss function and the triplet discrimination loss function; S205. Repeat steps S201 to S204 to perform iterative training until the remote sensing cross-modal retrieval model converges or reaches the preset number of iterations.

2. The remote sensing cross-modal retrieval method based on a deep ternary fusion sensing network according to claim 1, characterized in that, In step S201, the preprocessing of the input multimodal remote sensing data to be processed as input to the self-supervised modality enhancement module specifically includes: The remote sensing images are processed by cropping, resizing, color jittering, random rotation, and Gaussian blurring to obtain data-enhanced remote sensing images; the text data is enhanced using rule-based and back-translation methods to obtain fine-grained text data; the audio data is processed using time-shift, velocity-change cropping, and noise mixing methods to obtain data-enhanced audio data; then, the pre-processed multimodal remote sensing data and the pre-processed data are feature-aligned using an alignment loss function; the formula for calculating the alignment loss function is as follows: ; In the formula, Represents the alignment loss function. This represents the total number of samples in the multimodal remote sensing data. This represents the feature vector of the multimodal remote sensing data before preprocessing. This represents the preprocessed data feature vector. This represents the norm calculation of the data feature vector.

3. The remote sensing cross-modal retrieval method based on a deep ternary fusion sensing network according to claim 1, characterized in that, In step S202, the visual Transformer network includes two image feature extraction units with identical structures connected sequentially. Each image feature extraction unit includes a cascaded first normalization layer, a window multi-head self-attention module, a second normalization layer, and a multilayer perceptron. The input of the visual Transformer network serves as the input of the first normalization layer in the first image feature extraction unit. The output of the first normalization layer serves as the input of the window multi-head self-attention module. The output of the window multi-head self-attention module is added to the input of the first normalization layer, which serves as the input of the second normalization layer. The output of the second normalization layer serves as the input of the multilayer perceptron. The output of the multilayer perceptron is added to the input of the second normalization layer, which serves as the input of the second image feature extraction unit. The second image feature extraction unit is connected in the same way as the first image feature extraction unit, and the output of the second image feature extraction unit serves as the output of the visual Transformer network. In the visual Transformer network, the input remote sensing image is first segmented into several image patches. Then, a fully connected layer is used to linearly map each image patch to obtain a feature vector, and the average value of the feature vectors of each image patch is calculated. This yields the embedding vector of the remote sensing image, represented as... , Representing image blocks The feature vectors are obtained; finally, the image features of the remote sensing image are extracted through a window multi-head self-attention module and a multilayer perceptron regularization. ; Image features of the remote sensing image The expression is as follows: ; In the formula, This represents the image features extracted using a visual Transformer network. Presentation layer normalization layer, This indicates a multi-head self-attention module for windows. This represents a multilayer perceptron. This represents the embedding vector of a remote sensing image.

4. The remote sensing cross-modal retrieval method based on a deep ternary fusion sensing network according to claim 1, characterized in that, In step S202, the BERT extraction unit first performs word segmentation on each sentence in the input text to obtain several words, and then uses the embedding layer to linearly map the several words into embedding vectors. Then, calculate the average of several word embedding vectors. This allows us to obtain the word sequence features of each sentence in the text data, represented as... , Each represents a number of words The embedding vectors are obtained; finally, the word sequence features of each sentence are input into the BERT model to extract the text features of the text data. ; The text features of the text data The expression is as follows: ; In the formula, This represents the text features extracted by the BERT extraction unit. This indicates a multi-head attention module. This represents the word sequence characteristics of each sentence.

5. The remote sensing cross-modal retrieval method based on a deep ternary fusion sensing network according to claim 1, characterized in that, In step S202, the convolutional extraction network includes a pre-emphasis unit, a time frame unit, a Hamming window function, a Mel spectrogram conversion unit, and a convolutional neural network connected in sequence; wherein, the input of the convolutional extraction network serves as the input of the pre-emphasis unit, the output of the pre-emphasis unit serves as the input of the time frame unit, the output of the time frame unit serves as the input of the Hamming window function, the output of the Hamming window function serves as the input of the Mel spectrogram conversion unit, the output of the Mel spectrogram conversion unit serves as the input of the convolutional neural network, and the output of the convolutional neural network serves as the output of the convolutional extraction network; In the convolutional extraction network, the input audio data is first pre-emphasized to obtain a continuous audio signal with enhanced audio data. Next, the continuous audio signal is processed into a series of discrete time frames using a time frame unit. For each time frame, a Hamming window is used for spectral analysis, and a Fast Fourier Transform (FFT) is applied to convert the time-domain signal in the time frame into a frequency-domain signal. Then, a Mel spectrogram conversion unit converts the frequency-domain signal into a Mel frequency scale and generates a Mel spectrogram. Finally, the Mel spectrogram is used as input to the convolutional neural network to extract the audio features of the audio data. .

6. The remote sensing cross-modal retrieval method based on a deep ternary fusion sensing network according to claim 1, characterized in that, Step S203, after generating multimodal feature embeddings through fusion, further includes enhancing the similarity of features between semantically identical samples in the multimodal feature embeddings using a contrastive loss function in the self-supervised modality enhancement module; the formula for calculating the contrastive loss function is: ; In the formula, This represents the contrastive loss function. This represents the multimodal feature embedding vector of sample j. This represents a multimodal feature embedding that has the same semantics as sample j but has undergone data augmentation. This represents a similarity measurement function.

7. The remote sensing cross-modal retrieval method based on a deep ternary fusion sensing network according to claim 1, characterized in that, In step S204, the total loss function for training is: ; in, Represents the total loss function. Represents the triplet loss function. This represents the ternary discriminant loss function.

8. The remote sensing cross-modal retrieval method based on a deep ternary fusion sensing network according to claim 7, characterized in that, In step S204, the triplet loss function is: ; In the formula, Data representing a specific modality, Represents negative sample data. This represents the cosine similarity function.

9. The remote sensing cross-modal retrieval method based on a deep ternary fusion sensing network according to claim 7, characterized in that, In step S204, the ternary discrimination loss function is: ; In the formula, , and Let represent the image-text discrimination loss function, the image-audio discrimination loss function, and the audio-text discrimination loss function, respectively.

Citation Information

Patent Citations

  • Remote sensing image cross-modal retrieval method based on language and visual detail feature fusion

    CN116775922A

  • Cross-modal fine-grained retrieval method based on multi-channel fusion

    CN118113888A