A method for generating remote sensing image descriptions based on contrastive learning pre-training
By combining comparative learning pre-training technology and recurrent neural networks, a Chinese remote sensing image data set is constructed, which solves the shortcomings in the feature extraction and description generation of remote sensing image in the prior art, and achieves high-accuracy Chinese remote sensing image description generation.
Patent Information
- Application Number
- CN202311132295.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-04
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-09-04
AI Technical Summary
When the existing remote sensing image description generation method uses ImageNet pre-trained convolutional neural network to extract features, the algorithm performance is degraded due to the inconsistent distribution of natural images and remote sensing image data, and the Chinese description cannot be generated, which limits domestic applications.
A Chinese remote sensing image data set is constructed using a method based on contrast learning pre-training. The query encoder and key encoder are obtained through contrast learning pre-training, and robust visual features are extracted, and a recurrent neural network is used to establish a mapping between image features and language features to generate a Chinese description.
It effectively improves the model's understanding and extraction of remote sensing image features, improves the accuracy of description generation, solves the inductive bias problem, and realizes Chinese description generation, which is suitable for domestic applications.
Smart Images

Figure CN117173418B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method for generating descriptions of remote sensing images. Background Art
[0002] Generating descriptions of remote sensing images is a method that uses computer vision and natural language processing technologies to automatically generate text descriptions for remote sensing images. Remote sensing images are image data of the Earth's surface obtained by satellites, drones, or other sensors. In many fields, such as environmental monitoring, agriculture, urban planning, etc., remote sensing images provide crucial information, but processing large amounts of image data and understanding the content therein remains a challenge. Traditionally, the interpretation and analysis of remote sensing images need to be carried out manually by professionals, which is a time-consuming and laborious task. With the progress of computer vision and natural language processing, researchers have begun to explore how to use machine learning and deep learning technologies to automatically generate text descriptions for remote sensing images. By using the technology of generating descriptions of remote sensing images, it is possible to automatically interpret and understand remote sensing images, thereby accelerating the analysis and monitoring process of the Earth's surface. This technology helps professionals in various fields to more effectively utilize remote sensing data and provides important support for decision-making. However, there are still some challenges, such as the improvement of description generation related to the complex semantic information in remote sensing images and the generalization ability of the model, which require further research and exploration.
[0003] The remote sensing image description generation task not only requires the algorithm to understand elements, relationships, positions, scenes, etc. in the remote sensing image, but also requires the algorithm to describe them in natural language that conforms to human grammar rules. A common approach is to pair the remote sensing image with a large-scale labeled remote sensing image dataset and use a convolutional neural network (CNN) to extract image features. These features can include information such as color, texture, object shape, etc. Then, a recurrent neural network (RNN) or its variants (such as long short-term memory network - LSTM) is used to generate a text description related to the image. Zhang et al. introduced an attribute attention mechanism in the literature "ZHANG X, WANG X, TANG X, et al. Description generation for remote sensing images using attribute attention mechanism[J]. Remote Sensing, 2019, 11(6):612." to guide the attention calculation in the decoding process using features from different CNN layers. Lu et al. proposed a novel sound active attention framework for more specific description generation in the literature "LU X, WANG B, ZHENG X. Sound active attention framework for remote sensing image captioning[J]. IEEE Transactions on Geoscience and Remote Sensing, 2019, 58(3):1985–2000." by capturing the objects of interest to the interpreter to generate sentences. Wang et al. proposed a GLCM (Global–local captioning model) for remote sensing image captioning in the literature "WANG Q, HUANG W, ZHANG X, et al. GLCM: Global–local captioning model for remote sensing image captioning[J]. IEEE Transactions on Cybernetics, 2022." based on the attention mechanism, making full use of global and local features to generate more accurate descriptions.
[0004] However, the above remote sensing image description generation methods all use convolutional neural networks pre-trained on ImageNet to extract visual features. Due to problems such as inconsistent data distributions and large differences in image elements between natural image data and remote sensing image data, using convolutional neural networks pre-trained on ImageNet to extract remote sensing image features is not a suitable choice, which will introduce inductive bias and lead to a decline in algorithm performance. In addition, most of the datasets for remote sensing image description generation tasks are described in English, resulting in existing algorithms being unable to generate Chinese descriptions, which is not conducive to the application of remote sensing image description generation algorithms in China. Summary of the Invention
[0005] In order to overcome the deficiencies of the prior art, the present invention provides a method for generating remote sensing image descriptions based on contrastive learning pre-training. First, a Chinese dataset is constructed, then contrastive learning pre-training is carried out to obtain a query encoder and a key encoder, followed by visual feature extraction. Finally, using the extracted visual features, a mapping between image features and language features is established through a recurrent neural network and decoded; the method of the present invention effectively improves the model's understanding and extraction of remote sensing image features, and further improves the accuracy of the model's generation of remote sensing image descriptions.
[0006] The technical solution adopted by the present invention to solve its technical problems includes the following steps:
[0007] Step 1: Construction of the Chinese dataset;
[0008] Obtain a publicly available remote sensing image dataset with English text descriptions, use a machine translation algorithm to translate it into Chinese descriptions, and calibrate the translated Chinese descriptions to eliminate the incorrect text introduced by the machine translation algorithm.
[0009] Step 2: Contrastive learning pre-training;
[0010] Step 2-1: Construct an unlabeled multi-modal, multi-season remote sensing image dataset as the pre-training dataset D, which contains N samples, and each sample is represented as x; Γ represents a set of random image enhancement operations, including operations such as random rotation, random scaling, random cropping, random flipping, and Gaussian blur of the image;
[0011] Step 2-2: For each input sample x i , apply two different image enhancement operations randomly sampled from Γ to x i to obtain two positive sample pairs q and k with enhanced perspectives + ; after obtaining the positive sample pairs with different enhanced perspectives, construct a feature extraction network E;
[0012] Step 2-3: Based on the MoCo framework, use the query encoder E(·,θ q ) and the key encoder E(·,θk ), the result of one perspective enhancement is passed through the query encoder, and the enhancement result of the other perspective is passed through the key encoder, thereby obtaining two representation vectors f in the contrast representation space respectively. q and The process is expressed as follows:
[0013] f q =E(q,θ q )
[0014]
[0015] Among them, the parameters of the key encoder are consistent with those of the query encoder during initialization;
[0016] Step 2-4: During the training process, the query encoder uses a gradient update strategy, and the key encoder uses an exponential average momentum update to make the parameters of the key encoder close to the query encoder: θ k : = mθ k +(1-m)θ q , where m is the momentum parameter, which is a hyperparameter close to 1; the vector representation obtained after encoding is stored in a dynamic dictionary. When an iteration is completed, the representation vector at the head of the team is dequeued and the representation vector obtained by the encoder at the current moment is Join the team;
[0017] In the dynamic dictionary, the representation obtained by the key encoder from the view obtained from other sample enhancements constitutes the negative sample j≠i; project the data into the contrast representation space through the feature extraction network E, so that the distance between positive samples is closer to the distance between negative samples;
[0018] We use the following formula as the contrast loss function to get closer to f q and distance, pull away f q and Distance:
[0019]
[0020] Among them, τ is a temperature hyperparameter used to control the similarity between sample pairs; the similarity between sample pairs is calculated using the vector inner diameter;
[0021] Step 3: Visual feature extraction;
[0022] Initialize the visual feature extractor G in the current model with the parameters of the query encoder and key encoder pre-trained in step 2, and freeze the parameters of the first three stages so that they do not participate in downstream task training; g , using the attention mechanism to focus on and enhance the semantic features to obtain fa ; Specifically as follows:
[0023] Given an input image input ∈ C×H×W, use the pre-trained extractor G to obtain the visual feature vector f g :
[0024] f g = G(input, θ trainable , θ frozen )
[0025] where θ trainable represents the parameter to be updated by the gradient descent algorithm, and θ frozen represents the fixed parameter that will not be updated;
[0026] G uses a ResNet network structure composed of multiple repeated residual blocks. The calculation process of the residual structure is as follows:
[0027] z (l) = x (l-1) + F(x (l-1) )
[0028] where x (l-1) represents the output of the (l - 1)-th layer, and F(·) represents a non-linear function composed of a convolutional layer - batch normalization - activation function;
[0029] Flatten the feature vector f g ∈ c×h×w and perform positional embedding to convert it to f s ∈ c×hw. Apply the attention mechanism to enhance the flattened feature vector f s to obtain the visual feature vector f a :
[0030] MH(x) = Concat(h1, h2, …, h n )W O
[0031]
[0032] where Concat represents the tensor concatenation operation, represent the learnable linear transformation matrices corresponding to the query, key, and value respectively, and W O represents the transformation matrix after merging all attention heads; h1, h2, …, h n represent different attention heads;
[0033] Step 4: Natural language decoding;
[0034] Using the extracted visual features, establish a mapping between image features and language features through a recurrent neural network and decode; in the decoding stage, first input the image features into the initial state of the decoder, and then at each time step, the decoder generates a probability distribution to represent the next generated word or word sequence at the current time step; the generation process is iteratively performed until a special end symbol is generated or a preset maximum generation length is reached.
[0035] Further, the specific steps of step 4 are as follows:
[0036] Use the Transformer decoder to decode the visual feature vector f a into the natural language description of the remote sensing image;
[0037] For a given target description sequence S = [s1, s2, …, s T , first perform word embedding representation and positional encoding on the input sequence to obtain word vector representations; then perform masked self-attention calculation on the sequence to obtain sequence feature information h′:
[0038]
[0039] h′ = FFN(LN(h′ + WeS))
[0040] where, W e represents a learnable word vector encoding matrix, and W q , W k , W v represent the weight matrices corresponding to query, key, and value respectively, mask represents a mask matrix, which sets the word vectors after the t-th moment to zero; subsequently, perform cross-modal multi-head attention calculation on the sequence feature information h and the visual feature vector f a to obtain semantic attention feature f c , and finally input it into a fully connected layer for word prediction:
[0041] p t = softmax(Wf c + b)
[0042] where, p i,j represents the probability of the i-th word generated by the model at the j-th moment, and the index corresponding to the maximum value is taken to obtain the word at the j-th moment; W represents the weight matrix of the fully connected layer, and b represents the bias term.
[0043] Further, train the model using cross-entropy loss:
[0044]
[0045] where, T represents the length of the generated description, d vocabDenotes the vocabulary size, s i,j Denotes the label of the j-th word at the i-th time step of the target description, 0 or 1.
[0046] Preferably, the remote sensing image dataset is SSL4EO-S12.
[0047] The beneficial effects of the present invention are as follows:
[0048] The remote sensing image description generation method proposed by the present invention makes full use of a large amount of unlabeled remote sensing data, effectively improves the model's understanding and extraction of remote sensing image features, and further improves the accuracy of the model's generation of remote sensing image descriptions. Using contrastive learning to pre-train the visual feature extraction network avoids the problems of high cost and long time-consuming for manual annotation of large-scale datasets. Pre-training on a large-scale multi-modal and multi-season remote sensing image dataset can enable the network to learn more robust and unbiased image representations, enhance the distinctiveness and generalization of features, and thus improve the accuracy of the generated image descriptions. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is the overall structure diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0050] The present invention will be further described below in conjunction with the drawings and embodiments.
[0051] The present invention discloses a remote sensing image description generation method based on contrastive learning pre-training. The method aims to solve the inductive bias problem introduced by natural image pre-training through contrastive learning pre-training technology, thereby improving the accuracy of remote sensing image description generation. At the same time, by constructing a Chinese remote sensing image description dataset, the model can generate Chinese descriptions to meet the actual application requirements.
[0052] Its technical solution includes the following steps:
[0053] 1. Construction of the Chinese dataset. Obtain a publicly available remote sensing image dataset with English text descriptions, use a machine translation algorithm to translate it into Chinese descriptions, perform manual calibration on the translated Chinese descriptions to eliminate the incorrect text introduced by the machine translation algorithm, and further enrich the description content according to the image content.
[0054] 2. Contrastive learning pre-training. Use Google Earth Engine to crawl and construct a large-scale unlabeled multi-modal, multi-season remote sensing image dataset SSL4EO-S12. Through contrastive learning, by constructing positive and negative sample pairs, utilize the large-scale unlabeled remote sensing image dataset. It maximizes the similarity of positive sample pairs and minimizes the similarity of negative sample pairs to learn robust feature representations, enabling the network to capture semantic information in remote sensing images, enhancing the distinctiveness and discriminability of features. Secondly, the contrast of multi-modal data allows the network to learn shared feature representations of multiple modalities, providing a more comprehensive and integrated feature expression. In addition, by training with images of different seasons as positive sample pairs in the pre-training stage, the network extracts consistent features from multi-season images, enhancing seasonal adaptability.
[0055] 3. Visual feature extraction. Visual feature extraction is a process of extracting features from the input remote sensing images based on the pre-trained network. Through the robust feature representations learned by the pre-trained network, further process and refine these features in the visual feature extraction stage to obtain more semantically informative and discriminative feature representations. Use the pre-trained convolutional neural network as the backbone network for feature extraction. Capture low-level features such as texture, edges, and shapes of the image through convolutional layers, and use pooling operations to reduce the dimension of features and retain the main information. In addition, introduce an attention mechanism to enhance the features of important regions and strengthen their representation signals. Introduce residual connections to solve the problems of gradient vanishing and explosion in the training process of deep networks, and accelerate the training and convergence process of the network.
[0056] 4. Natural language decoding. Use the visual features extracted in the above step 3 to establish a mapping between image features and language features through a recurrent neural network and decode. In the decoding stage, first input the image features into the initial state of the decoder. Subsequently, at each time step, the decoder generates a probability distribution to represent the next word or word sequence generated at the current time step. The generation process is iterated until a special end symbol is generated or the preset maximum generation length is reached.
[0057] Embodiment:
[0058] The present invention proposes a method for generating remote sensing image descriptions based on contrastive learning pre-training. Referring to Figure 1 , the specific implementation steps of the present invention are as follows:
[0059] Step 1, construction of Chinese dataset.
[0060] Collect publicly available English description remote sensing image datasets, and use machine translation software to translate the English descriptions corresponding to each image into Chinese descriptions to obtain an initial Chinese description dataset. Then, manually correct the translated Chinese descriptions. Specifically, when annotating text for each image, follow the following strategies and steps: 1) Clearly identify the ground object elements that the task focuses on, and ensure that these ground object elements are described when annotating; 2) The annotation text should use easy-to-understand and concise language so that people can quickly understand the image content; 3) Avoid overly subjective or ambiguous language, and ensure the use of accurate, unified, and standardized vocabulary; 4) A description text should contain at least six words; 5) Encourage the use of novel vocabulary and flexible expressions.
[0061] Step 2, contrastive learning pre-training.
[0062] Given a pre-training dataset D, containing N samples, each sample is represented as x, and τ represents a set of random image augmentation operations, including operations such as random rotation, random scaling, random cropping, random flipping, Gaussian blur, etc. For each input sample x i , for x i Apply two different image augmentation operations randomly sampled from τ to obtain positive sample pairs q and k with two augmented perspectives + . After obtaining positive sample pairs with different augmented perspectives, construct a feature extraction network E. Specifically, based on the MoCo framework, this invention uses a query encoder E(·, θ q ) and a key encoder E(·, θ k ), pass the result of one perspective augmentation through the query encoder, and the result of the other perspective augmentation through the key encoder, and then obtain two representation vectors f q and in the contrastive representation space respectively. The process can be expressed in the following form:
[0063] f q = E(q, θ q )
[0064]
[0065] Among them, the parameters of the key encoder are the same as those of the query encoder during initialization. During the training process, the query encoder adopts a normal gradient update strategy, while the key encoder adopts an exponential average momentum update, making the parameters of the key encoder slowly approach those of the query encoder: θ k := mθ k + (1 - m)θ q, where m is the momentum parameter, which is a hyperparameter very close to 1. The vector representation obtained after encoding is stored in a dynamic dictionary. When an iteration is completed, the representation vector at the head of the queue is dequeued and the representation vector obtained by the encoder at the current moment is In the dynamic dictionary, the negative samples are represented by the key encoder from the view obtained by augmenting other samples. j≠i. The data is projected into the contrast representation space through the feature extraction network E, so that the distance between positive samples is relatively close and the distance between negative samples is relatively far. Specifically, the following formula is used as the contrast loss function to bring f closer q and distance, pull away f q and Distance:
[0066]
[0067] Among them, τ is a temperature hyperparameter used to control the similarity between sample pairs. By adjusting the τ value, the model can control whether to focus more on positive sample pairs or negative sample pairs during training. The similarity between sample pairs is calculated using the vector inner diameter.
[0068] Step 3: Visual feature extraction.
[0069] The encoder parameters pre-trained in step 2 are used to initialize the visual feature extractor G in the current model. At the same time, the parameters of the first three stages are frozen and do not participate in the downstream task training to maintain the extractor's ability to extract low-level features. The feature f extracted by G is g , using the attention mechanism to focus on and enhance the semantic features to obtain f a .
[0070] Specifically, given an input image input∈C×H×W, the pre-trained extractor G is used to obtain the visual feature vector f g :
[0071] f g =(input,θ trainable ,θ frozen )
[0072] Among them, θ trainable represents the parameters that will be updated by the gradient descent algorithm, θ frozen In the present invention, G uses a ResNet network structure composed of multiple repeated residual blocks, and the residual structure calculation process is as follows: (l) =x (l-1) +F(x (l-1) ), where x (l-1)denotes the output of the (l-1)-th layer, and F(·) denotes a non-linear function composed of a convolutional layer - batch normalization - activation function.
[0073] Next, the feature vector f g ∈ c×h×w is flattened and position-embedded to be transformed into f s ∈ c×hw. The attention mechanism is applied to the flattened feature vector f s to enhance it and obtain the visual feature vector f a :
[0074] MH(x) = Concat(h1, h2, …, h n )W O
[0075]
[0076] where Concat represents the tensor concatenation operation and W represents the learnable linear transformation matrix.
[0077] Step 4, natural language decoding.
[0078] For the visual feature vector f a extracted in Step 3, the Transformer decoder is used to decode it into the natural language description of the remote sensing image. For a given target description sequence S = [s1, s2, …, s T , first, the input sequence is represented by word embedding and position encoding to obtain the word vector representation. Then, masked self-attention calculation is performed on the sequence to obtain the sequence feature information h:
[0079]
[0080] h = FFN(LN(h + WeS))
[0081] where W e represents the learnable word vector encoding matrix, and W q , W k , W v represent the weight matrices corresponding to the query, key, and value respectively, mask represents the mask matrix, which sets the word vectors after the current time step to zero. Subsequently, the sequence feature information h and the visual feature vector f a are subjected to cross-modal multi-head attention calculation to obtain the semantic attention feature f c , and finally, it is input into a fully connected layer for word prediction:
[0082] p i,j = softmax(Wf c + b)
[0083] where p i,jIt represents the probability of the i-th word generated by the representative model at the j-th moment. The index corresponding to the maximum value is taken to obtain the word at the j-th moment; W represents the weight matrix of the fully connected layer, and b represents the bias term.
[0084] To optimize the model to output the expected description, cross-entropy loss is used to train the model:
[0085]
[0086] The effects of the present invention can be further illustrated by the following experiments.
[0087] 1. Experimental conditions
[0088] The present invention conducts simulation experiments using Python in the system environment of Linux version 5.0.0-23-generic (buildd@lgw01-amd64-030) (gcc version 7.4.0 (Ubuntu 7.4.0-1ubuntu1~18.04.1)) and under the NVIDIA GeForce RTX 3090 graphics card.
[0089] The Chinese dataset in the experiment is constructed based on the publicly available English remote sensing image description dataset NWPU-Captions. For contrastive learning pre-training, the large-scale unlabeled multi-modal and multi-season remote sensing images come from the publicly available SSL4EO-S12 dataset. During the contrastive learning pre-training process, the stochastic gradient descent optimization algorithm is used, with a weight decay of 0.0001, a momentum of 0.9, an initial learning rate of 0.03, and a total of 100 training rounds. During the training process of the remote sensing image description task, the Adam optimizer is used. The initial learning rates of the visual feature extractor and the natural language decoder are 1e-4 and 5e-4 respectively, with a decay rate of 0.8, decaying once every four rounds, and a total of 50 training rounds, and the batch size is 32.
[0090] 2. Experimental content
[0091] The present invention uses BLEU (Bilingual Evaluation Understudy) to evaluate the accuracy of the algorithm description generation. The BLEU metric is based on the n-gram matching method, which compares the n-grams in the model generation result with the n-grams in the reference description, and calculates the number of matching n-grams to evaluate the accuracy and diversity of the algorithm description content generation. The commonly used values of n are 1, 2, 3, and 4. The specific calculation process is as follows:
[0092]
[0093] Among them, the text generated by the model is denoted as candidate, and the reference text provided by the dataset is denoted as reference. BLEU-1, BLEU-2, BLEU-3, and BLEU-4 can be respectively defined according to different values of n.
[0094] The experimental results on the remote sensing image dataset are shown in Table 1.
[0095] Table 1 Experimental Results
[0096]
[0097] The comparison of the BLEU-1, BLEU-2, BLEU-3, and BLEU-4 metrics of the present invention with the show-attend-tell method on the constructed Chinese remote sensing image description dataset verifies the effectiveness of the present invention. Generally speaking, the present invention realizes the generation of Chinese remote sensing image descriptions, improves the ability of traditional methods to extract remote sensing image features, and further improves the accuracy of description generation.
Claims
1. A method for generating remote sensing image descriptions based on contrastive learning pre-training, characterized in that, It includes the following steps: Step 1: Construction of Chinese dataset; Obtain a publicly available remote sensing image dataset with English text descriptions, translate it into Chinese descriptions using a machine translation algorithm, and calibrate the translated Chinese descriptions to eliminate the incorrect text introduced by the machine translation algorithm; Step 2: Contrastive learning pre-training; Step 2-1: Construct an unlabeled multi-modal and multi-season remote sensing image dataset as the pre-training dataset D, which contains N samples, and each sample is represented as x; Γ represents a set of random image enhancement operations, including operations such as random rotation, random scaling, random cropping, random flipping, and Gaussian blur of images; Step 2-2: For each input sample x i , for x i , apply two different image enhancement operations randomly sampled from Γ to obtain positive sample pairs q and k with two enhanced perspectives + ; After obtaining positive sample pairs from different enhanced perspectives, construct a feature extraction network E; Step 2-3: Based on the MoCo framework, use the query encoder E(·, θ q ) and the key encoder E(·, θ k ) respectively. Pass the result of enhancing one perspective through the query encoder and the result of enhancing the other perspective through the key encoder, and then obtain two representation vectors f q and in the contrastive representation space respectively. The process is expressed in the following form: f q = E(q, θ q ) Among them, the parameters of the key encoder are the same as those of the query encoder during initialization; Step 2-4: During the training process, the query encoder adopts a gradient update strategy, and the key encoder adopts an exponential moving average momentum update to make the parameters of the key encoder approach those of the query encoder: θ k := mθ k +(1 - m)θ q , where m is the momentum parameter, a hyperparameter close to 1; the vector representations obtained after encoding are stored in a dynamic dictionary. When an iteration is completed, the representation vector at the head of the queue is dequeued, and the representation vector obtained by the encoder at the current moment is enqueued; In the dynamic dictionary, the representations obtained by the key encoder from the views enhanced from other samples constitute negative samples. j ≠ i; project the data into the contrastive representation space through the feature extraction network E, so that the distance between positive samples is closer than the distance between negative samples. Bring f closer by using the following as the contrastive loss function q and and increase the distance between f q and : Among them, τ is a temperature hyperparameter used to control the similarity between sample pairs; the similarity between sample pairs is calculated using the vector inner diameter; Step 3: Visual feature extraction; Initialize the parameters of the query encoder and key encoder pre-trained in step 2 for the visual feature extractor G in the current model, and freeze the parameters of the first three stages without participating in the downstream task training; for the feature f extracted by G g , use the attention mechanism to focus on and enhance the semantic features to obtain f a ; specifically as follows: Given an input image $input \in \mathbb{C} \times H \times W$, the visual feature vector $f$ is obtained using the pre-trained extractor $G$. g : f g =(input, θ trainable , θ frozen ) Among them, θ trainable represents the parameter to be updated by the gradient descent algorithm, and θ frozen represents the fixed parameter that will not be updated; G uses a ResNet network structure composed of multiple repeated residual blocks, and the calculation process of the residual structure is as follows: z (l) = x (l-1) + F(x (l-1) ) where x (l-1) represents the output of the (l-1)-th layer, and F(·) represents a non-linear function composed of a convolutional layer, batch normalization, and an activation function; Flatten the feature vector f g ∈ c×h×w and perform position embedding to convert it to f s ∈ c×hw, and apply the attention mechanism to the flattened feature vector f s to enhance it and obtain the visual feature vector f a : MH(x) = Concat(h1, h2, …, h n )W O Among them, Concat represents the tensor concatenation operation, respectively represent the learnable linear transformation matrices corresponding to the query, key, and value, W O represents the transformation matrix after merging all attention heads; h1, h2,... h n represent different attention heads; Step 4: Natural language decoding; Using the extracted visual features, establish a mapping between image features and language features through a recurrent neural network and decode; in the decoding stage, first input the image features into the initial state of the decoder, and then at each time step, the decoder will generate a probability distribution to represent the next generated word or word sequence at the current time step; the generation process will be iteratively carried out until a special end symbol is generated or the preset maximum generation length is reached.
2. The method for generating a remote sensing image description based on contrastive learning pre-training according to claim 1, wherein, The specific content of Step 4 is as follows: Use a Transformer decoder to decode the visual feature vector f a into a natural language description of the remote sensing image; For a given target description sequence S = [s1, S2,... s T , first, perform word embedding representation and positional encoding on the input sequence to obtain word vector representations; then, perform masked self-attention calculation on the sequence to obtain sequence feature information h': h′ = FFN(LN(h′ + W e S)) Among them, W e represents a learnable word vector encoding matrix, and W q , W k , and W v represent the weight matrices corresponding to the query, key, and value respectively. Mask represents a mask matrix that zeros out the word vectors after the t-th moment; subsequently, the sequence feature information h and the visual feature vector f a are subjected to cross-modal multi-head attention calculation to obtain the semantic attention feature f c , and finally, it is input into a fully connected layer for word prediction: p t = softmax(Wf c + b) Among them, p i,j represents the probability of the i-th word generated at the j-th moment of the model. The index corresponding to the maximum value is used to obtain the word at the j-th moment; W represents the weight matrix of the fully connected layer, and b represents the bias term.
3. A method for generating remote sensing image descriptions based on contrastive learning pre-training according to claim 1, characterized in that, Train the model using cross-entropy loss: where T represents the length of the generated description, d vocab represents the vocabulary size, s i,j represents the label of the j-th word at the i-th time step of the target description, 0 or 1.
4. A method for generating remote sensing image descriptions based on contrastive learning pre-training according to claim 1, characterized in that, The remote sensing image dataset is SSL4EO-S12.
Citation Information
Patent Citations
Superobject information-based remote sensing image target extraction method, device, electronic apparatus, and medium
WO2020232905A1
News event search method and system based on multi-level image-text semantic alignment model
WO2023093574A1