Remote sensing image description generation method based on auxiliary language enhancement
By building a target language generator and auxiliary language generator, combining high-level and deep semantic features, and using multilingual annotation data, the problem of insufficient context perception ability in remote sensing image description generation is solved, and more accurate text description generation is achieved.
Patent Information
- Application Number
- CN202510455306.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-12
AI Technical Summary
Existing remote sensing image description generation methods have poor perception of the context when generating text descriptions, resulting in low accuracy of text descriptions.
The target language generator and auxiliary language generator are built based on the Transformer model. Through weight interaction and attention mechanisms, high-level semantic features and deep semantic features are combined to extract shared visual features, and multi-language annotation data enriches the model training data, improving the context perception ability of text generators.
It significantly improves the accuracy of remote sensing image description generation, adapts to application needs in different language environments, and improves the language modeling ability and text description accuracy of language generators.
Smart Images

Figure CN120472307A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision processing, and in particular to a remote sensing image description generation method based on auxiliary language enhancement. Background Art
[0002] Image captioning is a cross-disciplinary task involving computer vision, natural language processing, and deep learning. It aims to identify semantic information from images. Image semantic information typically includes objects, their attributes, and their relationships. Textual descriptions are generated by combining computer vision and natural language processing techniques. Remote sensing images refer to geographic information captured by high-altitude platforms such as satellites, aircraft, or drones. Remote sensing image captioning applies these techniques to remote sensing images to generate natural language text describing the image scene, its ground objects, their attributes, and the relationships between them. Remote sensing image captioning provides a method for converting complex geographic information from image modality to textual modality, making it easier for non-experts to understand remote sensing images. This field has great potential in practical applications. It can generate real-time textual descriptions for drone-captured images used in traffic control and rescue scenarios, and can also provide valuable textual references for monitoring surface environmental changes and forest and grassland cover.
[0003] Common remote sensing image caption generation methods follow an encoder-decoder architecture, where the encoder extracts image features and the decoder establishes a relationship between these features and the text to generate a caption. Existing methods lack contextual awareness when generating captions, resulting in poor caption accuracy. Summary of the Invention
[0004] The main purpose of this application is to provide a remote sensing image description generation method based on auxiliary language enhancement, aiming to solve the problem of poor accuracy of remote sensing image text description in existing methods.
[0005] To achieve the above-mentioned objectives, the present application provides a method for generating remote sensing image descriptions based on auxiliary language enhancement, comprising: extracting high-level semantic features and deep semantic features of remote sensing images respectively; extracting inherent visual features based on the high-level semantic features; fusing the deep semantic features and the inherent visual features to obtain shared visual features; inputting the shared visual features into a pre-trained language generator to obtain a text description of the remote sensing image; wherein, a method for constructing the language generator comprises: constructing a target language generator and an auxiliary language generator respectively based on a Transformer model; and performing weighted interaction on the target language generator and the auxiliary language generator to obtain a language generator.
[0006] Optionally, the high-level semantic features include a first high-level semantic feature and a second high-level semantic feature, and the method for extracting the high-level semantic features includes: inputting the remote sensing image into the residual network, using the features output by the penultimate residual block group of the residual network as the first high-level semantic feature, and using the features output by the last residual block group of the residual network as the second high-level semantic feature.
[0007] Optionally, based on high-level semantic features, inherent visual features are extracted, including: performing convolution operations on the first high-level semantic features through the original convolution layer, the center difference convolution layer, the angular difference convolution layer, the horizontal difference convolution layer and the vertical difference convolution layer, and performing vector splicing operations on the results of all convolution operations to obtain spliced features; fusing the spliced features with the second high-level semantic features to obtain first fused features; optimizing the fused features through the pixel attention layer and the channel attention layer, respectively, to obtain first features and second features respectively; fusing the first features and the second features to obtain the second fused features; optimizing the first fused features and the second fused features through the spatial attention layer, and performing convolution operations to obtain inherent visual features.
[0008] Optionally, the inherent visual features are determined as follows:
[0009] x3=Concat(F VC (x1),F CDC (x1),F ADC (x1),F HDC (x1),F VDC (x1))
[0010]
[0011] Where x1 represents the first high-level semantic feature, x2 represents the second high-level semantic feature, x3 represents the concatenated feature, x4 represents the inherent visual feature, VC represents the original convolution, CDC represents the center difference convolution, ADC represents the angle difference convolution, HDC represents the horizontal difference convolution, VDC represents the vertical difference convolution, Concat represents the vector concatenation operation, represents the element-by-element addition of a vector, ⊙ represents the element-by-element multiplication of a vector, σ represents the Sigmod activation function, PA represents pixel attention, SA represents spatial attention, and CA represents channel attention.
[0012] Optionally, weighted interaction is performed on the target language generator and the auxiliary language generator to obtain the language generator, including: performing parameter interaction on each structural block in the target language generator and the corresponding structural block in the auxiliary language generator through a weighted interaction path to obtain parameters of the language generator; wherein the parameter interaction method in the weighted interaction path is to perform weighted summation on the parameters of the target language generator and the auxiliary language generator.
[0013] Optionally, the parameters include Q, K, and V of the attention layer.
[0014] Optionally, the weighted interaction path is:
[0015]
[0016] The parameter interaction mode in the weight interaction path is:
[0017] W T ' LG =W TLG +λW ALG
[0018] In the formula, TLG represents the target language generator, ALG represents the auxiliary language generator, Block represents the basic building block of Transformer, k represents the sequence number of Block, Bridge represents the weighted interaction path, and W TLG Represents the parameter, W ALG Indicates the parameters of ALG, W T ' LG represents the reparameterized parameter, and λ represents the scaling factor of the weight interaction.
[0019] Optionally, the method for extracting deep semantic features includes: extracting features from remote sensing images through a CLIP network to obtain deep semantic features.
[0020] Optionally, the method for acquiring shared visual features includes: fusing deep semantic features and inherent visual features through an attention mechanism to obtain shared visual features.
[0021] To achieve the above-mentioned objectives, the present application also provides a remote sensing image description generation device based on auxiliary language enhancement, comprising: a first feature extraction module, used to extract high-level semantic features and deep semantic features of the remote sensing image respectively; a second feature extraction module, used to extract inherent visual features based on high-level semantic features; a feature fusion module, used to fuse deep semantic features and inherent visual features to obtain shared visual features; and a text generation module, used to input the shared visual features into a pre-trained language generator to obtain a text description of the remote sensing image.
[0022] Compared with the prior art, the present invention has the following advantages:
[0023] The remote sensing image description generation method based on auxiliary language enhancement of the present invention performs weighted interaction on the target language generator and the auxiliary language generator to obtain a language generator. The auxiliary language of the auxiliary language generator can effectively supplement the training data of the target language generator and provide diversified language expressions, thereby improving the text generator's ability to perceive context, significantly enhancing the language modeling ability of the language generator, and further improving the accuracy of the generated language description; introducing an attention mechanism to extract inherent visual features from high-level semantic information as the basis for cross-language visual alignment, it can dynamically guide the model to focus on important visual feature areas (shared visual features) in the image based on the current output (second high-level semantic features) and accumulated empirical knowledge (first high-level semantic features), ensuring that the generated image description is more accurate; according to actual application requirements, the model can flexibly set the target language to Chinese or English, thereby adapting to remote sensing image description tasks in different language environments and meeting actual application requirements in multilingual scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a flowchart of a remote sensing image description generation method based on auxiliary language enhancement in this application;
[0025] Figure 2 This is a schematic diagram of the training process of a remote sensing image description generation method based on auxiliary language enhancement in this application;
[0026] Figure 3 This is a structural diagram of a language-independent feature enhancement module in a remote sensing image description generation method based on auxiliary language enhancement in this application;
[0027] Figure 4 This is a structural diagram of a parameter weight interaction module in a remote sensing image description generation method based on auxiliary language enhancement in this application.
[0028] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0029] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0030] The current method of generating remote sensing image descriptions only uses data sources annotated in English, which leads to deficiencies in the use of multilingual data, and further leads to poor accuracy in the generation of target language descriptions of remote sensing images. In order to solve this problem, the present invention proposes a remote sensing image description generation method based on auxiliary language enhancement. For an image, different languages are used for annotation. Although the text representation is different, they all correspond to the same potential image semantic features. This correspondence can enrich the training data of the model, thereby enhancing the extraction of rich semantic features by the visual feature extractor. The differences in grammatical structure of different languages can provide a variety of sentence construction methods, thereby improving the text generator's ability to perceive context, which has not been fully considered in previous studies. Therefore, the present invention considers annotations in different languages, which not only helps the model utilize rich data types, but also realizes cross-language information sharing and transfer in multilingual learning, thereby improving the text generator's ability to perceive context, and further improving the accuracy of the generation of target language descriptions of remote sensing images.
[0031] The first embodiment of the present invention provides a remote sensing image description generation method based on auxiliary language enhancement, such as Figure 1 As shown, the specific steps include:
[0032] Step S1, extracting high-level semantic features and deep-level semantic features of remote sensing images respectively;
[0033] Among them, the high-level semantic features include first high-level semantic features and second high-level semantic features, such as object categories and scene backgrounds. The specific method for extracting high-level semantic features includes: inputting the remote sensing image into the residual network, using the features output by the penultimate residual block group of the residual network as the first high-level semantic features, and using the features output by the last residual block group of the residual network as the second high-level semantic features.
[0034] For example, the residual network can be ResNet-152, which consists of 152 convolution layers, batch normalization, ReLU activation functions, and fully connected layers. Given a remote sensing image I∈C×H×W, ResNet-152 is used to extract high-level semantic features of the remote sensing image, including the visual feature outputs x1, x2 of the third and fourth layers, where x1, x2 = CNN(I). At the same time, the CLIP network is used to extract features from the remote sensing image to obtain deep semantic features, x clip = CLIP(I), which can help understand complex semantic relationships. The CLIP network is pre-trained using the ViT-L / 14 backbone network, a large-scale version of the Visual Transformer (ViT) model. It uses a 14×14 image tile size and a self-attention mechanism to embed images and text into a common semantic space.
[0035] Step S2: extracting inherent visual features, such as image shape, texture, and color, based on high-level semantic features. Specifically, the inherent visual features can be extracted based on high-level semantic features through a language-independent feature enhancement module, as follows.
[0036] Step S21, performing convolution operations on the first high-level semantic features through the original convolution layer, the center difference convolution layer, the angular difference convolution layer, the horizontal difference convolution layer, and the vertical difference convolution layer, and performing vector splicing operations on the results of all the convolution operations to obtain spliced features;
[0037] x3=Concat(F VC (x1),F CDC (x1),F ADC (x1),F HDC (x1),F VDC (x1))
[0038] Where x1 represents the first high-level semantic feature, x2 represents the second high-level semantic feature, x3 represents the concatenated feature, VC represents the original convolution, CDC represents the center difference convolution, ADC represents the angle difference convolution, HDC represents the horizontal difference convolution, and VDC represents the vertical difference convolution.
[0039] Step S22, fusing the splicing feature with the second high-level semantic feature to obtain a first fused feature;
[0040]
[0041] Where x2 represents the second high-level semantic feature, represents element-wise addition of vectors;
[0042] Step S23, feature optimization is performed on the fused features through the pixel attention layer and the channel attention layer respectively, to obtain the first feature and the second feature respectively;
[0043] Step S24, fusing the first feature and the second feature to obtain a second fused feature;
[0044] In step S25, the first fusion feature and the second fusion feature are optimized through the spatial attention layer, and a convolution operation is performed to obtain inherent visual features.
[0045] Specifically, the attention weight distribution of the first fusion feature is obtained by performing feature optimization on the first fusion feature and the second fusion feature through the spatial attention layer;
[0046]
[0047] Through the convolution layer, the attention weight distribution of the first fusion feature is used to perform a convolution operation on the second high-level semantic feature, the splicing feature and the first fusion feature to obtain the inherent visual feature.
[0048]
[0049] Where x4 represents the inherent visual features, Concat represents the vector concatenation operation, ⊙ represents the element-wise multiplication of the vector, σ represents the Sigmod activation function, PA represents pixel attention, SA represents spatial attention, and CA represents channel attention.
[0050] In this embodiment, inherent visual features such as image shape, texture, and color are extracted as the basis for cross-language visual alignment, providing a visual basis for the training of the target language generator and the auxiliary language generator.
[0051] Step S3, fusing the deep semantic features and the inherent visual features to obtain shared visual features;
[0052] Specifically, the deep semantic features and inherent visual features are fused through the attention mechanism to obtain shared visual features. The specific calculation process is as follows:
[0053] x fusion =Concat(SelfAttention(x4),x clip )
[0054] x share =x fusion ☉σ(x fusion )
[0055] Where x clip represents deep semantic features, x fusion Represents x clip The fusion feature of x and x4 after fusion through the self-attention mechanism, x share Represents shared visual features.
[0056] In step S4, the shared visual features are input into a pre-trained language generator to obtain a text description of the remote sensing image. The method for constructing the language generator includes: constructing a target language generator and an auxiliary language generator based on the Transformer model respectively; and performing weighted interaction on the target language generator and the auxiliary language generator to obtain the language generator.
[0057] Specifically, the language generator generation method involves performing parameter interaction between each structural block in the target language generator and the corresponding structural block in the auxiliary language generator through a parameter-heavy interaction module, also known as a weighted interaction pathway, to obtain the language generator's parameters. These parameters include Q, K, and V of the attention layer. The parameter interaction in the weighted interaction pathway involves a weighted summation of the parameters of the target and auxiliary language generators. The specific calculation method is as follows.
[0058] The weighted interaction pathway is:
[0059]
[0060] The parameter interaction formula in the weighted interaction path is:
[0061] W T ' LG =W TLG +λW ALG
[0062] In the formula, TLG represents the target language generator, ALG represents the auxiliary language generator, Block represents the basic building block of Transformer, k represents the sequence number of Block, Bridge represents the weighted interaction path, and W TLG Represents the parameter, W ALG Indicates the parameters of ALG, W T ' LG represents the reparameterized parameter, and λ represents the scaling factor of the weight interaction.
[0063] It is worth noting that both the target language generator and the auxiliary language generator include a visual encoder and a text decoder. The visual encoder extracts visual representations that match a specific language, while the text decoder converts these visual representations into fluent and coherent text. By sharing the mapping relationship between visual and text features and performing decoding, an attention mechanism is introduced. The initial language generation models of the target language generator and the auxiliary language generator are identical. Before training begins, the target language and auxiliary language are determined. Images labeled with the target language and auxiliary language are fed into the two initial language generation models, respectively, to generate the target language generator and auxiliary language generator. When training the language generators, the shared visual features are fed as queries into the pre-trained target language generator and pre-trained auxiliary language generator, respectively. The experience and knowledge of the corresponding generators are used to guide attention allocation and prioritize key information. A probability distribution is generated at each time step to predict the next word or word sequence. This generation process continues iteratively until an end symbol is generated or the maximum generation length is reached. Corresponding training is performed to obtain a pre-trained target language generator and a pre-trained auxiliary language generator; the parameters of the pre-trained target language generator and the pre-trained auxiliary language generator are input into a heavy parameter weight interaction module to perform parameter interaction to obtain a language generator.
[0064] The following cross entropy loss function is minimized during the training process:
[0065]
[0066] Where, L XE (θ) represents the cross entropy loss function, θ represents the parameters of the model, T represents the length of the sequence, The output at time t is represents all outputs from time 1 to time t-1, log represents the natural logarithm function, represents the conditional probability distribution.
[0067] In this embodiment, a parameter sharing mechanism is used to promote information flow and collaborative optimization between the target language generator and the auxiliary language generator; the language structure perception ability of the auxiliary language generator can be transferred to the target language generator, thereby enhancing the language understanding and generation ability of the target language generator and improving model performance.
[0068] Example 1
[0069] This experiment was carried out using Python in the Linux version 5.0.0-23-generic (buildd@lgw01-amd64-030) (gccversion 7.4.0 (Ubuntu 7.4.0-1ubuntu1~18.04.1)) system environment and NVIDIA GeForce RTX3090 graphics card.
[0070] The experimental dataset used is the publicly available bilingual remote sensing image description dataset (BRSC-RSICD) in Chinese and English. This dataset contains 10,921 remote sensing images with a resolution of 224×224. Each image is accompanied by five English and Chinese descriptions. The data is split into training, validation, and test sets in an 8:1:1 ratio. The model was initialized with a learning rate of 4e-4 and trained using the Adam optimizer for 30 epochs with a batch size of 32.
[0071] This experiment uses BLEU (Bilingual Evaluation Understudy) to evaluate the accuracy of the algorithm's description generation. The core of this algorithm is to calculate the accuracy of sentences by comparing the number of evaluation sentences and reference sentences (N-grams), and combined with the BP penalty factor, the calculation method is as follows:
[0072]
[0073] Among them, P N is the precision of each order of N-gram, h k (c i ) indicates that the kth N-gram appears in the sentence c to be evaluated i The number of times in Indicates the number of times the kth N-gram appears in the standard reference sentence, h k (s ij ) represents the number of times a certain N-gram appears in multiple standard reference sentences, w n is the precision weight of each N-gram, l c is the length c of the sentence to be evaluated i , l s is the length of the annotation sentence.
[0074] BLEU-1, BLEU-2, BLEU-3, and BLEU-4 can be defined based on N. The experimental results on the test set are as follows. For the generation results with English as the target language, refer to Table 1. For the generation results with Chinese as the target language, refer to Table 2.
[0075] Table 1 Experimental results in English
[0076]
[0077] Table 2 Chinese experimental results
[0078]
[0079] As can be seen from the results in Tables 1 and 2, Example 1 outperforms the MG method in BLEU-1, BLEU-2, BLEU-3, and BLEU-4 scores. In particular, when Chinese is used as the target language, the evaluation results of the Chinese description are significantly improved, further verifying the effectiveness of the present invention.
[0080] In general, the present invention implements a remote sensing image description generation technology based on auxiliary language enhancement, improves the language modeling capability of the model, and thus enhances the accuracy of description generation.
[0081] A second embodiment of the present invention provides a remote sensing image description generation device based on auxiliary language enhancement, including: a first feature extraction module, used to extract high-level semantic features and deep semantic features of the remote sensing image respectively; a second feature extraction module, used to extract inherent visual features based on the high-level semantic features; a feature fusion module, used to fuse the deep semantic features and the inherent visual features to obtain shared visual features; and a text generation module, used to input the shared visual features into a pre-trained language generator to obtain a text description of the remote sensing image.
[0082] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A remote sensing image description generation method based on auxiliary language enhancement, characterized in that: include: Extract high-level semantic features and deep semantic features of remote sensing images respectively; Extracting inherent visual features based on the high-level semantic features; fusing the deep semantic features and the inherent visual features to obtain shared visual features; Inputting the shared visual features into a pre-trained language generator to obtain a text description of the remote sensing image; The method for constructing the language generator includes: Build the target language generator and auxiliary language generator based on the Transformer model; The target language generator and the auxiliary language generator are weighted to interact with each other to obtain a language generator.
2. The remote sensing image description generation method based on auxiliary language enhancement according to claim 1, characterized in that: The high-level semantic features include a first high-level semantic feature and a second high-level semantic feature, and the method for extracting the high-level semantic features includes: The remote sensing image is input into a residual network, the feature output by the penultimate residual block group of the residual network is used as the first high-level semantic feature, and the feature output by the last residual block group of the residual network is used as the second high-level semantic feature.
3. The remote sensing image description generation method based on auxiliary language enhancement according to claim 2, characterized in that: The extracting of inherent visual features based on the high-level semantic features includes: Performing convolution operations on the first high-level semantic features through the original convolution layer, the center difference convolution layer, the angular difference convolution layer, the horizontal difference convolution layer, and the vertical difference convolution layer, respectively, and performing vector splicing operations on the results of all the convolution operations to obtain spliced features; Fusing the splicing feature with the second high-level semantic feature to obtain a first fused feature; Feature optimization is performed on the fused features through a pixel attention layer and a channel attention layer respectively, to obtain a first feature and a second feature respectively; Fusing the first feature and the second feature to obtain a second fused feature; The first fusion feature and the second fusion feature are optimized through the spatial attention layer, and a convolution operation is performed to obtain inherent visual features.
4. The remote sensing image description generation method based on auxiliary language enhancement according to claim 3 is characterized in that: The inherent visual features are determined as follows: x3=Concat(F VC (x1),F CDC (x1),F ADC (x1),F HDC (x1),F VDC (x1)) Where x1 represents the first high-level semantic feature, x2 represents the second high-level semantic feature, x3 represents the concatenated feature, x4 represents the inherent visual feature, VC represents the original convolution, CDC represents the center difference convolution, ADC represents the angle difference convolution, HDC represents the horizontal difference convolution, VDC represents the vertical difference convolution, Concat represents the vector concatenation operation, represents the element-by-element addition of a vector, ⊙ represents the element-by-element multiplication of a vector, σ represents the Sigmod activation function, PA represents pixel attention, SA represents spatial attention, and CA represents channel attention.
5. The remote sensing image description generation method based on auxiliary language enhancement according to claim 1, characterized in that: The step of performing weighted interaction on the target language generator and the auxiliary language generator to obtain a language generator includes: Through the weighted interaction path, each structural block in the target language generator and the corresponding structural block in the auxiliary language generator are respectively subjected to parameter interaction to obtain the parameters of the language generator; The parameter interaction mode in the weighted interaction path is to perform weighted summation on the parameters of the target language generator and the auxiliary language generator.
6. The remote sensing image description generation method based on auxiliary language enhancement according to claim 5 is characterized in that: The parameters include Q, K and V of the attention layer.
7. The remote sensing image description generation method based on auxiliary language enhancement according to claim 5 is characterized in that: The weighted interaction path is: The parameter interaction mode in the weight interaction path is: IN' TLG =In TLG +λW ALG In the formula, TLG represents the target language generator, ALG represents the auxiliary language generator, Block represents the basic building block of Transformer, k represents the sequence number of Block, Bridge represents the weighted interaction path, and W TLG Represents the parameter, W ALG represents the parameters of ALG, W′ TLG represents the reparameterized parameter, and λ represents the scaling factor of the weight interaction.
8. The remote sensing image description generation method based on auxiliary language enhancement according to claim 1 is characterized in that: The method for extracting the deep semantic features includes: The CLIP network is used to extract features from remote sensing images and obtain deep semantic features.
9. The remote sensing image description generation method based on auxiliary language enhancement according to claim 1, characterized in that: The method for acquiring the shared visual features includes: The deep semantic features and inherent visual features are fused through the attention mechanism to obtain shared visual features.
10. A remote sensing image description generation device based on auxiliary language enhancement, characterized in that: include: The first feature extraction module is used to extract high-level semantic features and deep-level semantic features of the remote sensing image respectively; A second feature extraction module is used to extract inherent visual features based on the high-level semantic features; A feature fusion module, configured to fuse the deep semantic features and inherent visual features to obtain shared visual features; The text generation module is used to input the shared visual features into a pre-trained language generator to obtain a text description of the remote sensing image.