Remote sensing image description method based on multi-modal feature fusion

Through the combination of multimodal feature fusion and contrast learning mechanisms, features are extracted in combination with ResNet-152 and YOLOv8, and feature optimization is used to use CLIP model and graph attention network to perform feature optimization, the problem of insufficient semantics of remote sensing image description method in complex environments is solved, and high-quality remote sensing image description is achieved.

CN119942342APending Publication Date: 2025-05-06GUILIN UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202510058332.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing remote sensing image description methods are difficult to effectively capture multi-scale and diversified information in complex environments, resulting in insufficient semantics and low description quality.

Method used

Multimodal feature fusion method is adopted, scene-level and object-level features are extracted in combination with ResNet-152 and YOLOv8, a comparison learning mechanism and CLIP model are introduced, and feature fusion and optimization are performed through graph attention network and multi-head attention mechanism, and remote sensing image description is generated using a priori knowledge-guided Transformer model.

Benefits of technology

It significantly improves the accuracy and quality of remote sensing image description, enhances the model's robustness to complex scenes and multiple objects, and achieves more comprehensive feature expression and semantic consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942342A_ABST
    Figure CN119942342A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image description method based on multi-modal feature fusion, and the method comprises the three steps: S1, extracting features: extracting scene-level features through ResNet-152, extracting object-level features through YOLOv8, and carrying out the optimization through comparative learning; s2, feature enhancement: extracting features, processing the extracted features through a graph attention network and a multi-head attention mechanism, and performing enhancement in combination with CLIP features; s3, model training: the features are substituted into a transformer, and multiple times of training are carried out to enhance the accuracy; and S4, image description: substituting the remote sensing image and finally generating accurate description. The method is mainly designed for solving the problems that semantic information of multi-scale and multi-type targets is difficult to fully capture in a complex environment, efficient alignment between images and text description is difficult to achieve, and the requirement for target recognition and description precision is continuously improved in an existing remote sensing image description technology. The method aims at solving the problems that in remote sensing image analysis, semantic consistency of diversified scene information, fine-grained target features and text description is insufficient and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and remote sensing image processing, and in particular to remote sensing image analysis and automatic description generation technology. More specifically, the present invention relates to a remote sensing image description method of multimodal feature fusion, which combines a deep learning model with multi-scale feature extraction, contrastive learning, a graph attention network, a multi-head attention mechanism, and a CLIP model to improve the automatic description capability of remote sensing images, and is widely used in the fields of remote sensing image automatic annotation, environmental monitoring, disaster warning, etc. Background Art

[0002] With the continuous development of remote sensing technology, remote sensing images have become an important tool for obtaining geographic information, environmental monitoring, resource surveys and other fields. However, the automatic analysis and understanding of remote sensing images still face many challenges, especially in the generation of image descriptions. Existing remote sensing image description methods usually rely on traditional convolutional neural network (CNN) models to extract image features, but in complex environments, traditional methods are difficult to effectively capture multi-scale and diverse information in remote sensing images. Especially for remote sensing images containing multiple targets and scene information, existing methods often have problems such as insufficient semantics and low description quality.

[0003] In recent years, with the continuous advancement of deep learning technology, especially the introduction of the YOLO series models and the CLIP model, image feature extraction and multimodal feature fusion have become possible. As an advanced target detection model, YOLOv8 has excellent multi-scale target recognition capabilities and can better cope with complex remote sensing image scenes. The CLIP model effectively improves the correlation between images and text by mapping them to a common semantic space. Although these technologies have significantly improved the effect of image description generation, there is still the problem of how to fully integrate different features and optimize the semantic matching between images and text.

[0004] 1. A remote sensing image description method based on multi-scale information extraction and fusion is described in Chinese patent document CN202410476892.6, which uses the resnet50 convolutional neural network to extract image features. A light feature correction factor is introduced in the image feature generation process to reduce the impact of non-ideal light on object recognition. The regularized intermediate feature map and sentence embedding vector are sent to the proposed three-stage improved Transformer module to extract and fuse the multi-scale information of the image and generate an image description.

[0005] 2. A remote sensing image description method based on multi-category label assistance is recorded in Chinese patent document CN202311101768.3. The encoder module uses Resnet50 pre-trained on ImageNet to extract full-text features of the input remote sensing image after preprocessing and deduplication; the auxiliary multi-label classification module is used to obtain the predicted significant target category after multi-classification judgment of the input full-text features of the image through SoftMax; the LSTM decoder module is used to output image description information after learning the mapping relationship between the image features and the syntactic features in the input information.

[0006] 3. A method for generating wide-area remote sensing descriptions based on target detection is described in Chinese patent document CN201911143698.1. First, a remote sensing image is acquired; a training sample set and a test sample set are constructed; the remote sensing image is processed using a Faster-RCNN network model; the target is clustered using a K-means clustering algorithm; the segmented image is processed using a ResNet101 network model; the corresponding image description is obtained using LSTM; and the target detection result is detected again to see if it is in the description, thereby obtaining the final result.

[0007] The above methods all use traditional models such as ResNet for feature extraction. Although they have shown some performance in capturing the global features of remote sensing images, they still have many shortcomings: the feature extraction strategy is single, and only models such as ResNet50 or ResNet101 are used, which cannot effectively capture multi-scale and multi-target features; the combination of global and local information in complex remote sensing scenes is not fully considered, resulting in incomplete semantic expression; in addition, in terms of feature fusion and semantic matching, most of them use simple methods, which makes it difficult to fully utilize multimodal information. These problems lead to the lack of description accuracy and semantic richness of existing methods when facing complex remote sensing images.

[0008] In view of the above shortcomings, the present invention proposes a multimodal feature fusion remote sensing image description method that integrates YOLOv8, ResNet-152, contrastive learning mechanism, graph attention network and CLIP model. In the feature extraction stage, this method combines ResNet-152 to capture global scene information and uses YOLOv8 to extract multi-target local features; in the feature optimization stage, the contrastive learning mechanism is introduced to improve the semantic distinction ability of image features; in the feature fusion stage, the CLIP model is used to enhance the correlation between image and text; at the same time, the graph attention network and the Transformer model guided by prior knowledge are used to further enhance the interactive expression ability of contextual information. Through the collaborative innovation of the above technologies, the present invention significantly improves the accuracy and quality of remote sensing image description, and provides an efficient and comprehensive solution for the automated analysis and application of remote sensing images. Summary of the invention

[0009] The object of the present invention is to provide a remote sensing image description method of multimodal feature fusion, which specifically comprises the following steps:

[0010] Step S1. Feature extraction: using a multimodal feature extraction module, this module is used to extract scene-level features and object-level features;

[0011] Step S2. Feature enhancement: using the feature enhancement module, further enhancing the features generated by the multimodal feature extraction module to improve its effectiveness in the remote sensing image description task;

[0012] Step S3. Model training: Use the Transformer guided by prior knowledge, integrate domain prior knowledge into the Transformer, model the enhanced features and start training on the remote sensing dataset;

[0013] Step S4. Image description: Use the model with the best training effect and substitute it into the remote sensing image to generate a remote sensing image description.

[0014] The feature extraction steps of step S1 are as follows:

[0015] Step S11: ResNet-152 network is used to extract scene-level features from remote sensing images. ResNet-152 is a deep convolutional neural network that introduces residual connections, which enables the network to train deeper models and effectively alleviate the gradient vanishing problem, thereby significantly improving the feature extraction capability.

[0016] (1)

[0017] Where I is the input remote sensing image, is the extracted scene-level feature.

[0018] Step S12: Using the YOLOv8 network to extract object-level features from the remote sensing image. YOLOv8 integrates multi-scale feature grouping and fusion mechanisms based on traditional target detection methods, and can more accurately capture the semantics and edge information of the target;

[0019] (2)

[0020] Where I is the input remote sensing image, , is the number of objects detected, is the dimension of each object feature.

[0021] Step S13: Contrastive learning enhancement: For the extracted scene-level features and object-level features, independent contrastive learning modules are designed to enhance them respectively;

[0022] Step S14: Comparative learning module design, constructing positive and negative sample pairs, and optimizing feature representation by utilizing scene-level or object-level feature similarities or differences;

[0023] Step S15: By comparing the loss function, the model generates highly correlated representations between semantically similar features, while reducing the interference of irrelevant features, thereby improving the distinguishing ability of features;

[0024] (3)

[0025] in, , are the embedding vectors of object-level features and scene-level features, respectively, sim(.,.) is the similarity function (e.g. cosine similarity), and T is the temperature hyperparameter.

[0026] The feature enhancement steps of step S2 are as follows:

[0027] Step S21: using the global features extracted by the CLIP model, fusing them with the multimodal features, and enhancing the richness and consistency of feature representation through semantic alignment and spatial association;

[0028] (4)

[0029] in, , is the dimension of the CLIP feature.

[0030] (5)

[0031] (6)

[0032] in, and are object-level and scene-level features, respectively. is the global feature extracted by CLIP, , , is the weight coefficient, which is used to balance the contribution of each feature.

[0033] Step S22: construct an attention mechanism based on the graph structure to further optimize the feature representation by modeling the contextual relationship between the scene and the object;

[0034] (7)

[0035] (8)

[0036] (9)

[0037] in, and is the node feature, a and W are trainable parameters, N(i) is the set of neighbor nodes of node i, is the activation function.

[0038] Step S23: Combine the multi-head attention mechanism to strengthen the information interaction between different feature dimensions and improve the global modeling ability of features;

[0039] (10)

[0040] (11)

[0041] in, , , , is the linear change matrix of each head, is the output linear change matrix.

[0042] The model training steps of step S3 are as follows:

[0043] Step S31: embedding common patterns, structural relationships and semantic rules in the remote sensing field into the attention mechanism of the Transformer to guide the feature interaction process. Domain prior knowledge is integrated into the Transformer through position encoding and semantic embedding to enhance the model's understanding of the unique patterns of remote sensing images;

[0044] Step S32: Utilize the encoder-decoder structure to further strengthen the connection between input features and output descriptions, generate high-quality description sentences that conform to actual semantics, and perform multiple training on remote sensing datasets.

[0045] The image description steps of step S4 are as follows:

[0046] Step S41: Use the model with the best training effect, substitute it into the remote sensing image and finally generate an accurate remote sensing image description. Score the generated description according to different evaluation indicators (such as BLEU, ROUGE, CIDEr) to ensure the accuracy and robustness of the description.

[0047] The present invention has the following beneficial effects and advantages:

[0048] (1) Multimodal feature extraction optimization: This paper combines ResNet-152 and YOLOv8 to extract scene-level and object-level features, and uses multi-scale feature grouping and fusion strategies to capture multi-level information, making the feature expression more comprehensive and accurate.

[0049] (2) Significant contrastive learning enhancement effect: The present invention introduces contrastive learning modules for scene-level and object-level features respectively, which significantly improves the feature discrimination ability and generalization, and enhances the robustness of the model to complex scenes and multiple objects.

[0050] (3) Feature enhancement and fusion capability improvement: Through the effective fusion of graph attention network, multi-head attention mechanism and CLIP features, the present invention greatly improves the semantic consistency and interaction capability between features, providing high-quality input for description generation.

[0051] (4) Effective use of prior knowledge: By embedding domain prior knowledge in Transformer, the present invention enhances the semantic understanding ability of the model. The generated description sentences are more in line with actual application requirements and have higher accuracy and expressiveness.

[0052] (5) Wide application value: The present invention has wide application potential in the fields of remote sensing image interpretation, land use classification, disaster monitoring, etc., and can provide intelligent and efficient solutions for remote sensing image analysis and description. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0054] Figure 1 :A step diagram of a remote sensing image description method of multimodal feature fusion according to the present invention.

[0055] Figure 2 :This is a flow chart of the implementation of the present invention based on the multimodal feature fusion model

[0056] From left to right, the modules are the multi-level feature extraction module, the feature enhancement module, and the prior knowledge-guided Transformer (Transformer forPKA).

[0057] Figure 3 : RSICD (Remote Sensing Image Captioning Dataset) dataset used in this invention DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0059] Example:

[0060] Combination Figure 1 The steps of the remote sensing image description method based on multimodal feature fusion of the present invention are described as follows:

[0061] Dataset preparation and preprocessing

[0062] This example uses the public RSICD (Remote Sensing Image Captioning Dataset) dataset Figure 3 The RSICD dataset contains about 10,921 remote sensing images, with an average of 5 English descriptions per image. The dataset is divided into a training set of 8,000 images, a validation set of 1,000 images, and a test set of 1,921 images. Each image is scaled (256×256) and normalized to fit the input requirements of the pre-trained model.

[0063] Step S1: Feature extraction

[0064] The multimodal feature extraction module (MFE module) is used to extract scene-level features and object-level features from remote sensing images. The specific steps include the following:

[0065] Step S11: scene-level feature extraction

[0066] The ResNet-152 network is used to extract scene-level features from remote sensing images and obtain a 2048-dimensional global feature representation. As a deep convolutional neural network, ResNet-152 can effectively capture the high-level semantic information of images.

[0067] Step S12: Object-level feature extraction

[0068] The YOLOv8 network is used to extract object-level features from remote sensing images. As an advanced target detection model, YOLOv8 has excellent multi-scale target recognition capabilities and can accurately detect 5 to 20 targets in complex remote sensing image scenes, generating a 512-dimensional feature vector for each target.

[0069] Step S2: Feature enhancement

[0070] The feature enhancement module (FE module) is used to further enhance the features generated by the MFE module to improve its effectiveness in remote sensing image description tasks. The specific steps include the following:

[0071] Step S21: CLIP model feature fusion

[0072] The global features (512 dimensions) extracted by the CLIP model are fused with multimodal features. Through semantic alignment and spatial association, the richness and consistency of feature representation are enhanced, and the effective fusion of multimodal features in a shared semantic space is achieved.

[0073] Step S22: Construction of graph structure attention mechanism

[0074] We build an attention mechanism based on graph structure to further optimize feature representation by modeling the contextual relationship between scenes and objects. We use graph attention network (GAT) to model the context of scene-level features.

[0075] Step S23: Multi-head attention mechanism application

[0076] Combined with the multi-head attention mechanism (8 attention heads), the information interaction between different feature dimensions is strengthened, and the global modeling ability of features is improved. The multi-head attention mechanism can capture multiple relationships between features in parallel and improve the richness of feature representation.

[0077] Step S3: Model training

[0078] Use the Transformer guided by prior knowledge to model the enhanced features and start training on the remote sensing dataset. The specific steps include the following:

[0079] Step S31: Prior knowledge embedding

[0080] The common patterns, structural relationships and semantic rules in the field of remote sensing are embedded into the attention mechanism of Transformer to guide the feature interaction process. The semantic understanding ability of the model is enhanced by pre-defined ground object types and their spatial distribution patterns.

[0081] Step S32: Encoder-Decoder Structure Training

[0082] The Transformer with encoder-decoder structure is used as the description generation model to further strengthen the connection between input features and output descriptions. The Adam optimizer is used in the training process, the initial learning rate is set to 1e-4, and the batch size is 32. The training is completed for about 20 epochs, and the training is stopped early when the F1 value of the validation set no longer improves to prevent overfitting.

[0083] Step S4: Image description generation

[0084] Use the best trained model to generate descriptions for new remote sensing images. This includes the following sub-steps:

[0085] Step S41: Description generation

[0086] Substitute the remote sensing image into the model and let the trained Transformer model generate the corresponding remote sensing image description. Score the generated description according to different evaluation indicators (such as BLEU, ROUGE, CIDEr, etc.) to ensure the accuracy and robustness of the description.

[0087] Experimental results and analysis

[0088] The experimental results on the RSICD dataset show that the remote sensing image description method based on multimodal feature fusion proposed in this paper is superior to existing methods in multiple evaluation indicators. For example, in the BLEU-4 indicator, this method is improved by about 3 percentage points compared with the traditional ResNet+LSTM method; in the CIDEr indicator, it is improved by about 5 percentage points. In addition, it can be seen from the ablation experiment that each module, such as the contrastive learning module, the graph attention network, and the CLIP feature fusion, has made significant contributions to the improvement of the final performance.

[0089] It should be noted that, in this application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0090] The above description is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest range consistent with the principles and novel features applied for herein.

Claims

1. A remote sensing image description method based on multimodal feature fusion, characterized in that: The following steps are involved: S1: Extract features using ResNet-152 and YOLOv8 and optimize through contrastive learning; S2: The features are enhanced by the graph attention network and multi-head attention mechanism combined with CLIP features; S3: Substitute the features into the transformer guided by prior knowledge for training; S4: Finally, the best trained model is used to generate descriptions.

2. The remote sensing image description method of multimodal feature fusion according to claim 1, characterized in that: The features extracted in step S1 are specifically as follows: The ResNet-152 model is used to extract scene-level features of remote sensing images. This feature mainly captures the global information of the image and has strong expressive power: (1) Where I is the input remote sensing image, is the extracted scene-level features; The YOLOv8 model is used to extract object-level features of remote sensing images. This feature focuses on local targets in the image and can handle multi-scale targets in the image: (2) Where I is the input remote sensing image, , is the number of objects detected, is the dimension of each object feature.

3. The remote sensing image description method of multimodal feature fusion according to claim 1, characterized in that: The contrastive learning optimization in step S1 is specifically as follows: By comparing scene-level features and object-level features, we can optimize feature representation and improve the semantic understanding ability of the model by maximizing the similarity between images and text descriptions and minimizing the similarity between images and irrelevant texts: (3) in, , are the embedding vectors of object-level features and scene-level features, respectively, sim(.,.) is the similarity function (e.g. cosine similarity), and T is the temperature hyperparameter.

4. The remote sensing image description method of multimodal feature fusion according to claim 1, characterized in that: The step S2 is specifically as follows in combination with the CLIP feature: The optimized scene-level and object-level features are fused with the features extracted by the CLIP model to form a multimodal feature representation, thereby improving the accuracy of image description: (4) in, , is the dimension of the CLIP feature; (5) (6) in, and are object-level and scene-level features, respectively. is the global feature extracted by CLIP, , , is the weight coefficient, which is used to balance the contribution of each feature.

5. The remote sensing image description method of multimodal feature fusion according to claim 1, characterized in that: The processing of the graph attention network and the multi-head attention mechanism in step S2 is as follows: The fused scene-level features and object-level features are respectively input into the graph attention network and the multi-head attention mechanism module for enhancement. The enhanced features can effectively capture the contextual information in the image: (7) (8) (9) in, and is the node feature, a and W are trainable parameters, N(i) is the set of neighbor nodes of node i, is the activation function; (10) (11) in, , , , is the linear change matrix of each head, is the output linear change matrix.

6. The remote sensing image description method of multimodal feature fusion according to claim 1, characterized in that: The specific steps of model training in step S3 are as follows: By embedding common patterns, structural relationships, and semantic rules in the remote sensing field into the attention mechanism of the Transformer, the feature interaction process is guided; domain prior knowledge is integrated into the Transformer through position encoding and semantic embedding to enhance the model's understanding of the unique patterns of remote sensing images; The encoder-decoder structure is used to further strengthen the connection between input features and output descriptions, generate high-quality descriptive sentences that conform to actual semantics, and conduct multiple training on remote sensing datasets.

7. The remote sensing image description method of multimodal feature fusion according to claim 1, characterized in that: The description generated in step S4 is specifically as follows: The model with the best performance in the training phase is used to input the optimized multimodal features (including scene-level features, object-level features, and CLIP fusion features) into the decoder part of the model to generate a text description consistent with the input remote sensing image content; the generated description can not only accurately reflect the main content of the image, but also reflect the detailed features in the image, such as the target type, spatial distribution, texture information, and scene semantics in the image; The description generated by the model fully considers the particularity of remote sensing images and reflects the diversity and contextual relevance of image targets in the description; for example, for remote sensing images containing complex scenes, the description can mention the overall information and local details of the scene to ensure the completeness and logic of the description.

Citation Information

Patent Citations

  • A Wide-Switch Remote Sensing Description Generation Method Based on Target Detection

    CN110929640B

  • Remote sensing image description method based on multi-category label assistance

    CN117953363A

  • Remote sensing image description method based on multi-scale information extraction and fusion

    CN118334520A

Cited By

  • A Scene-Guided Zero-Shot Remote Sensing Image Text Description Generation Method and System

    CN122676511A