Remote sensing image-text retrieval method based on enhanced data set and multilevel attention module

By building an enhanced data set and introducing multi-level attention (MLLA) module, the problem of insufficient processing capabilities for image deformation, damage, distortion and noise in remote sensing image text retrieval is solved, and the search effect and robustness are significantly improved.

CN119961478APending Publication Date: 2025-05-09GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411894851.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-21
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing remote sensing image graphic search technology is not ideal when dealing with image deformation, damage, distortion and noise, and lacks the ability to capture image details and local features, resulting in low retrieval accuracy and information loss and inconsistency problems during multimodal learning.

Method used

By building an enhanced data set and introducing a multi-level attention (MLLA) module, the model's ability to capture image details and local features is enhanced, and the model's robustness to image deformation, damage, distortion and noise is improved.

Benefits of technology

It significantly improves the image and text retrieval effect of remote sensing images, improves the model's ability to capture image details and local features, and enhances the robustness of image deformation, damage, distortion and noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961478A_ABST
    Figure CN119961478A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image and text retrieval method based on an enhanced data set and a multi-level attention (MLL) module, and relates to a remote sensing image and text retrieval method. The method comprises the following steps: firstly, carrying out data preprocessing, carrying out normalization processing on original remote sensing image data, and constructing an enhanced data set by adopting various transformation technologies; secondly, on the basis of an existing RemoteCLIP model, a multi-level linear attention module is integrated, and an EnhanceMLLA-RemoteCLIP model is constructed; then, performing fine tuning training on the model by using an enhanced data set to improve the capturing capability of the model on image details and local features; and finally, the trained EnhanceMLLA-RemoteCLIP model is applied to a remote sensing image to be retrieved, so that accurate and effective image-text retrieval is realized. By introducing an enhanced data set and a multi-level attention mechanism, the robustness of the model to image deformation, damage, distortion and noise is remarkably enhanced, the retrieval capability of the remote sensing image is improved, and retrieval of various complex remote sensing image data is more efficient and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image retrieval, and in particular to a remote sensing image graphic and text retrieval method based on an enhanced data set and a multi-level layer attention (MLLA) module. Background Art

[0002] With the continuous development of remote sensing technology, remote sensing images are increasingly used in military monitoring, urban planning, environmental protection and other fields. Researchers at home and abroad have been committed to improving the image and text retrieval capabilities of remote sensing images. Abroad, researchers mainly focus on the application of deep learning, convolutional neural networks (CNNs) and generative adversarial networks (GANs) in remote sensing image processing. NASA and the European Space Agency (ESA) have invested a lot of resources in the collection and processing of remote sensing image data, and promoted the development of related technologies through public competitions and data sets (such as SpaceNet).

[0003] At present, patented technologies mainly use the following methods to realize image and text retrieval of remote sensing images:

[0004] 1. Traditional machine learning-based methods, such as support vector machine (SVM) and random forest (RF), match images by extracting manual features.

[0005] 2. Deep learning methods based on convolutional neural networks (CNNs), which extract high-level features of images by training large-scale data sets to achieve image-text matching.

[0006] 3. Use Generative Adversarial Networks (GANs) for data augmentation to address common deformation, damage, distortion, and noise problems in remote sensing images.

[0007] 4. Apply multimodal learning to improve the accuracy of retrieval by combining multimodal information of images and texts.

[0008] In the field of patent technology, there have been a number of patents that attempt to improve the processing and retrieval capabilities of remote sensing images through different means. Patent CN201810744473.0 proposes a remote sensing image semantic generation method based on a fast regional convolutional neural network to solve the problem that the existing technology cannot obtain the relationship between targets in the image, nor the relationship between the target and the image as a whole. Patent CN202310835877.1 is based on a remote sensing image and text retrieval method guided by visual semantic alignment to solve the common visual-semantic imbalance problem in remote sensing image and text retrieval. In addition, patent CN202410008193.9 introduces a cross-modal remote sensing image and text retrieval method that fuses local features with global features, aiming to correct global information through local information and supplement local information with global information, thereby improving the accuracy of retrieval. However, these existing technologies still have the following shortcomings in practical applications:

[0009] 1. The problems of image deformation, damage, distortion and noise have not been completely solved, resulting in unsatisfactory retrieval results.

[0010] 2. Existing methods are not capable of capturing image details and local features, which affects the accuracy of retrieval.

[0011] 3. Although multimodal learning methods help improve retrieval results, they still face challenges when integrating information from different modalities, such as information loss and inconsistency.

[0012] In order to solve these problems and improve the image and text retrieval capabilities of remote sensing images, this paper proposes a remote sensing image image and text retrieval method based on enhanced datasets and multi-level attention (MLLA) modules. By building an enhanced dataset and introducing a multi-level attention mechanism, the model's ability to capture image details and local features is improved, thereby improving the retrieval effect. This method combines international cutting-edge technology with domestic research breakthroughs, aiming to provide a more accurate and robust solution for remote sensing image image and text retrieval. Summary of the invention

[0013] The purpose of the present invention is to provide a remote sensing image graphic and text retrieval method based on an enhanced dataset and an MLLA module, which realizes efficient and accurate graphic and text retrieval of remote sensing images by enhancing the dataset and applying a multi-level attention mechanism.

[0014] Data preprocessing: The original remote sensing image data is normalized and enhanced through a variety of transformations (such as color enhancement, contrast enhancement, brightness enhancement, sharpness enhancement, image flipping, rotation, cropping, scaling, shearing and adding noise, etc.) to obtain the enhanced image and construct an enhanced data set.

[0015] Model construction: Based on the existing RemoteCLIP model, a multi-level attention (MLLA) module is added to build the EnhanceMLLA-RemoteCLIP model. The MLLA module includes a linear attention layer (Linear Attention) and a multi-layer perceptron (MLP), which captures and processes local features of the image through a multi-level self-attention mechanism.

[0016] Model training: Use the enhanced dataset to fine-tune the EnhanceMLLA-RemoteCLIP model to enhance the model's ability to capture image details and local features.

[0017] Image and text retrieval: The remote sensing image to be retrieved is input into the trained EnhanceMLLA-RemoteCLIP model, image features are extracted and matched with text features to achieve accurate image and text retrieval.

[0018] The data preprocessing steps are as follows:

[0019] Normalization: Normalize the data according to formula (1):

[0020] (1)

[0021] In the formula, is the value to be normalized, is the maximum value among the collected data features. It is the smallest value among the collected data features;

[0022] Data enhancement: Perform multiple transformations on the normalized image to generate an enhanced data set. Specific image enhancement methods include but are not limited to the following: color enhancement, sharpness enhancement, brightness enhancement, contrast enhancement, image flipping and rotation, cropping, scaling, shearing, and adding noise.

[0023] The step model construction includes the following steps:

[0024] Constructing MLLA module: The MLLA module captures and processes local features of the image through a multi-level self-attention mechanism. It is mainly composed of a linear attention module, a multi-layer perceptron module, a forget gate, a block design, and a position encoding.

[0025] Model integration: Based on the existing RemoteCLIP model, the MLLA module is added to integrate the two and build the EnhanceMLLA-RemoteCLIP model.

[0026] The step model training includes the following steps:

[0027] Fine-tune the EnhanceMLLA-RemoteCLIP model using the enhanced dataset;

[0028] During the model training process, the model parameters are optimized through cross-validation and early stopping mechanism;

[0029] Use the validation set to evaluate model performance, select the best model and save it.

[0030] The step of image and text retrieval includes the following steps:

[0031] Input the remote sensing image to be retrieved into the trained EnhanceMLLA-RemoteCLIP model to extract image features;

[0032] Match image features with text features to generate retrieval results.

[0033] The present invention has the following beneficial effects and advantages:

[0034] 1. By constructing an enhanced dataset, the diversity and generalization of the data are enhanced, and the robustness of the model to image deformation, damage, distortion and noise is improved;

[0035] 2. By introducing the multi-level attention (MLLA) module, the model's ability to capture image details and local features is enhanced, significantly improving the image and text retrieval effect of remote sensing images;

[0036] 3. Through mixed precision training and gradient scaling technology, the model training process is accelerated, the use of video memory is reduced, and the training efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a flow chart of a remote sensing image text retrieval method based on an enhanced dataset and a multi-level attention (MLLA) module of the present invention.

[0038] Figure 2 This is a data set information table according to an embodiment of the present invention.

[0039] Figure 3 This is a comparison chart of the changes between the enhanced dataset and the original dataset.

[0040] Figure 4 This is a diagram of the fine-tuning training situation of an embodiment of the present invention.

[0041] Figure 5 This is a performance comparison chart of the model trained by the embodiment of the present invention and the original model. DETAILED DESCRIPTION

[0042] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0043] In the following description, many specific implementation details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0044] Example:

[0045] Combination Figure 1 The specific implementation steps of a remote sensing image text retrieval method based on an enhanced data set and a multi-level attention module are described. The present invention uses RSICD, RSITMD, and UCM as experimental data sets. The specific data set information is as follows: Figure 2 shown.

[0046] Step 1) Data preprocessing and building enhanced datasets

[0047] The original data sets RSICD, RSITMD, and UCM were preprocessed and the data were normalized to ensure the uniformity of data specifications for subsequent experiments. Then, enhancement processing was performed. The specific enhancement methods are as follows:

[0048] Color enhancement: enhance the color of the image using random coefficients;

[0049] Contrast enhancement: Use random coefficients to enhance the contrast of the image;

[0050] Brightness enhancement: Use random coefficients to enhance the brightness of the image;

[0051] Sharpness enhancement: Use random coefficients to enhance the sharpness of the image;

[0052] Image flip: randomly choose horizontal flip or vertical flip;

[0053] Rotation: Randomly select 90, 180 or 270 degree rotation;

[0054] Cropping: Randomly crop a portion of the image and then resize it back to its original size;

[0055] Scale: Randomly scale the image;

[0056] Shear: randomly adds shear transformations to the image;

[0057] Add Noise: Randomly add noise to the image.

[0058] Constructing an enhanced dataset requires ensuring that each enhancement method appears in a balanced manner so that the model can achieve better training results under each enhancement method.

[0059] Combination Figure 3 Illustrates the changes in some images in the dataset after enhancement.

[0060] Step 2) Construction of EnhanceMLLA-RemoteCLIP model

[0061] Import the RemoteCLIP original model and build the MLLA module in Python, and combine the two. Use Torch to implement the linear attention module, multi-layer perceptron module, forget gate, block design, and position encoding module combination, and adjust the hyperparameters (such as embedding dimension, hidden layer dimension, number of layers, etc.) as needed.

[0062] Step 3) Model fine-tuning training

[0063] The model constructed using step 2) is trained on the enhanced data set of step 1), and the training process and control are implemented in code. In this embodiment, the batch_size is set to 16 and the learning rate is 5e-6. A smaller learning rate can help the training process be smoother, especially suitable for fine-tuning models. At the same time, the cross entropy loss function is used to calculate the loss between image features and text features to optimize model parameters; the training process is accelerated and the use of video memory is reduced through mixed precision training and gradient scaling technology; the learning rate scheduler and early stopping mechanism are used for training optimization to avoid overfitting; the validation set is used to evaluate the model performance, and the best model is saved based on the validation loss.

[0064] Combination Figure 4 Illustrates the loss during fine-tuning of a trained model.

[0065] Step 4) Remote sensing image and text retrieval application

[0066] Use the best trained EnhanceMLLA-RemoteCLIP model saved in step 3) for dataset retrieval, which improves the robustness to image deformation, damage, distortion and noise.

[0067] Combination Figure 5 It shows that the trained model EnhanceMLLA-RemoteCLIP has improved the image retrieval effect compared with the original model.

[0068] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image text retrieval method based on enhanced dataset and multi-level attention module, characterized in that The steps include: Step 1) Normalize the original remote sensing image data and obtain the enhanced image through various transformations (such as color enhancement, contrast enhancement, brightness enhancement, sharpness enhancement, image flipping, rotation, cropping, scaling, shearing and adding noise, etc.), thereby constructing an enhanced data set; Step 2) Based on the existing RemoteCLIP model, a multi-level attention (MLLA) module is added to build the EnhanceMLLA-RemoteCLIP model; Step 3) Use the enhanced dataset to fine-tune the EnhanceMLLA-RemoteCLIP model to enhance the model's ability to capture image details and local features; Step 4) Input the remote sensing image to be retrieved into the trained EnhanceMLLA-RemoteCLIP model, extract image features and match them with text features to achieve accurate image and text retrieval.

2. The remote sensing image text retrieval method based on enhanced data set and multi-level attention module according to claim 1, characterized in that: Step 1) The enhancement methods include but are not limited to: color enhancement: using random coefficients to enhance the color of the image; contrast enhancement: using random coefficients to enhance the contrast of the image; brightness enhancement: using random coefficients to enhance the brightness of the image; sharpness enhancement: using random coefficients to enhance the sharpness of the image; image flipping: randomly selecting horizontal flipping or vertical flipping; Rotation: Randomly select 90 degrees, 180 degrees or 270 degrees for rotation; Cropping: Randomly crop a portion of the image and then resize it back to its original size; Scale: Randomly scale the image; Shear: randomly adds shear transformations to the image; Add noise: Randomly add noise to the image; use the new image data obtained by the above enhancement methods to construct an enhanced dataset.

3. The remote sensing image text retrieval method based on enhanced data set and multi-level attention module according to claim 1, characterized in that: Step 2) The MLLA module includes a linear attention layer (Linear Attention) and a multi-layer perceptron (MLP), which captures and processes the local features of the image through a multi-level self-attention mechanism.

4. The remote sensing image text retrieval method based on enhanced data set and multi-level attention module according to claim 1, characterized in that: Step 3) The fine-tuning training uses a smaller learning rate for training; uses a cross entropy loss function to calculate the loss between image features and text features to optimize model parameters; accelerates the training process and reduces video memory usage through mixed precision training and gradient scaling technology; uses a learning rate scheduler and an early stopping mechanism for training optimization to avoid overfitting; and uses a validation set to evaluate model performance, and saves the best model based on the validation loss.

5. The remote sensing image text retrieval method based on enhanced data set and multi-level attention module according to claim 1, characterized in that: Step 4) The retrieval object dataset is an enhanced dataset, so as to compare the retrieval accuracy and robustness of each model for the enhanced image.

Citation Information

Patent Citations

  • A method for semantic generation of remote sensing images based on fast region convolutional neural networks

    CN108960330B

  • Remote sensing image-text retrieval method based on guiding visual semantic alignment

    CN117009569A

  • A cross-modal remote sensing image and text retrieval method based on the fusion of local and global features

    CN117520589B