Remote sensing image long subtitle generation method, system and device based on interactive wavelet transform and Transform and medium

By using interactive wavelet transformation and Transformer methods in remote sensing image subtitle generation, a high-quality image-text-to-dataset is constructed and image-text comparison learning is optimized, which solves the problem of image-text alignment and achieves the generation of more accurate and rich subtitle descriptions.

CN120111162APending Publication Date: 2025-06-06XIDIAN UNIV

Patent Information

Application Number
CN202510225548.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively align images and text in the generation of remote sensing image subtitles, and it is difficult to accurately extract complex features such as the positional relationship and proportion of land objects, resulting in relatively single and inaccurate descriptions.

Method used

Using an interactive wavelet transform and Transformer method, the visual representation part can learn to extract the most relevant visual information to the text by constructing a high-quality remote sensing image-text pair dataset and using the interactive wavelet transform module to optimize the pre-training objectives of image-text comparison learning.

Benefits of technology

It realizes a more comprehensive mining of complex semantic content of remote sensing images, generates more accurate and rich subtitle descriptions, and overcomes the problems of incomplete and inaccurate semantic information representation in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111162A_ABST
    Figure CN120111162A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image long subtitle generation method, system and device based on interactive wavelet transform and Transform and a medium, and the method comprises the steps: carrying out the semantic segmentation of an obtained remote sensing image data set through employing a semantic segmentation model, and generating the proportion information of various ground features in different directions in a remote sensing image; inputting the image into a large language model, generating a text according to specific requirements, and constructing an image-text pair; the image-text pairs are comprehensively examined, and obviously wrong image-text pairs are removed; arranging and storing the reviewed image-text pair as a remote sensing image-text pair data set; building a remote sensing image long caption generation network which comprises an image encoder, an interactive wavelet transform module and a language model; training an interactive wavelet transform module; training a remote sensing image long caption generation network; evaluating the performance; the system, the equipment and the medium are used for implementing the method. The method has the advantages that the data set quality is improved, the model understanding and generating capability is enhanced, and the subtitle generating accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image subtitle generation, and in particular relates to a remote sensing image long subtitle generation method, system, device and medium based on interactive wavelet transform and Transformer. Background Art

[0002] Remote sensing image caption generation is an important task in the field of remote sensing. It aims to give accurate and rich text descriptions to remote sensing images so as to better understand and utilize the information in remote sensing images. It has important application value in many fields such as environmental monitoring, disaster warning, agricultural production, urban planning, etc. For example, in environmental monitoring, accurate image captions can help quickly identify environmental problems such as deforestation and water pollution; in disaster warning, it can timely describe the location and extent of the disaster and provide key information for rescue operations.

[0003] As a key part of remote sensing image subtitle generation, long text generation faces many challenges. Due to the complexity of remote sensing images, which contain multiple types of objects, complex spatial structures, and rich texture information, long texts are often required to accurately describe them. At the same time, images and texts belong to different modalities and there is a modality gap, which makes it difficult to effectively align image and text features when generating long texts, resulting in limited semantic coherence and accuracy. In addition, long text generation aims to extract more visual feature descriptions. In addition to the main object attribute information, it also needs to include information such as the location relationship of objects and the proportion of various objects. However, existing technologies are insufficient in extracting such information. With the continuous development of remote sensing technology, the demand for accuracy and richness of image subtitle generation is becoming increasingly urgent.

[0004] The patent application document with the publication number CN113192030A discloses a remote sensing image description generation method and system, which uses deep learning technology to extract multi-level visual features of the remote sensing image to be described; based on the multi-level visual features of the remote sensing image to be described, the spatial attention mechanism and the channel attention mechanism are used to obtain the multi-level features of the remote sensing image to be described; based on the multi-level visual features of the remote sensing image to be described, the contextual features of the remote sensing image to be described are obtained by using the contextual attention module; based on the multi-level features and contextual features of the remote sensing image to be described, the high-level semantic features of the remote sensing image to be described are obtained by using the visual sentinel adaptive mechanism; the high-level semantic features of the remote sensing image to be described are input into the trained language model to obtain the description sentence of the remote sensing image to be described. The disadvantage of this method is that there is no special processing for the modal gap between the image and the text in the process. For complex remote sensing images that require long text descriptions, it is impossible to align the features between the remote sensing image and the text, and it is difficult to generate accurate and detailed descriptions.

[0005] The patent application document with the publication number CN115035508A discloses a topic-guided Transformer remote sensing image caption generation method, which mainly solves the problem that the description generated by the prior art is single and cannot accurately represent the semantic information in the image. Its implementation scheme is: build a topic encoder composed of Transformer and topic vectors, and pre-train on the classification data set; build a semantic decoder composed of a random mask layer, an embedding layer, a Transformer decoder and a cascade of soft-max layers; connect the topic encoder and the semantic decoder to obtain a remote sensing image caption generation network; set training parameters, and iteratively train the remote sensing image caption generation network with the standard RSICD data set; use the trained remote sensing image caption generation network to generate caption descriptions. The topic-guided approach of this method may limit the comprehensiveness of feature extraction. The topic vector can only capture features related to the preset topic. Some complex features in remote sensing images that are not closely related to the topic but are very important for accurate caption generation cannot be effectively extracted, and it is difficult to ensure its language richness in the long text caption generation task.

[0006] In summary, the existing technologies may not be comprehensive and accurate enough in representing semantic information, and it is difficult to fully capture the complex semantic content of remote sensing images, especially the detailed semantic information involving multiple objects and their relationships, resulting in relatively simple descriptions. There are deficiencies in dealing with the modal gap between images and texts and alignment issues when generating long texts, and it is impossible to ensure the consistency of image features and text semantics, which is prone to semantic discontinuity or inaccuracy. Summary of the invention

[0007] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a method, system, device and medium for generating long subtitles for remote sensing images based on interactive wavelet transform and Transformer. By constructing a high-quality remote sensing image-text pair data set and using it to train an interactive wavelet transform module, the pre-training target of image-text contrast learning is optimized, so that the visual representation part can learn to extract the visual information most relevant to the text and better capture the semantic relationship between the image and the text. The present invention can more comprehensively mine the complex semantic content of remote sensing images, including detailed information of various landforms and their relationships, thereby generating more accurate and rich subtitle descriptions, overcoming the problem of incomplete and inaccurate representation of semantic information in the prior art.

[0008] In order to achieve the above object, the technical solution adopted by the present invention is:

[0009] A remote sensing image long subtitle generation method based on interactive wavelet transform and Transformer comprises the following steps:

[0010] Step 1: obtain a remote sensing image dataset, use a semantic segmentation model to perform semantic segmentation on the acquired remote sensing image dataset, and generate information on the proportion of various types of objects in different directions in the remote sensing image;

[0011] Step 2, input the proportion information of various types of objects in different directions in the remote sensing image generated in step 1 into the large language model, generate text according to specific requirements, and construct image-text pairs;

[0012] Step 3: conduct a comprehensive review of the image-text pairs constructed in step 2 and remove the image-text pairs that are obviously wrong;

[0013] Step 4, organizing and saving the image-text pairs reviewed in step 3 to be used as a remote sensing image-text pair dataset;

[0014] Step 5: Build a remote sensing image long caption generation network, including: image encoder, interactive wavelet transform module and language model;

[0015] The image encoder is used to convert the input image into high-dimensional visual features to provide visual information; the output high-dimensional visual features and the input text are input into the interactive wavelet transform module;

[0016] The interactive wavelet transform module is located between the image encoder and the language model. It is connected to the language model using a fully connected layer for information filtering and interaction. It filters irrelevant information from the high-dimensional visual features extracted by the frozen image encoder and enhances the visual representation related to the input text. The interactive wavelet transform module outputs a multi-scale visual feature representation, referred to as a visual representation, which is a high-dimensional feature vector.

[0017] The language model is used to generate long subtitles of remote sensing images. It receives the multi-scale visual feature representation output by the interactive wavelet transform module, generates a text description corresponding to the image content, and finally outputs the long subtitles of the remote sensing image.

[0018] Step 6, using the remote sensing image-text pair dataset obtained in step 4 to train the interactive wavelet transform module built in step 5;

[0019] Step 7, connecting the language model in step 5 with the interactive wavelet transform module trained in step 6, guiding the language model to generate text based on visual features for feature alignment training, and obtaining a trained remote sensing image long subtitle generation network;

[0020] Step 8: Evaluate the performance of the remote sensing image long subtitle generation network trained in step 7, and calculate the matching scores between subtitles generated by different models and remote sensing images through the Long-CLIP model.

[0021] The specific method of step 1 is:

[0022] A remote sensing image dataset is obtained, and a semantic segmentation model is used to perform semantic segmentation on the remote sensing image dataset. The remote sensing image dataset is divided into six major categories, including water bodies, buildings, roads, woodlands, grasslands, and farmlands. The remote sensing image is divided into five rectangular areas, namely, upper left, upper right, lower left, lower right, and middle. The proportion information of various types of objects in different directions in the remote sensing image is automatically generated according to grammatical rules. That is, in the remote sensing image, in the [upper left] area: [water bodies] account for [A]%, [buildings] account for [B]%, [roads] account for [C]%, [woodlands] account for [D]%, [grassland] accounts for [E]%, and [farmland] accounts for [F]%; the same description is made for the remaining four areas of upper right, lower left, lower right, and middle.

[0023] In step 3, the following three criteria are used to determine whether the image-text pair is obviously wrong:

[0024] Error judgment of the description of the proportion of land objects: whether the quantity information of various land objects mentioned in the text generated by the large language model matches the proportion information of various land objects in different directions in the remote sensing image generated in step 1. The basis for judging whether they match each other includes but is not limited to: land object categories with a proportion of less than 10% cannot use "largely present" or "dominant"; land object categories with a proportion of more than 50% cannot use "smallly present" or "somewhat visible";

[0025] Semantic logic error judgment: Check whether the text generated by the large language model violates basic geographic logic or semantic contradictions. The basis for judging whether there are violations of basic geographic logic or semantic contradictions includes but is not limited to: the text describes a certain area as a large area of ​​water, but at the same time describes the area as having a large amount of farmland without a reasonable transition or explanation, or describes a certain area as having a large number of buildings while there are no roads in the area without a reasonable explanation or description;

[0026] Verification of language expression accuracy: The text generated by the large language model is checked based on standard English grammar, vocabulary usage specifications, and the accuracy of professional terminology in the remote sensing field. If there are obvious grammatical errors or errors in the use of professional terminology, it is judged as a language expression error.

[0027] In the step 5, the high-dimensional visual features output by the image encoder are input into the interactive wavelet transform module together with the input text. The interactive wavelet transform module includes a parallel image converter and a text converter, and shares weights; there are N repeated specific units in both the image converter and the text converter, and the specific unit in the image converter includes a wavelet transform layer, a first layer of addition and normalization, a multi-head self-attention layer, a second layer of addition and normalization, a feedforward layer, and a third layer of addition and normalization; the specific unit input in the image converter consists of two parts: a high-dimensional visual feature and a learnable visual cue word. After the learnable visual cue word is processed by the wavelet transform layer, the first layer of addition and normalization operation is performed with the learnable visual cue word that has not been processed by the wavelet transform layer, and the result is then input into the multi-head self-attention layer together with the high-dimensional visual feature, and the output of the multi-head self-attention layer is subjected to a second layer of addition and normalization operation with the result of the first layer. The output result is processed by the feedforward layer, and then the third layer of addition and normalization operation is performed with the result of the second layer of addition and normalization. After N cycles, the output of the image converter is obtained; the specific unit in the text converter includes a wavelet transform layer, a first layer of addition and normalization, a feedforward layer, and a second layer of addition and normalization; the input of the specific unit in the text converter is the input text, and the input text is processed by the wavelet transform layer and then the first layer of addition and normalization operation is performed. The output result is processed by the feedforward layer and then the second layer of addition and normalization operation is performed with the result of the first layer of addition and normalization. After N cycles, the output of the text converter is obtained; the output of the image converter and the output of the text converter are compared and learned by image and text to update the learnable parameters in the interactive wavelet transform module. The interactive wavelet transform module outputs a multi-scale visual feature representation, referred to as visual representation, which is a high-dimensional feature vector.

[0028] The step 6 comprises:

[0029] Step 6.1, the wavelet transform layer in the interactive wavelet transform module uses db4 wavelet transform to process the input embedding;

[0030] Step 6.2, during the training process, the image-text contrastive learning loss (ITC) is calculated using a bidirectional objective;

[0031] L ITC =L g2t +L t2g

[0032] Among them, L g2t is the contrast loss from image to text, L t2g is the contrast loss from text to image; its specific calculation formula is as follows:

[0033]

[0034] Among them, B represents the batch size, that is, the number of image-text pairs contained in a batch of data, u is an index in the batch, indicating the image sample currently being processed, It is the feature representation of the image sample with index u after being processed by the image transformer based on wavelet transform. It is the feature representation of the text sample with index k after being processed by the text transformer based on wavelet transform. This means that k is the index of the text sample in the batch corresponding to the image sample u, and τ is the temperature parameter of the softmax function.

[0035] The specific method of step 7 is:

[0036] Step 7.1, connect the language model in step 5 and the interactive wavelet transform module trained in step 6 through a fully connected layer, and the fully connected layer adjusts the visual representation output by the interactive wavelet transform module to the same dimension as the text embedding of the language model to obtain the visual representation embedding;

[0037] Step 7.2, add the visual representation embedding obtained in step 7.1 as a prefix to the text embedding to form a new embedding sequence and input it into the language model in step 5, guide the language model to generate text based on visual features for further feature alignment training, and use language modeling loss to further perform feature alignment training on the fully connected layer and wavelet transform layer during the training process. The trained remote sensing image long subtitle generation network is obtained.

[0038] The present invention also provides a remote sensing image long subtitle generation system based on interactive wavelet transform and Transformer, comprising:

[0039] The module for generating information on the proportion of ground objects is used to obtain a remote sensing image dataset, perform semantic segmentation on the acquired remote sensing image dataset using a semantic segmentation model, and generate information on the proportion of various ground objects in different locations in the remote sensing image;

[0040] The image-text pair construction module is used to input the proportion information of various types of objects in different directions in the remote sensing image into the large language model, generate text according to specific requirements, and construct image-text pairs;

[0041] The error rejection module is used to conduct a comprehensive review of image-text pairs and reject those with obvious errors;

[0042] A remote sensing image-text pair dataset generation module is used to organize and save the reviewed image-text pairs to be used as a remote sensing image-text pair dataset;

[0043] Remote sensing image long subtitle generation network construction module, used to build remote sensing image long subtitle generation network, including: image encoder, interactive wavelet transform module and language model;

[0044] An interactive wavelet transform module training module, used to train an interactive wavelet transform module using a remote sensing image-text pair dataset;

[0045] The remote sensing image long subtitle generation network training module is used to connect the language model with the trained interactive wavelet transform module, guide the language model to generate text based on visual features for feature alignment training, and obtain the trained remote sensing image long subtitle generation network;

[0046] The performance evaluation module is used to evaluate the performance of the trained remote sensing image long subtitle generation network, and calculate the matching scores between subtitles generated by different models and remote sensing images through the Long-CLIP model.

[0047] The present invention also provides a remote sensing image long subtitle generation device based on interactive wavelet transform and Transformer, comprising:

[0048] Memory: a computer program for storing the above-mentioned method for generating long subtitles for remote sensing images based on interactive wavelet transform and Transformer, which is a computer-readable device;

[0049] Processor: used to implement the method for generating long subtitles of remote sensing images based on interactive wavelet transform and Transformer when executing the computer program.

[0050] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the method for generating long subtitles for remote sensing images based on interactive wavelet transform and Transformer.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] 1. The present invention generates information on the proportion of various types of objects in different directions in remote sensing images and combines them with a large language model to present image information more accurately and in detail. This innovation provides a high-quality data information foundation for subsequent data set production and solves the problem that existing public data sets are not detailed enough to fully and accurately describe image content.

[0053] 2. The present invention uses wavelet transform layers to process input visual representations and text representations. It mainly plays the role of information filtering and interaction, enhancing the model's ability to understand image content and generate text descriptions. This innovation improves the alignment effect of image and text features when generating long texts, thereby improving semantic coherence and accuracy.

[0054] 3. The present invention abandons the traditional text similarity evaluation index, which has limitations when processing complex language output and is difficult to accurately quantify model performance. The present invention adopts a new evaluation idea, no longer relying solely on text similarity, but starting from the perspective of image and text matching. The Long-CLIP model is introduced to calculate the matching score. It can effectively capture the deep semantic association and fine-grained feature matching between images and texts, and is highly adapted to the remote sensing image subtitle generation task, providing a more effective method for evaluating model performance.

[0055] In summary, the present invention combines the multi-directional feature ratio information with a large language model to construct a data set, uses the cross-modal feature alignment technology based on interactive wavelet transform to achieve precise feature alignment between data of different modalities, and introduces the Long-CLIP model deep matching evaluation system to perform deep matching and comprehensive evaluation on the generated subtitles and images, thereby achieving high-quality output of remote sensing image subtitle generation, which has the advantages of improving the quality of the data set, enhancing the model's understanding and generation capabilities, and improving the accuracy of subtitle generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is an implementation flow chart of the present invention.

[0057] Figure 2 It is a partitioning method for calculating the proportion of various types of objects based on the semantic segmentation results when making a dataset.

[0058] Figure 3 It is a structural diagram of the interactive wavelet transform module and the frozen image encoder.

[0059] Figure 4 It is to guide the language model to generate text based on visual features and further perform feature alignment training using the model structure diagram.

[0060] Figure 5 It is the subtitle content generated by the model of the present invention and the remote sensing image used. DETAILED DESCRIPTION

[0061] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0062] The technical problem to be solved by the present invention is the difficulty in effectively aligning images and texts and accurately extracting complex features (such as the position relationship and proportion of objects) faced by remote sensing image subtitle generation tasks when generating long texts.

[0063] The present invention uses a high-quality remote sensing image-text pair dataset to solve the problem of the remote sensing image-text pair dataset being insufficient in accurately and richly describing the content of remote sensing images. A remote sensing image long subtitle generation method based on interactive wavelet transform and Transformer is used to better extract complex information and align cross-modal information. It relies on the design of interactive wavelet transform to better align remote sensing images and text. Interactive wavelet transform can extract multi-scale features of remote sensing images and text in the wavelet domain, thereby improving the accuracy and richness of remote sensing image subtitle generation.

[0064] like Figure 1 As shown, a remote sensing image long subtitle generation method based on interactive wavelet transform and Transformer includes the following steps:

[0065] Step 1: obtain a high-quality remote sensing image dataset, use a semantic segmentation model to perform semantic segmentation on the obtained high-quality remote sensing image dataset, and generate information on the proportion of various types of objects in different directions in the remote sensing image;

[0066] The specific method of step 1 is:

[0067] To obtain high-quality remote sensing image datasets, we first identified multiple reliable data sources, including high-resolution satellite images (such as GF3 and GF7) and professional geographic image platforms such as Google Maps. We constructed screening conditions based on key features of remote sensing images, such as the image resolution range (0.8m-1.2m), the image shooting time range (within the past year), and the geographic coordinate range of the image coverage area, so that the automated data collection program can accurately capture image data that meets the requirements from the selected data source according to these conditions. We used the existing high-performance semantic segmentation model DeepLab-V3-Plus to perform semantic segmentation on the remote sensing image dataset and divided it into six major categories: water bodies, buildings, roads, woodlands, grasslands, and farmlands. Figure 2 As shown, the partitioning method divides the remote sensing image into five rectangular areas, namely: upper left, upper right, lower left, lower right, and middle, and automatically generates the proportion information of various types of objects in the remote sensing image in different directions according to grammatical rules. That is, in the remote sensing image, in the [upper left] area: [water body] accounts for [A]%, [building] accounts for [B]%, [road] accounts for [C]%, [forest land] accounts for [D]%, [grassland] accounts for [E]%, and [farmland] accounts for [F]%; the same description is made for the remaining four areas of upper right, lower left, lower right, and middle.

[0068] Step 2: Input the proportion information of various types of objects in different directions in the remote sensing image generated in step 1 into the ChatGPT-3.5 large language model, guide it to generate text according to specific requirements (for example, use English to describe the content of the picture and summarize and generalize the information of the whole picture) and construct image-text pairs;

[0069] Step 3: conduct a comprehensive review of the image-text pairs constructed in step 2 and remove the image-text pairs that are obviously wrong;

[0070] In step 3, the following three criteria are used to determine whether the image-text pair is obviously wrong:

[0071] Error judgment on the description of the proportion of land features: whether the quantity information of various land features mentioned in the text generated by the large language model matches the proportion information of various land features in different directions in the remote sensing image generated in step 1. The basis for judging whether they match each other includes but is not limited to: the land feature categories with a proportion of less than 10% cannot be described as "existing in large numbers" or "dominant", and the land feature categories with a proportion of more than 50% cannot be described as "existing in small numbers" or "somewhat visible";

[0072] Semantic logic error judgment: Check whether the text generated by the large language model violates basic geographic logic or semantic contradictions. The basis for judging whether there are violations of basic geographic logic or semantic contradictions includes but is not limited to: the text describes a certain area as a large area of ​​water, but at the same time describes the area as having a large amount of farmland without a reasonable transition or explanation, or describes a certain area as having a large number of buildings while there are no roads in the area without a reasonable explanation or description; such situations are judged as semantic logic errors;

[0073] Verification of language accuracy: The text generated by the large language model is checked based on standard English grammar, vocabulary usage norms, and the accuracy of professional terminology in the remote sensing field. If there are obvious grammatical errors (such as inconsistency between subject and predicate, misuse of part of speech, etc.) or incorrect use of professional terminology (such as mistakenly writing "woodland" as "wood"), it is judged as a language expression error.

[0074] Step 4, organizing and saving the image-text pairs reviewed in step 3 to be used as a remote sensing image-text pair dataset;

[0075] Step 5: Build a remote sensing image long subtitle generation network, including three main parts: image encoder, interactive wavelet transform module and language model.

[0076] like Figure 3As shown, in step 5, the image encoder is used to convert the input image into high-dimensional visual features to provide visual information; the image encoder adopts the pre-trained visual Transformer model ViT-L / 14 derived from CLIP; the present invention uses a frozen image encoder to process the input image to obtain high-dimensional visual features.

[0077] The interactive wavelet transform module is located between the image encoder and the language model, and is connected to the language model using a fully connected layer for information filtering and interaction; the high-dimensional visual features extracted by the frozen image encoder are used to filter irrelevant information and enhance the visual representation related to the text; the high-dimensional visual features output by the image encoder are input into the interactive wavelet transform module together with the input text, and the interactive wavelet transform module includes parallel image converters and text converters, and shares weights; there are N repeated specific units in both the image converter and the text converter, and the specific unit in the image converter includes a wavelet transform layer, a first layer of addition and normalization, a multi-head self-attention layer, a second layer of addition and normalization, a feedforward layer, and a third layer of addition and normalization; the input of the specific unit in the image converter consists of two parts: a high-dimensional visual feature and a learnable visual cue word. After the learnable visual cue word is processed by the wavelet transform layer, the first layer of addition and normalization operation is performed with the learnable visual cue word that has not been processed by the wavelet transform layer, and the result is then compared with the high-dimensional The visual features are input into the multi-head self-attention layer together, the output of the multi-head self-attention layer is added to the result of the first layer and normalized, and the second layer addition and normalization operation is performed. After the output result is processed by the feedforward layer, the third layer addition and normalization operation is performed with the result of the second layer addition and normalization. After N cycles, the output of the image converter is obtained; the specific unit in the text converter includes a wavelet transform layer, a first layer addition and normalization, a feedforward layer, and a second layer addition and normalization; the input of the specific unit in the text converter is the input text, the input text is processed by the wavelet transform layer, the first layer addition and normalization operation is performed, the output result is processed by the feedforward layer, and the second layer addition and normalization operation is performed with the result of the first layer addition and normalization. After N cycles, the output of the text converter is obtained; the output of the image converter and the output of the text converter are used to update the learnable parameters in the interactive wavelet transform module through image-text comparison learning, and the interactive wavelet transform module outputs a multi-scale visual feature representation, referred to as a visual representation, which is a high-dimensional feature vector;

[0078] The language model is used to generate long subtitles for remote sensing images. It receives the multi-scale visual feature representation output by the interactive wavelet transform module, generates a text description corresponding to the image content, and finally outputs the long subtitles for the remote sensing images. The language model uses the decoder-only OPT-6.7B language model with unsupervised training.

[0079] Step 6, using the remote sensing image-text pair dataset obtained in step 4 to train the interactive wavelet transform module built in step 5; Figure 3 As shown, the interactive wavelet transform module is connected to the frozen image encoder and pre-trained using the remote sensing image-text pair dataset. The purpose is to train the interactive wavelet transform module, which enables the visual representation part to learn to extract the visual representation information most relevant to the text, mainly by optimizing the image-text contrast learning pre-training objective.

[0080] The step 6 comprises:

[0081] Step 6.1, the wavelet transform layer in the interactive wavelet transform module uses db4 wavelet transform to process the input embedding. The specific implementation of db4 wavelet transform is taken as an example of a two-dimensional signal S(mn), where m and n represent the indexes of its row and column respectively;

[0082] Calculate the two-dimensional approximation coefficient A:

[0083] Apply a low-pass filter h to the two-dimensional signal S(mn) Lo Perform row convolution to get S Lo,row (m,n);

[0084] S Lo,row (m,n)=S(m,n)*h Lo

[0085] Among them, h Lo Represents the low-pass filter coefficient, the row convolution result S Lo,row (m,n), use the same low-pass filter to perform column convolution and get S Lo,col (m,n);

[0086] S Lo,col (m,n)=S Lo,row (m,n)*h Lo

[0087] For S Lo,col (m,n) is downsampled to obtain the two-dimensional approximation coefficient A;

[0088] A(i,j)=S Lo,col (2i,2j);

[0089] Calculate the horizontal detail factor D H :

[0090] Apply a high-pass filter h to the two-dimensional signal S(mn) Hi Perform row convolution to get S Hi,row (m,n);

[0091] S Hi,row(m,n)=S(m,n)*h Hi

[0092] Among them, h Hi Represents the high-pass filter coefficient, the row convolution result S Hi,row (m,n), use a low-pass filter to perform column convolution and get S Hi,col (m,n);

[0093] S Hi,col (m,n)=S Hi,row (m,n)*h Lo

[0094] For S Hi,col (m,n) is downsampled to obtain the horizontal detail coefficient D H ;

[0095] D H (i,j)=S Hi,col (2i,2j);

[0096] Calculate the vertical detail factor D V :

[0097] Apply a high-pass filter h to the two-dimensional signal S(mn) Hi Perform column convolution to get S Hi,col (m,n);

[0098] S Hi,col (m,n)=S(m,n)*h Hi

[0099] Column convolution result S Hi,col (m,n), use a low-pass filter to perform row convolution and get S Hi,row (m,n);

[0100] S Hi,row (m,n)=S Hi,col (m,n)*h Lo

[0101] For S Hi,row (m,n) is downsampled to obtain the vertical detail coefficient D V ;

[0102] D V (i,j)=S Hi,row (2i,2j);

[0103] Calculate the diagonal detail factor D D :

[0104] Apply a high-pass filter h to the two-dimensional signal S(mn) Hi Perform row convolution to get S Hi,col(m,n);

[0105] S Hi,row (m,n)=S(m,n)*h Hi

[0106] Column convolution result S Hi,row (m,n), use a high-pass filter to perform column convolution and get S Hi,row (m,n);

[0107] S Hi,col (m,n)=S Hi,row (m,n)*h Hi

[0108] For S Hi,col (m,n) is downsampled to obtain the diagonal detail coefficient D D ;

[0109] D D (i,j)=S Hi,col (2i,2j);

[0110] The steps of two-dimensional wavelet inverse transform are as follows:

[0111] For the two-dimensional approximation coefficient A, the horizontal detail coefficient D H , vertical detail factor D V , diagonal detail coefficient D D Up-sampling is performed separately, taking the two-dimensional approximation coefficient as an example;

[0112]

[0113] Horizontal detail coefficient D H , vertical detail factor D V , diagonal detail coefficient D D Perform the same operation to obtain the two-dimensional approximate coefficients after upsampling Horizontal detail factor after upsampling Vertical detail factor after upsampling Diagonal detail coefficient after upsampling

[0114] The two-dimensional approximation coefficients after upsampling Horizontal detail factor Vertical detail factor Diagonal detail factor Perform inverse wavelet transform;

[0115] For the two-dimensional approximation coefficients after upsampling Use a low-pass filter h′ Lo After column convolution, the result is filtered using a low-pass filter h′ LoPerform row convolution to obtain low-frequency reconstruction subband

[0116] For the horizontal detail coefficient after upsampling Use a low-pass filter h′ Lo After column convolution, the result is filtered using a high-pass filter h′ Hi Perform row convolution to obtain horizontal high-frequency reconstruction subbands

[0117] For the vertical detail coefficient after upsampling Use a low-pass filter h′ Lo After row convolution, the result is filtered using a high-pass filter h′ Hi Perform column convolution to obtain vertical high-frequency reconstruction subbands

[0118] For the diagonal detail coefficient after upsampling Use a high pass filter h′ Hi After column convolution, the result is filtered using a high-pass filter h′ Hi Perform row convolution to obtain diagonal high-frequency reconstruction subbands

[0119] Reconstruct the low frequency subband Horizontal high frequency reconstruction subband Vertical high frequency reconstruction subband Diagonal high frequency reconstruction subband Add together to obtain the restored two-dimensional signal S;

[0120]

[0121] The input visual representation and text representation are processed through wavelet transform and inverse transform operations to capture multi-scale features and local information in images and text. Image-text contrastive learning aims to maximize the mutual information between images and text by comparing the similarities of positive and negative image-text pairs in a batch of data, thereby learning a joint representation. In this process, the features of the image information processed by the interactive wavelet transform module are aligned with the features of the text information processed.

[0122] Step 6.2, during the training process, the image-text contrastive learning loss (ITC) is calculated using a bidirectional objective;

[0123] L ITC =L g2t +L t2g

[0124] Among them, L g2t is the contrast loss from image to text, L t2g is the contrast loss from text to image; its specific calculation formula is as follows:

[0125]

[0126] Among them, B represents the batch size, that is, the number of image-text pairs contained in a batch of data, u is an index in the batch, indicating the image sample currently being processed, It is the feature representation of the image sample with index u after being processed by the image transformer based on wavelet transform. It is the feature representation of the text sample with index k after being processed by the text transformer based on wavelet transform. This means that k is the index of the text sample in the batch corresponding to the image sample u, and τ is the temperature parameter of the softmax function.

[0127] Step 7, connecting the language model in step 5 with the interactive wavelet transform module trained in step 6, guiding the language model to generate text based on visual features to further perform feature alignment training, and obtaining a trained remote sensing image long subtitle generation network;

[0128] The specific method of step 7 is:

[0129] Step 7.1, such as Figure 4 As shown, the language model in step 5 is connected to the interactive wavelet transform module trained in step 6 through a fully connected layer. The fully connected layer adjusts the visual representation output by the interactive wavelet transform module to the same dimension as the text embedding of the language model to obtain the visual representation embedding;

[0130] Step 7.2, add the visual representation embedding obtained in step 7.1 as a prefix to the text embedding to form a new embedding sequence and input it into the language model in step 5, guide the language model to generate text based on visual features for further feature alignment training, and use language modeling loss to further perform feature alignment training on the fully connected layer and wavelet transform layer during the training process. The trained remote sensing image long subtitle generation network is obtained.

[0131] This design allows controlling the information flow between visual and textual information while maintaining interaction between the two. Due to the structural characteristics of the interactive wavelet transform module, the image encoder cannot interact directly with the text tokens generated by the language model. Although the visual representation embedding is added as a prefix, the loss is essentially still a sequence-to-sequence language modeling loss, and the visual representation embedding does not participate in the calculation of the loss.

[0132] Step 8, evaluate the performance of the remote sensing image long subtitle generation network trained in step 7. Considering the limitations of traditional text similarity evaluation indicators in processing complex and rich language outputs, it is difficult to accurately quantify model performance. In this experiment, the Long-CLIP model is used to calculate the matching scores between subtitles generated by different models and remote sensing images. The Long-CLIP model has demonstrated excellent capabilities in the field of multimodal information processing, and can effectively capture the deep semantic associations and fine-grained feature matching between images and texts, and is highly adapted to the remote sensing image subtitle generation task involved in the present invention. Figure 5 As shown, the subtitle content generated by the model of the present invention is highly matched with the remote sensing image content.

[0133] Simulation experiment

[0134] 1. Simulation experiment conditions:

[0135] The hardware test platform of the simulation experiment of the present invention is: CPU 48 cores 192G; graphics card is Ascend 910PremiumA 32G

[0136] The software platform of the simulation experiment of the present invention is: operating system: EulerOS2.0 (SP8); online platform: ModelArts; notebook integrated development environment, Python 3.8.19; hardware platform: Ascend 910A

[0137] 2. Simulation content and result analysis:

[0138] Table 1 is the detailed information of the remote sensing image-text pair dataset; Table 2 is the average image-text matching calculated with the help of the Long-CLIP model; Table 3 is the long caption generation model for different remote sensing images Figure 5 Generated long subtitle description.

[0139] The simulation experiment of the present invention is conducted on a remote sensing image-text pair data set that is collected, produced and collated by the user. The detailed information of the data set used in the present invention is shown in Table 1.

[0140] Table 1 Detailed information of remote sensing image-text pair dataset

[0141]

[0142]

[0143] As shown in Table 1, the dataset contains 496 pairs of remote sensing image-text data, with the text presented in English. The images come from GF3, GF7 optical remote sensing satellite images and Google Earth, and the text is generated with the help of ChatGPT-3.5 under the guidance of semantic segmentation results. The average length of the text is 233 words. The long text length provides sufficient information for the model to learn the correspondence between complex semantics and image details, laying a solid data foundation for subsequent research on image-text cross-modal tasks.

[0144] This experiment uses a self-made data set to compare the performance changes after replacing the interactive wavelet transform module with the self-attention module and the interactive wavelet transform module with the interactive Fourier transform module. The present invention obtains 20 remote sensing images and uses three different remote sensing image long subtitle generation models to generate a total of 60 image-text pairs, and calculates the matching degree between the subtitles generated by different models and the images through the Long-CLIP model, and obtains the average matching degree between the subtitles generated by each model and the image as shown in Table 2. Figure 5 The generated long subtitles are shown in Table 3.

[0145] Table 2 Average image-text matching calculated using the Long-CLIP model

[0146] Model Average matching degree Self-Attention Model 0.8279 Interactive Fourier Transform Model 0.8196 Model of the present invention 0.8432

[0147] Table 3 Different remote sensing image long caption generation models Figure 5 Generated long subtitle description

[0148]

[0149]

[0150] As can be seen from Table 2, in the comparative experiment, the model of the present invention showed obvious advantages. Compared with the average matching degree of 0.8279 of the self-attention model and the average matching degree of 0.8196 of the interactive Fourier transform model, the average matching degree of 0.8432 of the model of the present invention is the highest. This shows that the interactive wavelet transform module adopted by the present invention can more effectively extract and process image features, so that the generated subtitles are more accurately matched with the image content in semantics and details. A higher matching degree means that the model of the present invention can capture more key information when understanding remote sensing images and generating corresponding long subtitles, thereby generating a text description that is more in line with the actual situation of the image.

[0151] From Table 3, we can see that different models have Figure 5The generated long subtitle descriptions differ in content richness and accuracy. The long subtitles generated by the self-attention model clearly describe the main objects in different directions in the image, such as the lake in the upper left corner, the residential area in the lower left corner, the boat on the right, and the small river in the lower right corner. There is a certain spatial layout description of the image scene, but the overall focus is on the presentation of discrete objects. The long subtitles generated by the interactive Fourier transform model are relatively concise, mainly focusing on residential areas and water bodies, and also have a certain description of the areas around the water bodies, but in comparison, the information is relatively limited, and the expression is relatively fragmented, lacking a coherent description of the overall scene. The long subtitles generated by the model of the present invention not only describe in detail the land cover types at different locations of the image, such as developed areas, bare land, waters, etc., but also explain the proportion of each area, grasp the main features of the image as a whole, and the content is richer and more comprehensive. In the understanding and description of the content of remote sensing images, it can provide more in-depth and broad information, and more accurately reflect the geographical information presented by remote sensing images. This further proves that the model of the present invention has a stronger ability to parse and express complex image information in the task of generating long subtitles for remote sensing images.

[0152] The key points and protection points of the present invention are:

[0153] 1. Solve the problems faced by remote sensing image subtitle generation tasks when generating long texts, such as difficulty in effectively aligning images and texts, and difficulty in accurately extracting complex features (such as the position relationship and proportion of ground objects).

[0154] 2. Collect high-resolution remote sensing image data from multiple reliable sources, use the semantic segmentation model DeepLab-V3-Plus to segment the remote sensing images into six major categories, and generate proportional descriptions of the partitioned objects. With the help of ChatGPT-3.5's API interface service, generate text descriptions based on the proportion information of various objects in different directions in the remote sensing images, and then modify the output results through manual verification to produce a high-quality remote sensing image-text pair dataset.

[0155] 3. Use image-text pairs to pre-train the interactive wavelet transform module, by optimizing the image-text contrast learning pre-training objective, the key wavelet transform layer uses db4 wavelet transform to process the input embedding to capture the multi-scale features and local information in the image and text.

[0156] The present invention also provides a remote sensing image long subtitle generation system based on interactive wavelet transform and Transformer, comprising:

[0157] The module for generating information on the proportion of ground objects is used to obtain the remote sensing image data set in step 1, perform semantic segmentation on the acquired remote sensing image data set using a semantic segmentation model, and generate information on the proportion of various ground objects in different positions in the remote sensing image;

[0158] The image-text pair construction module is used to implement step 2 in which the proportion information of various types of objects in different directions in the remote sensing image generated in step 1 is input into the large language model, and the text is generated according to specific requirements to construct the image-text pair;

[0159] An error elimination module is used to implement a comprehensive review of the image-text pairs constructed in step 2 in step 3, and eliminate the image-text pairs with obvious errors;

[0160] A remote sensing image-text pair data set generation module is used to implement step 4 to organize and save the image-text pairs reviewed in step 3 to be used as a remote sensing image-text pair data set;

[0161] A remote sensing image long subtitle generation network construction module is used to implement the construction of a remote sensing image long subtitle generation network in step 5, including: an image encoder, an interactive wavelet transform module and a language model;

[0162] An interactive wavelet transform module training module is used to implement the interactive wavelet transform module built in step 5 using the remote sensing image-text pair data set obtained in step 4 in step 6;

[0163] The remote sensing image long subtitle generation network training module is used to realize the connection of the language model in step 5 with the interactive wavelet transform module trained in step 6 in step 7, guide the language model to generate text based on visual features for feature alignment training, and obtain the trained remote sensing image long subtitle generation network;

[0164] The performance evaluation module is used to evaluate the performance of the remote sensing image long subtitle generation network trained in step 7 in step 8, and calculate the matching scores between subtitles generated by different models and remote sensing images through the Long-CLIP model.

[0165] The present invention also provides a remote sensing image long subtitle generation device based on interactive wavelet transform and Transformer, comprising:

[0166] Memory: a computer program for storing the above-mentioned method for generating long subtitles for remote sensing images based on interactive wavelet transform and Transformer, which is a computer-readable device;

[0167] Processor: used to implement the method for generating long subtitles of remote sensing images based on interactive wavelet transform and Transformer when executing the computer program.

[0168] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the method for generating long subtitles for remote sensing images based on interactive wavelet transform and Transformer.

Claims

1. A remote sensing image long subtitle generation method based on interactive wavelet transform and Transformer, characterized in that: The following steps are involved: Step 1: obtain a remote sensing image dataset, use a semantic segmentation model to perform semantic segmentation on the acquired remote sensing image dataset, and generate information on the proportion of various types of objects in different directions in the remote sensing image; Step 2, input the proportion information of various types of objects in different directions in the remote sensing image generated in step 1 into the large language model, generate text according to specific requirements, and construct image-text pairs; Step 3: conduct a comprehensive review of the image-text pairs constructed in step 2 and remove the image-text pairs that are obviously wrong; Step 4, organizing and saving the image-text pairs reviewed in step 3 to be used as a remote sensing image-text pair dataset; Step 5: Build a remote sensing image long caption generation network, including: image encoder, interactive wavelet transform module and language model; The image encoder is used to convert the input image into high-dimensional visual features to provide visual information; the output high-dimensional visual features and the input text are input into the interactive wavelet transform module; The interactive wavelet transform module is located between the image encoder and the language model. It is connected to the language model using a fully connected layer for information filtering and interaction. It filters irrelevant information from the high-dimensional visual features extracted by the frozen image encoder and enhances the visual representation related to the input text. The interactive wavelet transform module outputs a multi-scale visual feature representation, referred to as a visual representation, which is a high-dimensional feature vector. The language model is used to generate long subtitles of remote sensing images. It receives the multi-scale visual feature representation output by the interactive wavelet transform module, generates a text description corresponding to the image content, and finally outputs the long subtitles of the remote sensing image. Step 6, using the remote sensing image-text pair dataset obtained in step 4 to train the interactive wavelet transform module built in step 5; Step 7, connecting the language model in step 5 with the interactive wavelet transform module trained in step 6, guiding the language model to generate text based on visual features for feature alignment training, and obtaining a trained remote sensing image long subtitle generation network; Step 8, evaluate the performance of the remote sensing image long subtitle generation network trained in step 7, and calculate the matching scores between subtitles generated by different models and remote sensing images through the Long-CLIP model.

2. According to claim 1, a remote sensing image long subtitle generation method based on interactive wavelet transform and Transformer is characterized in that: The specific method of step 1 is: obtaining a remote sensing image data set, performing semantic segmentation on the remote sensing image data set using a semantic segmentation model, and dividing the remote sensing image data set into six major categories, including: water bodies, buildings, roads, woodlands, grasslands, and farmlands; dividing the remote sensing image into five rectangular areas, namely: upper left, upper right, lower left, lower right, and middle; and automatically generating information on the proportion of various types of objects in the remote sensing image in different directions according to grammatical rules, that is, in the remote sensing image, in the [upper left] area: [water bodies] account for [A]%, [buildings] account for [B]%, [roads] account for [C]%, [woodlands] account for [D]%, [grassland] accounts for [E]%, and [farmland] accounts for [F]%; the same description is made for the remaining four areas of upper right, lower left, lower right, and middle.

3. The method for generating long subtitles of remote sensing images based on interactive wavelet transform and Transformer according to claim 1, characterized in that: In step 3, the following three criteria are used to determine whether the image-text pair is obviously wrong: Error judgment of the description of the proportion of land objects: whether the quantity information of various land objects mentioned in the text generated by the large language model matches the proportion information of various land objects in different directions in the remote sensing image generated in step 1. The basis for judging whether they match each other includes but is not limited to: land object categories with a proportion of less than 10% cannot use "largely present" or "dominant"; land object categories with a proportion of more than 50% cannot use "smallly present" or "somewhat visible"; Semantic logic error judgment: Check whether the text generated by the large language model violates basic geographic logic or semantic contradictions. The basis for judging whether there are violations of basic geographic logic or semantic contradictions includes but is not limited to: the text describes a certain area as a large area of ​​water, but at the same time describes the area as having a large amount of farmland without a reasonable transition or explanation, or describes a certain area as having a large number of buildings while there are no roads in the area without a reasonable explanation or description; Verification of language expression accuracy: The text generated by the large language model is checked based on standard English grammar, vocabulary usage specifications, and the accuracy of professional terminology in the remote sensing field. If there are obvious grammatical errors or errors in the use of professional terminology, it is judged as a language expression error.

4. The method for generating long subtitles of remote sensing images based on interactive wavelet transform and Transformer according to claim 1, characterized in that: In step 5: The high-dimensional visual features output by the image encoder are input into the interactive wavelet transform module together with the input text. The interactive wavelet transform module includes parallel image converters and text converters, and shares weights. There are N repeated specific units in both the image converter and the text converter. The specific unit in the image converter includes a wavelet transform layer, a first layer of addition and normalization, a multi-head self-attention layer, a second layer of addition and normalization, a feedforward layer, and a third layer of addition and normalization. The input of the specific unit in the image converter consists of two parts: a high-dimensional visual feature and a learnable visual cue word. After the learnable visual cue word is processed by the wavelet transform layer, the first layer of addition and normalization operation is performed with the learnable visual cue word that has not been processed by the wavelet transform layer. The result is then input into the multi-head self-attention layer together with the high-dimensional visual feature. The output of the multi-head self-attention layer is subjected to a second layer of addition and normalization operation with the result of the first layer of addition and normalization. The output result is processed by the feedforward layer, and then the third layer addition and normalization operation is performed with the result of the second layer addition and normalization. After N cycles, the output of the image converter is obtained; the specific unit in the text converter includes a wavelet transform layer, a first layer addition and normalization, a feedforward layer, and a second layer addition and normalization; the input of the specific unit in the text converter is the input text, and the input text is processed by the wavelet transform layer and then the first layer addition and normalization operation is performed. The output result is processed by the feedforward layer and then the second layer addition and normalization operation is performed with the result of the first layer addition and normalization. After N cycles, the output of the text converter is obtained; the output of the image converter and the output of the text converter are compared and learned by image and text to update the learnable parameters in the interactive wavelet transform module. The interactive wavelet transform module outputs a multi-scale visual feature representation, referred to as visual representation, which is a high-dimensional feature vector.

5. The method for generating long subtitles of remote sensing images based on interactive wavelet transform and Transformer according to claim 1, characterized in that: The step 6 comprises: Step 6.1, the wavelet transform layer in the interactive wavelet transform module uses db4 wavelet transform to process the input embedding; Step 6.2, during the training process, the image-text contrastive learning loss (ITC) is calculated using a bidirectional objective; L ITC =L g2t +L t2g Among them, L g2t is the contrast loss from image to text, L t2g is the contrast loss from text to image; its specific calculation formula is as follows: Among them, B represents the batch size, that is, the number of image-text pairs contained in a batch of data, u is an index in the batch, indicating the image sample currently being processed, It is the feature representation of the image sample with index u after being processed by the image transformer based on wavelet transform. It is the feature representation of the text sample with index k after being processed by the text transformer based on wavelet transform. This means that k is the index of the text sample in the batch corresponding to the image sample u, and τ is the temperature parameter of the softmax function.

6. The method for generating long subtitles of remote sensing images based on interactive wavelet transform and Transformer according to claim 1, characterized in that: The specific method of step 7 is: Step 7.1, connect the language model in step 5 and the interactive wavelet transform module trained in step 6 through a fully connected layer, and the fully connected layer adjusts the visual representation output by the interactive wavelet transform module to the same dimension as the text embedding of the language model to obtain the visual representation embedding; In step 7.2, the visual representation embedding obtained in step 7.1 is added as a prefix to the text embedding to form a new embedding sequence and input into the language model in step 5, guiding the language model to generate text based on visual features for further feature alignment training. The training process uses language modeling loss to perform further feature alignment training on the fully connected layer and the wavelet transform layer to obtain a trained remote sensing image long subtitle generation network.

7. A remote sensing image long subtitle generation system based on interactive wavelet transform and Transformer based on the method according to any one of claims 1 to 6, characterized in that: include: The module for generating information on the proportion of ground objects is used to obtain a remote sensing image dataset, perform semantic segmentation on the acquired remote sensing image dataset using a semantic segmentation model, and generate information on the proportion of various ground objects in different locations in the remote sensing image; The image-text pair construction module is used to input the proportion information of various types of objects in different directions in the remote sensing image into the large language model, generate text according to specific requirements, and construct image-text pairs; The error rejection module is used to conduct a comprehensive review of image-text pairs and reject those with obvious errors; A remote sensing image-text pair dataset generation module is used to organize and save the reviewed image-text pairs to be used as a remote sensing image-text pair dataset; Remote sensing image long subtitle generation network construction module, used to build remote sensing image long subtitle generation network, including: image encoder, interactive wavelet transform module and language model; An interactive wavelet transform module training module, used to train an interactive wavelet transform module using a remote sensing image-text pair dataset; The remote sensing image long subtitle generation network training module is used to connect the language model with the trained interactive wavelet transform module, guide the language model to generate text based on visual features for feature alignment training, and obtain the trained remote sensing image long subtitle generation network; The performance evaluation module is used to evaluate the performance of the trained remote sensing image long subtitle generation network, and calculate the matching scores between subtitles generated by different models and remote sensing images through the Long-CLIP model.

8. A remote sensing image long subtitle generation device based on interactive wavelet transform and Transformer, characterized in that: include: Memory: a computer program storing a method for generating long subtitles of remote sensing images based on interactive wavelet transform and Transformer as described in any one of claims 1 to 6, which is a computer-readable device; Processor: used to implement the method for generating long subtitles for remote sensing images based on interactive wavelet transform and Transformer as described in any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement a remote sensing image long subtitle generation method based on interactive wavelet transform and Transformer as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Remote sensing image description generation method and system

    CN113192030A

  • Remote sensing image subtitle generation method of Transform based on theme guidance

    CN115035508A

Cited By

  • Anaphora video target segmentation method and system based on wavelet correction learning

    CN120931929A