Image description generation method and system for remote sensing data product image
By integrating the SEDC module and the ResNet101 network to extract the features of the remote sensing data product images, and building a corpus with the BERT model. Using the Transformer model to generate image descriptions, the problems of difficulty in generating content diversity and redundancy control in the existing technology are solved, and the quality and accuracy of generated text are significantly improved.
Patent Information
- Application Number
- CN202510071657.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to ensure the richness of semantics and logical coherence in remote sensing data images at the same time. Especially in diversified scenarios, the diversity and redundancy control of generated content have become key issues, resulting in low quality of generated text.
Using an image description generation method for remote sensing data product images, a corpus built by fusing the image encoder of the SEDC module and a BERT model, combined with the Transformer model, generate multi-angle image descriptions for product images. The method includes acquiring remote sensing data products to form product images, extracting image features through the SEDC module and the ResNet101 network, constructing a corpus and generating image descriptions through the Transformer model.
It significantly improves the accuracy and relevance of image descriptions, makes the generated description closer to the actual sea surface temperature gradient situation, and improves the model's ability to extract remote sensing image features and the quality of generated text.
Smart Images

Figure CN119992326A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing, and in particular relates to an image description generation method and system for remote sensing data product images. Background Art
[0002] Image captioning is an interdisciplinary subject between natural language processing and computer vision. It aims to generate grammatically correct and semantically consistent descriptions based on visual content. The application scope of this technology is very wide, covering multiple scenarios from daily life to professional fields. Although many studies have been devoted to optimizing the effect of image description generation in recent years, generating natural and coherent image descriptions still faces great challenges. Current models often find it difficult to ensure both semantic richness and logical coherence at the same time, especially in diverse scenarios, where the diversity and redundancy control of generated content become key issues. In order to ensure the accuracy of the generated content, sentences must not only have semantic consistency and logical coherence, but also try to avoid the generation of redundant information. Only when these conditions are met can image description generation technology play its maximum role in various application scenarios.
[0003] Remote sensing data images are image data of the earth's surface or the atmosphere, ocean and other environments acquired from a long distance through remote sensing technology. As a key data source in the fields of environmental monitoring, resource management and disaster warning, remote sensing images occupy an indispensable position in modern research. Among them, the ocean is an important object captured by remote sensing, and its complex dynamic characteristics (such as fluctuations, flows, etc.) make it particularly important to generate natural language descriptions for these images. By generating accurate natural language descriptions, not only can the interpretability of the data be effectively improved, but also the ease of use of the data can be significantly improved, making it more valuable in practical applications.
[0004] The core advantage of remote sensing data image text description generation technology is that it enables non-professional users to quickly understand the key information in the image without having a professional background in remote sensing or image processing. The biggest feature of this technology is that it converts the complex visual information in remote sensing images into easy-to-understand text descriptions through natural language, thereby lowering the threshold for image data interpretation. In the past, remote sensing data images usually required professionals to interpret them in combination with technical backgrounds, which often required some knowledge thresholds and analytical capabilities. Through text descriptions, users can directly understand the main content of the image through language, such as ocean gradient analysis, geographic location and other key information, greatly improving the convenience and efficiency of information acquisition.
[0005] At present, the technology of generating text description of remote sensing data images is gradually focusing on how to effectively model the multimodal association between images and texts, especially the ability to handle multi-scale and multi-target in complex remote sensing scenes. One of the current implementation technologies is a memory-guided transformer based on spatial channel attention, which generates image captions through the integration of CNN and Transformer, and deeply understands the multi-scale, multi-shape, and multi-target semantic knowledge of remote sensing; another is to measure the representation of images and captions by embedding them in the same semantic space, and propose a sentence collective representation method to improve the generation performance of the model; another is to introduce an attribute attention mechanism so that the model can better capture the correspondence between the semantic information in the image and the specific object; another is to use the Transformer network with multi-scale residual connections for image description generation, showing good performance on remote sensing datasets.
[0006] The above methods have achieved good results in image description tasks, but these methods are still insufficient in utilizing feature information. Marine remote sensing data images usually include longitude, latitude and time dimensions. Each pixel of the data corresponds to a specific geographical location and time point, and contains a temperature value. There are also high requirements for generating text, specifically to ensure that the generated description can describe the physical phenomena in the image; use geographic coordinates or geographic features to accurately describe the spatial position; clearly describe the content of satellite-captured information, etc. The current model cannot learn different text representations for the same features or multiple features of the same image, resulting in low quality of generated text.
[0007] Therefore, there is an urgent need in this field for an image description generation method for remote sensing data product images that can achieve dense description tasks and improve the quality of generated text. Summary of the invention
[0008] In view of the above deficiencies in the prior art, an object of the present invention is to provide a method and system for generating image descriptions for remote sensing data product images.
[0009] The present invention provides an image description generation method for remote sensing data product images, comprising:
[0010] S1: Acquire remote sensing data products to form product images;
[0011] S2: convolve the product image through the image encoder integrated with the SEDC module to obtain image features;
[0012] S3: Using the BERT model, a corpus is constructed based on the collected original description texts;
[0013] S4: Generate multi-angle image descriptions for the product image through a Transformer model according to the image features and the corpus to obtain a generation result.
[0014] According to the image description generation method for remote sensing data product images provided by the present invention, step S2 further comprises:
[0015] S21: Select the backbone network;
[0016] S22: performing convolution on the product image through the backbone network to obtain a first feature map;
[0017] S23: Establish a SEDC module, and perform convolution on the first feature map through the SEDC module to obtain image features.
[0018] According to the image description generation method for remote sensing data product images provided by the present invention, the backbone network in step S21 is a ResNet101 network.
[0019] According to the image description generation method for remote sensing data product images provided by the present invention, step S23 further includes:
[0020] S231: Perform a dilated convolution on the first feature map to obtain a second feature map;
[0021] S232: Perform global average pooling on the second feature map to obtain a third feature map;
[0022] S233: Outputting a weight vector according to the third feature map through a fully connected layer;
[0023] S234: According to the weight vector, multiply the third feature map by the second feature map to obtain image features.
[0024] According to an image description generation method for remote sensing data product images provided by the present invention, in step S23, the expression for processing the first feature map by the SEDC module is:
[0025]
[0026] in, is the output image feature, x is the first feature map of the input SEDC module, DilatedConv(·) is the dilated convolution operation, σ(·) is the sigmoid function, ReLU is the ReLU function, λ1 is the weight matrix of the first fully connected layer, λ2 is the weight matrix of the second fully connected layer, H is the height value of the input image, W is the width value of the input image, c is the channel index value of the input image, h is the height index value of the input image, and w is the width index value of the input image.
[0027] According to the image description generation method for remote sensing data product images provided by the present invention, step S3 further comprises:
[0028] S31: Collect original description text;
[0029] S32: Segmenting the original description text to obtain a vocabulary sequence including a plurality of vocabulary units;
[0030] S33: vectorizing the vocabulary sequence to obtain an optimized vectorized representation;
[0031] S34: Constructing a corpus based on the vectorized representation.
[0032] According to the image description generation method for remote sensing data product images provided by the present invention, step S32 further includes:
[0033] S321: segmenting the original text to obtain basic word blocks;
[0034] S322: performing fine-grained segmentation on the basic word chunks according to a preset vocabulary to obtain fine-grained word chunks;
[0035] S323: adding sentence markers to the original text sentence including multiple fine-grained word blocks, and generating digital IDs for all fine-grained word blocks and sentence markers to obtain a vocabulary sequence including multiple vocabulary units.
[0036] According to the image description generation method for remote sensing data product images provided by the present invention, step S33 further comprises:
[0037] S331: Convert each vocabulary unit into a vector representation, and introduce CLS and SEP tags to mark the start of the sentence and divide it into clauses, so as to obtain the basic vector representation of the sentence;
[0038] S332: Introduce position encoding into the basic vectorized representation of the description angle marker words of the original text sentence to distinguish the original text sentences described from different angles and obtain an optimized vectorized representation.
[0039] According to the image description generation method for remote sensing data product images provided by the present invention, step S4 further comprises:
[0040] S41: preset target description text;
[0041] S42: Based on the image features, the Transformer model is trained through the corpus and the target description text to obtain a generation model;
[0042] S43: Generate an image description for the product image using the generation model to obtain a generation result.
[0043] The present invention also provides an image description generation system for remote sensing data product images, which is used to execute the image description generation method for remote sensing data product images as described in any one of the above items, including:
[0044] A collection module is used to obtain remote sensing data products to form product images;
[0045] An image feature acquisition module, wherein the image feature acquisition module is configured as an image encoder integrated with a SEDC module, and is used to perform convolution on the product image to obtain image features;
[0046] A corpus construction module, wherein the corpus construction module is configured as a BERT model, and is used to construct a corpus according to the collected original description texts;
[0047] A generation module is configured as a Transformer model, and is used to generate multi-angle image descriptions for product images according to the image features and the corpus to obtain generation results.
[0048] The present invention provides an image description generation method and system for remote sensing data product images. By combining an image encoder and a corpus constructed by a BERT model, the present invention can more accurately capture the image features of product images, and combine rich text description information to generate a description text that is highly matched with the image content, which helps to improve the accuracy and relevance of the image description, and makes the generated description closer to the actual sea surface temperature gradient. Secondly, the present invention uses a ResNet101 network combined with a SEDC module for feature extraction, which can effectively improve the model's ability to extract remote sensing image features. The fused SEDC module can learn richer image features and enhance the model's generalization ability, so that it can cope with product images under different conditions. In addition, the present invention can construct a more sophisticated and rich corpus by performing fine-grained word segmentation and vectorization processing on the original description text. At the same time, the present invention introduces strategies such as position encoding and SEP identification to further optimize the vectorized representation, so that the Transformer model can fully consider the contextual information and angle differences of the text description when generating image descriptions, thereby improving the quality and diversity of the generated text. The method provided by the present invention realizes an automated processing flow from obtaining remote sensing data products to obtaining data product images to generating image descriptions, which greatly saves the time cost of manual annotation and description, is of great significance for the analysis and application of remote sensing data products, and helps to promote the rapid development of fields such as marine environment monitoring and climate change research.
[0049] The present invention provides a new solution for the description generation of remote sensing data product images by combining remote sensing image processing technology and natural language processing technology, and especially provides new ideas and methods for the analysis and application of other similar ocean surface temperature front data. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings are only used to illustrate specific embodiments and are not considered to limit the present invention. In the entire drawings, the same reference symbols represent the same components. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0051] Figure 1 A flow chart of an image description generation method for remote sensing data product images provided by an embodiment of the present invention;
[0052] Figure 2 A schematic diagram of the structure of an image description generation system for remote sensing data product images provided by an embodiment of the present invention.
[0053] Reference numerals:
[0054] 100, collection module; 200, image feature acquisition module; 300, corpus construction module; 400, generation module. DETAILED DESCRIPTION
[0055] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in the field without making creative work should fall within the scope of protection of the present invention.
[0056] Furthermore, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts disclosed in the present invention.
[0057] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the orientation or position relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inside", "outside", etc. is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. The terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a connection between the two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0058] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of methods and systems consistent with some aspects of the present invention as detailed in the appended claims.
[0059] The embodiments of the present invention are described below with reference to the accompanying drawings.
[0060] like Figure 1 As shown, the present invention provides an image description generation method for remote sensing data product images, comprising:
[0061] S1: Acquire remote sensing data products to form product images.
[0062] It should be noted that the image description generation method for remote sensing data product images of the present invention is applicable to the image description generation of all remote sensing data products. In this embodiment, the sea surface temperature gradient is taken as an example for description.
[0063] In a specific embodiment, the present invention studies remote sensing observation data from a certain sea area, and constructs a picture-text pair dataset of ocean gradient remote sensing data based on the South China Sea, and trains and evaluates on this dataset, which contains training images and four corresponding text descriptions.
[0064] The method for obtaining sea surface temperature gradient remote sensing data is to first obtain the sea surface temperature through remote sensing inversion, and then calculate the temperature gradient to obtain the sea surface temperature gradient remote sensing data. The temperature gradient can express the front. The specific ocean gradient remote sensing image uses a color gradient from blue to red to intuitively show the intensity of the temperature gradient in different areas. High gradient areas (such as orange to red) clearly indicate drastic changes in sea surface temperature, while low gradient areas (blue) indicate that the sea water temperature changes are small. Data visualization is used to help quickly identify marine phenomena. The same image contains multi-angle description content, namely geographical location, temperature change, cause, and impact.
[0065] S2: Convolve the product image through the image encoder integrated with the SEDC module to obtain image features.
[0066] Wherein, step S2 further comprises:
[0067] S21: Select the backbone network.
[0068] Among them, the backbone network in step S21 is a ResNet101 network.
[0069] Furthermore, the image encoder module selected by the present invention is based on the ResNet101 architecture. The ResNet101 architecture is a deep and efficient convolutional neural network that is widely used in image feature extraction tasks. The module uses multi-level convolution and pooling operations to gradually extract visual features of different abstract levels from the input image, gradually transitioning from low-level features (such as edges and textures) to high-level features (such as shapes and object composition).
[0070] S22: Convolve the product image through the backbone network to obtain a first feature map.
[0071] Specifically, the encoder captures features of different scales through stacked convolution blocks (including multiple 1×1 and 3×3 convolution operations), and uses adaptive average pooling to achieve feature dimensionality reduction and global feature integration. The choice ensures that the encoder can obtain global spatial feature representation while maintaining fine-grained information.
[0072] S23: Establish a SEDC module, and perform convolution on the first feature map through the SEDC module to obtain image features.
[0073] In the final stage of step S2, the present invention integrates the SE-Dilated Convolution module. After the last feature map of layer 4 of the network model, the module integrates the dilated convolution to effectively expand the receptive field of the network, so that the model can capture a wider range of contextual information.
[0074] Wherein, step S23 further comprises:
[0075] S231: Perform a dilated convolution on the first feature map to obtain a second feature map.
[0076] Different from traditional convolution operations, dilated convolution allows the introduction of holes in the convolution kernel, thereby expanding the coverage of the convolution operation without increasing the number of parameters, allowing the model to examine the input features from a broader perspective.
[0077] S232: Perform global average pooling on the second feature map to obtain a third feature map.
[0078] Furthermore, step S232 mainly performs a Squeeze operation, i.e., a global average pooling layer, which aims to compress the feature map of each channel into a single value, with the purpose of compressing the information of the spatial dimension (height and width) into the channel dimension, so that subsequent operations can perform feature adjustments based on the correlation between channels.
[0079] The introduction of the adaptive average pooling layer makes the output feature dimension of the encoder fixed, thereby enhancing the robustness and versatility of the network when processing input images of different sizes. By design, the adaptive average pooling layer consistently outputs a predefined, fixed-size feature map regardless of the size of the input image. This property ensures the consistency of the feature representation passed to subsequent layers, thereby decoupling the output of the encoder from the resolution of the input image.
[0080] S233: Output a weight vector according to the third feature map through a fully connected layer.
[0081] Specifically, the obtained weight vector is used to adjust the importance of each channel, that is, to weight the features of each channel. The purpose of step S233 is to enhance the features that are useful for the task and suppress the unimportant features.
[0082] Furthermore, the present invention also integrates the SE attention mechanism in the process of dilated convolution. By learning the importance weight of each channel, different channel features are adaptively weighted, so that the network can more flexibly adapt to feature information of different scales and improve the model's ability to understand complex scenes, thereby providing a more powerful feature foundation for subsequent tasks.
[0083] S234: According to the weight vector, multiply the third feature map by the second feature map to obtain image features.
[0084] Furthermore, the weight vector output in step S234 is used for the multiplication operation of step S234, that is, the weight vector is multiplied by the feature map output by the filter to adjust the feature strength of each channel, thereby re-weighting the feature map, so that the network can adaptively adjust the importance of different channels.
[0085] After the Scale operation, the feature map is passed to the output layer, and finally the feature value obtained after processing by the SE-DilatedConvolution module is output, that is, the above-mentioned image feature.
[0086] Among them, in step S23, the expression for processing the first feature map by the SEDC module is:
[0087]
[0088] in, is the output image feature, x is the first feature map of the input SEDC module, DilatedConv(·) is the dilated convolution operation, σ(·) is the sigmoid function, ReLU is the ReLU function, λ1 is the weight matrix of the first fully connected layer, λ2 is the weight matrix of the second fully connected layer, H is the height value of the input image, W is the width value of the input image, c is the channel index value of the input image, h is the height index value of the input image, and w is the width index value of the input image.
[0089] In step S2, for the ResNet101 architecture that introduces the SEDC module, the present invention retrains and optimizes it specifically for the custom data set, so that the model can better adapt to the feature extraction requirements of specific tasks and can learn more targeted feature expressions, thereby generating a more professional model with better final performance in specific application fields.
[0090] S3: Through the BERT model, a corpus is built based on the collected original description text.
[0091] The purpose of step S3 is to build a corpus based on BERT, which can be generally summarized into two key steps: word segmentation and vectorization.
[0092] Wherein, step S3 further comprises:
[0093] S31: Collect original description text.
[0094] S32: Segmenting the original description text to obtain a vocabulary sequence including a plurality of vocabulary units.
[0095] Wherein, step S32 further comprises:
[0096] S321: segmenting the original text to obtain basic word blocks;
[0097] S322: performing fine-grained segmentation on the basic word chunks according to a preset vocabulary to obtain fine-grained word chunks;
[0098] S323: adding sentence markers to the original text sentence including multiple fine-grained word blocks, and generating digital IDs for all fine-grained word blocks and sentence markers to obtain a vocabulary sequence including multiple vocabulary units.
[0099] Specifically, step S32 and the corresponding processes S321 to S323 correspond to the word segmentation process, which is the first step in corpus construction. Its core goal is to convert the original text into a token sequence that can be processed by the BERT model.
[0100] Furthermore, the present invention uses the bert-base-uncased model and adopts the WordPiece word segmentation algorithm. The entire word segmentation process can be broken down into the following steps: First, the input sentence is divided into chunks according to spaces, punctuation marks, and control characters, and all letters are converted to lowercase; secondly, WordPiece segmentation is performed to perform more fine-grained segmentation on each chunk obtained by basic word segmentation. Specifically, the algorithm starts from the beginning of the chunk and tries to find the longest matching substring in the vocabulary. If found, the substring is used as a token and the process is repeated for the rest of the chunk. If not found, the entire chunk is marked as an unknown token `[UNK]`, and participle adjectives such as "ed, ing" and so on will be prefixed with ## in front of these tokens to indicate that it is a subword; secondly, for special tokens, such as the beginning, punctuation, and end of a sentence, the word segmenter will add special tokens at the beginning and end of the sequence; finally, all the tokens that have been segmented are mapped to corresponding digital IDs according to the vocabulary to form the final token ID sequence that is input into the BERT model.
[0101] S33: vectorize the vocabulary sequence to obtain an optimized vectorized representation.
[0102] Wherein, step S33 further comprises:
[0103] S331: Convert each vocabulary unit into a vector representation, and introduce CLS and SEP tags to mark the start of the sentence and divide it into clauses, so as to obtain the basic vector representation of the sentence;
[0104] S332: Introduce position encoding into the basic vectorized representation of the description angle marker words of the original text sentence to distinguish the original text sentences described from different angles and obtain an optimized vectorized representation.
[0105] Specifically, step S33 and the corresponding process from S331 to S332 correspond to the vectorization process, which is the second step of corpus construction. Its core goal is to use the BERT model to convert the token ID sequence into a vector representation containing semantic information.
[0106] Furthermore, firstly, each token ID will be converted into a corresponding token Embedding vector. Secondly, in order to distinguish different sentences, the present invention introduces Segment Embedding to identify the sentence to which each token belongs. Subsequently, the present invention studies the description of different angles and the vocabulary of sentences with different descriptions to achieve dense description tasks, and has a position vectorization representation Position Embedding to better let the model understand the position of each token in the sentence, and obtain the final input vectorization representation Input Embedding as the input of the BERT model. Finally, after multiple layers of encoding, the BERT model will output an output vectorization representation OutputEmbedding sequence with the same length as the input sequence, in which each vector corresponds to a token in the input sequence and contains rich semantic information of the token and its context. These mutually related vectors provide a key semantic representation basis for subsequent corpus applications.
[0107] S34: Constructing a corpus based on the vectorized representation.
[0108] After step S3, based on the basic framework for building a corpus, a corpus is constructed, which provides high-quality, semantically rich vectorized data for subsequent text generation of various natural language processing tasks.
[0109] S4: Generate multi-angle image descriptions for the product image through a Transformer model according to the image features and the corpus to obtain a generation result.
[0110] The key to the image description generation model is how to convert the extracted image features into natural language descriptions. The present invention uses the encoded image features and the generated partial text as input, and finally iteratively predicts and generates a complete dense text description through the Transformer decoder.
[0111] Wherein, step S4 further comprises:
[0112] S41: preset target description text;
[0113] S42: Based on the image features, the Transformer model is trained through the corpus and the target description text to obtain a generation model;
[0114] S43: Generate an image description for the product image using the generation model to obtain a generation result.
[0115] Through the corpus construction in step S3, the text required for the dense description task has been constructed, so in steps S41 to S43, it is first necessary to convert the target text, that is, the correct description corresponding to the image, into a form that the model can understand. The encoded image features serve as the memory input of the decoder, providing the visual information required to generate text. These feature vectors contain rich image semantic information and guide the direction of text generation.
[0116] The structure of the Transformer decoder is as follows: the decoder is composed of 6 identical decoding layers stacked together, each of which contains three core sublayers: self-attention layer, encoder-decoder attention layer, and feedforward neural network. The self-attention layer captures the dependencies between tokens in the target text, calculates the attention weight between each token and all other tokens, and performs weighted averaging, thereby integrating the contextual information within the sentence into the representation of each token. In the process of generating text, in order to prevent the model from obtaining future information in advance, the decoder uses a masking mechanism to mask the sequence information after the current token.
[0117] Subsequently, the encoder-decoder attention mechanism acts as a bridge connecting images and texts. By calculating the attention weights between the generated text and image features, the decoder can dynamically focus on the relevant areas in the image when generating each word. The multi-layer perceptron (MLP) in the last layer, which is composed of multiple fully connected layers with nonlinear activation functions, enables it to learn the complex mapping between attention-weighted features and vocabulary probability distributions.
[0118] This intricate interplay between attention and MLP ensures that the generated description is both linguistically fluent and closely aligned with the visual content of the input image, ensuring that the generated description closely fits the image content, ultimately generating an image description.
[0119] like Figure 2 As shown, the present invention also provides an image description generation system for remote sensing data product images, comprising:
[0120] A collection module 100 is used to obtain remote sensing data products to form product images;
[0121] An image feature acquisition module 200, which is configured as an image encoder integrated with a SEDC module, and is used to perform convolution on the product image to obtain image features;
[0122] A corpus construction module 300, wherein the corpus construction module 300 is configured as a BERT model, and is used to construct a corpus according to the collected original description text;
[0123] The generation module 400 is configured as a Transformer model, and is used to generate multi-angle image descriptions for product images according to the image features and the corpus to obtain generation results.
[0124] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0125] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0126] The present invention provides an image description generation method and system for remote sensing data product images, and constructs a new model named Ocean-T-Cap (Ocean-Transformer-Cap). The Ocean-T-Cap model of the present invention is improved based on the encoder-decoder generation model architecture. The model uses the ResNet-101 model encoder to extract the feature vector of the image, and then fuses the image features and text features through the element-by-element interactive attention mechanism, and finally generates a natural language description through the Transformer decoder.
[0127] Based on the data set selected corresponding to step S1, the present invention also performs a model evaluation on the Ocean-T-Cap model.
[0128] In the image description generation task, the generated text content of the model usually uses statistical principles to evaluate the similarity between the generated text and the target text, such as METEOR, ROUGE, etc., which are often used in text translation tasks. METEOR evaluates the quality of the generated text through a multi-level matching method. First, it performs exact matching and more flexibly captures the semantic similarity between the candidate text and the reference text through sentence changes and synonym recognition. ROUGE mainly evaluates the quality of the generated text based on n-gram matching, focusing on measuring the overlap between the candidate text and the reference text at the short sentence level. ROUGE includes many variants, such as ROUGE-N (evaluating n-gram matching) and ROUGE-L (based on the longest common subsequence). Perplexity is an evaluation of the probability distribution of text data by the language model, which indicates the "perplexity" of the model when predicting the next short sentence. If the model's perplexity is low, it means that the model is more confident in the generation of the text; on the contrary, a higher perplexity indicates that the model has greater uncertainty in the generation process. BLEU is a commonly used automated evaluation indicator that measures the quality of generation by comparing the similarity between the generated text and a set of reference texts. In summary, the present invention can more comprehensively evaluate the quality of text description in the image-to-text task, taking into account both the similarity of short sentences and the accuracy of semantic expression.
[0129] After the model is trained, the generation results are quantitatively evaluated through a series of standardized automatic evaluation indicators. Specifically, the present invention is compared with the CNN+LSTM text generation architecture, and the comparison results are shown in Table 1.
[0130] Table 1 Evaluation of image-generated text generation
[0131]
[0132] As shown in Table 1, the present invention deeply compares the performance characteristics of the two image description models, Ocean-T-Cap and CNN+LSTM. The experimental results show that the Ocean-T-Cap model has a significant advantage in text generation quality. The Ocean-T-Cap model shows a clear leading advantage in all evaluation indicators, including BLEU, METEOR, and ROUGE scores, which shows that the method of the present invention and the corresponding Ocean-T-Cap model can generate more accurate, fluent, and more matching descriptions with the reference text. In addition, the Ocean-T-Cap model shows lower perplexity in all tasks, indicating that it is more accurate in predicting the next word and the model is more stable.
[0133] In addition, in order to evaluate the quality of the model-generated content, the present invention compared the OCEAN-T-CAP model-generated results with the reference content. Although the generated text has some weaknesses, namely, it incorrectly identifies the exact location of the highest temperature gradient (placing it on the west coast instead of the southeast coast of Vietnam) and shifts the focus from vortex formation to broader offshore transport effects, it also shows some advantages.
[0134] Notably, the proposed OCEAN-T-CAP model does accurately capture the relative positions of higher gradients, indicating that it understands their general spatial distribution. Furthermore, the generated text successfully identifies the key mechanisms driving these gradients—the interaction of coastal upwelling and eddy mixing—and presents them in a concise and fluent manner. Although the model lacks in reproducing subtle details such as the curved belt pattern, these limitations are offset by its ability to grasp and expand on the core information, demonstrating its ability to expand on specialized concepts while still retaining a basic grasp of key details.
[0135] The present invention provides a method and system for generating image descriptions for remote sensing data product images, which have the following advantages:
[0136] The present invention integrates the SEDC module and the ResNet101 network to construct an efficient image encoder, which not only significantly improves the feature extraction capability of remote sensing data product images, but also further broadens the network's field of view through the synergy of the dilated convolution in the SEDC module and the SE attention mechanism, and deepens the model's understanding of the complex marine environment, thereby ensuring that the extracted image features are both accurate and targeted.
[0137] In the process of building the corpus, the present invention introduced advanced word segmentation and vectorization technology based on the BERT model. By implementing fine-grained word segmentation processing and cleverly integrating optimization methods such as position encoding, the present invention successfully improved the semantic richness of text data, laid a solid foundation for subsequent natural language processing tasks, and provided higher-quality and more valuable input materials.
[0138] Combining image features with an optimized corpus, the present invention uses the powerful capabilities of the Transformer model to generate highly accurate image descriptions for product images. The Transformer model, with its unique self-attention mechanism and encoder-decoder attention mechanism, can accurately convert image features into natural and fluent language descriptions, making the generated descriptions highly consistent with the image content, greatly enhancing the readability and practicality of the descriptions.
[0139] The present invention greatly shortens the time of manual annotation and description and significantly improves processing efficiency by realizing full automated processing from acquiring remote sensing data products to obtaining data product images and then to generating image descriptions. This is of great significance for meeting the needs of analysis and application of large-scale remote sensing data products, and provides strong support for promoting rapid development in fields such as marine environment monitoring and climate change research.
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the above embodiments, or replace some of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily thought of by those skilled in the art within the technical scope disclosed in the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for generating image descriptions for remote sensing data product images, characterized in that: include: S1: Acquire remote sensing data products to form product images; S2: convolve the product image through the image encoder integrated with the SEDC module to obtain image features; S3: Using the BERT model, a corpus is constructed based on the collected original description texts; S4: Generate multi-angle image descriptions for the product image through a Transformer model according to the image features and the corpus to obtain a generation result.
2. The method for generating image description for remote sensing data product images according to claim 1, characterized in that: Step S2 further comprises: S21: Select the backbone network; S22: performing convolution on the product image through the backbone network to obtain a first feature map; S23: Establish a SEDC module, and perform convolution on the first feature map through the SEDC module to obtain image features.
3. The method for generating image description for remote sensing data product images according to claim 2, characterized in that: The backbone network in step S21 is a ResNet101 network.
4. The method for generating image description for remote sensing data product images according to claim 2, characterized in that: Step S23 further comprises: S231: Perform a dilated convolution on the first feature map to obtain a second feature map; S232: Perform global average pooling on the second feature map to obtain a third feature map; S233: Outputting a weight vector according to the third feature map through a fully connected layer; S234: According to the weight vector, multiply the third feature map by the second feature map to obtain image features.
5. The method for generating image description for remote sensing data product images according to claim 2, characterized in that: In step S23, the SEDC module processes the first feature map as follows: in, is the output image feature, x is the first feature map of the input SEDC module, DilatedConv(·) is the dilated convolution operation, σ(·) is the sigmoid function, ReLU is the ReLU function, λ1 is the weight matrix of the first fully connected layer, λ2 is the weight matrix of the second fully connected layer, H is the height value of the input image, W is the width value of the input image, c is the channel index value of the input image, h is the height index value of the input image, and w is the width index value of the input image.
6. The method for generating image description for remote sensing data product images according to claim 1, characterized in that: Step S3 further comprises: S31: Collect original description text; S32: Segmenting the original description text to obtain a vocabulary sequence including a plurality of vocabulary units; S33: vectorizing the vocabulary sequence to obtain an optimized vectorized representation; S34: Constructing a corpus based on the vectorized representation.
7. The method for generating image description for remote sensing data product images according to claim 6, characterized in that: Step S32 further includes: S321: segmenting the original text to obtain basic word blocks; S322: performing fine-grained segmentation on the basic word chunks according to a preset vocabulary to obtain fine-grained word chunks; S323: adding sentence markers to the original text sentence including multiple fine-grained word blocks, and generating digital IDs for all fine-grained word blocks and sentence markers to obtain a vocabulary sequence including multiple vocabulary units.
8. The method for generating image description for remote sensing data product images according to claim 6, characterized in that: Step S33 further comprises: S331: Convert each vocabulary unit into a vector representation, and introduce CLS and SEP tags to mark the start of the sentence and divide it into clauses, so as to obtain the basic vector representation of the sentence; S332: Introduce position encoding into the basic vectorized representation of the description angle marker words of the original text sentence to distinguish the original text sentences described from different angles and obtain an optimized vectorized representation.
9. The method for generating image description for remote sensing data product images according to claim 1, characterized in that: Step S4 further comprises: S41: preset target description text; S42: Based on the image features, the Transformer model is trained through the corpus and the target description text to obtain a generation model; S43: Generate an image description for the product image using the generation model to obtain a generation result.
10. An image description generation system for remote sensing data product images, used to execute the image description generation method for remote sensing data product images according to any one of claims 1 to 9, characterized in that: include: A collection module is used to obtain remote sensing data products to form product images; An image feature acquisition module, wherein the image feature acquisition module is configured as an image encoder integrated with a SEDC module, and is used to perform convolution on the product image to obtain image features; A corpus construction module, wherein the corpus construction module is configured as a BERT model, and is used to construct a corpus according to the collected original description texts; A generation module is configured as a Transformer model, and is used to generate multi-angle image descriptions for product images according to the image features and the corpus to obtain generation results.