A method for generating remote sensing images from text
By introducing a hierarchical prototype layer and confidence-driven dynamic prototype learning strategy, the problem of difficulty in generating high-quality remote sensing images in the prior art is solved, and higher image quality and detail richness are achieved.
Patent Information
- Application Number
- CN202411364238.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-09-29
AI Technical Summary
The prior art is difficult to generate high-quality, high-resolution remote sensing images, and generating real and detailed remote sensing images from text descriptions is still challenging.
Using a hierarchical prototype layer and confidence-driven dynamic prototype learning strategy, through VQGAN and dynamic hierarchical prototype blocks, we gradually learn and adapt to more prototypes, and improve the richness and accuracy of feature representation of remote sensing images.
The quality and details of the generated remote sensing images are significantly improved, the model shows higher robustness and accuracy when processing complex data, and the generated image content is closer to the real remote sensing image.
Smart Images

Figure CN119379855B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - field of remote sensing technology and artificial intelligence, and specifically relates to a method for generating remote sensing images from text. Background Art
[0002] Remote sensing images are regarded as the "photos" of the earth and are important earth observation data. It can generate a wide range of real - time data remotely and plays a key role in fields such as land and urban planning, environmental monitoring, and military target recognition. With the progress of technology, substantial progress has been made in the understanding of remote sensing images, and the progress of geographic information systems has promoted the visualization of remote sensing data. In addition, the development of machine learning and artificial intelligence, which transforms text descriptions into highly semantically related images, connecting natural language with computer vision, has also promoted the development of artificial intelligence in "understanding".
[0003] As widely recognized in this field, there are significant differences between natural images and remote sensing images in terms of view and content. Remote sensing images present complex and diverse data, requiring a broad understanding of the interactions between various geographical aspects and phenomena. Natural images are recorded from a horizontal perspective, where the foreground and background are clearly visible in the human line of sight, and the main subject in the image often occupies most of the area and is easy to identify. On the other hand, remote sensing images are recorded from a vertical view (bird's - eye view), containing a large number of objects and more complex situations, where the foreground and background are difficult to distinguish, and the size of the main subject in the image is similar to that of the background or even hidden in the background, making it difficult for even humans to detect. These factors make it more difficult to generate remote sensing images than natural images. In addition, with the explosive growth of remote sensing data, it is still a challenge for traditional remote sensing image processing technologies to create the desired content without cumbersome manual operations.
[0004] Despite recent successes in text-to-natural image generation, text-to-high-resolution remote sensing image generation remains challenging. Bejiga et al. proposed the first work on text-to-remote sensing image generation, where conditional GAN was applied to generate very low spatial resolution grayscale remote sensing images from text descriptions of ancient geographical regions. They later improved the text encoding by using a pre-trained Doc2Vec encoder to capture different levels of information in the input text. However, the grayscale remote sensing images generated by these methods have very low spatial resolution and miss many details. Chen et al. further proposed a text-based deep supervised GAN to generate satellite images with a spatial size of 128×128 and used the generated images to augment the training set for change detection tasks. Zhao et al. proposed a structured GAN to synthesize remote sensing images from text with a spatial size of 256×256, emphasizing that structural information is an important factor in evaluating the fidelity of the generated images. Xu et al. proposed using a modern Hopfield network to generate high-resolution remote sensing images from text. However, given the difficulty in identifying and understanding foreground elements in remote sensing images and the complex and diverse spatial distributions of different ground objects, as well as the fact that text descriptions and remote sensing images are distinct modal information, the performance of the above methods is limited, and it is still very difficult to generate high-quality and realistic remote sensing images from text descriptions. Summary of the Invention
[0005] To overcome the deficiencies of the above prior arts, the purpose of the present invention is to provide a method for generating remote sensing images from text, which improves the richness and accuracy of remote sensing image feature representation by introducing a hierarchical prototype layer and a confidence-driven dynamic prototype learning strategy. The model can gradually learn and adapt to more prototypes, enabling the model to exhibit higher robustness and accuracy when processing complex data. To achieve the above technical objectives, the present invention adopts the following technical solutions.
[0006] A method for generating remote sensing images from text specifically includes the following steps:
[0007] S1: Prepare the training dataset RSICD, and obtain text descriptions and corresponding real remote sensing images;
[0008] S2: Train a vector quantization generative adversarial network (VQGAN);
[0009] S3: Encode the text-image pairs through text and image encoders;
[0010] S4: Concatenate the text and image tokens;
[0011] S5: Input the joint concatenated sequence into a dynamic hierarchical prototype block to extract features;
[0012] S6: Adopt a dynamic prototype learning strategy during the training process;
[0013] S7: Use the trained model to generate remote sensing images.
[0014] Further, the specific steps of step S1 are as follows: Extract the text description of each picture into a separate txt file. The input is the original json file of the RSICD dataset, and the output is a txt file containing all text descriptions, all training sample file names, and all test sample file names.
[0015] Further, the specific steps of step S2 are as follows: Pre-generate a discrete numerical where At each encoding position of search for the nearest code in q to generate a variable of the same dimension. Then, use a CNN Encoder to encode based on the already numerically discretized z
[0016]
[0017] The self-supervised loss is as follows:
[0018]
[0019] Among them, the first term in the above formula is the reconstruction loss, and sg(·) is the gradient termination operation. In addition, the adversarial loss in GAN is added, and its loss function can be expressed as:
[0020]
[0021] To sum up:
[0022]
[0023] Further, the specific steps of step S3 are as follows: (1) First, regard each character in the text as an independent token, count the frequencies of all character pairs, select the most frequent characters for merging to form new tokens, and finally repeat the above steps until the predetermined number of tokens is reached. Let n be the maximum length of the input sentence. In the case where the number of words in the input text description is less than n, use 0 as a placeholder to fill the empty tokens, convert the input text to text tokens using Byte Pair Encoding (BPE), and then use a pre-trained word embedding model to convert the text tokens to vector representations. (2) Use a pre-trained image encoder to convert the input image into a series of image tokens, and then convert the image tokens to vector representations. Through the above steps, the text and image information are converted into a unified token representation.
[0024] Further, the specific steps of step S4 are as follows: Concatenate the text tokens and image tokens together in the order of text tokens first and image tokens second to form a combined token sequence, and then add positional encodings to each token to help the model understand the relative positional relationship between the tokens.
[0025] Further, the specific steps of step S5 are as follows: Input the combined token sequence into the LayerScale and PreNorm layers in the first dynamic hierarchical prototype block. The token sequence after normalization and scaling is then input into the Hopfield layer with a temperature parameter introduced. The Hopfield layer stores and retrieves representative prototypes by minimizing the energy function, and then the token sequence processed by the Hopfield layer is input into the hierarchical prototype layer.
[0026] After the token sequence enters the hierarchical prototype layer, it first enters the first Hopfield layer to store and retrieve information, then enters the first self-attention layer to capture the dependencies between different positions in the sequence, then enters the second Hopfield layer to further refine the feature representation, and finally enters the second self-attention layer to enhance the interaction between features. During the forward propagation process, the input data is processed sequentially through these hierarchical structures, and finally outputs the feature representation combined with each layer. The formula is as follows:
[0027] HierarchicalPrototypeLayer(x)=SA 2 (H 2 (SA 1 (H 1 (x)))) (5)
[0028] Where, H 1 and H 2 represent the first and second Hopfield layers respectively, and SA 1 and SA 2 represent the first and second self-attention layers respectively.
[0029] Multiple hierarchical prototype layers are introduced in each dynamic hierarchical prototype block. Each dynamic hierarchical prototype block contains two hierarchical prototype layers and two standard Hopfield layers and self-attention layers. The formula is as follows:
[0030]
[0031] Where, P represents the dynamic hierarchical prototype block, LS represents LayerScale, PN represents PreNorm, HL represents the Hopfiled layer, HPL represents the hierarchical prototype layer, SA represents the self-attention layer, and n blk represents the number of prototypes.
[0032] The token sequence processed by the hierarchical prototype layer is input into the LayerScale and PreNorm layers again for normalization and scaling. Then, the normalized and scaled token sequence is input into the self-attention layer to calculate the similarity between the query, key, and value vectors, capturing the long-range dependencies in the token sequence. Subsequently, the token sequence processed by the first dynamic hierarchical prototype block is successively input into the subsequent dynamic hierarchical prototype blocks.
[0033] Furthermore, step S6 is specifically as follows: During training, the model generates images regularly (every 100 steps) and calculates their confidence. The confidence of the generated images is the average of the maximum values of the prediction probabilities of the model for the generated images. First, the model generates an image based on the input text. Then, the model calculates the Logits value of the generated image and converts the Logits value into a probability distribution through the softmax function. The formula for the softmax function is:
[0034]
[0035] where, z i refers to the i-th logits value. Take the maximum value in the probability distribution as the confidence and calculate its average. The formula for the confidence is:
[0036]
[0037] where, N is the number of samples, and logits i is the logits value of the i-th sample.
[0038] Furthermore, step S7 is specifically as follows: After being processed by multiple dynamic hierarchical prototype blocks, the model generates predicted image tokens These tokens represent the visual features of the generated image, and the generation process can be expressed as: where, f is the generation function of the model, and S is the jointly concatenated sequence. Then, calculate the cross-entropy loss between the predicted image tokens and the original tokens. The cross-entropy loss is used to measure the difference between the predicted tokens and the true tokens, and the cross-entropy loss is defined as:
[0039]
[0040] where, y i is the probability distribution of the original image tokens, is the probability distribution of the predicted image tokens. The predicted image tokens are input into the decoder network. The decoder network converts the image token sequence into actual image pixel values through the learned features, and then through a series of post-processing steps, generates the final high-quality remote sensing image.
[0041] Through the above technical solutions, the method for generating remote sensing images from text provided by the present invention has the following advantages and effects:
[0042] 1. The method of the present invention designs a hierarchical prototype layer, which combines multiple Hopfield layers and self-attention layers to capture richer feature representations. Specifically, these hierarchical structures of the hierarchical prototype layer can capture more complex features and relationships. In particular, when multiple hierarchical prototype layers are added, the impact on the model's generation ability is particularly significant, manifested as a significant improvement in the quality and details of the generated images.
[0043] 2. The method of the present invention designs a dynamic prototype learning strategy to overcome the drawback of the fixed number of prototypes in traditional methods, which cannot adapt to the model's requirements during the training process. Our method dynamically adjusts the number of prototypes according to confidence and training stages during the training process, enabling the model to have the ability of adaptive learning. Experimental results show that the method of dynamically adjusting the number of prototypes is superior to the method of fixed number of prototypes in multiple evaluation metrics, and the quality of the generated images is significantly improved.
[0044] 3. After introducing a temperature parameter into the Hopfield layer, the method of the present invention enables it to more flexibly adjust the smoothness of the normalization exponential function, thus performing the memory and retrieval processes more stably. Description of the Drawings
[0045] Figure 1 is the overall flowchart of a method for generating remote sensing images from text according to the present invention;
[0046] Figure 2 is the algorithm flowchart of a method for generating remote sensing images from text according to the present invention;
[0047] Figure 3 are different variants of the dynamic hierarchical prototype block of a method for generating remote sensing images from text according to the present invention;
[0048] Figure 4 is the remote sensing image generated by a method for generating remote sensing images from text according to the present invention. Detailed Embodiments
[0049] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following will describe the specific implementation of the present invention in further detail by combining examples or the prior art and the accompanying drawings.
[0050] The present invention provides a text-to-image generation algorithm based on text descriptions. The overall flowchart is as Figure 1 , and the algorithm flowchart is as Figure 2 , and the specific implementation is as follows:
[0051] Step 1: Prepare the training dataset RSICD, and obtain remote sensing images and corresponding detailed sentence descriptions;
[0052] The RSICD dataset was initially collected for the remote sensing image captioning task. There are 30 scene classes in the dataset, which contains a total of 10,921 aerial remote sensing images with different resolutions. The size of each image is 224×224 pixels, and each image is accompanied by 5 text descriptions. 8,734 text-image pairs obtained by splitting it are used as the training set of this method, while the remaining 2,187 text-image pairs are used as the test set of this method. And all images are adjusted to a resolution of 256×256.
[0053] Step 2: Train the Vector Quantized Generative Adversarial Network (VQGAN). The model consists of four parts, namely CNN Encoder, CNN Decoder, Codebook, and CNN Discriminator;
[0054] In the second step, we use the Adam optimizer with an initial learning rate of 1e-3 to train the VQGAN model. To optimize the learning process, an "exponential" learning rate decay strategy is adopted, that is, at the end of each training epoch, the learning rate is multiplied by 0.996, and the entire training process is carried out for 1000 epochs. Once the VQGAN is trained, the weight parameters in the encoder network, decoder network, and codebook will be fixed.
[0055] Step 3: Encode the text-image pairs through the text and image encoders, and then concatenate the text and image tokens;
[0056] Text and image embedding: Create embedding layers for text and image respectively to convert them into high-dimensional vector representations. Assume the number of text tokens is N t , and the number of image tokens is E i , and the embedding dimension is d. The embedding matrices are respectively:
[0057] E t = Embedding(N t , d)
[0058] E i = Embedding(N i , d)
[0059] Position embedding: Create position embeddings for text and image to capture the position information in the sequence. Assume the length of the text sequence is N i ,, and the length of the image sequence is L i , then the position embedding matrices are respectively:
[0060] E i= Embedding(N i , d)
[0061]
[0062] VQGAN Freezing: Freeze VQGAN (Vector Quantization Generative Adversarial Network) to prevent it from being updated during training:
[0063] set_requires_grad(Gan, False)
[0064] Step 4, input the joint splicing sequence into the dynamic hierarchical prototype block to extract features;
[0065] First, initialize a dynamic hierarchical prototype block for hierarchical prototype learning in the latent space. Assume the number of prototypes is (N p ), then the parameters of the dynamic hierarchical prototype block module are:
[0066] W p = DPPrototypeBlock(d, L t + L i , n_block, heads, dim_head, N p )
[0067] The hierarchical prototype layer includes four main parts: Hopfield layer: used to store and retrieve information; self-attention layer: used to capture dependencies between different positions in the sequence; the second Hopfield layer: further refine the feature representation; the second self-attention layer: enhance the interaction between features. During the forward propagation process, the input data is processed sequentially through these hierarchical structures, and finally the output combines the feature representations of each layer.
[0068] Combining these definitions, we can represent the forward propagation process of the hierarchical prototype layer as follows: First, input the data x, which is processed by the first Hopfield layer:
[0069] x 1 = Hopfield 1 (x) = HopfieldLayer(x, num_prototype)
[0070] At this step, num_prototype represents the number of prototypes, and the Hopfield layer maps the input data to a high-dimensional space and stores it as prototypes.
[0071] Next, the output x 1 of the first Hopfield layer is processed by the first attention layer:
[0072] x2 = SelfAttention 1 (x 1 ) = SelfAttention(x 1 , dim_head)
[0073] The self-attention layer calculates the similarity between each token in the input sequence and other tokens through heads of the dim_head dimension to aggregate the context information of the text.
[0074] Then, the outputs of the first Hopfield layer and the first sub-attention layer are added together and passed to the second Hopfield layer:
[0075]
[0076] The second Hopfield layer further processes the input data to capture more complex patterns and features.
[0077] Finally, the output x of the second Hopfield layer 3 is processed by the second self-attention layer:
[0078]
[0079] The second self-attention layer further aggregates the context information to enhance the expressive power of the model. Finally, the outputs of all layers are added together to obtain the final output:
[0080] output = x 1 + x 2 + x 3 + x 4
[0081] Step 5, generate images and calculate the confidence during and after the training process;
[0082] Text processing: Ensure that the input text is within the specified length range and convert it into an embedding representation. Assume the input text is T, where P t represents the position of each word in the text:
[0083] T emb = E t (T) + P t
[0084] Generation process: The model gradually generates image tokens. At each step, logits are calculated based on the current text and image tokens, and new tokens are generated through sampling.
[0085] L = softmax(W p · S)
[0086] The above formula represents passing the current sequence S through the weight matrix W p to calculate logits, then converting them into a probability distribution through the softmax function, and then sampling to select the next token and adding it to the sequence. This process is repeated until a complete sequence of image tokens is generated.
[0087] Image decoding: New data is generated by learning the latent distribution of the data, and the generated image tokens are converted into the final image through the decoder. The decoding process is as follows:
[0088] I decoded = vae.decode(I)
[0089] where vae.decode represents the decoder function of the VAE. The role of the decoder is to convert the representation I in the latent space into the final image I decoded .
[0090] Calculating the probability distribution: By calculating the logits of the generated image and using the softmax function to calculate the probability, the confidence of the generated image is evaluated.
[0091] First, by calling the compute_logits function and passing in the text and image data, the formula for Logits is:
[0092] logits = f(T, I)
[0093] T represents the input text data, which is usually embedded in a vector space after preprocessing. I represents the input image data, which can be a single image or a set of multiple images, and is usually converted into a feature vector after preprocessing.
[0094] logits are the raw scores before applying the activation function in the output layer of the neural network. Assuming logits are L, and converting them into a probability distribution P through the softmax function, the formula for calculating the probability distribution is:
[0095] P = softmax(L)
[0096] Evaluating the confidence: Confidence represents the degree of certainty of the model that the generated image belongs to a certain class. The confidence is evaluated by selecting the maximum value in the probability distribution P, and the formula is:
[0097] confidence = max(P)
[0098] The method for generating remote sensing images from text proposed by the present invention is trained on 1 NVIDIA RTX A6000 and 1 NVIDIA RTX 4090 graphics card. The number of dynamic hierarchical prototype blocks is fixed at 10, the model is trained for 1000 epochs, and the batch size is set to 16. The results show that, as Figure 4 , the remote sensing images generated by the method of the present invention have better quality, and the detail information and structure of the ground objects in the images are improved, and the image content is closer to real remote sensing images.
[0099] Table 1
[0100]
[0101] Figure 3 shows variants of various dynamic hierarchical prototype blocks. As can be seen from Table 1, the hierarchical prototype layer of the present invention significantly improves the performance of the model. No matter which layer of the basic prototype block it is introduced into, it can improve the accuracy and robustness of the model. In particular, after adding multiple hierarchical prototype layers, the quality and details of the generated images are significantly improved, verifying its effectiveness and applicability in the task of generating remote sensing images from text.
[0102] Table 2
[0103]
[0104] Table 2 shows the quantitative results of the impact of different dynamic prototype learning strategies on the model. We classify the dynamic prototype learning strategies into four categories: 1. Dynamic prototype learning strategy-a: A confidence-based prototype adjustment strategy. 2. Dynamic prototype learning strategy-b: A training progress-based prototype adjustment strategy. 3. Dynamic prototype learning strategy-c: A fixed interval-based prototype adjustment strategy. 4. Dynamic prototype learning strategy-d: A phased prototype adjustment strategy. The results show that the model with the dynamic prototype learning strategy is superior to the baseline model in all evaluation metrics. Specifically, the cross-entropy loss of the model is lower, indicating that it converges better during the training process, and the IS and FID scores are also greatly improved.
[0105] In view of the problem that it is difficult to generate high-quality and realistic remote sensing images from text descriptions, the present invention proposes a method for generating remote sensing images from text. Experiments were conducted on the RSICD dataset, and significant improvements were observed in the IS index, FID index, and zero-shot classification accuracy OA. Overall, our model performs excellently in generating image quality and handling unseen data, especially leading far ahead in terms of zero-shot classification accuracy. In addition, our model is different from traditional GAN-based methods and instead adopts a Transformer-based method. The Transformer-based method has unique advantages in processing sequence data and capturing long-range dependencies, which may be one of the reasons for the excellent performance of our model in the zero-shot classification task.
[0106] The above embodiments are only used to illustrate the present invention and are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the present invention, and the patent protection scope of the present invention shall be defined by the claims.
[0107] The technical content not described in detail in the present invention is well-known technology.
Claims
1. A method for generating remote sensing images from text, characterized in that: The steps include: S1, prepare the training dataset RSICD, obtain text descriptions and corresponding real remote sensing images; S2. Train a vector quantization generative adversarial network VQGAN; S3, encoding the text-image pair by a text and image encoder; S4, splicing the text and image tags; S5, input the joint concatenation sequence into the dynamic hierarchical prototype block to extract features; specifically, input the joint tag sequence into the first dynamic hierarchical prototype block for standardization and scaling; the standardized and scaled tag sequence is input into the Hopfield layer that introduces the temperature parameter; then the tag sequence processed by the Hopfield layer is input into the hierarchical prototype layer, and the formula of the hierarchical prototype layer is as follows: HierarchicalPrototypeLayer(x)=SA2(H2(SA1(H1(x))))(1) H1 and H2 represent the first and second Hopfield layers, respectively. SA1 and SA2 represent the first and second self-attention layers, respectively. Each dynamic hierarchical prototype block contains two hierarchical prototype layers and two standard Hopfield layers and self-attention layers, and the formula is as follows: Among them, P represents the dynamic hierarchical prototype block, LS represents LayerScale, PN represents PreNorm, HL represents the Hopfiled layer, HPL represents the hierarchical prototype layer, SA represents the self-attention layer, and n blk Indicates the number of prototypes; then the tag sequence processed by the first dynamic hierarchical prototype block is sequentially input into the subsequent dynamic hierarchical prototype blocks; S6, the training process adopts a dynamic prototype learning strategy; S7. Use the trained model to generate remote sensing images.
2. The method for generating remote sensing images from text according to claim 1, characterized in that: In step S2, a discrete value is generated in advance. exist Each encoding position goes to Find the code closest to it, and then use CNN Encoder to encode it:
3. The method for generating remote sensing images from text according to claim 1, characterized in that: In the step S3, each character in the text is first regarded as an independent token, the frequencies of all character pairs are counted, the most frequent characters are selected for merging to form a new token, and the above steps are repeated until a predetermined number of tokens is reached; let n be the maximum length of the input sentence, and when the number of words in the input text description is less than n, 0 is used as a placeholder to fill the empty token, and then the pre-trained word embedding model is used to convert the text token into a vector representation; Use a pre-trained image encoder to convert the input image into image tokens, and then convert the image tokens into vector representations.
4. The method for generating remote sensing images from text according to claim 1, characterized in that: In step S4, the text mark and the image mark are concatenated together in the order of the text mark first and the image mark second to form a joint mark sequence, and then a position code is added to each mark.
5. The method for generating remote sensing images from text according to claim 1, characterized in that: In step S6, during the training process, the model will periodically generate images and calculate their confidences; first, the model generates images based on the input text, then the model calculates the Logits value of the generated image, and then converts the Logits value into a probability distribution through a normalized exponential function. The formula of the normalized exponential function is: Among them, z i Refers to the i-th logits value, takes the maximum value in the probability distribution as the confidence, and calculates its average value. The confidence calculation formula is: Among them, N is the number of samples, logits i is the logits value of the i-th sample; The number of prototypes is dynamically adjusted according to the confidence and training stage: if the confidence of the generated image is lower than the set confidence threshold, the number of prototypes is increased; when the number of prototypes has reached the maximum value and the training has reached a certain stage, the number of prototypes is reduced.
6. The method for generating remote sensing images from text according to claim 1, characterized in that: In step S7, after processing through multiple hierarchical prototype layers, the model generates a predicted image tag The generation process can be expressed as: Among them, f is the generation function of the model, S is the joint concatenation sequence, and then the cross entropy loss between the predicted image label and the original label is calculated. The cross entropy loss is defined as: Among them, y i is the probability distribution of the original image label, It is the probability distribution of predicted image labels; the predicted image labels are converted into image pixel values through the decoder network to generate the final high-quality remote sensing image.
Citation Information
Patent Citations
Remote sensing image content description method based on variational self-attention reinforcement learning
CN111126282A
Remote sensing image description generation method based on comparative learning pre-training
CN117173418A