Remote Sensing Image Super-Resolution Method and Product Based on Diffusion Model and Multimodal Large Language Model
By combining pre-trained stable diffusion model and multimodal large language model, the multi-level semantic information enhances the super-resolution method of remote sensing images is solved, and the existing technology lacks detailed capture capabilities at high magnification is achieved, achieving high-quality super-resolution image generation and faster training process.
Patent Information
- Application Number
- CN202510221647.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Existing remote sensing image super-resolution methods are difficult to capture complex details at high magnification, and there are insufficient generalization ability and training time of the diffusion model.
The pre-trained stable diffusion model is used to combine the multimodal large language model. By obtaining the content description information and texture description information of low-resolution remote sensing images, combining global category information, and adaptively integrating prior information using the cross attention module to generate high-quality super-resolution remote sensing images.
Excellent performance in multiple image quality on different data sets is achieved, the quality and resolution of super-resolution images are improved, and training time is shortened.
Smart Images

Figure CN119722462B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a remote sensing image super-resolution method and product based on a diffusion model and a multimodal large language model. Background Art
[0002] Remote sensing images (RSIs) allow the observation of objects on Earth from space, providing a valuable source of information for monitoring the Earth's surface. In recent years, remote sensing images have been widely used in various fields, including environmental monitoring, resource exploration, and land cover classification. However, the acquisition of high-resolution (HR) remote sensing images is often hindered by hardware limitations. And due to atmospheric conditions, the image quality and resolution are impaired. Therefore, there is a strong need for effective methods to improve low-resolution (LR) images and enhance human perception to meet the requirements of various downstream practical applications.
[0003] Remote sensing image super-resolution (RSI SR) involves using computational methods and techniques to improve the resolution of images. The main goal of super-resolution is to reconstruct high-resolution images with rich texture details from low-resolution images, so as to be able to obtain high-resolution images beyond the limitations of various factors. Therefore, remote sensing image super-resolution has become an active research topic in the remote sensing field. Due to the rapid development of deep learning technologies, super-resolution models based on deep neural networks, such as models based on convolutional neural networks (CNNs), models based on transformers, hybrid Transformer-CNN models, and models based on Mamba, etc., have all made significant progress. Although these discriminative methods using feed-forward convolutional networks may be sufficient at low magnification factors in super-resolution, they often fail to capture the complex details required at high magnification factors. To recover visually convincing details, deep generative adversarial networks (GANs) have been applied to super-resolution tasks. Generative adversarial networks (GANs) utilize the adversarial optimization between the generator and the discriminator to leverage the interaction between these two components and drive the generator towards the direction of recovering real images. Generative adversarial networks have demonstrated convincing image generation results. However, existing methods often encounter some limitations because generative adversarial networks require carefully designed regularization and optimization techniques to control optimization instability and mode collapse.
[0004] Recently, the Denoising Diffusion Probability Model (DDPM) has attracted great attention and recognition in image-to-image translation. DDPM includes a diffusion process and a reverse process. The diffusion process injects noise into an image by transferring the data distribution to the latent variable distribution, while the reverse process generates high-quality images by denoising. Due to its principled and well-defined probabilistic diffusion mechanism, DDPM not only overcomes the training instability and mode collapse problems often encountered in Generative Adversarial Networks (GANs), but also surpasses the limitations of the Peak Signal-to-Noise Ratio (PSNR) loss function commonly used in discriminative models. Recent work has introduced diffusion probability models for efficient Remote Sensing Image Super-Resolution (RSI SR), which trains diffusion models from scratch on specific RSI SR datasets. Although it generates more realistic details, due to training on specific datasets, the generalization ability of its model is still not ideal, and the training time of the model is quite long. Summary of the Invention
[0005] This application provides a remote sensing image super-resolution method and product based on a diffusion model and a multi-modal large language model to at least partially solve the above problems.
[0006] In a first aspect of this application, a remote sensing image super-resolution method based on a diffusion model and a multi-modal large language model is provided. The method includes:
[0007] Obtaining content description information and texture description information of a low-resolution remote sensing image based on a multi-modal large language model;
[0008] Obtaining the global category of the low-resolution remote sensing image based on a classifier;
[0009] Processing the low-resolution remote sensing image, the content description information and texture description information of the low-resolution remote sensing image, and the global category of the low-resolution remote sensing image based on a Stable Diffusion model to obtain a super-resolution remote sensing image.
[0010] Optionally, processing the low-resolution remote sensing image, the content description information and texture description information of the low-resolution remote sensing image, and the global category of the low-resolution remote sensing image through a Stable Diffusion model to obtain a super-resolution remote sensing image includes:
[0011] Processing the low-resolution remote sensing image based on a Variational Autoencoder to obtain a latent representation;
[0012] Encoding the content description information and texture description information of the low-resolution remote sensing image based on a text encoder to obtain description encoding information, and the description encoding information interacts with the latent representation through a cross-attention module;
[0013] Encode the global category of the low-resolution remote sensing image based on the category-aware encoder to obtain category encoding information, and modulate the intermediate feature map in the residual block of the U-Net architecture of the stable diffusion model through spatial feature transformation;
[0014] Obtain a noise-free latent representation based on the U-Net architecture, and obtain a super-resolution remote sensing image based on the variational autoencoder.
[0015] Optionally, the classifier and the stable diffusion model are jointly trained based on training samples. During the joint training process, update the parameters of the classifier, update the parameters of the category-aware encoder to be trained, and freeze the other parameters of the pre-trained stable diffusion model. The training samples are sample low-resolution remote sensing image images carrying global category labels.
[0016] Optionally, the joint training process includes the following steps:
[0017] Encode the sample low-resolution remote sensing image image into a sample latent vector through a variational autoencoder , within t steps, gradually introduce noise into the sample latent vector through the diffusion process of the stable diffusion model to generate a sample noisy latent vector , the stable diffusion model integrates the following parameters to predict the added noise: sample low-resolution remote sensing image 、sample noisy latent vector 、sample content description information and sample texture description information , update the parameters of the category-aware encoder by minimizing the difference between the predicted noise obtained based on the stable diffusion model and the actual noise introduced during the diffusion process. The loss function is expressed as follows:
[0018] ;
[0019] Among them, , represents the predicted noise of the stable diffusion model, represents the category-aware encoder to be trained, represents the pre-trained text encoder with frozen parameters, cls represents sample content description information and sample texture description information;
[0020] Train the classifier to be trained based on the sample low-resolution remote sensing image image with global category labels. The loss function is expressed as follows:
[0021] ;
[0022] where N is the number of sample low-resolution remote sensing image images, and M is the number of global categories. y ic denotes the true label of the i th sample low-resolution remote sensing image image. p ic denotes the predicted probability that the i th sample belongs to the c th class;
[0023] The loss function for the joint training of the classifier and the stable diffusion model is expressed as follows:
[0024] ;
[0025] where L denotes the total loss, L D denotes the diffusion loss, L C denotes the classification loss, λ denotes the balance coefficient.
[0026] Optionally, the content description information and texture description information of the low-resolution remote sensing image are encoded based on a text encoder to obtain description encoding information, including:
[0027] The text features of the content description information and texture description information of the low-resolution remote sensing image are respectively extracted through a pre-trained text encoder :
[0028] ;
[0029] The text features of the content description information and texture description information of the low-resolution remote sensing image are concatenated to obtain description encoding information c :
[0030] c = concat( f con , f tex ).
[0031] Optionally, the class-aware encoder includes noise schedule parameters, and the signal-to-noise ratio at each time step is provided during the joint training process of the noise schedule parameters.
[0032] The second aspect of the present application provides a remote sensing image super-resolution device based on a diffusion model and a multi-modal large language model. The remote sensing image super-resolution device based on a diffusion model and a multi-modal large language model includes:
[0033] A description acquisition module for acquiring content description information and texture description information of a low-resolution remote sensing image based on a multi-modal large language model;
[0034] A category acquisition module for acquiring the global category of the low-resolution remote sensing image based on a classifier;
[0035] A super-resolution processing module for processing the low-resolution remote sensing image, the content description information and texture description information of the low-resolution remote sensing image, and the global category of the low-resolution remote sensing image based on a stable diffusion model to obtain a super-resolution remote sensing image.
[0036] The third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the remote sensing image super-resolution method based on a diffusion model and a multi-modal large language model as described in the first aspect of the present invention.
[0037] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the remote sensing image super-resolution method based on a diffusion model and a multi-modal large language model as described in the first aspect of the present invention.
[0038] The fifth aspect of the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, it implements the steps in the remote sensing image super-resolution method based on a diffusion model and a multi-modal large language model as described in the first aspect of the present invention.
[0039] Existing diffusion models are usually trained from scratch on specific remote sensing datasets and only learn the mapping from LR to HR, resulting in limited generalization ability and only overly smoothed results. To address these issues, this application proposes a remote sensing image super-resolution method based on diffusion models and multi-modal large language models. This method uses a pre-trained Stable Diffusion model, which is trained on a large-scale dataset, can achieve stable and detailed super-resolution effects, and has a faster convergence rate. In addition, this application also uses a multi-modal large language model (MLLM) to enhance the model's semantic understanding ability and provide detailed content and texture descriptions. Furthermore, this application introduces an additional classifier that provides global information and concise signals about the LR. Combining the advantages of MLLM can provide content description information and texture description information about low-resolution remote sensing images, and these prior information are adaptively integrated into the diffusion model through the cross-attention module. Finally, by combining the powerful generation ability of the Stable Diffusion model with the semantic understanding ability of MLLM, the method proposed in this application performs excellently in reconstructing remote sensing images and achieves excellent performance in multiple image quality aspects on different datasets. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions of this application, the drawings required for use in the description of this application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0041] Figure 1 It is a schematic diagram of high-level semantic information at different levels of a remote sensing image of the remote sensing image super-resolution method based on diffusion models and multi-modal large language models provided by this application;
[0042] Figure 2 It is a flowchart of the steps of the remote sensing image super-resolution method based on diffusion models and multi-modal large language models provided by this application;
[0043] Figure 3 It is a schematic diagram of the overall architecture of the Stable Diffusion model in the remote sensing image super-resolution method based on diffusion models and multi-modal large language models provided by this application;
[0044] Figure 4 It is a schematic diagram of the model framework of the remote sensing image super-resolution method based on diffusion models and multi-modal large language models provided by this application;
[0045] Figure 5It is a schematic diagram of the super-resolution results of representative methods under different categories of the remote sensing image super-resolution method provided by this application, which is based on a diffusion model and a multi-modal large language model. Detailed implementation manners
[0046] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0047] StableSR is an image super-resolution technology based on a pre-trained diffusion model, aiming to utilize the powerful generation ability of the diffusion model to improve the resolution and quality of real-world images. The following is the core principle of StableSR:
[0048] 1. Utilizing the pre-trained diffusion model, StableSR is based on the pre-trained Stable Diffusion model. By fine-tuning a lightweight time-aware encoder and a controllable feature wrapping module (CFW), it injects low-resolution image information into the diffusion model.
[0049] 2. Time-aware encoder, the time-aware encoder is responsible for generating time-aware features, which are adaptively modulated in different iterations of the diffusion model, so as to better retain the generated prior knowledge. This design not only improves the training efficiency but also enhances the super-resolution performance.
[0050] 3. Controllable feature wrapping module, to address the randomness and information loss problems of the diffusion model, StableSR introduces a controllable feature wrapping module. This module adjusts the output of the diffusion model through the residual connection of multi-scale intermediate features, thereby achieving a balance between fidelity and authenticity.
[0051] 4. Progressive aggregation sampling strategy, to support super-resolution tasks of arbitrary resolutions, StableSR adopts a progressive aggregation sampling strategy. This strategy divides the image into overlapping blocks and uses a Gaussian kernel for fusion in each diffusion iteration, thus achieving a smoother boundary transition.
[0052] StableSR retains the pre-trained diffusion prior without sacrificing its network structure and weights. It only needs to fine-tune some lightweight modulation layers for the super-resolution task. Since it inherits the generation ability of the stable diffusion model and does not need to make any assumptions about degradation, it realizes open-domain image super-resolution. Most current methods attempt to use better and larger models to adapt to the mapping between low-resolution and high-resolution images, regarding the image super-resolution problem as a pure low-level vision restoration problem.
[0053] In this application, a diffusion-based remote sensing image super-resolution method guided by a multi-modal large language model (MLLM) is introduced, which combines the powerful generative ability of the diffusion model and the semantic understanding ability of the MLLM. Previous related research required training the diffusion model from scratch on specific remote sensing super-resolution datasets, while the method of this application utilizes the prior knowledge of the pre-trained stable diffusion model on large datasets. This enables the model of this application to achieve powerful and detailed super-resolution effects and has a faster convergence speed.
[0054] In addition, this application believes that high-level semantic information is crucial for low-level super-resolution tasks, especially in supplementing low-resolution images. Figure 1 The high-level semantic information at different levels of remote sensing images is demonstrated, providing key clues for the super-resolution task. Specifically, in this application, the category information is extracted by a classifier, and the corresponding prior extractor is the classifier. The content information and texture information are extracted by the multi-modal large language model, and the corresponding prior extractor is the multi-modal large language model. The multi-modal large language model shows powerful capabilities in multi-modal understanding. Therefore, this application uses the MLLM to generate descriptive information (including content description information and texture description information) of the low-resolution image as semantic priors. These descriptive information contain both semantic content and detailed texture information, which are cleverly integrated into the super-resolution model to guide the diffusion model to generate high-fidelity and realistic images. This application adopts cross-attention to effectively combine the semantic prior and the generative prior.
[0055] In addition, this application further attempts to introduce global category semantic information (specifically, the category to which the main subject of the picture belongs) during the super-resolution (SR) process. Instead of using a pre-trained classifier, this application proposes a jointly optimized image classifier to achieve generalization. The global category information, as global semantics, is used together with the embedding of the low-resolution image as an embedding in the latent space. This condition can be generated at multiple scales during the denoising process, thereby improving the overall quality of the output image. This application uses three publicly available remote sensing datasets: AID, DOTA, and DIOR to comprehensively evaluate the effectiveness of the method proposed in this application. Through a large number of experiments, it is proved that the method of this application is always superior to existing methods.
[0056] Specifically, the diffusion-based remote sensing image super-resolution method guided by a multimodal large language model proposed in this application combines the advantages of generating priors in StableDiffusion and semantic priors in multimodal large language models to improve the quality of super-resolution images. In this application, the advantages of multimodal large language models and classifiers are utilized to extract multi-level semantic information. The MLLM provides detailed content and texture descriptions, while the classifier provides valuable global category information. These multi-level semantic information are adaptively integrated in the diffusion model through the cross-attention module. This application also designs a category-aware encoder to insert category information into the U-Net.
[0057] Specifically, this application provides a flowchart of the steps of a remote sensing image super-resolution method based on a diffusion model and a multimodal large language model, as Figure 2 shown. Specifically, the remote sensing image super-resolution method based on a diffusion model and a multimodal large language model includes the following steps:
[0058] S201, obtaining content description information and texture description information of the low-resolution remote sensing image based on the multimodal large language model.
[0059] S202, obtaining the global category of the low-resolution remote sensing image based on the classifier.
[0060] S203, processing the low-resolution remote sensing image, the content description information and texture description information of the low-resolution remote sensing image, and the global category of the low-resolution remote sensing image based on the StableDiffusion model to obtain a super-resolution remote sensing image.
[0061] The rapid development of deep learning has led to the development of numerous RSI SR methods. These deep learning-based methods have proven to be superior to traditional techniques, marking a significant advancement in this field. These methods can be divided into three major categories: discriminative-based methods, GAN-based methods, and diffusion-based methods.
[0062] In the embodiments of the present invention, the StableDiffusion model is adopted as the basic diffusion framework. Compared with the diffusion model, the StableDiffusion model has a faster training speed. Traditional diffusion models need to input full-size images into the U-Net during the reverse diffusion process, which results in significant slowdowns when dealing with large-size images and time steps. However, StableDiffusion addresses this drawback by introducing an autoencoder to compress the image into a low-dimensional representation, called the latent representation. This compression enables StableDiffusion to achieve a faster processing speed, making it a more efficient alternative.
[0063] The Stable Diffusion model (SD) is a deep learning architecture based on the Latent Diffusion Model (LDM), mainly used to generate high-quality images according to text prompts. Its network architecture mainly includes the following three core components: the Variational Autoencoder (VAE), U-Net, and Text Encoder. Among them, the role of the VAE is to compress the input image from the pixel space to a low-dimensional latent space. During training, the encoder of the VAE converts an image (such as 512×512×3) into a latent representation (such as 64×64×4), which significantly reduces computational and memory requirements. During the generation process, the decoder of the VAE restores the denoised latent representation to the final image. U-Net is the core component of the Stable Diffusion model, responsible for predicting and removing noise in the latent space. It adopts an encoder-decoder structure, combining Residual Network (ResNet) and Vision Transformers. During the generation process, U-Net receives the noise latent representation and text embedding as inputs, gradually predicts and removes the noise, and finally generates an image that conforms to the text description.
[0064] The Stable Diffusion model usually uses a pre-trained CLIP text encoder to process the input text prompts. CLIP converts the text into embedding vectors, which serve as conditional information to guide the denoising process of U-Net, ensuring that the generated image is consistent with the text description.
[0065] In addition, the Stable Diffusion model also introduces a Scheduler to control the process of adding and removing noise. This architecture significantly improves the generation efficiency and image quality by performing diffusion and denoising in the latent space.
[0066] The overall architecture of the Stable Diffusion model is as Figure 3 shown. First, the input image is encoded using the Variational Autoencoder (VAE) to obtain a latent representation . Then, Gaussian noise is iteratively added, and after t steps, a noisy latent representation z t is obtained. The goal of the training process is to estimate the Gaussian noise t using the UNet architecture based on text or image c at each step .
[0067] UNet combines residual blocks and cross-attention modules at various spatial resolutions to improve its predictions. Text-based prompts are processed through the CLIP text encoder to generate multimodal features, which are then integrated into the cross-attention layers of UNet. Additionally, the time progression of the diffusion steps is embedded into UNet through residual blocks. The optimization process can be expressed as:
[0068] (1)
[0069] where, represents the original Gaussian noise, represents the predicted noise, and iterative refinement leads to obtaining a noise-free latent representation, denoted as which is then reconstructed to the pixel domain using a VAE-based decoder D to obtain the output image . To incorporate text context into the image synthesis process, StableDiffusion employs a cross-attention module. The latent features of the image, denoted as F, and the text-encoded features obtained from the CLIP text encoder, denoted as F c , are transformed through projection layers. This transformation produces a query Q = W Q .F , a key K = W K ·F c , and a value W = W V ·F c , where W Q 、W K 、W V represent the weight matrices for query, key, and value projections respectively. The attention module aggregates the value features through a weighted aggregation process, defined as:
[0070] (2)
[0071] where d represents the output dimension of the key and query features. Then, the latent features are refined based on the results obtained from the attention block.
[0072] In the StableDiffusion model, the cross-attention module (Cross-Attention Mechanism) ensures that the generated image is closely related to the input text description by combining text embeddings with the latent representation of the image.
[0073] The working principle of the cross-attention module specifically includes: The StableDiffusion model uses a pre-trained CLIP model to convert the input text into an embedding vector. These text embeddings interact with the latent representation of the image (generated by the VAE encoder) through the cross-attention layer. In the cross-attention layer, each word of the text embedding calculates an attention score with each position (or patch) of the image latent representation, thereby establishing a semantic association between the text and the image. The cross-attention module generates attention scores by calculating the similarity between the text word embeddings and the image latent representation. These scores represent the importance of each word in the text to each position in the image. For example, in the prompt "a woman with green eyes", the cross-attention module will associate "green" and "eyes" with the corresponding parts in the image, thus ensuring that the generated woman has green eyes.
[0074] In this way, the cross-attention module guides the image generation process to ensure that the generated image is semantically consistent with the text description. It not only focuses on the keywords in the text but also can understand the grammatical structure and context information. For example, by analyzing the cross-attention map, researchers found that certain grammatical structures (such as the relationship between adjectives and nouns) affect the effect of image generation.
[0075] The cross-attention module plays a crucial role in the StableDiffusion model. It enables the model to accurately generate images according to the text description and avoid generating content unrelated to the text. It also provides support for the interpretability of the model. By analyzing the attention map, it can be understood how the model maps the text to the image. In addition, the cross-attention module is also the focus of model fine-tuning (such as LoRA). By adjusting the weights of these layers, stylized generation can be efficiently achieved.
[0076] The cross-attention module ensures that the generation result is highly consistent with the input prompt through the semantic association between the text and the image.
[0077] When training the diffusion model, first, the input image is converted into a latent representation through the encoder of the VAE, and then Gaussian noise is gradually added to the latent representation. This process is called the diffusion process.
[0078] The denoising process is the core of the StableDiffusion model. The goal is to gradually remove the noise through a denoising network (usually U-Net) and restore the original latent representation. U-Net learns the distribution of the noise, predicts the noise at each time step, and gradually restores the clear latent representation.
[0079] In the embodiment of the present invention, a remote sensing image super-resolution framework named based on the diffusion model and the multimodal large language model is proposed to address the challenges of remote sensing super-resolution. As Figure 4As shown in the figure, the remote sensing image super-resolution framework based on the diffusion model and the multi-modal large language model consists of three main modules: the Stable Diffusion module, the MLLM Semantic Prior module, and the Class-Aware Encoder module. First, this application studies the formulation of the super-resolution problem, aiming to lay a solid foundation for the method of this application. Next, this application outlines the remote sensing image super-resolution framework based on the diffusion model and the multi-modal large language model, and details the implementation details of the Stable Diffusion module, the MLLM Semantic Prior module, and the Class-Aware Encoder module.
[0080] Remote sensing image super-resolution is the process of improving the resolution and detail level of remote sensing images. Traditional image super-resolution tasks involve learning the mapping from low-resolution images x to high-resolution images y The mathematical expression is as follows:
[0081] (3)
[0082] Given a low-resolution input, the conditional probability density function of the high-resolution image is usually modeled as a standard Gaussian distribution or a Laplace distribution.
[0083] Among them, the mean is the true super-resolution high-definition image. Since image super-resolution is an ill-posed problem, a single low-resolution image may correspond to multiple high-resolution images, and the model may learn the average of multiple high-resolution images. In addition, the changes in the imaging characteristics of different sensors may prevent the model from reaching the optimal state, resulting in blurred output. Considering the above problems and inspired by the multi-modal large language model (MLLM), this application proposes to introduce additional prior information to enrich the information of low-resolution images, thereby transforming the one-to-many problem into a relatively deterministic mapping. This application introduces the following modeling method:
[0084] (4)
[0085] where con represents additional conditions, that is, prior knowledge. In this application, the condition conIt includes three aspects of description: category, content, and texture. Image category descriptions represent what the image depicts or its theme. They provide basic information about the image content, facilitating the understanding and classification of images. Image content descriptions involve the presence, location, and characteristics of objects or scenes, including their categories, shapes, sizes, etc. It also includes the depiction of people, actions, or emotional expressions in the image. Additionally, it includes environmental backgrounds and scene features, such as indoor or outdoor settings, natural landscapes, or urban landscapes, as well as other relevant features or specific context information related to the image. These descriptions help to more accurately understand and convey the content of the image. Image texture descriptions involve capturing surface features, details, and texture patterns, including features such as smoothness, roughness, texture density, and texture orientation.
[0086] It also covers color features within the image, including brightness, saturation, and hue. Texture and color descriptions provide detailed information about the visual characteristics of the image, helping to more precisely depict and express the appearance and visual perception of the image. Obviously, these multi-level prior descriptions play a crucial role in supplementing detailed information for the super-resolution model. By decomposing the image into category, content, and texture components, this modeling method provides better guidance for the reconstruction results, making the reconstruction results have higher fidelity and reasonable diversity.
[0087] Specifically, step S203 includes the following sub-steps:
[0088] S2031, Process the low-resolution remote sensing image based on the variational autoencoder to obtain a latent representation.
[0089] S2032, Encode the content description information and texture description information of the low-resolution remote sensing image based on the text encoder to obtain description encoding information, and the description encoding information interacts with the latent representation through a cross-attention module.
[0090] S2033, Encode the global category of the low-resolution remote sensing image based on the category-aware encoder to obtain category encoding information, and the category encoding information modulates the intermediate feature map in the residual block of the U-Net architecture of the stable diffusion model through spatial feature transformation.
[0091] S2034, Obtain a noiseless latent representation based on the U-Net architecture, and obtain a super-resolution remote sensing image based on the variational auto-decoder.
[0092] As Figure 4 shown, it shows the overall framework of the remote sensing image super-resolution framework proposed in this application based on the diffusion model and the multi-modal large language model. The diffusion model and the multi-modal large language model are used to address existing challenges. The entire framework consists of three main components: the multi-modal large language model, the stable diffusion model, and the category-aware encoder.
[0093] This framework integrates multiple priors into image super-resolution, including semantic priors generated by MLLMs, priors of pre-trained Stable Diffusion models (SD), and global class priors. It encodes the semantic and class priors to extract features and then adaptively injects them into the denoising process through cross-attention to modulate the latent representation of SD.
[0094] Specifically, this application uses a pre-trained text-to-image (T2I) SD model as the backbone network to improve the training stability and convergence speed. Additionally, in the semantic prior extraction stage, this application uses state-of-the-art MLLMs to obtain content description information and texture description information for low-resolution (LR) images as semantic priors. Then, these priors are input into the CLIP text encoder, resulting in two types of embeddings.
[0095] To effectively combine the semantic priors obtained in the first stage with a well-defined diffusion model, this application designs a cross-attention module. This module promotes the interaction between the semantic priors generated in the first stage and the priors of the SD model.
[0096] Furthermore, this application introduces a class-aware encoder to modulate the U-Net features, aiming to enhance the class awareness of the pre-trained SD model. Finally, the remote sensing image super-resolution framework based on the diffusion model and the multimodal large language model outputs high-quality super-resolution images.
[0097] In recent years, significant progress has been made in the fields of language understanding and generation through the development of large language models (LLMs). These models are trained on vast text corpus datasets, enabling them to generate contextually relevant text. The success of large language models has sparked exploration into incorporating the visual modality into large language models, giving rise to multimodal large language models (MLLMs). In the field of multimodal large language models, images are typically processed using pre-trained visual encoders (such as vision transformers) and combined with text annotation embeddings within the large language model framework. This fusion of visual and text information extends the semantic understanding capabilities of large language models, making multimodal large language models perform well in generating responses based on visual inputs.
[0098] Therefore, this application uses LLaVA to perceive low-resolution (LR) images and extract semantically meaningful information. LLaVA effectively utilizes comprehensive annotations including captions and bounding boxes to generate high-quality captions as well as question-answer pairs directly related to the images. In the method of this application, the goal is to enhance the semantic understanding of the multimodal language model from the perspectives of content and texture by formulating appropriate instructions. For example, the instruction is: #Task: Remote Sensing Image Content Description. Create a concise description that depicts the main content of the provided remote sensing image. Focus on the key features and keep the description within 100 words. #Guidelines: Identify the main subject or object. Describe the scene type and any significant details. Be accurate and concise. #Example Output: "The satellite image shows a densely populated urban area with a central park surrounded by high-rise buildings and a road network. A river can be seen in the suburbs."
[0099] The output obtained is: A satellite image shows a suburban community characterized by single-family homes with well-defined lots and gardens. The area includes intersecting streets, driveways, and some green spaces. Cars are parked along the roads, indicating residential activity.
[0100] Another example: The instruction is: #Task: Remote Sensing Image Texture Description. Create a short description that captures the texture of the provided remote sensing image. Highlight the visual patterns and surface features within 100 words. #Guidelines: - Describe the texture patterns and surface features. - Mention any tonal or color variations that affect the texture. - Convey the essence of the texture in vivid language. #Example Output: "The image features a blend of textures: a patchwork of farmland, the smooth surface of a lake, and the rough texture of distant mountains."
[0101] The answer is: In this aerial view, a lush courtyard brings soft and irregular shapes. The intersecting roads provide a linear contrast, while the tree line under the shade enriches the scene with deeper tones, integrating the artificial order with the natural vitality.
[0102] In the RSI SR task, the integration of auxiliary image data is crucial for achieving high-quality output. Therefore, this application adds a novel class-aware encoder to the diffusion architecture. The encoder is designed to provide important context clues about the class of the subject depicted in the low-resolution image to the diffusion model.
[0103] The method proposed in this application first uses a classifier to extract relevant features from the low-resolution image. To improve the prediction accuracy, this application uses ResNet-50 as the classifier of this application, and trains it with the low-resolution image as the input to enhance the fidelity and context understanding ability of the super-resolution process. The objective function is the cross-entropy loss, as follows:
[0104] (5)
[0105] Where N is the number of sample low - resolution remote sensing image images, M is the number of global classes, y ic denotes the true label of the i th sample low - resolution remote sensing image image, p ic denotes the predicted probability that the i th sample of the classifier to be trained belongs to the c th class;
[0106] Using the output of the class - aware encoder, the intermediate feature maps in the residual blocks of the U - Net architecture are modulated through spatial feature transformation. In addition, this application integrates temporal information into the encoder to enhance the overall quality of the generated images.
[0107] In the early steps of the process, when the generated images contain limited information and low signal - to - noise ratio, the stronger constraints provided by the class - aware encoder provide more useful guidance for the diffusion process, improving the understanding of the information to be reconstructed. In contrast, as the diffusion process progresses and more image information is reconstructed, the constraints from the LR images can be adjusted to be less dominant. This flexibility is facilitated by treating the noise schedule as a hyperparameter, which provides the signal - to - noise ratio for each time step during training. By considering the temporal aspect, this application strikes a balance in leveraging the guidance of the LR images throughout the diffusion process. In addition, introducing temporal conditioning during the diffusion process helps balance the use of low - resolution images. By combining the class - aware encoder and integrating temporal information, the method of this application not only improves the quality of image reconstruction but also outperforms existing state - of - the - art super - resolution models, as demonstrated by the comprehensive experimental evaluations of this application.
[0108] In this application, two consecutive cross - attention layers are deployed to integrate additional semantic priors, enhancing the model's ability to utilize rich context information. These priors are generated by the MLLM and incorporated into the diffusion model.
[0109] Given an image x with aspect - ratio distortion, a multi - modal language model (MLLM) generates a content description and a texture description . This application uses the pre - trained text encoder in CLIP to extract text features:
[0110] (6)
[0111] For simplicity, this application is along the sequence dimension c= concat( f con , f tex ) directly concatenates these two types of features. Then, the present application enables the interaction between the text features and the latent features of the diffusion model. For the k+1 layer x k+1 of the latent features, cross-attention is adopted as follows:
[0112] (7)
[0113] In this way, after providing sufficient semantic prior descriptions, cross-attention adaptively assigns weights to each word. It allows for the adaptive selection of semantic features. The effectiveness of this setting is further confirmed by ablation studies.
[0114] In the present application, the classifier and the Stable Diffusion model are jointly trained based on training samples. During the joint training process, the parameters of the classifier are updated, the parameters of the category-aware encoder to be trained are updated, and the other parameters of the pre-trained Stable Diffusion model are frozen. The training samples are low-resolution remote sensing image images carrying global category labels.
[0115] Specifically, in the present application, during the joint training process of the model, first, the LR image is encoded into a latent vector . Subsequently, within t steps of random numbers in the range of , noise is gradually introduced into this vector through the diffusion process to generate a noisy latent state . The remote sensing image super-resolution framework based on the diffusion model and the multi-modal large language model integrates the following parameters to predict the added noise: the low-resolution image , the noisy latent state , the content prompt and the texture prompt, to predict the added noise.
[0116] The optimization objective of this model is to minimize the difference between the predicted noise and the actual noise introduced during the diffusion process. This objective can be expressed as follows:
[0117] (8)
[0118] where and represents the noise prediction of the Stable Diffusion model. is the category-aware encoder to be optimized. is a frozen CLIP encoder. In addition, the final loss function of the model in the present application is designed as follows:
[0119] (9)
[0120] Among them, L represents the total loss, L D represents the diffusion loss, L C represents the classification loss, λ represents the balance coefficient. To reduce the computational overhead and training time, all parameters of the stable diffusion model are frozen during training.
[0121] In the test phase, the present application adopts a classifier-free guidance strategy to improve the quality of the generated images without additional training. This strategy involves using the negative prompts of the classifier-free to guide the diffusion model. At each step, predictions are made using the prompts generated from the multimodal large language model. Meanwhile, the prompt is replaced with the negative prompt, denoted as c neg . The prediction results generated using the two prompts are fused to obtain the final output.
[0122] In the experiment, a combination of negative words such as "blurry, dotty, noisy, unclear, low resolution, over-smoothed" is used as the negative prompt to generate higher-quality images. This method enables the present application to utilize the power of negative prompts without a separate classifier, ultimately improving the entire image generation process.
[0123] In the remote sensing super-resolution task, the present application uses three popular datasets to evaluate the performance of the model of the present application, namely AID, DOTA, and DIOR. According to the experimental settings, the training set contains 3000 images randomly selected from the AID dataset, and the size of each image is 640×640. Specifically, the present application selects 100 images from each of the 30 categories of AID to construct the training set. In addition, the present application creates a test set for each category, with each category containing 10 images, ensuring no overlap with the training set, resulting in a total of 300 test images. In addition, the present application also uses subsets of the DOTA and DIOR datasets for testing, containing 700 and 1000 images respectively. The resolution of these images is 512×512. Therefore, the test set of the present application contains a total of 2000 images. In the simulation experiment of the present application, the bicubic interpolation method is used for image degradation.
[0124] In this application, the pre-trained SD-v1.5 model is adopted as the text-to-image (T2I) model. The AdamW optimizer is used to fine-tune the remote sensing image super-resolution framework model based on the diffusion model and the multi-modal large language model, with a weight decay of 0.01 and 100,000 iterations. The training parameters include a batch size of 32 and a learning rate of 5e-5, and all operations are carried out on an NVIDIA A100 GPU. For image generation, this application uses DDPM sampling over 50 time steps, and the balance coefficient λ is set to 0.05.
[0125] To comprehensively evaluate the performance of the remote sensing image super-resolution method, this application uses four metrics commonly used in the SRI SR task to evaluate the quality of the generated images: Frechet Inception Distance (FID), Learned Perceptual Image Patch Similarity (LPIPS), Deep Image Structure and Texture Similarity (DISTS), and the widely recognized Peak Signal-to-Noise Ratio (PSNR). Specifically, FID is used to capture the similarity between the generated images and the real images. LPIPS and DISTS are used for perceptual quality assessment. PSNR is used to evaluate the fidelity of image reconstruction. In addition, considering the lack of real images in actual experiments, this application also adopts the Natural Image Quality Evaluator (NIQE) as a no-reference metric. This metric enables this application to gain an in-depth understanding of the perceptual quality and high-frequency details of the generated images.
[0126] This application conducts a comparative analysis of the remote sensing image super-resolution framework model based on the diffusion model and the multi-modal large language model of this application with the state-of-the-art (SOTA) single-image super-resolution methods, including EDSR, RCAN, HAT-L, MSRGAN, ESRGAN, SPSR, SR3, IRSDE, and EDiffSR. These methods are carefully selected in this application to ensure a comprehensive evaluation. Specifically, EDSR, RCAN, and HAT-L are discriminative model-based models that adopt wide convolutional neural networks (CNNs), Transformer architectures, and their variants, respectively. While MSRGAN is a generative adversarial network (GAN)-based method, which is a modified version of SRGAN that removes the batch normalization (BN) layer to avoid artifacts and uses the same perceptual loss as ESRGAN
[77] . SPSR combines a carefully designed gradient loss to preserve structural details and has shown good performance. On the other hand, SR3 first adopts a denoising diffusion probabilistic model to iteratively improve the noisy output. IRSDE and EDiffSR are recently proposed diffusion methods.
[0127] 1) Quantitative comparison: Based on the performance of different methods on various categories of the AID dataset. In this application, FID is used as the quantitative evaluation metric. A lower FID indicates a smaller difference between the generated image and the real image, meaning a higher quality of the generated image and a great similarity to the real image. Generally speaking, the generation methods based on GANs are significantly superior to those based on discriminative models. This is because the adversarial training method of GANs helps to generate more texture details. In addition, the diffusion model has an advantage over the GAN method. The fundamental reason is that the diffusion model has a stronger distribution modeling ability and can better capture the distribution of high-dimensional images. According to the experimental results, it can be determined that the remote sensing image super-resolution framework based on the diffusion model and the multimodal large language model in this application performs best in most categories. It is worth noting that the remote sensing image super-resolution framework based on the diffusion model and the multimodal large language model ranks among the top two in almost all categories, which makes the method in this application reach the SOTA level in terms of the average FID of all categories. Specifically, the method in this application is 1.56 higher than the second-ranked method EDiffSR in terms of FID, which verifies the effectiveness of the method proposed in this application. Although both EDiffSR and the method in this application belong to the diffusion model, the reason why the method in this application performs better is that it utilizes more prior information, including the pre-trained diffusion model, the text and content descriptions of the image, and the category information. These prior knowledge can supplement the missing information in the LR input and provide richer input conditions, thus improving the performance.
[0128] In addition, this application evaluated the FID, LPIPS, DISTS, and NIQE metrics of each method on three datasets (AID, DOTA, and DIOR). The experimental results show that the method in this application achieved the best performance in all three metrics. It is worth noting that the GAN-based methods always achieved the best results in terms of LPIPS. The reason behind this may be that these methods incorporated the LPIPS loss during training. Among all the diffusion-based methods, the method in this application is superior to other methods. DISTS is used to measure the structure and details of an image, and the method in this application achieved the lowest DISTS score. This indicates that the method in this application can preserve the image structure and restore fine details.
[0129] 2) Qualitative comparison: To visually evaluate the effectiveness of different methods, this application presented a visualization of the super-resolution images. Figure 5 The super-resolution results of representative methods under different categories are presented. In Figure 5Among them, the image samples from the first row to the last row are respectively: railwaystation_17, playground_249, school_20, resort_20, farmland_89, square_156, viaduct_109, storagetanks_93. The images of each column from left to right are processed by Bicubic, HAT-L, ESRGAN, EDiffSR and the method of this application. It can be seen that generally speaking, the method of this application has achieved the best results in retaining the structure and restoring the texture, and is closer to the real image. For example, in the playground_249 sample, the method of this application can restore the line markings on the football field. In contrast, the discriminative method and the GAN-based method destroy the structure of the image, and the line markings generated by the diffusion model-based EDiffSR method appear blurred. The same phenomenon has also been observed in other categories in this application. The method of this application performs more prominently on images with high structural texture. This application attributes this to the fact that the text description provides prior knowledge for the diffusion model, which helps to retain the structure.
[0130] It can be seen that this application proposes a remote sensing image super-resolution method based on diffusion for multi-modal large language models. Its core idea is to introduce model prior and semantic prior to alleviate the ill-posed one-to-many problem in image super-resolution. To enhance the robustness of the model, this application directly inherits the pre-trained stable diffusion model on a large-scale dataset. This application only fine-tunes a small number of parameters to adapt the model to the remote sensing super-resolution task. To provide more image information, this application uses a multi-modal large language model to generate content and texture descriptions, providing clues for the model to fill in texture details. This application also combines image category information to enable the model to perform advanced understanding on low-resolution inputs. This application has conducted experiments on three datasets, and the results consistently demonstrate the effectiveness of the proposed method.
[0131] Based on the same inventive concept, this application also provides a remote sensing image super-resolution device based on a diffusion model and a multi-modal large language model. The remote sensing image super-resolution device based on a diffusion model and a multi-modal large language model includes:
[0132] A description acquisition module, configured to acquire content description information and texture description information of a low-resolution remote sensing image based on a multi-modal large language model;
[0133] A category acquisition module, configured to acquire the global category of the low-resolution remote sensing image based on a classifier;
[0134] A super-resolution processing module for processing the low-resolution remote sensing image, the content description information and texture description information of the low-resolution remote sensing image, and the global category of the low-resolution remote sensing image based on a stable diffusion model to obtain a super-resolution remote sensing image.
[0135] Optionally, the super-resolution processing module includes:
[0136] A first processing sub-module for processing the low-resolution remote sensing image based on a variational autoencoder to obtain a latent representation;
[0137] A first encoding sub-module for encoding the content description information and texture description information of the low-resolution remote sensing image based on a text encoder to obtain description encoding information, and the description encoding information interacts with the latent representation through a cross-attention module;
[0138] A second encoding sub-module for encoding the global category of the low-resolution remote sensing image based on a category-aware encoder to obtain category encoding information, and the category encoding information modulates the intermediate feature map in the residual block of the U-Net architecture of the stable diffusion model through a spatial feature transformation;
[0139] A second processing sub-module for obtaining a noiseless latent representation based on the U-Net architecture and obtaining a super-resolution remote sensing image based on a variational auto-decoder.
[0140] Optionally, the classifier and the stable diffusion model are jointly trained based on training samples. During the joint training process, the parameters of the classifier are updated, the parameters of the category-aware encoder to be trained are updated, and other parameters of the pre-trained stable diffusion model are frozen. The training samples are sample low-resolution remote sensing image images carrying global category labels.
[0141] Optionally, the joint training process includes the following steps:
[0142] Encoding the sample low-resolution remote sensing image image into a sample latent vector through a variational autoencoder , within t steps, gradually introducing noise into the sample latent vector through the diffusion process of the stable diffusion model to generate a sample noisy latent vector , the stable diffusion model integrates the following parameters to predict the added noise: the sample low-resolution remote sensing image , the sample noisy latent vector , the sample content description information and the sample texture description information , updating the parameters of the category-aware encoder by minimizing the difference between the predicted noise obtained based on the stable diffusion model and the actual noise introduced during the diffusion process. The loss function is expressed as follows:
[0143] ;
[0144] Among them, , represents the predicted noise of the stable diffusion model, represents the class-aware encoder to be trained, represents the pre-trained text encoder with frozen parameters, cls represents the sample content description information and sample texture description information;
[0145] Train the classifier to be trained based on the sample low-resolution remote sensing image with global class labels, and the loss function is expressed as follows:
[0146] ;
[0147] Among them, N is the number of sample low-resolution remote sensing images, and M is the number of global classes. y ic represents the i th true label of the sample low-resolution remote sensing image, p ic represents the i th sample of the classifier to be trained belonging to the c th class prediction probability;
[0148] The loss function for the joint training of the classifier and the stable diffusion model is expressed as follows:
[0149] ;
[0150] Among them, L represents the total loss, L D represents the diffusion loss, L C represents the classification loss, λ represents the balance coefficient.
[0151] Optionally, the second encoding sub-module is specifically used for:
[0152] Extract the content description information of the low-resolution remote sensing image and texture description information respectively through the pre-trained text encoder
[0153] ;
[0154] For the content description information of the low-resolution remote sensing image and texture description information Concatenate the text features to obtain the description coding information c :
[0155] c = concat( f con , f tex ).
[0156] Optionally, the category-aware encoder includes a noise schedule parameter that provides the signal-to-noise ratio for each time step during the joint training process.
[0157] Based on the same inventive concept, the present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the steps in the remote sensing image super-resolution method based on the diffusion model and the multi-modal large language model as described in any of the above embodiments.
[0158] Based on the same inventive concept, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the remote sensing image super-resolution method based on the diffusion model and the multi-modal large language model as described in any of the above embodiments.
[0159] Based on the same inventive concept, the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the steps in the remote sensing image super-resolution method based on the diffusion model and the multi-modal large language model as described in any of the above embodiments.
[0160] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0161] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0162] This application is described with reference to the flowcharts and / or block diagrams of methods, terminal devices (apparatus), and computer program products according to this application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable terminal devices generate a device for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0163] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0164] These computer program instructions can also be loaded onto a computer or other programmable terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0165] Although the preferred embodiments of this application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of this application.
[0166] Finally, it should also be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising said element.
[0167] The above has introduced in detail a remote sensing image super-resolution method based on a diffusion model and a multi-modal large language model provided by the present invention. Specific examples are used in this application to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A remote sensing image super-resolution method based on a diffusion model and a multimodal large language model, characterized in that: The method comprises: Obtain content description information and texture description information of low-resolution remote sensing images based on a multimodal large language model; Obtaining a global category of the low-resolution remote sensing image based on a classifier; Processing the low-resolution remote sensing image, content description information and texture description information of the low-resolution remote sensing image, and the global category of the low-resolution remote sensing image based on a stable diffusion model to obtain a super-resolution remote sensing image; The low-resolution remote sensing image, the content description information and the texture description information of the low-resolution remote sensing image, and the global category of the low-resolution remote sensing image are processed by a stable diffusion model to obtain a super-resolution remote sensing image, including: Processing the low-resolution remote sensing image based on a variational autoencoder to obtain a potential representation; Encoding the content description information and the texture description information of the low-resolution remote sensing image based on a text encoder to obtain description encoding information, wherein the description encoding information interacts with the latent representation through a cross-attention module; Encoding the global category of the low-resolution remote sensing image based on a category-aware encoder to obtain category encoding information, wherein the category encoding information modulates an intermediate feature map in a residual block of a U-Net architecture of a stable diffusion model through a spatial feature transform; A noise-free latent representation is obtained based on the U-Net architecture, and a super-resolution remote sensing image is obtained based on a variational autodecoder; The classifier and the stable diffusion model are obtained by joint training based on training samples. During the joint training process, the parameters of the classifier are updated, the parameters of the category-aware encoder to be trained are updated, and other parameters of the pre-trained stable diffusion model are frozen. The training samples are sample low-resolution remote sensing images carrying global category labels. The joint training process includes the following steps: Encode the sample low-resolution remote sensing image into a sample latent vector through a variational autoencoder , within t steps, the noise is gradually introduced into the sample latent vector through the diffusion process of the stable diffusion model to generate the sample noise latent vector , the stable diffusion model integrates the following parameters to predict the added noise: sample low-resolution remote sensing image , sample noise latent vector , sample content description information and sample texture description information , the parameters of the category perceptual encoder are updated based on the minimization of the difference between the predicted noise obtained by the stable diffusion model and the actual noise introduced in the diffusion process. The loss function is expressed as follows: ; in, , represents the prediction noise of the stable diffusion model, represents the category-aware encoder to be trained, represents a pre-trained text encoder with frozen parameters, cls Represents sample content description information and sample texture description information; c Represents description coding information, which is composed of the content description information of the low-resolution remote sensing image and texture description information The text features of are concatenated; The classifier to be trained is trained based on sample low-resolution remote sensing images with global category labels, and the loss function is expressed as follows: ; Among them, N is the number of sample low-resolution remote sensing images, M is the number of global categories, y ic Indicates i The true labels of the sample low-resolution remote sensing images, p ic Represents the classifier to be trained i The samples belong to c The predicted probability of each category; The loss function of the joint training of the classifier and the stable diffusion model is expressed as follows: ; in, L represents the total loss, L D represents the diffusion loss, L C represents the classification loss, λ Represents the balance coefficient.
2. The remote sensing image super-resolution method based on diffusion model and multimodal large language model according to claim 1, characterized in that: The content description information and texture description information of the low-resolution remote sensing image are encoded based on a text encoder to obtain description encoding information, including: Through pre-trained text encoder Extract the content description information of the low-resolution remote sensing image respectively and texture description information Text features: ; Content description information of the low-resolution remote sensing image and texture description information The text features are concatenated to obtain the description encoding information c : c = concat( f con , f tex )。 3. The remote sensing image super-resolution method based on diffusion model and multimodal large language model according to claim 1, characterized in that: The class perceptual encoder includes noise schedule parameters that provide a signal-to-noise ratio value for each time step during a joint training process.
4. A remote sensing image super-resolution device based on a diffusion model and a multimodal large language model, characterized in that: The remote sensing image super-resolution device based on the diffusion model and the multimodal large language model includes: Description acquisition module, used to obtain content description information and texture description information of low-resolution remote sensing images based on a multimodal large language model; A category acquisition module, used for acquiring the global category of the low-resolution remote sensing image based on a classifier; A super-resolution processing module, for processing the low-resolution remote sensing image, the content description information and the texture description information of the low-resolution remote sensing image, and the global category of the low-resolution remote sensing image based on a stable diffusion model to obtain a super-resolution remote sensing image; The super-resolution processing module comprises: A first processing submodule is used to process the low-resolution remote sensing image based on a variational autoencoder to obtain a potential representation; A first encoding submodule is used to encode the content description information and the texture description information of the low-resolution remote sensing image based on a text encoder to obtain description encoding information, and the description encoding information interacts with the potential representation through a cross attention module; A second encoding submodule is used to encode the global category of the low-resolution remote sensing image based on a category-aware encoder to obtain category encoding information, wherein the category encoding information modulates an intermediate feature map in a residual block of a U-Net architecture of a stable diffusion model through a spatial feature transformation; A second processing submodule is used to obtain a noise-free potential representation based on the U-Net architecture and obtain a super-resolution remote sensing image based on a variational autodecoder; The classifier and the stable diffusion model are obtained by joint training based on training samples. During the joint training process, the parameters of the classifier are updated, the parameters of the category-aware encoder to be trained are updated, and other parameters of the pre-trained stable diffusion model are frozen. The training samples are sample low-resolution remote sensing images carrying global category labels. The joint training process includes the following steps: Encode the sample low-resolution remote sensing image into a sample latent vector through a variational autoencoder , within t steps, the noise is gradually introduced into the sample latent vector through the diffusion process of the stable diffusion model to generate the sample noise latent vector , the stable diffusion model integrates the following parameters to predict the added noise: sample low-resolution remote sensing image , sample noise latent vector , sample content description information and sample texture description information , the parameters of the class perception encoder are updated based on the minimization of the difference between the predicted noise obtained by the stable diffusion model and the actual noise introduced in the diffusion process. The loss function is expressed as follows: ; in, , represents the prediction noise of the stable diffusion model, represents the category-aware encoder to be trained, represents the pre-trained text encoder with frozen parameters, cls represents the sample content description information and sample texture description information; c Represents description coding information, which is composed of the content description information of the low-resolution remote sensing image and texture description information The text features of are concatenated; The classifier to be trained is trained based on sample low-resolution remote sensing images with global category labels, and the loss function is expressed as follows: ; Among them, N is the number of sample low-resolution remote sensing images, M is the number of global categories, yic represents the true label of the i-th sample low-resolution remote sensing image, and pic represents the predicted probability that the i-th sample of the classifier to be trained belongs to the c-th category; The loss function of the joint training of the classifier and the stable diffusion model is expressed as follows: ; Among them, L represents the total loss, LD represents the diffusion loss, LC represents the classification loss, and λ represents the balance coefficient.
5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the remote sensing image super-resolution method based on the diffusion model and the multimodal large language model described in any one of claims 1 to 3 is implemented.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the remote sensing image super-resolution method based on a diffusion model and a multimodal large language model as described in any one of claims 1 to 3 is implemented.
7. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps in the remote sensing image super-resolution method based on a diffusion model and a multimodal large language model described in any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method fusing diffusion model and converter
CN118691826A
Remote sensing image generation method and device based on multi-condition controllable diffusion model
CN118982597A