A remote sensing image ground object sample conversion method based on a diffusion model
By combining a diffusion model and a seasonally invariant feature network, the problem of ground feature characteristics in the conversion of remote sensing image samples in different seasons was solved, achieving high-quality remote sensing image generation and sample library expansion, and improving the accuracy of intelligent interpretation of remote sensing images.
Patent Information
- Application Number
- CN202510224621.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Existing methods are insufficient to effectively address the conversion between remote sensing image samples from different seasons, especially failing to fully consider the complex characteristics of ground features in remote sensing images, thus limiting the development of intelligent interpretation technology for remote sensing images.
A remote sensing image seasonal transformation network is constructed using a diffusion model. Combined with a seasonal invariant feature network, remote sensing samples for the desired season are generated using multi-source remote sensing image samples from different time series through image perception compression, latent diffusion, and seasonal invariant feature learning modules.
It fully preserves remote sensing ground feature information, meets the needs of intelligent interpretation tasks, generates high-quality remote sensing image samples, expands the seasonal sample library, and improves the accuracy of intelligent interpretation of remote sensing images.
Smart Images

Figure CN120088650B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the application of deep learning technology in the generation and conversion of high-resolution remote sensing images, and specifically to a method for converting remote sensing image ground cover samples based on a diffusion model, which performs seasonal conversion on remote sensing image ground cover samples using a diffusion model. Background Technology
[0002] Currently, as a core driving force of the new round of technological and industrial revolution, artificial intelligence technology, especially deep learning, has been deeply coupled with remote sensing applications, and breakthrough progress has been made in remote sensing intelligent interpretation methods based on deep learning (Reference 2). As a data-driven technology, deep learning relies on large-scale, high-quality labeled datasets, and numerous publicly available sample datasets have emerged in the remote sensing community (References 3-7). Domestic and international researchers have conducted very beneficial explorations into the openness and sharing of samples (References 1, 8, 9). However, the conversion of ground cover samples in remote sensing intelligent interpretation remains a challenging area. In the field of computer vision, sample simulation generation and conversion methods based on generative adversarial network models have provided a reference for the conversion of remote sensing intelligent interpretation samples. However, existing methods do not fully consider the characteristics of remote sensing and are difficult to meet the needs of remote sensing intelligent interpretation tasks.
[0003] In the synthesis of remote sensing image samples from different seasons, the inconsistencies between the remote sensing images to be interpreted and existing sample data in terms of imaging season, resolution, imaging method, and spectral characteristics severely restrict the development of high-precision intelligent interpretation technology for remote sensing images. With the emergence of generative models, many scholars at home and abroad have begun to explore methods for image sample synthesis. These methods mainly use deep learning models from the field of computer vision for transfer and extension, such as classic networks like generative adversarial networks (references 10, 11) and autoencoders (references 12, 13), as well as the rapidly developing diffusion model (reference 14). This model generates new images through a diffusion process (adding noise) and a reverse diffusion process (removing noise). In the DALL-E2 model proposed by Ramesh et al. (reference 15), the CLIP model (reference 16) is combined with the diffusion model to learn robust representations of semantic and stylistic images, generating images with diverse styles and very high resolution and realism.
[0004] Although existing models have been well-established in the field of natural image generation, most of these models are designed for spatial, geometric, and textural features in natural images, making it difficult to effectively model the characteristics of ground objects in multi-source remote sensing image samples. Existing methods for synthesizing images between different seasons mostly address seasonal transitions between natural images, failing to fully consider the complex characteristics of ground objects in remote sensing images and thus struggling to meet the diverse needs of remote sensing image databases.
[0005] References:
[0006] [1] Gong Jianya, Xu Yue, Hu Xiangyun, Jiang Liangcun, Zhang Mi, et al. Current status and research of intelligent interpretation sample database for remote sensing images [J]. Acta Geodaetica et Cartographica Sinica, 2021, 50(8): 1013-1022. DOI:10.11947 / j.AGCS.2021.20210085.
[0007] [2] Zhou Peicheng, Cheng Gong, Yao Xiwen, Han Junwei. Machine learning paradigm in high-resolution remote sensing image interpretation. Journal of Remote Sensing. 2021, 25(1):182-197.
[0008] [3]Xia GS, Hu J, Hu F, Shi B, Bai X, Zhong Y, Zhang L, Lu
[0009] [4]Xia GS,Bai
[0010] [5]Cheng G, Han J, Lu X. Remote sensing image scene classification: Benchmark and state of the art. Proceedings of the IEEE.2017,105(10):1865-83.
[0011] [6]Helber P,Bischke B,Dengel A,Borth D.Eurosat:A novel dataset anddeep learning benchmark for land use and land cover classification.IEEEJournal of Selected Topics in Applied Earth Observations and RemoteSensing.2019,12(7):2217-26.
[0012] [7]Li K,Wan G,Cheng G,Meng L,Han J.Object detection in optical remotesensing images:Asurvey and a new benchmark.ISPRS Journal of Photogrammetryand Remote Sensing.2020a,159:296-307.
[0013] [8]Alemohammad H.The case for open-access ML-ready geospatialtraining data.In 2021IEEE International Geoscience and Remote SensingSymposium IGARSS2021(pp.1146-1148).IEEE.[9]McKinstry A,Boydell O,Le Q,PreetI,Hanafin J,Fernandez M,Warde A,Kannan V,Griffiths P.AI-Ready TrainingDatasets for Earth Observation:Enabling FAIR data principles for EO trainingdata.InEGU General Assembly Conference Abstracts 2021(pp.EGU21-12384).
[0014]
[10] Creswell A,White T,Dumoulin V,Arulkumaran K,Sengupta B,BharathAA.Generative adversarial networks:An overview.IEEE Signal ProcessingMagazine.2018,35(1):53-65.
[0015]
[11] Goodfellow,I.J.,Pouget-Abadie,J.,Mirza,M.,Xu,B.,Warde-Farley,D.,Ozair,S.,Courville,A.,and Bengio,Y.(2014).Generative adversarial nets.InNIPS’2014
[0016]
[12] Kingma DP,Welling M.An introduction to variationalautoencoders.arXiv preprint arXiv:1906.02691.2019.
[0017]
[13] Kingma DP,Welling M.Auto-encoding variational bayes.arXivpreprint arXiv:1312.6114.2013.
[0018]
[14] Sehwag,V.,Hazirbas,C.,Gordo,A.,Ozgenel,F.,&Ferrer,C.C.(2022).Generating High Fidelity Data from Low-density Regions using DiffusionModels.2022IEEE / CVF Conference on Computer Vision and Pattern Recognition(CVPR),11482-11491.
[0019]
[15] Ramesh,A.,Dhariwal,P.,Nichol,A.,Chu,C.,&Chen,M.(2022).Hierarchical Text-Conditional Image Generation with CLIP Latents.ArXiv,abs / 2204.06125.
[0020]
[16] Radford, A., Kim, JW, Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning Transferable Visual Models From Natural Language Supervision. International Conference on Machine Learning. Summary of the Invention
[0021] To address the need for conversion between remote sensing image samples from different seasons and the problems existing in current remote sensing image land cover conversion methods, this invention employs a diffusion model to construct a remote sensing image seasonal conversion network. Furthermore, it adds a season-invariant feature network to accommodate the different characteristics of land covers in different seasons, fully utilizing multi-source remote sensing image samples from different time series to generate and synthesize remote sensing samples for the desired season. The method described in this invention fully considers the characteristics of relevant land covers in different seasons, preserving as much remote sensing land cover information as possible to meet the needs of intelligent interpretation tasks.
[0022] The technical solution provided by this invention is as follows:
[0023] Firstly, a method for converting remote sensing image ground cover samples based on a diffusion model is provided, comprising the following steps:
[0024] Acquire remote sensing data, perform preprocessing, and obtain preprocessed data;
[0025] A neural network is constructed; the neural network uses a stable diffusion model as its basic network and includes an image-aware compression module, a latent diffusion model module, and a seasonally invariant feature learning module; wherein, the image-aware compression module includes an encoder. and decoder Used to compress images while preserving important information; the latent diffusion model module includes a diffusion module and a temporal denoising autoencoder with a cross-attention mechanism layer for data denoising; the seasonal invariant feature learning module includes a feature extraction network, a domain encoder, a category classifier, and a domain adaptation classifier for extracting seasonal invariant features from images as conditions to guide image generation.
[0026] The neural network is trained in stages using preprocessed data to obtain the final neural network.
[0027] The remote sensing imagery samples and the seasonal descriptive text information to be converted are input into the final neural network, which outputs the seasonal conversion results.
[0028] In one possible implementation, the remote sensing data includes remote sensing images with seasonal information and remote sensing images with semantic labels; the preprocessing includes cropping and partitioning into training, validation, and test sets.
[0029] In one possible implementation, the image-aware compression module is a pre-trained variational encoder for images. Using encoder The image x is encoded into the latent representation space using the following formula:
[0030]
[0031] in, This represents a three-channel RGB image, where H represents the image height, W represents the image width, and 3 represents the image channel dimension.
[0032] After the image conversion process is completed, the image is reconstructed from the latent representation space using a decoder. The specific formula is as follows:
[0033]
[0034] in, Indicates the reconstruction of the image. This indicates the decoder.
[0035] Furthermore, the latent diffusion model module is positioned between the encoder and decoder of the image perceptual compression module, and includes a diffusion module and a temporal denoising autoencoder with a cross-attention mechanism layer; the diffusion module is used to superimpose Gaussian noise onto the image; the temporal denoising autoencoder with a cross-attention mechanism layer includes a cross-attention downsampling module, a downsampling module, a cross-attention module, an upsampling module, and a cross-attention upsampling module connected in sequence to form a Unet; the cross-attention downsampling module and the cross-attention upsampling module are connected in a skip connection; the upsampling module and the downsampling module are connected in a skip connection.
[0036] The objective function L of the potential diffusion model moduleLDM for:
[0037]
[0038] Among them, the temporal denoising autoencoder with cross-attention mechanism layer ∈ θ (z t ,t) is a time-conditional Unet network, ∈ θ For temporal denoising autoencoders, t is the time step, z is the time step. t The noisy image at time step t; Let Unet represent the expected value, and ∈ represent the model Unet with predicted noise. Indicates a Gaussian distribution. Represents a given Gaussian distribution The noise.
[0039] Furthermore, the seasonally invariant feature learning module includes a feature extraction network and a domain encoder τ. θ Category classifiers and domain-adaptive classifiers;
[0040] The feature extraction network is a Unet network;
[0041] The domain encoder τ θ This is used to map conditional control information y from different modes into an intermediate representation τ. θ (y) is integrated into the intermediate layer of the Unet network in the temporal denoising autoencoder with cross-attention mechanism layer, so as to guide the image conversion through the semantic labels of remote sensing images and the text description of target images.
[0042] The category classifier is composed of multiple convolutional layers connected together, used to extract semantic features of the image, so that the feature extraction network can learn seasonally invariant semantic features.
[0043] The domain adaptation classifier is a domain classification network composed of multiple linear layers, used to determine the seasonal information of an image, enabling the feature extraction network to determine the season of the image and thus extract seasonally invariant feature information.
[0044] In one possible implementation, the method of training the neural network in stages using preprocessed data includes:
[0045] Training is performed separately in the source and target domains: the input image is encoded by a feature extraction network to extract image features; the extracted image features are classified by a category classifier and a domain classifier, where the category classifier determines the land cover category of each pixel in the input image, and the domain classifier determines whether the input image belongs to the source or target domain; the loss of the two classifiers is calculated, and the weights of the seasonally invariant feature learning module are updated through backpropagation;
[0046] Training the latent diffusion model: using an encoder The forward propagation process involves encoding the image into a latent space and adding noise to it. The reverse propagation process trains the diffusion model by encoding the relevant text describing the target season from the image using a pre-trained CLIP text encoder. Season-invariant features of the image are extracted using a season-invariant feature learning module. Conditional control information is mapped and incorporated into the intermediate layer of a UNet network with a cross-attention mechanism layer for temporal denoising autoencoder. The temporal denoising autoencoder with the cross-attention mechanism layer predicts noise in the latent space, resulting in a representation of the target season image in the latent space after denoising. This representation is then processed by the decoder. Reconstructing images from the latent representation space yields seasonally altered images;
[0047] The loss function is calculated between the seasonally transformed image and the actual image of the target season, and the weights of the neural network are updated through backpropagation.
[0048] Secondly, a remote sensing image ground feature sample conversion device based on a diffusion model is provided, comprising:
[0049] The acquisition module is used to acquire remote sensing data, perform preprocessing, and obtain preprocessed data.
[0050] The building module is used to construct a neural network; the neural network uses a stable diffusion model as its base network and includes an image-aware compression module, a latent diffusion model module, and a season-invariant feature learning module; wherein, the image-aware compression module includes an encoder. and decoder Used to compress images while preserving important information; the latent diffusion model module includes a diffusion module and a temporal denoising autoencoder with a cross-attention mechanism layer for data denoising; the seasonal invariant feature learning module includes a feature extraction network, a domain encoder, a category classifier, and a domain adaptation classifier for extracting seasonal invariant features from images as conditions to guide image generation.
[0051] The training module is used to train the neural network in stages using preprocessed data to obtain the final neural network.
[0052] The output module is used to input remote sensing image land cover samples and seasonal descriptive text information to be converted into the final neural network, and output the seasonal conversion result.
[0053] Thirdly, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the remote sensing image ground feature sample conversion method based on the diffusion model as described in the first aspect.
[0054] Fourthly, a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the remote sensing image ground feature sample conversion method based on a diffusion model as described in the first aspect.
[0055] Fifthly, a computer program product includes a computer program that, when executed by a processor, implements the remote sensing image ground feature sample conversion method based on a diffusion model as described in the first aspect.
[0056] The present invention has the following beneficial effects:
[0057] (1) The method described in this invention uses a stable diffusion model as the basic generation model and utilizes the powerful generation capability of the diffusion model to generate remote sensing images of different seasons.
[0058] (2) The method described in this invention constructs a seasonally invariant feature extraction network, which fully considers the characteristics of relevant land features in different seasons during the seasonal transition process, ensuring sufficient remote sensing land feature information.
[0059] (3) The method described in this invention uses a transfer learning approach based on a stable diffusion model to add an additional domain-adaptive classifier to constrain the quality of the generated images, thereby generating better quality remote sensing image samples.
[0060] (4) The method described in this invention expands the seasonal sample library through seasonal transformation to meet the needs of intelligent interpretation tasks. Attached Figure Description
[0061] Figure 1 This is a schematic flowchart of a remote sensing image ground cover sample conversion method based on a diffusion model, as described in an embodiment of the present invention.
[0062] Figure 2 This is a schematic diagram of the neural network model described in an embodiment of the present invention;
[0063] Figure 3 A schematic diagram of a temporal denoising autoencoder with a cross-attention mechanism layer;
[0064] Figure 4 This is a schematic diagram of a feature extraction network;
[0065] Figure 5 This is a rendering of the seasonal transition, in which... Figure 5 (a) in the image is the converted winter image. Figure 5 (b) in the image is the converted summer image. Figure 5 (c) in the image represents the input condition image before conversion;
[0066] Figure 6 This is a schematic diagram of a remote sensing image ground feature sample conversion device based on a diffusion model.
[0067] Figure 7 This is a schematic diagram of the electronic device described in this invention. Detailed Implementation
[0068] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0069] To address the need for conversion between remote sensing image samples from different seasons and the problems existing in current remote sensing image feature conversion methods, this invention provides a remote sensing image feature sample conversion method based on a diffusion model. The method employs a diffusion model to construct a remote sensing image seasonal conversion network and adds a season-invariant feature learning module to address the different characteristics of features in different seasons. It fully utilizes multi-source remote sensing image samples from different time series to generate remote sensing samples for the desired season. This method fully considers the characteristics of relevant features in different seasons, preserving remote sensing feature information as much as possible and meeting the needs of intelligent interpretation tasks.
[0070] like Figure 1 As shown in the figure, this embodiment of the invention provides a method for converting remote sensing image ground cover samples based on a diffusion model, the steps of which are as follows:
[0071] Step 1: Select remote sensing data, perform preprocessing, and obtain preprocessed data.
[0072] In one possible implementation, the remote sensing data includes remote sensing images with seasonal information and remote sensing images with semantic labels; the preprocessing includes cropping and partitioning into training, validation, and test sets.
[0073] For example, the datasets used in this embodiment are the SEN1-2 dataset with seasonal information, the Postdam and Vaihingen datasets with semantic labels, and the datasets are pre-processed to divide them into 80% as training set, 10% as validation set, and 10% as test set for experiments.
[0074] (1) For SEN1-2, this dataset was created using GEE (Google Earth Engine). GEE was used to select sampling points around the world and sampled four seasons in 2017: spring (March 1, 2017 - May 30, 2017), summer (June 1, 2017 - August 31, 2017), autumn (September 1, 2017 - November 30, 2017), and winter (December 1, 2016 - February 28, 2017). SAR images of the Sentinel-1 (SEN1) C-band and RGB image data of the Sentinel-2 (SEN2) 4, 3, and 2 bands were obtained. The image size is 256*256, and there are a total of 282,384 RGB and SAR image pairs. In this embodiment, only some seasonal RGB images are selected as the dataset for constructing the reference sample library.
[0075] (2) The Vaihingen dataset contains aerial images from Vaihingen, Germany, a relatively small village with many detached buildings and small multi-story buildings. The images in this dataset cover an area of approximately 1 square kilometer, with a resolution of 9 cm and an average size of 2494×2064 pixels. They include five foreground object classes (impermeable surfaces, buildings, low vegetation, trees, and cars) and one background class. In this embodiment, 33 images were divided into 3300 images of 512×512 pixels each for the experiment.
[0076] (3) The Postdam dataset contains aerial images of Potsdam, a typical historical city with large building complexes, narrow streets, and dense settlements. This two-dimensional semantically labeled dataset covers an area of approximately 2.3 square kilometers, with an image resolution of 5 cm, and consists of 38 images of 6000×6000 pixels. The dataset provides NIR-RGB channels, with the corresponding categories being the same as those in the Vaihingen dataset. This embodiment uses only the RGB three-channel images from the Postdam dataset, dividing them into 1000×1000 pixel images to obtain 3660 images for conversion.
[0077] Step 2: Construct a neural network; the neural network uses a stable diffusion model as the basic network, which includes three network modules: an image perception compression module, a latent diffusion model module, and a seasonally invariant feature learning module.
[0078] In one possible implementation, the neural network structure is as follows: Figure 2As shown, the image-aware compression module is a pre-trained variational encoder, comprising an encoder and a decoder. This module processes the original image, ignoring high-frequency information and retaining only important and fundamental features. This significantly reduces the computational complexity of the training and sampling phases.
[0079] It should be noted that the image perception compression module can learn a potential representation space that is perceptually equivalent to the image space. Its advantage is that only one general autoencoder model needs to be trained, which can be used to train different diffusion models and be used on different tasks.
[0080] Specifically, for a given image Using encoder The image is encoded into the latent representation space using the following formula:
[0081]
[0082] in, This represents a three-channel RGB image, where H represents the image height, W represents the image width, and 3 represents the image channel dimension.
[0083] After the image conversion process is completed, a decoder is used to reconstruct the image from the latent representation space. The specific formula is as follows:
[0084]
[0085] in, Indicates the reconstruction of the image. This indicates the decoder.
[0086] In one possible implementation, the potential diffusion model module includes a diffusion module and a temporal denoising autoencoder with a cross-attention mechanism layer.
[0087] Furthermore, the diffusion module is used to add noise to the image, including using a Markov chain to progressively add Gaussian noise to the original data to obtain a noisy image that follows a Gaussian distribution. The specific formula is:
[0088]
[0089] Where, x t Let q(x) be the image at time t. t |x t-1 ) is x t The probability distribution of the noisy image at time step follows a Gaussian distribution. It follows a Gaussian distribution, where I is the diagonal covariance matrix and β is the distribution. t The variance used for each time step t.
[0090] Understandably, a temporal denoising autoencoder ∈ θ (x t ,t); t=1…T, where x t It represents the noise of the input image x at time t, where T represents the number of time steps. ∈ θ This is a temporal denoising autoencoder, which learns the data distribution by gradually denoising normally distributed variables, based on the input x. t To predict the corresponding denoised result.
[0091] The objective function of the potential diffusion model module is as follows:
[0092]
[0093] Among them, the temporal denoising autoencoder with cross-attention mechanism layer ∈ θ (z t ,t) is a time-conditional Unet network, ∈ θ For temporal denoising autoencoders, t is the time step, z is the time step. t The noisy image at time step t; Let Unet represent the expected value, and ∈ represent the model Unet with predicted noise. Indicates a Gaussian distribution. Represents a given Gaussian distribution The noise.
[0094] See Figure 3 The structure is a temporal denoising autoencoder with a cross-attention mechanism layer. Control information is integrated into each downsampling and upsampling layer of UNet through a cross-attention layer mapping. This includes a cross-attention downsampling module, a downsampling module, a cross-attention module, an upsampling module, and a cross-attention upsampling module that are connected in sequence to form UNet. The cross-attention downsampling module and the cross-attention upsampling module are connected in a skip connection. The upsampling module and the downsampling module are connected in a skip connection.
[0095] For example, the cross-attention downsampling module is a cross-attention mechanism layer; the downsampling module is a convolutional layer; the cross-attention module is a cross-attention mechanism layer; the upsampling module is a convolutional layer; and the cross-attention upsampling module is a cross-attention mechanism layer.
[0096] The implementation of the cross-attention mechanism layer is as follows:
[0097]
[0098] Where Attention represents the attention vector, Q, K, and V are the information to be queried, the vector being queried, and the value obtained; the superscript T represents the transpose matrix; and d represents the dimension of the hidden layer. As an intermediate representation of UNet, This represents the feature encoder, where i represents the feature vector. Let N represent the real number field, and N represent the feature dimension. The dimension of the hidden layer is represented by i, the feature vector is represented by i, and ∈ represents the model Unet that predicts noise. is the learnable projection matrix, and d represents the dimension of the hidden layer.
[0099] It should be noted that the structure of the Unet model (i.e., the model predicting noise) is a cross-attention downsampling module, a downsampling module, a cross-attention module, an upsampling module, and a cross-attention upsampling module; the cross-attention downsampling module and the cross-attention upsampling module are skipped connections; the upsampling module and the downsampling module are also skipped connections. That is, ∈ θ (z t The structure ,t) is consistent with ∈, where ∈ represents the truth value, and ∈ θ (z t ,t) represents the value predicted by the model. After calculating the loss, the model parameters are optimized and updated.
[0100] In the latent diffusion model module, given the latent representation space z0 obtained from image encoding, the diffusion module then obtains z. t , z t Temporal denoising autoencoder with cross-attention mechanism layer ∈ θ Get z t-1 , z t-1 Then, after passing through the same temporal denoising autoencoder ∈ θ Repeat this process t times to obtain z0. z0 represents the latent spatial features at time 0.
[0101] The latent diffusion model described in this embodiment incorporates a pre-trained image-aware compression model, which includes an encoder. and a decoder Encoders can be used during training. The input image x is encoded to obtain z, and then z is obtained through a diffusion process. t This allows the model to learn in the latent representation space, enabling it to focus on more important semantic information in the data. Training in a low-dimensional space also makes computation more efficient. Since the forward propagation process of diffusion is fixed, it can be done during training from... To obtain z efficiently t Furthermore, the learned samples only need to be processed by the decoder. A single decoding operation is sufficient to generate an image in the image space.
[0102] In one possible implementation, the seasonally invariant feature learning module includes a feature extraction network, a domain encoder, a category classifier, and a domain adaptation classifier.
[0103] Furthermore, the feature extraction network is a Unet network, such as... Figure 4 As shown, it includes an upsampling layer, a downsampling layer, and a skip connection layer; the upsampling layer and the downsampling layer are skip connected. Both the upsampling layer and the downsampling layer are composed of convolutional layers.
[0104] Furthermore, the domain encoder consists of 5 convolutional layers, each with a 3×3 convolutional kernel, used to encode seasonal image features.
[0105] Furthermore, the category classifier consists of four convolutional layers connected together, each with a 3×3 convolutional kernel, used to extract semantic features of the image, enabling the feature extraction network to learn seasonally invariant semantic features.
[0106] Furthermore, the domain adaptation classifier is a domain classification network composed of multiple linear layers (fully connected layers), with the parameter of each linear layer being the number of feature channels. The domain adaptation classifier is used to determine the seasonal information of the image, enabling the feature extraction network to determine the season to which the image belongs, thereby extracting seasonally invariant feature information.
[0107] Data flow direction: such as Figure 2 As shown, on the one hand, the original image extracts features through a feature extraction network to obtain feature representations, and then the feature representations are integrated into a temporal denoising autoencoder with a cross-attention mechanism layer by a domain encoder. θ In the middle; on the other hand, the text is encoded by the domain encoder τ θ Integrating into a temporal denoising autoencoder with a cross-attention mechanism layer θ In the middle, the domain encoder encodes the seasonal features of the image to determine the season of the input image; the category classifier determines the semantic category of the image and learns seasonally invariant features.
[0108] Understandably, given the complexity of multiple ground features in remote sensing imagery, a conditional mechanism is introduced, specifically a conditional temporal denoising autoencoder, to achieve ∈ θ (z t The diffusion model is extended by using (t,y), where y is the introduced conditional control information. This allows the process of synthesizing remote sensing images to be guided and controlled by the conditional control information y.
[0109] Specifically, a temporal noise predictor is implemented on a UNet network with a cross-attention mechanism layer for a temporal denoising autoencoder. θ (z t ,t,y), and through the domain encoder τ θMapping conditional control information y from different modalities into intermediate representations (M represents the feature dimension of different modalities, d represents the dimension of the hidden layer, and τ represents the representation of the domain encoder). The final model can then integrate conditional control information into the intermediate layer of the temporal denoising autoencoder UNet with a cross-attention mechanism layer through a cross-attention layer mapping. Ultimately, it guides the seasonal or modal transformation of the image through the semantic labels of the remote sensing image and the textual description of the target image. The conditional control information y includes seasonally invariant features and textual descriptions.
[0110] This invention adds a domain adaptation classifier to the stable diffusion model using transfer learning to constrain the quality of the generated images, thereby generating higher quality remote sensing image samples.
[0111] Step 3: Train the neural network in stages using the preprocessed data to obtain the final neural network.
[0112] The training process is divided into two stages: (1) training the seasonal invariant feature learning module; (2) training the latent diffusion model module to perform seasonal transformation of images. During training, seasonal invariant features and text descriptions are added as conditions to guide the seasonal transformation of images during denoising.
[0113] In one possible implementation, step 3 includes the following sub-steps:
[0114] S3.1, Training is performed in the source domain and the target domain respectively:
[0115] (1) The input image is encoded by a feature extraction network to extract image features;
[0116] (2) The extracted image features are classified by a category classifier and a domain adaptation classifier; wherein, the category classifier determines the land cover category to which each pixel of the input image belongs, and the domain classifier determines whether the input image belongs to the source domain or the target domain.
[0117] (3) Calculate the loss of the two classifiers and update the weights of the seasonally invariant feature module through backpropagation.
[0118] S3.2, Training the latent diffusion model module:
[0119] (1) Using an encoder The image is encoded into the latent space, and noise is added to it to complete the forward diffusion process;
[0120] (2) The reverse process of training the diffusion model is to encode the relevant text of the target season description that the image needs to be converted into using the pre-trained CLIP text encoder; the seasonal invariant feature learning module is used to extract the seasonal invariant features of the image, and the conditional control information is mapped and integrated into the intermediate layer of the UNet with the cross attention mechanism layer through the cross attention mechanism layer.
[0121] (3) The temporal denoising autoencoder with cross-attention mechanism layer predicts noise in the latent space and obtains the representation of the target seasonal image in the latent space after denoising.
[0122] (4) After the decoder Reconstructing images from the latent representation space yields seasonally altered images;
[0123] (5) Calculate the loss function between the seasonally transformed image and the target seasonal real image, and update the weights of the neural network through backpropagation.
[0124] It should be noted that, considering the number of network parameters and the amount of training data, this embodiment adopts a fine-tuning approach, using remote sensing data to fine-tune the pre-trained stable diffusion model.
[0125] Step 4: Input the remote sensing image ground feature samples and the seasonal description text information to be converted into the final neural network, and output the seasonal conversion result.
[0126] In one possible implementation, after the neural network is trained, the network weights are saved to a ckpt file. With the help of the ckpt file, seasonal transformation can be performed directly using the input remote sensing imagery and the seasonal description text information to be transformed, and the final result can be output. Figure 5 This is a rendering of the seasonal transition, in which... Figure 5 (a) in the image is the converted winter image. Figure 5 (b) in the image is the converted summer image. Figure 5 (c) in the image represents the input condition image before conversion;
[0127] The following describes a remote sensing image ground feature sample conversion device based on a diffusion model provided by the present invention. The remote sensing image ground feature sample conversion device based on a diffusion model described below can be referred to in correspondence with the remote sensing image ground feature sample conversion method based on a diffusion model described above.
[0128] Figure 6 This is a schematic diagram of the structure of the remote sensing image ground cover sample conversion device based on the diffusion model provided in an embodiment of the present invention, as shown below. Figure 6 As shown, it includes: acquisition module 61, construction module 62, training module 63, and output module 64, wherein:
[0129] The acquisition module 61 is used to acquire remote sensing data and perform preprocessing.
[0130] Module 62 is used to construct a neural network; the neural network uses a stable diffusion model as its basic network and includes an image perceptual compression module, a latent diffusion model module, and a season-invariant feature learning module; wherein, the image perceptual compression module includes an encoder. and decoder Used to compress images while preserving important information; the latent diffusion model module includes a diffusion module and a temporal denoising autoencoder with a cross-attention mechanism layer for data denoising; the seasonal invariant feature learning module includes a feature extraction network, a domain encoder, a category classifier, and a domain adaptation classifier for extracting seasonal invariant features from images as conditions to guide image generation.
[0131] Training module 63 is used to train the neural network in stages to obtain the final neural network;
[0132] Output module 64 is used to input remote sensing image ground feature samples and seasonal description text information to be converted into the final neural network and output the seasonal conversion result.
[0133] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740. The processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions from the memory 730 to execute a remote sensing image feature sample conversion method based on a diffusion model.
[0134] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0135] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute a remote sensing image ground feature sample conversion method based on a diffusion model provided by the above methods.
[0136] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform a diffusion model-based remote sensing image ground feature sample conversion method provided by the above methods.
[0137] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for converting remote sensing image ground cover samples based on a diffusion model, characterized in that, Includes the following steps: Acquire remote sensing data, perform preprocessing, and obtain preprocessed data; A neural network is constructed; the neural network uses a stable diffusion model as its basic network and includes an image-aware compression module, a latent diffusion model module, and a seasonally invariant feature learning module; wherein, the image-aware compression module includes an encoder. and decoder The image perceptual compression module is used to compress images while preserving important information. The latent diffusion model module includes a diffusion module and a temporal denoising autoencoder with a cross-attention mechanism layer for data denoising. The seasonal invariant feature learning module includes a feature extraction network, a domain encoder, a category classifier, and a domain adaptation classifier, used to extract seasonal invariant features from the image as conditions to guide image generation. The latent diffusion model module is positioned between the encoder and decoder of the image perceptual compression module and includes a diffusion module and a temporal denoising autoencoder with a cross-attention mechanism layer. The diffusion module is used to superimpose Gaussian noise onto the image. The temporal denoising autoencoder with a cross-attention mechanism layer includes a cross-attention downsampling module, a downsampling module, a cross-attention module, an upsampling module, and a cross-attention upsampling module connected sequentially to form a Unet. The cross-attention downsampling module and the cross-attention upsampling module are skipped connections; the upsampling module and the downsampling module are skipped connections. The feature extraction network is a Unet network; the domain encoder Used to transfer conditional control information of different modes Mapping to intermediate representation The Unet network, which incorporates a cross-attention mechanism layer into the intermediate layer of a temporal denoising autoencoder with a cross-attention mechanism layer, guides image transformation through semantic labels of remote sensing images and text descriptions of target images. The category classifier, composed of multiple convolutional layers, is used to extract semantic features of the images, enabling the feature extraction network to learn seasonally invariant semantic features. The domain adaptation classifier, a domain classification network composed of multiple linear layers, is used to determine the seasonal information of the images, enabling the feature extraction network to determine the season of the images and thus extract seasonally invariant feature information. The neural network is trained in stages using preprocessed data to obtain the final neural network. The remote sensing imagery samples and the seasonal descriptive text information to be converted are input into the final neural network, which outputs the seasonal conversion results.
2. The method for converting remote sensing image ground cover samples based on a diffusion model according to claim 1, characterized in that, The remote sensing data includes remote sensing images with seasonal information and remote sensing images with semantic labels; the preprocessing includes cropping and dividing the training set, validation set and test set.
3. The remote sensing image ground cover sample conversion method based on diffusion model according to claim 1, characterized in that, The image-aware compression module is a pre-trained variational encoder for images. Using an encoder Image x Encoding to the latent representation space, the formula is: in, ; This represents a three-channel RGB image, where H represents the image height, W represents the image width, and 3 represents the image channel dimension. After the image conversion process is completed, the image is reconstructed from the latent representation space using a decoder. The specific formula is as follows: in, Indicates the reconstruction of the image. This indicates the decoder.
4. The method for converting remote sensing image ground cover samples based on a diffusion model according to claim 1, characterized in that, Objective function of the potential diffusion model module for: Among them, the temporal denoising autoencoder with cross-attention mechanism layer For time-conditional Unet networks, For timing denoising autoencoders, t For time steps, for t Noisy image at each time step; Expressing expectations, The model Unet represents the predicted noise. Indicates a Gaussian distribution. ) represents a given Gaussian distribution The noise.
5. The method for converting remote sensing image ground cover samples based on a diffusion model according to claim 1, characterized in that, The method of training a neural network in stages using preprocessed data includes: Training is performed separately in the source and target domains: the input image is encoded by a feature extraction network to extract image features; the extracted image features are classified by a category classifier and a domain classifier, where the category classifier determines the land cover category of each pixel in the input image, and the domain classifier determines whether the input image belongs to the source or target domain; the loss of the two classifiers is calculated, and the weights of the seasonally invariant feature learning module are updated through backpropagation; Training the latent diffusion model: using an encoder The forward propagation process involves encoding the image into a latent space and adding noise to it. The reverse propagation process trains the diffusion model by encoding the relevant text describing the target season from the image using a pre-trained CLIP text encoder. Season-invariant features of the image are extracted using a season-invariant feature learning module. Conditional control information is mapped and incorporated into the intermediate layer of a UNet network with a cross-attention mechanism layer for temporal denoising autoencoder. The temporal denoising autoencoder with the cross-attention mechanism layer predicts noise in the latent space, and after denoising, the representation of the target season image in the latent space is obtained. Finally, the image is decoded. Reconstructing images from the latent representation space yields seasonally altered images; The loss function is calculated between the seasonally transformed image and the actual image of the target season, and the weights of the neural network are updated through backpropagation.
6. A remote sensing image ground feature sample conversion device based on a diffusion model, characterized in that, To implement the method according to any one of claims 1-5, comprising: The acquisition module is used to acquire remote sensing data, perform preprocessing, and obtain preprocessed data. The building module is used to construct a neural network; the neural network uses a stable diffusion model as its base network and includes an image-aware compression module, a latent diffusion model module, and a season-invariant feature learning module; wherein, the image-aware compression module includes an encoder. and decoder The first module is used to compress images while preserving important information; the second module is a latent diffusion model module, which includes a diffusion module and a temporal denoising autoencoder with a cross-attention mechanism layer, used for data denoising; the third module is a seasonal invariant feature learning module, which includes a feature extraction network, a domain encoder, a category classifier and a domain adaptation classifier, used to extract seasonal invariant features from images as conditions to guide image generation. The training module is used to train the neural network in stages using preprocessed data to obtain the final neural network. The output module is used to input remote sensing image land cover samples and seasonal descriptive text information to be converted into the final neural network, and output the seasonal conversion result.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the remote sensing image ground feature sample conversion method based on the diffusion model as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the remote sensing image ground feature sample conversion method based on the diffusion model as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the remote sensing image ground feature sample conversion method based on the diffusion model as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Ground object target identification method and device based on deep learning segmentation network
CN112101309A
Conditional diffusion model-based confrontation sample purification method
CN119293511A