High-resolution remote sensing image semantic segmentation method based on diffusion model
Through the diffusion model-based autoencoder and denoising U-Net architecture, the problem of difficult segmentation of small targets in high-resolution remote sensing images is solved, high-precision semantic segmentation is achieved in complex scenes, and the generalization ability and segmentation effect of the model are improved.
Patent Information
- Application Number
- CN202510587884.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-09-19
AI Technical Summary
Existing remote sensing image semantic segmentation methods have difficulty processing small targets in high-resolution images and lack sufficient annotated data, resulting in limited model generalization capabilities, especially insufficient segmentation accuracy and robustness in complex scenes.
An autoencoder and denoising U-Net architecture based on a diffusion model is adopted. Through a staged training strategy, the autoencoder is first trained to obtain feature representation, and then the conditional encoding module and denoising U-Net are jointly trained. The noise injection module and multi-scale feature extraction are used for image segmentation, and the cross-attention mechanism is combined to realize conditionally guided diffusion generation.
The accuracy and robustness of multi-classification segmentation of high-resolution remote sensing images have been improved, especially the ability to accurately identify small targets in complex background and noisy environments, thereby improving the accuracy and stability of semantic segmentation.
Smart Images

Figure CN120673054A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of space optical remote sensing, and in particular relates to a semantic segmentation method for high-resolution remote sensing images based on a diffusion model. Background Art
[0002] Since the 20th century, remote sensing technology has been widely used in fields such as Earth observation, resource management, and environmental monitoring. With the explosive growth of remote sensing data and the continuous improvement of image resolution, efficient and accurate interpretation of remote sensing images has become a critical task in fields such as Geographic Information Systems (GIS), urban planning, and agricultural monitoring. Semantic segmentation, a technique for classifying each pixel in an image, accurately identifies and segments objects within the image and is a key tool for achieving in-depth understanding and analysis of remote sensing images. Semantic segmentation of remote sensing images involves classifying each pixel in a remote sensing image into specific feature categories, such as buildings, roads, water bodies, and vegetation, based on their characteristics. This not only provides accurate feature classification but also determines the spatial distribution of these categories. This enables precise interpretation of remote sensing images at a fine-grained level, empowering powerful analytical capabilities for numerous application scenarios and providing high-frequency, high-precision geographic information support for the implementation of major national strategies such as national land space planning, ecological and environmental governance, and disaster prevention and mitigation.
[0003] With the development of remote sensing technology, the resolution and complexity of images have continued to increase, placing higher demands on semantic segmentation algorithms. Semantic segmentation methods can be roughly divided into traditional methods and deep learning-based methods. Traditional remote sensing image semantic segmentation algorithms can be further divided into rule-based and machine learning-based methods. Rule-based methods include threshold segmentation, region growing, watershed algorithms, and graph cuts; machine learning-based methods include maximum likelihood estimation, support vector machines (SVMs), and decision trees. These methods typically require manual feature extraction. The core idea is to use low-level features such as image texture, color, and shape and prior knowledge to classify pixels. Although these methods can achieve effective segmentation to a certain extent, they have difficulty handling complex scenes, have poor resistance to noise and complex backgrounds, and perform poorly in high-dimensional data, resulting in limited segmentation accuracy.
[0004] In recent years, with the upgrading of computer hardware and software, semantic segmentation methods based on deep learning have become mainstream, and semantic segmentation has entered a new period of development. Semantic segmentation methods based on deep learning automatically learn the multi-level features of images, can be trained end-to-end on large-scale datasets, and perform excellently in feature extraction and classification. Deep learning-based methods enable these models to extract high-dimensional spatial information in complex scenes and perform accurate pixel-level classification, significantly improving the accuracy and robustness of segmentation. According to the model architecture, semantic segmentation can be divided into: semantic segmentation based on Convolutional Neural Networks (CNN) and semantic segmentation based on Transformer.
[0005] Convolutional neural networks (FCNs) are the foundational model for deep learning applied to image processing and are widely used for semantic segmentation of remote sensing images. Fully Convolutional Neural Networks (FCNs) innovatively replace the fully connected layers in CNNs with convolutional layers, enabling end-to-end pixel-level classification. Building on FCNs, the encoder-decoder architecture further advances semantic segmentation. The encoder, typically composed of a series of convolutional and pooling layers, extracts deep image features, gradually reduces the resolution, and compresses the image information into a high-level semantic representation. The decoder, through upsampling and convolutional layers, gradually restores the compressed feature maps to the original image resolution, recovering spatial information and outputting a refined segmentation result. Representative methods include SegNet and U-Net. Dilated convolutions increase the receptive field by inserting holes between convolution kernel elements, allowing the network to capture a wider range of contextual information without increasing the number of parameters or reducing resolution. The DeepLab series of models exemplify the application of dilated convolutions. DeepLabv1 addresses the spatial information loss problem caused by multi-layer pooling in traditional CNNs by introducing dilated convolutions to maintain feature map resolution and improve segmentation accuracy. DeepLabv2 further proposes Atrous Spatial Pyramid Pooling (ASPP), which uses dilated convolutions with different dilation rates to process feature maps in parallel, fusing multi-scale information to meet the needs of segmenting objects of different sizes. DeepLabv3+ combines an encoder-decoder structure with dilated convolutions, using dilated convolutions in the encoder to obtain multi-scale context and the decoder to restore boundary details. Multi-scale strategies are also an important means to improve segmentation performance. Methodologically, on the one hand, the input image is scaled to different scales through the image pyramid, and then fed into the network separately, fusing the prediction results at each scale. On the other hand, within the network, similar to the ASPP module, dilated convolutions with different dilation rates or parallel branches are used to process feature maps, extracting multi-scale features and improving the robustness and accuracy of semantic segmentation in complex scenes.
[0006] The Transformer model has been gradually introduced into the field of computer vision due to its success in natural language processing. Remote sensing images usually have rich spatial information and complex scene structures, in which the spatial distribution and mutual relationship between different ground objects are crucial for accurate semantic segmentation. In the semantic segmentation of remote sensing images, the Transformer can capture long-range dependencies in the image through its powerful global modeling capabilities. Specifically, through the self-attention mechanism, the Transformer model can comprehensively consider the information of all other pixels in the image when processing each pixel in the image, thereby making a more accurate judgment on the semantic category of each pixel. This global modeling capability enables the Transformer model to avoid segmentation errors caused by local information limitations when processing remote sensing images containing large-scale ground objects and complex geographical scenes, thereby significantly improving the accuracy of segmentation.
[0007] Image semantic segmentation technology continues to evolve within the deep learning framework. Based on different learning strategies, it can be broadly categorized into two approaches: those based on generative adversarial networks (GANs) and those based on diffusion models. These two approaches are driving the development of semantic segmentation technology from different perspectives, each demonstrating unique advantages and application potential. GANs consist of two core components: a generator and a discriminator. The generator's primary task is to generate samples similar to real data based on an input noise vector, while the discriminator is responsible for determining whether the input data originates from the real data distribution or is generated by the generator. In semantic segmentation tasks, the generator strives to produce high-quality segmentation results, while the discriminator distinguishes the generated segmentation results from the ground-truth segmentation labels. Performance is improved through continuous adversarial training between the two. This adversarial training mechanism enables GANs to effectively enhance the detail and consistency of semantic segmentation models. The application of GANs in semantic segmentation of remote sensing images is primarily focused on refining and optimizing the generated segmentation results, particularly when working with small sample sizes or poorly labeled data.
[0008] Applying diffusion models to semantic segmentation tasks has brought new insights to the field. Diffusion models are a type of generative method based on probabilistic models that have demonstrated strong capabilities in image generation tasks in recent years. By gradually learning the image generation process, diffusion models are able to better capture the complex structures and details in remote sensing images. Although the application of diffusion models to semantic segmentation is still in the exploratory stage, a number of research results have been impressive. For example, when processing remote sensing images in high-noise environments or scenes with complex backgrounds, diffusion models can effectively overcome noise interference, identify target objects in complex backgrounds, and achieve excellent segmentation results.
[0009] Although deep learning methods have made significant progress in semantic segmentation, the task of semantic segmentation of remote sensing images still faces multiple challenges. Remote sensing images usually have extremely high resolution. In high-resolution images, accurate segmentation of small targets is very difficult, requiring the model to have stronger detail discrimination capabilities. The acquisition of high-quality pixel-level label data is expensive, especially in the annotation of large-scale remote sensing images. The lack of sufficient labeled data limits the generalization ability of the model. The methods mentioned above have their own advantages and are suitable for different application scenarios. Together, they are driving the continuous development of remote sensing image analysis technology. With the deepening of research, we can expect to see these methods applied in more complex scenarios in the future, and continue to combine with other technologies to further improve the performance and application scope of remote sensing image semantic segmentation, and continue to promote the development of remote sensing image analysis technology. Summary of the Invention
[0010] The present invention proposes a semantic segmentation method for high-resolution remote sensing images based on a diffusion model, which can achieve high-precision segmentation of high-resolution remote sensing images.
[0011] The technical solutions for implementing the present invention are as follows:
[0012] A semantic segmentation method for high-resolution remote sensing images based on a diffusion model. The specific process is as follows:
[0013] The first stage of training: the high-resolution remote sensing image is annotated with the type of each pixel to generate a labeled image, and the labeled image is used to train the autoencoder including the encoder and decoder;
[0014] Second stage training: Based on the autoencoder, a conditional diffusion model is loaded and trained, including a noise injection module, a conditional encoding module, and a denoising U-Net. The autoencoder parameters are fixed during training. The noise injection module adds noise to the latent variables output by the encoder. The conditional encoding module is used to extract multi-scale features from the high-resolution image. The denoising U-Net is guided by the multi-scale features to predict image components and noise components and reconstruct the image. The decoder is used to decode the reconstructed image.
[0015] Image semantic segmentation: Use the trained network to perform semantic segmentation on high-resolution remote sensing images.
[0016] Optionally, the present invention further includes performing normalization processing on the label image so that the pixel values of the label image are distributed within the interval [-1, 1].
[0017] Optionally, the present invention calculates cross entropy loss and adversarial loss during the first stage of training.
[0018] Optionally, the noise injection module of the present invention adds noise to the latent variable output by the encoder, and the latent variable after adding noise is z t ;
[0019]
[0020] Among them, z0 is the initial variable output by the encoder, f t is an explicit transition function, and the initial noise n is generated according to the size of the latent variable z0.
[0021] Optionally, the initial noise n of the present invention is:
[0022] n=Sample(N(0,I),shape(z0))
[0023] Among them, the Sample function represents sampling from the specified distribution, shape(z0) represents the shape of the latent variable z0, and the noise n obeys the standard normal distribution, that is, n~N(0,I).
[0024] Optionally, the conditional coding module of the present invention uses Swin Transformer to process the high-resolution remote sensing image layer by layer, and extracts multi-scale features f1, f2, f3, and f4 layer by layer.
[0025] Optionally, the multi-scale feature guidance process of the present invention is:
[0026] Assume that the feature of a certain level of U-Net is h, and the conditional feature is f = {f1, f2, f3, f4}. The calculation process of the cross attention mechanism at each level can be expressed as:
[0027] h′=CrossAttention(h,f)
[0028] Among them, the CrossAttention function represents the cross attention calculation, and the features h at each level of U-Net are updated to h′.
[0029] Optionally, during the second stage training of the present invention, the total loss function includes: L1 and L2 losses of the image component and the noise component, L1 loss of the reconstructed image, and cross entropy loss and FocalLoss loss of the final prediction result.
[0030] Optionally, after the second stage of training, the present invention selects the best model based on the performance indicators on the validation set for the semantic segmentation task of optical remote sensing images.
[0031] Optionally, during the second stage training of the present invention, the latent variables generated by the encoder in the autoencoder are further preprocessed before being input into the diffusion model. The preprocessing is: multiplying the latent variables by a scaling factor of 0.3.
[0032] Beneficial effects
[0033] Based on the theoretical framework of the diffusion probability model, this method innovatively proposes a diffusion-based semantic segmentation method for multi-classification tasks. The core architecture of this method consists of three key components: first, an autoencoder is used to represent the latent features of the input image; second, a denoising U-Net network is designed to gradually predict the image and noise separately during the diffusion process and recover the semantic information; finally, a conditional encoding module is introduced to enhance the model's ability to express specific semantic features. A phased training strategy is adopted for these three modules: first, the autoencoder is pre-trained separately to obtain an effective feature representation space, and then the denoising U-Net and conditional encoding modules are jointly trained to achieve accurate semantic segmentation. This modular design not only ensures training stability but also successfully extends the diffusion model to complex multi-class semantic segmentation tasks. Experiments on multiple public datasets verify the effectiveness and superiority of this method. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0035] Figure 1 A schematic diagram of the overall structure of a high-resolution remote sensing image semantic segmentation method based on a diffusion model provided by the present invention;
[0036] Figure 2 This is a structural diagram of the decoupled diffusion model provided by the present invention. DETAILED DESCRIPTION
[0037] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0038] It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments may be combined with each other; and, based on the embodiments in this disclosure, all other embodiments obtained by persons of ordinary skill in the art without creative work are within the scope of protection of this disclosure.
[0039] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.
[0040] The embodiment of the present application provides a semantic segmentation method for high-resolution remote sensing images based on a diffusion model. It applies the diffusion probability model to multi-classification semantic segmentation tasks for the first time, and has significant advantages in processing complex scenes and small target segmentation, bringing a new research paradigm to the field of semantic segmentation. The present invention designs a method based on implicit space diffusion model. The network uses the implicit diffusion model (Latent Diffusion Model, LDM) to map the image to the latent space through the encoder of the autoencoder to generate latent variables, and then applies the diffusion probability model (Diffusion Probabilistic Model, DPM) to the latent variables for noise addition. At the same time, the original image is conditionally encoded using the conditional coding module and introduced into the denoising U-Net to realize conditional guided diffusion generation. Finally, the generated latent variables are mapped back to the original image space through the decoder of the autoencoder to realize multi-classification segmentation. The present invention verifies the effectiveness and robustness of the proposed method by conducting experiments on multiple public optical remote sensing data sets. The specific process of the method is as follows:
[0041] The first stage of training: the high-resolution remote sensing image is annotated with the type of each pixel to generate a labeled image, and the labeled image is used to train the autoencoder including the encoder and decoder;
[0042] Second stage training: Based on the autoencoder, a conditional diffusion model is loaded and trained, including a noise injection module, a conditional encoding module, and a denoising U-Net. The autoencoder parameters are fixed during training. The noise injection module adds noise to the latent variables output by the encoder. The conditional encoding module is used to extract multi-scale features from the high-resolution image. The denoising U-Net is guided by the multi-scale features to predict image components and noise components and reconstruct the image. The decoder is used to decode the reconstructed image.
[0043] Image semantic segmentation: Use the trained network to perform semantic segmentation on high-resolution remote sensing images.
[0044] Each step of the above process of this embodiment is described in detail below:
[0045] Step 1: First stage training: Label the type of each pixel in the high-resolution remote sensing image to generate a label image, and use the label image to train the autoencoder including the encoder and decoder.
[0046] Normalization of label images: The input dimension in the encoder branch of the autoencoder is R 1×H×W The label image GT is annotated with the land use type of each pixel in the remote sensing image. H and W are the height and width of the label image, respectively. To eliminate data distribution differences between different regions, the label image GT is normalized so that its pixel values are distributed within the interval [-1, 1].
[0047] Feature extraction and dimensionality reduction: An autoencoder consists of two parts: an encoder and a decoder. The encoder compresses high-dimensional input data into low-dimensional feature representations, while the decoder reconstructs these low-dimensional feature representations into high-dimensional output data. The encoder typically consists of multiple convolutional layers and downsampling layers, while the decoder consists of multiple deconvolutional layers and upsampling layers.
[0048] The normalized labeled image enters a feature extraction network consisting of multiple convolutional and downsampling layers. In this network, the convolutional layers use convolution kernels to accurately extract image features, while the downsampling layers progressively reduce the size of the feature map, efficiently compressing the originally high-dimensional labeled image into a low-dimensional latent space, resulting in the latent space feature representation z.
[0049] The encoder part of the self-encoder is Encode, the label image G∈R 1×H×W As input, after being processed by the encoder, the latent space feature representation is obtained The encoding process can be expressed by the following formula:
[0050] z=Encode(GT) (1)
[0051] Sampling: It is usually assumed that the spatial feature representation z obtained after encoding obeys a certain probability distribution, such as Gaussian distribution. Let z obey the mean μ and variance σ 2 Gaussian distribution, that is, z~N(μ,σ 2 ). Use the reparameterization technique to sample from this feature distribution to obtain the latent variable z0.
[0052] The core idea of the reparameterization technique is to represent the sampling process as a deterministic transformation, so that the sampling process can be back-propagated. Through the reparameterization technique, the problem of sampling from a non-standard Gaussian distribution can be transformed into the problem of sampling from a standard normal distribution. The specific formula is as follows:
[0053] z0=μ+σ·∈ (2)
[0054] Here, ∈ is a random variable sampled from the standard normal distribution N(0,1).
[0055] Reconstruction: The decoder part of the autoencoder is Decode, which takes the latent variable z0 as the input of the decoder. The decoder restores z0 to the size of the original feature map Y∈R through multi-level upsampling and feature reconstruction operations. N×H×W , N is the number of categories, corresponding to N types of features. The reconstruction process is expressed as follows:
[0056] Y=Decode(z0) (3)
[0057] Loss calculation during the first-stage autoencoder training: To ensure that the generated images are visually highly similar to real images and have sufficient detail quality and realism, cross entropy loss and adversarial loss are used. The adversarial loss uses the discriminator to evaluate the authenticity of the reconstructed image Y, thereby further improving the detail and realism of the generated images.
[0058] Step 2: Second stage training: Based on the autoencoder, load the diffusion model, conditional encoding module and denoising U-Net, set the autoencoder parameters to non-trainable, and train the diffusion model in the latent space;
[0059] In this step, based on the pre-trained autoencoder model trained in the first phase, the conditional diffusion model is loaded, including the noise injection module, the conditional encoding module, and the denoising U-Net. The autoencoder parameters are set to non-trainable. This is to maintain the feature extraction capability of the autoencoder in subsequent training, focusing on training the diffusion model and the conditional encoding module. The training process is as follows:
[0060] Autoencoder encoding and preprocessing: The autoencoder uses the encoder to encode the label image GT to obtain the spatial feature representation z, and then samples the latent variable z0 from the feature distribution. To make it more suitable for subsequent model processing, z0 is preprocessed by multiplying the latent variable z0 by a scaling factor of 0.3 before being input into the diffusion model.
[0061] Noise injection module for noise addition: In the image generation and processing process based on the diffusion model, noise addition and prediction are key links. Their purpose is to simulate the noise addition and restoration of data during the diffusion process, so that the model can learn the distribution characteristics of the data. Figure 2 As shown, this embodiment adopts a decoupled diffusion model. The decoupled forward diffusion process is controlled by a combination of explicit transition probabilities and standard Wiener processes, corresponding to the mapping from image to zero and the mapping from zero to noise, respectively. The mapping from image to zero is controlled by explicit transition probabilities, while the mapping from zero to noise is controlled by the standard Wiener process. The forward diffusion process formula of the decoupled diffusion model is expressed as follows:
[0062] q(z t |z0)=N(z0+∫0f t dt,tI) (5)
[0063] Among them, z0 is the initial variable, z t is the latent variable after adding noise, f t is the explicit transition function, and I is the identity matrix. The initial noise n is generated according to the size of the latent variable z0, where the noise n follows a standard normal distribution, i.e., n~N(0,I). The process of generating noise can be expressed as:
[0064] n=Sample(N(0,I),shape(z0)) (6)
[0065] Here, the Sample function represents sampling from the specified distribution, and shape(z0) represents the shape of the latent variable z0. Adding the generated initial noise n to the latent variable z0 will give the noisy latent variable z t , its mathematical expression is:
[0066]
[0067] The conditional encoding module performs multi-scale feature extraction: the original input optical remote sensing image is X, and its dimension is R 3 ×H×W , where 3 represents the three channels of the image (red, green, and blue). The conditional encoding module uses Swin Transformer to process the image X layer by layer, extracting multi-scale features f1, f2, f3, and f4 layer by layer, whose dimensions are The formula is as follows:
[0068] f1,f2,f3,f4=Swin Transformer(X) (4)
[0069] These features correspond to different levels of image resolution and contain feature information at different scales. Low-resolution features (such as f4) contain global semantic information about the image, helping to identify major objects and scene categories in the image; whereas high-resolution features (such as f1) retain more detailed information, which is important for accurately segmenting object boundaries and identifying small targets. The conditional encoding module uses the Swin Transformer to conditionally encode the original input optical remote sensing image X. This module can effectively extract multi-scale features, providing rich and comprehensive feature information for subsequent tasks such as image semantic segmentation, thereby improving the performance and accuracy of the model.
[0070] Denoising U-Net prediction: Use denoising U-Net to predict the latent variable z t and time step t for joint prediction. The denoising U-Net is an encoder-decoder structure, and there is a jump connection between the encoder and decoder, which can effectively capture feature information of different scales. In order to enhance the conditional controllability of the diffusion process, it is necessary to inject the semantic information of the original remote sensing image into the denoising process to achieve conditional guidance, that is, to use the conditional encoding module to extract the multi-scale features f1, f2, f3, f4 of the image to be segmented, and integrate them into the features of each level of the U-Net encoder and decoder through the cross-attention mechanism. The cross-attention mechanism allows the model to establish connections between different feature representations, thereby effectively using conditional information to guide the denoising process. Suppose the feature of a certain level of U-Net is h, and the conditional feature is f = {f1, f2, f3, f4}. The calculation process of the cross-attention mechanism at each level can be expressed as:
[0071] h′=CrossAttention(h,f) (9)
[0072] The CrossAttention function represents a cross-attention calculation. It achieves information fusion by calculating the attention score between the query, key, and value, and performing a weighted summation of features. After the cross-attention mechanism is applied, the U-Net feature h at each level is updated to h′, which incorporates the conditional feature f at each level. By incorporating the feature information of the original remote sensing image into the denoising process, it can provide conditional guidance for the denoising process, effectively improving the semantic consistency and detail quality of the final segmentation result.
[0073] For example, cultivated land usually has regular shapes and neat textures, while woodlands have irregular distributions and complex texture features. Texture, structure and other features in the conditional information can guide the denoising U-Net to more accurately identify and segment different land types during the prediction process, thereby improving the accuracy and reliability of the segmentation results. Under the guidance of the conditional information, the denoising U-Net predicts the image component C from the latent variables through different decoder branches. pred and noise component n pred , then the prediction process can be expressed as:
[0074] C pred ,n pred =Net(z n ,t) (8)
[0075] At a certain time step t, the denoising U-Net will be based on the input noise hidden variable z n The noise and corresponding image components that should be removed at the current moment are predicted. This process simulates the reverse denoising process of the diffusion model, and by continuously predicting and removing noise, the original image information is gradually restored.
[0076] The predicted image component C pred and noise component n pred Perform integration and reconstruct image x rec , and x rec Input the decoder of the autoencoder for decoding.
[0077] The decoder of the autoencoder performs image decoding: The decoder of the autoencoder is used to restore the low-dimensional features in the latent space to the high-dimensional image in the original space. It consists of multiple layers, usually including deconvolution layers (transposed convolution layers), upsampling layers, and nonlinear activation function layers. These layers work together to gradually amplify and reconstruct the input feature map to restore the output with the original image size and semantic information, and obtain the final semantic segmentation prediction result X pred ∈R N ×H×W , N is the number of categories. For X pred Each pixel point (i, j) in has a corresponding probability value on N channels, indicating the possibility that the pixel belongs to each land use type.
[0078] To obtain a specific category label for each pixel, the probability values of each pixel in N channels are compared, and the category corresponding to the channel with the largest probability value is selected as the category label of the pixel. This result is a prediction of the land use type in the optical remote sensing image, and each pixel is assigned a corresponding category label.
[0079] Loss calculation and back propagation: Training the second stage latent space diffusion model requires supervision of both data and noise components. Calculate the image component C pred and noise component n pred L1, L2 loss, reconstruct the image x rec The L1 loss and the final prediction result X pred The cross-entropy loss and focal loss are combined to form the final total loss. Backpropagation is performed based on the total loss to update model parameters and optimize network performance. For example, a large cross-entropy loss indicates low classification accuracy, and adjusting model parameters through backpropagation can improve classification accuracy.
[0080] After training is complete, the best model is selected based on performance metrics on the validation set (such as mIoU, Accuracy, etc.). For example, the mean intersection over union (mIoU) of the models is calculated on the validation set, and the model with the highest mIoU is selected for the semantic segmentation task of optical remote sensing images.
[0081] Step 3: Use the trained network to achieve semantic segmentation of high-resolution remote sensing images
[0082] The image to be segmented is input into the trained neural network, and the neural network automatically performs the following sampling and prediction process.
[0083] After the training process of step 1 and step 2, the reverse denoising process formula of the decoupled diffusion model can be obtained as follows:
[0084]
[0085] The sampling phase begins with generating initial noise. This is assumed to follow a standard normal distribution, as the standard normal distribution is both random and universal, providing a rich information foundation for subsequent denoising. This initial noise serves as the starting point for subsequent denoising.
[0086] After generating the initial noise, the denoising operation is then performed step by step according to a pre-set time step. The time step gradually decreases from the maximum value (corresponding to the state with the highest noise level) to the state with the noise-free original image. At each time step, a prediction is made based on the current image, the current time step, and the condition information to obtain the image component and the noise component. After the predicted image component and noise component are obtained, the current image is updated.
[0087] The update method is as shown in Equation 10, gradually removing noise. As the time step progresses, the noise gradually decreases, and the image gradually recovers meaningful semantic information. When the time step is reduced to 0, the denoised image is obtained.
[0088] Next, the denoised image is input into the decoder for decoding to obtain the final semantic segmentation prediction result.
[0089] Taking the semantic segmentation of land use types in optical remote sensing images as an example, starting with completely random initial noise, the image is continuously adjusted at each time step based on conditional information and model predictions, gradually removing the noise. As the denoising process progresses, the image gradually becomes clearer from a blurry and noisy state, ultimately achieving accurate semantic segmentation of land use types, clearly displaying the corresponding land feature categories in different areas of the image, such as cultivated land, forest land, and water bodies.
[0090] The overall network structure described in the above steps is as follows Figure 1 As shown in the figure, based on the theoretical framework of the diffusion probability model, an innovative semantic segmentation method based on the diffusion model is proposed for multi-classification tasks. Its core architecture contains three key components: first, the autoencoder is used to represent the latent features of the input image; second, a denoising U-Net network is designed to gradually predict the image and noise during the diffusion process to restore the semantic information; third, a conditional coding module is introduced to enhance the model's ability to express specific semantic features. A phased training strategy is adopted, first pre-training the autoencoder separately to obtain an effective feature representation space, and then jointly training the denoising U-Net and conditional coding modules to achieve accurate semantic segmentation. This modular design not only ensures training stability, but also successfully extends the diffusion model to complex multi-classification semantic segmentation tasks. The effectiveness and superiority of this method have been verified by multiple public datasets.
[0091] like Figure 2 As shown in the figure, a decoupled diffusion model is designed based on the principles of probabilistic diffusion models and implemented in latent space by a denoising U-Net. The decoupled forward diffusion process is controlled by a combination of explicit transition probabilities and a standard Wiener process, corresponding to the image-to-zero mapping and the zero-to-noise mapping, respectively. The image-to-zero mapping is controlled by the explicit transition probabilities, while the zero-to-noise mapping is controlled by the standard Wiener process. The denoising U-Net predicts the image and noise from latent variables. Training this model requires supervision of both the data and noise components. The corresponding derivation ensures that the backward diffusion process can smoothly generate latent variables.
[0092] Compared with the prior art, the embodiments of the present application have the following effects:
[0093] This method uses an autoencoder architecture to achieve image feature compression and reconstruction. Specifically, the encoder maps high-dimensional remote sensing image labels into a low-dimensional latent space through multi-layer convolution and downsampling operations, generating latent variable representations with rich semantic information. After the diffusion model generates latent variables, the decoder gradually reconstructs the latent variables into segmentation results in the original image space through symmetrical upsampling and deconvolution operations. This autoencoder-based compression-reconstruction framework not only reduces the modeling complexity of the diffusion model but also successfully achieves accurate multi-category segmentation of high-resolution remote sensing images, providing a new solution for remote sensing image interpretation.
[0094] This method leverages the principles of diffusion probability models to design a decoupled diffusion model, implemented in latent space by a denoising U-Net. The decoupled forward diffusion process is controlled by a combination of explicit transition probabilities and a standard Wiener process, representing the image-to-zero mapping and the zero-to-noise mapping, respectively. The denoising U-Net predicts the image and noise separately from the latent variables. Training the decoupled diffusion model requires supervision of both the data and noise components. Therefore, through corresponding derivation, the latent variables can be smoothly generated during the backward diffusion process.
[0095] This method designs a conditional encoding module to encode deep features of the input original image and apply this information to the diffusion process, achieving conditionally guided generation. During the diffusion process, these conditional features interact with the latent features of the diffusion model through a cross-attention mechanism, effectively incorporating prior knowledge such as the original image's structural information and texture features into the diffusion process. Guided by the conditional encoding module, the diffusion model can produce more accurate and consistent segmentation results, demonstrating enhanced robustness in dealing with complex backgrounds and noise interference.
[0096] The purpose of the present invention is to break through the limitations of traditional semantic segmentation methods and innovatively apply the Diffusion Probabilistic Model to multi-classification semantic segmentation tasks for the first time. In particular, it shows significant advantages when dealing with complex scenes and small target segmentation tasks, and provides a new research paradigm for the field of semantic segmentation. In response to the problems of blurred boundaries and low accuracy in small target segmentation in complex scene segmentation of existing deep learning methods, the present invention converts the image segmentation problem into a conditional generation process by designing an end-to-end segmentation framework based on a diffusion model. Specifically, the method utilizes the gradual denoising characteristics of the diffusion model, converts the true segmentation map into a noise distribution through a forward diffusion process, and then reconstructs an accurate segmentation result from the noise through a reverse denoising process.
[0097] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A semantic segmentation method for high-resolution remote sensing images based on a diffusion model, characterized in that: The specific process is: The first stage of training: the high-resolution remote sensing image is annotated with the type of each pixel to generate a labeled image, and the labeled image is used to train the autoencoder including the encoder and decoder; Second stage training: Based on the autoencoder, a conditional diffusion model is loaded and trained, including a noise injection module, a conditional encoding module, and a denoising U-Net. The autoencoder parameters are fixed during training. The noise injection module adds noise to the latent variables output by the encoder. The conditional encoding module is used to extract multi-scale features from the high-resolution image. The denoising U-Net is guided by the multi-scale features to predict image components and noise components and reconstruct the image. The decoder is used to decode the reconstructed image. Image semantic segmentation: Use the trained network to perform semantic segmentation on high-resolution remote sensing images.
2. The high-resolution remote sensing image semantic segmentation method based on the diffusion model according to claim 1, characterized in that: It also includes normalizing the label image so that the pixel values of the label image are distributed within the interval [-1,1].
3. The high-resolution remote sensing image semantic segmentation method based on the diffusion model according to claim 1, characterized in that: During the first stage of training, cross entropy loss and adversarial loss are calculated.
4. The high-resolution remote sensing image semantic segmentation method based on the diffusion model according to claim 1, characterized in that: The noise injection module adds noise to the latent variable output by the encoder, and the latent variable after adding noise is z t ; Among them, z0 is the initial variable output by the encoder, f t is an explicit transition function, and the initial noise n is generated according to the size of the latent variable z0.
5. The high-resolution remote sensing image semantic segmentation method based on the diffusion model according to claim 4, characterized in that: The initial noise n is: n=Sample(N(0,I),shape(z0)) Among them, the Sample function represents sampling from the specified distribution, shape(z0) represents the shape of the latent variable z0, and the noise n obeys the standard normal distribution, that is, n~N(0,I).
6. The method for semantic segmentation of high-resolution remote sensing images based on a diffusion model according to claim 4, characterized in that: The conditional coding module uses Swin Transformer to process high-resolution remote sensing images layer by layer and extract multi-scale features f1, f2, f3, and f4 layer by layer.
7. The method for semantic segmentation of high-resolution remote sensing images based on a diffusion model according to claim 4, characterized in that: The multi-scale feature guidance process is: Assume that the feature of a certain level of U-Net is h, and the conditional feature is f = {f1, f2, f3, f4}. The calculation process of the cross attention mechanism at each level can be expressed as: h′=CrossAttention(h,f) Among them, the CrossAttention function represents the cross attention calculation, and the features h at each level of U-Net are updated to h′.
8. The method for semantic segmentation of high-resolution remote sensing images based on a diffusion model according to claim 4, characterized in that: During the second stage of training, the total loss function includes: L1 and L2 losses of the image component and noise component, L1 loss of the reconstructed image, and cross entropy loss and FocalLoss loss of the final prediction result.
9. The high-resolution remote sensing image semantic segmentation method based on diffusion model according to claim 1, characterized in that: After the second stage of training, the best model is selected based on the performance indicators on the validation set for the semantic segmentation task of optical remote sensing images.
10. The high-resolution remote sensing image semantic segmentation method based on diffusion model according to claim 1, characterized in that: During the second stage of training, the latent variables generated by the encoder in the autoencoder are preprocessed before being input into the diffusion model. The preprocessing is: multiplying the latent variables by a scaling factor of 0.3.
Citation Information
Cited By
Complex structural part scribed line segmentation method, device and system and storage medium
CN120564003A
Remote sensing image semantic segmentation method and device based on diffusion model
CN120912891A
Implicit diffusion super-resolution assisted remote sensing image small target detection and identification method
CN121147498A
AIGC image fine granularity detection method based on adaptive layering mechanism
CN121236569A
Aigc image fine-grained detection method based on adaptive hierarchical mechanism
CN121236569B