Remote sensing image road segmentation method based on diffusion pre-training

By adopting a diffusion pre-training method in the remote sensing image road segmentation technology, and using the denoising diffusion probability model for self-supervised learning, the problems of insufficient accuracy and dependence on labeled data in the existing technology are solved, and the road segmentation effect with high accuracy and robustness is achieved.

CN120047683AActive Publication Date: 2025-05-27NANJING UNIV OF SCI & TECH

Patent Information

Application Number
CN202411904256.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-27
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

The existing remote sensing image road segmentation technology has shortcomings in accuracy, robustness and dependence on labeled data. Especially in intensive prediction tasks, insufficient generalization capabilities of the model and labeling errors are likely to lead to misjudgment.

Method used

Using a diffusion pre-training method, the pre-trained data set and fine-tuning data set are constructed, and the segmented backbone model is built using the UNet network, and the denoising diffusion probability model is used for self-supervised learning in the pre-training stage to reduce the dependence on the annotated data.

Benefits of technology

It significantly improves the segmentation accuracy of road areas in remote sensing images, improves the robustness of the model in complex contexts, reduces the need for a large amount of labeled data, and realizes effective learning with fewer labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047683A_ABST
    Figure CN120047683A_ABST
Patent Text Reader

Abstract

The invention proposes a remote sensing image road segmentation method based on diffusion pre-training, and the method comprises the steps: constructing a segmentation backbone model based on a UNet network, enabling a coding part to comprise a ConvBlock and four coding blocks, enabling a decoding part to comprise four decoding blocks and a prediction head, and outputting a road segmentation image after a remote sensing image passes through the ConvBlock, the four coding blocks, the four decoding blocks and the prediction head; based on the de-noising diffusion probability model, noise is added to the remote sensing image in the pre-training data set to obtain a remote sensing image with noise, the remote sensing image with noise is input into the segmentation backbone model, a reconstructed image is output by a fourth decoding block, and the segmentation backbone model pre-training process is completed by taking the minimum values of the reconstructed image and the original remote sensing image as the target; and loading a segmentation backbone model weight obtained by pre-training, and finishing segmentation backbone model fine tuning based on the fine tuning data set for an actual remote sensing image road segmentation task. According to the method, the segmentation precision can be ensured under the condition of fewer labels, and complex backgrounds and noise interference can be effectively processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer vision and deep learning technologies, and particularly to a method for remote sensing image road segmentation based on diffusion pre-training. Background Art

[0002] Remote sensing image road segmentation algorithms are mainly applied in fields such as urban planning, traffic monitoring, and disaster assessment. Accurate road segmentation is crucial for aspects such as traffic flow prediction, urban expansion planning, and public facility layout. Automated road extraction can greatly improve the efficiency and accuracy of planning. Traditional fully supervised methods have the problem of relying heavily on image annotation. Especially in tasks such as remote sensing image road segmentation with intensive prediction, deep learning models usually face challenges such as difficult dataset annotation and insufficient model generalization ability. The generalization ability of the model may be affected by the quality of the labels, and annotation errors will lead to misjudgment of the model. In addition, obtaining high-quality annotation data is not only costly but also very time-consuming, especially in intensive prediction tasks such as segmentation, making this problem particularly prominent. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for remote sensing image road segmentation based on diffusion pre-training to solve the deficiencies of existing road segmentation technologies in processing remote sensing images, especially the challenges in terms of accuracy, robustness, and dependence on annotation data.

[0004] The technical solution for achieving the purpose of the present invention is as follows: A method for remote sensing image road segmentation based on diffusion pre-training includes the following contents and steps:

[0005] Step 1: Construct a pre-training dataset and a fine-tuning dataset, where the pre-training dataset only contains remote sensing images of complete road maps, and the fine-tuning dataset contains remote sensing images and corresponding road annotations;

[0006] Step 2: Based on the UNet network, construct a segmentation backbone model. The encoding part includes a ConvBlock and 4 encoding blocks, and the decoding part includes 4 decoding blocks and a prediction head. After the remote sensing image passes through the ConvBlock, 4 encoding blocks, 4 decoding blocks, and a prediction head, a road segmentation picture is output;

[0007] Step 3: Based on the denoising diffusion probability model, add noise to the remote sensing images in the pre-training dataset to obtain noisy remote sensing images, input them into the segmentation backbone model, and the reconstructed image is output by the fourth decoding block. With the goal of minimizing the value between the reconstructed image and the original remote sensing image, complete the pre-training process of the segmentation backbone model;

[0008] Step 4: Load the weights of the segmentation backbone model obtained by pre-training and complete the fine-tuning of the segmentation backbone model based on the fine-tuning dataset;

[0009] Step 5: Input the remote sensing image to be measured into the fine-tuned segmentation backbone model to complete the actual remote sensing image road segmentation task.

[0010] Further, in Step 1, construct a pre-training dataset and a fine-tuning dataset. The specific method is as follows:

[0011] Collect images from multiple public datasets, crop and remove the unannotated images without a complete roadmap, and construct a pre-training dataset;

[0012] Collect images from multiple public datasets, select the data with images and corresponding road annotations, and construct a fine-tuning dataset.

[0013] Further, in Step 2, construct a segmentation backbone model based on the UNet network. The encoding part includes a ConvBlock and 4 encoding blocks, and the decoding part includes 4 decoding blocks and a prediction head. After the remote sensing image passes through the ConvBlock, 4 encoding blocks, 4 decoding blocks, and a prediction head, a road segmentation picture is output. The specific method is as follows:

[0014] Given a remote sensing image This input passes through the ConvBlock and becomes Then it passes through the first encoding block and becomes a feature map And so on, passing through the second, third, and fourth encoding blocks respectively to obtain the feature maps X 2 ,X 3 ,X 4 , in this process, the number of channels gradually increases and the size of the feature map gradually decreases; correspondingly, the feature map X 4 is concatenated with the feature map X 3 and then input into the fourth decoding block to obtain the feature map Y 4 ,Y 4 is concatenated with X 2 and then input into the third decoding block to obtain the feature map Y 3 ; and so on, the output of the remaining decoding blocks is Y 2 ,Y 1 ; finally, the feature map Y 1 passes through a 1×1Conv and outputs a road segmentation picture;

[0015] The encoding block is composed of a PatchMerging and an SSFBlock in series, and the decoding block is composed of a PatchExpand and an SSFBlock in series;

[0016] a. Spectral-Spatial Fusion Block (SSFBlock)

[0017] PatchMerging performs operations on the feature map X of the (l - 1)-th layer l-1After the operation, the feature map is obtained In the l-th SSFBlock, the feature map First, it goes through layer normalization (LayerNorm) and spectral convolution operation, and the obtained feature map is added to the feature map to obtain the feature map After that, the feature map goes through layer normalization and MLP operation to obtain The obtained output is added to the feature map to obtain the output X of this layer l ;

[0018] b. Spectral convolution (SP-Conv)

[0019] The feature map of the l-th layer After normalization, the feature map is obtained In the spectral convolution module, the feature map is divided into two tensors along the channel dimension where X' l is directly input into the adaptive feature selection and enhancement (AFSE) layer to obtain the feature map X′ l ' After DCT transformation, it is input into the AFSE layer and then inverse-transformed to obtain the feature map Finally, and are concatenated along the channel to obtain the spectral convolution output

[0020] c. AFSE layer

[0021] The AFSE layer is composed of a detail enhancement branch (DEB), a meta-attention branch (MAB), and a linear transformation branch (LTB) in parallel;

[0022] DEB is a depthwise separable convolution that enhances local detail features by sequentially combining depthwise convolution and pointwise convolution; LTB is a parallel branch of DEB that performs linear transformation in the channel dimension through 1×1 convolution; For the feature map X' l , the basic structure composed of DEB and LTB is expressed as follows:

[0023]

[0024] Among them, represents 1×1 convolution, represents 3×3 depthwise convolution;

[0025] The MAB is a meta-attention module designed to refine semantic information through an attention mechanism similar to squeeze-and-excitation to generate semantic-aware weights for re-weighting and recalibrating features from the DEB and LTB branches; for a given feature map The MAB process is described as follows:

[0026]

[0027] where GAP(·) and GMP(·) represent global average pooling and global max pooling operations respectively, and FC m (·) represents the fully connected operation;

[0028] Finally, the output of the AFSE layer is expressed as:

[0029]

[0030] Furthermore, in step 3, based on the denoising diffusion probability model, noise is added to the remote sensing images in the pre-training dataset to obtain noisy remote sensing images, which are input into the segmentation backbone model. The reconstructed images are output by the fourth decoding block. With the goal of minimizing the value between the reconstructed images and the original remote sensing images, the pre-training process of the segmentation backbone model is completed. The specific method is as follows:

[0031] The clear remote sensing image x 0 First, it undergoes DCT transformation to the frequency domain to obtain the frequency spectrum image Then, noise is gradually added to obtain noisy frequency spectrum images Next, the inverse DCT is performed on the noisy frequency spectrum images to obtain the noisy remote sensing images x 1 ,…,x t-1 ,x t ,…x T , and then they are respectively input into the segmentation backbone model. The reconstructed images are output by the fourth decoding block. The reconstructed images are With the goal of minimizing the difference between the reconstructed images and the original image x 0 , the pre-training process of the segmentation backbone model is completed;

[0032] When pre-training the segmentation backbone model, the mean squared error (MSE) is selected as the loss function to measure the difference between the restored image and the real image. During the optimization process, the AdamW optimizer is used, the initial learning rate is set to 1e-4, and a dynamic learning rate scheduling strategy is adopted to ensure the stability of the training process and accelerate convergence. The maximum number of diffusion steps for training is 1000, and the batch size is 128.

[0033] Furthermore, in step 4, the weights of the pre-trained segmentation backbone model are loaded, and the segmentation backbone model is fine-tuned based on the fine-tuning dataset. The specific method is as follows:

[0034] Fine-tune the segmented backbone model given in Step 2 and the fine-tuning dataset made in Step 1, load the pre-trained weights of the segmented backbone model obtained in Step 3, and apply it to the labeled fine-tuning dataset for fine-tuning. During fine-tuning, the prediction head of the model is initialized to 0, and the other parts remain with the original pre-trained weights unchanged;

[0035] During the fine-tuning process, the mean squared error (MSE) is used as the loss function, and the AdamW optimizer is used to update the parameters. The batch size is set to 8.

[0036] Further, in Step 5, input the remote sensing image to be measured into the fine-tuned segmented backbone model to complete the actual remote sensing image road segmentation task. The specific method is as follows:

[0037] First, perform a cropping operation on the input remote sensing image. Subsequently, input the processed image into the segmented backbone model for inference. Through the Sigmoid activation function, output the probability value of each pixel belonging to the road area. Finally, generate the final road segmentation result according to the predicted probability value and save it as a grayscale image, where black represents the background and white represents the road.

[0038] A remote sensing image segmentation system based on diffusion pre-training implements the described remote sensing image segmentation method based on diffusion pre-training to achieve remote sensing image segmentation based on diffusion pre-training.

[0039] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the described remote sensing image segmentation method based on diffusion pre-training to achieve remote sensing image segmentation based on diffusion pre-training.

[0040] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the described remote sensing image segmentation method based on diffusion pre-training to achieve remote sensing image segmentation based on diffusion pre-training.

[0041] Compared with the prior art, the significant advantages of the present invention are as follows:

[0042] 1) Based on the improved UNet model, combined with the spectral-spatial fusion block (SSFBlock), it enhances the model's ability in local and global feature extraction, significantly improving the segmentation accuracy of the road area in the image.

[0043] 2) In the pre-training stage, a frequency-domain diffusion model is used for self-supervised learning, effectively improving the model's robustness in complex backgrounds and solving the challenge of few annotations. Description of the Drawings

[0044] Figure 1 This is the flowchart of road segmentation for remote sensing images of the present invention.

[0045] Figure 2 This is the flowchart of frequency-domain diffusion pre-training of the present invention.

[0046] Figure 3 This is the overall structure diagram of the backbone model for road segmentation of the present invention.

[0047] Figure 4 This is the schematic diagram of the SSFBlock module of the backbone model for road segmentation of the present invention.

[0048] Figure 5 This is the schematic diagram of the spectral convolution module used in the backbone model for road segmentation of the present invention

[0049] Figure 6 This is the test effect diagram of the embodiment of the present invention Detailed implementation manners

[0050] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0051] By introducing a diffusion self-supervised pre-training strategy and transferring it to the frequency-domain perspective, the present invention can effectively improve the accuracy of road segmentation in remote sensing images and significantly reduce the need for a large amount of labeled data; to further improve the segmentation effect, the present invention designs a suitable backbone network, which combines the advantages of the UNet framework and innovatively integrates local and global features in the spatial domain and the discrete cosine transform (DCT) domain.

[0052] As Figure 1 shown, the method for road segmentation of remote sensing images based on diffusion pre-training includes the following steps:

[0053] Step 1, constructing a pre-training dataset and a fine-tuning dataset:

[0054] The dataset includes a pre-training dataset and a fine-tuning dataset. Among them, the pre-training dataset only includes remote sensing images, while the fine-tuning dataset includes remote sensing images and corresponding road annotations. The images for pre-training and fine-tuning are all from public datasets.

[0055] For the pre-training dataset, pictures from multiple public datasets (including SpaceNet3, OSM, Ottawa, RoadTracer) are collected, a total of 3,846 pictures, and then they are cropped, and after removing the images without a complete route map, there are approximately 50,000 unlabeled images; the fine-tuning dataset uses public datasets (such as DeepGlobe, Massachusetts) with remote sensing images and corresponding road annotations, which are divided into a training set and a validation set according to 8:2.

[0056] All training data are cropped into 256×256 image patches, and data augmentation by random flipping is applied.

[0057] Step 2, construct a road segmentation backbone model:

[0058] As Figure 3 shown, the backbone model is based on the UNet structure, adopts a symmetric encoder and decoder design, and realizes the fusion of shallow and deep features through skip connections. The encoding part consists of a stem block (ConvBlock) composed of a 3×3 convolution and four encoding blocks, while the corresponding decoding part includes four decoding blocks and a lightweight prediction head (1×1Conv).

[0059] Given a remote sensing image This input passes through the ConvBlock and becomes Then it passes through the first encoding block to become a feature map And so on, passing through the second, third, and fourth encoding blocks respectively to obtain the feature maps X 2 , X 3 , X 4 , in this process, the number of channels gradually increases and the size of the feature maps gradually decreases. Correspondingly, the feature map X 4 is concatenated with the feature map X 3 and input into the fourth decoding block to obtain the feature map Y 4 , Y 4 is concatenated with Y 2 and input into the third decoding block to obtain the feature map Y 3 . And so on, the remaining decoding blocks output Y 2 , Y 1 . Finally, the feature map Y 1 passes through 1×1Conv and outputs the road segmentation picture.

[0060] The encoding block is composed of a PatchMerging and an SSFBlock in series, and the decoding block is composed of a PatchExpand and an SSFBlock in series.

[0061] a spectral-spatial fusion block (SSFBlock)

[0062] The SSFBlock, as Figure 4 shown, consists of a spectral convolution (SP-Conv) layer and a multi-layer perceptron (MLP) layer, and adopts the residual connection structure in Transformer. PatchMerging operates on the feature map X l-1 and obtains the feature map In the l-th SSFBlock, the feature map First, through layer normalization (LayerNorm) and spectral convolution operations, the obtained feature map is added to the feature map to obtain the feature map After that, the feature map goes through layer normalization and MLP operations to obtain The obtained output is added to the feature map to obtain the output X of this layer l . Its formula is expressed as follows:

[0063]

[0064] b Spectral convolution (SP-Conv)

[0065] As Figure 5 shown, the spectral convolution (SP-Conv) includes two parallel branches in the frequency domain and the spatial domain. The feature map of the l-th layer after normalization obtains the feature map In the spectral convolution module, the feature map is divided into two tensors along the channel dimension where X' l is directly input into the Adaptive Feature Selection and Enhancement (AFSE) layer to obtain the feature map X' l ′ after DCT transformation is input into the AFSE layer and then inverse-transformed to obtain the feature map Finally, and are concatenated along the channel to obtain the spectral convolution output

[0066] c AFSE layer

[0067] The AFSE layer is composed of a Detail Enhancement Branch (DEB), a Meta-Attention Branch (MAB), and a Linear Transformation Branch (LTB) in parallel.

[0068] The DEB is a depthwise separable convolution that enhances local detail features by sequentially combining depthwise convolution and pointwise convolution. The LTB is a parallel branch of the DEB that performs linear transformation in the channel dimension through 1×1 convolution. For the feature map X′ l , the basic structure composed of the DEB and the LTB can be expressed as follows:

[0069]

[0070] Among them, represents 1×1 convolution, represents 3×3 depthwise convolution.

[0071] The MAB is a meta-attention module designed to refine semantic information through an attention mechanism similar to squeeze-and-excitation to generate semantic-aware weights for re-weighting and recalibrating features from the DEB and LTB branches. For a given feature map The MAB process can be described as follows:

[0072]

[0073] where GAP(·) and GMP(·) denote global average pooling and global max pooling operations respectively, and FC m (·) represents the fully connected operation.

[0074] Finally, the output of the AFSE layer is expressed as:

[0075]

[0076] Step 3, Self-supervised Diffusion Pretraining

[0077] The pre-training framework draws on the classical denoising diffusion probabilistic model (DDPM), but extends it to the frequency domain, proposing a frequency-domain diffusion pre-training method. Its training framework is as Figure 2 shown. A clear remote sensing image x 0 is first transformed to the frequency domain by DCT to obtain a spectrogram and then noise is gradually added to obtain noisy spectrograms Subsequently, the noisy spectrograms are respectively subjected to iDCT to obtain noisy remote sensing images x 1 , …, x t-1 , x t , …x T , and then they are respectively input into the backbone model, and the reconstructed image is output by the fourth decoding block. The reconstructed image is With the goal of minimizing the difference between the reconstructed image and the original image x 0 , the pre-training process of the backbone model is completed.

[0078] Generally speaking, the pre-training architecture can be divided into four processes: discrete cosine transform, forward diffusion adding noise, reverse denoising, and inverse discrete cosine transform. This process is unsupervised training, does not rely on image labels, and only uses the information of the image itself for training.

[0079] During pre-training, the mean squared error (MSE) is selected as the loss function to measure the difference between the restored image and the real image. During the optimization process, the AdamW optimizer is used, the initial learning rate is set to 1e-4, and a dynamic learning rate scheduling strategy is adopted to ensure the stability of the training process and accelerate convergence. The maximum number of diffusion steps for training is 1000, the batch size is 128, and the optimization parameters are used to effectively improve the model performance and reduce overfitting.

[0080] Step 4, fine-tuning for road segmentation:

[0081] Fine-tune using the segmentation model given in Step 2 and the fine-tuning dataset made in Step 1, load the pre-trained model weights obtained in Step 3, and apply them to the labeled fine-tuning dataset for fine-tuning. During fine-tuning, the prediction head of the model (i.e., the Figure 3 1×1 Conv block) will be initialized to 0, while the other parts of the model will keep the original pre-trained weights unchanged.

[0082] During the training process, the mean squared error (MSE) is used as the loss function, and the AdamW optimizer is used for parameter update. The batch size is set to 8 to ensure effective fine-tuning under the condition of limited samples.

[0083] Step 5, inference and prediction:

[0084] After completing the fine-tuning, a fine-tuned segmentation model is obtained. In this step, the actual remote sensing image to be measured is input into the segmentation network fine-tuned in Step 4. First, the input remote sensing image is cropped. Subsequently, the processed image is input into the model such as Figure 3 for inference. Through the Sigmoid activation function, the model outputs the probability value of each pixel belonging to the road area. Finally, the final road segmentation result is generated according to the predicted probability value and saved as a grayscale image, where black represents the background and white represents the road.

[0085] The present invention also proposes a remote sensing image segmentation system based on diffusion pre-training, which implements the remote sensing image segmentation method based on diffusion pre-training to achieve remote sensing image segmentation based on diffusion pre-training.

[0086] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the remote sensing image segmentation method based on diffusion pre-training is implemented to achieve remote sensing image segmentation based on diffusion pre-training.

[0087] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the remote sensing image segmentation method based on diffusion pre-training is implemented to achieve remote sensing image segmentation based on diffusion pre-training.

[0088] In summary, by introducing diffusion self-supervised learning, the method of the present invention does not require image annotation during pre-training and only requires a small amount of annotation during fine-tuning. It can effectively learn the road features in remote sensing images with less labeled data. The pre-trained model can improve the segmentation accuracy, accelerate convergence, and significantly enhance the performance of the model in practical applications.

[0089] Embodiment

[0090] To verify the effectiveness of the proposed solution of the present invention, the Massachusetts dataset was selected for the following simulation experiments:

[0091] 1) Construct the pre-training dataset and the fine-tuning dataset

[0092] Collect 3,846 remote sensing images with complete road maps from the public datasets of SpaceNet3, OSM, Ottawa, and RoadTracer, and then crop them to obtain approximately 50,000 unlabeled remote sensing images to construct the pre-training dataset; divide the Massachusetts dataset into a training set and a validation set according to 8:2 to obtain the fine-tuning dataset.

[0093] During pre-training and fine-tuning, all images were cropped into image patches of 256×256, and the data augmentation method of random flipping was applied.

[0094] 2) Construct the backbone model for road segmentation

[0095] The backbone model is based on the UNet structure, adopts a symmetric encoder and decoder design, and realizes the fusion of shallow and deep features through skip connections. The encoding part consists of a stem block (ConvBlock) composed of a 3×3 convolution and 4 encoding blocks, while the corresponding decoding part includes 4 decoding blocks and a lightweight prediction head (1×1Conv).

[0096] 3) Self-supervised diffusion pre-training

[0097] During pre-training, the mean squared error (MSE) was selected as the loss function to measure the difference between the restored image and the real image. During the optimization process, the AdamW optimizer was used, the initial learning rate was set to 1e-4, and a dynamic learning rate scheduling strategy was adopted to ensure the stability of the training process and accelerate convergence. The maximum number of diffusion steps for training was 1000, the batch size was 128, and the optimization parameters were used to effectively improve the model performance and reduce overfitting.

[0098] 4) Fine-tuning for road segmentation

[0099] For fine-tuning, the segmentation model given in step 2 and the fine-tuning dataset made in step 1 were used, and the pre-trained model weights obtained in step 3 were loaded and applied to the labeled fine-tuning dataset for fine-tuning. During fine-tuning, the prediction head of the model (i.e., Figure 3 the 1×1Conv block) would be initialized to 0, while the other parts of the model would remain with the original pre-trained weights unchanged.

[0100] During the training process, the mean squared error (MSE) was used as the loss function, and the AdamW optimizer was used for parameter update. The batch size was set to 8 to ensure effective fine-tuning under the condition of limited samples.

[0101] The fine-tuning results are shown in Table 1.

[0102] Table 1 Comparison of fine-tuning results

[0103] Method Acc P R mIoU UNet 78.04 63.23 70.32 76.39 DDPM 77.14 84.24 79.35 83.86 Ours-model 78.81 85.88 78.81 83.29 Ours-pretrain 79.34 86.66 79.56 84.46

[0104] 5) Inference and prediction

[0105] Input the actual remote sensing image to be measured into the segmentation network fine-tuned in step 4. First, perform the operation of cropping the input remote sensing image. Subsequently, input the processed image into the model as Figure 3 for inference. Through the Sigmoid activation function, the model outputs the probability value of each pixel belonging to the road area. Finally, as Figure 6 shown, generate the final road segmentation result according to the predicted probability value and save it as a grayscale image, where black represents the background and white represents the road.

[0106] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0107] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A remote sensing image road segmentation method based on diffusion pre-training, characterized in that: Includes the following contents and steps: Step 1: construct a pre-training dataset and a fine-tuning dataset, where the pre-training dataset only contains remote sensing images of the complete road map, and the fine-tuning dataset contains remote sensing images and corresponding road annotations; Step 2: Build a segmentation backbone model based on the UNet network. The encoding part includes a ConvBlock and 4 encoding blocks, and the decoding part includes 4 decoding blocks and a prediction head. After the remote sensing image passes through ConvBlock, 4 encoding blocks, 4 decoding blocks and a prediction head, the road segmentation picture is output; Step 3: Based on the denoising diffusion probability model, add noise to the remote sensing image in the pre-training data set to obtain a noisy remote sensing image, input it into the segmentation backbone model, and output the reconstructed image by the fourth decoding block. The goal is to minimize the value of the reconstructed image and the original remote sensing image, and complete the pre-training process of the segmentation backbone model. Step 4: Load the pre-trained segmentation backbone model weights and complete segmentation backbone model fine-tuning based on the fine-tuning dataset; Step 5: Input the remote sensing image to be tested into the fine-tuned segmentation backbone model to complete the actual remote sensing image road segmentation task.

2. The remote sensing image road segmentation method based on diffusion pre-training according to claim 1 is characterized in that: Step 1: Build a pre-training dataset and a fine-tuning dataset. The specific method is as follows: Collect images from multiple public datasets, crop and remove the unlabeled images without complete roadmaps, and build a pre-training dataset. Collect images from multiple public datasets, select data with images and corresponding road annotations, and build a fine-tuning dataset.

3. The remote sensing image road segmentation method based on diffusion pre-training according to claim 1 is characterized in that: Step 2: Build a segmentation backbone model based on the UNet network. The encoding part includes ConvBlock and 4 encoding blocks, and the decoding part includes 4 decoding blocks and a prediction head. After the remote sensing image passes through ConvBlock, 4 encoding blocks, 4 decoding blocks and a prediction head, the road segmentation picture is output. The specific method is as follows: Given a remote sensing image The input is transformed into Then it passes through the first encoding block and becomes a feature map Similarly, after the second, third, and fourth encoding blocks, feature maps X2, X3, and X4 are obtained. In this process, the number of channels gradually increases and the size of the feature map gradually decreases. Correspondingly, feature map X4 is concatenated with feature map X3 and then input into the fourth decoding block to obtain feature map Y4. Y4 is concatenated with X2 and then input into the third decoding block to obtain feature map Y3. Similarly, the outputs of the remaining decoding blocks are Y2 and Y1. Finally, feature map Y1 passes through 1×1Conv to output a road segmentation picture. The encoding block consists of a PatchMerging and a SSFBlock in series, and the decoding block consists of a PatchExpand and a SSFBlock in series; a. Spectral-Spatial Fusion Block (SSFBlock) PatchMerging for the feature map X of the l-1th layer l-1 After the operation, the feature map is obtained In the lth SSFBlock, the feature map After layer normalization (LayerNorm) and spectral convolution operations, the feature map obtained With feature map Add together to get the feature map After that, the feature map After layer normalization and MLP operation, we get The output and feature map obtained Add together to get the output X of this layer l ; b. Spectral Convolution (SP-Conv) Feature map of layer l After normalization, the feature map is obtained In the spectral convolution module, the feature map Split into two tensors by channel dimension where X' l Directly input the adaptive feature selection and enhancement (AFSE) layer to obtain the feature map X' l 'After DCT transformation, input into AFSE layer, and then perform inverse transformation to obtain feature map at last, and Splice along the channel to get the spectral convolution output c. AFSE layer The AFSE layer consists of a detail enhancement branch (DEB), a meta-attention branch (MAB), and a linear transformation branch (LTB) in parallel; DEB is a depth-wise separable convolution that enhances local detail features by sequentially combining depth-wise convolution and point-wise convolution; LTB is a parallel branch of DEB that performs linear transformation in the channel dimension through 1×1 convolution; for the feature map X' l , the basic structure composed of DEB and LTB is described as follows: in, represents 1×1 convolution, Represents a 3×3 depth convolution; MAB is a meta-attention module that aims to refine semantic information through a squeeze-excitation-like attention mechanism to generate semantic-aware weights for reweighing and recalibrating features from DEB and LTB branches; for a given feature map The MAB process is described as follows: Among them, GAP(·) and GMP(·) represent the global average pooling and global maximum pooling operations respectively, and FC m (·) indicates a full connection operation; Finally, the output of the AFSE layer is expressed as:

4. The remote sensing image road segmentation method based on diffusion pre-training according to claim 1 is characterized in that: Step 3: Based on the denoising diffusion probability model, add noise to the remote sensing image in the pre-training data set to obtain a noisy remote sensing image, input it into the segmentation backbone model, and output the reconstructed image by the fourth decoding block. The goal is to minimize the value of the reconstructed image and the original remote sensing image to complete the segmentation backbone model pre-training process. The specific method is: The clear remote sensing image x0 is first transformed into the frequency domain by DCT to obtain the spectrum diagram Then gradually add noise to obtain the spectrum with noise Then, the spectrum with noise is analyzed separately. Perform iDCT to obtain the noisy remote sensing image x1,…,x t-1 ,x t ,…x T , and then input them into the segmentation backbone model respectively, and the reconstructed image is output by the fourth decoding block. The reconstructed image is The goal is to minimize the difference between the reconstructed image and the original image x0, and complete the pre-training process of the segmentation backbone model; When pre-training the segmentation backbone model, the mean square error (MSE) is selected as the loss function to measure the difference between the restored image and the real image. During the optimization process, the AdamW optimizer is used, the initial learning rate is set to 1e-4, and a dynamic learning rate scheduling strategy is adopted to ensure the stability of the training process and accelerate convergence. The maximum diffusion steps of training are 1000 and the batch size is 128.

5. The remote sensing image road segmentation method based on diffusion pre-training according to claim 1, characterized in that: Step 4: Load the pre-trained segmentation backbone model weights and complete the segmentation backbone model fine-tuning based on the fine-tuning dataset. The specific method is as follows: Fine-tune the segmentation backbone model given in step 2 and the fine-tuning dataset produced in step 1, load the pre-trained segmentation backbone model weights trained in step 3, and apply them to the annotated fine-tuning dataset for fine-tuning. During fine-tuning, the prediction head of the model is initialized to 0, and the other parts keep the original pre-trained weights unchanged; During fine-tuning, the mean square error (MSE) was used as the loss function, the AdamW optimizer was used for parameter update, and the batch size was set to 8.

6. The remote sensing image road segmentation method based on diffusion pre-training according to claim 1, characterized in that: Step 5: Input the remote sensing image to be tested into the fine-tuned segmentation backbone model to complete the actual remote sensing image road segmentation task. The specific method is as follows: First, the input remote sensing image is cropped, and then the processed image is input into the segmentation backbone model for inference. Through the Sigmoid activation function, the probability value of each pixel belonging to the road area is output. Finally, the final road segmentation result is generated according to the predicted probability value and saved as a grayscale image, where black represents the background and white represents the road.

7. A remote sensing image segmentation system based on diffusion pre-training, characterized in that: Implement the remote sensing image segmentation method based on diffusion pre-training described in any one of claims 1 to 6 to achieve remote sensing image segmentation based on diffusion pre-training.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the remote sensing image segmentation method based on diffusion pre-training according to any one of claims 1 to 6 is implemented to achieve remote sensing image segmentation based on diffusion pre-training.

9. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the remote sensing image segmentation method based on diffusion pre-training according to any one of claims 1 to 6 is implemented to achieve remote sensing image segmentation based on diffusion pre-training.

Citation Information

Patent Citations

  • Remote sensing image cloud detection method based on spectral feature guidance and spatial-spectral convolution

    CN117079135A

  • Remote sensing image semantic segmentation method fusing diffusion model and converter

    CN118691826A

  • Boundary-optimized remote sensing image semantic segmentation method and apparatus, and device and medium

    WO2023077816A1

Cited By

  • Intelligent fish tank monitoring and management system based on edge AI and cloud edge cooperation

    CN121259542A