Remote sensing image segmentation method based on improved DeepLabV3 +

Through the improved DeepLabV3+ model and the lightweight MobileNetV2 network, combined with channel attention and spatial attention mechanism, the problem of remote sensing image segmentation method in the prior art is solved, and the segmentation accuracy and stability of complex geographic scenes is achieved.

CN120107600APending Publication Date: 2025-06-06CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510316487.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing remote sensing image semantic segmentation method based on deep learning is not stable and accurate enough when dealing with factors such as lighting changes, shadows and noise, and is not effective for complex land scenes and multi-category segmentation tasks.

Method used

The improved DeepLabV3+ model is adopted, and a lightweight MobileNetV2 network is used as the backbone network, and channel attention and spatial attention mechanisms are introduced during the encoding and decoding stages to enhance feature representation and boundary segmentation accuracy.

Benefits of technology

The accuracy and stability of remote sensing image segmentation are improved, especially when dealing with complex geographic scenes and multi-category segmentation tasks, the model's utilization efficiency of multi-scale features and the recognition ability of geographic boundaries is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107600A_ABST
    Figure CN120107600A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image segmentation method based on improved DeepLabV3 +, and belongs to the field of images, and the method comprises the steps: S1, obtaining different land images through the aerial photographing of an unmanned aerial vehicle, and obtaining a data set; s2, performing semantic annotation and segmentation on the image in the data set, and dividing the whole data set into a training set, a test set and a verification set; s3, constructing a segmentation model based on the improved DeepLabV < 3 + >; s4, training the improved DeepLabv3 + network model by adopting the training set; s5, obtaining an optimal segmentation model; s6, inputting the test set image into the trained improved DeepLabv3 + network model to obtain an optimal segmentation image; the method is based on the improved DeepLabv3 + network model, the training parameter quantity is small, the precision is high, the edge segmentation is finer, and the hole problem can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of high-resolution remote sensing image land cover segmentation, and in particular to a remote sensing image segmentation method based on improved DeepLabV3+. Background Art

[0002] The development of remote sensing technology enables us to obtain a large amount of high-resolution image data. Remote sensing images are usually collected through remote sensing satellites, aerial photography aircraft, drones and other platforms. These platforms are equipped with various types of sensors that can capture electromagnetic radiation on the earth's surface and convert them into digital images. High-resolution remote sensing images have great application value in ecological environment monitoring, agricultural production, urban planning and land use management. As an important task in the field of computer vision, remote sensing image semantic segmentation aims to assign each pixel in the image to a predefined semantic category. This task is of great significance for understanding image content, object recognition, geographic information analysis and other aspects. Due to the characteristics of high spatial resolution, rich spectral information and large data volume of remote sensing images, it has become an urgent need to quickly and accurately perform automatic object recognition and information extraction on remote sensing images.

[0003] With the rapid advancement of computing performance, deep learning methods have begun to attract people's attention. Researchers have found that deep neural networks can automatically learn features in large data sets and extract deep semantic features of images. Compared with traditional machine learning methods, deep learning methods have stronger representation capabilities, better generalization capabilities, and higher accuracy in remote sensing image semantic segmentation tasks. However, remote sensing image semantic segmentation methods based on deep learning also have some limitations to a certain extent. First, they are sensitive to factors such as lighting changes, shadows, and noise, which may result in unstable and inaccurate segmentation results. Secondly, they are not effective for complex ground scenes and multi-category segmentation tasks. Summary of the invention

[0004] In order to solve the above problems, based on this, this patent proposes a remote sensing image segmentation method based on improved DeepLabV3+. The present invention provides the following technical solutions: S1. Use drone aerial photography technology to collect remote sensing images of different lands; S2. Perform preprocessing operations such as feature extraction, filtering, and denoising on the collected remote sensing images; S3. Build a remote sensing image segmentation model based on improved DeepLabV3+ in the model training stage; S4. Set different training parameters, train the model, and obtain the optimal model; S5. Input the test set remote sensing image into the segmentation model for testing; S6. Obtain the best segmented image.

[0005] According to S2, the collected images are pre-segmented to fit the subsequent model, divided into training set and test set, and labeled.

[0006] According to S3, the remote sensing image segmentation model based on the improved DeepLabV3+ is divided into the encoding stage and the decoding stage. The encoding stage extracts multi-scale semantic features to capture the global context of the image. The decoding stage fuses shallow details with deep semantics, restores spatial resolution, and improves boundary segmentation accuracy.

[0007] In view of the poor performance in complex terrain scenes and multi-category segmentation tasks, an improved lightweight MobileNetV2 network is used as the backbone network. The feature map resolution is gradually reduced through downsampling operations (such as convolution step size and pooling), while the number of channels is increased.

[0008] The deep features output by ASPP are upsampled by 4 times (e.g., from 1 / 16 to 1 / 4 of the original image size) through bilinear interpolation or transposed convolution. The upsampled high-level features are combined with shallow features in semantics and details and spliced ​​along the channel dimension.

[0009] Introducing spatial attention and channel attention in the bottleneck layer of the MobileNetV2 network to enhance the effectiveness of feature representation. By learning pixel-level attention weights, background noise is suppressed and target areas (such as farmland and road boundaries) are highlighted.

[0010] Adding spatial attention after the deep separable convolution in the bottleneck layer can further refine the spatial distribution of the feature map and make up for the problem of detail loss caused by the limited receptive field of deep convolution. The spatial attention mechanism captures local and global spatial dependencies through convolution or global pooling operations to prevent small objects (such as telephone poles and small buildings) from being ignored in deep convolution.

[0011] Adding channel attention after the extended convolution of the bottleneck layer can optimize the correlation between feature channels and avoid the dilution of semantic information caused by point-by-point convolution (1×1Conv). While increasing the amount of computation by a small amount, the model's efficiency in utilizing multi-scale features is significantly improved.

[0012] The calculation formula of channel attention is: in, is the feature map, is the channel attention weight matrix, , To generate a 2D feature map using two pooling methods in the channel dimension, is the activation function, , is the weight.

[0013] The calculation formula for spatial attention is: in, is the feature map, is the spatial attention weight matrix, , To generate a 2D feature map using two pooling methods in the channel dimension, is the activation function, , is the weight, To perform a 7x7 convolution operation.

[0014] Use 1x1 convolution to increase nonlinear expression capabilities through linear activation (ReLU6), providing a richer feature space for subsequent deep convolution.

[0015] Use 3x3 convolution to perform spatial feature extraction on the expanded channels.

[0016] Use residuals to skip the attention module and directly connect the input and output to alleviate the gradient vanishing problem and keep shallow details.

[0017] use The activation function is due to The idea of ​​random regularization is added to improve the accuracy of the network, improve the fitting ability of the model, and achieve better performance in some deep neural networks.

[0018] Set different parameters to debug the model, and obtain the optimal parameter segmentation model by comparing different models. Input the divided test set images into the model to obtain the final segmented image.

[0019] Compared with the prior art, the present invention has the following beneficial effects: Deep learning models require a lot of storage space and computing resources in semantic segmentation, which has always been a pain point in the field of semantic segmentation. The present invention uses a lightweight improved MobileNetV2 network as the backbone network. This flexibility enables the model to adapt to different application requirements and computing resources, balancing performance and efficiency.

[0020] The present invention uses channel attention and spatial attention to further improve the fusion of multi-scale information, suppress background noise, and highlight the boundaries and key bands of objects. The attention mechanism allows the model to focus on certain specific areas or features when processing images, thereby ignoring irrelevant information. This mechanism can be implemented through spatial attention modules and channel attention modules, which enhances the network's ability to recognize key features and helps the model better understand the complex content of remote sensing images.

[0021] This model uses the ISPRS Vaihingen dataset and the Potsdam dataset as data sources, which improves the data quality and ensures the high accuracy of the model in the semantic segmentation task. It enriches the number of samples in each category in the dataset to make it more balanced, reduces the model's dependence on high-frequency categories, and enhances the model's ability to recognize objects of different sizes. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic diagram of the steps of a remote sensing image segmentation method based on improved DeepLabV3+ of the present invention; Figure 2 It is a DeepLabV3+ network structure diagram of a remote sensing image segmentation method based on improved DeepLabV3+ of the present invention; Figure 3 It is a network structure diagram of an improved MobileNetV2 of a remote sensing image segmentation method based on an improved DeepLabV3+ of the present invention; Figure 4 A bottleneck layer structure diagram of a remote sensing image segmentation method based on improved DeepLabV3+ of the present invention; Figure 5 The present invention provides a channel attention and spatial attention mechanism structure diagram of a remote sensing image segmentation method based on an improved DeepLabV3+. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0024] The embodiments of the present invention are further described in detail below with reference to the accompanying drawings.

[0025] like Figure 1 As shown, the present invention discloses a remote sensing image segmentation method based on improved DeepLabV3+, the method comprising the following steps: S1. Use drone aerial photography technology to collect remote sensing images of different lands; S2. Perform preprocessing operations such as feature extraction, filtering, and denoising on the collected remote sensing images; S3. Build a remote sensing image segmentation model based on improved DeepLabV3+ in the model training stage; S4. Set different training parameters, train the model, and obtain the optimal model; S5. Input the test set remote sensing image into the segmentation model for testing; S6. Obtain the best segmented image.

[0026] According to step S1, the remote sensing image ISPRS Vaihingen dataset and Potsdam dataset used in the present invention.

[0027] According to step S2, the remote sensing images in the data set are subjected to image processing such as feature extraction and Gaussian noise, and the images are segmented into adaptive sizes and divided into a training set, a test set, and a validation set according to a 7:2:1 ratio.

[0028] According to step S3, the improved DeepLabV3+ model is divided into an encoding stage and a decoding stage.

[0029] In the encoding stage, the improved MobileNetV2 network is used as the backbone network, which cooperates with the multi-scale hole convolution of ASPP to improve the recognition ability of objects of different sizes.

[0030] MobileNetV2 is a lightweight convolutional neural network designed for mobile and embedded systems. It reduces the number of model parameters and computational complexity through deep separable convolution technology, and enhances the learning ability of the network by introducing residual connections. The network structure of MobileNetV2 includes a pre-trained convolutional base and a specific classifier layer that can be fine-tuned for different tasks.

[0031] In order to enhance the network's ability to recognize key features, channel attention mechanism and spatial attention mechanism are introduced in MobileNetV2.

[0032] The channel attention mechanism performs global average pooling and global maximum pooling on the shallow feature information in the channel dimension, thereby obtaining two single-dimensional feature vectors. The feature vectors are weighted by the fully connected layer, and finally the feature information in the channel domain is enhanced. The specific expression is in, is the feature map, is the channel attention weight matrix, , To generate a 2D feature map using two pooling methods in the channel dimension, is the activation function, , is the weight.

[0033] The spatial attention mechanism obtains a 2D feature map through average pooling and maximum pooling, and then distributes the weights through convolution, ultimately achieving feature information enhancement of changes in remote sensing images in the spatial dimension. The specific expression is: in, is the feature map, is the spatial attention weight matrix, , To generate a 2D feature map using two pooling methods in the channel dimension, is the activation function, , is the weight, To perform a 7x7 convolution operation.

[0034] By using 1×1 point-by-point convolution in the bottleneck layer of the improved MobileNetV2 network, the number of channels of the feature map can be effectively reduced, thereby significantly reducing the computational complexity and improving the efficiency of training and reasoning, which is particularly suitable for processing large-scale image data sets; secondly, point-by-point convolution provides a flexible parameter control mechanism. By adjusting the number of output channels, the depth of the feature map can be precisely controlled, making the model easier to train and less prone to overfitting; finally, although the point-by-point convolution itself is linear, a nonlinear activation function is usually added after it, thereby introducing nonlinear mapping, which helps to learn more complex image features and semantic information and improve the model's perception ability.

[0035] Using 3×3 depthwise separable convolution in the bottleneck layer can decompose the traditional convolution into depthwise convolution in the spatial dimension (only processing single-channel features) and point-by-point convolution in the channel dimension (adjusting the number of channels). While maintaining the feature extraction capability, it greatly reduces the amount of computation and parameters (compared to traditional convolution, it reduces the amount of parameters and computation by about 89%), thus achieving a lightweight model. Its efficient structure is particularly suitable for scenarios such as remote sensing image segmentation, which can not only retain the edge details of objects (such as roads and building contours), but also enhance semantic expression through channel fusion.

[0036] use The activation function can better model the nonlinear relationship of input data, improve the model's fitting ability, and achieve better performance in some deep neural networks. Its smooth characteristics help reduce the gradient explosion or gradient vanishing problems that occur during network training, which in turn helps improve the training stability and convergence speed of the model.

[0037] The present invention uses three commonly used indicators for remote sensing image scene classification to evaluate performance, namely: loss rate, which is an indicator to measure the difference between the model prediction result and the actual label, and the calculation formula is as follows: in is the total number of samples, is the true label, is the predicted value, Typically this is cross entropy or mean squared error.

[0038] Overall accuracy is used to measure the proportion of samples predicted correctly by the model to the total number of samples. The calculation formula is as follows: Among them, TP is the number of samples correctly predicted by the model as positive, TN is the number of samples correctly predicted by the model as negative, FP is the number of samples incorrectly predicted by the model as positive (i.e., "false positives"), and FN is the number of samples incorrectly predicted by the model as negative (i.e., "false negatives").

[0039] Precision is used to predict the correct proportion of samples in the positive category. The calculation formula is as follows: Among them, TP is the number of samples correctly predicted by the model as positive, and FP is the number of samples incorrectly predicted by the model as positive (i.e., "false positives").

[0040] All experiments in this paper were completed under the PyTorch1.3.1 deep learning framework, and the execution environment was a 64-bit Windows 10 operating system, equipped with an NVIDIA GeForce RTX 2080 graphics card, 8G / B memory, and accelerated by CUDA11.8. During the training process, the model was iterated 300 times.

[0041] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0042] The above description is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest range consistent with the principles and novel features applied for herein.

Claims

1. A remote sensing image segmentation method based on improved DeepLabV3+, characterized in that: S1. Collect remote sensing images; S2. Perform data preprocessing; S3. Build a segmentation model based on improved DeepLabV3+; S4. Perform model training with different parameters; S5. Obtain the optimal segmentation model; S6. Test the remote sensing image segmentation image to obtain a segmented image.

2. The remote sensing image segmentation method based on improved DeepLabV3+ according to claim 1, characterized in that: In S1, the collected remote sensing images are divided into image training, label training set, image test set and label test set.

3. The remote sensing image segmentation method based on improved DeepLabV3+ according to claim 1, characterized in that: In S2, preprocessing operations such as denoising and filtering are performed on the remote sensing image.

4. The remote sensing image segmentation method based on improved DeepLabV3+ according to claim 1, characterized in that: The segmentation model based on improved DeepLabV3+ in S3 is divided into an encoding stage and a decoding stage.

5. The remote sensing image segmentation method based on improved DeepLabV3+ according to claim 4, characterized in that: The encoding stage consists of an improved MobileNetV2 network, an ASPP network and a 1x1 convolution module.

6. The remote sensing image segmentation method based on improved DeepLabV3+ according to claim 4, characterized in that: The decoding stage consists of 1x1 convolution, 3x3 convolution, connection layer, 4x upsampling and 2x upsampling modules.

7. The remote sensing image segmentation method based on improved DeepLabV3+ according to claim 5, characterized in that: The improved MobileNetV2 network consists of a convolutional layer, a 1x1 convolution, an average pooling layer, and 17 bottleneck layers.

8. The remote sensing image segmentation method based on improved DeepLabV3+ according to claim 7, characterized in that: The bottleneck layer consists of an expansion layer, a depth convolution, a stationary point convolution, an inverted residual module, a channel attention module and a spatial attention module.

9. The remote sensing image segmentation method based on improved DeepLabV3+ according to claim 8, characterized in that: The channel attention and spatial attention are composed of global average pooling, global maximum pooling, convolution and activation function modules.

10. The remote sensing image segmentation method based on improved DeepLabV3+ according to claim 1, characterized in that: The training process is divided into building a segmentation model, performing parameter optimization, and performing image testing to obtain the optimal segmentation result.