Lightweight Remote Sensing Image Semantic Segmentation Method Based on Improved Deeplabv3+

Through the improved DeepLab v3+ network model, combining lightweight networks and specific modules, the complexity and noise sensitivity problems of remote sensing image segmentation are solved, and the semantic segmentation effect with high precision and robustness is achieved.

CN115984850BActive Publication Date: 2025-07-25HEFEI ZHENGZE LINGJUN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310116702.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-15
Publication Date
2025-07-25
Estimated Expiration
2043-02-15

AI Technical Summary

Technical Problem

The existing remote sensing image segmentation method requires manual threshold determination, the process is complex, the noise is sensitive, the segmentation effect and generalization ability are poor, and it is difficult to achieve high accuracy in complex environments.

Method used

The improved DeepLab v3+ network model is adopted, and the lightweight network MobileNet v2 is used as the backbone network, combining HDC module, strip pooling module and standardized attention mechanism NAM, and the model is optimized through data enhancement and cross-entropy loss function to improve segmentation accuracy and robustness.

Benefits of technology

It realizes small training parameters, high accuracy, and delicate edge segmentation. It is suitable for semantic segmentation in complex environments, improves the average intersect ratio and average pixel accuracy, and is suitable for a variety of segmentation objects such as fields, buildings, and waters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984850B_ABST
    Figure CN115984850B_ABST
Patent Text Reader

Abstract

The present invention relates to a lightweight remote sensing image semantic segmentation method based on improved DeepLabv3+, including: obtaining images in different scenarios through remote sensing satellites to obtain a data set; selecting segmentation objects in the data set, performing semantic annotation and segmentation on the segmentation objects, performing image enhancement, and dividing the entire data set into a training set, a test set, and a validation set after image enhancement; constructing an improved DeepLabv3+ network model, and training the improved DeepLabv3+ network model using the training set; inputting the test images in the test set into the trained improved DeepLabv3+ network model, selecting the segmentation objects as fields, building complexes, or water areas, and saving the resulting images of semantic segmentation. The present invention is based on an improved DeepLabv3+ network model, with a small number of training parameters, high accuracy, more delicate edge segmentation, and effectively improves the hole problem; the system of the present invention targets multiple segmentation objects such as fields, building complexes, and water areas, is convenient for use in different scenarios, and is intelligent and convenient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image semantic segmentation, and in particular to a lightweight remote sensing image semantic segmentation method based on improved Deeplabv3+. Background Art

[0002] With the continuous improvement of the spectral resolution of remote sensing sensors, people's understanding of the spectral attributes and characteristics of ground objects has also deepened continuously. Many ground object characteristics hidden in narrow spectral ranges have been gradually discovered, greatly accelerating the development of remote sensing technology and making hyperspectral remote sensing one of the important research directions in the field of remote sensing in the 21st century. Compared with traditional measurement methods, it has the advantages of high temporal resolution, high spatial resolution, and low labor cost, and can collect and analyze geographical data more accurately and dynamically. Therefore, it is widely used in many fields such as urban planning, land and crop monitoring, and weather forecasting.

[0003] Among them, object-based segmentation is a key technology in remote sensing image processing. For object segmentation, some traditional segmentation methods require manual determination of thresholds, and the extraction process is relatively complex. Moreover, they are sensitive to noise and outliers in the image, and the segmentation effect and generalization ability are poor. Facing remote sensing images in complex environments, traditional image segmentation technologies are also difficult to achieve high accuracy and good effects, and cannot meet the requirements in practical applications. Summary of the Invention

[0004] To solve the defect of poor segmentation effect, the purpose of the present invention is to provide a lightweight remote sensing image semantic segmentation method based on improved Deeplabv3+ with a small number of training parameters, high accuracy, and more delicate edge segmentation.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A lightweight remote sensing image semantic segmentation method based on improved Deeplabv3+, the method includes the following steps in sequence:

[0006] (1) Data acquisition: Obtain images in different scenarios through remote sensing satellites and convert them into three-channel images in jpg format to obtain a data set;

[0007] (2) Data preprocessing: Select segmentation objects in the data set, perform semantic annotation and segmentation on the segmentation objects, perform image enhancement, and divide the entire data set into a training set, a test set, and a validation set after image enhancement;

[0008] (3) Creation and training of the network model: Construct an improved DeepLab v3+ network model, and use the training set to train the improved DeepLab v3+ network model to obtain the trained improved DeepLab v3+ network model and its parameters;

[0009] (4) Testing phase: Input the test images in the test set into the trained improved DeepLab v3+ network model. Select the segmentation objects as fields, building complexes or waters, and save the resulting images of semantic segmentation.

[0010] The specific content of step (1) is as follows: Collect satellite remote sensing images of different scenes including building complexes, farmlands, and waters respectively. Convert the multi-channel hyperspectral images into three-channel jpg format images to obtain a data set for feature extraction.

[0011] The specific content of step (2) includes the following steps:

[0012] (2a) Select the segmentation objects, use the image annotation tool Lableme to mark the building complexes, farmlands, and waters in the images. After marking, perform segmentation and divide the images into images of 500×500.

[0013] (2b) Perform data augmentation on the segmented images, randomly perform horizontal or vertical flipping, add random noise, expand the data set, and divide the expanded data set into a training set, a test set, and a validation set in the ratio of 7:2:1 in sequence.

[0014] The specific content of step (3) includes the following steps:

[0015] (3a) Build an improved DeepLab v3+ network model, and use the lightweight network MobileNet v2 as the backbone network.

[0016] (3b) In the Encoder of the coding area, continuously use three dilated convolutions with dilation rates of 1, 2, and 3 as the HDC module to replace the original convolutions in the Atrous Spatial Pyramid Pooling (ASPP). To ensure that the receptive field remains unchanged, use one HDC module to replace the convolution with a dilation rate of 6, use two HDC modules to replace the convolution with a dilation rate of 12, and use three HDC modules to replace the convolution with a dilation rate of 18. When performing convolution, cover the square area of the low-level feature layer to improve the hole problem caused by the grid effect.

[0017] The HDC module is a dilated convolution that follows the HDC principle. Define the formula for the maximum distance between two non-zero elements:

[0018]

[0019] where M i is the maximum distance between two non-zero elements in the i-th layer, and r i is the dilation rate of the i-th layer. To avoid the loss caused by the grid effect, it is necessary to satisfy that the maximum distance M i between two non-zero elements in the i-th layer is ≤ the convolution kernel size K;

[0020] (3c) In the Atrous Spatial Pyramid Pooling with Strip Pooling (ASPP), the original global average pooling module is replaced by a strip pooling module to build the dependencies between channels through vertical pooling and horizontal pooling respectively, and information is collected from different spatial dimensions;

[0021] During vertical pooling, the pixel values of each column in the feature map x are added and then averaged, and the output y after vertical strip pooling v ∈R W is a row vector of 1xW:

[0022]

[0023] During horizontal pooling, the pixel values of each row in the feature map x are added and then averaged, and the output y after horizontal strip merging h ∈R H is a column vector of H×1:

[0024]

[0025] where the input feature x ∈ R C×H×W is the input tensor, C represents the number of channels, H represents the height of the feature map, and W represents the width of the feature map; the output y h ∈R H is a column vector of H×1, and the output y v ∈R W is a row vector of 1xW, and i and j respectively represent the i-th row and the j-th column of the feature map;

[0026] To obtain the output z ∈ R C×H×W containing more useful global priors, an expansion operation is performed to obtain y h ∈R C×H and y v ∈R C×W , and they are combined to obtain y ∈ R C×H×W :

[0027]

[0028] The output z is:

[0029] z = Scale(x, σ(f(y))) (5)

[0030] where Scale() represents element-wise multiplication, σ is the Sigmoid function, and f is a 1×1 convolution;

[0031] (3d) In the Decoder of the decoding area, the shallow features of the 4th and 7th layers are extracted from the backbone network, and the normalization-based attention mechanism NAM is applied. Then, a ResNet50 module is constructed, using a convolutional block that first reduces the dimension and then increases the dimension, and replacing the 3×3 convolution with a dilated convolution with a dilation rate of 4 to enrich the detailed features;

[0032] The normalization-based attention mechanism NAM uses the scale factor of batch normalization to represent the importance of channels through sparse weight penalties and standard deviations. The normalization-based attention mechanism NAM includes a channel attention sub-module and a spatial attention sub-module. The channel attention mechanism is as follows:

[0033]

[0034] Among them, represents the mean of the mini-batch B; σ B represents the standard deviation of B; γ and β represent the scaling factor and displacement respectively; ε is a small number to avoid division by zero; Bin represents the input value of B, and Bout represents the output value of B

[0035] Normalize the information in the channel dimension of the input feature map, and the final output feature obtained after applying the weight is as follows:

[0036] M c = sigmoid(W γ (BN(F1))) (7)

[0037] Among them, γ is the scaling factor, M C represents the output feature, W γ represents the weight of this channel, and F1 represents the input feature map;

[0038] The spatial attention sub-module uses the same normalization method for each pixel in the space of the input feature map, and the final output feature is:

[0039] M s = sigmoid(W λ (BN s (F2))) (8)

[0040] Among them, λ is the scaling factor of each channel, M s represents the output feature, W λ represents the weight of this channel, and F2 represents the input feature map;

[0041] (3e) Select cross - entropy as the loss of the algorithm, and mean Intersection over Union (mIoU) and mean Pixel Accuracy (mPA) as evaluation metrics. Evaluate the training effect of the network model from two aspects: the proportion of correctly predicted pixels in the union of predicted pixels and ground - truth pixels, and the proportion of correct pixels in the total number of pixels.

[0042] The cross - entropy is:

[0043]

[0044] where y i is the ground - truth value of a certain pixel, which is 0 or 1 in the binary classification task; is the predicted value of a certain pixel; n is the sample size for each loss calculation;

[0045] The calculation formulas for mean Intersection over Union (mIoU) and mean Pixel Accuracy (mPA) are as follows:

[0046]

[0047]

[0048] where TP represents True Positive, that is, the model predicts a positive example and it is actually a positive example; FP represents False Positive, the model predicts a positive example but it is actually a negative example; FN represents False Negative, the model predicts a negative example but it is actually a positive example.

[0049] Step (4) specifically includes the following steps:

[0050] (4a) Save the pre - trained model parameters for different segmentation objects;

[0051] (4b) Input the test images in the test set into the trained improved DeepLab v3+ network model. Select the segmentation objects as fields, building groups or waters. After outputting the corresponding segmentation results, save the result images of semantic segmentation.

[0052] From the above technical solutions, the beneficial effects of the present invention are as follows: First, based on the improved DeepLab v3+ network model, the present invention has a small number of training parameters, high accuracy, more delicate edge segmentation, and effectively improves the hole problem; Second, the system of the present invention is applicable to multiple segmentation objects such as fields, building groups, and waters, which is convenient for use in different scenarios, intelligent and convenient; Third, the method of the present invention has good robustness and can be applied to semantic segmentation in complex environments. In the actual segmentation of binary classification and multi - classification, the mean Intersection over Union (mIoU) and mean Pixel Accuracy (mPA) are both improved compared with the original model. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is the flowchart of the method of the present invention;

[0054] Figure 2 This is the comparison diagram of the original DeepLab v3+ network structure in the present invention;

[0055] Figure 3 This is the improved DeepLab v3+ network structure diagram in the present invention;

[0056] Figure 4 This is the schematic diagram of the HDC module in the present invention;

[0057] Figure 5 This is the schematic diagram of the strip pooling module in the present invention;

[0058] Figure 6 This is the schematic diagram of the improved ResNet50 module in the present invention;

[0059] Figure 7 This is the channel attention mechanism diagram in the present invention;

[0060] Figure 8 This is the spatial attention mechanism diagram in the present invention. Detailed implementation manners

[0061] As Figure 1 shown, a lightweight remote sensing image semantic segmentation method based on improved Deeplabv3+ includes the following steps in sequence:

[0062] (1) Data collection: Obtain images in different scenarios through remote sensing satellites, and convert them into three-channel images in jpg format to obtain a data set;

[0063] (2) Data preprocessing: Select segmentation objects in the data set, perform semantic annotation and segmentation on the segmentation objects, perform image enhancement, and divide the entire data set into a training set, a test set, and a validation set after image enhancement;

[0064] (3) Creation and training of the network model: Build an improved DeepLab v3+ network model, and use the training set to train the improved DeepLab v3+ network model to obtain the trained improved DeepLab v3+ network model and its parameters;

[0065] (4) Test stage: Input the test images in the test set into the trained improved DeepLab v3+ network model, select the segmentation objects as fields, building groups, or waters, and save the result images of semantic segmentation.

[0066] The specific content of step (1) is: Collect satellite remote sensing images of different scenes including building groups, farmlands, and waters respectively, and convert the multi-channel hyperspectral images into three-channel jpg format images to obtain a data set for feature extraction.

[0067] Step (2) specifically includes the following steps:

[0068] (2a) Select the segmentation object, use the image annotation tool Lableme to mark the building complex, farmland, and water area in the image, and perform segmentation after marking to divide the image into images of 500×500;

[0069] (2b) Perform data augmentation on the segmented image, randomly perform horizontal or vertical flipping, add random noise, expand the dataset, and divide the expanded dataset into a training set, a test set, and a validation set in the ratio of 7:2:1 in sequence.

[0070] Step (3) specifically includes the following steps:

[0071] (3a) Build an improved DeepLab v3+ network model, and use the lightweight network MobileNet v2 as the backbone network;

[0072] (3b) In the Encoder of the coding area, continuously use three dilated convolutions with dilation coefficients of 1, 2, and 3 as the HDC module to replace the original convolution in the Atrous Spatial Pyramid Pooling (ASPP). To ensure that the receptive field remains unchanged, use one HDC module to replace the convolution with a dilation rate of 6, use two HDC modules to replace the convolution with a dilation rate of 12, and use three HDC modules to replace the convolution with a dilation rate of 18. When convolving, cover the square area of the low-level feature layer to improve the hole problem caused by the grid effect;

[0073] The HDC module is a dilated convolution that follows the HDC principle. Define the formula for the maximum distance between two non-zero elements:

[0074]

[0075] where M i is the maximum distance between two non-zero elements in the i-th layer, and r i is the dilation coefficient of the i-th layer. To avoid the loss caused by the grid effect, it is necessary to satisfy that the maximum distance M i between two non-zero elements in the i-th layer ≤ the convolution kernel size K;

[0076] (3c) In the Atrous Spatial Pyramid Pooling (ASPP), use a strip pooling module to replace the original global average pooling module, and respectively construct the dependence between channels through vertical pooling and horizontal pooling to collect information from different spatial dimensions;

[0077] During vertical pooling, add the pixel values of each column in the feature map x and then calculate the mean value. The output y v ∈R W of the vertical strip pooling is a row vector of 1xW:

[0078]

[0079] During horizontal pooling, the pixel values of each row in the feature map x are added and then averaged, and the output y after horizontal strip merging h ∈R H is a column vector of H×1:

[0080]

[0081] Among them, the input feature x∈R C×H×W is the input tensor, C represents the number of channels, H represents the height of the feature map, and W represents the width of the feature map; the output y h ∈R H is a column vector of H×1, and the output y v ∈R W is a row vector of 1xW, and i and j respectively represent the i-th row and the j-th column of the feature map;

[0082] To obtain the output z∈R containing more useful global priors C×H×W , after performing an expansion operation, y is obtained h ∈R C×H and y v ∈R C×W , combined to obtain y∈R C×H×W :

[0083]

[0084] The output z is:

[0085] z = Scale(x, σ(f(y))) (5)

[0086] Among them, Scale() represents element-wise multiplication, σ is the Sigmoid function, and f is a 1×1 convolution;

[0087] (3d) In the decoder region Decoder, the 4th and 7th shallow features are extracted from the backbone network and the normalized attention mechanism NAM is applied. Then, a ResNet50 module is constructed, and a convolutional block with dimensionality reduction first and then dimensionality increase is used, and the 3×3 convolution in it is replaced with a dilated convolution with a dilation rate of 4 to enrich the detailed features;

[0088] The normalized attention mechanism NAM uses the scale factor of batch normalization to represent the importance of channels through sparse weight penalties and standard deviations. The normalized attention mechanism NAM includes a channel attention sub-module and a spatial attention sub-module, as Figure 7 、 8 shown. The channel attention mechanism is:

[0089]

[0090] Among them, represents the mean of small batch B; σ B represents the standard deviation of B; γ and β represent the scaling factor and displacement respectively; e is a small number to avoid a denominator of 0; Bin represents the input value of B, and Bout represents the output value of B

[0091] Normalize the information in the channel dimension of the input feature map, and the final output feature after applying weights is as follows:

[0092] M c = sigmoid(W γ (BN(F1))) (7)

[0093] Among them, γ is the scaling factor, M C represents the output feature, W γ represents the weight of this channel, and F1 represents the input feature map;

[0094] The spatial attention sub-module uses the same normalization method for each pixel in the space of the input feature map, and the final output feature is:

[0095] M s = sigmoid(W λ (BN s (F2))) (8)

[0096] Among them, λ is the scaling factor of each channel, M s represents the output feature, W λ represents the weight of this channel, and F2 represents the input feature map;

[0097] (3e) Select cross-entropy as the loss of the algorithm, and mean intersection over union mIoU and mean pixel accuracy mPA as evaluation metrics. Evaluate the training effect of the network model from two perspectives: the proportion of the correctly predicted pixels in the union of the predicted pixels and the ground truth pixels, and the proportion of the correct pixels in the total pixels; the cross-entropy is:

[0098]

[0099] Among them, y i is the ground truth of a certain pixel, and the ground truth is 0 or 1 in the binary classification task; is the predicted value of a certain pixel; n is the sample size for each loss calculation; the calculation formulas for mean intersection over union mIoU and mean pixel accuracy mPA are as follows:

[0100]

[0101]

[0102] Among them, TP represents a true positive example, that is, the model predicts a positive example and it is actually a positive example; FP represents a false positive example, the model predicts a positive example but it is actually a negative example; FN represents a false negative example, the model predicts a negative example but it is actually a positive example.

[0103] The specific steps of step (4) include the following steps:

[0104] (4a) Save the pre-trained model parameters for different segmentation objects;

[0105] (4b) Input the test images in the test set into the trained improved DeepLab v3+ network model. Select the segmentation objects as fields, building complexes or waters. After outputting the corresponding segmentation results, save the result images of semantic segmentation.

[0106] As Figure 2 、 3 shown, the present invention uses the lightweight network MoibleNet v2 as the backbone network to reduce the number of model parameters; in the enhanced feature extraction network, the standard dilated convolution in the original network is replaced with the HDC module to effectively solve the grid effect, and the traditional spatial average pooling is replaced with the strip pooling module to enrich local details; in the decoder, the network residual module of ResNet50 is added after the fusion of low-level features to further obtain rich target edge feature information at the low level; and the NAM attention mechanism is added to enhance the semantic information of the shallow layer and enhance the ability of the network to capture the feature information of small target objects.

[0107] The following Table 1 shows the structural parameters of MobileNet v2 in the present invention, which is mainly composed of depthwise separable convolutions. Although the accuracy slightly decreases when used as a feature extraction network for training, its inverted residual structure greatly improves the network performance, reduces the number of parameters, and improves the network efficiency. Among them, Input represents the number of input channels for each layer, Operator includes depthwise separable convolution (bottleneck), ordinary convolution (conv2d), average pooling (avgpool), t represents the magnification factor of the 1×1 convolution for dimensionality increase in the inverted residual structure, c is the number of output channels, n represents the number of repetitions of bottleneck, and s represents the stride;

[0108] Table 1

[0109] Layer Input Operator t c n s r 1 3 Conv2d - 32 1 2 1 2 32 bottleneck 1 16 1 1 1 3 16 bottleneck 6 24 2 2 1 4 24 bottleneck 6 32 3 2 1 5 32 bottleneck 6 64 4 1 1 6 64 bottleneck 6 96 3 1 1 7 96 bottleneck 6 160 3 1 4 8 160 bottleneck 6 320 1 1 1

[0110] As Figure 4 shown, the HDC module follows the dilated convolution of the HDC principle. Specifically, it continuously uses 3 dilated convolutions with dilation rates of 1, 2, and 3. As Figure 5 shown, horizontal and vertical strip pooling operations can be used to collect remote contexts from different spatial dimensions. AsFigure 6 As shown, a convolutional block that first reduces the dimension and then increases the dimension is used to further refine the semantic information of the low-level feature map, and the 3×3 convolution in ResNet50 is replaced with a dilated convolution with a dilation rate of 4.

[0111] As shown in Table 2 below, the mIou and mPA parameter values before and after the improvement on three datasets are presented.

[0112] Table 2

[0113]

[0114] In summary, the present invention is based on an improved DeepLab v3+ network model, with a small number of training parameters, high precision, finer edge segmentation, and effectively improving the hole problem; the system of the present invention is applicable to various segmentation objects such as fields, building groups, and waters, facilitating use in different scenarios, and being intelligent and convenient.

Claims

1. A lightweight remote sensing image semantic segmentation method based on improved Deeplabv3+, characterized in that: The method includes the following steps in sequence: (1) Data acquisition: Obtain images in different scenarios through a remote sensing satellite, convert them into three-channel images in jpg format, and obtain a dataset; (2) Data preprocessing: Select segmentation objects in the dataset, perform semantic annotation and segmentation on the segmentation objects, perform image enhancement, and divide the entire dataset into a training set, a test set, and a validation set after image enhancement; (3) Creation and training of the network model: Construct an improved DeepLab v3+ network model, and use the training set to train the improved DeepLab v3+ network model to obtain the trained improved DeepLab v3+ network model and its parameters; (4) Test phase: Input the test images in the test set into the trained improved DeepLab v3+ network model, select the segmentation objects as fields, building complexes, or waters, and save the result images of semantic segmentation; Step (3) specifically includes the following steps: (3a) Construct an improved DeepLab v3+ network model, and use the lightweight network MobileNet v2 as the backbone network; (3b) In the Encoder of the encoding area, continuously use three dilated convolutions with dilation coefficients of 1, 2, and 3 as the HDC module to replace the original convolution in the Atrous Spatial Pyramid Pooling (ASPP). To ensure that the receptive field remains unchanged, use one HDC module to replace the convolution with a dilation rate of 6, use two HDC modules to replace the convolution with a dilation rate of 12, and use three HDC modules to replace the convolution with a dilation rate of 18. When convolving, cover the square area of the low-level feature layer to improve the hole problem caused by the grid effect; (3c) In the Atrous Spatial Pyramid Pooling (ASPP), use a strip pooling module to replace the original global average pooling module, construct the dependence between channels through vertical pooling and horizontal pooling respectively, and collect information from different spatial dimensions; (3d) In the Decoder of the decoding area, extract the 4th and 7th shallow feature layers from the backbone network, apply the Normalization-based Attention Mechanism (NAM), then construct a ResNet50 module, use a convolutional block with dimensionality reduction first and then dimensionality increase, and replace the 3×3 convolution with a dilated convolution with a dilation rate of 4 to enrich the detailed features; (3e) Select cross-entropy as the loss of the algorithm, and the mean Intersection over Union (mIoU) and the mean Pixel Accuracy (mPA) as evaluation metrics to evaluate the training effect of the network model from two perspectives: the proportion of the correctly predicted pixels in the union of the predicted pixels and the true pixels, and the proportion of the correct pixels in the total pixels; 2. The lightweight remote sensing image semantic segmentation method based on improved Deeplabv3+ according to claim 1, characterized in that: Step (1) specifically refers to: Collect satellite remote sensing images of different scenes including building complexes, farmlands, and waters respectively, convert the multi-channel hyperspectral images into three-channel jpg format images, and obtain a dataset for feature extraction; 3. The lightweight remote sensing image semantic segmentation method based on the improved Deeplabv3+ according to claim 1, characterized in that: Step (2) specifically includes the following steps: (2a) Select segmentation objects, use the image annotation tool Lableme to mark the building complexes, farmlands, and waters in the image, perform segmentation after marking, and divide the image into images of 500×500; (2b) Perform data augmentation on the segmented images, randomly flip them horizontally or vertically, add random noise, expand the dataset, and sequentially divide the expanded dataset into a training set, a test set, and a validation set at a ratio of 7:2:

1.

4. The lightweight remote sensing image semantic segmentation method based on the improved Deeplabv3+ according to claim 1, characterized in that: The construction of the improved DeepLab v3+ network model specifically refers to: The HDC module is atrous convolution following the HDC principle, and the formula for the maximum distance between two non-zero elements is defined as: Among them, M i is the maximum distance between two non-zero elements in the i-th layer, and r i is the dilation coefficient of the i-th layer; to avoid losses caused by the grid effect, it is necessary to satisfy that the maximum distance M i between two non-zero elements in the i-th layer is ≤ the convolution kernel size K; During vertical pooling, the pixel values of each column in the feature map x are added up and then averaged, and the output y after vertical strip pooling v ∈R W is a 1xW row vector: During horizontal pooling, the pixel values of each row in the feature map x are added and then averaged, and the output y after horizontal strip merging h ∈R H is a column vector of H×1: Among them, the input feature \(x\in\mathbb{R}\) C×H×W is the input tensor, \(C\) represents the number of channels, \(H\) represents the height of the feature map, and \(W\) represents the width of the feature map; the output \(y\) h \(\in\mathbb{R}\) H is a column vector of \(H\times1\), and the output \(y\) v \(\in\mathbb{R}\) W is a row vector of \(1\times W\), and \(i\) and \(j\) respectively represent the \(i\)-th row and \(j\)-th column of the feature map; To obtain an output \(z\in\mathbb{R}\) containing a more useful global prior C×H×W , after performing an expansion operation, we get \(y\) h \(\in\mathbb{R}\) C×H and \(y\) v \(\in\mathbb{R}\) C ×W , and by combining them, we obtain \(y\in\mathbb{R}\) C×H×W : The output z is: z = Scale(x, σ(f(y))) (5) Among them, Scale() represents element-wise multiplication, σ is the Sigmoid function, and f is a 1×1 convolution; The normalized attention mechanism NAM uses the scale factor of batch normalization to represent the importance of channels through sparse weight penalties and standard deviations. The normalized attention mechanism NAM includes a channel attention sub-module and a spatial attention sub-module. The channel attention mechanism is: Among them, represents the mean of the small batch B; σ Β represents the standard deviation of B; γ and β represent the scaling factor and displacement respectively; ε is a small number to avoid a zero denominator; Bin represents the input value of B, and Bout represents the output value of B Normalize the information in the channel dimension of the input feature map, and the final output feature after applying weights is as follows: M c = sigmoid(W γ (BN(F1))) (7) where γ is the scaling factor, M C represents the output feature, W γ represents the weight of this channel, and F1 represents the input feature map; The spatial attention sub-module uses the same normalization method for each pixel in the space of the input feature map, and the final output feature is: M s = sigmoid(W λ (BN s (F2))) (8) where λ is the scaling factor for each channel, M s represents the output feature, W λ represents the weight of this channel, and F2 represents the input feature map; The cross-entropy is: where y i is the true value of a certain pixel, and in a binary classification task, the true value is 0 or 1; is the predicted value of a certain pixel; n is the sample size for each loss calculation; The calculation formulas for the mean intersection over union mIoU and the mean pixel accuracy mPA are as follows: Among them, TP represents true positives, that is, the model predicts a positive example and it is actually a positive example; FP represents false positives, the model predicts a positive example but it is actually a negative example; FN represents false negatives, the model predicts a negative example but it is actually a positive example.

5. The lightweight remote sensing image semantic segmentation method based on the improved Deeplabv3+ according to claim 1, characterized in that: The specific steps of step (4) include the following steps: (4a) Save the pre-trained model parameters for different segmentation objects; (4b) Input the test images in the test set into the trained improved DeepLab v3+ network model. Select the segmentation objects as fields, building complexes, or waters. After outputting the corresponding segmentation results, save the result images of semantic segmentation.