A Remote Sensing Image Semantic Segmentation Method and System Based on Improved DeepLabV3+ Network
By adopting the R-Drop regularization method in remote sensing image segmentation, the poor generalization performance and overfitting problems of traditional DeepLabV3+ network in complex backgrounds are solved, and higher segmentation accuracy and consistency are achieved.
Patent Information
- Application Number
- CN202210677113.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-15
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-06-15
AI Technical Summary
The existing remote sensing image segmentation method has poor generalization performance in complex backgrounds and is difficult to apply on a large scale. The traditional DeepLabV3+ network has overfitting problems during training, resulting in model inconsistency and performance degradation.
The R-Drop regularization method is used instead of Dropout, and the generalization ability of the model is enhanced by randomly deleting hidden units in each small batch training, combining cross entropy and symmetric KL divergence.
It improves the accuracy and consistency of remote sensing image segmentation, improves pixel accuracy and average interleaving ratio, and enhances the segmentation effect of the model in complex environments.
Smart Images

Figure CN115035418B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image segmentation, and relates to a remote sensing image segmentation method, in particular to a semantic segmentation method and system for remote sensing images based on an improved DeepLabV3+ network. Background Art
[0002] Remote sensing image segmentation algorithms refer to predicting each pixel in a remote sensing image, which is a pixel-level classification algorithm and can be widely applied to many application scenarios such as land planning, environmental monitoring, and disaster assessment, and has great application value. Traditional image segmentation methods mainly use manually designed classifiers based on low-level features of images such as color and texture to segment images, and then label the semantics of the segmented images. Such as pixel-level clustering segmentation methods, pixel-level threshold segmentation methods, pixel-level decision tree classification methods, etc. These algorithms have achieved the requirements of image segmentation to a certain extent, but have high requirements for manually designed feature extractors and poor generalization performance for data sets, and it is difficult to be widely applied to general scenarios with complex backgrounds.
[0003] In recent years, with the rapid development of computer hardware, especially the improvement of GPU computing power, it has greatly promoted the progress of artificial intelligence and also provided great impetus for the development of computer vision. Semantic segmentation is a basic task in computer vision. With the powerful computing power of GPU, image segmentation methods based on deep learning can quickly segment remote sensing images and accurately extract useful information. Currently, this method has become the mainstream method in the field of remote sensing image segmentation. Semantic segmentation architectures usually have different forms and can generally be understood as an encoder-decoder network. Among them, the encoder is usually a pre-trained classification network such as ResNet to extract image features. For the decoder, its main role is reflected in the mapping of discriminable features, which can realize the mapping from semantics to the pixel space, thereby obtaining dense classification, and this is also the function that semantic segmentation needs to achieve.
[0004] DeepLabV3+ is a network model with good performance in semantic segmentation. It mainly extracts feature resolution by arbitrarily controlling the encoder, and can also achieve a balance between efficiency and accuracy. Applying the MobileNetV2 network model to semantic segmentation, a depthwise separable convolutional network is used in the decoding module, which enhances the execution efficiency of the encode-decode in this way. The DeepLabV3+ network adopts the Dropout method to avoid overfitting problems during training. Dropout randomly ignores some neurons during training. During the forward propagation process, the contributions of these ignored neurons to downstream neurons temporarily disappear, and during the backward propagation, these neurons will not have any weight updates. However, this operation will cause the sub-models generated after ignoring some neurons each time to be different, making the trained model have a certain degree of randomness to a certain extent. It is a combination constraint of multiple sub-models, which affects the performance of the network model. Summary of the Invention
[0005] In view of the above problems existing in the prior art, the present invention provides a method and system for semantic segmentation of remote sensing images based on an improved DeepLabV3+ network. The present invention uses the R-Drop regularization method to replace the Dropout method used in the original DeepLabV3+ network. R-Drop further regularizes the model space, surpasses Dropout, and can further improve the generalization ability of the model, thereby enabling effective segmentation of remote sensing urban road images.
[0006] The technical solutions adopted by the present invention are as follows:
[0007] A method for semantic segmentation of remote sensing images based on an improved DeepLabV3+ network, comprising the following steps:
[0008] Step S1. Obtain a remote sensing road data set and perform preprocessing;
[0009] Step S2. Build an improved DeepLabV3+ semantic segmentation network based on the Pytorch environment;
[0010] Step S3. Use the training data and validation data obtained in Step S1 to train the improved DeepLabV3+ semantic segmentation network model;
[0011] Step S4. Input the test data obtained in Step S1 into the improved DeepLabV3+ semantic segmentation network model in Step S3 to obtain the semantic segmentation result of the remote sensing road image.
[0012] Further, the specific steps of Step S1 include the following steps:
[0013] S11. Download or create a remote sensing dataset from an open-source dataset website;
[0014] S12. Place the original image files and label files in different folders respectively, which were originally in one folder;
[0015] S13. Randomly divide the data in the dataset into training data, validation data, and test data according to the ratio of 2:1:1. The list files of the file names after division are stored in the path where the project is located, namely train.txt, val.txt, and test.txt respectively.
[0016] Further, the step S2 specifically includes the following steps:
[0017] S21. Improving the DeepLabV3+ semantic segmentation network model can be divided into an encoder module and a decoder module;
[0018] S22. In the encoder module, MobileNetV2 is used as the backbone network to extract shallow features and deep features of remote sensing images;
[0019] S23. Use the Atrous Spatial Pyramid Pooling module (also known as the ASPP module, ASPP is the English abbreviation of Atrous Spatial Pyramid Pooling) to perform further feature extraction operations on the deep features obtained in S21. The Atrous Spatial Pyramid Pooling module consists of a 1×1 convolution, three dilated convolutions with dilation rates of 6, 12, and 18 respectively, and an ImagePooling (global average pooling) module. The three dilated convolutions are used to capture receptive field information at different scales and capture feature information at different scales. Global average pooling and the 1×1 convolutional layer are used to extract features;
[0020] S24. Use the concatenate feature fusion method to stack the feature layers with different receptive fields obtained in step S23. At this time, the number of input channels is 5 times the original input channel number, and the number of channels can be reduced to the original value through a 1×1 convolutional layer to obtain deep features;
[0021] S25. In the decoder module, after adjusting the number of channels of the shallow features obtained in step S22 using a 1×1 convolution, perform concatenate feature fusion with the result after quadruple upsampling of the deep feature layer obtained in step S24;
[0022] S26. Use two 3×3 convolutional layers to refine the feature fusion result obtained in step S25, and then perform quadruple upsampling to obtain the segmentation prediction map.
[0023] Further, the step S3 specifically includes the following steps:
[0024] S31. Set the initial parameters of the training model as follows:
[0025] Initial learning rate, i.e., learning rate: 0.014;
[0026] Weight decay, i.e., weight decay: 0.0005;
[0027] Momentum, i.e., momentum: 0.9;
[0028] The batch size is determined according to the video memory size of the server for actual training;
[0029] S32. During the training process, adopt the R-Drop regularization method, that is: in each mini-batch training, each data sample undergoes two forward passes, and each pass is implemented by different sub-models by randomly deleting some hidden units.
[0030] The specific process is as follows: The training data is The goal of training is to learn a model P w (y i |x i ), where n is the number of training samples, (x i , y i ) is the labeled data pair, x i is the input data, y i is the label, and the loss of each sample is the cross-entropy:
[0031] L i =-logP w (y i |x i )
[0032] In the case of using the R-Drop regularization method, it can be considered that the sample passes through two slightly different models, denoted as and The final loss of the model is divided into two parts. One part is the conventional cross-entropy:
[0033]
[0034] The other part is the symmetric KL divergence between the two models, and its role is to make the outputs of the models passing through different Dropouts twice as consistent as possible:
[0035]
[0036] The final loss of the network model is the weighted sum of the above two losses:
[0037]
[0038] Among them, α is the weight of the auxiliary loss, which is set to 1, and the cross-entropy is used as the loss function.
[0039] S33. Calculate the gradient according to the loss function obtained in step S32, and use the stochastic gradient descent method as the optimizer to update the weight values and bias values of the neural network.
[0040] S34. Introduce Pixel accuracy (PA) and Mean Intersection over Union (MIoU) to evaluate the performance of the model. PA represents the proportion of the number of pixels with correct predicted categories in the total number of pixels, and MIoU represents the accuracy of the network model in segmenting images. The higher the MIoU value, the better the image segmentation effect. The calculation methods are as follows:
[0041]
[0042]
[0043] In the above formula, TP (True Positive) represents that the model prediction is correct, that is, both the model prediction and the actual are positive examples; FP (False Positive) represents that the model prediction is incorrect, that is, the model predicts this category as a positive example, but actually this category is a negative example; FN (False Negative) represents that the model prediction is incorrect, that is, the model predicts this category as a negative example, but actually this category is a positive example; TN (True Negative) represents that the model prediction is correct, meaning that both the model prediction and the actual are negative examples; N represents the number of categories, and the subscript i represents the i-th category.
[0044] S35. Repeat the training process of steps S32 - S24. After each round of training, use the validation dataset to evaluate the network model, save the model according to the optimal MIoU result, and stop training until the number of iterations reaches the set value, and save the trained model.
[0045] Furthermore, step S4 specifically includes the following steps:
[0046] S41. Load the model trained in step S3, and read in the test images and labels of the test data obtained in step S1.
[0047] S42. Calculate the index score and save the test results.
[0048] The present invention also discloses a remote sensing image semantic segmentation system based on an improved DeepLabV3+ network, including the following modules:
[0049] Data Classification Module: Obtain the remote sensing road dataset and perform preprocessing. The data in the dataset is divided into training data, validation data, and test data;
[0050] Model Building Module: Build an improved DeepLabV3+ semantic segmentation network model based on the Pytorch environment;
[0051] Training Module: Use the training data and validation data obtained by the data classification module to train the improved DeepLabV3+ semantic segmentation network model;
[0052] Obtaining Segmentation Result Module: Input the test data obtained by the data classification module into the improved DeepLabV3+ semantic segmentation network model of the training module to obtain the semantic segmentation result of the remote sensing road image.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] A method for semantic segmentation of remote sensing images based on an improved DeepLabV3+ network. Compared with the method based on the traditional DeepLabV3+ network model, since the R-Drop regularization method is adopted in the present invention, the outputs of two sub-models randomly selected from dropout for each data sample during training can be regularized. The present invention can not only reduce the degree of freedom of the network model parameters, but also alleviate the inconsistency between the training and inference stages, and enhance the generalization ability. Description of the Drawings
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0056] Figure 1 It is a schematic flowchart of a method for semantic segmentation of remote sensing images based on an improved DeeplabV3+ model provided in Embodiment 1 of the present invention.
[0057] Figure 2 It is a schematic diagram of the R-Drop regularization method provided in Embodiment 1 of the present invention.
[0058] Figure 3 It is a diagram of the semantic segmentation result of the remote sensing road image provided in Embodiment 1 of the present invention.
[0059] Figure 4 It is a block diagram of a system for semantic segmentation of remote sensing images based on an improved DeepLabV3+ network provided in Embodiment 2 of the present invention. Detailed implementation manners
[0060] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.
[0061] Embodiment 1
[0062] As Figure 1 shown, this embodiment provides a remote sensing image semantic segmentation method based on an improved DeepLabV3+ model, which specifically includes the following steps:
[0063] Step S1. Obtain a remote sensing road dataset and perform preprocessing. In this embodiment, the DeepGlobe Road Extraction Dataset dataset downloaded from the open-source dataset website kaggle.com is used, and 2000 remote sensing road RGB satellite images with a size of 1024×1024 are randomly selected from it and randomly divided into training data, validation data, and test data according to a ratio of 2:1:1.
[0064] Step S2. Build an improved DeepLabV3+ semantic segmentation network based on the Pytorch environment. In this embodiment, MobileNetV2 is selected as the backbone network of the DeepLabV3+ semantic segmentation network to extract shallow features and deep features; the deep features will be input into the ASPP module to obtain multi-scale feature layers with different receptive fields through 5 different operations such as convolution, dilated convolution, and global average pooling. After concatenate stacking processing, the number of channels is reduced to the original value through 1×1 convolution to obtain deep features and input them into the decoder module. In the decoder module of the network model, the number of channels of the shallow features input from the encoder module is adjusted, the deep features are upsampled by 4 times, and then the results of the two are concatenate stacked. After the stacking is completed, the stacked result is subjected to two 3×3 depthwise separable convolutions and 4 times upsampling to restore to the original image size, that is, the predicted segmentation map of the remote sensing image is obtained.
[0065] Step S3. Use the training data and validation data obtained in Step S1 to train the improved DeepLabV3+ semantic segmentation network model. In order to verify the feasibility of the network designed by the present invention and the recognition effect of the path in a complex environment, the network is programmed and trained and tested. The specific experimental environment and configuration are shown in Table 1:
[0066] Table 1 Experimental environment and configuration
[0067]
[0068]
[0069] And set the initial parameters of the training model as shown in Table 2:
[0070] Table 2 Initial Parameter Settings
[0071]
[0072] After setting the above parameters, training can be carried out. During the training process, the R-Drop regularization method is used instead of the Dropout method used in the original DeepLabV3+ network. Specifically, in each mini-batch training, each data sample undergoes two forward passes, and each pass is implemented by different sub-models by randomly deleting some hidden units. The schematic diagram of the R-Drop regularization method is as Figure 2 shown.
[0073] The specific process is as follows:
[0074] The training data is The goal of training is to learn a model P w (y i |x i ), where n is the number of training samples, (x i , y i ) is the labeled data pair, x i is the input data, y i is the label, and the loss of each sample is the cross-entropy:
[0075] L i =-logP w (y i |x i )
[0076] In the case of using the R-Drop regularization method, it can be considered that the sample passes through two slightly different models, denoted as and The final loss of the model is divided into two parts. One part is the conventional cross-entropy:
[0077]
[0078] The other part is the symmetric KL divergence between the two models, and its role is to make the outputs of the models passing through different Dropouts twice as consistent as possible:
[0079]
[0080] The final loss of the network model is the weighted sum of the above two losses:
[0081]
[0082] Among them, α is the weight of the auxiliary loss, which is set to 1, and the cross-entropy loss function is adopted.
[0083] Pixel accuracy (PA) and Mean Intersection over Union (MIoU) are introduced to evaluate the performance of the model. PA represents the proportion of the number of pixels with correct predicted categories to the total number of pixels, and MIoU represents the accuracy of the network model in segmenting images. The higher the MIoU value, the better the image segmentation effect. The calculation methods are as follows:
[0084]
[0085]
[0086] In the formula, TP (True Positive) represents that the model prediction is correct, that is, both the model prediction and the actual are positive examples; FP (False Positive) represents that the model prediction is incorrect, that is, the model predicts this category as a positive example, but actually this category is a negative example; FN (False Negative) represents that the model prediction is incorrect, that is, the model predicts this category as a negative example, but actually this category is a positive example; TN (True Negative) represents that the model prediction is correct, meaning that both the model prediction and the actual are negative examples; N represents the number of categories, and the subscript i represents the i-th category.
[0087] During the training stage, the stochastic gradient descent method is used as the optimizer to calculate the updated weight values and bias values of the convolutional neural network; after each round of training, the validation dataset is used to evaluate the network model, and the model is saved according to the optimal MIoU result. Training stops after 300 iterations, and the trained model is saved.
[0088] Step S4. Input the test data obtained in step S1 into the improved DeepLabV3+ semantic segmentation network model to obtain the semantic segmentation result of the remote sensing road image, and the result graph is as Figure 3 shown.
[0089] In addition to conducting experiments on the improved DeepLabV3+ semantic segmentation network model, the present invention also trains a corresponding model with the DeepLabV3+ algorithm on the selected remote sensing road dataset and compares it with the performance of the algorithm of the present invention. The performance of the two algorithms on the remote sensing road dataset is shown in Table 3:
[0090] Table 3 Performance comparison of two models on the remote sensing road dataset
[0091] Model method PA(%) MIoU(%) DeepLabV3+ 97.3721 73.8854 The present invention 97.6744 76.8213
[0092] The remote sensing semantic segmentation method that improves the DeepLabV3+ network proposed by the present invention has been improved in pixel accuracy and provides nearly 3 points in the mean intersection over union. The segmentation effect of the method proposed by the present invention on images is significantly better than that of the original DeepLabV3+ algorithm.
[0093] Example 2
[0094] As Figure 4 shown, this embodiment discloses a remote sensing image semantic segmentation system based on an improved DeepLabV3+ network, including the following modules:
[0095] Data classification module: Obtain a remote sensing road data set and perform preprocessing. The data in the data set is divided into training data, validation data, and test data;
[0096] Model building module: Build an improved DeepLabV3+ semantic segmentation network model based on the Pytorch environment;
[0097] Training module: Use the training data and validation data obtained by the data classification module to train the improved DeepLabV3+ semantic segmentation network model;
[0098] Module for obtaining segmentation results: Input the test data obtained by the data classification module into the improved DeepLabV3+ semantic segmentation network model of the training module to obtain the semantic segmentation result of the remote sensing road image.
[0099] Other contents of this embodiment can refer to Embodiment 1.
[0100] The above are only the preferred embodiments of the present invention, and are not limitations to the present invention in other forms. Those skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A semantic segmentation method for remote sensing images based on an improved DeepLabV3+ network, characterized in that It includes the following steps: S1. Obtain the remote sensing road dataset and perform preprocessing. The data in the dataset is divided into training data, validation data, and test data; S2. Build an improved DeepLabV3+ semantic segmentation network model based on the Pytorch environment; S3. Use the training data and validation data obtained in step S1 to train the improved DeepLabV3+ semantic segmentation network model; S4. Input the test data obtained in step S1 into the improved DeepLabV3+ semantic segmentation network model in step S3 to obtain the semantic segmentation result of the remote sensing road image; Step S2 specifically includes the following steps: S21. The improved DeepLabV3+ semantic segmentation network model is divided into an encoder module and a decoder module; S22. In the encoder module, MobileNetV2 is used as the backbone network to extract shallow features and deep features of the remote sensing image; S23. Use the spatial pyramid pooling module to perform further feature extraction operations on the deep features obtained in S21; The spatial pyramid pooling module consists of a 1×1 convolution, three dilated convolutions with dilation rates of 6, 12, and 18 respectively, and an ImagePooling module. The three dilated convolutions are used to capture receptive field information of different scales and capture feature information of different scales. The ImagePooling module and the 1×1 convolutional layer are used to extract features; S24. Use the concatenate feature fusion method to stack the feature layers with different receptive fields obtained in step S23. At this time, the input channel number is 5 times the original input channel number, and the channel number is reduced to the original value through a 1×1 convolutional layer to obtain deep features; S25. In the decoder module, after adjusting the channel number of the shallow features obtained in step S22 using a 1×1 convolution, perform concatenate feature fusion with the result after 4-fold upsampling of the deep feature layer obtained in step S24; S26. Use two 3×3 convolutional layers to refine the feature fusion result obtained in step S25, and then perform four-fold upsampling to obtain the segmentation prediction map; Step S3 specifically includes the following steps: S31. Set the initial parameters of the training model as follows: Initial learning rate, i.e., learning rate: 0.014; Weight decay, i.e., weight decay: 0.0005; Momentum, i.e., momentum: 0.9; S32. During the training process, the R-Drop regularization method is adopted, that is: in each mini-batch training, each data sample undergoes two forward passes, and each pass is processed by different sub-models by randomly deleting some hidden units; specifically as follows: the training data is The goal of training is to learn a model P w (y i |x i ), where n is the number of training samples, (x i , y i ) is the labeled data pair, x i is the input data, y i is the label, and the loss of each sample is the cross-entropy: L i = -logP w (y i |x i ) In the case of using the R-Drop regularization method, it is considered that the samples pass through two slightly different models, denoted as and The final loss of the model is divided into two parts. One part is the conventional cross-entropy: Another part is the symmetric KL divergence between the two models: The final loss of the network model is the weighted sum of the above two losses: Where α is the weight of the auxiliary loss, set to 1, and the loss function uses cross-entropy; S33. Calculate the gradient according to the loss function obtained in step S32, and use the stochastic gradient descent method as the optimizer to update the weight values and bias values of the neural network; S34. The Pixel Accuracy (PA) and the Mean Intersection over Union (MIoU) are introduced to evaluate the performance of the model. PA represents the proportion of the number of pixels with correct predicted classes to the total number of pixels, and MIoU represents the accuracy of the network model in segmenting images. The higher the MIoU value, the better the image segmentation effect. The calculation methods are as follows: In the formula, TP represents correct model prediction, that is, both the model prediction and the actual are positive examples; FP represents incorrect model prediction, that is, the model predicts this class as a positive example, but actually this class is a negative example; FN represents incorrect model prediction, that is, the model predicts this class as a negative example, but actually this class is a positive example; TN represents correct model prediction, meaning that both the model prediction and the actual are negative examples; N represents the number of classes, and the subscript i represents the i-th class; S35. Repeat the training process of steps S32 - S34. After each round of training, use the validation data to evaluate the network model, and save the model according to the optimal MIoU result. Stop training until the number of iterations reaches the set value, and save the trained model.
2. The remote sensing image semantic segmentation method based on the improved DeepLabV3+ network according to claim 1, wherein Step S1 specifically includes the following steps: S11. Download or create a remote sensing image dataset from an open-source dataset website; S12. Place the original image files and label files in different folders respectively, which were originally in one folder; S13. Randomly divide the data in the dataset into training data, validation data, and test data according to the ratio of 2:1:
1. The file name list files after division are stored in the path where the project is located, namely train.txt, val.txt, and test.txt respectively.
3. The remote sensing image semantic segmentation method based on the improved DeepLabV3+ network according to claim 1, wherein Step S4 specifically includes the following steps: S41. Load the model trained in step S3, and read the test images and labels of the test data obtained in step S1; S42. Calculate the metric scores and save the test results.
4. A remote sensing image semantic segmentation system based on an improved DeepLabV3+ network, characterized in that, It includes the following modules: Data Classification Module: Obtain a remote sensing road dataset and perform preprocessing. The data in the dataset is divided into training data, validation data, and test data; Model Building Module: Build an improved DeepLabV3+ semantic segmentation network model based on the Pytorch environment; Training Module: Use the training data and validation data obtained by the Data Classification Module to train the improved DeepLabV3+ semantic segmentation network model; Obtain Segmentation Result Module: Input the test data obtained by the Data Classification Module into the improved DeepLabV3+ semantic segmentation network model of the Training Module to obtain the semantic segmentation result of the remote sensing road image; The Model Building Module is specifically as follows: The improved DeepLabV3+ semantic segmentation network model is divided into an encoder module and a decoder module; In the encoder module, MobileNetV2 is used as the backbone network to extract shallow features and deep features of the remote sensing image; The spatial pyramid pooling module is used to perform further feature extraction operations on the obtained deep features; the spatial pyramid pooling module consists of a 1×1 convolution, three dilated convolutions with dilation rates of 6, 12, and 18 respectively, and an ImagePooling module. The three dilated convolutions are used to capture receptive field information at different scales and capture feature information at different scales. The ImagePooling module and the 1×1 convolutional layer are used to extract features; The concatenate feature fusion method is used to stack the feature layers with different receptive fields obtained. At this time, the input channel number is 5 times the original input channel number, and the channel number is reduced to the original value through a 1×1 convolutional layer to obtain deep features; In the decoder module, after adjusting the channel number of the obtained shallow features using a 1×1 convolution, the result is concatenated with the deep feature layer obtained in step S24 after 4-fold upsampling for feature fusion; Two 3×3 convolutional layers are used to refine the obtained feature fusion result, and then 4-fold upsampling is performed to obtain the segmentation prediction map; The training module is as follows: S31. Set the initial parameters of the training model as follows: Initial learning rate, that is, learning rate: 0.014; Weight decay, that is, weight decay: 0.0005; Momentum, that is, momentum: 0.9; S32. During the training process, the R-Drop regularization method is adopted, that is: in each mini-batch training, each data sample undergoes two forward passes, and each pass is processed by different sub-models by randomly deleting some hidden units; specifically as follows: The training data is The goal of training is to learn a model P w (y i |x i ), where n is the number of training samples, (x i , y i ) is the labeled data pair, x i is the input data, y i is the label, and the loss of each sample is the cross-entropy: L i = -logP w (y i |x i ) In the case of using the R-Drop regularization method, it is considered that the sample passes through two slightly different models, denoted as and The final loss of the model is divided into two parts. One part is the conventional cross-entropy: Another part is the symmetric KL divergence between the two models: The final loss of the network model is the weighted sum of the above two losses: Among them, α is the weight of the auxiliary loss, set to 1, and the loss function uses cross-entropy; S33. Calculate the gradient according to the loss function obtained in step S32, and use the stochastic gradient descent method as the optimizer to update the weight values and bias values of the neural network; S34. Introduce pixel accuracy PA and mean intersection over union MIoU to evaluate the performance of the model. PA represents the proportion of the number of pixels with correct predicted categories in the total number of pixels, and MIoU represents the accuracy of the network model in segmenting images. The higher the MIoU value, the better the image segmentation effect; the calculation methods are as follows: In the formula, TP represents correct model prediction, that is, both the model prediction and the actual are positive examples; FP represents incorrect model prediction, that is, the model predicts this category as a positive example, but actually this category is a negative example; FN represents incorrect model prediction, that is, the model predicts this category as a negative example, but actually this category is a positive example; TN represents correct model prediction, meaning that both the model prediction and the actual are negative examples; N represents the number of categories, and the subscript i represents the i-th category; S35. Repeat the training process of steps S32 - S34. After each round of training, use the validation data to evaluate the network model, save the model according to the optimal MIoU result, and stop training until the number of iterations reaches the set value, and save the trained model.
Citation Information
Patent Citations
Improved remote sensing image segmentation method based on DeepLabV3
CN114387518A