Remote Sensing Image Semantic Segmentation Method Based on Adaptive Multi-Scale Feature Pyramid Network
By introducing an adaptive multi-scale feature pyramid network and a hollow space convolution pooling pyramid module in the semantic segmentation of remote sensing images, the problem of insufficient feature extraction capabilities in the prior art is solved, and efficient segmentation of targets of different scales is achieved.
Patent Information
- Application Number
- CN202210477728.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The existing high-score remote sensing image semantic segmentation method is insufficient in the feature extraction ability, and blindly expanding the receptive field cannot effectively solve the problem.
A remote sensing image semantic segmentation method based on adaptive multi-scale feature pyramid network is designed. By introducing pyramid networks into the encoder to fusion multi-scale features, and using hollow space convolution pooling pyramid modules and switchable hollow convolutions in the decoder, the receptive field is adaptively adjusted to adapt to different target features.
The feature extraction capability of semantic segmentation of remote sensing images is improved, and the features of large and small targets can be effectively captured, which significantly improves the segmentation effect.
Smart Images

Figure CN114758134B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a remote sensing image semantic segmentation method based on an adaptive multi-scale feature pyramid network. Background Art
[0002] Semantic segmentation of high-resolution remote sensing images is also known as classification of ground object elements. Classification of ground object elements is a pixel-level remote sensing image classification with finer granularity, and it has practical guiding significance in carrying out national geographical conditions monitoring, "red line of cultivated land", "ecological red line", etc. based on remote sensing images. For the semantic segmentation of remote sensing images, currently, manual methods are mainly used to extract and divide ground object elements. This way of processing remote sensing data is inefficient and costly. Now, there is an urgent need to replace it with a more accurate and faster automated method. Semantic segmentation of high-resolution remote sensing images based on deep learning has attracted extensive attention and research in recent years. However, due to the disordered distribution and complex types of ground objects in remote sensing images, regions with different target sizes need to adapt different receptive fields to extract features. At the same time, factors such as the scale change of targets in remote sensing images also pose great challenges to the semantic segmentation of remote sensing images. Therefore, it is a feasible improvement scheme to adaptively select receptive fields of variable sizes according to different target features.
[0003] Chinese Patent Publication No.: CN 113837191 A; Publication Date: December 24, 2021, discloses a cross-satellite remote sensing image semantic segmentation method based on bidirectional unsupervised domain adaptation fusion, including source domain-target domain image bidirectional conversion model training, image conversion model bidirectional converter parameter selection, source domain-target domain image bidirectional conversion, source domain and pseudo-target domain semantic segmentation model training, source domain and target domain class segmentation probability generation, and fusion. The invention uses source-target and target-source bidirectional domain adaptation to probabilistically fuse class segmentation on the source domain and the target domain, improving the accuracy and robustness of the cross-satellite remote sensing image semantic segmentation model. Further, through bidirectional semantic consistency loss and converter parameter selection, the influence brought by the unstable converter effect in the image bidirectional conversion model is avoided.
[0004] Chinese Patent Publication No.: CN 113392845 A; Publication Date: September 14, 2021, discloses a deep learning remote sensing image semantic segmentation method and system based on U-NET. The method includes: correcting and reconstructing the initial remote sensing data, classifying and preprocessing to construct a remote sensing sample library, making prediction and classification based on the segmentation network model to obtain a basic training data set, and enhancing it to obtain a target training data set; using the trained segmentation network model to process the remote sensing image to obtain the target segmentation result. The invention uses atmospheric correction and radiation correction to eliminate the radiation error, eliminates interference data, and then performs super-resolution reconstruction on the remote sensing image, effectively avoiding the interference of external factors and reducing the resolution requirement for the remote sensing image; and the segmentation network model can adaptively learn the features of different targets and achieve multi-target segmentation in the same remote sensing image. Therefore, when changing the segmentation target, only the corresponding data set needs to be used to retrain the segmentation network model, without manually reconstructing features and algorithms, greatly reducing the workload.
[0005] Chinese Patent Publication No.: CN 109409240 B; Publication Date: February 11, 2022, discloses a SegNet remote sensing image semantic segmentation method combined with random walk, which is divided into a SegNet initial segmentation step and a random walk optimization segmentation step. In the SegNet initial segmentation step, the initial semantic segmentation image and category intensity information are output through SegNet; in the random walk optimization segmentation step, first, the random walk seed region is selected, and according to the classification intensity information output by SegNet, the classification significance index of different categories is calculated, and the threshold is set to select the seed regions of different categories; second, according to the original image gradient and the classification intensity information of SegNet, the calculation of the undirected edge weight is carried out; in the third step, starting from the seed region and combining the undirected edge weight, random walk is carried out on the entire initial segmentation image, and finally the optimized segmentation result on the entire image is obtained. The invention performs random walk on the entire image to achieve prediction error and control, greatly reducing edge burrs and patchy classification errors, and completing high-precision remote sensing image semantic segmentation.
[0006] In summary, there are many limitations in the existing high-resolution remote sensing image semantic segmentation methods, which are mainly manifested in:
[0007] (1) The existing semantic segmentation networks lack the ability to extract features of targets with different scales in remote sensing images.
[0008] (2) It is not advisable to blindly optimize the segmentation effect by expanding the receptive field. A reasonable way is to use a variable-size receptive field to adaptively switch and select according to the features of different objects. Summary of the Invention
[0009] Objective of the Invention: Aiming at the problems existing in the prior art, the present invention proposes a remote sensing image semantic segmentation method based on an adaptive multi-scale feature pyramid network. The designed segmentation network adopts an encoder-decoder structure. In the encoder part, a pyramid network is introduced to fuse low-level features and high-level features. In the decoder part, first, an atrous spatial pyramid pooling module is introduced to fuse multi-scale features to process objects of different sizes. Then, a switchable atrous convolution is designed to adaptively switch to the appropriate receptive field features according to the features while expanding the receptive field. Finally, it is fused with the upsampled features to further enhance the segmentation effect of the network.
[0010] Technical Solution: To achieve the objective of the present invention, the technical solution adopted by the present invention is: a remote sensing image semantic segmentation method based on an adaptive multi-scale feature pyramid network, and the method includes the following steps:
[0011] (1) Construct a dataset for remote sensing image semantic segmentation, make corresponding sample annotations, and divide each type of remote sensing image into a training set Train and a test set Test according to a certain proportion;
[0012] (2) Construct an adaptive multi-scale feature pyramid network, and use the training set of remote sensing image data to train the adaptive multi-scale feature pyramid network;
[0013] (3) Set training parameters, construct a loss function, use the training set to train the constructed adaptive multi-scale feature pyramid network, and update the network parameters until the parameter values converge; the convergence condition is that the value of the loss function no longer decreases;
[0014] (4) Input the test set into the trained network to obtain the semantic segmentation result map of the test set.
[0015] Furthermore, the method for constructing the dataset for remote sensing image semantic segmentation in step (1) is as follows:
[0016] (1.1) Divide the remote sensing image semantic segmentation dataset Image = [Image 1 , …, Image i , …, Image N , and make the corresponding sample label dataset Label = [Label 1 , …, Label i , …, Label N , where N represents the total number of samples of remote sensing images, Image i represents the i-th remote sensing image, and Label i represents the label corresponding to the i-th remote sensing image;
[0017] (1.2) Divide the remote sensing image dataset into a training set and a test set. For the remote sensing images in the dataset, randomly select M images from them to construct the training set, and the remaining N - M remote sensing images to construct the test set. Then in the training set, there are: TrainImage = [TrainImage 1 , …, TrainImage i , …, TrainImage M , TrainLabel = [TrainLabel 1 , …, TrainLabel i , …, TrainLabeL M . In the test set, there are TestImage = [TestImage 1 , …, TestImage i , …, TestImage N-M , TestLabel = [TestLabel 1 , …, TestLabel i , …, TestLabeL N-M . Among them, TrainImage i represents the i-th remote sensing image for training, TrainLabel i represents the label corresponding to the i-th remote sensing image for training, TestImage i represents the i-th remote sensing image for testing, and TestLabel i represents the label corresponding to the i-th remote sensing image for testing.
[0018] Further, the step (2) constructs an adaptive multi-scale feature pyramid network, specifically as follows:
[0019] (2.1) First, build a convolutional neural network based on ResNet50 and remove the fully connected layers therein. Take the outputs of the convolutional layers at the end of each stage except the first stage in ResNet50 as the input features of the pyramid network, which are set as four groups, namely C 2 , C 3 , C 4 , C 5 respectively;
[0020] (2.2) Feed the four groups of features C 2 , C 3 , C 4 , C 5 obtained in step (2.1) into the pyramid network. Through the feature fusion of the pyramid network, feature maps of different sizes are obtained respectively, denoted as P 2 , P 3 , P4 , P 5 ;
[0021] Furthermore, through the feature fusion of the pyramid network in step (2.2), feature maps of different sizes are obtained, denoted as P 2 , P 3 , P 4 , P 5 , specifically as follows:
[0022] (2.2.1) The feature C 5 is subjected to channel reduction through a 1×1 convolution to obtain the feature P 5 , specifically:
[0023] P 5 = δ(G 1×1conv (C 5 ))
[0024] where δ(·) represents the rectified linear unit function, and G 1×1conv (·) represents a convolution operation of size 1×1.
[0025] (2.2.2) The feature P 5 is upsampled by a factor of 2 and fused with the feature C 4 after being reduced in dimension through a 1×1 convolution to obtain the feature P 4 , specifically:
[0026]
[0027] where represents element-wise addition, and Up 2× (·) represents upsampling by a factor of 2.
[0028] (2.2.3) The feature P 4 is upsampled by a factor of 2 and fused with the feature C 3 after being reduced in dimension through a 1×1 convolution to obtain the feature P 3 , specifically:
[0029]
[0030] (2.2.4) The feature P 3 is upsampled by a factor of 2 and fused with the feature C 2 after being reduced in dimension through a 1×1 convolution to obtain the feature P 2 , specifically:
[0031]
[0032] (2.3) The four groups of features P 2 , P3 , P 4 , P 5 As the output of the encoder, it is fed into the decoder;
[0033] (2.4) Send the feature P 5 into the atrous spatial pyramid pooling part in the decoder to construct multi-scale features, obtaining the feature U 5 , specifically:
[0034]
[0035] where represents stacking features in the channel dimension, represents a 3×3 convolution operation with a dilation rate of 1, represents a 3×3 convolution operation with a dilation rate of 3, represents a 3×3 convolution operation with a dilation rate of 5, G 2×2pool (·) represents an average pooling operation of size 2×2.
[0036] (2.5) Send the features P 2 , P 3 , P 4 into the switchable atrous convolution for processing respectively;
[0037] Furthermore, in step (2.5), the features P 2 , P 3 , P 4 are sent into the switchable atrous convolution for processing respectively, obtaining four groups of processed feature maps, denoted as P 2 ″, P 3 ″, P 4 ″, specifically as follows:
[0038] (2.5.1) Perform global average pooling on the feature P 2 , then pass through a 1×1 convolution, and then fuse with the original feature to obtain the feature P 2 ′. Specifically:
[0039]
[0040] where G GAP (·) represents global average pooling.
[0041] Then this feature is processed by the switchable atrous convolution to obtain the output feature P 2 ″. Specifically:
[0042] α(P 2 ′) = δ(G 1×1conv (G GAP (P 2′)))
[0043]
[0044] (2.5.2) Perform global average pooling on feature P 3 , then perform 1×1 convolution, and then fuse it with the original feature to obtain feature P 3 ′. Specifically:
[0045]
[0046] Then process this feature with switchable dilated convolution to obtain the output feature P 3 ″. Specifically:
[0047] α(P 3 ′) = δ(G 1×1conv (G GAP (P 3 ′)))
[0048]
[0049] (2.5.3) Perform global average pooling on feature P 4 , then perform 1×1 convolution, and then fuse it with the original feature to obtain feature P 4 ′. Specifically:
[0050]
[0051] Then process this feature with switchable dilated convolution to obtain the output feature P 4 ″. Specifically:
[0052] α(P 4 ′) = δ(G 1×1conv (G GAP (P 4 ′)))
[0053]
[0054] (2.6) Upsample feature U 5 by a factor of 2 and fuse it with the feature P 4 ″ that has been processed by switchable dilated convolution, and then perform 3×3 convolution to obtain feature U 4 , specifically:
[0055]
[0056] where G 3×3conv (·) represents the 3×3 convolution operation, and Up 2× (·) represents the upsampling operation by a factor of 2.
[0057] (2.7) Upsample the feature U by a factor of 2 and fuse it with the feature P processed by the switchable atrous convolution, and then obtain the feature U through a 3×3 convolution. 4 Specifically: 3 3 3 2
[0058]
[0059] (2.8) Upsample the feature U by a factor of 2 and fuse it with the feature P processed by the switchable atrous convolution, and then obtain the feature U through a 3×3 convolution. 3 Specifically: 2 2
[0060]
[0061] (2.9) Further upsample the feature U by a factor of 4, and then perform a 1×1 convolution to obtain the output segmentation result map O. Specifically:
[0062] 2 O = Softmax(G 1×1conv (Up 4× (U 2 )))
[0063]
[0064]
[0065] 4× wherein, Up(·) represents the 4-fold upsampling process, and Softmax(·) represents the softmax activation function.
[0066] Advantageous effects: Compared with the prior art, the technical solution of the present invention has the following advantageous technical effects: (1) The present invention constructs multi-scale features by introducing a pyramid network in the encoder part, and further enhances the multi-scale information by introducing an atrous spatial pyramid pooling module in the decoder.
[0067] Figure 1 (2) The present invention designs a switchable atrous convolution module, which can adaptively select the receptive field of the required scale size according to different features, so that the network can effectively capture the features of large target objects in the remote sensing image while fully capturing the features of small target objects. Brief Description of the Drawings
[0067] Figure 1 is the network architecture diagram of the adaptive multi-scale feature pyramid built.
[0068] Figure 2 is the calculation flow chart of the switchable atrous convolution. Detailed Embodiments
[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Among them, the described embodiments are some, rather than all, of the embodiments of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention.
[0070] Embodiment 1
[0071] Reference Figure 1 , this embodiment provides a remote sensing image semantic segmentation method based on an adaptive multi-scale feature pyramid network, which specifically includes the following steps:
[0072] (1) Construct a dataset for remote sensing image semantic segmentation, and make corresponding sample annotations. Divide each type of remote sensing image into a training set Train and a test set Test according to a certain proportion;
[0073] (1.1) Divide the remote sensing image semantic segmentation dataset Image = [Image 1 , …, Image i , …, Image N , and make the corresponding sample label dataset Label = [Label 1 , …, Label i , …, Label N , where N represents the total number of samples of remote sensing images, Image i represents the i-th remote sensing image, and Label i represents the label corresponding to the i-th remote sensing image;
[0074] (1.2) Divide the remote sensing image dataset into a training set and a test set. For the remote sensing images in the dataset, randomly select M images from them to construct the training set, and the remaining N - M remote sensing images to construct the test set. Then in the training set, there are: TrainImage = [TrainImage 1 , …, TrainImage i , …, TrainImage M , TrainLabel = [TrainLabel 1 , …, TrainLabel i , …, TrainLabel M , and in the test set, there are TestImage = [TestImage 1 , …, TestImage i , …, TestImageN-M , TestLabel = [TestLabel 1 , …, TestLabel i , …, TestLabeL N-M . Among them, TrainImage i represents the i-th remote sensing image for training, TrainLabel i represents the label corresponding to the i-th remote sensing image for training, TestImage i represents the i-th remote sensing image for testing, TestLabel i represents the label corresponding to the i-th remote sensing image for testing.
[0075] Among them, each remote sensing image for training and testing is segmented into a size of 256×256.
[0076] (2) Construct an adaptive multi-scale feature pyramid network and use the remote sensing image data training set to train the adaptive multi-scale feature pyramid network;
[0077] (2.1) First, build a convolutional neural network based on ResNet50 and remove the fully connected layers therein. The output of the convolutional layers at the end of each stage of ResNet50 except the first stage is used as the input feature of the pyramid network, and a total of four groups are set as C 2 , C 3 , C 4 , C 5 . Among them, the tensor size of feature C 2 is 64×64×256, the tensor size of feature C 3 is 32×32×512, the tensor size of feature C 4 is 16×16×1024, and the tensor size of feature C 5 is 8×8×2048.
[0078] (2.2) Send the four groups of features C 2 , C 3 , C 4 , C 5 obtained in step (2.1) into the pyramid network. Through the feature fusion of the pyramid network, feature maps of different sizes are obtained respectively, denoted as P 2 , P 3 , P 4 , P 5 . Among them, the tensor size of feature P 5 is 8×8×256, the tensor size of feature P 4 is 16×16×256, the tensor size of feature P 3 is 32×32×256, feature P2 The tensor size is 64×64×256.
[0079] Furthermore, through the feature fusion of the pyramid network in step (2.2), feature maps of different sizes are obtained, denoted as P 2 、P 3 、P 4 、P 5 , specifically as follows:
[0080] (2.2.1) The feature C 5 is subjected to channel dimensionality reduction through a 1×1 convolution to obtain the feature P 5 , specifically:
[0081] P 5 = δ(G 1×1conv (C 5 ))
[0082] where δ(·) represents the rectified linear unit function, and G 1×1conv (·) represents a 1×1 convolution operation.
[0083] (2.2.2) The feature P 5 is upsampled by a factor of 2 and fused with the feature C 4 that has been dimensionally reduced through a 1×1 convolution to obtain the feature P 4 , specifically:
[0084]
[0085] where represents element-wise addition, and Up 2× (·) represents upsampling by a factor of 2.
[0086] (2.2.3) The feature P 4 is upsampled by a factor of 2 and fused with the feature C 3 that has been dimensionally reduced through a 1×1 convolution to obtain the feature P 3 , specifically:
[0087]
[0088] (2.2.4) The feature P 3 is upsampled by a factor of 2 and fused with the feature C 2 that has been dimensionally reduced through a 1×1 convolution to obtain the feature P 2 , specifically:
[0089]
[0090] (2.3) The four groups of features P 2 、P3 , P 4 , P 5 As the output of the encoder, it is fed into the decoder;
[0091] (2.4) Send the feature P 5 into the atrous spatial pyramid pooling part in the decoder to construct multi-scale features, and obtain the feature U 5 . The tensor size of the feature U 5 is 8×8×256, specifically:
[0092]
[0093] where represents stacking features in the channel dimension, represents a 3×3 convolution operation with a dilation rate of 1, represents a 3×3 convolution operation with a dilation rate of 3, represents a 3×3 convolution operation with a dilation rate of 5, G 2×2pool (·) represents an average pooling operation with a size of 2×2.
[0094] (2.5) Send the features P 2 , P 3 , P 4 into the switchable atrous convolution for processing respectively;
[0095] Furthermore, in step (2.5), the features P 2 , P 3 , P 4 are respectively sent into the switchable atrous convolution for processing, and four groups of processed feature maps are obtained, which are respectively denoted as P 2 ″, P 3 ″, P 4 ″, specifically as follows:
[0096] (2.5.1) Perform global average pooling on the feature P 2 , then pass through a 1×1 convolution, and then fuse with the original feature to obtain the feature P 2 ′. Specifically:
[0097]
[0098] where G GAP (·) represents global average pooling.
[0099] Then this feature is processed by the switchable atrous convolution to obtain the output feature P 2 ″. Specifically:
[0100] α(P 2 ′) = δ(G 1×1conv (GGAP (P 2 ′)))
[0101]
[0102] (2.5.2) Global average pooling is performed on the feature P 3 , followed by a 1×1 convolution, and then it is fused with the original feature to obtain the feature P 3 ′. Specifically:
[0103]
[0104] Then this feature is processed by switchable atrous convolution to obtain the output feature P 3 ″. Specifically:
[0105] α(P 3 ′) = δ(G 1×1conv (G GAP (P 3 ′)))
[0106]
[0107] (2.5.3) Global average pooling is performed on the feature P 4 , followed by a 1×1 convolution, and then it is fused with the original feature to obtain the feature P 4 ′. Specifically:
[0108]
[0109] Then this feature is processed by switchable atrous convolution to obtain the output feature P 4 ″. Specifically:
[0110] α(P 4 ′) = δ(G 1×1conv (G GAP (P 4 ′)))
[0111]
[0112] (2.6) Upsample the feature U 5 by a factor of 2 and fuse it with the feature P 4 ″ that has been processed by switchable atrous convolution, and then perform a 3×3 convolution to obtain the feature U 4 , and the tensor size of the feature U 4 is 16×16×256. Specifically:
[0113]
[0114] Among them, G 3×3conv(·) represents a 3×3 convolution operation, Up 2× (·) represents a 2x upsampling operation.
[0115] (2.7) Upsamples the feature U 4 by 2x and fuses it with the feature P 3 processed by the switchable dilated convolution, and then passes through a 3×3 convolution to obtain the feature U 3 , and the tensor size of the feature U 3 is 32×32×256, specifically:
[0116]
[0117] (2.8) Upsamples the feature U 3 by 2x and fuses it with the feature P 2 processed by the switchable dilated convolution, and then passes through a 3×3 convolution to obtain the feature U 2 , and the tensor size of the feature U 2 is 64×64×256, specifically:
[0118]
[0119] (2.9) Further upsamples the feature U 2 by 4x, then performs a 1×1 convolution to obtain the output segmentation result map O, and the size of the segmentation result map O is 256×256, which is the same as the size of the input image, specifically:
[0120] O = Softmax(G 1×1conv (Up 4× (U 2 )))
[0121] where Up 4× (·) represents the 4x upsampling process, and Softmax(·) represents the softmax activation function.
[0122] (3) Set the training parameters, construct the loss function, and use the training set to train the constructed adaptive multi-scale feature pyramid network, and update the network parameters until the parameter values converge; the convergence condition is that the loss function value no longer decreases;
[0123] Furthermore, setting the training parameters and constructing the loss function in step (3) are specifically as follows:
[0124] In this implementation, the set training parameters include: the batch size is set to 16, the initial learning rate is set to 2.5×10 -2 , the weight decay coefficient is 10 -4 , and the parameter optimization uses the stochastic gradient descent method.
[0125] (4) Input the test set into the trained network to obtain the semantic segmentation result map of the test set.
[0126] The above has schematically described the present invention and its embodiments. This description is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention. The actual structure and method are not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, creatively design structural ways and embodiments similar to this technical solution, they all fall within the protection scope of the present invention.
Claims
1. Remote sensing image semantic segmentation method based on adaptive multi-scale feature pyramid network, characterized in that, the method comprises the following steps: (1) Construct a multi-class remote sensing image data set, and make corresponding sample annotations. Divide each type of remote sensing image into a training set Train and a test set Test according to a ratio; (2) Construct an adaptive multi-scale feature pyramid network, and use the training set of remote sensing image data to train the adaptive multi-scale feature pyramid network; (3) Set training parameters, construct a loss function, use the training set to train the constructed adaptive multi-scale feature pyramid network, and update the network parameters until the parameter values converge; the convergence condition is that the loss function value no longer decreases; (4) Input the test set into the trained network to obtain the semantic segmentation result map of the test set; In step (2), the method for constructing the adaptive multi-scale feature pyramid network is as follows: (2.1) First, build a convolutional neural network based on ResNet50 and remove the fully connected layers therein; use the outputs of the convolutional layers at the end of each stage except the first stage in ResNet50 as the input features of the pyramid network, which are divided into four groups and set as C 2 , C 3 , C 4 , C 5 ; (2.2) Feed the four groups of features C 2 、C 3 、C 4 、C 5 into the pyramid network; through the feature fusion of the pyramid network, feature maps of different sizes are obtained respectively, denoted as P 2 、P 3 、P 4 、P 5 ; In step (2.2), through the feature fusion of the pyramid network, feature maps of different sizes are obtained, denoted as P 2 , P 3 , P 4 , P 5 , specifically as follows: (2.2.1) Feature C 5 is dimension-reduced in channels through 1×1 convolution to obtain Feature P 5 , specifically: P 5 = δ(G 1×1conv (C 5 )) where δ(·) represents the rectified linear unit function, and G 1×1conv (·) represents a 1×1 convolution operation; (2.2.2) Upsample the feature P 5 by a factor of 2 and fuse it with the feature C that has been reduced in dimension by a 1×1 convolution 4 to obtain the feature P 4 , specifically as follows: Among them represents element-wise addition, Up 2× (·) represents 2x upsampling; (2.2.3) Upsample the feature P by a factor of 2 and fuse it with the feature C that has been reduced in dimension by 1×1 convolution to obtain the feature P 4 Specifically, it is as follows: 3 Upsample the feature P by a factor of 2 and fuse it with the feature C that has been reduced in dimension by 1×1 convolution to obtain the feature P 3 Specifically, it is as follows: (2.2.4) Upsample the feature P 3 by a factor of 2, and fuse it with the feature C after dimensionality reduction by 1×1 convolution 2 to obtain the feature P 2 , specifically as follows: (2.3) Take the four groups of features P obtained in step (2.2) 2 、P 3 、P 4 、P 5 as the output of the encoder and input it into the decoder; ( 2.4) Send feature P 5 into the atrous spatial pyramid pooling part in the decoder to construct multi-scale features, obtaining feature U 5 , specifically: Among them represents that features are stacked in the channel dimension, represents a 3×3 convolution operation with a dilation rate of 1, represents a 3×3 convolution operation with a dilation rate of 3, represents a 3×3 convolution operation with a dilation rate of 5, G 2×2pool (·) represents an average pooling operation of size 2×2; (2.5) Input feature P 2 、P 3 、P 4 into switchable dilated convolutions for processing respectively; Further, in step (2.5), feature P 2 , P 3 , P 4 are respectively input into the switchable dilated convolution for processing, and four groups of processed feature maps are obtained, which are respectively denoted as P 2 ″, P 3 ″, P 4 ″, specifically as follows: (2.5.1) Perform global average pooling on feature P 2 then perform 1×1 convolution, and then fuse it with the original feature to obtain feature P 2 '; specifically: Among them, G GAP (·) represents global average pooling; Then the feature is processed by switchable dilated convolution to obtain the output feature P 2 ″; specifically: α(P 2 ′) = δ(G 1×1conv (G GAP (P 2 ′))) (2.5.2) Perform global average pooling on feature P 3 Then, perform 1×1 convolution, and then fuse it with the original feature to obtain feature P 3 ′; specifically: Then the feature is processed by switchable dilated convolution to obtain the output feature P 3 ″; specifically: α(P 3 ′) = δ(G 1×1conv (G GAP (P 3 ′))) (2.5.3) Perform global average pooling on feature P 4 then perform 1×1 convolution, and then fuse it with the original feature to obtain feature P 4 '; specifically: Then the feature is processed by switchable dilated convolution to obtain the output feature P 4 ″; specifically: α(P 4 ′) = δ(G 1×1conv (G GAP (P 4 ′))) (2.6) Upsample feature U by a factor of 2 and fuse it with feature P after being processed by switchable dilated convolution 5 ″, and then obtain feature U 4 through a 3×3 convolution, specifically as follows: 4 Among them, G 3×3conv (·) represents a 3×3 convolution operation, Up 2× (·) represents a 2-fold upsampling operation; (2.7) Upsample feature U by a factor of 2 and fuse it with feature P after being processed by switchable dilated convolution 4 to obtain feature U after passing through a 3×3 convolution, specifically: 3 Upsample feature U by a factor of 2 and fuse it with feature P after being processed by switchable dilated convolution 3 to obtain feature U after passing through a 3×3 convolution, specifically: (2.8) Upsample the feature U 3 by a factor of 2 and fuse it with the feature P after being processed by the switchable dilated convolution 2 , and then obtain the feature U after passing through a 3×3 convolution 2 , specifically: (2.9) Feature U 2 Perform 4x upsampling on it, and then perform 1x1 convolution to obtain the output segmentation result map O, specifically as follows: O = Softmax(G 1×1conv (Up 4× (U 2 ))) Among them, Up 4× (·) represents a 4-fold upsampling process, and Softmax(·) represents the softmax activation function.
2. The remote sensing image semantic segmentation method based on the adaptive multi-scale feature pyramid network according to claim 1, and the method for dividing the training set and the test set in step (1) is as follows: (1.1) Divide the remote sensing image semantic segmentation dataset Image = [Image 1 , …, Image i , …, Image N , and create the corresponding sample label dataset Label = [Label 1 , …, Label i , …, Label N , where N represents the total number of samples of the remote sensing images, Image i represents the i-th remote sensing image, and Label i represents the label corresponding to the i-th remote sensing image; (1.2) Divide the remote sensing image dataset into a training set and a test set. For the remote sensing images in the dataset, randomly select M images from them to construct the training set, and the remaining N - M remote sensing images to construct the test set. Then in the training set, there are: TrainImage = [TrainImage 1 , …, TrainImage i , …, TrainImage M , TrainLabel = [TrainLabel 1 , …, TrainLabel i , …, TrainLabel M . In the test set, there are TestImage = [TestImage 1 , …, TestImage i , …, TestImage N-M , TestLabel = [TestLabel 1 , …, TestLabel i , …, TestLabel N-M ; wherein, TrainImage i represents the i-th remote sensing image for training, TrainLabel i represents the label corresponding to the i-th remote sensing image for training, TestImage i represents the i-th remote sensing image for testing, TestLabel i represents the label corresponding to the i-th remote sensing image for testing.
Citation Information
Patent Citations
A semantic segmentation method for remote sensing images combining random walks and SegNet
CN109409240B
Deep learning remote sensing image semantic segmentation method and system based on U-NET
CN113392845A
Cross-satellite remote sensing image semantic segmentation method based on bidirectional unsupervised domain adaptive fusion
CN113837191A
High-resolution remote sensing image classification method based on novel feature pyramid depth network
CN110728192A
Remote sensing image semantic segmentation method based on pyramid segmentation attention module
CN113807210A