A U-Net++ remote sensing semantic segmentation method based on visual basic model optimization
By introducing basic visual model coding and channel attention mechanism in U-Net++ remote sensing semantic segmentation method, combined with the advantages of Transformer and CNN, the problem of insufficient prediction accuracy of semantic segmentation of remote sensing images is solved, and the segmentation effect is significantly improved.
Patent Information
- Application Number
- CN202510296459.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The prior art is difficult to effectively combine the advantages of Transformer and CNN to deeply mine remote sensing image feature information, resulting in insufficient prediction accuracy of semantic segmentation of remote sensing images.
U-Net++ remote sensing semantic segmentation method based on visual basic model optimization is adopted to obtain global semantic information through visual basic coding module, and local spatial information is obtained in combination with U-Net++ model coding module, and feature fusion enhancement is used to utilize channel attention mechanism.
The prediction accuracy of semantic segmentation of remote sensing images was improved, the average recall rate was improved by 0.82%, the average accuracy rate was improved by 1.95%, the average interchange ratio was improved by 1.40%, and the F1 value was increased by 1.54%.
Smart Images

Figure CN119810454B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a U-Net++ remote sensing semantic segmentation method based on visual basic model optimization. Background Art
[0002] Remote sensing image semantic segmentation technology refers to the process of classifying and labeling each pixel in the remote sensing image according to its most likely land feature category, thereby dividing the image into different semantic areas. Traditional deep learning remote sensing semantic segmentation is mainly based on convolutional neural networks (CNNs), among which U-Net++ is widely used due to its excellent network structure and excellent representation in spatial position. However, due to the locality of convolution operations, U-Net++ is difficult to establish the ability of global semantic interaction and long-distance context association. The introduction of Transformer not only enables deep learning to pay more attention to the global information of input data, but also removes invalid information and strengthens valid information by calculating the mapping relationship between input data vectors. However, since it is completely based on the self-attention mechanism, there will be a certain loss of position information, so the semantic segmentation method is still mainly based on the CNN architecture. Until the release of the Visual Basic Model, semantic segmentation methods began to evolve from CNN to Transformer. As a large visual basic model, SAM has attracted widespread attention due to its powerful general segmentation capabilities. However, since SAM's training objects are natural images and it lacks prior knowledge of remote sensing images, it is difficult to deeply mine the feature information of remote sensing images. Therefore, how to effectively combine the advantages of Transform and CNN, and use SAM's powerful general segmentation capabilities to make it suitable for remote sensing image semantic segmentation tasks and improve prediction accuracy is a key issue that needs to be urgently addressed. Summary of the invention
[0003] The purpose of the present invention is to overcome the above problems and provide a U-Net++ remote sensing semantic segmentation method based on visual base model optimization, which can improve prediction accuracy, give full play to the powerful prior knowledge of visual base model, and establish global semantic interaction and long-distance context association.
[0004] To achieve the above objectives, the present invention adopts the following technical solutions.
[0005] A U-Net++ remote sensing semantic segmentation method based on visual basic model optimization of the present invention comprises the following steps:
[0006] S1. Visual basic coding: The remote sensing image is passed through the visual basic coding module to obtain the visual feature map X0. The visual basic coding module includes a block embedding module, a position embedding module, a basic block module, and a neck module, among which:
[0007] The block embedding module segments, flattens, and linearly embeds the remote sensing image to obtain the feature map;
[0008] The position embedding module adds the position vector feature of the feature map processed by the block embedding module to the feature map and inputs it into the basic block module;
[0009] The basic block module includes 16 transformer blocks, which are composed of multi-head attention mechanism MSA and multi-layer perceptron MLP blocks alternately. Normalized LN is used in each transformer block. k Before application, residual connection is applied in each transformer bloc k Afterwards, the MLP is applied, which consists of two layers with GELU nonlinearity;
[0010] The neck module reduces the number of feature maps output by the transformer block from 768 to 256 through two layers of convolution, obtaining the visual feature map X0;
[0011] S2. U-Net++ encoding: The U-Net++ encoding module is formed by stacking 5 basic modules. Each basic module contains two convolution layers. The convolution layer of the first basic module has 16 convolution kernels. The convolution kernels of the convolution layer of each basic module from the second to the fifth basic module are twice the number of convolution kernels of the previous basic module. The convolution layer uses a normalization layer and an activation layer to further extract the deep feature information of the remote sensing image by downsampling the remote sensing image feature map output by the previous basic module as the input of the next basic module. After downsampling, the size of the feature map is reduced by 1 times; the 3-band remote sensing image with a size of 1024×1024 is encoded by U-Net++ to obtain the remote sensing image feature map X 0,0 , remote sensing image feature map X1,0, remote sensing image feature map X2,0, remote sensing image feature map X3,0, remote sensing image feature map X4,0;
[0012] S3, feature fusion enhancement: The feature fusion enhancement module connects the visual feature map X0 obtained in step S1 with the remote sensing image feature map X4,0 obtained in step S2, and then performs feature fusion enhancement on the connected feature maps through the channel attention mechanism. After reprojection integration, the scale of the original feature map is restored to obtain the enhanced feature map X1;
[0013] S4. Decoding: In the first step, the remote sensing image feature map X0,0 is connected with the feature map after upsampling the remote sensing image feature map X1,0 to obtain the feature map X0,1. The remote sensing image feature map X1,0 is connected with the feature map after upsampling the remote sensing image feature map X2,0 to obtain the feature map X1,1. The remote sensing image feature map X2,0 is connected with the feature map after upsampling the remote sensing image feature map X3,0 to obtain the feature map X2,1. The remote sensing image feature map X3,0 is connected with the upsampling feature map X4,0 to obtain the feature map X4,1. The feature map after upsampling the strong feature map X1 is feature-connected to obtain the feature map X3,1; in the second step, the remote sensing image feature map X0,0, feature map X0,1 and the feature map after upsampling the feature map X1,1 are feature-connected to obtain the feature map X0,2, the remote sensing image feature map X1,0, feature map X1,1 and the feature map after upsampling the feature map X2,1 are feature-connected to obtain the feature map X1,2, the remote sensing image feature map X2,0, feature map X2,1 and the feature map X3,1 are feature-connected The sampled feature maps are connected to obtain feature map X2,2; the third step is to connect the remote sensing image feature map X0,0, feature map X0,1, feature map X0,2 and the feature map after upsampling the feature map X1,2 to obtain feature map X0,3, and the remote sensing image feature map X1,0, feature map X1,1, feature map X1,2 and the feature map after upsampling the feature map X2,2 to obtain feature map X1,3; the fourth step is to connect the remote sensing image feature map X0,0, feature map The feature map X0,1, feature map X0,2, feature map X0,3 and the upsampled feature map X1,3 are connected to obtain a feature map X0,4 with a size of 1024×1024. Each pixel value of the feature map X0,4 is multiplied by a 1×1 convolution parameter value to obtain a prediction result X with a size of 1024×1024 and a pixel value between [0-1]. Each pixel value in the prediction result X represents the predicted probability that the pixel is the target object.
[0014] In the above-mentioned U-Net++ remote sensing semantic segmentation method based on visual basic model optimization, the operation steps of the block embedding module described in step S1 are as follows:
[0015] (1) Image segmentation: A 1024×1024 3-band remote sensing image is divided into 64×64 16×16 3-band remote sensing image blocks through a convolution layer with a convolution kernel size of 16×16 and a stride of 16;
[0016] (2) Flattening: flatten each remote sensing image block into a 768 one-dimensional vector;
[0017] (3) Linear embedding: Through a fully connected layer, the flattened remote sensing image block vector is mapped to a high-dimensional space to obtain a 768×64×64 feature map.
[0018] The above-mentioned U-Net++ remote sensing semantic segmentation method based on visual basic model optimization, wherein: the channel attention mechanism described in step S3 includes global information embedding, adaptive recalibration and reweighting, the global information embedding compresses the spatial feature encoding on each channel into a global feature, adopts global average pooling to achieve, and outputs a vector of dimension 1×1×C; after the adaptive recalibration obtains the 1×1×C global feature, a fully connected layer is added, and the importance of each channel is predicted, and the importance of different channels is obtained and then reweighted; the reweighting is the multiplication of the channel weight 1×1×C and the original feature vector W×H×C, and the weight values of each channel calculated by the channel attention mechanism are multiplied by the two-dimensional matrix of the corresponding channel of the original feature map to obtain the output result.
[0019] Compared with the prior art, the present invention has obvious beneficial effects. From the above technical scheme, it can be seen that the present invention adopts a dual encoding module, namely a visual basic model encoding module and a U-Net++ model encoding module, which are respectively used to obtain a visual feature map with global semantic information and a remote sensing image feature map with local spatial information, so as to improve the utilization rate of remote sensing image information. The feature fusion enhancement module of the present invention adopts a channel attention mechanism to learn the dependencies between features in the remote sensing image channel, overcome the limitations of the convolutional neural network in capturing long-distance dependencies, and strengthen U-Net++'s understanding of the relationship between local spatial features and global semantic features, thereby solving the noise interference problem in remote sensing images and improving the robustness of U-Net++; the present invention improves the U-Net++ network structure through the powerful prior knowledge of the visual basic model. Under the same conditions, the method of the present invention improves the average recall rate of the original U-Net++ prediction results by 0.82%, the average accuracy rate by 1.95%, the average intersection-over-union ratio by 1.40%, and the F1 value by 1.54%. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a structural schematic diagram of the present invention;
[0021] Figure 2 It is a schematic diagram of the structure of the visual basic coding module of the present invention;
[0022] Figure 3 It is a schematic diagram of the structure of the U-Net++ encoding module of the present invention;
[0023] Figure 4 It is a structural schematic diagram of the feature fusion enhancement module of the present invention;
[0024] Figure 5 It is a schematic diagram of the structure of a decoding module of the present invention;
[0025] Figure 6This is a visualization diagram of the prediction results of the present invention. DETAILED DESCRIPTION
[0026] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0027] Example:
[0028] The present invention is a U-Net++ remote sensing semantic segmentation method based on visual basic model optimization (see Figure 1 ) includes the following steps:
[0029] S1, Visual Basic Coding: The remote sensing image is passed through the Visual Basic Coding module to obtain the visual feature map X0. The Visual Basic Coding module (see Figure 2 ) includes a block embedding module, a position embedding module, a basic block module, and a neck module, wherein:
[0030] The block embedding module segments, flattens, and linearly embeds the remote sensing image to obtain a feature map. The specific steps are as follows: (1) Image segmentation: a 1024×1024 three-band remote sensing image is divided into 64×64 16×16 three-band remote sensing image blocks through a convolution layer with a convolution kernel size of 16×16 and a stride of 16; (2) Flattening: each remote sensing image block is flattened into a 768 one-dimensional vector; (3) Linear embedding: a fully connected layer is used to map the flattened remote sensing image block vector to a high-dimensional space to obtain a 768×64×64 feature map for the subsequent Transformer encoder;
[0031] The position embedding module is a learnable parameter matrix, which adds the position vector features of the feature map processed by the block embedding module to the feature map and then inputs it into the basic block module;
[0032] The basic block module includes 16 transformer blocks, which are composed of multi-head attention mechanism MSA and multi-layer perceptron MLP blocks alternately. Normalized LN is used in each transformer block. k Before application, residual connection is applied in each transformer bloc k After that, the MLP consists of two layers with GELU nonlinearity, which is expressed as:
[0033]
[0034]
[0035]
[0036] Where, X 1 represents the original features of the input basic block module, X0 represents the output results of the features after different transformer blocks, ;
[0037] The neck module reduces the number of feature maps output by the transformer block from 768 to 256 through two layers of convolution to obtain a visual feature map. ;
[0038] S2, U-Net++ encoding: U-Net++ encoding module (see Figure 3 ) is formed by stacking 5 basic modules, which are used for remote sensing image feature map X 0,0 , remote sensing image feature map X1,0, remote sensing image feature map X2,0, remote sensing image feature map X3,0, remote sensing image feature map X4,0 extraction, each basic module contains two convolution layers, the convolution layer of the first basic module is 16 convolution kernels, the convolution kernels of the convolution layer of each basic module from the second to the fifth basic modules are twice the number of convolution kernels of the previous basic module, the convolution layer uses the normalization (BN) layer and the activation (LeakyReLu) layer, the remote sensing image feature map output by the previous basic module is further extracted through downsampling (MaxPool2d) The deep feature information of the remote sensing image is used as the input of the next basic module. After downsampling, the size of the feature map is reduced by 1 times; the 3-band remote sensing image with a size of 1024×1024 is encoded by U-Net++ and obtained in turn , , , , ;
[0039] S3. Feature Fusion Enhancement: The feature fusion enhancement module fuses the visual feature map extracted by the visual basic encoding module and the remote sensing image feature map extracted by the U-Net++ encoding module, so that the improved U-Net++ can understand the interdependence between the feature maps in the remote sensing image channel, so as to build the global and local relationship between the feature maps. Feature Fusion Enhancement Module (see Figure 4 ) Connect the visual feature map X0 obtained in step S1 with the remote sensing image feature map X4,0 obtained in step S2, and then perform feature fusion enhancement on the connected feature maps through the channel attention mechanism. After reprojection and integration, restore the original feature map scale to obtain the enhanced feature map X1. The process is expressed as:
[0040]
[0041]
[0042] represents connection, SE represents channel attention mechanism, represents reprojection, X1 is the integrated feature map;
[0043] The channel attention mechanism includes global information embedding, adaptive recalibration and reweighting. The global information embedding compresses the spatial feature encoding on each channel into a global feature, which is implemented by global average pooling, and the output dimension is a vector of 1×1×C; after the adaptive recalibration obtains the 1×1×C global feature, a fully connected layer is added, and the importance of each channel is predicted, and the importance of different channels is obtained before reweighting; the reweighting is the multiplication of the channel weight 1×1×C and the original feature vector W×H×C, and the weight values of each channel calculated by the channel attention mechanism are multiplied by the two-dimensional matrix of the corresponding channel of the original feature map to obtain the output result;
[0044] S4. Decoding: The decoding module upsamples the remote sensing image feature map X1,0, the remote sensing image feature map X2,0, the remote sensing image feature map X3,0 and the enhanced feature map X1 through the upsampling of the U-Net++ decoder to obtain feature maps, and uses dense feature connection and jump feature connection to connect the upsampled feature maps with the feature maps of the same size in the decoding module to achieve deep fusion of feature maps; each upsampling will double the size of the feature map to retain more contextual information and detail features, including the following four steps: The first step is to connect the remote sensing image feature map X0,0 with the remote sensing image feature map X1,0 after upsampling. The feature map X0,1 is obtained by feature connection. The remote sensing image feature map X1,0 is connected with the feature map after upsampling the remote sensing image feature map X2,0 to obtain the feature map X1,1. The remote sensing image feature map X2,0 is connected with the feature map after upsampling the remote sensing image feature map X3,0 to obtain the feature map X2,1. The remote sensing image feature map X3,0 is connected with the feature map after upsampling the enhanced feature map X1 to obtain the feature map X3,1. In the second step, the remote sensing image feature map X0,0, the feature map X0,1 and the feature map after upsampling the feature map X1,1 are connected to obtain the feature map X0,2. The feature map X1,0, the feature map X1,1 and the feature map X2,1 are sampled and connected to obtain the feature map X1,2. The remote sensing image feature map X2,0, the feature map X2,1 and the feature map X3,1 are sampled and connected to obtain the feature map X2,2. In the third step, the remote sensing image feature map X0,0, the feature map X0,1, the feature map X0,2 and the feature map X1,2 are sampled and connected to obtain the feature map X0,3. The remote sensing image feature map X1,0, the feature map X1,1, the feature map X1,2 and the feature map X2,2 are sampled and connected to obtain the feature map X2,2. Feature map X1,3; Step 4: Connect the remote sensing image feature map X0,0, feature map X0,1, feature map X0,2, feature map X0,3 and the upsampled feature map X1,3 to obtain a feature map X0,4 of size 1024×1024. Through a 1×1 convolution (the 1×1 convolution parameter is adaptively updated during the training process), multiply each pixel value of the feature map X0,4 by the 1×1 convolution parameter value to obtain a prediction result X of size 1024×1024 with a pixel value between [0-1]. Each pixel value in the prediction result X represents the predicted probability that the pixel is the target object. Dense feature connection and jump feature connection not only improve the network's ability to capture complex structures and detailed information, but also improve segmentation accuracy. The process is expressed as:
[0045]
[0046]
[0047]
[0048] in represents reprojection, represents feature connection, represents upsampling, represents a 1×1 convolution, , .
[0049] Comparison of model training and evaluation experiments
[0050] The water surface on the remote sensing image is labeled to obtain the water surface vector sample, and the water surface raster label is obtained after binarization processing. The pixel value of the water surface labeled area in the label is 1, and the pixel value of other areas is 0; the remote sensing image and the water surface raster label are sliced to obtain 1024×1024 three-band remote sensing image slices and water surface label slices; the obtained slices are divided into training set, validation set, and test set in a ratio of 6:2:2 to obtain the semantic segmentation sample data set;
[0051] Use the training set and validation set to train the remote sensing image semantic segmentation model, freeze the visual basic encoding module, train the remaining modules in the network model, and use the Adam optimization function to update the model parameters. The core formula of the Adam optimization function is as follows:
[0052]
[0053] in is the bias-corrected first-order moment estimate (mean), is the second moment estimate (variance), represents the learning rate, represents the weight decay parameter, To update the model parameters, Represents a very small positive number to prevent the denominator from being 0.
[0054] The remote sensing water surface semantic segmentation dataset used has 4282 training sets and 1428 validation sets. A total of 20 rounds of training are used, and a mixed segmentation loss is used, which is expressed as:
[0055]
[0056] In the formula, C is the number of classes observed in a given dataset, , represents the target labels and predicted probabilities of c classes and n pixels in the batch, and N represents the number of pixels in a batch.
[0057] The embodiment of the present invention was evaluated on 1428 test samples, with an average recall rate of 73.39%, an average accuracy rate of 58.33%, an average intersection-over-union ratio of 56.25%, and an F1 value of 65.00%; while the average recall rate of the original U-Net++ was 72.57%, the average accuracy rate was 56.38%, the average intersection-over-union ratio was 54.85, and the F1 value was 63.46%. Compared with before the improvement, the average recall rate of the improved method proposed in the present invention increased by 0.82%, the average accuracy rate increased by 1.95%, the average intersection-over-union ratio increased by 1.40%, and the F1 value increased by 1.54%. Figure 6 This is a visualization diagram of the prediction results, which includes, from left to right, the predicted image, the artificially drawn real water surface, the water surface predicted by the embodiment of the present invention, and the water surface predicted by the original U-Net++.
[0058] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still falls within the scope of the technical solution of the present invention.
Claims
1. A U-Net++ remote sensing semantic segmentation method based on visual basic model optimization, characterized by The following steps are involved: S1. Visual basic coding: The remote sensing image is passed through the visual basic coding module to obtain the visual feature map X0. The visual basic coding module includes a block embedding module, a position embedding module, a basic block module, and a neck module, among which: The block embedding module segments, flattens, and linearly embeds the remote sensing image to obtain the feature map; The position embedding module adds the position vector feature of the feature map processed by the block embedding module to the feature map and inputs it into the basic block module; The basic block module includes 16 transformer blocks, which are composed of multi-head attention mechanism MSA and multi-layer perceptron MLP blocks alternately. Normalized LN is used in each transformer block. k Before application, residual connection is applied in each transformer bloc k Afterwards, the MLP is applied, which consists of two layers with GELU nonlinearity; The neck module reduces the number of feature maps output by the transformer block from 768 to 256 through two layers of convolution, obtaining the visual feature map X0; S2. U-Net++ encoding: The U-Net++ encoding module is formed by stacking 5 basic modules. Each basic module contains two convolution layers. The convolution layer of the first basic module has 16 convolution kernels. The convolution kernels of the convolution layer of each basic module from the second to the fifth basic module are twice the number of convolution kernels of the previous basic module. The convolution layer uses a normalization layer and an activation layer to further extract the deep feature information of the remote sensing image by downsampling the remote sensing image feature map output by the previous basic module as the input of the next basic module. After downsampling, the size of the feature map is reduced by 1 times; the 3-band remote sensing image with a size of 1024×1024 is encoded by U-Net++ to obtain the remote sensing image feature map X 0,0 , remote sensing image feature map X1,0, remote sensing image feature map X2,0, remote sensing image feature map X3,0, remote sensing image feature map X4,0; S3, feature fusion enhancement: The feature fusion enhancement module connects the visual feature map X0 obtained in step S1 with the remote sensing image feature map X4,0 obtained in step S2, and then performs feature fusion enhancement on the connected feature maps through the channel attention mechanism. After reprojection integration, the scale of the original feature map is restored to obtain the enhanced feature map X1; S4. Decoding: In the first step, the remote sensing image feature map X0,0 is connected with the feature map after upsampling the remote sensing image feature map X1,0 to obtain the feature map X0,1. The remote sensing image feature map X1,0 is connected with the feature map after upsampling the remote sensing image feature map X2,0 to obtain the feature map X1,1. The remote sensing image feature map X2,0 is connected with the feature map after upsampling the remote sensing image feature map X3,0 to obtain the feature map X2,1. The remote sensing image feature map X3,0 is connected with the upsampling feature map X4,0 to obtain the feature map X4,1. The feature map after upsampling the strong feature map X1 is feature-connected to obtain the feature map X3,1; in the second step, the remote sensing image feature map X0,0, feature map X0,1 and the feature map after upsampling the feature map X1,1 are feature-connected to obtain the feature map X0,2, the remote sensing image feature map X1,0, feature map X1,1 and the feature map after upsampling the feature map X2,1 are feature-connected to obtain the feature map X1,2, the remote sensing image feature map X2,0, feature map X2,1 and the feature map X3,1 are feature-connected The sampled feature maps are connected to obtain feature map X2,2; the third step is to connect the remote sensing image feature map X0,0, feature map X0,1, feature map X0,2 and the feature map after upsampling the feature map X1,2 to obtain feature map X0,3, and the remote sensing image feature map X1,0, feature map X1,1, feature map X1,2 and the feature map after upsampling the feature map X2,2 to obtain feature map X1,3; the fourth step is to connect the remote sensing image feature map X0,0, feature map The feature map X0,1, feature map X0,2, feature map X0,3 and the upsampled feature map X1,3 are connected to obtain a feature map X0,4 with a size of 1024×1024. Each pixel value of the feature map X0,4 is multiplied by a 1×1 convolution parameter value to obtain a prediction result X with a size of 1024×1024 and a pixel value between [0-1]. Each pixel value in the prediction result X represents the predicted probability that the pixel is the target object.
2. The U-Net++ remote sensing semantic segmentation method based on visual basic model optimization according to claim 1, characterized in that: The operation steps of the block embedding module described in step S1 are: (1) Image segmentation: A 1024×1024 3-band remote sensing image is divided into 64×64 16×16 3-band remote sensing image blocks through a convolution layer with a convolution kernel size of 16×16 and a stride of 16; (2) Flattening: flatten each remote sensing image block into a 768 one-dimensional vector; (3) Linear embedding: Through a fully connected layer, the flattened remote sensing image block vector is mapped to a high-dimensional space to obtain a 768×64×64 feature map.
3. The U-Net++ remote sensing semantic segmentation method based on visual basic model optimization according to claim 1, characterized in that: The channel attention mechanism described in step S3 includes global information embedding, adaptive recalibration and reweighting. The global information embedding compresses the spatial feature encoding on each channel into a global feature, which is implemented by global average pooling, and the output dimension is a vector of 1×1×C; after the adaptive recalibration obtains the 1×1×C global feature, a fully connected layer is added, and the importance of each channel is predicted, and the importance of different channels is obtained before reweighting; the reweighting is the multiplication of the channel weight 1×1×C and the original feature vector W×H×C, and the weight values of each channel calculated by the channel attention mechanism are multiplied by the two-dimensional matrix of the corresponding channel of the original feature map to obtain the output result.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method based on context information and attention mechanism
CN110197182A
Image semantic segmentation method based on Transform visual upsampling module
CN113888744A