Remote Sensing Image Semantic Segmentation Method Based on Boundary Information Guided Multi-Information Fusion
Through the BGFNet three-branch network structure, the feature fusion of boundary branches and spatial branches is used to solve the problem of insufficient utilization of remote sensing image boundary information, which improves segmentation accuracy and reduces the computational complexity.
Patent Information
- Application Number
- CN202310756109.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-06-26
AI Technical Summary
The boundary information in semantic segmentation of remote sensing images does not fully play a role, the calculation complexity is high, and there is information redundancy in existing network designs.
The three-branch network structure BGFNet is adopted, and the boundary branch, semantic branch and spatial branch are used to integrate boundary features and spatial characteristics through the bidirectional boundary gated module and the boundary guide linear attention module, and predict it in combination with the channel attention mechanism.
The accuracy of semantic segmentation of remote sensing images is improved, while the computational complexity is reduced, and the balance between time and accuracy is achieved.
Smart Images

Figure CN116797792B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image semantic segmentation, and particularly to a remote sensing image semantic segmentation method based on boundary information-guided multi-information fusion. Background Art
[0002] The remote sensing image semantic segmentation task refers to extracting, encoding, and decoding the feature information of a given remote sensing image through a specific neural network, and performing pixel-level prediction to obtain the distribution position and range of each category in the image. Affected by its imaging target and characteristics, remote sensing images have the characteristics of complex target boundary feature information that is not easy to fit, high similarity between multiple categories, and large differences within a single category. At the same time, the imbalance in the distribution of remote sensing image data poses a great challenge to the feature extraction and encoding / decoding capabilities of neural networks.
[0003] The semantic segmentation algorithm model extracts feature information through methods such as convolution, and each pixel obtains a corresponding semantic label through encoding and decoding. The fully convolutional network (FCN, 2015, Long J, et al.) realizes end-to-end segmentation by using convolutional layers to replace fully connected layers, and is no longer restricted by the size of the input image. U-net (2015, Ronneberger O, et al.) proposed a general network structure of U-shaped encoding-decoding on this basis, which can make more full use of the encoded information in the decoding stage, greatly improving the performance of accurate recognition and positioning of small targets. However, the continuous reduction of the size of the feature map during the encoding process has a certain impact on the final classification result. Therefore, improving the multi-scale feature extraction ability of the encoder part has become the research focus. For example, DeepLabV2 (2017, Chen L C, et al.) extracts and fuses features in parallel through dilated convolutions with multiple different sampling rates, thereby improving the network's feature extraction ability. PSPNet (2017, Zhao H, et al.) obtains feature maps of different regions through pooling layers at multiple different levels, and splices feature maps of different scales to achieve multi-scale feature fusion. BiseNetv1 (2018, Changqian Yu, et al.) proposed to adopt a dual-branch structure. One is a spatial information path with a small stride, which retains spatial information to generate high-resolution feature maps; the other is a semantic information path with fast downsampling to obtain deep context information. And a new feature fusion module is designed to achieve a balance between speed and accuracy. STDCNet (2021, Mingyuan Fan, et al.) designed a new module more suitable for the semantic segmentation task on the basis of BiseNetv1. By gradually reducing the number of channels and restoring the number of channels of the feature map by splicing at the end of the module to complete feature extraction. At the same time, in order to make up for the loss caused by removing the spatial information branch, an auxiliary segmentation head is added after the third stage of the network to enhance the network's ability to learn detailed information.
[0004] Although the existing research results have achieved certain results in the semantic segmentation of remote sensing images, the accuracy is not high and the computational complexity is relatively high in the aspects of complex boundary information and small target segmentation of remote sensing images.
[0005] In PIDNet, the boundary information only interacts with other information at the final stage of the network and does not fully play its due role. At the same time, the module design of the micro-branch and ratio-branch leads to a large amount of information redundancy. Summary of the Invention
[0006] Based on this, it is necessary to provide a semantic segmentation method for remote sensing images that can use boundary information to guide multi-information fusion to address the above technical problems.
[0007] A remote sensing image semantic segmentation method based on boundary information-guided multi-information fusion, the method comprising:
[0008] Obtain a remote sensing image, and annotate the remote sensing image to obtain training samples.
[0009] Construct a remote sensing image semantic segmentation network, the remote sensing image semantic segmentation network including an input network, a feature extraction network, and a segmentation network; the input network is used to extract shallow features of the training samples by using convolution operations; the feature extraction network includes a boundary branch, a semantic branch, and a spatial branch, each branch including three stages, each stage of the boundary branch including a bidirectional boundary gating module, the spatial branch and the semantic branch including stacked convolution modules and class bottleneck structure modules, and a boundary-guided linear attention module further included after each stage in the spatial branch; the feature extraction network is used to extract boundary features and semantic features of each stage of the shallow features by using the boundary branch and the semantic branch respectively, and fuse the boundary features and semantic features of this stage into the features extracted by the spatial branch of this stage by using the BGA to obtain the spatial features of this stage; the segmentation network is used to make predictions based on the shallow features and the spatial features of three stages to obtain a predicted semantic segmentation result.
[0010] Train the remote sensing image semantic segmentation network according to the annotation of the training samples and the predicted semantic segmentation result obtained after inputting the training samples into the remote sensing image semantic segmentation network to obtain a trained remote sensing image semantic segmentation network.
[0011] Input the obtained remote sensing image to be segmented into the trained remote sensing image semantic segmentation network to obtain a semantic segmentation result of the remote sensing image to be segmented.
[0012] In one embodiment, the input network includes a Stem module and two stacked convolution modules; training the remote sensing image semantic segmentation network according to the training samples and the predicted semantic segmentation result obtained after inputting the training samples into the remote sensing image semantic segmentation network to obtain a trained remote sensing image semantic segmentation network, including:
[0013] Input the training samples into the input network to extract features of different dimensions through the Stem module, and perform feature extraction on the obtained features through two stacked convolution modules with different strides to obtain shallow features.
[0014] Input the shallow features into the feature extraction network to obtain the spatial features of each stage.
[0015] Input the shallow features and the spatial features of each stage into the segmentation network to obtain a predicted semantic segmentation result.
[0016] Backward train the remote sensing image semantic segmentation network according to the annotation of the training samples and the predicted semantic segmentation result to obtain a trained remote sensing image semantic segmentation network.
[0017] In one embodiment, input the shallow features into the feature extraction network to obtain spatial features at each stage, including:
[0018] Input the shallow features into the boundary branch to obtain first-stage boundary features, second-stage boundary features, and third-stage boundary features.
[0019] Input the shallow features into the semantic branch to obtain first-stage semantic features, second-stage semantic features, and third-stage boundary features.
[0020] Input the shallow features into the first stage of the spatial branch to obtain first intermediate spatial features.
[0021] Input the first intermediate spatial features, the first-stage semantic features, and the first-stage boundary features into the first boundary-guided linear attention module to obtain first-stage spatial features.
[0022] Input the first-stage spatial features into the second stage of the spatial branch to obtain second intermediate spatial features; input the second intermediate spatial features, the second-stage semantic features, and the second-stage boundary features into the second boundary-guided linear attention module to obtain second-stage spatial features.
[0023] Input the second-stage spatial features into the third stage of the spatial branch to obtain third intermediate spatial features; input the third intermediate spatial features, the third-stage semantic features, and the third-stage boundary features into the third boundary-guided linear attention module to obtain third-stage spatial features.
[0024] In one embodiment, each stage of the boundary branch includes a bidirectional boundary gating module; the bidirectional boundary gating module includes: a convolutional layer, a depthwise separable convolutional layer, and a gating mechanism; input the shallow features into the boundary branch to obtain first-stage boundary features, second-stage boundary features, and third-stage boundary features, including:
[0025] Input the shallow features into the bidirectional boundary gating module of the first stage of the boundary branch, and the first-stage boundary features obtained are:
[0026] x out = GELU(Cat(f x (x in ),fy (x in )))*f(x in )+x in
[0027] Among them, x in represents the input shallow feature, and x out represents the output of the bidirectional boundary gating module. f(·), f x (·), f y (·) respectively represent the extraction of conventional, horizontal, and vertical convolutional boundary information. f(x in ) = DWConv 3×3 (x in ), f x (x in ) = DWConv 3×1 (Conv 1×1 (x in ))), f y (x in ) = DWConv 1×3 (Conv 1×1 (x in )). Cat(·) represents the concatenation operation, and GELU(·) represents the gating mechanism.
[0028] Input the first-stage boundary feature into the first stage of the boundary branch to obtain the second-stage boundary feature.
[0029] Input the second-stage boundary feature into the first stage of the boundary branch to obtain the third-stage boundary feature.
[0030] In one embodiment, each of the first and second stages of the spatial branch includes a stacked convolutional module, and the third stage of the spatial branch includes a bottleneck-like structure module; input the shallow feature into the first stage of the spatial branch to obtain the first intermediate spatial feature, including:
[0031] Input the shallow feature into the stacked convolutional module in the first stage of the spatial branch to obtain the first intermediate spatial feature as:
[0032] x out ′ = Conv 3×3 (Conv 3×3 (x in )) + x in
[0033] Among them, x out ′ represents the first intermediate spatial feature, and x in represents the input shallow feature.
[0034] In one embodiment, inputting the second-stage spatial feature into the third stage of the spatial branch to obtain a third intermediate spatial feature, including:
[0035] Inputting the second-stage spatial feature into the third stage of the spatial branch to obtain a third intermediate spatial feature as:
[0036] x out ″ = Conv 1×1 (Conv 3×3 (Conv 1×1 (x in ′)))+x in ′
[0037] where x out ″ represents the third intermediate spatial feature, and x in ′ represents the second-stage spatial feature.
[0038] In one embodiment, the boundary-guided linear attention module includes: a convolutional layer, a Softmax function, and a Sigmoid function.
[0039] Inputting the first intermediate spatial feature, the first-stage semantic feature, and the first-stage boundary feature into the first boundary-guided linear attention module to obtain a first-stage spatial feature, including:
[0040] Inputting the first intermediate spatial feature, the first-stage semantic feature, and the first-stage boundary feature into the first boundary-guided linear attention module to obtain a first-stage spatial feature as:
[0041] x out ″′ = attn * x s + x s
[0042] attn = Soffmax(Resize(f c (x c )) + f b (x b )) * (1 - σ)
[0043] σ = Sigmoid(f b (x b ))
[0044] where x s , x b , x c respectively represent the first intermediate spatial feature, the first-stage boundary feature, and the first-stage semantic feature, x out ″′ is the first-stage spatial feature, and f b (·), fc (·) are the mappings of the input boundary information and spatial information respectively, σ is the score of the input boundary information, attn is the attention score calculated based on the semantic information and boundary information, and Resize(·) is an image size adjustment function that upsamples the output size of the semantic information to the same size as the boundary information size.
[0045] In one embodiment, the spatial features of each stage include the first node spatial feature, the second stage spatial feature, and the third stage spatial feature.
[0046] The segmentation network includes a spatial segmentation head and a channel attention mechanism.
[0047] Input the shallow features and the spatial features of each stage into the segmentation network to obtain the predicted semantic segmentation result, including:
[0048] Concatenate the shallow features, the first node spatial feature, the second stage spatial feature, and the third stage spatial feature and input them into the channel attention mechanism for feature fusion to obtain the fused feature.
[0049] Predict the fused feature through the spatial segmentation head to obtain the predicted semantic segmentation result.
[0050] In one embodiment, the segmentation network further includes a boundary auxiliary segmentation head and a spatial auxiliary segmentation head.
[0051] The boundary auxiliary segmentation head and the spatial auxiliary segmentation head are in the second stage of the boundary branch and the spatial branch of the feature extraction network.
[0052] Perform backpropagation training on the remote sensing image semantic segmentation network according to the annotation of the training sample and the predicted semantic segmentation result to obtain a trained remote sensing image semantic segmentation network, including:
[0053] Construct a total loss function, which is the weighted sum of the boundary auxiliary segmentation head loss function, the spatial auxiliary segmentation head loss function, and the spatial segmentation head loss function; the spatial auxiliary segmentation head loss function and the spatial segmentation head loss function adopt the cross-entropy loss function, and the boundary auxiliary segmentation head loss function adopts the binary cross-entropy loss function.
[0054] According to the annotation of the training sample, the predicted semantic segmentation result, and the total loss function, perform backpropagation training on the remote sensing image semantic segmentation network to obtain a trained remote sensing image semantic segmentation network.
[0055] In one embodiment, both the first stage and the second stage of the semantic branch include n stacked convolutional modules, and the third stage includes two bottleneck-like structure modules and one parallel aggregated pyramid pooling module.
[0056] The above-mentioned remote sensing image semantic segmentation method based on boundary information-guided multi-information fusion includes: obtaining a remote sensing image and annotating the remote sensing image to obtain training samples; constructing a remote sensing image semantic segmentation network, which includes an input network, a feature extraction network, and a segmentation network; the input network is used to extract shallow features of the training samples by means of convolutional operations; the feature extraction network includes a boundary branch, a semantic branch, and a spatial branch, and each branch includes three stages. Each stage of the boundary branch includes a bidirectional boundary gating module. The spatial branch and the semantic branch include stacked convolutional modules and bottleneck-like structure modules. A boundary-guided linear attention module is further included after each stage in the spatial branch; the feature extraction network is used to extract the boundary features and semantic features of each stage of the shallow features by means of the boundary branch and the semantic branch respectively, and fuse the boundary features and semantic features of this stage into the features extracted by the spatial branch of this stage by using BGA to obtain the spatial features of this stage; the segmentation network is used to make predictions based on the shallow features and the spatial features of the three stages to obtain a predicted semantic segmentation result; training the remote sensing image semantic segmentation network according to the annotation of the training samples and the predicted semantic segmentation result obtained after inputting the training samples into the remote sensing image semantic segmentation network to obtain a trained remote sensing image semantic segmentation network; inputting the obtained remote sensing image to be segmented into the trained remote sensing image semantic segmentation network to obtain the semantic segmentation result of the remote sensing image to be segmented. Using this method for remote sensing image semantic segmentation can improve the accuracy of semantic segmentation while reducing the computational complexity. Description of the Drawings
[0057] Figure 1 It is a schematic flowchart of the remote sensing image semantic segmentation method based on boundary information-guided multi-information fusion in one embodiment;
[0058] Figure 2 It is a schematic diagram of the remote sensing image semantic segmentation network structure in another embodiment;
[0059] Figure 3 It is a schematic diagram of the Stem module structure in another embodiment;
[0060] Figure 4 It is a schematic diagram of two stacked convolutional modules in another embodiment;
[0061] Figure 5 It is a schematic diagram of the bidirectional boundary gating module structure in another embodiment;
[0062] Figure 6Schematic diagram of the stacked convolutional module structure in another embodiment;
[0063] Figure 7 Schematic diagram of the bottleneck-like structure module in another embodiment;
[0064] Figure 8 Schematic diagram of the boundary-guided linear attention module structure in another embodiment;
[0065] Figure 9 Schematic diagram of the channel attention mechanism structure in another embodiment;
[0066] Figure 10 Schematic diagram of the parallel aggregation pyramid pooling module structure in another embodiment. Detailed implementation manners
[0067] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0068] Boundary-Guided Linear Attention Block (abbreviation: BGA block).
[0069] Bidirectional Boundary Gate Block (abbreviation: BBG block).
[0070] In one embodiment, as Figure 1 shown, a remote sensing image semantic segmentation method based on boundary information-guided multi-information fusion is provided, and the method includes the following steps:
[0071] Step 100: Obtain a remote sensing image, and annotate the remote sensing image to obtain training samples.
[0072] Specifically, the remote sensing image is an image including five land cover categories such as buildings, farmland, forests, grasslands and water areas.
[0073] Step 102: Construct a remote sensing image semantic segmentation network.
[0074] The remote sensing image semantic segmentation network includes an input network, a feature extraction network and a segmentation network.
[0075] The input network is used to extract shallow features of the training samples by using convolutional operations.
[0076] The feature extraction network includes a boundary branch, a semantic branch, and a spatial branch. Each branch consists of three stages. Each stage of the boundary branch includes a bidirectional boundary gating module. The spatial branch and the semantic branch include stacked convolution modules and class bottleneck structure modules. A boundary-guided linear attention module is also included after each stage in the spatial branch. The feature extraction network is used to extract the boundary features and semantic features of each stage of the shallow features respectively by the boundary branch and the semantic branch, and fuse the boundary features and semantic features of this stage into the features extracted by the spatial branch of this stage using BGA to obtain the spatial features of this stage.
[0077] The segmentation network is used to make predictions based on the shallow features and the spatial features of three stages to obtain the predicted semantic segmentation result.
[0078] Specifically, the remote sensing image semantic segmentation network is a three-branch network structure, abbreviated as: BGFNet.
[0079] In PIDNet, the boundary information only interacts with other information at the final stage of the network and does not fully play its due role. At the same time, the module design of the micro-branch and the ratio branch leads to a large amount of information redundancy. Therefore, a new three-branch network structure BGFNet for remote sensing image semantic segmentation network is proposed. The structure and parameter settings of the remote sensing image semantic segmentation network are as Figure 2 shown. The boundary branch uses a bidirectional boundary gating module (BBG module) to extract boundary features. The semantic branch and the spatial branch adopt the same modules as PIDNet. To retain more detailed information of the image to a greater extent, the spatial branch and the boundary branch do not downsample the feature maps. At the same time, only one module is used to extract feature information at each stage. After each stage, a boundary-guided linear attention module (BGA module) is used to fuse the features output by the spatial branch and the semantic branch at this stage into the spatial branch. Finally, the outputs of the spatial branch at each stage are fused through concatenation and channel attention mechanism and predicted through a segmentation head to output a mask image.
[0080] The bidirectional boundary gating module solves the problems of complex boundary information and high difficulty in small target segmentation in remote sensing images by enhancing the boundary information extraction ability.
[0081] By extracting different types of feature information in the image in three ways, a three-branch network structure BGFNet is designed, achieving a relative balance between time and accuracy; it is proposed to fuse features according to the characteristics of different branch feature extractions at each stage of the network to give full play to the role of boundary information, and a new boundary-guided attention fusion module is designed.
[0082] Step 104: Train the remote sensing image semantic segmentation network according to the annotation of the training samples and the predicted semantic segmentation results obtained by inputting the training samples into the remote sensing image semantic segmentation network, so as to obtain a trained remote sensing image semantic segmentation network.
[0083] Step 106: Input the obtained remote sensing image to be segmented into the trained remote sensing image semantic segmentation network to obtain the semantic segmentation result of the remote sensing image to be segmented.
[0084] In the above method for remote sensing image semantic segmentation guided by boundary information for multi-information fusion, the method includes: obtaining a remote sensing image and annotating the remote sensing image to obtain training samples; constructing a remote sensing image semantic segmentation network, where the remote sensing image semantic segmentation network includes an input network, a feature extraction network, and a segmentation network; the input network is used to extract the shallow features of the training samples by using convolutional operations; the feature extraction network includes a boundary branch, a semantic branch, and a spatial branch, each branch includes three stages, each stage of the boundary branch includes a bidirectional boundary gating module, and the spatial branch and the semantic branch include a stacked convolutional module and a class bottleneck structure module, and a boundary-guided linear attention module is further included after each stage in the spatial branch; the feature extraction network is used to extract the boundary features and semantic features of each stage of the shallow features by using the boundary branch and the semantic branch respectively, and fuse the boundary features and semantic features of this stage into the features extracted by the spatial branch of this stage by using BGA to obtain the spatial features of this stage; the segmentation network is used to make predictions based on the shallow features and the spatial features of the three stages to obtain the predicted semantic segmentation results; train the remote sensing image semantic segmentation network according to the annotation of the training samples and the predicted semantic segmentation results obtained by inputting the training samples into the remote sensing image semantic segmentation network to obtain a trained remote sensing image semantic segmentation network; input the obtained remote sensing image to be segmented into the trained remote sensing image semantic segmentation network to obtain the semantic segmentation result of the remote sensing image to be segmented. Using this method for remote sensing image semantic segmentation can improve the accuracy of semantic segmentation and reduce the computational complexity at the same time.
[0085] In one embodiment, the input network includes a Stem module and two stacked convolutional modules; Step 104 specifically includes the following steps:
[0086] Step 200: Input the training samples into the input network to extract features of different dimensions through the Stem module, and perform feature extraction on the obtained features through two stacked convolutional modules with different strides to obtain shallow features.
[0087] Specifically, the Stem module structure in Bisenetv2 is adopted. The input is downsampled by a factor of 1 / 2 through a convolution with a stride of 2, and then passes through a parallel convolution module and a max-pooling module to extract features of different dimensions, reducing information loss during the downsampling process. After concatenating the outputs of the two branches and passing through a 1×1 convolution (replacing the 3×3 convolution with a 1×1 convolution to reduce the number of parameters and computational complexity), a feature map of 1 / 4 is obtained. The Stem module structure is as shown in Figure 3 shown.
[0088] The strides of the two stacked convolution modules are different. Shallow features of the remote sensing image are extracted through the two stacked convolution modules with different strides for subsequent operations and the superposition of receptive fields. The structures of the two stacked convolution modules are as shown in Figure 4 shown.
[0089] Step 202: Input the shallow features into the feature extraction network to obtain spatial features at each stage.
[0090] Step 204: Input the shallow features and the spatial features at each stage into the segmentation network to obtain the predicted semantic segmentation results.
[0091] Step 206: Perform backpropagation training on the remote sensing image semantic segmentation network according to the annotations of the training samples and the predicted semantic segmentation results to obtain a trained remote sensing image semantic segmentation network.
[0092] In one embodiment, Step 202 specifically includes the following steps:
[0093] Step 300: Input the shallow features into the boundary branch to obtain the first-stage boundary features, the second-stage boundary features, and the third-stage boundary features.
[0094] Step 302: Input the shallow features into the semantic branch to obtain the first-stage semantic features, the second-stage semantic features, and the third-stage boundary features.
[0095] Step 304: Input the shallow features into the first stage of the spatial branch to obtain the first intermediate spatial features.
[0096] Step 306: Input the first intermediate spatial features, the first-stage semantic features, and the first-stage boundary features into the first boundary-guided linear attention module to obtain the first-stage spatial features.
[0097] Step 308: Input the first-stage spatial features into the second stage of the spatial branch to obtain the second intermediate spatial features; input the second intermediate spatial features, the second-stage semantic features, and the second-stage boundary features into the second boundary-guided linear attention module to obtain the second-stage spatial features.
[0098] Step 310: Input the second-stage spatial features into the third stage of the spatial branch to obtain the third intermediate spatial features; input the third intermediate spatial features, the third-stage semantic features, and the third-stage boundary features into the third boundary-guided linear attention module to obtain the third-stage spatial features.
[0099] In one embodiment, each stage of the boundary branch includes a bidirectional boundary gating module; the bidirectional boundary gating module includes: a convolutional layer, a depthwise separable convolutional layer, and a gating mechanism; Step 300 includes: inputting the shallow features into the bidirectional boundary gating module of the first stage of the boundary branch to obtain the first-stage boundary features as:
[0100] x out = GELU(Cat(f x (x in ), f y (x in )))*f(x in ) + x in (1)
[0101] where x in represents the input shallow features, x out represents the output of the bidirectional boundary gating module, f(·), f x (·), f y (·) respectively represent the extraction of conventional, horizontal, and vertical convolutional boundary information, f(x in ) = DWConv 3×3 (x in ), f x (x in ) = DWConv 3×1 (Conv 1×1 (x in ))), f y (x in ) = DWConv 1×3 (Conv 1×1 (x in ))), Cat(·) represents the concatenation operation, and GELU(·) represents the gating mechanism.
[0102] Input the first-stage boundary features into the first stage of the boundary branch to obtain the second-stage boundary features; input the second-stage boundary features into the first stage of the boundary branch to obtain the third-stage boundary features.
[0103] Specifically, affected by the density of the semantic segmentation task, there are high requirements for the small object recognition ability of the network, and at the same time, the boundary information plays a crucial role in feature fusion.
[0104] As a differential-based operator in the classical image processing field, the Sobel edge detection operator can detect the positions with the fastest gray-level changes in an image and is commonly used for image edge detection. The Sobel operator is divided into horizontal and vertical directions and is calculated using the following two templates respectively:
[0105] Horizontal direction Vertical direction
[0106] Inspired by it, a new boundary information extraction module, the bidirectional boundary gating module (BBG module), is proposed. While performing conventional convolution operations, a branch is paralleled to enhance the network's ability to extract boundary information. In this branch, depthwise separable convolutions with sizes of 3×1 and 1×3 are used to extract the boundary information in the horizontal and vertical directions of the feature map. At the same time, the excellent feature information selection ability of the gating mechanism is utilized to fuse the information of the two branches. The structure of the bidirectional boundary gating module is as shown in Figure 5 Shown.
[0107] In one embodiment, each of the first and second stages of the spatial branch includes a stacked convolution module, and the third stage of the spatial branch includes a bottleneck-like structure module; step 304 includes: inputting the shallow features into the stacked convolution module in the first stage of the spatial branch to obtain the first intermediate spatial features as:
[0108] x out ′ = Conv 3×3 (Conv 3×3 (x in )) + x in (3)
[0109] where, x out ′ represents the first intermediate spatial features, and x in represents the input shallow features.
[0110] Specifically, feature extraction, feature reuse, and the stacking of receptive fields are achieved through the stacked convolution module. Among them, Stride is the stride of the first 3×3 convolution. When S = 2, the output image resolution is 1 / 2 of the input. The structure of the stacked convolution module is as shown in Figure 6 Shown.
[0111] In one embodiment, step 310 includes: inputting the second-stage spatial features into the third stage of the spatial branch to obtain the third intermediate spatial features as:
[0112] x out ″ = Conv 1×1 (Conv 3×3 (Conv 1×1 (x in ′)))+x in ′ (4)
[0113] Among them, x out ″ represents the third intermediate space feature, and x in ′ represents the second-stage space feature.
[0114] Specifically, feature extraction, feature reuse, receptive field superposition, and channel number expansion are achieved through a bottleneck-like structure module. Among them, Stride is the convolution stride of 3×3. When S = 2, the output image resolution is 1 / 2 of the input; exp is the channel number expansion multiple. Preferably, the channel number expansion multiple is 2 and is achieved by the second 1×1 convolution. The structure of the bottleneck-like structure module is as Figure 7 shown.
[0115] In one embodiment, the boundary-guided linear attention module includes: a convolutional layer, a Softmax function, and a Sigmoid function; step 306 includes: inputting the first intermediate space feature, the first-stage semantic feature, and the first-stage boundary feature into the first boundary-guided linear attention module to obtain the first-stage space feature as:
[0116] x out ″′ = attn * x s + x s (5)
[0117] attn = Softmax(Resize(f c (x c )) + f b (x b )) * (1 - σ) (6)
[0118] σ = Sigmoid(f b (x b )) (7) Among them, x s , x b , x c respectively represent the first intermediate space feature, the first-stage boundary feature, and the first-stage semantic feature, x out ″′ is the first-stage space feature, f b (·), f c (·) are respectively the mappings of the input boundary information and space information, σ is the score of the input boundary information, attn is the attention score calculated according to the semantic information and boundary information, Resize(·) is an image size adjustment function, and the output size of the semantic information is upsampled to the same size as the boundary information size.
[0119] Specifically, in PIDNet (2022, Jiacong Xu, et al.), it is mentioned that the output size of the semantic branch in the dual-branch network structure is relatively small. When performing the final fusion output, an upsampling operation is usually adopted, with a multiple generally ranging from 8 to 16 times. During the upsampling process, this will cause the "submergence" of the target boundary. To solve this problem, the present invention designs a boundary-guided linear attention module. The boundary-guided linear attention module consists of a convolution, a Softmax, and a Sigmoid function. Guided by the boundary information, the outputs of the spatial branch and the semantic branch are fused with each other to obtain the final result. The outputs of the boundary branch and the semantic branch first pass through convolution and upsampling according to their own characteristics to map the number of channels and the image size to the same size for subsequent operations. Then, the boundary information is superimposed on the semantic information through addition to supplement its boundary features. The attention map of the semantic information branch is obtained using the Softmax function, and the attention map of the boundary information is obtained using the Sigmoid function. The Hadamard product of the two is used to obtain the attention map for fusion. The attention map is multiplied by the spatial information and then passed through a residual connection to obtain the final output. The structure of the boundary-guided linear attention module is as Figure 8 shown.
[0120] In one embodiment, the spatial features at each stage include the first-node spatial features, the second-stage spatial features, and the third-stage spatial features; the segmentation network includes a spatial segmentation head and a channel attention mechanism; step 204 includes: concatenating the shallow features, the first-node spatial features, the second-stage spatial features, and the third-stage spatial features and inputting them into the channel attention mechanism for feature fusion to obtain fused features; predicting the fused features through the spatial segmentation head to obtain the predicted semantic segmentation result.
[0121] Specifically, the structure of the spatial segmentation head is the same as that of the segmentation head in PIDNet.
[0122] The spatial segmentation head consists of a batch normalization layer, a ReLU activation function, a 3×1 convolutional layer, a batch normalization layer, a ReLU activation function, and a 1×1 convolutional layer.
[0123] The structure of the channel attention mechanism is as Figure 9 shown. In the channel attention mechanism, first, global average pooling is performed on each channel of the input feature map to obtain the average value on each channel. Then, a 1×1 convolution is used to map them into a weight vector. Finally, each channel of the input feature map is weighted to obtain a weighted feature map, and the original information is fully utilized through a residual connection. This can not only make the model pay more attention to the feature channels useful for segmentation but also better utilize the feature information, thereby improving the model performance.
[0124] In one embodiment, the segmentation network further includes a boundary-assisted segmentation head and a spatial-assisted segmentation head; the boundary-assisted segmentation head and the spatial-assisted segmentation head are in the second stage of the boundary branch and the spatial branch of the feature extraction network; step 206 includes: constructing a total loss function, which is a weighted sum of the boundary-assisted segmentation head loss function, the spatial-assisted segmentation head loss function, and the spatial segmentation head loss function; the spatial-assisted segmentation head loss function and the spatial segmentation head loss function adopt the cross-entropy loss function, and the boundary-assisted segmentation head loss function adopts the binary cross-entropy loss function; according to the annotation of the training samples, the predicted semantic segmentation result, and the total loss function, the remote sensing image semantic segmentation network is trained in reverse to obtain a trained remote sensing image semantic segmentation network.
[0125] Specifically, the boundary-assisted segmentation head and the spatial-assisted segmentation head have the same structure as the spatial segmentation head.
[0126] In the third stage of the network, auxiliary segmentation heads are added to the boundary and spatial branches to accelerate the training speed of the network and enhance the network stability. The spatial-assisted segmentation head and the spatial segmentation head adopt the cross-entropy loss function, denoted as l0 and l1 respectively, while the boundary-assisted segmentation head adopts the binary cross-entropy loss function, denoted as l2. The total loss function Loss is the weighted sum of l0, l1, and l2. As a preference, the corresponding weights are [ω0, ω1, ω2] = [0.4, 1.0, 20.0].
[0127] Loss = ω0l0 + ω1l1 + ω2l2 (8)
[0128] In one embodiment, both the first stage and the second stage of the semantic branch include n stacked convolutional modules, and the third stage includes two class bottleneck structure modules and 1 parallel aggregation pyramid pooling module.
[0129] Specifically, the structure of the parallel aggregation pyramid pooling module (Parallel Aggregation Pyramid Pooling Module, abbreviated as: PAPPM) is as Figure 9 shown.
[0130] The parallel aggregation pyramid pooling module is used to quickly obtain semantic information under different receptive fields through parallel multi-scale pooling modules, and then aggregate the multi-scale feature information through addition and convolution operations to improve the receptive field of the network and the ability to obtain global information, and also have better recognition ability for target objects of different sizes, thereby improving the performance of the network. At the same time, due to the interaction design between different scales, information loss can be avoided, so more information can be retained, thereby improving the accuracy and robustness of the network.
[0131] It should be understood that although Figure 1The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least a part of the steps in
[0132] In a verification embodiment, the experiment uses the open-source platform mmsegmentation to complete training and testing, and the specific environment configuration is shown in Table 1.
[0133] Table 1 Experimental environment configuration
[0134]
[0135] The dataset uses the GID5: Large-scale High-resolution Satellite Land Cover Dataset released by Wuhan University, which contains 5 land cover categories including buildings, farmland, forests, grasslands and water areas, and a total of 150 scenes of Gaofen-2 satellite remote sensing images with pixel-level annotations. Among them, the training set is 120 scenes of images, the validation set is 30 scenes of images, and the size of one scene of image is 6800×7200 pixels. Before the experiment, the images have been processed into 21,840 training set images and 5,460 validation set images with a size of 1024×1024.
[0136] During training, a data augmentation method of randomly cropping to a size of 512×512 is adopted, the batch size is 8, the initial learning rate is 0.01, the weight decay parameter is 0.0005, and the number of iterations is 40,950.
[0137] A fair comparison is made with STDCNet, BiseNet, Seaformer(2023, Qiang Wan, et al.), and PIDNet according to the experimental configuration. As shown in Table 2, on the GID5 dataset, the segmentation accuracy of BGFNet-m reaches up to 82.24%, and only STDC1 is close to us with an average intersection over union of 80.64%. And BGFNet-s achieves an average intersection over union of 77.93% with 7.4M parameters and 22.7 GFLOPs of computational volume, which is higher than other networks. It can be seen that the remote sensing image semantic segmentation network BGFNet of the present application has achieved an excellent balance between segmentation accuracy and model size.
[0138] Table 2 Experimental results on the GID5 dataset
[0139]
[0140] Note: The computational load is the test result on an image of size 1024×1024.
[0141] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0142] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A semantic segmentation method for remote sensing images based on boundary information-guided multi-information fusion, characterized in that, The method includes: Obtain a remote sensing image, and annotate the remote sensing image to obtain training samples; Construct a remote sensing image semantic segmentation network, which includes an input network, a feature extraction network, and a segmentation network; the input network is used to extract shallow features of the training samples by using convolution operations; the feature extraction network includes a boundary branch, a semantic branch, and a spatial branch, and each branch includes three stages. Each stage of the boundary branch includes a bidirectional boundary gating module, and the spatial branch and the semantic branch include a stacked convolution module and a class bottleneck structure module. A boundary-guided linear attention module is further included after each stage in the spatial branch; the feature extraction network is used to extract boundary features and semantic features of each stage of the shallow features by using the boundary branch and the semantic branch respectively, and use the boundary-guided linear attention module to fuse the boundary features and semantic features of this stage into the features extracted by the spatial branch of this stage to obtain the spatial features of this stage; the segmentation network is used to make predictions based on the shallow features and the spatial features of three stages to obtain a predicted semantic segmentation result; the bidirectional boundary gating module includes: a convolutional layer, a depthwise separable convolutional layer, and a gating mechanism; the boundary-guided linear attention module includes: a convolutional layer, a Softmax function, and a Sigmoid function; Train the remote sensing image semantic segmentation network according to the annotation of the training samples and the predicted semantic segmentation result obtained by inputting the training samples into the remote sensing image semantic segmentation network to obtain a trained remote sensing image semantic segmentation network; Input the obtained remote sensing image to be segmented into the trained remote sensing image semantic segmentation network to obtain the semantic segmentation result of the remote sensing image to be segmented.
2. The method according to claim 1, wherein The input network includes a Stem module and two stacked convolutional modules; Train the remote sensing image semantic segmentation network according to the training samples and the predicted semantic segmentation result obtained by inputting the training samples into the remote sensing image semantic segmentation network to obtain a trained remote sensing image semantic segmentation network, including: Input the training samples into the input network to extract features of different dimensions through the Stem module, and perform feature extraction on the obtained features through two stacked convolutional modules with different strides to obtain shallow features; Input the shallow features into the feature extraction network to obtain the spatial features of each stage; Input the shallow features and the spatial features of each stage into the segmentation network to obtain a predicted semantic segmentation result; Perform backpropagation training on the remote sensing image semantic segmentation network according to the annotation of the training samples and the predicted semantic segmentation result to obtain a trained remote sensing image semantic segmentation network.
3. The method according to claim 2, wherein Input the shallow features into the feature extraction network to obtain the spatial features of each stage, including: Input the shallow features into the boundary branch to obtain the first-stage boundary feature, the second-stage boundary feature, and the third-stage boundary feature; Input the shallow features into the semantic branch to obtain the first-stage semantic features, the second-stage semantic features, and the third-stage semantic features; Input the shallow features into the first stage of the spatial branch to obtain the first intermediate spatial features; Input the first intermediate spatial features, the first-stage semantic features, and the first-stage boundary features into the first boundary-guided linear attention module to obtain the first-stage spatial features; Input the first-stage spatial features into the second stage of the spatial branch to obtain the second intermediate spatial features; input the second intermediate spatial features, the second-stage semantic features, and the second-stage boundary features into the second boundary-guided linear attention module to obtain the second-stage spatial features; Input the second-stage spatial features into the third stage of the spatial branch to obtain the third intermediate spatial features; input the third intermediate spatial features, the third-stage semantic features, and the third-stage boundary features into the third boundary-guided linear attention module to obtain the third-stage spatial features.
4. The method according to claim 2, wherein Each stage of the boundary branch includes a bidirectional boundary gating module; Input the shallow features into the boundary branch to obtain the first-stage boundary features, the second-stage boundary features, and the third-stage boundary features, including: Input the shallow features into the bidirectional boundary gating module of the first stage of the boundary branch to obtain the first-stage boundary features as: x out = GELU(Cat(f x (x in ), f y (x in )))*f(x in ) + x in Among them, x in represents the input shallow features, and x out represents the output of the bidirectional boundary gating module. f(·), f x (·), f y (·) respectively represent the extraction of conventional, horizontal, and vertical convolutional boundary information. f(x in ) = DWConv 3×3 (x in ), f x (x in ) = DWConv 3×1 (Conv 1×1 (x in ))), f y (x in ) = DWConv 1×3 (Conv 1×1 (x in ))). Cat(·) represents the concatenation operation, and GELU(·) represents the gating mechanism; Input the first-stage boundary features into the first stage of the boundary branch to obtain the second-stage boundary features; Input the second-stage boundary features into the first stage of the boundary branch to obtain the third-stage boundary features.
5. The method according to claim 3, characterized in that, The first and second stages of the spatial branch each include a stacked convolution module, and the third stage of the spatial branch includes a class bottleneck structure module; Input the shallow features into the first stage of the spatial branch to obtain the first intermediate spatial features, including: Input the shallow features into the stacked convolution module of the first stage of the spatial branch to obtain the first intermediate spatial features as: x out ' = Conv 3×3 (Conv 3×3 (x in )) + x in Among them, x out ' represents the first intermediate spatial feature, and x in represents the input shallow feature.
6. The method according to claim 5, wherein Input the second-stage spatial features into the third stage of the spatial branch to obtain the third intermediate spatial features, including: Input the second-stage spatial features into the third stage of the spatial branch to obtain the third intermediate spatial features as: x out ″ = Conv 1×1 (Conv3×3(Conv 1×1 (x in ′)))+x in ′ where x out ″ represents a third intermediate space feature, and x in ′ represents a second stage space feature.
7. The method according to claim 3, characterized in that, Input the first intermediate spatial features, the first-stage semantic features, and the first-stage boundary features into the first boundary-guided linear attention module to obtain the first-stage spatial features, including: Input the first intermediate spatial features, the first-stage semantic features, and the first-stage boundary features into the first boundary-guided linear attention module to obtain the first-stage spatial features as: x out ″′ = attn * x s + x s attn = Softmax(Resize(f c (x c )) + f b (x b )) * (1 - σ) σ = Sigmoid(f b (x b )) Among them, x s , x b , x c represent the first intermediate space feature, the first stage boundary feature, and the first stage semantic feature respectively, x out ″′ is the first stage space feature, f b (·), f c (·) are the mappings of the input boundary information and space information respectively, σ is the score of the input boundary information, attn is the attention score calculated according to the semantic information and boundary information, and Resize(·) is the image size adjustment function.
8. The method according to claim 2, wherein The spatial features of each stage include the first-node spatial features, the second-stage spatial features, and the third-stage spatial features; The segmentation network includes a spatial segmentation head and a channel attention mechanism; Input the shallow features and the spatial features of each stage into the segmentation network to obtain the predicted semantic segmentation result, including: Concatenate the shallow features, the first node spatial features, the second-stage spatial features, and the third-stage spatial features, and input them into the channel attention mechanism for feature fusion to obtain fused features; Pass the fused features through the spatial segmentation head for prediction to obtain a predicted semantic segmentation result.
9. The method according to claim 8, wherein The segmentation network further includes a boundary auxiliary segmentation head and a spatial auxiliary segmentation head; The boundary auxiliary segmentation head and the spatial auxiliary segmentation head are in the second stage of the boundary branch and the spatial branch of the feature extraction network; According to the annotation of the training samples and the predicted semantic segmentation result, perform backpropagation training on the remote sensing image semantic segmentation network to obtain a trained remote sensing image semantic segmentation network, including: Construct a total loss function, which is a weighted sum of the boundary auxiliary segmentation head loss function, the spatial auxiliary segmentation head loss function, and the spatial segmentation head loss function; the spatial auxiliary segmentation head loss function and the spatial segmentation head loss function adopt the cross-entropy loss function, and the boundary auxiliary segmentation head loss function adopts the binary cross-entropy loss function; According to the annotation of the training samples, the predicted semantic segmentation result, and the total loss function, perform backpropagation training on the remote sensing image semantic segmentation network to obtain a trained remote sensing image semantic segmentation network.
10. The method according to claim 1, wherein Both the first stage and the second stage of the semantic branch include n stacked convolutional modules, and the third stage includes two class bottleneck structure modules and one parallel aggregation pyramid pooling module.