Multi-spectral Remote Sensing Image Segmentation Method Based on Band-Position Selection
The waveband-location selection method addresses the challenges of MSIs segmentation by leveraging spectral-spatial relationships, improving feature extraction and segmentation accuracy through adaptive selection and fusion techniques.
Patent Information
- Application Number
- CN202210817848.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-06-29
- Filing Date
- 2022-07-12
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-07-12
AI Technical Summary
When the existing multispectral remote sensing image segmentation method is difficult to effectively capture representative landform information when processing multiple spectral bands, and the existing methods fail to fully consider the correlation between bands and spatial position weight, resulting in poor segmentation effect, especially in complex scenarios and small target recognition.
The multi-spectral remote sensing image segmentation method based on band-position selection is adopted, and effective information is captured through the band-position adaptive selection module, spectral-space features are extracted in combination with the three-dimensional residual module, and feature recalibration is used to use adaptive feature fusion and channel attention module to perform feature recalibration, and pixel-level prediction is finally performed to generate high-quality segmentation results.
Effectively segment complex scenes in multi-spectral remote sensing images, enhance semantic segmentation results, and better extract small target features of scattered distributions, and the generated segmentation graph boundary contours are clearer, improving segmentation accuracy and detail retention capabilities.
Smart Images

Figure CN115187984B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of image segmentation, and particularly relates to a multispectral remote sensing image segmentation method based on band-position selection. Background Art
[0002] Multispectral Remote Sensing Images (MSIs) contain a large amount of ground object information, which is contained in multiple spectral bands of the image. MSIs can achieve ground object classification through semantic segmentation and play an important role in fields such as land cover, environmental change, and military monitoring.
[0003] Currently, there are two challenges in MSIs semantic segmentation. First, MSIs have a wide imaging range, with the characteristics of large within-class variance and small between-class variance, and it is difficult to obtain discriminative features. Second, MSIs have multiple spectral bands, and the amount of information contained in different bands varies greatly, making it more difficult to capture effective information. For the first challenge, modeling context information helps to segment complex scenes of MSIs. Currently, there are many methods for capturing context information of images. For example, global context information can be captured by stitching features of four different pyramid scales; or the medium- and long-range spatial relationships between objects can be modeled, and a relationship feature map containing spatial context information can be generated through element-wise dot product operations. These methods all use skip connections similar to U-Net to directly fuse features at different levels to generate multi-scale context features, but they are relatively rough and do not further explore the contribution degree of features at different levels. For the second challenge, existing MSIs semantic segmentation methods mostly use two-dimensional convolution to process the information of each band, without considering the influence of different bands of MSIs and their correlations on segmentation.
[0004] In addition, MSIs usually contain several to more than a dozen spectral bands, and how to capture representative ground object information from multiple bands is a problem worthy of consideration. For the classification of Hyperspectral Remote Sensing Images (HSIs), some scholars have studied the influence of band relationships on classification performance based on three-dimensional convolution methods. However, most current research on MSIs semantic segmentation focuses on relationship modeling in the channel or spatial dimension and does not consider the correlation between bands. Summary of the Invention
[0005] Aiming at the deficiencies in the prior art, the purpose of the present disclosure is to provide a multispectral remote sensing image segmentation method based on band-position selection, which can effectively cope with the segmentation of complex scenes of MSIs containing multiple spectral bands, can capture the spectral bands rich in effective information of MSIs, helps to enhance the semantic segmentation result, and can better extract the features of small targets distributed dispersedly.
[0006] To achieve the above object, the present disclosure provides the following technical solutions:
[0007] A multi-spectral remote sensing image segmentation method based on band-position selection, comprising the following steps:
[0008] S1: Obtain the multi-spectral remote sensing image to be segmented and input it into a multi-spectral remote sensing image segmentation model for encoding, and the encoding includes:
[0009] S11: Capture the spectral bands rich in effective information in the multi-spectral remote sensing image to obtain a band-position feature map;
[0010] S12: Extract spectral-spatial features from the band-position feature map to obtain a shallow low-level spectral-spatial feature map and a deep high-level spectral-spatial feature map;
[0011] S2: Decode the deep high-level spectral-spatial feature map, and the decoding includes:
[0012] S21: Perform adaptive feature fusion on the shallow low-level spectral-spatial feature map and the deep high-level spectral-spatial feature map to obtain a fused feature map;
[0013] S22: Perform feature recalibration on the fused feature map to obtain a channel feature map;
[0014] S23: Upsample the channel feature map to obtain a decoded map with restored spatial resolution, and perform pixel-level prediction on the decoded map to obtain a segmentation result map of the multi-spectral remote sensing image.
[0015] Preferably, in step S1, the multi-spectral remote sensing image segmentation model includes:
[0016] A band-position adaptive selection module for capturing the spectral bands rich in effective information in the multi-spectral remote sensing image to obtain a band-position feature map;
[0017] A three-dimensional residual module for extracting spectral-spatial features from the band-position feature map to obtain a shallow low-level spectral-spatial feature map and a deep high-level spectral-spatial feature map;
[0018] An adaptive feature fusion module for performing adaptive feature fusion on the shallow low-level spectral-spatial feature map and the deep high-level spectral-spatial feature map to obtain a fused feature map;
[0019] A channel attention module for performing feature recalibration on the fused feature map to obtain a channel feature map;
[0020] A prediction module, configured to upsample the channel feature map to obtain a decoded map with restored spatial resolution, and configured to perform pixel-level prediction on the decoded map to obtain a segmentation result map of the multi-spectral remote sensing image.
[0021] Preferably, step S11 includes:
[0022] S111: Provide global context information for each spectral band in the multi-spectral remote sensing image through global average pooling to obtain a band map;
[0023] S112: Obtain a band weight map of the multi-spectral remote sensing image based on the band map through one-dimensional convolution;
[0024] S113: Perform Hadamard product operation on the multi-spectral remote sensing image and the band weight map to obtain a three-dimensional feature map with different band weights;
[0025] S114: Perform a mean operation on the three-dimensional feature map in the spectral dimension to obtain a spatial feature map retaining full-band information;
[0026] S115: Obtain a position weight map of the multi-spectral remote sensing image based on the spatial feature map through two-dimensional convolution;
[0027] S116: Perform an element-wise multiplication operation on the position weight map and the three-dimensional feature map in the spatial dimension to obtain a band-position feature map containing spatial relationships.
[0028] Preferably, in step S112, the band weight map W n is obtained by the following formula:
[0029] W n = σ(f 3 (δ(F i )))
[0030] And
[0031]
[0032] where F i represents the i-th input multi-spectral remote sensing image, and F i ∈R N×H×W ,R N×H×W represents three-dimensional real space, F(:, :, :) represents the spectral dimension of the input multi-spectral remote sensing image, N, H, and W respectively represent the number of spectral bands, height, and width of the multi-spectral remote sensing image; h and w respectively represent the height of the h-th row and the width of the w-th column of the multi-spectral remote sensing image, and h ∈ H, 0 ≤ h < H, w ∈ W, 0 ≤ w < W; δ(·) represents a per-band global average pooling operation, and f 3 (·) represents a one-dimensional convolution operation with a convolution kernel size of 3.
[0033] Preferably, in step S115, the position weight map W l is obtained by the following formula:
[0034]
[0035] wherein, represents the obtained spatial feature map, and f 3×3 (·) represents a two-dimensional convolution operation with a convolution kernel size of 3×3, and σ(·) represents the Sigmoid function.
[0036] Preferably, the element-wise multiplication operation of the position weight map and the three-dimensional feature map in the spatial dimension is expressed as:
[0037] F l = F b ⊙W l
[0038] wherein, F l represents the obtained band-position feature map.
[0039] Preferably, in step S12, the spectral-spatial features are extracted from the band-position feature map through multiple three-dimensional residual modules. Among them, the calculation process of each three-dimensional residual module is as follows:
[0040]
[0041] wherein, x l represents the input of the l-th three-dimensional residual module, and x l+1 represents the output of the l-th three-dimensional residual module. represents the band-position adaptive selection layer, R(·) represents the residual function, and W l represents the weight matrix of the l-th three-dimensional residual module.
[0042] Step S21 includes the following steps:
[0043] S211: Perform adaptive feature fusion on the deep high-level spectral-spatial feature map and the shallow low-level spectral-spatial feature map to obtain a fused feature map;
[0044] S212: Update the weight values of the trainable parameter matrices W α , W β through backpropagation until the weight values reach the optimum.
[0045] Preferably, step S22 includes the following steps:
[0046] S221: Perform per-channel compression on the fused feature map to obtain a one-dimensional feature map;
[0047] S222: Learn the channel weight vector U of the one-dimensional feature map using one-dimensional convolution CAM to obtain a channel feature map containing different channel weights.
[0048] Preferably, step S23 includes the following steps:
[0049] S231: Upsample the channel feature map using transposed convolution to obtain a decoded map with restored spatial resolution;
[0050] S232: Map the features of the decoded map with restored spatial resolution to the sample category space using two-dimensional convolution, perform pixel-level prediction on the decoded map through the Softmax function to obtain the prediction probabilities of each pixel on the decoded map belonging to different categories, thereby obtaining the segmentation result.
[0051] Compared with the prior art, the beneficial effects brought by the present disclosure are as follows:
[0052] 1) Aiming at the characteristics that multi-spectral remote sensing images contain multiple spectral bands, a band-position adaptive selection module is proposed to select effective bands while modeling spatial position weights to capture the spatial context information of multi-spectral remote sensing images, which can effectively segment multi-spectral remote sensing images with high intra-class variance and low inter-class variance.
[0053] 2) A three-dimensional residual module is proposed, which can utilize the band correlation of multi-spectral remote sensing images to extract more discriminative spectral-spatial features.
[0054] 3) An adaptive feature fusion module is proposed to adaptively adjust the fusion ratio of feature maps at different levels through network learning. In addition, channel attention is extended to three-dimensional data to refine the fused feature maps, further improving the segmentation results of multi-spectral remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a flowchart of a multi-spectral remote sensing image segmentation method based on band-position selection provided by an embodiment of the present disclosure;
[0056] Figure 2 is a multi-spectral remote sensing image segmentation model based on band-position selection provided by an embodiment of the present disclosure;
[0057] Figure 3 is a schematic structural diagram of a band-position adaptive selection module provided by an embodiment of the present disclosure;
[0058] Figure 4 is a schematic structural diagram of a three-dimensional residual module provided by an embodiment of the present disclosure;
[0059] Figure 5It is a schematic structural diagram of a channel attention module provided by an embodiment of the present disclosure;
[0060] Figure 6 It is a multi-spectral remote sensing image provided by an embodiment of the present disclosure;
[0061] Figure 7 It is a band graph provided by an embodiment of the present disclosure;
[0062] Figure 8 It is a band weight graph provided by an embodiment of the present disclosure;
[0063] Figure 9 It is a three-dimensional feature map provided by an embodiment of the present disclosure;
[0064] Figure 10 It is a spatial feature map provided by an embodiment of the present disclosure;
[0065] Figure 11 It is a position weight graph provided by an embodiment of the present disclosure;
[0066] Figure 12 It is a band-position feature map provided by an embodiment of the present disclosure;
[0067] Figure 13 It is a shallow low-level spectral-spatial feature map provided by an embodiment of the present disclosure;
[0068] Figure 14 It is a deep high-level spectral-spatial feature map provided by an embodiment of the present disclosure;
[0069] Figure 15 It is a fused feature map provided by an embodiment of the present disclosure;
[0070] Figure 16 It is a channel feature map provided by an embodiment of the present disclosure;
[0071] Figure 17 It is the segmentation result of the Potsdam dataset;
[0072] Figure 18 It is the segmentation result of the Qinghai dataset;
[0073] Figure 19 It is the segmentation result of the Tibet Plateau dataset. Detailed implementation manners
[0074] Next, reference will be made to the appendix Figures 1 to 19Specific embodiments of the present disclosure are described in detail. Although specific embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0075] It should be noted that certain terms are used in the specification and claims to refer to specific components. Those skilled in the art should understand that technicians may use different nouns to refer to the same component. The specification and claims do not use the difference in nouns as a way to distinguish components, but use the difference in the functions of components as the criterion for distinction. As mentioned throughout the specification and claims, "comprising" or "including" is an open-ended term and should be interpreted as "including but not limited to". The subsequent description in the specification is the preferred implementation manner for implementing the present disclosure, but the description is for the purpose of the general principles of the specification and is not used to limit the scope of the present disclosure. The protection scope of the present disclosure shall be subject to what is defined by the appended claims.
[0076] To facilitate the understanding of the embodiments of the present disclosure, the following will further explain with specific embodiments as examples in conjunction with the drawings, and each drawing does not constitute a limitation on the embodiments of the present disclosure.
[0077] In one embodiment, as Figure 1 shown, the present disclosure provides a multi-spectral remote sensing image segmentation method based on band-position selection, including the following steps:
[0078] S1: Obtain the multi-spectral remote sensing image to be segmented and input it into the multi-spectral remote sensing image segmentation model for encoding, and the encoding includes:
[0079] S11: Use the band-position adaptive selection module to capture the spectral bands rich in effective information in the multi-spectral remote sensing image to obtain a band-position feature map;
[0080] S12: Use a plurality of three-dimensional residual modules to extract spectral-spatial features from the band-position feature map to obtain a shallow low-level spectral-spatial feature map and a deep high-level spectral-spatial feature map;
[0081] S2: Decode the deep high-level spectral-spatial feature map, and the decoding includes:
[0082] S21: Use the adaptive feature fusion module to adaptively fuse the shallow low-level spectral-spatial feature map and the deep high-level spectral-spatial feature map to obtain a fused feature map;
[0083] S22: Feature recalibration is performed on the fused feature map in the channel dimension using a channel attention module to obtain a channel feature map containing different channel weights;
[0084] S23: Upsample the channel feature map to obtain a decoded map with restored spatial resolution;
[0085] S24: Perform pixel-level prediction on the decoded map to obtain a segmentation result map of the multi-spectral remote sensing image.
[0086] The above embodiments constitute the complete technical solution of the present disclosure. The solution in this embodiment can effectively handle the segmentation of complex scenes containing multiple spectral bands, can capture the information-rich spectral bands in the multi-spectral remote sensing image, helps to enhance the semantic segmentation result, and can better extract the features of small targets distributed dispersedly. By modeling the band-spatial position weights, the boundary contour of the generated segmentation map is clearer.
[0087] In another embodiment, as Figure 2 shown, the multi-spectral remote sensing image segmentation model includes:
[0088] A band-position adaptive selection module for capturing the spectral bands rich in effective information in the multi-spectral remote sensing image to obtain a band-position feature map; as Figure 3 shown, the band-position adaptive selection module includes a band selection layer and a spatial position selection layer. Among them, the band selection layer includes: a global average pooling layer, a one-dimensional convolutional layer with a convolutional kernel size of 3, a Sigmoid function, and a Hadamard product function; the spatial position selection layer includes: a mean function, a two-dimensional convolutional layer with a convolutional kernel size of 3×3, a Sigmoid function, and a Hadamard product function.
[0089] A three-dimensional residual module for extracting spectral-spatial features from the band-position feature map to obtain a shallow low-level spectral-spatial feature map and a deep high-level spectral-spatial feature map; as Figure 4 shown, the three-dimensional residual module contains a band-position adaptive selection layer, two three-dimensional convolutional layers with a convolutional kernel size of 3×3×3, a Relu activation function, and a 1×1×1 three-dimensional convolutional layer for adjusting the size of the feature map;
[0090] An adaptive feature fusion module for adaptively fusing the deep high-level spectral-spatial feature map with the shallow low-level spectral-spatial feature map to obtain a fused feature map; the adaptive feature fusion module includes two trainable parameter matrices W α 、W β , for adaptively adjusting the fusion ratio of the high-level spectral-spatial feature map and the shallow low-level spectral-spatial feature map;
[0091] The channel attention module is used to recalibrate the features of the fused feature map to obtain a channel feature map; as Figure 5 shown, the channel attention module includes a global average pooling layer, a one-dimensional convolutional layer with a convolutional kernel size of 3, a sigmoid function, and a Hadamard product function.
[0092] The prediction module is used to upsample the channel feature map to obtain a decoded map with restored spatial resolution, and to perform pixel-level prediction on the decoded map to obtain a segmentation result map of the multi-spectral remote sensing image; the prediction module includes two transposed convolutional layers, a two-dimensional convolutional layer with a size of 3×3, and a Softmax function.
[0093] In this embodiment, the training process of the multi-spectral remote sensing image segmentation model is as follows:
[0094] 1. Divide the multi-spectral remote sensing image data into a training set and a test set;
[0095] 2. Input the training set data into the network for training, and use the band-position adaptive selection module to obtain the band-position feature map of the training data;
[0096] 3. Use the three-dimensional residual module to extract spectral-spatial features from the band-position feature map to obtain a shallow low-level spectral-spatial feature map and a deep high-level spectral-spatial feature map;
[0097] 4. Use the adaptive feature fusion module to adaptively fuse the deep high-level spectral-spatial feature map with the shallow low-level spectral-spatial feature map to obtain a fused feature map;
[0098] 5. Use the channel attention module to recalibrate the features of the fused feature map to obtain a channel feature map;
[0099] 6. Based on the channel feature map, use the prediction module to obtain a segmentation result map, calculate the model segmentation loss based on the segmentation result map, and update the model parameters through backpropagation;
[0100] 7. Repeat steps 2 to 6 until the number of training times reaches the maximum number of iterations, and the training is completed.
[0101] In another embodiment, step S11 includes:
[0102] S111: Provide global context information for each spectral band in the multi-spectral remote sensing image as shown in Figure 6 shown through global average pooling to obtain a band map as shown in Figure 7 shown;
[0103] In this step, since the input multi-spectral remote sensing image is three-dimensional, it is impossible to directly obtain the band weight map through the subsequent one-dimensional convolution operation. Therefore, it is necessary to compress the multi-spectral remote sensing image into one dimension (that is, only retain the spectral dimension in the multi-spectral remote sensing image, but at the same time, it is necessary to retain the global spatial information of each spectral band of the image as much as possible, so as to learn a band weight map that can accurately reflect the importance of different spectral bands). The role of global average pooling is to compress the input multi-spectral remote sensing image F ∈ R N×H×W to F ∈ R N , that is, add up the spectral values at different spatial positions on each spectral band of the multi-spectral remote sensing image and then take the average. This can provide global context information for each spectral band, and then the band weight map can be obtained through one-dimensional convolution operation.
[0104] S112: Based on the band map and through one-dimensional convolution or fully connected layer learning, obtain the band weight map W Figure 8 as shown, which can accurately reflect the importance of different spectral bands in the multi-spectral remote sensing image n ;
[0105] In this step, the band weight map W n is obtained through the following formula:
[0106] W n = σ(f 3 (δ(F i )))
[0107] And
[0108]
[0109] where F i represents the i-th input multi-spectral remote sensing image, and F i ∈ R N×H×W , R N×H×W represents the three-dimensional real number space, F(:, :, :) represents the spectral dimension of the input multi-spectral remote sensing image, N, H, and W respectively represent the number of spectral bands, height, and width of the multi-spectral remote sensing image; h and w respectively represent the height of the h-th row and the width of the w-th column of the multi-spectral remote sensing image, and h ∈ H, 0 ≤ h < H, w ∈ W, 0 ≤ w < W; δ(·) represents the per-band global average pooling operation, f 3 (·) represents a one-dimensional convolution operation with a convolution kernel size of 3, and the one-dimensional convolution formula is as follows:
[0110]
[0111] where represents the value at position (x + h) of the one-dimensional convolution input X 1d , represents the one-dimensional convolution output Y1d The value at position x, represents the one-dimensional convolutional kernel K 1d The value at position h, where H represents the one-dimensional convolutional kernel K 1d is the height of K, and h ∈ H, 0 ≤ h < H.
[0112] S113: Perform the Hadamard product operation on the multispectral remote sensing image and the band weight map to obtain a three-dimensional feature map F Figure 9 as shown, which is assigned different band weights b ;
[0113] In this step, the Hadamard product operation is expressed as:
[0114]
[0115] where F b represents the obtained three-dimensional feature map, and ⊙ represents the Hadamard product operation.
[0116] S114: Perform a mean operation on the three-dimensional feature map F b in the spectral dimension to obtain a spatial feature map that retains all-band information as Figure 10 shown;
[0117] In this step, the three-dimensional feature map F b can be mean-operated by the following formula:
[0118]
[0119] where represents performing a mean operation in the spectral dimension, and n represents the nth spectral band of the three-dimensional feature map F b ;
[0120] S115: Based on the spatial feature map, obtain the position weight map W1 of the multispectral remote sensing image through two-dimensional convolutional learning as Figure 11 shown;
[0121] In this step, the position weight map W l is obtained through the following formula:
[0122]
[0123] where represents the obtained spatial feature map, σ(·) represents the Sigmoid function, and f 3×3 (·) represents a two-dimensional convolution operation with a convolutional kernel size of 3×3, and the two-dimensional convolution formula is as follows:
[0124]
[0125] where Denote the two-dimensional convolution input X 2d The value at position (x + h, y + w), Denote the two-dimensional convolution output Y 2d The value at position (x, y), Denote the two-dimensional convolution kernel K 2d The value at position (h, w), where H and W respectively represent the height and width of the two-dimensional convolution kernel K 2d And h ∈ H, 0 ≤ h < H, w ∈ W, 0 ≤ w < W.
[0126] S116: Multiply the position weight map W l Element-wise with the three-dimensional feature map F b In the spatial dimension, that is:
[0127] F l = F b ⊙ W l
[0128] Where F l Represents the obtained band-position feature map.
[0129] Through the above formula, a band-position feature map containing spatial relationships as shown in Figure 12 Can be obtained.
[0130] In another embodiment, since the original ResNet network uses two-dimensional convolution to extract image features and does not consider the correlation between spectral bands, it downsamples by 32 times from the input image to the output feature map, resulting in the loss of a large amount of pixel information, which makes it more difficult to restore image details in the decoder stage. Therefore, in this embodiment, 9 three-dimensional residual modules are used in the encoding stage to downsample the band-position feature map by 16 times to extract spectral-spatial features. Compared with the original ResNet, the resolution of the extracted spectral-spatial feature map can be doubled, and more detailed information can be retained while effectively encoding the semantic information of multi-spectral remote sensing images.
[0131] Furthermore, in each three-dimensional residual module, a band-position adaptive selection layer is added to generate a weight vector to guide the original feature map to remove redundant features and enhance the expression of effective features. For each three-dimensional residual module, its calculation process is described as follows:
[0132]
[0133] Where x l Represents the input of the l-th three-dimensional residual module, and x l+1 Represents the output of the l-th three-dimensional residual module, Represents the band-position adaptive selection layer, and R(·) represents the residual function, Wl Represents the weight matrix of the l-th three-dimensional residual module.
[0134] In another embodiment, step S21 includes the following steps:
[0135] S211: Combine the shallow low-level spectral-spatial feature map as shown in Figure 13 with the deep high-level spectral-spatial feature map as shown in Figure 14 for adaptive feature fusion (which can be performed through three methods: addition, element-wise multiplication, and concatenation) to obtain the fused feature map as shown in Figure 15 ;
[0136] S212: Update the weight values of the trainable parameter matrices W α , W β through backpropagation until the weight values reach the optimal values when the loss function converges to the minimum.
[0137] In this embodiment, the shallow low-level spectral-spatial feature map contains a large amount of detailed information such as shapes and boundaries, while the deep high-level spectral-spatial feature map contains more semantic information that can classify pixel categories. However, the segmentation map restored using only the high-level feature map is too rough and loses a lot of detailed information, severely limiting the model performance. During the upsampling process, incorporating low-level detailed features can make the restored feature map more refined. The fusion of low-level detailed features and high-level semantic features helps improve the accuracy of image segmentation, but it is difficult to determine the contribution degree of different-level features to the semantic segmentation task, that is, it is difficult to determine the fusion ratio. Therefore, this embodiment uses two trainable parameter matrices W α , W β in the adaptive feature fusion module to adaptively adjust the fusion ratio of different-level features. Among them, W α is used to control the fusion weight of features in the shallow low-level spectral-spatial feature map, and W β is used to control the fusion weight of features in the deep high-level spectral-spatial feature map, and their definitions are as follows:
[0138]
[0139]
[0140] where W α and W β respectively control the fusion weights of the shallow low-level spectral-spatial feature map and the deep high-level spectral-spatial feature map, α h,w,n represents the weight value at position (h, w, n), with the initial value set to 0.5, and β h,w,n represents the weight value at position (h, w, n), with the initial value set to |1 - |α||, and |·| represents the absolute value operation.
[0141] Shallow low-level spectral-spatial feature map l low and deep high-level spectral-spatial feature map l high are defined as follows respectively:
[0142]
[0143]
[0144] where l h,w,n represents the eigenvalue at the position (h, w, n) of the feature map l.
[0145] l after upsampling high and l low remain the same size. There are multiple feature fusion strategies, but the performance gains of different fusion strategies in specific target tasks vary greatly. For this reason, this embodiment explores the effects of three fusion strategies, namely addition, element-wise multiplication, and concatenation, on the model segmentation performance. Addition fusion directly incorporates the low-level detail features extracted from the shallow layer into the high-level feature map to make up for the loss of detail information. That is to say, it will increase the amount of information in each dimension of the high-level feature map, thus making the restored feature map more accurate. Element-wise multiplication fusion is similar to enhancing the high-level features according to the low-level feature distribution. Concatenation fusion concatenates the low-level features and the high-level features together, which will cause an increase in the number of feature maps (channels). The feature map U after fusion with different strategies is expressed as follows:
[0146]
[0147] U⊙ = W α l low ⊙W β l high
[0148] U [·,·] = [W α l low , W β l high
[0149] U ⊙ and U [·,·] respectively represent the feature map U after fusion with the addition, element-wise multiplication, and concatenation strategies. The parameter matrices W α and W β update the parameter values at the corresponding positions as the network learns. The initial value α of W α is set to 0.5, and the parameter value β of W β is restricted to |1 - |α||. Update W α , W α , W β according to the following update rules:
[0150]
[0151]
[0152] η represents the learning rate, represents the gradient of the cross-entropy loss function with respect to W α and I (h,w,n) represents the identity matrix of size (h, w, n).
[0153] For semantic segmentation tasks, the importance of the information expressed by different channels (i.e., different feature maps) is different. The fused feature map U ∈ R C×N×W×H , where C represents the number of channels. Since 3D convolution generates a feature map with one more spectral dimension compared to 2D convolution, the existing channel attention cannot be directly utilized. Therefore, the feature map U is first compressed channel by channel, and then a 1D convolution is used to learn the channel weight vector to enhance the channels that contribute to the target task and suppress the ineffective channels:
[0154] U CAM = σ[G(U) · W u
[0155] G(·) represents compressing the feature map U channel by channel, W u is the weight matrix, [·] represents the dot product operation, and U CAM is the learned channel weight vector.
[0156] In another embodiment, step S22 includes the following steps:
[0157] S221: Compress the fused feature map channel by channel to obtain a 1D feature map;
[0158] In this step, since the fused feature map U is four-dimensional, i.e., U ∈ R C×N×W×H , where C represents the number of channels, it needs to be compressed to one dimension, i.e., U ∈ R C , in order to use 1D convolution to learn the channel weight vector.
[0159] S222: Use 1D convolution to learn the channel weight vector U cAM of the 1D feature map to obtain a channel feature map containing different channel weights as shown in Figure 16 .
[0160] In this step, the channel weight vector U CAM is calculated as follows:
[0161] U CAM = σ[G(U) · W u
[0162]
[0163] Among them, G(·) represents compressing the feature map U channel by channel, and W u is the weight matrix, [·] represents the dot product operation, N, H, and W respectively represent the number of spectral bands, height, and width of the feature map U, n, h, and w respectively represent the nth spectral band, the hth row height, and the wth column width of the feature map U, and n ∈ N, 0 ≤ n < N, h ∈ H, 0 ≤ h < H, w ∈ W, 0 ≤ w < W.
[0164] In another embodiment, step S23 includes the following steps:
[0165] S231: Upsample the channel feature map using deconvolution to obtain a decoded map with restored spatial resolution;
[0166] In this step, the deconvolution formula is as follows:
[0167]
[0168] Among them, Y represents the input of deconvolution, represents the output of deconvolution, f represents the filter composed of convolution kernels, N represents the number of convolution kernels, n ∈ N, 0 ≤ n < N.
[0169] S232: Map the features of the decoded map with restored spatial resolution to the sample category space using two-dimensional convolution, and perform pixel-level prediction on the decoded map through the Softmax function to obtain the prediction probabilities of each pixel on the decoded map belonging to different categories, thereby obtaining the segmentation result.
[0170] In this step, the two-dimensional convolution formula is as follows:
[0171]
[0172] Among them, represents the value at the position (x + h, y + w) of the two-dimensional convolution input X 2d in the two-dimensional convolution, represents the value at the position (x, y) of the two-dimensional convolution output Y 2d in the two-dimensional convolution. represents the value at the position (h, w) of the two-dimensional convolution kernel K 2d in the two-dimensional convolution, and H and W respectively represent the height and width of the two-dimensional convolution kernel K 2d in the two-dimensional convolution, and h ∈ H, 0 ≤ h < H, w ∈ W, 0 ≤ w < W.
[0173] The Softmax function expression is as follows:
[0174] Softmax(x i ) = p(y = c|xi )
[0175]
[0176] Among them, x i represents the pixel feature of the decoded image, and p(y = c|x i ) represents the probability that the pixel belongs to the c-th class.
[0177] Next, the present disclosure verifies the segmentation effect of the above method on multi-spectral remote sensing images through experiments.
[0178] 1. Dataset Introduction
[0179] The present disclosure selects three datasets including ISPRS Potsdam, Qinghai, and Tibet Plateau to train and test the above model. Among them, the ISPRS Potsdam dataset contains 38 images, each image with a size of 6000×6000 and a ground sampling distance of 5 cm. The Potsdam dataset contains four spectral bands and six land cover classes: impervious surfaces, buildings, low vegetation, trees, cars, and clutter / background. 22 images are used as the training set, and the remaining 16 are used as the test set. The Qinghai dataset is an image of the Qinghai Lake area captured by the Landsat 8 satellite, with a ground sampling distance of 30 m. The image size is 5549×5004 and contains 7 spectral bands. Four classes are manually annotated for the dataset: grassland, water body, cloud, and bare land. The image is cropped into 2530 image patches of size 256×256, of which 50% are used for training and the rest for testing. The Tibet Plateau dataset contains 12 images, each image with an average size of 5029×5022 and a ground sampling distance of 30 m, and contains seven spectral bands. The dataset contains six classes: bare land, construction land, ice and snow and cloud, water body, grassland, forest, and shrub. 6 images with ID numbers 1, 3, 5, 7, 9, and 11 are used as the training set, and the remaining 6 are used as the test set.
[0180] 2. Evaluation Index Design
[0181] In this embodiment, the indexes for evaluating the model performance include:
[0182] Overall classification accuracy (OA), which represents the ratio of the number of correctly classified pixels to the total number of pixels, and is calculated by the following formula:
[0183]
[0184] Among them, TP c , FP c , TN c and FNc They respectively represent the true positive, false positive, true negative, and false negative of class c, and N represents the number of classes.
[0185] The F1 score (Mean F1) is the weighted harmonic mean of the precision and recall for each class, and is used to comprehensively evaluate the ability of the model to accurately retrieve and fully retrieve. It is calculated by the following formula:
[0186]
[0187] Among them, precision represents precision, recall represents recall, and
[0188] The formula for calculating precision is as follows:
[0189]
[0190] The formula for calculating recall is as follows:
[0191]
[0192] The mean intersection over union (MIoU) represents the average of the ratios of the intersections to the unions of the predicted results and the ground truths for all classes of the model. MIoU can better measure the ability of the model to accurately segment objects and is calculated by the following formula:
[0193] The mean intersection over union is calculated by the following formula:
[0194]
[0195] Among them, TP c , FP c , TN c and FN c They respectively represent the true positive, false positive, true negative, and false negative of class c, and N represents the number of classes.
[0196] 3. Experimental settings
[0197] The model runs on the pytorch platform, uses the Adam optimizer, and the weight decay is set to 0.0005. The initial learning rate is set to 0.0003, and a learning rate decay strategy is adopted, decaying once every 30 epochs, and training for a total of 60 epochs. The batch size is set to 32.
[0198] This disclosure compares the above model with deep learning methods including U-Net, ResNet34, MAResU-Net, and DeepLabv3+, as shown in Table 1 specifically:
[0199] Table 1
[0200]
[0201] This table lists the structural comparison of different methods. Among them, ResNet34 uses the same decoding structure as U-Net to implement the image segmentation task. The feature extraction backbone of Deeplabv3+ is resnet101. MAResU-Net uses resnet34 as the feature extraction backbone. BLASeNet-A, BLASeNet-M, and BLASeNet-C respectively represent the use of addition, element-wise multiplication, and concatenation fusion strategies during feature fusion. All methods are implemented on a computer equipped with an Intel Xeon Silver 4214 CPU, 128GB of running memory, and an RTX3090 GPU.
[0202] 4. Analysis of Experimental Results
[0203] The experimental results on the Potsdam dataset are shown in Table 2:
[0204] Table 2
[0205]
[0206] As can be seen from Table 2, the OA score of ResNet34 has increased by 0.73% compared to U-Net, indicating that the residual block can reduce information loss during the convolution process and obtain higher segmentation accuracy. The segmentation result maps of different methods are as Figure 17 shown, where the label map represents the true annotation map of the land cover classes in the Potsdam dataset. MAResU-Net improves the segmentation results by introducing an attention mechanism in the skip connection, indicating that the application of the attention mechanism can effectively model context information and handle complex scenarios. The different fusion strategies of BLASeNet generally obtain similar experimental results. Generally speaking, the results of BLASeNet-A are better. Compared with MAResU-Net, OA / Mean F1 / MIoU have increased by 1.33% / 2.02% / 1.80% respectively. It shows that the method described in this disclosure has the ability to capture the spectral bands rich in effective information in multi-spectral remote sensing images, and at the same time can extract discriminative spectral-spatial features.
[0207] The experimental results on the Qinghai dataset are shown in Table 3:
[0208] Table 3
[0209]
[0210] As can be seen from Table 3, on the "Cloud" category, BLASeNet-A achieved an F1 score of 77.08%, which was 7.53% higher than that of the second place. This indicates that the method described in the present disclosure can simultaneously model the context information in the spatial and spectral domains, and can also accurately identify some small targets with dispersed distributions. MIoU is used to evaluate whether the model can accurately predict the shape of the target object. Therefore, MIoU can better measure the ability of the model to accurately segment objects. The MIoU score of BLASeNet-A reached 77.71, which was 3.42% higher than that of MAResU-Net. The segmentation result maps of different methods are as follows Figure 18 shown, where the label map represents the true annotation map of the land cover classes in the Qinghai dataset. It can be seen that the edge contours of the segmentation maps generated by the method described in the present disclosure are clearer.
[0211] The experimental results on the Tibetan Plateau dataset are shown in Table 4:
[0212] Table 4
[0213]
[0214]
[0215] As can be seen from Table 4, due to the use of dilated convolution in Deeplabv3+, a lot of detailed information was lost while expanding the receptive field, resulting in a large number of misclassifications, and even an F1 score of 0 was obtained on the small target "Construction land". Figure 19 The segmentation result maps of different methods are shown, where the label map represents the true annotation map of the land cover classes in the Tibet Plateau dataset. As Figure 19 can be seen, the method described in the present disclosure retains the detailed features well by modeling the band-spatial position weights, and the generated segmentation maps are more refined, capable of capturing obvious contour features. The discrimination ability between "Grass" and "Construction land" has been greatly improved, and there are also obvious improvements in retaining object boundaries.
[0216] The method described in the present disclosure can effectively handle the segmentation of complex scenes in remote sensing images containing multiple spectral bands. The experimental results on the Potsdam, Qinghai, and Tibetan Plateau datasets demonstrate the advantages of this method: it can capture the spectral bands rich in effective information in MSIs, which helps to enhance the semantic segmentation results, and can better extract the features of small targets with dispersed distributions. By modeling the band-spatial position weights, the generated segmentation maps have clearer boundary contours.
[0217] The technical solutions provided by the present disclosure have been introduced in detail in combination with specific embodiments. At the same time, the description of the above embodiments is only used to help understand the method and its core idea of the present disclosure. For those of ordinary skill in the art, according to the idea of the present disclosure, there will be changes in the specific implementation manners and application scopes. Therefore, the content of this specification should not be construed as a limitation to the present disclosure.
Claims
1. A multispectral remote sensing image segmentation method based on band-position selection, comprising the following steps: S1: Obtain the multispectral remote sensing image to be segmented and input it into the multispectral remote sensing image segmentation model for encoding, where the encoding includes: S11: Capture the spectral bands rich in effective information in the multispectral remote sensing image to obtain a band-position feature map; S12: Extract spectral-spatial features from the band-position feature map to obtain a shallow low-level spectral-spatial feature map and a deep high-level spectral-spatial feature map; S2: Decode the deep high-level spectral-spatial feature map, where the decoding includes: S21: Perform adaptive feature fusion on the shallow low-level spectral-spatial feature map and the deep high-level spectral-spatial feature map to obtain a fused feature map; S22: Perform feature recalibration on the fused feature map to obtain a channel feature map; S23: Upsample the channel feature map to obtain a decoded map with restored spatial resolution, and perform pixel-level prediction on the decoded map to obtain the segmentation result map of the multispectral remote sensing image; Among them, in step S1, the multispectral remote sensing image segmentation model includes: A band-position adaptive selection module for capturing the spectral bands rich in effective information in the multispectral remote sensing image to obtain a band-position feature map; A three-dimensional residual module for extracting spectral-spatial features from the band-position feature map to obtain a shallow low-level spectral-spatial feature map and a deep high-level spectral-spatial feature map; An adaptive feature fusion module for performing adaptive feature fusion on the shallow low-level spectral-spatial feature map and the deep high-level spectral-spatial feature map to obtain a fused feature map; A channel attention module for performing feature recalibration on the fused feature map to obtain a channel feature map; A prediction module for upsampling the channel feature map to obtain a decoded map with restored spatial resolution, and for performing pixel-level prediction on the decoded map to obtain the segmentation result map of the multispectral remote sensing image; Step S11 includes: S111: Provide global context information for each spectral band in the multispectral remote sensing image through global average pooling to obtain a band map; S112: Based on the band map, obtain the band weight map of the multispectral remote sensing image through one-dimensional convolution; S113: Perform Hadamard product operation on the multispectral remote sensing image and the band weight map to obtain a three-dimensional feature map with different band weights; S114: Perform a mean operation on the three-dimensional feature map in the spectral dimension to obtain a spatial feature map that retains full-band information; S115: Based on the spatial feature map, obtain the position weight map of the multispectral remote sensing image through two-dimensional convolution; S116: Perform an element-wise multiplication operation on the position weight map and the three-dimensional feature map in the spatial dimension to obtain a band-position feature map containing spatial relationships.
2. The method according to claim 1, wherein, In step S112, the band weight map is obtained through the following formula: ; And ; Among them, represents the i-th input multispectral remote sensing image, and , represents the three-dimensional real space, F(:,,) represents the spectral dimension of the input multispectral remote sensing image, N, H, and W respectively represent the number of spectral bands, height, and width of the multispectral remote sensing image; h and w respectively represent the height of the h-th row and the width of the w-th column of the multispectral remote sensing image, and , ; represents the per-band global average pooling operation, represents a one-dimensional convolution operation with a convolution kernel size of 3.
3. The method according to claim 1, wherein In step S115, the position weight map is obtained through the following formula: ; Among them, represents the obtained spatial feature map, represents a two-dimensional convolution operation with a convolution kernel size of 3×3, represents the Sigmoid function.
4. The method according to claim 1, wherein In step S116, the position weight map and the three-dimensional feature map are subjected to an element-wise multiplication operation in the spatial dimension, which is expressed as: ; Among them, represents the obtained band-position feature map.
5. The method according to claim 1, wherein In step S12, extracting spectral-spatial features from the band-position feature map is achieved through multiple three-dimensional residual modules, where the calculation process of each three-dimensional residual module is as follows: ; Among them, represents the input of the l-th 3D residual module, represents the output of the l-th 3D residual module, represents the band-position adaptive selection layer, represents the residual function, represents the weight matrix of the l-th 3D residual module.
6. The method according to claim 1, wherein Step S21 includes the following steps: S211: Perform adaptive feature fusion on the deep high-level spectral-spatial feature map and the shallow low-level spectral-spatial feature map to obtain a fused feature map; S212: Update the trainable parameter matrix through backpropagation , of the weight value until the weight value reaches the optimum.
7. The method according to claim 1, wherein Step S22 includes the following steps: S221: Perform channel-wise compression on the fused feature map to obtain a one-dimensional feature map; S222: Learn the channel weight vector of the one-dimensional feature map using one-dimensional convolution to obtain a channel feature map containing different channel weights.
8. The method according to claim 1, wherein, Step S23 includes the following steps: S231: Use transposed convolution to upsample the channel feature map to obtain a decoded map with restored spatial resolution; S232: Use two-dimensional convolution to map the features of the decoded map with restored spatial resolution to the sample category space, and perform pixel-level prediction on the decoded map through the Softmax function to obtain the prediction probabilities of each pixel on the decoded map belonging to different categories, thereby obtaining the segmentation result.