A semantic segmentation method for high-resolution remote sensing images based on multi-scale deep supervision

By adopting a multi-scale deep supervision network architecture in semantic segmentation of high-resolution remote sensing images, combining the feature extraction methods of spatial branches and semantic branches, the problem of extracting and utilization of multi-scale feature information in complex land objects is solved, and higher segmentation accuracy and training efficiency are achieved.

CN119559403BActive Publication Date: 2025-05-13WUXI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510132150.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract and utilize multi-scale feature information of complex land objects in semantic segmentation of high-resolution remote sensing images, resulting in insufficient discontinuity and accuracy of segmentation results.

Method used

A network architecture based on multi-scale deep supervision is adopted, and a spatial branch and semantic branch are designed, combined with spatial de-redundant residual module and residual attention Mamba, feature extraction and fusion are performed, and a large kernel grouping feature extraction module and multi-scale convolution attention module are used in the multi-scale feature aggregation part to suppress invalid features and activate effective features.

Benefits of technology

It improves the accuracy and training efficiency of semantic segmentation, can more effectively capture and utilize multi-scale feature information in remote sensing images, and improves the continuity and overall performance of segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559403B_ABST
    Figure CN119559403B_ABST
Patent Text Reader

Abstract

The present invention provides a high-resolution remote sensing image semantic segmentation method based on multi-scale deep supervision, comprising the following steps: step 1, obtaining ISPRS Potsdam data set and Vaihingen data set; step 2, building a multi-scale deep supervision network MDSNet, including a spatial branch, a semantic branch network and a multi-scale feature aggregation part; the spatial branch, the semantic branch network and the multi-scale feature aggregation part are all four-layer structures; step 3, inputting a training set into the multi-scale deep supervision network MDSNet for training, calculating a loss function and performing back propagation, updating network parameters, and obtaining a trained parameter model; step 4, inputting a test set into the trained parameter model, and outputting a segmentation map of a target image. The method proposed by the present invention improves the accuracy of semantic segmentation, as well as training efficiency and stability while improving the feature extraction capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a high-resolution remote sensing image semantic segmentation method based on multi-scale deep supervision. Background Art

[0002] With the rapid development of sensor technology and remote sensing platforms, a large number of high-resolution remote sensing images have been widely used in land cover analysis, urban planning, change detection, scene classification, environmental monitoring and other fields. High-resolution images contain repeated spatial detail information, which can clearly capture the boundaries of objects and reflect the complex semantic features of objects. Such fine-grained image data brings new opportunities and challenges to remote sensing semantic segmentation. Compared with image classification, semantic segmentation can provide more complex feature information. However, the effective extraction and utilization of complex features, complex object distribution, multi-scale changes and imbalance between categories significantly increase the difficulty of segmentation tasks.

[0003] With the development of artificial intelligence, the semantic segmentation of natural images has made rapid progress, and the accuracy of semantic segmentation methods based on deep learning is far superior to traditional methods. For the semantic segmentation of remote sensing images, semantic segmentation is more challenging than natural images, especially because objects of the same category may have significant differences in features, while objects of different categories may have similar features. The fully convolutional network (FCN) truly realizes pixel-level segmentation. Many researchers have proposed the method of fully convolutional network to solve the problem of semantic segmentation of high-resolution remote sensing images. Due to the limitations of convolution, it cannot observe the surrounding information (context information) as extensively as humans, but mainly judges the category by learning the texture of the object, but the lack of context learning will lead to misclassification; secondly, there are many shadows in remote sensing images. Large shadows will lead to small feature differences between different categories, increasing the difficulty of judgment. Small shadows will truncate the continuous features within a single object, resulting in discontinuity of the segmentation results. Convolutional neural networks have a certain depth of feature extraction layers. As the feature extraction layer goes deeper, the output high-level features have rich semantic information and good continuity, while shallow feature extraction has rich spatial information and more accurate target boundaries, but poor continuity. Therefore, it is particularly important to effectively utilize high-level features and low-level features. In recent years, attention mechanism has been widely used in computer vision tasks as an effective feature enhancement method, especially in semantic segmentation. The attention mechanism guides the model to focus on important areas and ignore irrelevant information by dynamically adjusting feature weights, thereby achieving effective utilization of contextual information. There are two ways to implement the attention mechanism, namely, the attention mechanism based on information aggregation and the attention mechanism based on unit similarity. The information aggregation-based method generates global contextual information by globally aggregating features of specific dimensions (such as channels, spaces, etc.). This information is mapped to generate attention weights, which are used to weight the original features. Multi-scale feature fusion is usually used to fuse high-level features and low-level features. The existing feature extraction mode and multi-scale information fusion method can capture contextual information and multi-scale information well, but there are still some defects. In the process of feature extraction, usually only the difference between local and global features is paid attention to, while the difference of the features themselves is ignored, that is, the effective extraction of semantic information and spatial information, and the complex feature information of high-resolution remote sensing images cannot be effectively utilized. In addition, most multi-scale information fusion methods only use simple operations such as splicing, addition, dot multiplication, etc. The fused features lack specificity and ignore the fusion differences of information at different scales. Summary of the invention

[0004] Purpose of the invention: The technical problem to be solved by the present invention is to provide a high-resolution remote sensing image semantic segmentation method based on multi-scale deep supervision in view of the shortcomings of the prior art, comprising the following steps: Step 1, obtaining the ISPRS Potsdam dataset and the Vaihingen dataset, and dividing the datasets into a training set, a validation set and a test set after uniform preprocessing.

[0005] Step 2: Build a multi-scale deep supervision network MDSNet (Multi-scale Deep Supervision Network, MDSNet), including spatial branches, semantic branch networks and multi-scale feature aggregation parts; the spatial branches, semantic branch networks and multi-scale feature aggregation parts are all four-layer structures.

[0006] Step 3: Input the training set into the multi-scale deep supervision network MDSNet for training, calculate the loss function and perform back propagation, update the network parameters, and obtain the trained parameter model.

[0007] Step 4: Input the test set into the trained parameter model and output the segmentation map of the target image.

[0008] Step 1 includes data augmentation on the Vaihingen dataset by random sampling and cropping.

[0009] Step 2 includes: the spatial branch includes a four-layer spatial de-redundant residual module, the outputs of the four-layer spatial de-redundant residual module are S1, S2, S3 and S4 from shallow to deep, and the output of the spatial de-redundant residual module is B×A×H×W four-dimensional features, where B is the batch size, A is the number of channels, H is the length, and W is the width.

[0010] The number of channels of the second-layer spatial de-redundant residual module is doubled compared with the number of channels of the first-layer spatial de-redundant residual module, and the length and width are halved; the number of channels of the third-layer spatial de-redundant residual module is doubled compared with the number of channels of the second-layer spatial de-redundant residual module, and the length and width are halved; the number of channels of the fourth-layer spatial de-redundant residual module is doubled compared with the number of channels of the third-layer spatial de-redundant residual module, and the length and width are halved.

[0011] S2, S3 and S4 are connected through channel reweighting, the degree of difference of each channel number is determined according to the learnable parameters of different feature maps, the importance of the channel is adjusted dynamically, and finally the output M4 of the channel reweighting connection is used as part of the input of the fourth layer in the multi-scale feature aggregation part.

[0012] Step 2 also includes: the semantic branch network includes four layers of residual attention Mamba, the residual attention Mamba is used to extract fine-grained features, the residual attention Mamba includes residual Transformer and Mamba, the residual Transformer focuses on global semantic modeling, thereby achieving comprehensive integration of multi-scale information, and the output of the four-layer residual attention Mamba is C1, C2, C3 and C4 from shallow to deep, and the number of channels and length and width changes are consistent with the spatial branch.

[0013] A double-layer convolution block is used before the spatial branch and the semantic branch to roughly extract features. Each convolution block contains a regular convolution, batch normalization, and ReLU activation function.

[0014] The multi-scale feature aggregation part includes a large kernel group feature extraction module and a multi-scale convolution attention module, wherein the large kernel group feature extraction module is used to selectively fuse and extract three types of features: spatial branch output S, upsampling and channel reweighted output M, and semantic branch output C. The large kernel group feature extraction module gradually combines feature mapping with attention coefficients. Since M comes from deep features, it is difficult for ordinary convolution to focus on features of a specific scale. In this regard, by introducing group convolution to integrate features from shallow layers (S, C) and deep layers, the grouping mechanism of group convolution limits the range of feature interaction, effectively reducing the interference of redundant features. At the same time, group convolution only performs convolution operations on channels within the group, significantly reducing the number of parameters and computational overhead; the features after group convolution are batch normalized and feature fused, and then the feature fusion results are subjected to ordinary convolution and spliced ​​with the initial spatial branch output S and semantic branch output C, and finally sent to the multi-scale convolution attention module.

[0015] The features output by the spatial branch and semantic branch networks are connected to the three-layer efficient upsampling convolution block using a four-layer multi-scale convolution feature fusion module. The outputs of the three-layer efficient upsampling convolution blocks are M1, M2 and M3 from shallow to deep. The output of each layer of the multi-scale convolution feature fusion module is input into the segmentation head for segmentation. The segmentation head includes a normal convolution block with a 3×3 convolution kernel, a normal convolution with a 1×1 convolution kernel and an upsampling. The segmentation head has four layers, corresponding to the four-layer multi-scale convolution feature fusion module.

[0016] Step 3 includes: transforming the feature dimension into The remote sensing image is sent into the multi-scale deep supervision network MDSNet, where b is batch_size, batch_size represents the total number of images trained in a batch; 3 represents the number of channels; the remote sensing image is sent into the double-layer convolution block, and the output feature dimension becomes , complete the rough extraction of the original image features.

[0017] The output of the double-layer convolution block is sent to the spatial branch and semantic branch network; for the spatial branch, the first-layer spatial redundancy removal residual module includes a left half and a right half residual part, wherein the left half includes horizontal and vertical convolutions, and the right half residual part is a residual structure; first The features of are horizontally and vertically convolved, and the convolution kernels are 3×1, 3×3 and 1×3 respectively, and the convolution results of three different convolution kernels of 3×1, 3×3 and 1×3 are concatenated to generate a size of The four-dimensional features containing the directional features are sent to the right half residual part of the first-layer spatial redundant residual module, and the four-dimensional features are sent to the convolution block without ReLU and the coordinate convolution block without ReLU. The convolution kernels of the convolution block without ReLU and the coordinate convolution block without ReLU are 1×1 and 3×3 respectively, and the output results are The four-dimensional features of the convolution block without ReLU are input into the wavelet transform convolution. The convolution kernel of the wavelet transform convolution is 5×5, and the output result is Finally, the output four-dimensional features of the convolution block without ReLU, the coordinate convolution block without ReLU and the wavelet transform convolution are added together, and then passed through a ReLU activation function to obtain the result: The four-dimensional feature S1; the coordinate convolution block removing ReLU includes coordinate convolution and batch normalization.

[0018] The processing process of the second, third and fourth layer spatial redundant removal residual modules is the same as that of the first layer spatial redundant removal residual module, and the output results are The four-dimensional features S2, The four-dimensional feature S3 and The four-dimensional feature S4 of is:

[0019] ,

[0020] .

[0021] in is the result after the feature expression is enhanced. It is the output result of the spatial redundancy removal residual module; For splicing operation, is a convolutional block, containing Regular convolution, batch normalization, and ReLU activation functions for large and small convolution kernels; is the activation function, To remove the convolution block of the ReLU activation function, for Wavelet transform convolution with large and small convolution kernels; is the coordinate convolution block, including Coordinate convolution, batch normalization and ReLU activation functions of large and small convolution kernels; To remove the coordinate convolution of the ReLU activation function, m and n are positive integers.

[0022] The multi-scale deep supervision network MDSNet uses the outputs S2, S3, and S4 of the last three feature extraction layers for feature fusion, upsamples the input feature M4 according to the standard of S4, and concatenates the upsampled feature maps according to the first dimension (in the python code, the first dimension is 1, [0, 1, 2, 3] is the sorting, that is, the number of C channels) to obtain the concatenated feature map, and adds the number of channels of the input features, sets a learnable weight parameter with the same number of channels, and normalizes the learnable weight parameter, multiplies the concatenated feature map by the normalized weight parameter, realizes the weighted operation of the input features, and obtains the weighted features, concatenates the weighted features, and obtains the convolution operation with a step size of 4 and a convolution kernel of 4×4. The four-dimensional feature M4.

[0023] For the semantic branch network, residual attention Mamba includes residual Transformer and Mamba. Residual Transformer directly models the global relationship of the input image through the self-attention mechanism, thereby capturing long-distance dependencies. At the same time, the residual link retains the local features extracted by the upper-layer Mamba (starting from the second layer of the semantic branch, Mamba is above the residual Transformer).

[0024] The residual Transformer includes an efficient attention block ETB, which uses an efficient multi-head self-attention mechanism to obtain the attention score. The result of the double-layer convolution block is input into the residual Transformer, and the output result is The four-dimensional features of are:

[0025] ,

[0026] ,

[0027] ,

[0028] .

[0029] in is the output of efficient multi-head self-attention, Q represents the query vector, K represents the key vector, V represents the numerical vector, represents an efficient multi-head self-attention mechanism after the fully connected layer, Normalize the instance. is a convolution with a convolution kernel of 1×1, and T represents transpose; is a fully connected feed-forward network layer, and x is the input feature; , is the weight matrix of the linear transformation; , is the bias term, is a nonlinear activation function, It is a residual connection structure for efficient multi-head self-attention; is the fully connected layer; y is the final output of the residual Transformer;

[0030] After the residual transformer extracts the global semantic features, y is input into the visual state space (VSS) module of the first layer residual attention Mamba to extract local semantic features of y. The VSS module includes a state space model (SSM). The output of the residual transformer is input into the residual attention Mamba to obtain The processing of the second, third and fourth layers of residual attention Mamba is the same as that of the first layer of residual attention Mamba, and the output results are The four-dimensional feature C2, The four-dimensional feature C3 and The four-dimensional feature C4 of is:

[0031] ,

[0032] ,

[0033] .

[0034] in Representation layer normalization, Indicates the Silu function activation operation, and They are the output of Silu function and the normalized output of LN layer, which serve as transition. is a depth-wise separable convolution, is the product operation, This is the final output result of the residual attention Mamba.

[0035] Finally, S1, S2, S3, S4 and C1, C2, C3, C4, M4 are sent to the multi-scale feature aggregation part. The input of the fourth layer of multi-scale feature aggregation is S4, C4 and M4. The three four-dimensional features S4, C4 and M4 are sent to the large kernel group feature extraction module. Each input undergoes a group convolution with a convolution kernel of 3×3 and batch normalization, and the output is The four-dimensional features of S4, C4 and M4 are then concatenated after group convolution and batch normalization, and then the ReLU activation function is passed, the convolution block of ReLU is removed, and the Sigmoid operation is performed. The output result is The characteristics of The features of are Hadamard products with S1 and C1, and the output is The four-dimensional features of The four-dimensional features of are input into the multi-scale convolutional attention module, and we get The four-dimensional feature M3 is used as one of the inputs of the third-layer multi-scale feature aggregation. The multi-scale feature aggregation process of the third, second, and first layers is the same as that of the fourth layer. The output results of the multi-scale feature aggregation of the third, second, and first layers are respectively The four-dimensional features of The four-dimensional features of The four-dimensional features of the multi-scale feature aggregation are sent to the segmentation head respectively, and the output results are The four-dimensional features of , 6 represents the six categories of segmentation.

[0036] Step 3 also includes: using the loss function to evaluate the final segmentation result The four-dimensional features of the second, third, and fourth layer segmentation heads are supervised as auxiliary losses, and the auxiliary losses are used to correct the main losses. The formula is:

[0037] ,

[0038] ,

[0039] ,

[0040] .

[0041] in, is the Dice Loss loss; X represents the pixel label of the real segmentation image, Y represents the pixel category of the segmentation image predicted by the multi-scale deep supervision network MDSNet, |X|∩|Y| is approximately the dot product between the pixels of the predicted image and the pixels of the real label image, and the dot product results are added, |X| and |Y| are respectively approximately the sum of the pixels in the corresponding images of the predicted image and the real label image; is the cross entropy loss; N is the number of pixels in the image, C is the number of categories, is the true label of the ith pixel, is a binary indicator function indicating whether the i-th pixel belongs to category c. If so, is 1, otherwise is 0; represents the predicted probability of the multi-scale deep supervision network MDSNet for the i-th pixel belonging to category c; The main loss is is the auxiliary loss output by the second-layer segmentation head; is the auxiliary loss output by the third-layer segmentation head; is the auxiliary loss output by the fourth-layer segmentation head; is the weighting coefficient, is the final total loss value.

[0042] Finally, the data set is trained and iterated according to step 3, with the number of iterations being 300. The weight parameter files of each iteration (files ending in .pth, the trained parameter model is the best parameter model, and the best parameter model is the model using the best weight parameter file best_epoch_weights.pth) are all saved in the diary. The best weight parameter file is determined by comparing the validation loss value and the training loss value during training. If the validation loss value and the training loss value become smaller at the same time, it is marked as the best weight parameter file.

[0043] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the described method.

[0044] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction is run on a computer, the steps of the method described are executed.

[0045] In view of the problems existing in the semantic segmentation of high-resolution remote sensing images mentioned in the background technology, the present invention proposes a method based on a multi-scale deep supervision network (MDSNet). First, different feature extraction branches are designed for spatial and semantic features respectively to make full use of the semantic and spatial information of remote sensing images; secondly, the proposed spatial de-redundancy residual module is added to the spatial branch, and wavelet transform and coordinate convolution are introduced to extract spatial features more specifically and pay better attention to edge features. Then, the proposed residual attention Mamba is added to the semantic branch to realize the semantic feature extraction at the global level; finally, the feature fusion part adopts a multi-scale feature fusion mechanism, and a large kernel grouping feature extraction module is designed to fuse the spatial branch, semantic branch features and deep-level features layer by layer, suppress invalid features, and activate effective features; in addition, this model adopts a deep supervision mechanism to add auxiliary supervision heads to the feature fusion layers at each stage to improve training efficiency.

[0046] The present invention has the following beneficial effects: (1) In view of the differences between semantic information and spatial information, a spatial feature extraction branch and a semantic feature extraction branch are designed respectively. Secondly, in view of spatial information, a spatial redundancy removal residual module is designed in the process of spatial feature extraction, and vertical and horizontal convolution kernels are introduced to obtain directional features in the vertical and horizontal directions, thereby capturing more spatial detail features. After that, the residual structure is connected, and coordinate convolution and wavelet transform convolution layers are introduced to further screen the spatial information, and the spatial branch outputs of the last three layers are re-weighted and connected for feature fusion.

[0047] (2) In the process of semantic feature extraction, residual attention Mamba is designed to focus more on local and global semantic context information. Finally, in the multi-scale feature fusion part, a multi-scale feature fusion mechanism is adopted, and a large kernel grouping feature extraction module is designed to fuse spatial branch and semantic branch features with deep features layer by layer, suppress invalid features, and activate valid features.

[0048] (3) A deep supervision mechanism is adopted to add auxiliary supervision heads to the feature fusion layers at each stage to improve training efficiency.

[0049] In summary, the method proposed in the present invention improves the feature extraction capability while improving the accuracy of semantic segmentation, as well as the training efficiency and stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.

[0051] Figure 1 An example of the International Society for Photogrammetry and Remote Sensing dataset is provided for the embodiment.

[0052] Figure 2 A structural diagram of a multi-scale deep supervision network provided for an embodiment.

[0053] Figure 3 A structural diagram of a spatial redundancy removal residual module provided in an embodiment.

[0054] Figure 4 A visualization diagram of feature expressiveness enhancement provided for an embodiment.

[0055] Figure 5 A structural diagram of the channel reweighted connection provided for an embodiment.

[0056] Figure 6 A structural diagram of the residual attention Mamba provided for the implementation example.

[0057] Figure 7A multi-scale feature aggregation structure is provided for an embodiment.

[0058] Figure 8 A diagram showing a visual qualitative comparison of the segmentation method provided in the embodiment with other methods on the Potsdam dataset.

[0059] Fig. 9 A histogram comparing the recall rates of each category of the segmentation method provided in the embodiment with other methods on the Vaihingen dataset. DETAILED DESCRIPTION

[0060] like Figure 2 As shown, the embodiment of the present invention provides a high-resolution remote sensing image semantic segmentation method based on multi-scale deep supervision, including the following steps: Step 1, obtaining the ISPRS Potsdam dataset and the Vaihingen dataset, the dataset sample image is as follows Figure 1 As shown in the figure (from left to right are the orthophoto of the Vaihingen dataset. The corresponding label of Vaihingen. The orthophoto of the Potsdam dataset. The corresponding label map of Potsdam), the datasets are uniformly preprocessed and divided into training set, validation set and test set in the ratio of 8:1:1. The Potsdam dataset contains 20102 remote sensing images with a spatial size of 256×256; the Vaihingen dataset contains 19335 remote sensing images with a spatial size of 256×256.

[0061] Step 2: Build MDSNet, including spatial branch, semantic branch and multi-scale feature aggregation part.

[0062] Step 3: Input the training set into MDSNet for training, calculate the loss function and perform back propagation, update the network parameters, and obtain the trained parameter model.

[0063] Step 4: Input the test set into the trained parameter model and output the segmentation map of the target image.

[0064] Step 1 includes: the acquisition of the data set is specifically: acquiring the ISPRS Potsdam data set and the Vaihingen data set, both of which are widely used high-resolution remote sensing image data sets. The International Society for Photogrammetry and Remote Sensing (ISPRS) Potsdam data set contains 38 high-resolution remote sensing images, all with a spatial size of 6000×6000 pixels, a spatial resolution of 5 cm, and bands including red (R), green (G), blue (B), and near infrared bands (NIR). The data set contains six categories that can be used for semantic segmentation. The ISPRS Vaihingen Dataset is released by the same organizer and has similar properties. It contains six categories and four optical bands. Unlike the Potsdam data set, the ground sample distance (GSD) is lower. The Vaihingen data set contains 33 high-resolution remote sensing images, with a spatial size of approximately 2500×2000 pixels and a spatial resolution of 9 cm.

[0065] The preprocessing and data set division are as follows: the Vaihingen data set has only 3135 images, which is too far from the capacity of the Potsdam data set. Using the same number of iterations is likely to cause overfitting problems. Therefore, the Vaihingen data set is enhanced by random sampling and cropping data enhancement, expanding the data set capacity to 19335 images, with an overlap rate of 0.7.

[0066] like Figure 2 As shown in the figure (SDCM is a spatial redundant removal residual module. CRC is a channel reweighted connection. ResT-Mamba is a residual attention Mamba. EUCB is an efficient upsampling convolution block. MCS-FFM is a multi-scale convolution feature fusion module. SH is a segmentation head. Principal Loss is the main loss. Auxiliary Loss is the auxiliary loss. CBR is a convolution block. Conv is a normal convolution. BN is batch normalization. ReLU is an activation function. Upsamling is upsampling. DWconv is a depth-separable convolution), step 2 includes: step 2.1: a double-layer convolution block is used before the spatial branch and the semantic branch to roughly extract features, and the convolution block contains one regular convolution, batch normalization and ReLU;

[0067] Step 2.2: Build a spatial branch, which consists of four layers of spatial de-redundant residual modules (Space De-redundant Convolution Module, SDCM). The output of each layer is S1, S2, S3 and S4 from shallow to deep. The number of channels is doubled, and the length and width are halved. As the spatial de-redundant residual module goes deeper, the spatial information gradually decreases. Therefore, S2, S3 and S4 are channel reweighted (Channel Reweight Concat, CRC). According to the difference between different feature maps, the importance of the channel is dynamically adjusted to better integrate the feature information of different numbers of channels. Finally, the output M4 of the channel reweighted connection is used as part of the fourth layer input in the multi-scale feature aggregation part, balancing the ratio of semantic information to spatial information.

[0068] Step 2.3: Build a semantic branch network, which includes a four-layer residual attention Mamba (ResT-Mamba). Residual attention Mamba is used to extract fine-grained features. Residual attention Mamba includes residual Transformer and Mamba. Residual Transformer focuses on global semantic modeling, thereby realizing the comprehensive integration of multi-scale information. The output of the four-layer residual attention Mamba is C1, C2, C3 and C4 from shallow to deep, and the number of channels and length and width changes are consistent with the spatial branch;

[0069] Step 2.4: The output features of the spatial branch and semantic branch of step 2.1 and step 2.2 are connected to the three-layer efficient up-convolution block (Efficient Up-convolution Block, EUCB) using the Multi-scale Convolutional Feature Fusion Module (MCS-FFM). The outputs of the efficient up-convolution block are M1, M2 and M3 from shallow to deep. The output of each layer of the multi-scale convolution feature fusion module is input into the segmentation head (SH) for segmentation. The segmentation head consists of a common convolution block with a 3×3 convolution kernel, a common convolution with a 1×1 convolution kernel, and an upsampling.

[0070] Step 3 includes: transforming the feature dimension into The remote sensing image is sent into MDSNet, where is batch_size, which is set to 1 here, indicating the total number of images in a batch training; 3 indicates the number of channels, is the length and width. Send it into the double-layer convolution block (CBR), and the output feature dimension becomes , complete the rough extraction of the original image features.

[0071] Secondly, the output of the double-layer convolutional block is sent to the spatial branch and the semantic branch. For the spatial branch, Figure 3 As shown in (CB is the convolution block without ReLU, CoCBR is the coordinate convolution block. CoCB is the coordinate convolution block without ReLU. CoordConv is the coordinate convolution. WTConv2d is the wavelet transform convolution), first The features are convolved horizontally and vertically, such as Figure 4 As shown, the convolution kernels are 3×1, 3×3, and 1×3, and the three are concatenated to generate a size of The four-dimensional features containing directional features are sent to the right half residual part of the spatial redundant residual module, and the four-dimensional features are sent to CB and CoCBR respectively, with convolution kernels of 1×1 and 3×3 respectively, and the output results are The four-dimensional features of CB are input into the wavelet transform convolution, the convolution kernel is 5×5, and the output result is Finally, the three four-dimensional features are added together and activated by a ReLU function. The result is The second, third and fourth layers are the same as the first layer, and the output results are The four-dimensional features S2, The four-dimensional feature S3 and The four-dimensional feature S4 of is:

[0072] ,

[0073] .

[0074] in is the result after the feature expression is enhanced. It is the output result of the spatial redundancy removal residual module; For splicing operation, is a convolutional block, containing Regular convolution, batch normalization, and ReLU activation functions for large and small convolution kernels; is the activation function, To remove the convolution block of the ReLU activation function, for Wavelet transform convolution with large and small convolution kernels; is the coordinate convolution block, including Coordinate convolution, batch normalization and ReLU activation functions of large and small convolution kernels; To remove the coordinate convolution of the ReLU activation function, m and n are positive integers.

[0075] Since the first layer of the feature extraction layer contains a large proportion of spatial information and a large amount of invalid edge information, using the features of this layer will cause an imbalance in the spatial semantic ratio of the subsequent multi-scale convolutional feature fusion module. At the same time, these invalid edge information will also cause more segmentation errors. Therefore, MDSNet uses the outputs of the last three feature extraction layers S2, S3, and S4 for feature fusion. Figure 5 As shown in the figure, since the lengths and widths of S2, S3 and S4 are different, the input features are upsampled according to the standard of S4, the upsampled feature maps are spliced ​​according to the first dimension, a learnable weight parameter with the same number of channels is set, and it is normalized. The spliced ​​feature map is multiplied by the normalized weight parameter to implement the weighted operation of the input features, these weighted features are spliced, and the final result is obtained by a convolution operation with a step size of 4 and a convolution kernel of 4×4. The result is Four-dimensional feature M4.

[0076] For semantic branches, such as Figure 6 As shown in the figure (MatMul is the product. Linear is the fully connected layer. SiLU is the activation function), the residual attention Mamba is mainly divided into two parts: residual transformer and Mamba. The left part is residual transformer and the right part is Mamba. The residual transformer is a new architecture that combines the local learning ability of the residual network with the global modeling ability of the transformer. It has strong global semantic feature extraction capabilities. Through the self-attention mechanism, it can directly model the global relationship of the input image, thereby capturing long-distance dependencies. At the same time, the residual link retains the local features extracted by the upper-layer Mamba, so that the model does not lose too much detail information when learning global features. The residual transformer is centered on the efficient attention block (ETB). In order to save computing costs, the efficient attention block uses an efficient multi-head self-attention mechanism to obtain the attention score. The result of the double-layer convolution is input into the residual transformer, and the output result is The four-dimensional features of . The specific formula is as follows:

[0077] ,

[0078] ,

[0079] ,

[0080] .

[0081] in is the output of efficient multi-head self-attention, Normalize the instance. The convolution kernel is 1×1, which is used to improve efficiency. It is a fully connected feed-forward network layer (FFN). , is the weight matrix of the linear transformation. , is the bias term, is a nonlinear activation function, It is a residual connection structure for efficient multi-head self-attention. is a fully connected layer. is the final output of the residual Transformer.

[0082] After the residual transformer extracts the global semantic features, it inputs them into the VSS module of Mamba to extract local semantic features, thus strengthening the semantic branch's ability to retain local features. Figure 6 As shown on the right, SSM is the core computing unit of the VSS module. The output of the upper layer residual transformer is input into Mamba to obtain The four-dimensional feature C1 of the second, third and fourth layers are consistent with the first layer process, and the output results are The four-dimensional feature C2, The four-dimensional feature C3 and The four-dimensional feature C4 of is:

[0083] ,

[0084] ,

[0085] .

[0086] Where y is the output feature of the residual Transformer. Representation layer normalization, Indicates the Silu function activation operation, is a depth-wise separable convolution, is the product operation, This is the final output result of the residual attention Mamba.

[0087] Finally, S1, S2, S3, S4 and C1, C2, C3, C4, M4 are sent to the multi-scale feature aggregation part. The input of the fourth layer of multi-scale feature aggregation is S4, C4 and M4, as shown in Figure 7As shown in the figure, the three four-dimensional features are sent to the large kernel group feature extraction module. Each input undergoes a group convolution with a convolution kernel of 3×3 and batch normalization, and the output is The four-dimensional features are then concatenated and processed by ReLU, CB and Sigmoid operations, and the output result is Finally, the result is Hadamard product with S1 and C1, and the output is The four-dimensional features of are finally input into the multi-scale convolutional attention module to obtain The four-dimensional feature M3 is used as one of the inputs of the third layer of multi-scale feature aggregation. The multi-scale feature aggregation process of the third, second, and first layers is the same as that of the fourth layer and is not listed here. The final output results are , , The multi-scale feature aggregation results are sent to the segmentation head, which consists of a common convolution block with a 3×3 convolution kernel, a common convolution with a 1×1 convolution kernel, and an upsampling. The output results are The four-dimensional features of , 6 represents the six categories of segmentation.

[0088] The loss function is used to supervise the final segmentation result and the remaining three layers of rough segmentation results as auxiliary losses, and the auxiliary losses are used to correct the main losses. Therefore, the main loss and auxiliary losses are supervised using the same loss function. The cross-entropy loss (CE Loss) and Dice Loss are weighted summed. Dice Loss has good robustness to unbalanced categories and can help the model learn the distinction between categories in a more balanced way. Since Dice Loss itself has a large value, as the training progresses, Dice Loss is prone to large oscillations. Therefore, the present invention introduces the cross-entropy loss function for neutralization. The formula is:

[0089] ,

[0090] ,

[0091] ,

[0092] .

[0093] in, is the Dice Loss loss; X represents the pixel label of the real segmentation image, Y represents the pixel category of the segmentation image predicted by the multi-scale deep supervision network MDSNet, |X|∩|Y| is approximately the dot product between the pixels of the predicted image and the pixels of the real label image, and the dot product results are added, |X| and |Y| are respectively approximately the sum of the pixels in the corresponding images of the predicted image and the real label image; is the cross entropy loss; N is the number of pixels in the image, C is the number of categories, is the true label of the ith pixel, is a binary indicator function indicating whether the i-th pixel belongs to category c. If so, is 1, otherwise is 0; Represents the predicted probability of the multi-scale deep supervision network MDSNet for the i-th pixel belonging to category c; the accuracy of the model prediction is measured by calculating the cross entropy between the predicted probability of each category and the actual label. The main loss is is the auxiliary loss output by the second-layer segmentation head; is the auxiliary loss output by the third-layer segmentation head; is the auxiliary loss output by the fourth-layer segmentation head; is the weighting coefficient used to balance the impact of the two loss functions, which is set to 1.0 here; is the final total loss value.

[0094] The visual qualitative comparison of the MDSNet provided in this embodiment with other methods on the Potsdam dataset is shown in the figure below: Figure 8 As shown, from left to right are the original image, label image, U-Net segmentation map, EIGNet segmentation map, DeepLabV3+ segmentation map, MAResU-Net segmentation map, CMLFormer segmentation map, A2FPN segmentation map, CGGLNet segmentation map, CMTFNet segmentation map, PSPNet segmentation map and MDSNet segmentation map. There are five groups of comparisons in the figure, and images containing most categories are selected to better observe the segmentation effect. Taking the remote sensing image in the third row as an example, the image contains all categories in the data set except the background. MDSNet can complete the semantic segmentation of the original image well, especially for the edge segmentation effect. It can be proved that MDSNet can still maintain a good segmentation effect during overall segmentation.

[0095] As shown in Table 1, the evaluation indicators are mean F1 score (mF1), mean Intersection over Union (mIoU), Overall Accuracy (OA) and mean Recall (mRecall). The quantitative results of the trunks of different branches of the Vaihingen dataset are used to verify the effectiveness of the SDCB proposed in the spatial branch and the ResT-Mamba proposed in the semantic branch. In the spatial branch, the present invention selects a five-stage double-layer convolution (Double Conv), and in the semantic branch, the present invention selects a five-stage context embedding block (Context Embedding Block, CEB). This module is proposed by BiSeNetV2 and uses global average pooling and residual connections to effectively embed global context information. For the spatial branch, it is particularly important to accurately extract spatial information, and the five-stage double-layer convolution not only causes feature redundancy, but also has no specificity for the extraction of spatial information. As the number of convolutional layers increases, the proportion of semantic information gradually increases. The coordinate convolution in the SDCB increases the module's perception of spatial information, and the wavelet transform increases the ability to extract edge details. The ablation study shows that when the semantic branch is CEB, after the spatial branch is changed from Double Conv to SDCB, the network's mIoU increases by 0.42%, mF1 increases by 0.25%, mRecall increases by 0.30%, and OA increases by 0.15%. This increase proves the effectiveness of SDCB for spatial feature extraction.

[0096] Table 1

[0097] For the semantic branch, the semantic branch needs a larger receptive field to capture high-level semantic information. CEB increases the receptive field and optimizes the extraction of global contextual semantic information, but ignores the extraction of local contextual semantic information. ResT-Mamba combines global and local contextual semantic information for extraction. The residual link retains the local features extracted by the upper-layer Mamba and optimizes the feature extraction process. The ablation study structure shows that when the spatial branch is Double Conv, after the semantic branch is replaced by ResT-Mamba from CEB, the network's mIoU increases by 0.20%, mF1 increases by 0.11%, mRecall increases by 0.16%, and OA increases by 0.12%. This fully proves the effectiveness of ResT-Mamba.

[0098] Secondly, in order to verify the effectiveness of each module, the present invention gradually adds a single module or mechanism through the baseline model. Since the effectiveness of SDCB and ResT-Mamba as the backbone of each branch has been proved in Table 2, this group of module ablation experiments uses SDCB as the backbone of the spatial branch, splicing + linear interpolation upsampling as the feature fusion part, and does not set the deep supervision mechanism. The results are shown in Table 2. The deep supervision mechanism combines auxiliary loss and main loss to avoid the gradient vanishing problem and improve segmentation efficiency. It can be seen from Table 3 that the baseline model using the deep supervision mechanism requires 100 iterations to reach 82.40% mIoU, while the baseline model using only the main loss requires 300 iterations to reach 82.38% mIoU. Since the random seed is not set, there is a slight difference in mIoU, which proves that the deep supervision mechanism improves the training efficiency.

[0099] Table 2

[0100] Table 3

[0101] ResT-Mamba has proven its effectiveness as a semantic branch in Table 2. It combines the residual Transformer and Mamba to extract global context semantic information while taking into account the extraction of local context semantic information, enhancing the extraction of multi-scale semantic information of the semantic branch, thereby improving the accuracy of semantic segmentation of remote sensing images. The ablation results in Table 3 show that compared with the baseline model without a semantic branch, the baseline model with ResT-Mamba has an increase of 2.30% in mIoU, 0.96% in mF1, 1.82% in mRecall, and 0.50% in OA. This fully proves the effectiveness of the semantic branch for the network.

[0102] MCS-FFM is the main module of the multi-scale feature aggregation part. Compared with ordinary network models, such as U-Net, the feature fusion part only needs to focus on the fusion of up to two features, while MDSNet needs to focus on the fusion of features of three different scales. For this reason, the present invention preliminarily fuses the features of three different scales through the large kernel grouping feature extraction module in MCS-FFM. Since the channel reweighted connection layer has not been added here, in order to fully verify the effectiveness of MCS-FFM, the output of the fourth layer of the spatial branch is input into MCS-FFM twice, thereby replacing the channel reweighted connection layer to complete the ablation experiment. The ablation results in Table 3 show that after using MCS-FFM to replace the feature fusion part of the baseline model, the network's mIoU increased by 1.93%, mF1 increased by 1.16%, mRecall increased by 0.76%, and OA increased by 0.88%. It fully proves the effectiveness of MCS-FFM for the network.

[0103] CRC is an important part of balancing the ratio of semantic features to spatial features in the input multi-scale convolutional feature fusion module, and also plays a role in strengthening edge segmentation. The ablation results in Table 3 show that after adding CRC, the network's mIoU increased by 0.46%, mF1 increased by 0.27%, mRecall increased by 0.20%, and OA increased by 0.26%. This fully proves the effectiveness of CRC in balancing the ratio of semantic spatial features and strengthening edge segmentation.

[0104] Table 3 shows the quantitative results of all the latest core methods on the Vaihingen dataset. The middle of Table 4 shows the intersection over union (IoU) and F1 score for each category, and the right side shows the overall evaluation indicators. MDSNet achieved the best results in average IoU, average F1 score and overall accuracy. There are mainly three structures. The first structure is the encoder-decoder structure, such as U-Net, CMLFormer, MAResU-Net and CMTFNet. As the originator of the encoder-decoder structure, U-Net has too many redundant features for the network due to the jump connection, which affects the segmentation results. The mIoU of MDSNet is 4.87% higher than that of U-Net, the mF1 is 2.89% higher, and the OA is 2.58% higher. Since CMLFormer, MAResU-Net and CMTFNet have made many improvements on U-Net, the advantages of these improvements cannot be fully reflected in 100 iterations, and their segmentation accuracy is also slightly lower than that of U-Net. The second structure is a dual-branch structure, such as EIGNet and CGGLNet. The spatial branch of EIGNet uses ordinary convolution, and the semantic branch uses ResNet18. The dual branches are not targeted enough, resulting in low indicators. The third structure is other mainstream networks, such as DeepLabV3+, PSPNet and A2FPN. All three cannot extract complex feature information in remote sensing images well, resulting in poor segmentation performance. In addition, the present invention adds other networks that use deep supervision mechanisms in the comparative experiments. For example, CGGLNet sets up a deep supervision mechanism. However, since CGGLNet uses an encoder-decoder structure, its various indicators are closer to MDSNet. However, since CGGLNet's spatial branch uses ordinary convolution and its semantic branch uses ResNet18, the dual branches are not targeted enough, resulting in its mIoU being 3.00% lower than MDSNet, mF1 being 1.77% lower, and OA being 1.62% lower; while MDSNet achieved 86.49% in mIoU, 92.68% in mF1, and 92.87% in OA, which quantitatively demonstrates the excellent segmentation performance of MDSNet. The recall rate is as follows: Fig. 9 As shown ( Fig. 9 The vertical axis in is the recall rate), MDSNet can achieve the best performance in all categories, indicating that MDSNet has the best segmentation performance on the Vaihingen dataset.

[0105] Table 4

[0106] Table 4 shows the quantitative results of all the latest core methods on the Potsdam dataset. From the data in the table, we can see that MDSNet is ahead of other models in all indicators, with mIoU reaching 82.43%, mF1 reaching 90.16%, and OA reaching 91.94%. It is the only model with mIoU exceeding 80% among many models, indicating that the model can accurately identify targets and will not miss too many real targets.

[0107] The present invention provides a high-resolution remote sensing image semantic segmentation method based on multi-scale deep supervision. There are many methods and ways to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.

Claims

1. A high-resolution remote sensing image semantic segmentation method based on multi-scale deep supervision, characterized in that: The following steps are involved: Step 1, obtain the ISPRS Potsdam dataset and the Vaihingen dataset, and divide the datasets into training set, validation set, and test set after uniform preprocessing; Step 2: Build a multi-scale deep supervision network MDSNet, including spatial branch, semantic branch network and multi-scale feature fusion part; The spatial branch, semantic branch network and multi-scale feature fusion part all have four-layer structures; Step 3: Input the training set into the multi-scale deep supervision network MDSNet for training, calculate the loss function and perform back propagation, update the network parameters, and obtain the trained parameter model; Step 4: Input the test set into the trained parameter model and output the segmentation map of the target image; Step 2 includes: The spatial branch includes a four-layer spatial redundant residual module, the outputs of the four-layer spatial redundant residual module are S1, S2, S3 and S4 from shallow to deep, and the output of the spatial redundant residual module is a four-dimensional feature of B×A×H×W, where B is the batch size, A is the number of channels, H is the length, and W is the width; The number of channels of the second-layer spatial de-redundant residual module is doubled compared with the number of channels of the first-layer spatial de-redundant residual module, and the length and width are halved; the number of channels of the third-layer spatial de-redundant residual module is doubled compared with the number of channels of the second-layer spatial de-redundant residual module, and the length and width are halved; the number of channels of the fourth-layer spatial de-redundant residual module is doubled compared with the number of channels of the third-layer spatial de-redundant residual module, and the length and width are halved; S2, S3 and S4 are connected through channel re-weighted connections, the degree of difference of each channel number is determined according to the learnable parameters of different feature maps, the importance of the channel is dynamically adjusted, and finally the output M4 of the channel re-weighted connection is used as the fourth layer input in the multi-scale feature fusion part; Step 2 also includes: the semantic branch network includes four layers of residual attention Mamba, the residual attention Mamba is used to extract fine-grained features, the residual attention Mamba includes residual Transformer and Mamba, the outputs of the four layers of residual attention Mamba are C1, C2, C3 and C4 from shallow to deep, and the number of channels and length and width changes are consistent with the spatial branch; A double-layer convolution block is used before the spatial branch and the semantic branch to roughly extract features. Each convolution block contains a regular convolution, batch normalization, and ReLU activation function. The multi-scale feature fusion part includes a large kernel group feature extraction module and a multi-scale convolution attention module, wherein the large kernel group feature extraction module is used to selectively fuse and extract three types of features, namely, spatial branch output S, upsampling and channel reweighted output M, and semantic branch output C; the large kernel group feature extraction module gradually combines feature mapping with attention coefficients, and integrates features from shallow and deep layers by introducing group convolution; batch normalization is performed on the features after group convolution, and feature fusion is performed, and then ordinary convolution is performed on the feature fusion results, and they are spliced ​​with the initial spatial branch output S and semantic branch output C, and finally sent to the multi-scale convolution attention module; The features output by the spatial branch and semantic branch networks are connected to the three-layer efficient upsampling convolution block using a four-layer multi-scale convolution feature fusion module. The outputs of the three-layer efficient upsampling convolution blocks are M1, M2 and M3 from shallow to deep. The output of each layer of multi-scale convolution feature fusion module is input into the segmentation head for segmentation. The segmentation head includes a common convolution block with a 3×3 convolution kernel, a common convolution with a 1×1 convolution kernel and an upsampling. The segmentation head has four layers, corresponding to the four-layer multi-scale convolution feature fusion module. Step 3 includes: sending the remote sensing image with feature dimension of b×3×256×256 into the multi-scale deep supervision network MDSNet, where b is batch_size, batch_size represents the total number of images trained in a batch; 3 represents the number of channels; sending the remote sensing image into the double-layer convolution block, the output feature dimension becomes b×64×256×256, and the rough extraction of the original image features is completed; The output of the double-layer convolution block is sent to the spatial branch and semantic branch network; for the spatial branch, the first-layer spatial redundant de-redundancy residual module includes a left half and a right half residual part, wherein the left half includes horizontal and vertical convolutions, and the right half residual part is a residual structure; firstly, the features of b×64×256×256 are horizontally and vertically convolved, and the convolution kernels are 3×1, 3×3 and 1×3 respectively, and the convolution results of three different convolution kernels of 3×1, 3×3 and 1×3 are concatenated to generate a four-dimensional feature of size b×192×128×128; the four-dimensional feature containing the directional feature is sent to the right half residual part of the first-layer spatial redundant de-redundancy residual module, and the four-dimensional feature is sent to the convolution block removing ReLU, the convolution block removing The coordinate convolution block of ReLU is removed, the convolution kernels of the convolution block without ReLU and the coordinate convolution block without ReLU are 1×1 and 3×3 respectively, and the output results are both four-dimensional features of b×64×128×128. The result of the convolution block without ReLU is input into the wavelet transform convolution, the convolution kernel of the wavelet transform convolution is 5×5, and the output result is a four-dimensional feature of b×64×128×128. Finally, the convolution block without ReLU, the coordinate convolution block without ReLU and the output four-dimensional features of the wavelet transform convolution are added, and a ReLU activation function is passed through, so as to obtain a four-dimensional feature S1 of b×64×128×128. The coordinate convolution block without ReLU includes coordinate convolution and batch normalization. The processing process of the second, third and fourth spatial redundant residual modules is the same as that of the first spatial redundant residual module. The output results are b×128×64×64 four-dimensional feature S2, b×256×32×32 four-dimensional feature S3 and b×512×16×16 four-dimensional feature S4. The specific formula is: F1=f concat (CBR 1×3 (CBR 1×1 ),CBR 3×1 (CBR 1×1 ),CBR 3×3 (CBR 1×1 )) F SDCM =ReLU(f concat (CB 1×1 (F1), WTcon2d 5×5 (CB 1×1 (F1)),CoCB 3×3 (CoCBR 3×3 (F1)))) Among them, F1 is the result after the feature expression is enhanced, F SDCM is the output result of the spatial redundancy removal residual module; f concat (·) is the splicing operation, CBR n×m is a convolution block, which contains a conventional convolution with a convolution kernel of size n×m, batch normalization, and ReLU activation function; ReLU(·) is the activation function, CB n×m To remove the convolution block of the ReLU activation function, WTcon2d m×m Wavelet transform convolution with a convolution kernel of size m×m; CoCBR m×m CoCB is a coordinate convolution block, which contains coordinate convolution of m×m convolution kernel, batch normalization and ReLU activation function; m×m To remove the coordinate convolution of the ReLU activation function, m and n are positive integers; The multi-scale deep supervision network MDSNet uses the outputs S2, S3, and S4 of the last three feature extraction layers for feature fusion, upsamples the input feature M4 according to the standard of S4, and splices the upsampled feature map according to the first dimension to obtain the spliced ​​feature map. By adding the number of channels of the input features, a learnable weight parameter with the same number of channels is set, and the learnable weight parameter is normalized, the spliced ​​feature map is multiplied by the normalized weight parameter to realize the weighted operation of the input features, and the weighted features are spliced. The weighted features are then concatenated, and a four-dimensional feature M4 of b×512×16×16 is obtained through a convolution operation with a step size of 4 and a convolution kernel of 4×4; For the semantic branch network, Residual Attention Mamba includes Residual Transformer and Mamba. Residual Transformer directly models the global relationship of the input image through the self-attention mechanism, thereby capturing long-distance dependencies. The residual Transformer includes an efficient attention block ETB, which uses an efficient multi-head self-attention mechanism to obtain the attention score. The result of the double-layer convolution block is input into the residual Transformer, and the output result is a four-dimensional feature of b×64×128×128. The formula is: FFN(x)=σ(xW1+b1)W2+b2 Where EMSA(Q,K,V) is the output of efficient multi-head self-attention, Q represents the query vector, K represents the key vector, V represents the numeric vector, EMSA(Linear(x)) represents the efficient multi-head self-attention mechanism after the fully connected layer, IN(·) represents instance normalization, Conv(·) represents the convolution with a kernel of 1×1, and T represents transpose; FFN(x) is a fully connected feedforward network layer, x represents the input feature; W1, W2 are the weight matrices of the linear transformation; b1, b2 are bias terms, σ is a nonlinear activation function, is the residual connection structure of efficient multi-head self-attention; Linear(·) is the fully connected layer; y is the final output of the residual Transformer; After the residual transformer extracts the global semantic features, y is input into the visual state space VSS module of the first layer of residual attention Mamba to extract local semantic features of y. The VSS module includes a state space module SSM. The output result of the residual transformer is input into the residual attention Mamba to obtain a four-dimensional feature C1 of b×64×128×128. The processing process of the second, third and fourth layers of residual attention Mamba is the same as that of the first layer of residual attention Mamba. The output results are respectively a four-dimensional feature C2 of b×128×64×64, a four-dimensional feature C3 of b×256×32×32 and a four-dimensional feature C4 of b×512×16×16. The specific formula is: X1=Silu(Linear(LN(y))) X2=LN(SSM(Silu(DWConv(Linear(LN(x)))))) Where LN(·) represents layer normalization, Silu(·) represents Silu function activation operation, X1 and X2 are the output of Silu function and the output of LN layer normalization respectively; DWConv(·) is depthwise separable convolution, is the product operation, X VSS The final output result of residual attention Mamba; Finally, S1, S2, S3, S4 and C1, C2, C3, C4, M4 are sent to the multi-scale feature fusion part together. The input of the fourth layer of multi-scale feature fusion is S4, C4 and M4. The three four-dimensional features S4, C4 and M4 are sent to the large kernel grouping feature extraction module. Each input undergoes a group convolution with a convolution kernel of 3×3 and batch normalization, and outputs b×256×16×16 four-dimensional features. Then, the results of group convolution and batch normalization of S4, C4 and M4 are concatenated, and the ReLU activation function is passed, the convolution block without ReLU and the Sigmoid operation are removed. The output result is b×1×16×16 features. Finally, the b×1×16×16 features are Hadamard-converted with S1 and C1. Product, output b×512×16×16 four-dimensional features, and finally input the b×512×16×16 four-dimensional features into the multi-scale convolutional attention module to obtain b×512×16×16 four-dimensional features M3 as the input of the third layer of multi-scale feature fusion. The multi-scale feature fusion process of the third, second and first layers is the same as that of the fourth layer. The output results of the multi-scale feature fusion of the third, second and first layers are b×256×32×32 four-dimensional features, b×128×64×64 four-dimensional features, and b×64×128×128 four-dimensional features respectively; the multi-scale feature fusion results are sent to the segmentation head respectively, and the output results are all b×6×256×256 four-dimensional features, and 6 represents the six categories of segmentation.

2. The method according to claim 1, characterized in that Step 1 includes data augmentation on the Vaihingen dataset by random sampling and cropping.

3. The method according to claim 2, characterized in that Step 3 also includes: using the loss function to supervise the four-dimensional features of the final segmentation result b×6×256×256 and the outputs of the second, third, and fourth layer segmentation heads as auxiliary losses, and using the auxiliary losses to correct the main losses. The formula is: l mian =λ aux =λ DiceLoss +λ CELoss λ=λ mian +δ(λ aux1 +λ aux2 +λ aux3 ) Among them, λ DiceLoss is the Dice Loss loss; X represents the pixel label of the real segmented image, Y represents the pixel category of the segmented image predicted by the multi-scale deep supervision network MDSNet, and C is the number of categories; λ CELoss is the cross entropy loss; N is the number of pixels in the image, y i is the true label of the ith pixel, y i,c is a binary indicator function indicating whether the i-th pixel belongs to category c. If so, y i,c is 1, otherwise y i,c is 0; p i,c represents the predicted probability of the multi-scale deep supervision network MDSNet for the i-th pixel belonging to category c; λ mian is the main loss, λ aux1 is the auxiliary loss of the second-layer segmentation head output; aux2 is the auxiliary loss of the third-layer segmentation head output; aux3 is the auxiliary loss output by the fourth-layer segmentation head; δ is the weighting coefficient, and λ is the final total loss value.

4. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 3.

5. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 3 are executed.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on double-branch feature fusion

    CN115797931A