Hybrid structure remote sensing image segmentation method based on state space model
By adopting a hybrid structure based on state space model in remote sensing image segmentation technology, combining the dual encoder structure of CNN and state space model, and introducing a window-based attention mechanism decoder, the bottleneck problem of the prior art when dealing with complex backgrounds and high-resolution images is solved, and more efficient semantic segmentation performance is achieved.
Patent Information
- Application Number
- CN202510216997.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Existing remote sensing image semantic segmentation technology has bottlenecks when dealing with complex backgrounds and high-resolution images, making it difficult to capture remote dependencies and global context information.
A hybrid structure remote sensing image segmentation method based on state space model is adopted, combining the dual encoder structure of CNN and state space model, and a window-based attention mechanism decoder is introduced to extract multi-scale features and capture global-scale feature dependencies.
It significantly improves the semantic segmentation performance of remote sensing images and can be better than other existing semantic segmentation models based on CNN, Transformer and Mamba under reasonable computing overhead.
Smart Images

Figure CN120147632A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing technology, and in particular to a method for segmenting remote sensing images with a hybrid structure based on a state space model. Background Art
[0002] In recent years, with the development of remote sensing technology and sensors, the acquisition of a large number of high-resolution remote sensing images has brought great challenges to the intelligent interpretation of remote sensing images. These images not only contain rich spatial details but also have great semantic potential. The semantic segmentation technology of remote sensing images has become an important tool in the fields of land use classification, urban planning, environmental monitoring, and disaster assessment by classifying the ground object types in remote sensing data pixel by pixel. However, due to the characteristics of multi-scale targets, complex scenes, and high resolution of remote sensing images, the semantic segmentation task faces many challenges.
[0003] Traditional convolutional neural networks (CNNs) have made significant progress in the semantic segmentation task of remote sensing images. For example, models such as U-Net and DeepLab have achieved good segmentation performance through hierarchical feature extraction and skip connection mechanisms. However, due to the receptive field limitation of CNNs, it is difficult for them to capture long-range dependencies and global context information, and there are bottlenecks in processing complex backgrounds or high-resolution images.
[0004] In recent years, Transformer based on the self-attention mechanism has shown excellent performance in semantic modeling. Models such as Swin Transformer and UNetFormer have achieved excellent performance in the semantic segmentation task. However, the high computational complexity and memory requirements of Transformer limit its application in high-resolution remote sensing images, especially in the deployment of resource-constrained devices. In addition, the training of Transformer depends on large-scale data, and its performance is limited on small datasets.
[0005] The state space model (SSM) provides a new idea for the semantic segmentation task through its linear computational complexity and long-range dependence modeling ability. Models such as Mamba and DenseMamba have shown potential in visual tasks. However, the adaptability of SSM in complex remote sensing scenes and its ability to express multi-scale features still need further research and optimization. Summary of the Invention
[0006] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a method for segmenting remote sensing images with a hybrid structure based on a state space model, which adopts a dual-encoder structure based on CNN and a state space model, and introduces a window-based attention mechanism decoder to significantly improve the semantic segmentation performance of remote sensing images.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A hybrid - structure remote - sensing image segmentation method based on a state - space model, comprising:
[0009] Obtain a remote - sensing image to be processed, input the remote - sensing image to be processed into a hybrid - structure remote - sensing image segmentation model, and obtain a segmentation result; the hybrid - structure remote - sensing image segmentation model is obtained by training with a training set, and the training set includes: remote - sensing images;
[0010] The hybrid - structure remote - sensing image segmentation model uses a residual network sub - model based on a convolutional neural network and an improved Mamba sub - model to form a dual - encoder structure for extracting multi - scale features, and a Transformer sub - model based on a hybrid attention convolutional module to form a decoder, which is used to capture feature dependencies in the global range through the multi - scale features, while extracting local features and capturing spatial information, and generating the segmentation result; wherein, the improved Mamba sub - model is the original Mamba sub - model introduced with a two - dimensional selective scanning module and a CSAM attention fusion module.
[0011] Optionally, the residual network sub - model includes:
[0012] A number of residual blocks for performing convolutional operations to generate multi - scale features;
[0013] A number of RCM modules for fusing the multi - scale features with different - scale features generated by the improved Mamba sub - model and inputting the fusion result into the decoder.
[0014] Optionally, generating different - scale features using the improved Mamba sub - model includes:
[0015] After performing layer normalization on the input features, divide them into two paths for subsequent processing:
[0016] The first path: sequentially pass through a linear layer, a depth convolution, a two - dimensional selective scanning module, and layer normalization for feature extraction;
[0017] The second path: pass through a CSAM attention fusion module and fuse the output result of the first path based on element - wise multiplication to generate different - scale features.
[0018] Optionally, the expression of the improved Mamba sub - model is:
[0019] Q x = f LN (x)
[0020] W x = f CSAM (Q x )
[0021] S x = f LN (f SS2D (D conv (f Linear (Q x )))))⊙W x
[0022]
[0023] where x represents the input feature, f LN (·) is the layer normalization operation, f CSAM (·) is the fusion attention operation, f Linear (·) is the linear transformation operation, D conv (·) is the depth convolution operation, ⊙ is the element-wise multiplication operation, f SS2D (·) is the 2D selective scanning operation, F x VSS represents the result obtained by processing the input feature x through the VSS block, Q x is the normalization result, W x is the attention operation result, S x is the result of the 2D selective scanning operation and the residual connection.
[0024] Optionally, the 2D selective scanning module includes:
[0025] A 2D selective scanning sub-module for forward and backward scanning of the input feature map in a cross-scanning manner. During the scanning process, the image is expanded in multiple directions to generate different sequences, and the different sequences are processed using the S6 operation.
[0026] Optionally, processing the different sequences using the S6 operation includes:
[0027] Obtaining the first change result, the second change result, and the third change result by respectively subjecting the input feature map to three linear changes, performing an element-wise exponential operation on the first change result to introduce non-linear dynamic characteristics, calculating the second change result using a matrix transformation, which is a composite operation of matrix exponentiation and inverse operation, further updating the intermediate state by combining the initial state and the input feature tensor to provide context information for the sequence, calculating each time step using a linear combination method, performing a state space vector operation on the third change result to obtain an output result, and integrating all time steps to generate a complete output sequence.
[0028] Optionally, the expression for processing the different sequences using the S6 operation is:
[0029] X O = Sse (x, O)
[0030]
[0031] Among them, O represents different directions, X represents the input feature, and S se (·) and S sm (·) respectively represent the scan expansion and scan merge operations, S6(·) represents the S6 operation, and X O represents respectively represent the results of four different scan directions of S6.
[0032] Optionally, the CSAM attention fusion module includes:
[0033] A channel attention module, which is used to extract channel features through global average pooling and max pooling operations, and further extract local features using a first depthwise separable convolution, and enhance important channels through a first sigmoid activation function;
[0034] A spatial attention module, which is used to combine the channel features, extract features through a first convolutional block, and obtain the key spatial positions in the feature map through a second sigmoid activation function;
[0035] A multi-scale convolution module, which is used to adopt second depthwise separable convolution kernels of different sizes, combine batch normalization and ReLU activation functions to extract different-scale features, and integrate the different-scale features using channel rearrangement and a second convolutional block;
[0036] Among them, the channel attention module, the spatial attention module, and the multi-scale convolution module are connected in sequence.
[0037] Optionally, generating the segmentation result using the decoder includes:
[0038] After normalizing the multi-scale features to standardize the feature distribution through batch normalization, divide them into two parallel processing paths for subsequent processing:
[0039] The first processing path: capture the feature dependencies in the global range through a window multi-head self-attention module, and at the same time use the window mechanism to reduce the computational complexity;
[0040] The second processing path: perform local feature extraction and capture of spatial information through a convolutional group module;
[0041] Merge the output results of the two processing paths, and after batch normalization again, input the merged features into a multi-layer perceptron layer for non-linear transformation, and further superimpose the multi-scale features and the non-linearly transformed features through residual connection to generate the segmentation result.
[0042] The beneficial effects of the present invention are as follows:
[0043] The present invention introduces a CVSS module based on the attention mechanism and visual semantic fusion (VSS) as an auxiliary encoder, which is combined with the CNN network to form a dual-encoder structure, thereby providing additional global perception information while minimizing the linear computational complexity. In addition, this paper also designs a brand-new regional convolution module (RCM) to fuse different features of the main encoder and the auxiliary encoder, thereby enhancing the feature representation ability.
[0044] To avoid increasing the model complexity, a structure based on the window self-attention mechanism is adopted in the decoder part. Through a series of benchmark tests and ablation experiments on the ISPRS Vaihingen and Potsdam datasets, UAVid dataset, and LoveDA dataset, the effectiveness and efficiency of CMT-UNet in remote sensing image segmentation tasks are verified. The experimental results show that CMT-UNet significantly outperforms other existing semantic segmentation models based on CNN, Transformer, and Mamba under reasonable computational overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1 It is the overall architecture diagram of a hybrid structure remote sensing image segmentation method based on the state space model in this embodiment;
[0047] Figure 2 It is the structure schematic diagram of the CVSS module in the auxiliary encoder of this embodiment;
[0048] Figure 3 It is the schematic diagram of the working principle of the two-dimensional selective scanning module (SS2D) in this embodiment;
[0049] Figure 4 It is the internal structure diagram of the channel and spatial attention fusion module (CSAM) in this embodiment;
[0050] Figure 5 It is the structure schematic diagram of the feature fusion module (RCM) in this embodiment;
[0051] Figure 6 It is the structure schematic diagram of the hybrid attention convolution module (HAC) decoder in this embodiment;
[0052] Figure 7 Visualization comparison diagrams of different methods in this embodiment for the Vaihingen dataset;
[0053] Figure 8 Visualization comparison diagrams of different methods in this embodiment for the Potsdam dataset. Detailed implementation manners
[0054] Next, the technical solutions in this embodiment will be clearly and completely described in conjunction with the accompanying drawings in this embodiment. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0055] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific implementation manners.
[0056] This embodiment discloses a hybrid structure remote sensing image segmentation method based on a state space model, including: obtaining a remote sensing image to be processed, inputting the remote sensing image to be processed into a hybrid structure remote sensing image segmentation model to obtain a segmentation result; the hybrid structure remote sensing image segmentation model is obtained by training using a training set, and the training set includes: remote sensing images; the hybrid structure remote sensing image segmentation model uses a residual network sub-model based on a convolutional neural network and an improved Mamba sub-model to form a dual encoder structure for extracting multi-scale features, and a Transformer sub-model based on a hybrid attention convolutional module to form a decoder for capturing global range feature dependencies through multi-scale features while extracting local features and capturing spatial information to generate a segmentation result; wherein, the improved Mamba sub-model is the original Mamba sub-model introduced with a two-dimensional selective scanning module and a CSAM attention fusion module.
[0057] Specifically:
[0058] This embodiment discloses a hybrid structure remote sensing image segmentation method based on a state space model, including the following steps:
[0059] S1: Obtain a remote sensing image and divide it into a training set, a validation set, and a test set;
[0060] S2: Construct a hybrid structure remote sensing image segmentation model based on a state space model;
[0061] S3: Use the training set and the validation set to train the hybrid structure remote sensing image segmentation model;
[0062] S4: Input the remote sensing image to be processed into the trained hybrid structure remote sensing image segmentation model to obtain the segmentation result.
[0063] Among them, in the hybrid structure of S2, a dual encoder based on the residual network of CNN and the improved Mamba model is designed, and the decoder is constructed by the Transformer model based on the hybrid attention convolution module. The network consists of three branches, and each branch contains four corresponding blocks for fusing the features extracted by the Mamba and CNN networks. After feature extraction and fusion are performed at four scales, multi-scale features are obtained, and these features are then input into the decoder through the region convolution module to generate the final prediction map.
[0064] Dual encoder design: The hybrid structure adopts a main encoder based on CNN and an auxiliary encoder based on the state space model, which are responsible for local and global feature extraction respectively.
[0065] Main encoder: ResNet50 is used as the basic network, which includes 4 residual blocks and 4 region convolution modules (RCM). The residual blocks are responsible for extracting multi-scale local features, and the RCM makes up for the deficiency of the main encoder in capturing global information by integrating the global features of the auxiliary branch.
[0066] Auxiliary encoder: The CVSS module is introduced, and the VSS block is adopted in the backbone network as the deep feature processing module. The VSS block enhances the semantic information extraction ability by maintaining the linear computational complexity while capturing the remote context information. After the input features are processed by layer normalization (LayerNorm), they are divided into two paths: The first path sequentially passes through the linear layer, depth convolution, two-dimensional selective scanning (SS2D) operation and layer normalization to extract fine-grained features. The second path passes through the channel and spatial attention module (CSAM), and is fused with the result of the first path by element-wise multiplication to output enhanced features.
[0067] The RCM merges the different scale features derived from the auxiliary branch into the main branch to make up for the limitations of the main branch in extracting global information. The outputs of the main encoder and the auxiliary encoder are first processed by layer normalization (LayerNorm) respectively. The window-based multi-head self-attention mechanism is used to model the long-range dependencies. The features of the auxiliary branch are extracted by the CVSS module with long dependencies, and the convolutional network is used to learn the local details.
[0068] Furthermore, the residual network sub-model includes: several residual blocks for performing convolutional operations to generate multi-scale features; several RCM modules for fusing the multi-scale features with the different scale features generated by the improved Mamba sub-model, and inputting the fusion result into the decoder.
[0069] Furthermore, generating different-scale features using the improved Mamba sub-model includes: after performing layer normalization on the input features, they are divided into two paths for subsequent processing: The first path: sequentially pass through a linear layer, a depth convolution, a two-dimensional selective scanning module, and layer normalization to extract features;
[0070] The second path: pass through the CSAM attention fusion module and fuse the output results of the first path based on element-wise multiplication to generate different-scale features.
[0071] Specifically:
[0072] Data is first fed into the system as input and undergoes initial batch normalization (Batch Norm) to standardize the feature distribution and help the model converge faster. The data is divided into two processing paths:
[0073] The first path passes through a window-based multi-head self-attention module to capture feature relationships and context information on a global scale.
[0074] The second path passes through a convolutional group module, which focuses on local feature extraction and spatial information capture.
[0075] Feature fusion: The results of the two paths are combined to complete feature integration.
[0076] Batch normalization: The combined features are batch-normalized again to further improve training stability.
[0077] Multi-layer perceptron (MLP) processing: Subsequently, the features pass through a multi-layer perceptron module for non-linear transformation to increase the model's expressive power.
[0078] Residual connection: Finally, the residual connection superimposes the initial input and the processed features as the final output. This mechanism helps alleviate the gradient vanishing problem and improve feature transmission efficiency.
[0079] Furthermore, the expression of the improved Mamba sub-model is:
[0080] Q x =f LN (x)
[0081] W x =f CSAM (Q x )
[0082] S x =f LN (f SS2D (Dconv (f Linear (Q x )))))⊙W x
[0083]
[0084] Among them, x represents the input feature, f LN (·) is the layer normalization operation, f CSAM (·) is the fusion attention operation, f Linear (·) is the linear transformation operation, D conv (·) is the depth convolution operation, ⊙ is the element-wise multiplication operation, f SS2D (·) is the 2D selective scanning operation, F x VSS represents the result obtained by processing the input feature x through the VSS block, Q x is the normalization result, W x is the attention operation result, S x is the result of the 2D selective scanning operation and the residual connection.
[0085] Furthermore, the 2D selective scanning module includes:
[0086] The 2D selective scanning sub-module is used to perform forward and backward scans on the input feature map in a cross-scanning manner. During the scanning process, the image is expanded in multiple directions to generate different sequences, and the S6 operation is used to process the different sequences.
[0087] Furthermore, using the S6 operation to process the different sequences includes:
[0088] The input feature map is respectively obtained with the first change result, the second change result, and the third change result through three linear changes. The element-wise exponential operation is performed on the first change result to introduce non-linear dynamic characteristics, and the second change result is calculated using matrix transformation. The matrix transformation is a composite operation of matrix exponentiation and inverse operation. Further, the intermediate state is updated by combining the initial state and the input feature tensor to provide context information for the sequence. Each time step is calculated using a linear combination method. The state space vector operation is performed on the third change result to obtain the output result, and all time steps are integrated to generate a complete output sequence, activating all pixels and significantly emphasizing the cross-type activation. The cross-scanning module scanning mechanism makes the central pixel most affected by the cross-direction pixels. For each pixel, its long-range content takes precedence over local information.
[0089] Specifically:
[0090] Two-dimensional Selective Scanning (SS2D): Cross-scanning enables forward and backward scanning of images. This process involves expanding the image in four directions to generate four different sequences, which are then processed by S6. The operation of S6 enables each element in the one-dimensional array to interact with all previously scanned samples through compressed hidden states, thereby capturing different features.
[0091] The implementation process of the S6 operation in SS2D mainly includes the following steps: First, the input feature tensor x, with shape [B, L, D], is passed through three linear transformations to obtain Δ, B, and C respectively. Subsequently, element-wise exponential operations are performed on Δ to introduce non-linear dynamic characteristics. Then, matrix B is calculated using matrix transformation, which involves a composite operation of matrix exponentiation and inverse operation, and this process can efficiently capture the long-range dependencies of the sequence. Next, the intermediate state h is updated by combining the initial state h 0 and the input x to provide context information for sequence generation. Finally, each time step y t is calculated using a linear combination, and all time steps are integrated to generate the complete output sequence y.
[0092] Furthermore, the expression for processing different sequences using the S6 operation is:
[0093] X O = S se (x, O)
[0094]
[0095] where O represents different directions, X represents the input feature, S se (·) and S sm (·) respectively represent the scan expansion and scan merging operations, S6(·) represents the S6 operation, and X O represents, respectively represent the results of the four different scan directions of S6.
[0096] Furthermore, the CSAM attention fusion module includes:
[0097] The channel attention module is used to extract channel features through global average pooling and max pooling operations, and further extract local features using the first depthwise separable convolution, and enhance important channels through the first sigmoid activation function;
[0098] The spatial attention module is used to combine channel features, extract features through the first convolutional block, and obtain the key spatial positions in the feature map through the second sigmoid activation function;
[0099] The multi-scale convolution module is used to extract features of different scales by using second-depth separable convolution kernels of different sizes, combined with batch normalization and ReLU activation functions, and integrate features of different scales through channel rearrangement and the second convolution block;
[0100] Among them, the channel attention module, the spatial attention module, and the multi-scale convolution module are connected in sequence.
[0101] Specifically:
[0102] The multi-scale attention mechanism (CSAM): combines the channel attention module and the spatial attention module to improve the segmentation ability for small targets and complex backgrounds. The CSAM module combines channel attention, spatial attention, and the multi-scale convolution module to improve the segmentation ability for complex backgrounds and small targets.
[0103] The channel attention module: extracts features through global average pooling, max pooling, and depthwise separable convolution, and combines the sigmoid activation function to enhance the expression of important channels.
[0104] The spatial attention module: combines the average pooling and max pooling results of the channels, and uses a 7×7 convolution to highlight the key positions in the feature map.
[0105] The multi-scale convolution module: further enriches the multi-scale feature expression ability through convolution kernels of multiple sizes combined with batch normalization and ReLU activation functions.
[0106] Furthermore, generating the segmentation result using the decoder includes:
[0107] After normalizing the multi-scale features to standardize the feature distribution through batch normalization, they are divided into two parallel processing paths for subsequent processing:
[0108] The first processing path: captures the feature dependencies in the global range through the window multi-head self-attention module, and at the same time uses the window mechanism to reduce the computational complexity;
[0109] The second processing path: performs local feature extraction and spatial information capture through the convolution group module;
[0110] The output results of the two processing paths are merged, and after batch normalization again, the merged features are input into the multi-layer perceptron layer for non-linear transformation. Further, the multi-scale features are superimposed with the non-linearly transformed features through residual connection to generate the segmentation result.
[0111] Specifically:
[0112] Construct the HAC decoder. This decoder combines the window-based multi-head self-attention module and the group convolution module to construct a lightweight Transformer decoder.
[0113] First, the feature distribution is standardized through Batch Normalization (Batch Norm), and then it is divided into two parallel processing paths. One path passes through the Windows-based Multihead Self-Attention module to capture feature dependencies in the global scope, while using the window mechanism to reduce computational complexity; the other path passes through the Convolutional Group module, which focuses on local feature extraction and capture of spatial information. The outputs of the two paths are then merged and passed through Batch Normalization again to enhance the effect of feature integration. Next, the fused features are fed into a Multi-Layer Perceptron (MLP) for non-linear transformation to further improve the expressive power of the model. Finally, the initial input and the processed features are superimposed through residual connections to alleviate the problem of gradient disappearance and improve the efficiency of feature transmission, and finally the output data is generated. This architecture combines the advantages of the self-attention mechanism and convolutional operations, can effectively capture multi-scale features, and at the same time maintain the stability of training and model performance.
[0114] As Figure 1 shown, this embodiment provides a hybrid structure remote sensing image segmentation method based on a state space model, including the following steps:
[0115] S1: Obtain a remote sensing image and divide it into a training set, a validation set, and a test set;
[0116] S2: Construct a hybrid structure remote sensing image segmentation model based on a state space model;
[0117] S3: Use the training set and the validation set to train the hybrid structure remote sensing image segmentation model;
[0118] S4: Input the remote sensing image to be processed into the trained segmentation model to obtain the segmentation result.
[0119] In this method, the segmentation model is designed as a dual encoder composed of a CNN-based residual network and an improved Mamba model, and the decoder is constructed by a Transformer model based on a Hybrid Attention Convolution module (HAC). The model contains three branches, and each branch includes four blocks for fusing the features extracted by the CNN and Mamba models. After feature extraction and fusion are performed at four scales, multi-scale features are generated, and these features are input into the decoder through a Region Convolution Module (RCM) to generate the final prediction map.
[0120] The present invention uses ResNet50 as the main encoder to learn local features. As Figure 1 shown, it contains 4 residual blocks and 4 Region Convolution Modules (RCM). These 4 residual blocks perform convolutional operations to generate multi-scale features, denoted as Compared with the auxiliary encoder, the main encoder can effectively extract features from remote sensing images through existing pre-trained models. Therefore, the present invention merges different scale features derived from the auxiliary branch into the main branch through RCM to make up for the limitation of the main branch in extracting global information. As Figure 5 shown, the outputs of the main encoder and the auxiliary encoder and are first processed by layer normalization (Layer Norm) respectively. This helps to accelerate network convergence and improve the stability of the training process. The main encoder extracts features through a CNN network. Therefore, on the basis of ensuring linear complexity, a window-based multi-head self-attention mechanism is adopted to model long-range dependencies. The features of the auxiliary branch are extracted by the CVSS module with long dependencies, so a convolutional network is used to learn local details.
[0121] The auxiliary encoder constructed by the present invention is as Figure 2 shown. The CVSS module is introduced into the auxiliary encoder, and the VSS block is adopted in the module backbone network as a deep feature processing module.
[0122] VSS block: It can capture long-range context information while maintaining linear complexity and enhance the refined extraction of semantic information. In the training stage, the VSS block conducts feature interaction with the semantic transformer to achieve joint training.
[0123] Feature processing flow: After the input features are processed by layer normalization (Layer Norm), they are divided into two paths:
[0124] The first path sequentially passes through a linear layer, a depth convolution, a two-dimensional selective scan (SS2D) operation, and layer normalization to extract fine-grained features;
[0125] The second path passes through a channel and spatial attention module (CSAM) and is fused with the first path by element-wise multiplication to obtain an enhanced feature output.
[0126] The formulation of the auxiliary encoder is as follows:
[0127] Q x = f LN (x) (1)
[0128] W x = f CSAM (Q x ) (2)
[0129] S x = f LN (f SS2D (D conv (f Linear (Qx )))))⊙W x (3)
[0130]
[0131] where \(x\) represents the input feature, \(f\) LN (·) is the layer normalization operation, \(f\) CSAM (·) is the fusion attention operation, \(f\) Linear (·) is the linear transformation operation, \(D\) conv (·) is the depth convolution operation, ⊙ is the element-wise multiplication operation, \(f\) SS2D (·) is the 2D selective scanning operation, \(F\) x VSS represents the result obtained by processing the input feature \(x\) through the VSS block.
[0132] The two-dimensional selective scanning (SS2D) structure. The SS2D module realizes the forward and backward scanning of the image through cross-scanning, effectively expanding the receptive field. The specific implementation is as Figure 3 shown. The SS2D module generates four sequences by expanding the image in four directions and uses the S6 operation to capture the feature information in different directions. Finally, the output image is restored to the original size through scan merging. The formula is as follows:
[0133] X O = S se (x, O) (5)
[0134]
[0135] where \(O\) represents different directions, \(O\in[1, 2, 3, 4]\). \(X\) represents the input feature. \(S\) se (·) and \(S\) sm (·) represent the scan expansion and scan merging operations respectively. \(S6(·)\) represents the S6 operation.
[0136] The pseudo-code of the S6 operation is shown in Table 1.
[0137] Table 1
[0138]
[0139] As Figure 4As shown, the CSAM attention fusion module includes a channel attention module, a spatial attention module, and a multi-scale convolution module. The channel attention module enhances the expression of important channels in the feature map through global average pooling, max pooling, depthwise separable convolution, and the sigmoid activation function. The spatial attention module combines the average pooling and max pooling results of the channels, passes through a 7x7 convolution and the sigmoid activation, and highlights important spatial positions in the feature map. The multi-scale convolution module uses depthwise separable convolution kernels of different sizes, combines batch normalization and the ReLU activation function, and further improves the multi-scale feature expression ability of the model through channel rearrangement and 1x1 convolution.
[0140] The HAC decoder constructed by the present invention is as Figure 6 shown. This decoder combines a window-based multi-head self-attention module and a grouped convolution module to construct a lightweight Transformer decoder.
[0141] First, the feature distribution is standardized through batch normalization (Batch Norm), and then it is divided into two parallel processing paths. One path passes through the window-based multi-head self-attention module to capture global feature dependencies, while using the window mechanism to reduce computational complexity; the other path passes through the convolutional group module to focus on local feature extraction and capture of spatial information. The outputs of the two paths are then merged and passed through batch normalization again to enhance the feature integration effect. Next, the fused features are fed into a multi-layer perceptron (MLP) for non-linear transformation to further improve the expression ability of the model. Finally, the initial input and the processed features are superimposed through residual connection to alleviate the gradient vanishing problem and improve the feature transmission efficiency, and finally the output data is generated. This architecture combines the advantages of the self-attention mechanism and convolutional operations, can effectively capture multi-scale features, and at the same time maintain the stability of training and model performance.
[0142] In the training stage of the present invention, the loss function adopts the combination of dice loss and cross-entropy loss , and its formula is as follows:
[0143]
[0144] where N and K represent the number of samples and the number of categories respectively. y (n) and represent the one-hot encoding of the true semantic label and the corresponding softmax output of the network respectively, and n ∈ [1, …, N]. Indicates the confidence that sample n belongs to class k.
[0145] In some embodiments, the embodiments of the present application provide an experimental comparison between the "remote sensing image segmentation method" of the present application and existing remote sensing image segmentation methods.
[0146] Experiment 1: Performance comparison on the UAVid test set
[0147] On the UAVid test set, the present invention was compared with a variety of efficient segmentation networks and evaluated based on the following performance metrics:
[0148] mIoU (mean intersection over union): Measures the segmentation accuracy of the prediction results;
[0149] GPU memory occupancy (MB): Reflects the resource requirements during network operation;
[0150] Complexity (FLOPs): Evaluates the computational complexity of the network;
[0151] Number of parameters (M): Measures the storage requirements of the model;
[0152] Speed (FPS): Measures the inference efficiency in frames per second.
[0153] Experiment 2: Study on the network stability of remote sensing images of different sizes
[0154] To evaluate the stability of the present invention under different input sizes, remote sensing images with square sizes (512×512, 1024×1024, 2048×2048) and rectangular sizes (512×1024, 1024×2048) were selected for training and testing, and the mIoU fluctuations of the segmentation results were analyzed.
[0155] Experiment 3: Quantitative and qualitative comparison on the UAVid test set
[0156] On the UAVid test set, using metrics such as mean intersection over union (mIoU), MeanF1 score (mean F1), and overall accuracy (OA), the method of the present invention was quantitatively compared with other advanced remote sensing image segmentation algorithms. In addition, visual diagrams of the segmentation results were provided to intuitively show the differences in segmentation performance.
[0157] Experimental conditions:
[0158] The experimental models of the present invention were all carried out on a server equipped with 4 NVIDIA Tesla A100 40GB GPUs and 2 Intel Xeon 6330 CPUs. The experimental environment was Ubuntu 22.04, the programming language was Python 3.10, and the deep learning framework was PyTorch 2.0.1. To accelerate convergence, the models were all trained using the AdamW optimizer. The basic learning rate was set to 6×10 -4 , and the cosine annealing strategy was adopted for learning rate adjustment.
[0159] During the training process on the UAVid dataset, the input image size was 1024×1024, and data augmentation techniques were used, including random vertical flipping, horizontal flipping, and random brightness adjustment. The training epoch was set to 60, and the batch size was 16. In the test phase, the test augmentation strategy of vertical flipping and horizontal flipping was used.
[0160] Evaluation metrics:
[0161] This application adopted the mean intersection over union (mIoU), Mean F1 score (mean F1), and overall accuracy (OA) as evaluation metrics. The mean intersection over union (mIoU) measures the accuracy of the prediction by calculating the ratio of the intersection and union between the prediction result and the ground truth. The F1 score comprehensively considers recall and precision and is used to evaluate the balance and integrity of the prediction result. The overall accuracy (OA) represents the proportion of correctly classified pixels in the total number of pixels.
[0162]
[0163] Among them, N is the number of semantic categories. TP i represents the true positive, FP i represents the false positive, FN i represents the false negative. Precision refers to the proportion of positive example samples correctly predicted by the model among all samples predicted as positive examples, and recall refers to the proportion of all actual positive example samples captured by the model among all actual positive example samples.
[0164] Experimental results and analysis:
[0165] As shown in Table 2, Table 2 shows the quantization comparison results with the state-of-the-art lightweight networks on the UAVid test set. Compared with multiple advanced segmentation networks, the present invention performs excellently in multiple key metrics:
[0166] Compared with the fastest model Fast-SCNN, the mIoU of the present invention has increased by 24.6%;
[0167] Compared with the UNetformer model with a similar number of parameters, the inference speed of the present invention reaches 106.9 FPS, and the mIoU is increased by 7.7%;
[0168] Compared with the Mamba structure, the mIoU of CMT-UNet is increased by 1.8%;
[0169] The mIoU of the present invention is 11.3% higher than that of the pure Transformer network Segmenter.
[0170] Further verified the effectiveness of the hybrid structure, CSAM and RCM modules in the present invention.
[0171] Table 2
[0172]
[0173] To evaluate the stability of the network, the present invention trained the CMT-UNet model under different input sizes, including square sizes (512×512, 1024×1024, 2048×2048) and rectangular sizes (512×1024, 1024×2048). The experimental results are summarized in Table 3, showing that the present invention performs stably under different input sizes, and the mIoU deviation is less than 0.7%. Among them, the input size of 1024×1024 achieved the highest mIoU on the UAVid validation set. In addition, the square input obtained a higher score than the rectangular input, but an overly large input size (such as 2048×2048) will reduce the mIoU of small targets.
[0174] Table 3
[0175]
[0176] The present invention trained multiple segmentation networks and conducted a detailed result comparison on the official Vaihingen test set. As shown in Table 4, the present invention has an advantage in the IoU performance of each category and achieved the best result of 69.1% in mIoU. Specifically, the invented remote sensing image segmentation method not only exceeds the efficient network ABCNet based on CNN by 6.8% in mIoU, but also is 3.3% and 1.8% higher than the latest hybrid Mamba network RS-Mamba and CM-UNet respectively. Especially in the difficult point of segmenting small targets in the category of 'person', the present invention leads other methods by at least 8.9% in the IoU index. Figure 7 Showed the comparison of the segmentation results between the present invention and the existing methods on the Vaihingen dataset. The visualization results prove the superior segmentation performance of the present invention for complex scenes. Figure 8 Showed the visualization comparison diagram between the present invention and the existing methods on the Potsdam dataset.
[0177] Table 4
[0178]
[0179] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0180] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A hybrid structure remote sensing image segmentation method based on a state space model, characterized in that: include: Acquire a remote sensing image to be processed, input the remote sensing image to be processed into a hybrid structure remote sensing image segmentation model, and obtain a segmentation result; The hybrid structure remote sensing image segmentation model is obtained by training with a training set, and the training set includes: a remote sensing image; The hybrid structure remote sensing image segmentation model uses a residual network sub-model based on a convolutional neural network and an improved Mamba sub-model to form a dual encoder structure for extracting multi-scale features, and a Transformer sub-model based on a hybrid attention convolution module to form a decoder for capturing global feature dependencies through the multi-scale features while extracting local features and capturing spatial information to generate the segmentation result; wherein the improved Mamba sub-model is the original Mamba sub-model that introduces a two-dimensional selective scanning module and a CSAM attention fusion module.
2. The hybrid structure remote sensing image segmentation method based on state space model according to claim 1, characterized in that: The residual network sub-model includes: Several residual blocks are used to perform convolution operations and generate multi-scale features; Several RCM modules are used to fuse the multi-scale features with the different-scale features generated by the improved Mamba sub-model, and input the fusion results to the decoder.
3. The hybrid structure remote sensing image segmentation method based on state space model according to claim 2 is characterized in that: The improved Mamba sub-model is used to generate different scale features including: After the input features are normalized, they are divided into two paths for subsequent processing: The first path: It passes through linear layers, depthwise convolution, 2D selective scanning modules, and layer normalization in sequence to extract features; The second path: Through the CSAM attention fusion module, the output results of the first path are fused based on element-wise multiplication to generate features of different scales.
4. The hybrid structure remote sensing image segmentation method based on state space model according to claim 1, characterized in that: The expression of the improved Mamba sub-model is: Q x =f LN (x) W x =f CSAM (Q x ) S x =f LN (f SS2D (D conv (f Linear (Q x )))))⊙W x Among them, x represents the input feature, f LN (·) is the layer normalization operation, f CSAM (·) is the fusion attention operation, f Linear (·) is a linear transformation operation, D conv (·) is the depthwise convolution operation, ⊙ is the element-wise multiplication operation, and f SS2D (·) is a 2D selective scanning operation, F x VSS represents the result obtained by processing the input feature x through the VSS block, Q x is the normalized result, W x is the result of attention operation, S x Select the sweep operation and residual connection results for 2D.
5. The hybrid structure remote sensing image segmentation method based on state space model according to claim 1, characterized in that: The two-dimensional selective scanning module comprises: The two-dimensional selective scanning submodule is used to scan the input feature map forward and backward by cross scanning. During the scanning process, the image is expanded in multiple directions to generate different sequences, and the different sequences are processed by S6 operation.
6. The hybrid structure remote sensing image segmentation method based on state space model according to claim 5, characterized in that: Processing the different sequences using the S6 operation includes: The input feature map is subjected to three linear changes to obtain the first change result, the second change result and the third change result respectively. The first change result is subjected to element-by-element exponential operation to introduce nonlinear dynamic characteristics, and the second change result is calculated using matrix transformation. The matrix transformation is a composite operation of matrix exponential and inverse operation. The initial state and the input feature tensor are further combined to update the intermediate state to provide context information for the sequence. Each time step is calculated using a linear combination method, and the state space vector operation is performed on the third change result to obtain the output result, and all time steps are integrated to generate a complete output sequence.
7. The hybrid structure remote sensing image segmentation method based on state space model according to claim 6 is characterized in that: The expression for processing the different sequences using the S6 operation is: X O =S se (x,O) Among them, O represents different directions, X represents input features, and S se (·) and S sm (·) represent scan expansion and scan merging operations, S6(·) represents S6 operation, X O express, Respectively represent the results of four different scanning directions of S6.
8. The hybrid structure remote sensing image segmentation method based on state space model according to claim 1, characterized in that: The CSAM attention fusion module includes: The channel attention module is used to extract channel features through global average pooling and maximum pooling operations, further extract local features using the first depthwise separable convolution, and enhance important channels through the first sigmoid activation function; A spatial attention module, for combining the channel features, performing feature extraction through the first convolution block, and obtaining key spatial positions in the feature map through the second sigmoid activation function; A multi-scale convolution module, for extracting features of different scales using second depthwise separable convolution kernels of different sizes in combination with batch normalization and ReLU activation function, and integrating the features of different scales using channel rearrangement and the second convolution block; Among them, the channel attention module, the spatial attention module and the multi-scale convolution module are connected in sequence.
9. The hybrid structure remote sensing image segmentation method based on state space model according to claim 1, characterized in that: Generating the segmentation result using the decoder includes: After the multi-scale features are batch normalized to standardize their distribution, they are divided into two parallel processing paths for subsequent processing: The first processing path: Through the window multi-head self-attention module, the global feature dependency is captured, and the window mechanism is used to reduce the computational complexity; The second processing path: local feature extraction and spatial information capture through the convolution group module; The output results of the two processing paths are merged and batch normalized again, and the merged features are input into the multi-layer perceptron layer for nonlinear transformation. The multi-scale features are further superimposed with the features after nonlinear transformation through residual connection to generate the segmentation result.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method based on multi-scale attention and double encoders
CN117496151A
Remote sensing image segmentation method based on dual-branch multi-scale feature fusion
CN118314353A
Optical remote sensing image segmentation method based on VMama model
CN118365882A
U-shaped multi-modal fusion segmentation method based on combination of graph neural network and Mamba model
CN119006813A
Double-feature fusion semantic segmentation system and method based on internet of things perception
WO2022227913A1
Cited By
Non-contact blood pressure detection method and system based on deep learning
CN120336833A
Pathological image colon gland segmentation method, device, equipment, medium and program product
CN120339630A
Remote sensing image building extraction method and system based on visual Mama model
CN120599504A
Mamb-Conv mixed multi-scale fusion remote sensing image color correction method and system
CN120807375A
A mamba-conv hybrid multi-scale fusion remote sensing image color correction method and system
CN120807375B