Remote sensing image segmentation method based on state space dual transformation and channel mixing
Through the method of state space dual transformation and channel mixing, the feature extraction and information interaction of remote sensing image segmentation are optimized, and the problem of high computing complexity in remote sensing image segmentation is solved, and efficient and accurate remote sensing image segmentation is achieved, which is suitable for edge devices and cloud processing.
Patent Information
- Application Number
- CN202510658433.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-21
AI Technical Summary
Existing deep learning methods have high computational complexity, limited global information interaction, and insufficient feature expression capabilities in remote sensing image segmentation. Especially when large-scale remote sensing data processing, computing resources are huge, making it difficult to apply to resource-constrained edge devices or real-time processing tasks.
The method of state space dual transformation and channel mixing is adopted, and the dimensionality reduction features are mapped to the hidden state space through non-causal state space dual transformation for channel mixing. Combined with the channel displacement enhancement mechanism, feature extraction and information interaction are optimized, and cross-modal information fusion capabilities are improved.
While reducing computing overhead, the accuracy and robustness of remote sensing image segmentation are improved, making it suitable for edge computing devices and cloud-based parallel processing tasks, and improving the edge accuracy and robustness of segmentation results.
Smart Images

Figure CN120472177A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing image segmentation method based on state space dual transformation and channel mixing. Background Art
[0002] Remote sensing image segmentation is a core task in many fields, including geographic information systems, environmental monitoring, and agricultural analysis. Traditional remote sensing image segmentation methods rely primarily on rule-based methods such as threshold segmentation, region growing, and edge detection. While these methods can achieve good segmentation results under specific conditions, they have limitations when dealing with complex backgrounds, multi-scale features, and high-resolution data.
[0003] In recent years, the rapid development of deep learning technologies, particularly models such as convolutional neural networks (CNNs) and visual transformers (ViTs), has provided a new path for remote sensing image segmentation. However, existing deep learning methods still face challenges such as high computational complexity, limited global information exchange, and insufficient feature expression capabilities. This is especially true when processing large-scale remote sensing data, as they require enormous computing resources and are difficult to apply to resource-constrained edge devices or real-time processing tasks.
[0004] Therefore, how to provide a remote sensing image segmentation method that can improve segmentation accuracy while reducing computational overhead is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0005] In response to the above research status, the present invention provides a remote sensing image segmentation method based on state-space dual transformation and channel mixing. On the one hand, it combines state-space modeling and hidden state mixing mechanism to optimize feature extraction and feature interaction processes, making information transmission more efficient; on the other hand, it introduces feature channel mixing and displacement enhancement mechanism to improve the cross-modal information fusion capability, making the segmentation results more accurate and robust.
[0006] The present invention provides a remote sensing image segmentation method based on state space dual transformation and channel mixing, comprising the following steps:
[0007] S1: Perform dimensionality reduction and feature extraction on the remote sensing input image through a series of convolution and linear transformation operations to obtain reduced-dimensionality features, compress the reduced-dimensionality features to generate hidden states, map the reduced-dimensionality features to the hidden state space using non-causal state space dual transformation, perform channel mixing in the hidden state space, and generate hidden state features;
[0008] S2: performing channel mixing after dimensionality reduction on the dimensionality reduction feature and the hidden state feature respectively; adding the dimensionality reduction feature and the hidden state feature to fuse the global information of the two; performing channel shift on the global information and performing channel shift on the dimensionality reduction feature and the hidden state feature after channel mixing respectively; performing channel mixing on the dimensionality reduction feature and the hidden state feature after channel shift again to extract local spatial information, and fusing them to generate the final enhanced feature;
[0009] S3: Perform a segmentation task based on the final enhanced features and output a segmentation result.
[0010] Preferably, the step of reducing the dimension of the input image in S1 includes:
[0011] Receive remote sensing input images, perform downsampling through four consecutive 3×3 convolutional layers, and generate feature maps after dimensionality reduction.
[0012] Preferably, before extracting features from the remote sensing input image in S1, the following steps are further included:
[0013] The feature map after dimensionality reduction is subjected to depthwise separable convolution processing, and is combined with the feature map after dimensionality reduction through residual connection. After normalization operation, the enhanced feature representation is obtained; feature extraction is performed on the enhanced feature representation.
[0014] Preferably, the step of compressing the dimensionality reduction feature to generate a hidden state in S1 includes:
[0015] The feature map after dimensionality reduction is transformed linearly to generate a projection matrix and a time step parameter. The discrete state is obtained by exponential transformation based on the projection matrix and the time step parameter, and the hidden state h is generated by combining the convolution operation result of the projection matrix with the Hadamard multiplication. - .
[0016] Preferably, the step of compressing the dimensionality reduction feature to generate a hidden state in S1 includes:
[0017] S11: Features after the dimensionality reduction features are enhanced After three linear transformations, the projection matrix and time step parameters are generated respectively:
[0018]
[0019] Where Linear represents the linear layer, ξ1, ξ3 are projection matrices, and ξ2 is the time step parameter;
[0020] S12: Use non-causal state space dual transformation for feature mapping and obtain discrete states through exponential transformation:
[0021]
[0022] Where, is the hidden state dynamic coefficient, I is the remote sensing input image;
[0023] S13: Combine the convolution operation results of the projection matrix and use Hadamard multiplication to generate the hidden state h - :
[0024] h - =θ⊙DWConv(ξ).
[0025] Preferably, the step of mapping the dimensionality reduction features to the hidden state space using the non-causal state space dual transformation in S1, performing channel mixing in the hidden state space, and generating the hidden state features comprises:
[0026] S14: Enhance the features after the dimensionality reduction features Map to the hidden state space and perform the non-causal state space dual transformation:
[0027]
[0028] Where h is the hidden state after mapping, which is used to capture long-range dependencies;
[0029] S15: Perform channel mixing on the mapped hidden state h and projection matrix ξ3 in the hidden state space to generate the hidden state feature h + :
[0030]
[0031] Where Sigmoid represents the sigmoid activation function.
[0032] Preferably, S2 comprises the following steps:
[0033] S21: Combine the dimension reduction feature φ and the hidden state feature h + Through 1×1 convolution layer dimensionality reduction, we get φ 1*1 , And mix the two channels;
[0034] S22: The channel-mixed features are respectively passed through 3×3 convolutional layers to extract local spatial information, and φ is obtained. 3*3 and
[0035] S23: Combine the dimension reduction feature φ and the hidden state feature h + Add them together to fuse the global information of the two, and then pass through a 3×3 depth-separable convolution layer and residual connection to obtain the feature φh :
[0036]
[0037] Where, Represents matrix addition operation;
[0038] S24: For feature φ h After channel displacement, they are respectively 3*3 and Add them together and get and
[0039] S25: and After channel mixing, they are added together through 1×1 convolution layers to obtain the final enhanced feature φ + , and serves as the input for subsequent segmentation tasks.
[0040] Preferably, the channel shift operation step includes: performing cyclic shift on the input features in the channel dimension direction, that is, moving the channel forward or backward by a certain step length to adjust the arrangement order of the input features in the channel dimension.
[0041] Preferably, said S3 includes a segmentation step of segmenting the model:
[0042] Restoring the final enhanced features step by step to the resolution of the remote sensing input image based on a step-by-step upsampling strategy;
[0043] Feature extraction and classification are performed on the restored feature map to obtain the predicted segmentation probability map P.
[0044] Preferably, S3 includes the step of training the segmentation model:
[0045] The segmentation model is trained by combining the composite loss function of Dice Loss and cross entropy loss:
[0046]
[0047]
[0048] Where α is the weight coefficient, i is the pixel index, and there are N pixels in total. i is the true label, P i is the predicted probability of the corresponding category, and ∈ is a parameter less than the preset value.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] Computational efficiency and resource utilization optimization: This method reduces the computational overhead during feature transformation through a state-space dual design of hidden state mixing, making global information interaction more efficient. At the same time, the introduction of channel mixing and channel shifting mechanisms ensures the effectiveness of feature fusion while avoiding the high computational complexity associated with traditional attention mechanisms. Therefore, this method can significantly reduce computational resource utilization and increase inference speed when processing large-scale, high-resolution remote sensing imagery, making it suitable for parallel processing tasks on edge computing devices and in the cloud.
[0051] Feature fusion and improved segmentation accuracy: This invention achieves efficient interaction of multiple features through modules such as channel mixing, channel shifting, and depthwise separable convolution, allowing different data to share information more fully and avoiding information redundancy and feature independence issues. At the same time, the combination of jump connections and multi-level feature fusion in the decoding process enhances the ability to capture detailed information and improves the edge accuracy and robustness of the segmentation results. Experiments have shown that this method has higher accuracy and generalization ability than traditional methods in remote sensing image segmentation tasks, especially in distinguishing complex land object categories.
[0052] In summary, the present invention further improves the accuracy and robustness of remote sensing image segmentation while ensuring efficient calculation, and provides a more efficient and accurate solution for remote sensing data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only embodiments of the present invention. Those skilled in the art can also derive other drawings based on the provided drawings without inventive effort.
[0054] Figure 1 Flowchart of a remote sensing image segmentation method based on state space dual transformation and channel mixing provided by an embodiment of the present invention;
[0055] Figure 2 1 is a schematic diagram of the state space dual transformation of hidden state mixing provided by an embodiment of the present invention;
[0056] Figure 3 This is a schematic diagram of the principle of feature channel mixing and displacement enhancement provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0058] The present invention discloses a remote sensing image segmentation method based on state space dual transformation and channel mixing. In order to solve the problems of complex object categories, large scale changes, and difficulty in obtaining global information in remote sensing images, the present invention combines state space models and hidden state mixing technology to achieve efficient global information interaction, and optimizes cross-modal information fusion through feature channel mixing and displacement enhancement mechanism to improve the segmentation effect of remote sensing images. Figure 1 As shown, the following steps are included:
[0059] S1: Reduce the dimension and extract features of the remote sensing input image through a series of convolution and linear transformation operations to obtain reduced-dimensional features, compress the reduced-dimensional features to generate hidden states, use non-causal state space dual transformation to map the reduced-dimensional features to the hidden state space, perform channel mixing in the hidden state space, and generate hidden state features;
[0060] S2: Reduce the dimension of the reduced dimensionality features and hidden state features respectively and then perform channel mixing; add the reduced dimensionality features and hidden state features to fuse the global information of the two; perform channel shift on the global information and then perform channel shift on the reduced dimensionality features and hidden state features after channel mixing respectively; the reduced dimensionality features and hidden state features after channel shift are again channel mixed to extract local spatial information, and then fuse them to generate the final enhanced features;
[0061] S3: Perform segmentation tasks based on the final enhanced features and output segmentation results.
[0062] In order to fully reflect the role of each step of the present invention, the feature extraction stage of the encoder is described as follows:
[0063] During the feature extraction phase, this embodiment introduces an efficient hidden state hybrid method to optimize global information interaction in remote sensing images. First, the input image undergoes a series of convolution operations to reduce dimensionality, reducing computational complexity and extracting key features. Subsequently, state-space transformation techniques are employed to map the original features into a hidden state space through linear transformations and state-space dual mappings, enabling feature interaction in a more compact representation.
[0064] In particular, this embodiment employs a non-causal state space transformation method to effectively capture long-range dependencies and enhance global information modeling capabilities. Furthermore, to further enhance feature expression capabilities, a hidden state mixing strategy is proposed, which performs channel mixing within the state space, thereby reducing computational complexity and improving computational efficiency. This design optimizes the model's computational resource consumption while ensuring efficient information exchange, making it more suitable for large-scale remote sensing data processing.
[0065] In order to further enhance the feature expression capability of remote sensing images, this embodiment introduces a feature channel mixing and displacement enhancement mechanism to optimize the fusion process of different modal information. In the feature mixing process, convolution dimensionality reduction is performed to enable feature information from different sources to interact in a more compact manner. Next, a channel mixing strategy is introduced to divide channel groups and perform cross-exchanges so that information from different modalities can be more effectively fused to avoid information redundancy and monotonous linear combinations. Compared with traditional splicing or weighted summation methods, the channel mixing method of this embodiment can promote efficient fusion of cross-modal features and improve segmentation accuracy and generalization ability.
[0066] Furthermore, to further enhance feature sharing, this embodiment provides an enhancement mechanism based on channel shifting. This method adjusts the arrangement of feature channels to enable more efficient interaction between features from different modalities, similar to the spatial attention mechanism, but without the added computational overhead. This operation allows for more efficient utilization of multi-source information in remote sensing images, improving final segmentation performance.
[0067] In one embodiment, Figure 2 As shown, the steps of reducing the dimensionality of the input image in S1 include:
[0068] Take the original input image I, whose size is H×W×3 (where H and W represent the image height and width respectively, and 3 represents the number of RGB channels). To extract features, it is first downsampled through four consecutive 3×3 convolutional layers, with the stride of each convolution layer set to 2 to reduce computational complexity and obtain key features. Finally, the feature map φ after dimensionality reduction is obtained:
[0069] φ=DWConv(I)
[0070] This feature map provides preliminary information expression for subsequent state space transformation.
[0071] In one embodiment, Figure 2 As shown, before feature extraction of the remote sensing input image in S1, the following steps are also included:
[0072] In order to further improve the feature expression ability, φ is processed by depthwise separable convolution (DWConv), and then residual connection is performed with the original φ, and the enhanced feature representation is obtained through the normalization layer (LayerNormalization)
[0073]
[0074] This process not only enhances the feature extraction capability, but also improves the stability and generalization ability of the network.
[0075] Multi-layer depthwise separable convolution (DWConv) and residual connection mechanisms are used to further optimize information flow and feature representation capabilities. Depthwise separable convolution can effectively reduce computational complexity while preserving local spatial information, while residual connections can maintain the information flow of original features, avoid the gradient vanishing problem, and improve model stability.
[0076] In one embodiment, the step of compressing the dimensionality reduction features to generate hidden states in S1 includes:
[0077] The feature map after dimensionality reduction is transformed linearly to generate the projection matrix and time step parameters. The discrete state is obtained by exponential transformation based on the projection matrix and time step parameters, and the hidden state h is generated by combining the convolution operation result of the projection matrix with Hadamard multiplication. - .
[0078] In one embodiment, Figure 2 As shown, the steps of compressing the dimensionality reduction features to generate hidden states in S1 include:
[0079] S11: Features after dimensionality reduction features are enhanced After three linear transformations, the projection matrix and time step parameters are generated respectively:
[0080]
[0081] Where Linear represents the linear layer, ξ1, ξ3 are projection matrices, and ξ2 is the time step parameter;
[0082] S12: Use non-causal state space dual transformation for feature mapping and obtain discrete states through exponential transformation:
[0083]
[0084] Where, is the hidden state dynamic coefficient, I is the remote sensing input image;
[0085] S13: Combine the convolution operation results of the projection matrix and use Hadamard multiplication to generate the hidden state h - :
[0086] h - =θ⊙DWConv(ξ).
[0087] The key to this step is to migrate traditional feature space transformation operations to the hidden state space, reducing computational complexity. This optimization is particularly suitable for high-resolution remote sensing image processing, effectively reducing memory usage and improving computational efficiency.
[0088] In one embodiment, in S1, the steps of mapping the reduced-dimensional features to the hidden state space using the non-causal state space dual transformation, performing channel mixing in the hidden state space, and generating the hidden state features include:
[0089] S14: Enhance the features after dimensionality reduction Map to the hidden state space and perform the non-causal state space dual transformation:
[0090]
[0091] Where h is the hidden state after mapping, which is used to capture long-range dependencies;
[0092] S15: Perform channel mixing on the mapped hidden state h and projection matrix ξ3 in the hidden state space to generate the hidden state feature h + :
[0093]
[0094] Where Sigmoid represents the sigmoid activation function.
[0095] By performing channel mixing in the hidden state space, this embodiment effectively reduces the computational overhead of the SSD layer and optimizes computational complexity while maintaining high expressiveness. This approach is particularly suitable for large-scale remote sensing data processing tasks on edge devices or in the cloud.
[0096] In one embodiment, Figure 3 As shown, S2 includes the following steps:
[0097] S21: Combine the dimension reduction feature φ and the hidden state feature h + Through 1×1 convolution layer dimensionality reduction, we get φ 1*1 , The two are then channel-mixed; channel mixing connects two data streams with different characteristics, and its effects are as follows:
[0098] Cross-modal feature interaction: By disrupting and rearranging channels of different modalities, information is effectively integrated to avoid information loss caused by separate processing.
[0099] Improve feature sharing capabilities: Break the feature independence of a single modality, enable complementary learning, and improve segmentation accuracy.
[0100] Enhanced information expression: Compared with direct splicing or addition operations, the channel mixing mechanism can more effectively promote the flow of cross-modal features and avoid the information redundancy caused by monotonous linear combinations.
[0101] In this embodiment, the specific steps of channel mixing include dividing the channels into multiple groups and performing channel cross-exchange within the groups to more efficiently fuse information from different channels. The rearranged features are used in subsequent convolution calculations to ensure that the distribution of the mixed features has stronger representation capabilities.
[0102] S22: The channel-mixed features are respectively passed through 3×3 convolutional layers to extract local spatial information, and φ is obtained. 3*3 and This operation helps to enhance the edge, texture and structural information of different features to adapt to the complex landform types in remote sensing images.
[0103] S23: Combine the dimension reduction feature φ and the hidden state feature h + The two features are added together to fuse the global information of the two, and then the feature φ is obtained through a 3×3 depth-separable convolution layer (to extract local correlation) and a residual connection (not only retaining the original features, but also improving training stability and avoiding gradient disappearance). h :
[0104]
[0105] Where, Represents matrix addition operation;
[0106] S24: For feature φ h After channel displacement, they are respectively 3*3 and Add them together and get and The introduction of channel shift operation promotes deep information interaction between features, enables more effective joint decision-making of multi-source data in remote sensing images, and enables different modalities to share information more fully. The role of channel shift:
[0107] Feature information transfer: The channel shift mechanism enables features of different modalities to influence each other rather than being calculated in isolation, thereby improving the fusion effect.
[0108] Modal perception enhancement: By adjusting the channel order, different modalities can capture each other's contextual information, similar to the idea of ShiftNet or SpatialShiftModule (SSM), to achieve lightweight information interaction.
[0109] Computationally efficient: Compared to additional attention mechanisms or complex transformations, channel shift does not introduce additional parameters and only adjusts the feature arrangement, so it has lower computational overhead.
[0110] S25: and After channel mixing, they are added together through 1×1 convolution layers to obtain the final enhanced feature φ + ,This feature carries multiple class fusion information and will be used as the input of the subsequent segmentation task, making the final segmentation result more robust and accurate.
[0111] In one embodiment, the channel shift operation step includes: performing cyclic shift on the input features in the channel dimension direction, that is, moving the channel forward or backward by a certain step length to adjust the arrangement order of the input features in the channel dimension.
[0112] In one embodiment, to optimize segmentation performance, S3 includes a training step for the segmentation model to ensure that the model addresses both class imbalance and pixel-level classification accuracy:
[0113] S301: Combine DiceLoss (for handling class imbalance) and cross entropy loss to train the segmentation model:
[0114]
[0115] Where α is the weight coefficient, i is the pixel index, and there are N pixels in total. i is the true label, P i is the predicted probability of the corresponding category, ∈ is a small value to prevent the denominator from being zero.
[0116] S302: Use the Adam optimizer for gradient update, and set the learning rate to an adaptive decay strategy (such as cosineannealing or poly decay) to ensure stable model convergence:
[0117]
[0118] Where η is the current learning rate, is the loss gradient.
[0119] In one embodiment, the trained and optimized model directly performs forward propagation on the input image during the inference phase and outputs a segmentation probability map P. Finally, the category label of each pixel is obtained through the argmax operation to obtain the final segmentation result:
[0120]
[0121] Where, is the final pixel-level classification result.
[0122] So far, the method of the present invention has completely realized the whole process from remote sensing image input to final segmentation result.
[0123] In the specific implementation, due to φ + In the first two steps, the image undergoes dimensionality reduction and feature transformation. In order to restore the spatial scale to match the input image, this embodiment provides a decoder-based step-by-step upsampling strategy.
[0124] The specific process is as follows:
[0125] S311: Deconvolution upsampling: Through a series of deconvolution operations, the features are gradually restored to the original resolution. The step size is set to 2 each time the upsampling is performed to maintain the continuity of the information flow.
[0126] S312: Skip Connection: During the decoding process, with the help of the U-Net structure, the shallow features φ generated in S1 are concatenated with the progressively upsampled features to enhance local detail information while reducing the blurring effect of high-level features.
[0127] S313: Fusion operation: During the upsampling process, a 3×3 convolutional layer (with BatchNormalization and ReLU) is used to further extract the fused features to ensure the adequacy of information expression.
[0128] S314: Reduce the number of channels to the number of categories C through a 1×1 convolution layer to obtain the predicted segmentation probability map P.
[0129] The above is a detailed introduction to the remote sensing image segmentation method based on state-space dual transformation and channel mixing provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
[0130] In this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. A remote sensing image segmentation method based on state space dual transformation and channel mixing, comprising the following steps: S1: Perform dimensionality reduction and feature extraction on the remote sensing input image through a series of convolution and linear transformation operations to obtain reduced-dimensionality features, compress the reduced-dimensionality features to generate hidden states, map the reduced-dimensionality features to the hidden state space using non-causal state space dual transformation, perform channel mixing in the hidden state space, and generate hidden state features; S2: performing channel mixing after dimensionality reduction on the dimensionality reduction feature and the hidden state feature respectively; adding the dimensionality reduction feature and the hidden state feature to fuse the global information of the two; performing channel shift on the global information and performing channel shift on the dimensionality reduction feature and the hidden state feature after channel mixing respectively; performing channel mixing on the dimensionality reduction feature and the hidden state feature after channel shift again to extract local spatial information, and fusing them to generate the final enhanced feature; S3: Perform a segmentation task based on the final enhanced features and output a segmentation result.
2. The remote sensing image segmentation method based on state space dual transformation and channel mixing according to claim 1, characterized in that: The step of reducing the dimension of the input image in S1 includes: Receive remote sensing input images, perform downsampling through four consecutive 3×3 convolutional layers, and generate feature maps after dimensionality reduction.
3. The remote sensing image segmentation method based on state space dual transformation and channel mixing according to claim 1, characterized in that: Before the feature extraction of the remote sensing input image in S1, the following steps are also included: The feature map after dimensionality reduction is subjected to depthwise separable convolution processing, and is combined with the feature map after dimensionality reduction through residual connection. After normalization operation, the enhanced feature representation is obtained; feature extraction is performed on the enhanced feature representation.
4. The remote sensing image segmentation method based on state space dual transformation and channel mixing according to claim 1, characterized in that: The step of compressing the dimension reduction feature to generate a hidden state in S1 includes: The feature map after dimensionality reduction is transformed linearly to generate a projection matrix and a time step parameter. The discrete state is obtained by exponential transformation based on the projection matrix and the time step parameter, and the hidden state h is generated by combining the convolution operation result of the projection matrix with the Hadamard multiplication. - .
5. The remote sensing image segmentation method based on state space dual transformation and channel mixing according to claim 4, characterized in that: The step of compressing the dimension reduction feature to generate a hidden state in S1 includes: S11: Features after the dimensionality reduction features are enhanced After three linear transformations, the projection matrix and time step parameters are generated respectively: Where Linear represents the linear layer, ξ1, ξ3 are projection matrices, and ξ2 is the time step parameter; S12: Use non-causal state space dual transformation for feature mapping and obtain discrete states through exponential transformation: Where, is the hidden state dynamic coefficient, I is the remote sensing input image; S13: Combine the convolution operation results of the projection matrix and use Hadamard multiplication to generate the hidden state h - : h - =θ☉DWConv(ξ).
6. The remote sensing image segmentation method based on state space dual transformation and channel mixing according to claim 5, characterized in that: In S1, the steps of mapping the dimensionality reduction features to the hidden state space using the non-causal state space dual transformation and performing channel mixing in the hidden state space to generate the hidden state features include: S14: Enhance the features after the dimensionality reduction features Map to the hidden state space and perform the non-causal state space dual transformation: Where h is the hidden state after mapping, which is used to capture long-range dependencies; S15: Perform channel mixing on the mapped hidden state h and projection matrix ξ3 in the hidden state space to generate the hidden state feature h + : Where Sigmoid represents the sigmoid activation function.
7. The remote sensing image segmentation method based on state space dual transformation and channel mixing according to claim 1, characterized in that: The S2 comprises the following steps: S21: Combine the dimension reduction feature φ and the hidden state feature h + Through 1×1 convolution layer dimensionality reduction, we get φ 1*1 , And mix the two channels; S22: The channel-mixed features are respectively passed through 3×3 convolutional layers to extract local spatial information, and φ is obtained. 3*3 and S23: Combine the dimension reduction feature φ and the hidden state feature h + Add them together to fuse the global information of the two, and then pass through a 3×3 depth-separable convolution layer and residual connection to obtain the feature φ h : Where, Represents matrix addition operation; S24: For feature φ h After channel displacement, they are respectively 3*3 and Add them together and get and S25: and After channel mixing, they pass through 1×1 convolution layers and then add until the final enhanced feature φ + , and serves as the input for subsequent segmentation tasks.
8. The remote sensing image segmentation method based on state space dual transformation and channel mixing according to claim 1, characterized in that: The channel shift operation step includes: performing cyclic shift on the input features in the channel dimension direction, that is, moving the channel forward or backward by a certain step length to adjust the arrangement order of the input features in the channel dimension.
9. The remote sensing image segmentation method based on state space dual transformation and channel mixing according to claim 1, characterized in that: S3 includes the segmentation steps of the segmentation model: Restoring the final enhanced features step by step to the resolution of the remote sensing input image based on a step-by-step upsampling strategy; Perform feature extraction and classification on the restored feature map to obtain the predicted segmentation probability map P.
10. The remote sensing image segmentation method based on state space dual transformation and channel mixing according to claim 1, characterized in that: S3 includes the training steps of the segmentation model: The segmentation model is trained by combining the composite loss function of Dice Loss and cross entropy loss: Where α is the weight coefficient, i is the pixel index, and there are N pixels in total. i is the true label, P i is the predicted probability of the corresponding category, and ∈ is a parameter less than the preset value.
Citation Information
Patent Citations
Medical image recovery method based on state space dual mechanism
CN118781002A
Remote sensing image water body segmentation method based on comparative learning and multi-modal fusion
CN118864865A
Pathological image classification method based on state space duality
CN119048825A
Visual feature representation learning method and system based on two-dimensional state space model
CN119131409A
Double-feature fusion semantic segmentation system and method based on internet of things perception
WO2022227913A1