Remote sensing image segmentation method and system based on convolution-state space fusion and position trigger
Through the convolution-state space fusion and position trigger method, the problems of blurred boundaries, omission of small targets and difficulty in multi-source information fusion in remote sensing image segmentation are solved, and efficient and accurate remote sensing image segmentation is achieved, which is suitable for fields such as urban planning, agricultural monitoring and disaster assessment.
Patent Information
- Application Number
- CN202510955535.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-11
AI Technical Summary
Existing remote sensing image segmentation technology has problems such as blurred boundaries, small targets being easily ignored, and difficulty in multi-source information fusion when processing high-resolution, multi-source and multi-modal images. In particular, it has shortcomings in maintaining the computational efficiency and robustness of the model.
A method based on convolution-state space fusion and position triggers is adopted. Through patch embedding and spatial position encoding, multiple convolution-state space fusion units (C-SSFU) and position triggers are combined. The cross-modal attention mechanism is used for multi-scale feature fusion. Adaptive pooling and fully connected layers are introduced to construct a supervised loss function to improve the segmentation performance of boundaries and small objects.
It significantly improves the accuracy of small target and boundary segmentation, takes into account both computational efficiency and robustness, reduces the complexity of global attention and convolution operations, and enhances the network's tolerance to atmospheric scattering and noise. It is suitable for multi-source remote sensing image segmentation in high-resolution scenes.
Smart Images

Figure CN120635462A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and more particularly to a remote sensing image segmentation method and system based on convolution-state space fusion and position triggers. Background Art
[0002] With the rapid development of remote sensing imaging technology and geographic information systems (GIS), remote sensing images are increasingly being used in areas such as urban planning, agricultural monitoring, disaster assessment, and land cover change analysis. Remote sensing image segmentation, as a key step in achieving intelligent object recognition and regional semantic extraction, has become a core task in remote sensing intelligent analysis. Its goal is to assign semantic category labels to each pixel in high-resolution remote sensing imagery, enabling accurate object extraction.
[0003] However, due to the characteristics of remote sensing images such as high resolution, strong texture, multi-source and multi-modality (such as multispectral, panchromatic, SAR, etc.), and complex scenes, traditional image segmentation methods face many challenges in this field, such as blurred boundaries, small targets that are easily ignored, and difficulty in fusing multi-source information.
[0004] In recent years, deep learning technology, especially convolutional neural networks (CNNs) and Transformers, has made significant progress in image segmentation tasks. CNN has become the mainstream model architecture for remote sensing image segmentation due to its powerful local perception capabilities and parameter sharing mechanism. However, due to its limited receptive field, it is difficult to fully capture long-distance dependencies, resulting in certain limitations when dealing with large-scale targets, inhomogeneous objects, and complex spatial structures. The Transformer architecture can model global contextual information through the self-attention mechanism and has gradually been introduced into remote sensing image analysis in recent years. However, it has high computational cost and large data requirements, and has insufficient expressive power when dealing with local details (such as small targets or edges).
[0005] In existing technologies, in order to improve the accuracy of remote sensing image segmentation, researchers have tried to combine multiple mechanisms, such as multi-scale feature fusion, attention mechanism, pyramid structure and residual connection.
[0006] However, these solutions still have the following shortcomings: (1) The feature extraction stage is not sensitive enough to local boundaries and small targets, which can easily lead to blurred boundaries and missed detection of small targets; (2) Multimodal fusion, such as the integration of multispectral and SAR, has information redundancy or semantic inconsistency problems; (3) There is a conflict between global modeling and local details, and there is a lack of an efficient fusion mechanism to unify the long-term and short-term dependencies of modeling.
[0007] Therefore, how to solve the problems of blurred boundaries, missing small targets and difficulty in multi-source information fusion in current remote sensing image segmentation while maintaining the computational efficiency and robustness of the model is an urgent problem that technical personnel in this field need to solve. Summary of the Invention
[0008] In view of this, the present invention provides a remote sensing image segmentation method and system based on convolution-state space fusion and position trigger to solve some of the technical problems mentioned in the background technology.
[0009] In order to achieve the above object, the present invention adopts the following technical solutions:
[0010] A remote sensing image segmentation method based on convolution-state space fusion and position trigger includes the following steps:
[0011] S1. Preprocess the original remote sensing image and convert the large-scale image into a structured representation of a unified scale through patch embedding and spatial position encoding, while preserving the spatial position information;
[0012] S2. Decouple the position-encoded patch vector sequence through multiple convolutional-state-space fusion units (C-SSFUs). The convolution branch extracts fine-grained edge and local structural features, while the state-space modeling branch captures long-range contextual information. The two branch features are then fused at multiple scales using a cross-modal attention mechanism. Position triggers are introduced multiple times to dynamically enhance edge and small object response regions, ultimately outputting the fused features.
[0013] S3. Through adaptive pooling and fully connected layers, the fused features are mapped to the final pixel-level classification probability map, which is compared with the true label map. A supervised loss function is constructed to improve the model's segmentation performance for minority categories and boundary areas, and the final segmentation result is obtained.
[0014] Preferably, the preprocessing includes: geometric correction and registration, atmospheric correction and denoising, and spectral normalization or standardization.
[0015] Preferably, the specific contents of Patch embedding and spatial position encoding are:
[0016] Divide the preprocessed image into non-overlapping patches;
[0017] Flatten each small block into a one-dimensional vector, and then map it to the D-dimensional feature space through a linear mapping layer;
[0018] Based on the spatial position information of each small block in the original image, a position code is added.
[0019] Preferably, the specific architecture of multiple convolution-state space fusion units C-SSFU and multiple introduction of position triggers is: the patch vector sequence with position encoding is sequentially input into the two serial convolution-state space fusion units C-SSFU in the first stage, the position trigger is introduced, and after the first position trigger, two C-SSFUs, a position trigger, four C-SSFUs, a position trigger and two C-SSFUs are serially connected in sequence.
[0020] Preferably, the specific content of the convolution-state space fusion unit C-SSFU is:
[0021] S21. Input features The channel grouping strategy is used to decouple features and obtain the first feature and the second feature
[0022] S22. Convolutional feature enhancement branch: first feature Perform deep feature extraction, and obtain local geometric priors through batch normalization layer, 3*3 convolution layer, batch normalization layer, ReLU activation function, 3*3 convolution layer, batch normalization layer, ReLU activation function, point-by-point convolution layer, and ReLU activation function
[0023] S23. State-space modeling branch: Second feature Long-range context modeling is performed, and features are obtained through layer normalization, linear layer, 3*3 depth-separable convolution layer, SiLU activation function, two-dimensional selective scanning block, and layer normalization. At the same time, the second feature After a linear layer and activation function, we get and After matrix multiplication and layer normalization, global semantics is obtained
[0024] S24. Through cross-modal attention fusion, dual-branch features are used for multi-spectral collaborative modeling:
[0025] Will and Perform splicing to obtain splicing features Bimodal interaction: leveraging local geometric priors in convolutional branches Global semantics for guiding state-space branching After the linear layer, we get Q. After the linear layer, we get K; spectral sensitive attention: Q is multiplied by K and then activated by SiLU function. Multiply to get features
[0026] S25. Residual feature refinement: Constructing feature enhancement pathways, features After channel shuffling and initial input Perform residual connection and output That is the output feature of the convolution-state space fusion unit C-SSFU.
[0027] Preferably, the specific content of introducing the position trigger to dynamically strengthen the edge and small target response area is:
[0028] A two-dimensional coordinate channel is added to the input feature map through the coordinate convolution layer to enhance spatial sensitivity; the GELU activation function introduces nonlinearity to improve expressiveness; the conventional convolution layer extracts spatial context; a multi-head mechanism is adopted, the features extracted by convolution in the current position trigger are used as queries, and the global semantic feature map of the previous stage or higher layer is used as keys and values. The attention weight is calculated by dot product, adaptively amplifying the responses of the areas where edges and small targets are located, while suppressing irrelevant background features. Finally, the weighted semantic information is fed back to the current position feature to achieve cross-scale information interaction and enhancement; finally, the convolution layer and the Sigmoid architecture are activated to generate a position trigger mask for pixel-level weighting of the feature map.
[0029] Preferably, the specific content of step S3 is:
[0030] S31. Input the obtained fusion feature map into the adaptive global pooling layer, perform average pooling in the spatial dimension, retain the channel dimension information, and compress spatial redundancy;
[0031] S32. The pooled feature vector is linearly mapped through a fully connected layer and then activated by an activation function to obtain the final segmentation probability map, i.e., the probability value of each pixel belonging to each category.
[0032] S33. Compare the obtained segmentation probability map with the real label map provided when inputting the neural network, and construct a joint loss function based on cross entropy loss and Dice loss
[0033] A remote sensing image segmentation system based on convolution-state space fusion and position trigger, based on the remote sensing image segmentation method based on convolution-state space fusion and position trigger, comprising a feature preprocessing module, a patch embedding and spatial position encoding module, a feature extraction and fusion module and an output module;
[0034] Feature preprocessing module, used to preprocess the original remote sensing image;
[0035] The patch embedding and spatial position encoding module is used to convert the pre-processed remote sensing images into a structured representation of a unified scale through patch embedding and spatial position encoding, while retaining the spatial position information;
[0036] The feature extraction and fusion module is used to decouple the position-encoded patch vector sequence through multiple convolutional-state-space fusion units (C-SSFUs). The convolution branch extracts fine-grained edge and local structural features, while the state-space modeling branch captures long-range contextual information. The two branch features are fused at multiple scales using a cross-modal attention mechanism. Position triggers are introduced multiple times to dynamically enhance edge and small target response areas, and the final fused features are output.
[0037] The output module is used to map the fused features into the final pixel-level classification probability map through adaptive pooling and fully connected layers, and compare it with the true label map to construct a supervised loss function to improve the model's segmentation performance for minority categories and boundary areas, and obtain the final segmentation result.
[0038] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a remote sensing image segmentation method based on convolution-state space fusion and position trigger.
[0039] A processing terminal includes a memory and a processor. The memory stores a computer program that can be run on the processor. When the processor executes the computer program, the remote sensing image segmentation method based on convolution-state space fusion and position trigger is implemented.
[0040] The above technical solution shows that, compared with the prior art, the present invention provides a remote sensing image segmentation method and system based on convolution-state space fusion and position trigger. While maintaining the computational efficiency of the model, the system effectively solves the common problems of blurred boundaries, missed small targets, and difficulty in multi-source information fusion in current remote sensing image segmentation through the collaborative design of local convolution, state space modeling, and position-sensitive control. The system has good application prospects and promotion value.
[0041] This invention significantly improves the accuracy of small object and boundary segmentation: the position trigger module dynamically amplifies the edge and small object response, and introduces residual detail reconstruction during the upsampling process, effectively alleviating the problems of class imbalance and boundary blur. The convolution-state space fusion unit (C-SSFU) takes into account both local geometric details and long-range dependencies within the same module, making the model more sensitive to the segmentation of small objects.
[0042] This invention balances computational efficiency and robustness: based on the design of patch embedding and intensity grouping, it greatly reduces the complexity of global attention and convolution operations, and is suitable for high-resolution scenes; multi-source preprocessing and progressive residual structure enhance the network's tolerance to interference such as atmospheric scattering, speckle noise, and cloud occlusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0044] Figure 1 Schematic diagram of the remote sensing image segmentation method based on convolution-state space fusion and position trigger provided by the present invention;
[0045] Figure 2 Schematic diagram of the convolution-state space fusion unit C-SSFU provided by the present invention. DETAILED DESCRIPTION
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0047] The embodiment of the present invention discloses a remote sensing image segmentation method based on convolution-state space fusion and position trigger, such as Figure 1 , including the following steps:
[0048] S1. Preprocess the original remote sensing image and convert the large-scale image into a structured representation of a unified scale through patch embedding and spatial position encoding, while preserving the spatial position information;
[0049] S2. Decouple the position-encoded patch vector sequence through multiple convolutional-state-space fusion units (C-SSFUs). The convolution branch extracts fine-grained edge and local structural features, while the state-space modeling branch captures long-range contextual information. The two branch features are then fused at multiple scales using a cross-modal attention mechanism. Position triggers are introduced multiple times to dynamically enhance edge and small object response regions, ultimately outputting the fused features.
[0050] S3. Through adaptive pooling and fully connected layers, the fused features are mapped to the final pixel-level classification probability map, which is compared with the true label map. A supervised loss function is constructed to improve the model's segmentation performance for minority categories and boundary areas, and the final segmentation result is obtained.
[0051] In order to further implement the above technical solution, the preprocessing content includes: geometric correction and alignment, atmospheric correction and denoising, and spectral normalization or standardization.
[0052] In this embodiment, geometric correction and registration: using the positioning information of the remote sensing platform and ground control points (GCPs), geometric correction and registration are performed on multi-temporal images / multi-source images to ensure one-to-one correspondence between image pixels;
[0053] Atmospheric correction and denoising: Perform atmospheric correction (such as 6S or DOS method) on multispectral data to remove the influence of atmospheric scattering; apply Lee filter to synthetic aperture radar (SAR) data to suppress speckle noise;
[0054] Spectral normalization / standardization: Normalize the spectral values of each band to [0,1] or standard normal distribution to reduce the impact of spectral differences between different sensors and different acquisition conditions on model training.
[0055] To further implement the above technical solution, the specific contents of patch embedding and spatial position encoding are as follows:
[0056] In order to take into account the global context and local texture features of high-resolution remote sensing images, the preprocessed images Division Non-overlapping patches of size P×P, where H and W are the spatial dimensions of the remote sensing image, i.e., height and width;
[0057]
[0058] Each small block x i Flattened to a one-dimensional vector Then it is mapped to the D-dimensional feature space through the linear mapping layer (fully connected layer);
[0059]
[0060] Add position encoding based on the spatial position information of each small block in the original image
[0061] In this embodiment, a learnable parameter or a sine / cosine function may be used:
[0062]
[0063] Patch embedding and spatial position encoding help to: reduce computational complexity. The small-block operation effectively reduces the computational burden of performing global convolution or attention directly on high-resolution images; retain local texture information. Each patch carries rich local texture, boundary and detail features of the ground object, providing a basis for subsequent segmentation.
[0064] In order to further implement the above technical solution, the specific architecture of multiple convolution-state space fusion units C-SSFU and multiple introduction of position triggers is as follows: the patch vector sequence with position encoding is input into the two serial convolution-state space fusion units C-SSFU in the first stage in sequence, and the position trigger is introduced. After the first position trigger, two C-SSFUs, a position trigger, four C-SSFUs, a position trigger and two C-SSFUs are serially connected in sequence to output the final fusion feature.
[0065] Each C-SSFU combines local convolution to capture detail boundaries with a global state space mechanism to model long-range dependencies, achieves multi-scale feature fusion, and generates feature representations.
[0066] In order to further implement the above technical solutions, Figure 2 , the specific content of the convolution-state space fusion unit C-SSFU is:
[0067] S21. Input features The channel grouping strategy is used to decouple features and obtain the first feature and the second feature
[0068] In this embodiment, H and W are the spatial dimensions of the remote sensing image, C is the number of channels, and the segmentation ratio is set to 1:1 to maintain computational symmetry;
[0069] S22. Convolutional feature enhancement branch: first feature Perform deep feature extraction, and obtain local geometric priors through batch normalization layer, 3*3 convolution layer, batch normalization layer, ReLU activation function, 3*3 convolution layer, batch normalization layer, ReLU activation function, point-by-point convolution layer, and ReLU activation function
[0070] In this embodiment, double 3×3 convolution is used to form a receptive field increasing structure, combined with batch normalization (BN) and ReLU activation function to enhance the ability to depict the edges of objects;
[0071] S23. State-space modeling branch: Second feature Long-range context modeling is performed, and features are obtained through layer normalization, linear layer, 3*3 depth-separable convolution layer, SiLU activation function, two-dimensional selective scanning block, and layer normalization. At the same time, the second feature After a linear layer and activation function, we get and After matrix multiplication and layer normalization, global semantics is obtained
[0072] S24. Through cross-modal attention fusion, dual-branch features are used for multi-spectral collaborative modeling:
[0073] Will and Splice and get Bimodal interaction: leveraging local geometric priors in convolutional branches Global semantics for guiding state-space branching After the linear layer, we get Q. After the linear layer, we get K; spectral sensitive attention: Q is multiplied by K and then activated by SiLU function. Multiply to get features
[0074] S25. Construct the feature enhancement path to refine the residual features. After channel shuffling and initial input Perform residual connection and output That is the output feature of the convolution-state space fusion unit C-SSFU;
[0075] In this embodiment, the feature enhancement pathway includes: a channel shuffling strategy, which performs periodic permutation (permutation period T = 4) on the spliced 2C channels to promote cross-dimensional interaction between multispectral bands and spatial features; a progressive residual, which uses random depth dropout to improve the model's robustness to noise such as cloud occlusion; and an identity mapping, which short-circuits the original features to retain low-frequency information in continuous areas such as farmland and water.
[0076] To further implement the above technical solution, we introduce a position trigger to dynamically enhance the edge and small target response areas to address the common problems of edge misclassification and small target loss in remote sensing segmentation. The specific content is as follows:
[0077] A two-dimensional coordinate channel is added to the input feature map through the coordinate convolution layer to enhance spatial sensitivity; the GELU activation function introduces nonlinearity to improve expression ability; the conventional convolution layer extracts spatial context; a multi-head mechanism is adopted, the features extracted by convolution in the current position trigger are used as query (Query), the global semantic feature map of the previous stage or higher layer is used as key (Key) and value (Value), the attention weight is calculated by dot product, the response of the edge and small target area is adaptively amplified, and irrelevant background features are suppressed. Finally, the weighted semantic information is fed back to the current position feature to achieve cross-scale information interaction and enhancement; finally, the position trigger mask M is generated through convolution layer and Sigmoid architecture activation. pos ∈[0, 1] H′×W′×1 , used to perform pixel-level weighting on feature maps.
[0078] In order to further implement the above technical solution, the specific content of step S3 is:
[0079] S31. Obtain the fused feature map The input is sent to the adaptive global pooling layer for average pooling in the spatial dimension, preserving the channel dimension information and compressing the spatial redundancy to adapt to the scale changes of image features of different input sizes;
[0080] S32. The pooled feature vector is linearly mapped through a fully connected layer and then passed through an activation function (such as Sigmoid or Softmax) to obtain the final segmentation probability map P, that is, the probability value of each pixel belonging to each category;
[0081] S33. Compare the obtained segmentation probability map P with the real label map Y provided when inputting the neural network, and construct a joint loss function based on cross entropy loss and Dice loss To improve the model's sensitivity to small objects and boundary areas;
[0082] In this embodiment, the network output probability map is P={P c,i,j}, the real label image is Y={Y c,i,j}, Y c,i,j ∈{0, 1}, where c=1,...,K, K represents the number of categories, (i, j) represents the pixel coordinates, and the joint loss function is:
[0083]
[0084] Among them, λ1 and λ2 are the weight coefficients of the two sub-losses, is the pixel-level cross entropy loss, Dice coefficient loss is used to alleviate the class imbalance problem;
[0085] Combining pixel-level cross entropy loss and Dice loss:
[0086] Cross Entropy Loss:
[0087]
[0088] Dice architecture loss:
[0089]
[0090] Here, ε is a small constant to prevent division by zero.
[0091] A remote sensing image segmentation system based on convolution-state space fusion and position trigger, based on a remote sensing image segmentation method based on convolution-state space fusion and position trigger, including a feature preprocessing module, a patch embedding and spatial position encoding module, a feature extraction and fusion module and an output module;
[0092] Feature preprocessing module, used to preprocess the original remote sensing image;
[0093] The patch embedding and spatial position encoding module is used to convert the pre-processed remote sensing images into a structured representation of a unified scale through patch embedding and spatial position encoding, while retaining the spatial position information;
[0094] The feature extraction and fusion module is used to decouple the position-encoded patch vector sequence through multiple convolutional-state-space fusion units (C-SSFUs). The convolution branch extracts fine-grained edge and local structural features, while the state-space modeling branch captures long-range contextual information. The two branch features are fused at multiple scales using a cross-modal attention mechanism. Position triggers are introduced multiple times to dynamically enhance edge and small target response areas, and the final fused features are output.
[0095] The output module is used to map the fused features into the final pixel-level classification probability map through adaptive pooling and fully connected layers, and compare it with the true label map to construct a supervised loss function to improve the model's segmentation performance for minority categories and boundary areas, and obtain the final segmentation result.
[0096] This paper introduces an end-to-end feature preprocessing module at the input stage to perform geometric correction, atmospheric correction, denoising, and spectral normalization on remote sensing images, effectively improving the usability and consistency of the original images. Subsequently, through patch embedding and spatial position encoding mechanisms, large-scale images are converted into a structured representation of a unified scale while preserving their spatial position information, reducing the difficulty of subsequent model processing.
[0097] During the feature extraction and fusion phase, a novel convolutional-state-space fusion unit (C-SSFU) was constructed. This module utilizes a channel decoupling design, using convolution branches to extract fine-grained edge and local structural features, and a state-space modeling branch to capture long-range contextual information. The two branch features are then fused in a multi-spectral, cross-modal attention mechanism to achieve semantic enhancement of the ground objects. Furthermore, "position triggers" are introduced multiple times into the backbone structure, dynamically strengthening the edge and small target response areas through mechanisms such as coordinate convolution and multi-head cross-attention, thereby addressing the issues of blurred boundaries and small target loss.
[0098] In the output stage, adaptive pooling and fully connected layers are introduced to map the fused features into the final pixel-level classification probability map, and a cross-entropy and Dice loss function is used to improve the model's segmentation performance for minority categories and boundary areas.
[0099] The overall architecture of the present invention not only takes into account the global context modeling capability, but also enhances the perception of object boundaries and details, and is suitable for high-precision semantic segmentation tasks of multi-source remote sensing images.
[0100] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a remote sensing image segmentation method based on convolution-state space fusion and position triggers.
[0101] A processing terminal includes a memory and a processor. The memory stores a computer program that can be run on the processor. When the processor executes the computer program, a remote sensing image segmentation method based on convolution-state space fusion and position trigger is implemented.
[0102] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences in architecture from other embodiments. Similar or identical parts between the various embodiments can be referenced. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For relevant parts, refer to the method description.
[0103] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A remote sensing image segmentation method based on convolution-state space fusion and position trigger, characterized in that: The following steps are involved: S1. Preprocess the original remote sensing image and convert the large-scale image into a structured representation of a unified scale through patch embedding and spatial position encoding, while preserving the spatial position information; S2. Decouple the position-encoded patch vector sequence through multiple convolutional-state-space fusion units (C-SSFUs). The convolution branch extracts fine-grained edge and local structural features, while the state-space modeling branch captures long-range contextual information. The two branch features are then fused at multiple scales using a cross-modal attention mechanism. Position triggers are introduced multiple times to dynamically enhance edge and small object response regions, ultimately outputting the fused features. S3. Through adaptive pooling and fully connected layers, the fused features are mapped to the final pixel-level classification probability map, which is compared with the true label map. A supervised loss function is constructed to improve the model's segmentation performance for minority categories and boundary areas, and the final segmentation result is obtained.
2. The remote sensing image segmentation method based on convolution-state space fusion and position trigger according to claim 1, characterized in that: Preprocessing includes: geometric correction and registration, atmospheric correction and denoising, and spectral normalization or standardization.
3. The remote sensing image segmentation method based on convolution-state space fusion and position trigger according to claim 1, characterized in that: The specific contents of patch embedding and spatial position encoding are as follows: Divide the preprocessed image into non-overlapping patches; Flatten each small block into a one-dimensional vector, and then map it to the D-dimensional feature space through a linear mapping layer; Based on the spatial position information of each small block in the original image, a position code is added.
4. The remote sensing image segmentation method based on convolution-state space fusion and position trigger according to claim 1, characterized in that: The specific architecture of multiple convolution-state space fusion units C-SSFU and multiple introduction of position triggers is as follows: the patch vector sequence with position encoding is input into the two serial convolution-state space fusion units C-SSFU in the first stage in sequence, and the position trigger is introduced. After the first position trigger, two C-SSFUs, a position trigger, four C-SSFUs, a position trigger and two C-SSFUs are serially connected in sequence.
5. The remote sensing image segmentation method based on convolution-state space fusion and position trigger according to claim 1, characterized in that: The specific contents of the convolution-state space fusion unit C-SSFU are: S21. Input features The channel grouping strategy is used to decouple features and obtain the first feature and the second feature S22. Convolutional feature enhancement branch: First feature Perform deep feature extraction, and obtain local geometric priors through batch normalization layer, 3*3 convolution layer, batch normalization layer, ReLU activation function, 3*3 convolution layer, batch normalization layer, ReLU activation function, point-by-point convolution layer, and ReLU activation function S23. State-space modeling branch: Second feature Long-range context modeling is performed, and features are obtained through layer normalization, linear layer, 3*3 depth-separable convolution layer, SiLU activation function, two-dimensional selective scanning block, and layer normalization. At the same time, the second feature After a linear layer and activation function, we get and After matrix multiplication, layer normalization is performed to obtain global semantics S24. Through cross-modal attention fusion, dual-branch features are used for multi-spectral collaborative modeling: Will and Perform splicing to obtain splicing features Bimodal interaction: leveraging local geometric priors in convolutional branches Global semantics for guiding state-space branching After the linear layer, we get Q. After the linear layer, K is obtained; spectral sensitive attention: Q is multiplied by K and then activated by SiLU function. Multiply to get features S25. Residual feature refinement: Constructing feature enhancement pathways, features After channel shuffling and initial input Perform residual connection and output That is the output feature of the convolution-state space fusion unit C-SSFU.
6. The remote sensing image segmentation method based on convolution-state space fusion and position trigger according to claim 1, characterized in that: The specific content of introducing position triggers to dynamically enhance the edge and small target response areas is as follows: The coordinate convolution layer adds a two-dimensional coordinate channel to the input feature map to enhance spatial sensitivity; the GELU activation function introduces nonlinearity to improve expressiveness; the conventional convolution layer extracts spatial context; A multi-head mechanism is adopted, and the features extracted by convolution in the current position trigger are used as queries. The global semantic feature map of the previous stage or higher layer is used as the key and value. The attention weight is calculated by dot product, and the response of the edge and small target areas is adaptively amplified, while irrelevant background features are suppressed. Finally, the weighted semantic information is fed back to the current position feature to achieve cross-scale information interaction and enhancement. Finally, the position trigger mask is generated through the convolution layer and Sigmoid architecture activation, which is used to perform pixel-level weighting on the feature map.
7. The remote sensing image segmentation method based on convolution-state space fusion and position trigger according to claim 1, characterized in that: The specific content of step S3 is: S31. Input the obtained fusion feature map into the adaptive global pooling layer, perform average pooling in the spatial dimension, retain the channel dimension information, and compress spatial redundancy; S32. The pooled feature vector is linearly mapped through a fully connected layer and then activated by an activation function to obtain the final segmentation probability map, i.e., the probability value of each pixel belonging to each category. S33. Compare the obtained segmentation probability map with the real label map provided when inputting the neural network, and construct a joint loss function based on cross entropy loss and Dice loss 8. A remote sensing image segmentation system based on convolution-state space fusion and position trigger, characterized in that: A remote sensing image segmentation method based on convolution-state space fusion and position trigger according to any one of claims 1 to 7, comprising a feature preprocessing module, a patch embedding and spatial position encoding module, a feature extraction and fusion module, and an output module; Feature preprocessing module, used to preprocess the original remote sensing image; The patch embedding and spatial position encoding module is used to convert the pre-processed remote sensing images into a structured representation of a unified scale through patch embedding and spatial position encoding, while retaining the spatial position information; The feature extraction and fusion module is used to decouple the position-encoded patch vector sequence through multiple convolutional-state-space fusion units (C-SSFUs). The convolution branch extracts fine-grained edge and local structural features, while the state-space modeling branch captures long-range contextual information. The two branch features are fused at multiple scales using a cross-modal attention mechanism. Position triggers are introduced multiple times to dynamically enhance edge and small target response areas, and the final fused features are output. The output module is used to map the fused features into the final pixel-level classification probability map through adaptive pooling and fully connected layers, and compare it with the true label map to construct a supervised loss function to improve the model's segmentation performance for minority categories and boundary areas, and obtain the final segmentation result.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the remote sensing image segmentation method based on convolution-state space fusion and position trigger according to any one of claims 1 to 7 is implemented.
10. A processing terminal comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the remote sensing image segmentation method based on convolution-state space fusion and position trigger as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Double-branch remote sensing image semantic segmentation method and device based on visual converter and Mama
CN119360026A
Semantic detail fusion and context enhancement remote sensing image segmentation method based on DeepLabv3 +
CN119672340A
Cross-scale semantic segmentation method, system and device for cloud and cloud shadow
CN120070462A
Hyperspectral remote sensing image classification method based on self-attention context network
WO2022073452A1
Boundary-optimized remote sensing image semantic segmentation method and apparatus, and device and medium
WO2023077816A1
Cited By
Image fine structure intelligent detection algorithm based on double-branch encoder
CN121074418A
Remote sensing image segmentation method based on frequency domain global channel perception and cross-channel attention fusion
CN121305177A
Method and system for classifying few-sample hyperspectral remote sensing images based on multi-path evolution
CN121505452A