Mama-based edge refinement remote sensing image semantic change detection method
By constructing a Mamba-based semantic change detection method for edge refinement remote sensing images, using twin encoder and difference module to extract features, combined with deep edge supervision strategy, the problem of edge roughness in the existing method is solved, and high-precision change detection and semantic segmentation effects are achieved.
Patent Information
- Application Number
- CN202510643660.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing semantic change detection methods do not pay enough attention to the detailed optimization of boundary regions in the feature extraction and fusion links, resulting in rough edges of the predicted change graph, limiting its performance in high-precision boundary division application scenarios.
Mamba-based semantic change detection method for edge refinement remote sensing images, including twin encoder extracting multi-scale features, difference module extracting different features, edge refinement decoder enhances edge information, and combining the loss function strategy of deep edge supervision and changing area supervision to improve the learning ability of edge information.
It significantly improves the boundary accuracy and semantic segmentation effect of change detection, can accurately judge the change area, solves the problem of rough edges in the existing methods, and improves the performance of the model in high-precision boundary division application scenarios.
Smart Images

Figure CN120495843A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a Mamba-based edge thinning remote sensing image semantic change detection method, belonging to the technical field of remote sensing image change detection. Background Art
[0002] Semantic change detection, a key research area in remote sensing and computer vision, plays an irreplaceable role in practical applications such as urban planning, disaster assessment, environmental monitoring, and land use analysis. By analyzing bi-temporal or multi-temporal remote sensing images, semantic change detection can identify changes in land cover types and assign semantic labels to them, providing critical information for decision-making. Daudt et al. proposed various supervised learning methods based on fully convolutional neural networks for semantic change detection in remote sensing images. They also proposed a network architecture that can simultaneously perform change detection and land cover mapping, utilizing predicted land cover information to assist change detection. Ding et al. decoupled semantic segmentation and change detection from the two subtasks, achieving deep integration of the two subtasks through deep feature fusion. They also designed Siam-SR and Cot-SR blocks to enhance the semantic representation of each temporal branch and model the semantic correlation between features in the two temporal branches. Chen et al. applied the Mamba architecture to change detection in remote sensing images and designed three different architectures for the change detection tasks of binary change detection, semantic change detection, and building damage assessment.
[0003] Existing semantic change detection methods fail to adequately optimize the details of boundary regions during feature extraction and fusion, lacking dedicated edge refinement mechanisms. Some prioritize global semantics while ignoring local boundary accuracy, while others rely on simple fusion without addressing boundary discontinuities. This results in rough edges in the predicted change maps, limiting their performance in high-precision boundary delineation applications. Rough edges not only degrade the visualization quality of the changed regions but can also lead to misjudgments of the change extent, particularly in tasks requiring precise boundary information, such as urban expansion monitoring or post-disaster reconstruction. Summary of the Invention
[0004] In order to solve the problem of rough edges in existing methods, the present invention aims to provide a method for detecting semantic changes in remote sensing images using edge thinning based on Mamba, comprising the following steps:
[0005] S1, constructs a twin encoder based on the visual state space model to extract multi-scale features of dual-temporal remote sensing images;
[0006] S2, builds a Mamba-based difference module to extract the difference features between the two-phase features output by the twin encoder;
[0007] The Mamba-based difference module includes:
[0008] The feature maps F1 and F2 of different time phases extracted from the twin backbone composed of the visual state space model are subtracted and the absolute value is taken. The absolute value is passed into the visual state space model to obtain the feature map F s ; Feature maps F1 and F2 are obtained after the linear layer and the depth convolution module to obtain feature maps X time1 and X time2 ;
[0009] The linear layer is used to generate matrices B, C and Δ from the two feature maps, where the matrix C generated by the two feature maps is exchanged to achieve information interaction between features, which is expressed as:
[0010]
[0011] The output feature map Y time1 and Y time2 Perform layer normalization, pass it into the linear layer and then perform element-wise addition with the original input feature maps F1 and F2, and add them to the feature map F s Perform element-wise multiplication to obtain the feature map F fus1 and feature map F fus2 ; For the feature map F fus1 and feature map F fus2 Perform channel dimension splicing, use convolution with a kernel size of 1 to halve the channel dimension and perform batch normalization and ReLU function activation to obtain the output feature map F diff , expressed as:
[0012] F fus1 =(Linear(LN(Y time1 ))+F1)×F s
[0013] F fus2 =(Linear(LN(Y time2 ))+F2)×F s
[0014] F diff =ReLU(BN(Conv 1×1 (Concat[F fus1 ,F fus2 ])))
[0015] Among them, LN is layer normalization, Linear represents linear layer, Concat represents channel dimension concatenation, BN represents batch normalization, and ReLU represents activation function;
[0016] S3, inputs the difference features into the edge-refined visual state space decoder, combines it with the edge enhancer to enhance the edge information, and generates a refined semantic change map of edge information;
[0017] S4, adopts a deep edge supervision strategy to calculate multi-level edge loss through edge labels to optimize network training effects;
[0018] S5, fuses the deep edge loss, semantic segmentation loss and change detection loss;
[0019] S6, inputs the dual-temporal remote sensing image into the trained model and outputs the semantic change detection results.
[0020] Furthermore, the dataset is SECOND, which consists of 4662 pairs of images with a resolution of 512×512; these images are divided into training set, validation set, and test set; the training set is randomly flipped and cropped, and all these images are normalized before being input into the network.
[0021] Furthermore, the twin Mamba encoder is a hierarchical network using a visual state space, that is, a 2D selective scan is performed on the input feature map; specifically, a convolution operation with a kernel size of 4 and a stride of 4 is performed on the input image with a resolution of 512×512, and the resolution of the input image is changed to 128×128, and the number of channels is changed from 3 to 128; the scaled feature map is linearly changed to twice the dimension of the original size, and a depth-wise separable convolution with a kernel size of 3 is performed on it; the feature map is subjected to a 2D selective scan to extract features; the feature map is linearly changed twice to restore the original dimension; the previous operation is iterated N times to complete the feature extraction of the layer, N varies according to the different levels, the backbone network has a total of four levels, and N of each layer is set to 2, 2, 15, 2, and the features after the extraction of the layer are doubled by block merging to pass into the next level of the backbone network until the end.
[0022] Furthermore, the Mamba-based difference module has a total of four levels according to the dual-phase features mentioned in the above-mentioned twin Mamba encoder. Each level subtracts the dual-phase features, takes the absolute value, and passes it into the visual state space model; the feature map is passed into the linear layer and the deep convolution layer for processing; the linear layer is used to generate the matrix B, C and Δ of the original input feature map, so that the matrix C generated by the two feature maps is exchanged and a selective state space mechanism is performed to achieve information interaction between features; the output feature map is layer-normalized, passed into the linear layer, and element-wise added with the original input feature map; the feature map is multiplied with the feature map extracted from the visual state space model; the feature map is dimensionally spliced, and a convolution with a convolution kernel size of 1 is used to halve the channel dimension to obtain the output difference feature.
[0023] Furthermore, the edge-refined visual state space has an input feature that is a feature map after upsampling by the decoder through the upper layer. The feature map is passed into the visual state space model for initial decoding, and the feature map is layer-normalized and convolved; the feature map is subjected to expansion and corrosion operations, and the feature map including edge information is obtained by subtracting the corrosion from the expanded feature map; the feature map containing edge information is multiplied with the feature map after initial decoding to enhance the edge information of the feature map; the feature map is processed using two consecutive convolutions and trained in combination with edge labels; the feature map and the feature map with enhanced edge information are element-wise added, and the feature map is processed using channel attention; the feature map processed by channel attention is multiplied with the feature map for initial decoding, and element-wise added with the feature map with enhanced edge information; the feature map is input into the spatial attention module, and multiplied with the feature map processed by channel attention and the feature map after initial decoding; the feature map is added to the feature map after initial decoding to obtain the output feature map.
[0024] Furthermore, the strategy of boundary supervision training uses the Canny operator to extract edges from the data labels in the training set as edge labels; downsamples the edge labels to the scales corresponding to different levels of the decoder; combines binary cross entropy with Dice loss, and sets the balance factor to 0.5; and sets the weights of the edge loss function at different levels to 0.25, 0.25, 0.5, and 0.75;
[0025] Furthermore, in the strategy of integrating the deep edge loss function, semantic segmentation loss and change detection loss, the loss function of the change detection part is set to be a combination of the cross entropy function and the Dice loss; and the loss function of the semantic segmentation part is set to be a combination of the cross entropy and the Lovasz-softmax function.
[0026] Compared with the prior art, the present invention has at least the following beneficial effects:
[0027] The present invention proposes a Mamba-based edge-refined remote sensing image semantic change detection model for change detection in remote sensing images. The model uses a visual state space model to build a semantic change detection infrastructure, which can simultaneously capture multi-scale features from small to large scales. The present invention proposes a Mamba-based difference module for extracting difference features between the backbones composed of the visual state space. Through the Mamba-based difference module, the features extracted by the Mamba backbone encoder can be effectively extracted for difference features, and the auxiliary model can accurately determine the change area in different scenarios. The present invention proposes a decoder composed of an edge-refined visual state space. The decoder effectively integrates traditional image processing algorithms with deep learning methods to improve the edge refinement of the semantic change map, solving the problem that existing change detection models often focus more on overall accuracy and pay insufficient attention to details of the edge part. The present invention proposes a loss function that combines deep boundary supervision training with change area supervision. Through deep boundary supervision training, by strengthening the model's learning of target edge information, the network is prompted to focus on capturing subtle differences between the target and the background, which can effectively improve the model's ability to depict edge information. The effectiveness of the method of the present invention is verified on the remote sensing image change detection dataset SECOND. In order to verify the effectiveness of the method of the present invention, it is compared with the existing methods on multiple evaluation indicators, including accuracy, intersection over union ratio, separation kappa coefficient and F1 score. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The embodiments of the present invention will be described with reference to the following drawings, and the objectives, technical features and advantages of the present invention will become more apparent:
[0029] Figure 1 This is a flowchart of an implementation method for detecting semantic changes in remote sensing images using Mamba-based edge thinning.
[0030] Figure 2 This is an overall framework diagram of a Mamba-based edge thinning remote sensing image semantic change detection method of the present invention;
[0031] Figure 3 This is a schematic diagram of the difference module based on Mamba in the present invention;
[0032] Figure 4 Schematic diagram of the visual state space module for edge refinement of the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, advantages and technical methods of the present invention more clear, the following is described in conjunction with specific examples. The specific examples described here are only used to explain the present invention and are not used to limit the present invention.
[0034] like Figure 1FIG. 1 is a flowchart of a method for detecting semantic changes in remote sensing images using Mamba-based edge thinning according to an embodiment of the present invention, including steps S1 to S7.
[0035] S1, the twin Mamba encoder backbone network composed of dual-temporal data input visual state space, extracts features of the dual-temporal image in sequence through different layers of the backbone network to obtain multi-level features;
[0036] In this example, the model is trained and tested on a public dataset, the SECOND dataset. This dataset is described below:
[0037] SECOND is a semantically annotated change detection dataset. It collects 4,662 aerial image pairs from multiple platforms and various sensors to ensure dataset diversity. These image pairs are distributed across cities such as Hangzhou, Chengdu, and Shanghai. Each image has been annotated at the pixel level. The annotation task for the SECOND dataset was performed by a team of experts, ensuring high accuracy of the annotations. Regarding the categories of change in the SECOND dataset, the focus is on six major land cover types: non-vegetated surface, trees, low vegetation, water, buildings, and playgrounds. These types are often found in natural or human-induced geographic changes. In the SECOND dataset, non-vegetated surface primarily includes impervious surface and bare ground. These six land cover categories in the SECOND dataset generate 30 common change categories (including those where no change occurred). By randomly selecting image pairs, the SECOND dataset presents the actual distribution of land cover categories during changes.
[0038] like Figure 2 As shown, this is the overall framework diagram of the Mamba-based edge refinement remote sensing image semantic change detection method according to an embodiment of the present invention. In this example of the present invention, the Mamba-based twin encoder backbone network is divided into four parts, namely, feature extraction at different levels.
[0039] The Twin Mamba encoder is a hierarchical network that uses the visual state space, that is, it performs a 2D selective scan on the input feature map; more specifically, it performs a convolution operation with a kernel size of 4 and a stride of 4 on the input image with a resolution of 512×512, changing the resolution of the input image to 128×128 and the number of channels from 3 to 128; the reduced-scale feature map is linearly transformed to double its dimension, and a depth-wise separable convolution with a kernel size of 3 is performed on it; the feature map is subjected to a 2D selective scan to further refine the features; the feature map is linearly transformed twice to restore the original dimension; the previous operation is iterated N times to complete the feature extraction of the layer, N varies according to the different levels, the backbone network has a total of four levels, and N of each layer is set to 2, 2, 15, and 2. After the features are extracted at the layer, the number of channels is doubled by block merging and then passed to the next layer of the backbone network until the end.
[0040] In S3, the dual-phase features extracted by the backbone network at different levels and branches are input into the Mamba-based difference module to obtain multi-level difference features.
[0041] like Figure 3 As shown, a Mamba-based difference module structure diagram of an embodiment of the present invention is provided. Step S3 specifically includes steps S31-S33:
[0042] S31, feature maps F1 and F2 are respectively extracted from different phases of the twin backbone composed of the visual state space model. The feature maps F1 and F2 are subtracted and the absolute value is taken. After being passed into the visual state space model, the feature map F is obtained. s ; Feature maps F1 and F2 are respectively passed through the linear layer and the depth convolution module to obtain the feature map X time1 and X time2 ;
[0043] S32, using a linear layer to generate matrices B, C and Δ from the two feature maps, where the matrix C generated by the two feature maps is exchanged to achieve information interaction between features, expressed as:
[0044]
[0045] S33, output feature map Y time1 and Y time2 Perform layer normalization, pass it into the linear layer and perform element-wise addition with the original input feature map F1 and feature map F2, and add it to the feature map F s Perform element-wise multiplication to obtain the feature map F fus1 and F fus2 ; For the feature map F fus1 and F fus2Perform channel dimension splicing, use convolution with a convolution kernel size of 1 to halve the channel dimension of the spliced feature map, perform batch normalization and ReLU function activation on it to obtain the output feature map F diff , expressed as:
[0046] F fus1 =(Linear(LN(Y time1 ))+F1)×F s
[0047] F fus2 =(Linear(LN(Y time2 ))+F2)×F s
[0048] F diff =ReLU(BN(Conv 1×1 (Concat[F fus1 ,F fus2 ])))
[0049] Among them, LN represents layer normalization, Linear represents linear layer, Concat represents channel dimension concatenation, BN represents batch normalization, and ReLU represents activation function.
[0050] S4, the difference features extracted by the Mamba-based difference module are passed into the edge-refined visual state space for layer-by-layer decoding:
[0051] like Figure 4 As shown, the present invention provides an edge-refined visual state space structure diagram, and the specific steps of step S4 include S41-S46:
[0052] S41, the feature map F m Use the state space module for feature extraction, perform layer normalization and convolution to obtain the feature map F v , expressed as:
[0053] F v =Conv(LN(VSS(F in )))
[0054] S42, output feature map F v Dilate and erode operations are performed respectively; the edge information is extracted by subtracting the eroded feature map from the dilated feature map, which is expressed as:
[0055] F e =Dilate(F v )-Erode(F v )
[0056] S43, the features after edge information extraction are convolved and combined with the feature map F v Perform element-wise multiplication to strengthen the feature map F v The edge information effect is expressed as:
[0057] F m =Conv(F e )×F v
[0058] S44, the feature map F e Use two consecutive convolutions as the output of the edge map and combine them with edge labels for supervised training; use the feature map F after two convolutions e With the feature map F m Perform element-level addition; input the processed feature map into the channel attention module CA to obtain the output feature map F CA , expressed as:
[0059] F CA =CA(Conv(Conv(F e ))+F m )
[0060] S45, the feature map F CA With the feature map F v Perform element-wise multiplication with the feature map F m Add and input into the spatial attention module SA to obtain the feature map F SA , expressed as:
[0061] F SA =SA(F CA ×F v +F m )
[0062] S46, the feature map F SA With the feature map F CA And the feature map F v Multiply, and feature map F v After adding, we get the output feature map F out , expressed as:
[0063] F out =F v +F SA ×F CA ×F v .
[0064] S5, calculate the loss of edge maps at different levels in the edge refinement decoder, and use deep edge supervision for training. The specific steps of step S5 include S51-S52:
[0065] S51, using the Canny operator to extract edges from the change image labels in the training data set, and using them as edge labels for training, to guide the model decoder to focus on feature learning of the edge parts during training;
[0066] S52 combines binary cross entropy and Dice loss for edge supervision; different weights are set for edge loss at different levels for deep edge supervision training, expressed as:
[0067]
[0068] L edge =αL BCE +(1-α)L Dice
[0069]
[0070] Among them, label i Represents the i-th pixel value on the label, p i Represents the i-th pixel value predicted by the model, ε is a small number to prevent division by zero error, α represents the balance factor to balance BCE loss and Dice loss, set α = 0.5; for deep supervision, the weight λ between the edge loss functions of different layers i , set λ i =[0.25,0.25,0.5,0.75].
[0071] S6, multiplying the decoded prediction map with the semantic segmentation map to perform pixel-level remote sensing image semantic change detection and generate a semantic change prediction map after the detection is completed, iteratively training and saving the model parameters of the best results.
[0072] S7, input the dual-temporal remote sensing images of the test set into the remote sensing image semantic change detection model to obtain the prediction of the changed ground objects.
[0073] In this embodiment, semantic segmentation and change detection are combined. For the change detection part, this embodiment adopts a method of combining the cross entropy function with the Dice loss. The loss function of this part is expressed as:
[0074] L cd =L ce +L dice
[0075] The loss function of the semantic segmentation part is a combination of cross entropy and Lovasz-softmax function. The loss function of this part is expressed as:
[0076]
[0077] Among them, T1 and T2 represent remote sensing images at different imaging times, and the overall loss function is expressed as:
[0078]
[0079] To verify the effectiveness of our proposed method, we compared it with current mainstream change detection methods. The SECOND dataset was used, and the evaluation metrics were precision, intersection over union (IoU), separation kappa coefficient, and F1 score. The comparison results are shown in Table 1.
[0080] Table 1 Comparison of the proposed method with mainstream change detection methods
[0081]
[0082] From Table 1, it can be seen that the method proposed in the present invention has a significant improvement in all indicators compared with the current mainstream change detection methods, which verifies the effectiveness of the method proposed in the present invention.
[0083] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various modifications or variations within the scope of the claims without affecting the essence of the present invention. The above preferred features may be used in any combination as long as they do not conflict with each other.
Claims
1. A semantic change detection method for remote sensing images based on edge thinning in Mamba, characterized by: The following steps are involved: S1, constructs a twin encoder based on the visual state space model to extract multi-scale features of dual-temporal remote sensing images; S2, builds a Mamba-based difference module to extract the difference features between the two-phase features output by the twin encoder; The Mamba-based difference module includes: The feature maps F1 and F2 of different time phases extracted from the twin backbone composed of the visual state space model are subtracted and the absolute value is taken, and the absolute value is passed into the visual state space model to obtain the feature map F s ; Feature maps F1 and F2 are respectively passed through the linear layer and the depth convolution module to obtain the feature map X time1 and X time2 ; The linear layer is used to generate matrices B, C and Δ from the two feature maps, where the matrix C generated by the two feature maps is exchanged to achieve information interaction between features, which is expressed as: The output feature map Y time1 and Y time2 Perform layer normalization, pass it into the linear layer and then perform element-wise addition with the original input feature maps F1 and F2, and add them to the feature map F s Perform element-wise multiplication to obtain the feature map F fus1 and feature map F fus2 ; For the feature map F fus1 and feature map F fus2 Perform channel dimension splicing, use convolution with a kernel size of 1 to halve the channel dimension and perform batch normalization and ReLU function activation to obtain the output feature map F diff , expressed as: F fus1 =(Linear(LN(Y time1 ))+F1)×F s F fus2 =(Linear(LN(Y time2 ))+F2)×F s F diff =ReLU(BN(Conv 1×1 (Concat[F fus1 ,F fus2 ]))) Among them, LN is layer normalization, Linear represents linear layer, Concat represents channel dimension concatenation, BN represents batch normalization, and ReLU represents activation function; S3, inputs the difference features into the edge-refined visual state space decoder, combines it with the edge enhancer to enhance the edge information, and generates a refined semantic change map of edge information; S4, adopts a deep edge supervision strategy to calculate multi-level edge loss through edge labels to optimize network training effects; S5, fuses the deep edge loss, semantic segmentation loss and change detection loss; S6, inputs the dual-temporal remote sensing image into the trained model and outputs the semantic change detection results.
2. The Mamba-based edge thinning remote sensing image semantic change detection method according to claim 1, characterized in that: The twin encoder based on the visual state space model includes: Perform 2D selective scanning on the input feature map: Specifically, the input image with a resolution of 512×512 is convolved with a kernel size of 4 and a stride of 4. The resolution of the input image is changed to 128×128, and the number of channels is increased from 3 to 128. The feature map after scale reduction is linearly transformed to double the original dimension, and a depthwise separable convolution with a kernel size of 3 is performed on it. The feature map is subjected to 2D selective scanning to refine the features. The feature map after 2D selection scanning is linearly transformed twice to restore the dimension; the feature extraction is completed by iterating N times, where N varies according to the different levels. The backbone network has four levels, and N of each layer is set to 2, 2, 15, and 2. After the features extracted at this level are doubled in number of channels by block merging, they are passed to the next level of the backbone network until the end.
3. The Mamba-based edge thinning remote sensing image semantic change detection method according to claim 1, characterized in that: The edge-refined visual state space decoder comprises: The feature map F m Use the state space module for processing, perform layer normalization and convolution to obtain the feature map F v , expressed as: F v =Conv(LN(VSS(F in ))) For the output feature map F v Use edge extractor to get feature F that enhances edge effect e ; The features after extracting edge information are convolved with the feature map F v Perform element-wise multiplication to achieve enhanced feature map F v The edge information effect is expressed as: F m =Conv(F e )×F v For the feature map F e Do convolution to form the output result of edge supervision, and combine the output result with F m Perform fusion and calculate the channel dimension weight of the fusion feature to obtain the feature map F CA ; The feature map F CA With the feature map F v Perform element-wise multiplication with the feature map F m Add, input into the spatial attention module SA, output feature map F SA , expressed as: F SA =SA(F CA ×F v +F m ) The feature map F SA With the feature map F CA And the feature map F v Multiply, and feature map F v After adding, we get the output feature map F out , expressed as: F out =F v +F SA ×F CA ×F v 。 4. The method for detecting semantic changes in remote sensing images using Mamba-based edge thinning according to claim 1, wherein: The deep edge supervision strategy includes: The Canny operator is used to extract the edges of the change image labels in the training dataset and used as edge labels for training. This guides the decoder of the model to focus on the feature learning of the edge parts during training. Binary cross entropy and Dice loss are combined for edge supervision. Different weights are set for edge losses at different levels for deep edge supervision training, which can be expressed as: L edge =αL BCE +(1-α)L Dice Among them, label i Represents the i-th pixel value on the label, p i Represents the i-th pixel value predicted by the model, ε is a small number used to prevent division by zero errors, α represents the balance factor to balance the BCE loss and Dice loss, and α is set to 0.5; for the weight λ between the edge loss functions of different levels in deep supervision i , set λ i =[0.25,0.25,0.5,0.75].
5. The method for detecting semantic changes in remote sensing images using Mamba-based edge thinning according to claim 1, wherein: The deep edge loss function, semantic segmentation loss and change detection loss fusion include: The change detection part is obtained by combining the cross entropy function with the Dice loss. The loss function of this part is expressed as: L cd =L ce +L dice The loss function of the semantic segmentation part is obtained by combining the cross entropy and the Lovasz-softmax function. The loss function of this part is expressed as: Among them, T1 and T2 represent remote sensing images at different imaging times, and the overall loss function is defined as follows:
6. The Mamba-based edge thinning remote sensing image semantic change detection method according to claim 3, characterized in that: The edge extractor comprises: For the output feature map F v Dilate and Erode operations are performed respectively; the edge features are obtained by subtracting the eroded feature map from the dilated feature map, which is expressed as: F e =Dilate(F v )-Erode(F v )。 7. The method for detecting semantic changes in remote sensing images using Mamba-based edge thinning according to claim 3, wherein: The channel dimension weights of the calculated fusion features are obtained by the feature map F CA ,include: The feature map F e Use two consecutive convolutions to form the output edge feature map, and combine it with the edge label for supervised training; use the feature map F after two convolutions e With the feature map F m Perform element-level addition; input the processed feature map into the channel attention module CA to obtain the output feature map F CA , expressed as: F CA =CA(Conv(Conv(F e ))+F m )。
Citation Information
Patent Citations
Remote sensing image change detection method using multi-level difference feature adaptive fusion
CN117576567A
Change detection method based on semantic prior and fusion space positioning
CN119478542A
Remote sensing image semantic change detection method and device based on Mamba model
CN119580258A
Transitive and commutative multimodal models and uses
WO2024223621A1
Cited By
Multi-task unsupervised change detection method fusing image domain alignment and segmentation network
CN121725373A
Multi-task unsupervised change detection method fusing graph domain alignment and segmentation network
CN121725373B
Intelligent interpretation method for double-time-phase remote sensing image
CN121982453A