Double-flow remote sensing image change detection method fused with Mmba enhancement
Through lightweight MobileNet backbone network and Mamba sequence modeling, combined with semantic segmentation aggregation module and channel interaction module, the accuracy and efficiency of remote sensing image change detection are improved, and the problems of high computational complexity and insufficient feature expression in the existing methods are solved.
Patent Information
- Application Number
- CN202510787795.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
When using complex scenarios, existing remote sensing image change detection methods are difficult to accurately capture local changes, and the calculation complexity is high. The traditional attention mechanism limits the efficiency of the model, resulting in limited detection accuracy and efficiency.
The lightweight MobileNet backbone network is adopted, combined with Mamba sequence modeling and semantic segmentation aggregation module, and the feature enhancement module FRM and channel interaction module CIM are used to improve global feature modeling capabilities and enhance feature detection in changing areas.
It improves the accuracy and robustness of change detection, reduces the calculation amount, improves the model operation efficiency, highlights the characteristics of the change area, and suppresses background interference.
Smart Images

Figure CN120298906A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and particularly to a dual-stream remote sensing image change detection method enhanced by Mamba fusion. Background Art
[0002] With the rapid development of remote sensing technology, change detection plays an important role in fields such as urban planning and disaster monitoring. Traditional change detection methods mainly rely on manual feature extraction and threshold judgment, with low efficiency and limited accuracy. In recent years, deep learning has made remarkable progress in the field of change detection. However, existing methods mostly use traditional attention mechanisms for feature enhancement, with high computational complexity and difficulty in capturing long-range dependencies. At the same time, the temporal feature modeling of dual-temporal images is insufficient, resulting in limited recognition accuracy for change regions, low change detection accuracy in edge regions, and difficulty in accurately locating the boundaries of change regions.
[0003] To solve the above problems, researchers have proposed various improvement schemes. For example, using lightweight backbone networks to reduce computational complexity and using multi-scale feature fusion to improve detection accuracy. However, these methods still fail to well solve the problems of feature expression and temporal modeling, and still face challenges in accuracy and efficiency in practical applications. Existing change detection models show limitations when dealing with complex scenes, especially when dealing with high-resolution remote sensing images with rich details, it is difficult to accurately capture local change features. In addition, the computational overhead of traditional attention mechanisms is large, which limits the efficiency of the model in practical applications. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a dual-stream remote sensing image change detection method enhanced by Mamba fusion with simple algorithm and high detection accuracy.
[0005] The technical solution of the present invention to solve the above technical problems is: a dual-stream remote sensing image change detection method enhanced by Mamba fusion, comprising the following steps:
[0006] S1, Image acquisition: Using a remote sensing image acquisition device, obtain remote sensing images of the same area at different times, that is, dual-temporal remote sensing images, to obtain a remote sensing image data set;
[0007] S2, Data preprocessing: Preprocess the obtained dual-temporal remote sensing images and divide them into a training set and a test set;
[0008] S3. Model Establishment: Design a dual-stream change detection network model RFAMNet enhanced by Mamba fusion. RFAMNet uses lightweight MobileNet as the backbone network, constructs a dual-stream feature preprocessor LightWeightEncoder and a feature enhancement module FRM, and then generates a change detection map through a semantic segmentation aggregation module SSAM, a channel interaction module CIM, and a global decoder Decoder. The feature enhancement module FRM includes Mamba sequence modeling to enhance global feature modeling and edge perception processing capabilities. The specific process of model establishment is as follows:
[0009] S31: Construct a dual-stream feature preprocessor LightWeight Encoder, input the dual-temporal remote sensing images into the lightweight MobileNet, and extract multi-scale feature representations;
[0010] S32: Construct a feature enhancement module FRM, integrating Mamba sequence modeling to dynamically capture long-range dependencies and edge perception;
[0011] S33: Establish a semantic segmentation aggregation module SSAM to integrate multi-scale differential features and mine global semantic information;
[0012] S34: Establish a channel interaction module CIM to enhance the interaction between features of different sizes and highlight the features of the changed area;
[0013] S35: Construct a global decoder Decoder to restore the spatial resolution and generate a change detection map;
[0014] S4. Model Training: Use the training set data to train the constructed dual-stream change detection network model until the entire model converges, and save the optimal model;
[0015] S5. Model Validation and Inference: Input the test set data into the trained optimal model, predict the changed areas in the test set, calculate the intersection over union between the predicted changed areas and the true changed areas , and perform statistics according to all categories of the changed areas to obtain the average intersection over union , and evaluate the prediction accuracy of change detection.
[0016] For the above dual-stream remote sensing image change detection method enhanced by Mamba fusion, in step S1, the obtained remote sensing image dataset S is:
[0017]
[0018] Among them, represents the image data of the area in the period, represents Region Image data of a certain period.
[0019] For the above-mentioned dual-stream remote sensing image change detection method fused with Mamba enhancement, the specific process of step S2 is as follows:
[0020] S21: Image annotation;
[0021] Annotate the target categories of the changed regions in the remote sensing image dataset to obtain the annotated remote sensing image dataset as:
[0022]
[0023] Among them, represents the remote sensing image , The change region category label map for comparison;
[0024] S22: Image segmentation;
[0025] For Use the sliding window method to segment the image into an image and label dataset of size 1024×1024:
[0026]
[0027]
[0028] Among them, represents The th 1024×1024 sample image segmented by the sliding window method, represents The th 1024×1024 sample image segmented by the sliding window method, is the comparison change region label map corresponding to , ;
[0029] S23: Dataset division: Randomly divide the sample dataset into a training set, a validation set, and a test set according to a ratio of 8:1:1;
[0030] S24: Data preprocessing: Normalize and perform data augmentation on the original dual-temporal remote sensing images.
[0031] For the above-mentioned dual-stream remote sensing image change detection method fused with Mamba enhancement, in step S22, the specific operation of the sliding window method is as follows:
[0032] First, divide the remote sensing image and the like into multiple unit grids of 1024×1024 pixels. If the remote sensing image cannot be divided proportionally, padding is performed, and the step size of each unit grid is denoted as ; then sample the labeled area through a sliding window while retaining the relative position information of the window center; the obtained sliding window sampled image is :
[0033]
[0034] where 、 represent the abscissa and ordinate of the center position of the sliding window, and represents the sliding step size;
[0035] For the edge area of the image, the mirror padding method is adopted to ensure that the edge area is completely sampled, and the size of the padded image remains 1024×1024.
[0036] For the above-mentioned dual-stream remote sensing image change detection method enhanced by fused Mamba, the specific process of step S31 is as follows:
[0037] The image of 、 The image of and their corresponding label data After passing through the dual-stream feature preprocessor Lightweight Encoder, preprocessed features are extracted, and then the preprocessed features extract multi-scale features through the lightweight MobileNet to obtain the global feature map where is The global feature map obtained from the image of is The global feature map obtained from the image of
[0038] The dual-stream feature preprocessor Lightweight Encoder performs feature extraction through the lightweight MobileNet. This network adopts the Weights-Sharing mechanism and consists of a Conv3×3 convolution operation, BatchNorm normalization, and ReLU activation function; MobileNet extracts multi-scale features through depthwise separable convolution DSConv and outputs multi-scale feature representations , , are features of 5 different scales, where ;
[0039] The above-mentioned dual-stream remote sensing image change detection method enhanced by fusing Mamba, the specific process of step S32 is as follows:
[0040] Input the preprocessed multi-scale features into the Feature Enhancement Module (FRM). Taking a specific hierarchical feature as the boundary, the features lower than this specific hierarchical feature and are first downsampled to the target resolution through max pooling, then the channels are aligned through Conv1×1 convolution, and local features are extracted through DSConv3×3 depthwise separable convolution; then the features higher than this specific hierarchical feature and are processed through Conv1×1 convolution to align channels, DSConv3×3 depthwise separable convolution, BatchNorm normalization, and ReLU activation function, and bilinearly upsampled to the target resolution; then the features of each level are respectively input into the MambaBlock module for sequence modeling enhancement; for serialization processing, the two-dimensional feature map is rearranged into a one-dimensional sequence , , , is in the real number field, and
[0041]
[0042] where , , , are learnable parameters, is the input signal, is the hidden state, and k represents the position in the feature sequence;
[0043] Then the output sequence is inverse mapped to a two-dimensional feature map , , ; by normalizing the enhanced features and connecting them with the input residuals, the enhanced feature of the th level is output. The enhanced features of each level are concatenated along the channel dimension, and are successively passed through Conv1×1 convolution and DSConv3×3 depthwise separable convolution to further fuse cross-scale feature information. The fused features are connected with the original features of the target level by residuals, and finally the enhanced th level feature is output. , .
[0044] For the above-mentioned dual-stream remote sensing image change detection method enhanced by fused Mamba, the specific process of step S33 is as follows:
[0045] Calculate the difference of the dual-temporal multi-scale features enhanced by the feature enhancement module FRM. Through the formula to obtain the -level difference feature , where represents the operation of taking the absolute value after element-wise subtraction, represents the -level feature in the period, the -level feature in the period. Take the 3rd-level difference feature as the low-level difference feature and the 4th-level difference feature as the high-level difference feature
[0046] Input the 4th-level difference feature and the 5th-level difference feature into the semantic segmentation aggregation module SSAM. First, perform bilinear upsampling on to make have the same resolution as , and then directly add them element-wise to obtain the coarse global semantic feature; then expand the number of channels through Conv3×3 convolution operation and evenly divide it into four sub-features on the channel dimension. Each sub-feature performs multi-scale feature learning through dilated convolution with different dilation rates; finally, splice the four sub-features after multi-scale feature learning, and perform feature fusion through Conv1×1 convolution and DSConv3×3 depthwise separable convolution to obtain the global semantic feature .
[0047] For the above-mentioned dual-stream remote sensing image change detection method enhanced by fused Mamba, the specific process of step S34 is as follows:
[0048] Input the obtained low-level difference feature , high-level difference feature and global semantic feature into the channel interaction module CIM for processing; first perform global average pooling GAP and Conv1×1 convolution operation on , , respectively to obtain the corresponding vectors for calculating attention weights , and ; Then, for , and , perform dimensional reshaping so that is respectively matrix-multiplied with , to obtain two similarity matrices, add the two similarity matrices and generate attention weights through the sigmoid function ; Reshape and perform matrix multiplication with , then process it through a Conv1×1 convolution, batch BatchNorm normalization, and ReLU activation function, and establish a residual connection with to obtain enhanced difference features .
[0049] For the above-mentioned method for change detection of fused Mamba-enhanced dual-stream remote sensing images, the specific process of step S35 is as follows:
[0050] Input , and into the decoding unit DU for processing; First, respectively pass and through a Conv1×1 convolution, and then perform bilinear upsampling so that and have the same resolution as . Add , and element-wise and input it into a DSConv3×3 depthwise separable convolution, and finally output the prediction result of the current layer after passing through a Conv1×1 convolution.
[0051] For the above-mentioned method for change detection of fused Mamba-enhanced dual-stream remote sensing images, the specific process of step S4 is as follows:
[0052] S41: Given the dual-temporal remote sensing images of the training set and their corresponding ground truth label data;
[0053] S42: Input the paired remote sensing images into RFAMNet, and after being processed by RFAMNe, output the prediction result of change detection ;
[0054] S43: Adopt the binary cross-entropy loss function and the dice loss function to construct the total loss function as:
[0055]
[0056] Among them represents the layer of the decoder, corresponding to to the predicted result after decoding the differential features; and respectively represent the binary cross-entropy loss function and the dice loss function of the decoder layer;
[0057] The formula of the binary cross-entropy loss function is:
[0058]
[0059] Among them and are the width and height of the input remote sensing image respectively, represents the true label of the pixel at coordinate and represents the predicted value of the pixel at coordinate ;
[0060] The formula of the dice loss function is:
[0061]
[0062] Among them represents the size of the intersection of the true label and the predicted result , represents the size of , represents the size of ;
[0063] Minimize the total loss function through the backpropagation algorithm, and perform iterative optimization training on the model parameters until the value converges and the model reaches the optimal;
[0064] During the iterative optimization training process, use the validation set to verify the training accuracy of the model in real time and save the model weights with the highest accuracy.
[0065] The beneficial effects of the present invention are as follows: The method proposed by the present invention improves the accuracy compared with other algorithms. Specifically, 1) a lightweight MobileNet is used as the backbone network for feature extraction, reducing the model parameters and computational amount, and improving the running efficiency of the model while ensuring the detection accuracy; 2) the MambaBlock and the semantic segmentation aggregation module SSAM are introduced to enhance the global feature modeling ability under linear complexity. By fusing feature information at different levels and multi-scale feature learning, the detection accuracy and robustness of the changed area are improved; 3) the channel interaction module CIM highlights the features of the changed area by mining the correlation of different-level features in the channel dimension, suppresses background interference, and further improves the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is the overall flowchart of the present invention.
[0067] Figure 2 It is the structural block diagram of the dual-stream change detection network model RFAMNet of the present invention.
[0068] Figure 3 It is the structural schematic diagram of the feature enhancement module FRM of the present invention.
[0069] Figure 4 is area Schematic diagrams of 5 remote sensing images in the period.
[0070] Figure 5 is area Schematic diagrams of 5 remote sensing images in the period.
[0071] Figure 6 It is the schematic diagram of the predicted image of the comparative changed area obtained by using the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0072] The present invention will be further described below with reference to the drawings and embodiments.
[0073] As Figure 1 shown, a dual-stream remote sensing image change detection method enhanced by Mamba fusion includes the following steps:
[0074] S1, image acquisition: Use a remote sensing image acquisition device to obtain remote sensing images of the same area in different periods, that is, dual-temporal remote sensing images, to obtain a remote sensing image dataset.
[0075] The obtained remote sensing image dataset S is:
[0076]
[0077] Among them, Indicate Area The image data of the period, Indicate Area The image data of the period.
[0078] S2, Data preprocessing: Preprocess the acquired multi-temporal remote sensing images and divide them into a training set and a test set.
[0079] The specific process of the step S2 is as follows:
[0080] S21: Image annotation;
[0081] Annotate the target categories of the changed areas in the remote sensing image dataset to obtain the annotated remote sensing image dataset It is:
[0082]
[0083] Among them, Indicates the remote sensing image , The change area category label map for comparison;
[0084] S22: Image segmentation;
[0085] For Use the sliding window method to segment the image into an image and label dataset of size 1024×1024:
[0086]
[0087]
[0088] Among them, Indicates The th 1024×1024 sample image segmented by the sliding window method, Indicates The th 1024×1024 sample image segmented by the sliding window method, Is the , Corresponding comparison change area label map;
[0089] The specific operation of the sliding window method is as follows:
[0090] First, divide the remote sensing image into multiple unit grids of 1024×1024 pixels in equal proportion. If the remote sensing image cannot be divided in equal proportion, padding is performed. The step size of each unit grid is denoted as ; Then, sample the labeled regions through a sliding window while retaining the relative position information of the window center; the sampled images obtained by the sliding window are :
[0091]
[0092] where 、 represent the abscissa and ordinate of the center position of the sliding window, represents the sliding step size;
[0093] For the edge regions of the image, use the mirror filling method to ensure that the edge regions are fully sampled, and the size of the filled image remains 1024×1024;
[0094] To improve the generalization ability and robustness of the model, perform data augmentation on the segmented samples. The data augmentation methods include: horizontal flipping and vertical flipping, randomly cropping the image while maintaining the original resolution, adjusting the brightness, contrast, and saturation of the image, etc. For each pair of dual-temporal remote sensing images, perform the same data augmentation operations to ensure that the change information is not affected by the augmentation operations;
[0095] S23: Dataset division: Randomly divide the sample dataset into a training set, a validation set, and a test set according to the ratio of 8:1:1;
[0096] S24: Data preprocessing: Perform normalization and data augmentation on the original dual-temporal remote sensing images.
[0097] S3. Model establishment: Design a dual-stream change detection network model RFAMNet that integrates Mamba enhancement. RFAMNet uses the lightweight MobileNet as the backbone network, constructs a dual-stream feature preprocessor LightWeightEncoder and a feature enhancement module FRM, and then generates a change detection map through a semantic segmentation aggregation module SSAM, a channel interaction module CIM, and a global decoder Decoder. The feature enhancement module FRM includes Mamba sequence modeling, enhancing the global feature modeling and edge perception processing capabilities. The specific process of model establishment is as follows:
[0098] S31: Construct a dual-stream feature preprocessor LightWeight Encoder, input the dual-temporal remote sensing images into the lightweight MobileNet, and extract multi-scale feature representations.
[0099] The specific process of step S31 is:
[0100] image of the 、 image of the and their corresponding tag data The preprocessed features are extracted through the two-stream feature preprocessor Lightweight Encoder , and then the preprocessed features extract multi-scale features through the lightweight MobileNet to obtain the global feature map , where is the global feature map obtained from the image in the is the global feature map obtained from the image in the
[0101] The two-stream feature preprocessor Lightweight Encoder performs feature extraction through the lightweight MobileNet. This network adopts the Weights-Sharing mechanism and consists of a Conv3×3 convolution operation, BatchNorm normalization, and ReLU activation function; MobileNet extracts multi-scale features through depthwise separable convolution DSConv and outputs multi-scale feature representations , , are features of 5 different scales, where .
[0102] S32: Construct a feature enhancement module FRM, integrate Mamba sequence modeling, and dynamically capture long-range dependencies and edge awareness
[0103] The specific process of step S32 is as follows
[0104] Input the preprocessed multi-scale features into FRM. Taking a specific hierarchical feature as the boundary, features lower than this specific hierarchical feature and are first downsampled to the target resolution through max pooling, and then the channels are aligned through a Conv1×1 convolution and local features are extracted through a DSConv3×3 depthwise separable convolution; then features higher than this specific hierarchical feature and are processed through a Conv1×1 convolution to align channels, a DSConv3×3 depthwise separable convolution, BatchNorm normalization, and ReLU activation function, and bilinearly upsampled to the target resolution; then the features of each level are respectively input into the MambaBlock module for sequence modeling enhancement; for serialization processing, the two-dimensional feature map is rearranged into a one-dimensional sequence , , , is in the real number domain All are dimensions; then state space modeling is performed, and the long-range dependencies are captured by parameterized state space model SSM, the formula is:
[0105]
[0106] in , , , is a learnable parameter, is the input signal, is the hidden state, k represents the position in the feature sequence;
[0107] Then output the sequence Inverse mapping to a two-dimensional feature map , , ; By normalizing the enhanced features and connecting them with the input residual, the output Level Enhancement Features , the enhanced features of each level The concatenation is done by channel dimension, and the cross-scale feature information is further fused through Conv1×1 convolution and DSConv3×3 deep separable convolution in sequence. The fused features are residually connected with the original features of the target level, and the enhanced features are finally output. The period Level Features , , .
[0108] S33: Establish a semantic segmentation aggregation module SSAM to integrate multi-scale difference features and mine global semantic information.
[0109] The specific process of step S33 is:
[0110] The difference calculation is performed on the dual-phase multi-scale features enhanced by the feature enhancement module FRM, using the formula Get the first Level difference characteristics ,in It represents the operation of taking the absolute value after element-by-element subtraction. express The period Level features, express The period Level 3 difference features As a low-level differential feature , Level 4 difference features As a high-level differentiating feature ;
[0111] Input the 4th - level differential features and the 5th - level differential features into the semantic segmentation aggregation module SSAM. First, perform bilinear upsampling to make the resolution the same as that, and then directly add them element - by - element to obtain the rough global semantic features; then expand the number of channels through a Conv3×3 convolution operation, and evenly divide them into four sub - features in the channel dimension. Each sub - feature respectively performs multi - scale feature learning through dilated convolutions with different dilation rates; finally, splice the four sub - features after multi - scale feature learning, and perform feature fusion through Conv1×1 convolution and DSConv3×3 depth - separable convolution to obtain the global semantic features .
[0112] S34: Establish a channel interaction module CIM to enhance the interaction between features of different sizes and highlight the features in the changing regions.
[0113] The specific process of step S34 is as follows:
[0114] Input the obtained low - level differential features , high - level differential features and global semantic features into the channel interaction module CIM for processing; first perform global average pooling GAP and Conv1×1 convolution operations on , , respectively to obtain the corresponding vectors , and for calculating attention weights; then reshape , and so that is respectively matrix - multiplied with , to obtain two similarity matrices, add the two similarity matrices and generate attention weights through the sigmoid function; reshape and perform matrix - multiplication with , then process it through Conv1×1 convolution, batch BatchNorm normalization, and ReLU activation function, and establish a residual connection with to obtain the enhanced differential features .
[0115] S35: Construct a global decoder Decoder to restore the spatial resolution and generate a change detection map.
[0116] The specific process of step S35 is as follows:
[0117] Input , and into the decoding unit DU for processing; first, and are respectively subjected to Conv1×1 convolution, and then bilinearly upsampled to make and have the same resolution as . Then, add , and element-wise and input the result into a DSConv3×3 depthwise separable convolution, and finally output the prediction result of the current layer after Conv1×1 convolution.
[0118] S4. Model training: Use the training set data to train the constructed two-stream change detection network model until the entire model converges, and save the optimal model.
[0119] The specific process of the said step S4 is as follows:
[0120] S41: Given the bi-temporal remote sensing images of the training set and their corresponding ground truth label data;
[0121] S42: Input the paired remote sensing images into the RFAMNet, and after the processing of the RFAMNet, output the prediction result of change detection ;
[0122] S43: Use the binary cross-entropy loss function and the dice loss function to construct the total loss function as:
[0123]
[0124] where represents the layer of the decoder, corresponding to the prediction result after decoding the difference features from to ; , respectively represent the binary cross-entropy loss function and the dice loss function of the layer of the decoder;
[0125] The formula of the binary cross-entropy loss function is:
[0126]
[0127] where and are the width and height of the input remote sensing image respectively, represents the coordinate of the true label of the pixel at, represents the coordinate of the predicted value of the pixel at;
[0128] Dice loss function formula is:
[0129]
[0130] where represents the true label and the prediction result of the intersection size, represents the size of, represents the size of;
[0131] Minimize the total loss function through the backpropagation algorithm, iteratively optimize and train the model parameters until the
[0132] value converges and the model reaches the optimal;
[0133] In the iterative optimization training process, use the validation set to verify the model training accuracy in real time and save the model weights with the highest accuracy. S5, model validation and inference: Input the test set data into the trained optimal model, predict the changed areas in the test set, calculate the intersection over union between the predicted changed areas and the true changed areas,
[0134]
[0135] where represents the number of classes of change detection, represents the predicted changed area, represents the changed area of the true label.
[0136] Example
[0137] Step S1: Adopt the WHU dataset, which contains two aerial images with a resolution of 0.2 m / pixel, 15354×32507. The dataset contains 22000 independent buildings, that is, a single category. Thus, the operation of image acquisition is completed.
[0138] Step S2: Based on the remotely sensed images of the same area in two different periods annotated in the WHU dataset , the sliding window method is used to segment the image, and 32×15 = 480 images with a size of 1024×1024 are obtained.
[0139]
[0140]
[0141]
[0142] According to the ratio of 8:1:1, the training set Train, the validation set Val, and the test set Test are divided, with 384, 48, and 48 images respectively, and are denoted as:
[0143]
[0144]
[0145]
[0146] So far, the operation of dataset preprocessing is completed.
[0147] Step S3, design a dual-stream remote sensing image change detection model RFAMNet that integrates Mamba enhancement, as Figure 2 shown, the specific steps are as follows:
[0148] S31: Construct a dual-stream feature preprocessor LightWeight Encoder, input the dual-temporal remote sensing images into the lightweight MobileNet, and extract multi-scale feature representations.
[0149] Images of the 、 Images of the and their corresponding label data Pass through the dual-stream feature preprocessor Lightweight Encoder to extract preprocessed features , and then the preprocessed features Extract multi-scale features through the lightweight MobileNet to obtain the global feature map , where is The global feature map obtained from the images of the is The global feature map obtained from the images of the
[0150] The dual-stream feature preprocessor, the Lightweight Encoder, extracts features through the lightweight MobileNet. This network adopts the Weights-Sharing mechanism and consists of Conv3×3 convolution operations, BatchNorm normalization, and ReLU activation functions. MobileNet extracts multi-scale features through depthwise separable convolutions (DSConv) and outputs multi-scale feature representations. , , where ; After Conv convolution and downsampling operations, where the dimension becomes 512×512×48, the dimension becomes 256×256×96, the dimension becomes 128×128×192, the dimension becomes 64×64×384, the dimension becomes 32×32×768.
[0151] S32: Construct the Feature Reinforcement Module (FRM), as Figure 3 shown, integrating Mamba sequence modeling to dynamically capture long-range dependencies and edge awareness.
[0152] Input the preprocessed multi-scale features into the Feature Reinforcement Module (FRM). Bounded by a specific hierarchical feature , features lower than this specific hierarchical feature and are first downsampled to the target resolution through max pooling and then aligned in channels through Conv1×1 convolution and local features are extracted through DSConv3×3 depthwise separable convolution; then features higher than this specific hierarchical feature and are processed through Conv1×1 convolution for channel alignment, DSConv3×3 depthwise separable convolution, BatchNorm normalization, and ReLU activation functions, and bilinearly upsampled to the target resolution, with the dimension becoming 128×128×192.
[0153] Then, the features of each level are respectively input into the MambaBlock module for sequence modeling enhancement; for serialization processing, the two-dimensional feature map is rearranged into a one-dimensional sequence , , , is in the real number field, are all dimensions; then state space modeling is performed to capture long-range dependency relationships through the parameterized state space model (SSM). The formula is:
[0154]
[0155] wherein 、 、 、 are learnable parameters, is the input signal, is the hidden state, and k represents the position in the feature sequence;
[0156] Then, the output sequence is inversely mapped to a two-dimensional feature map , , ; by normalizing the enhanced features and connecting them with the input residuals, the enhanced feature at the th level is output. The enhanced features at each level are concatenated along the channel dimension, and are successively passed through a Conv1×1 convolution and a DSConv3×3 depthwise separable convolution to further fuse cross-scale feature information. The fused features are connected with the original features of the target level by residuals, and finally the enhanced -th period -th level feature , , is output.
[0157] The difference between the dual-temporal multi-scale features enhanced by the feature enhancement module FRM is calculated. The -th level difference feature is obtained through the formula ,where represents the operation of taking the absolute value after element-wise subtraction, represents the -th level feature in the -th period, represents the -th level feature in the -th period. The 3rd level difference feature is used as the low-level difference feature ,and the 4th level difference feature is used as the high-level difference feature ; ;
[0158] The 4th level difference feature and the 5th level difference feature are input into the semantic segmentation aggregation module SSAM. First, is bilinearly upsampled so that the resolution of is the same as that of , and then the two are directly added element-wise to obtain the coarse global semantic feature; then, the number of channels is expanded through a Conv3×3 convolution operation and evenly divided into four sub-features along the channel dimension Each sub - feature performs multi - scale feature learning through dilated convolutions with different dilation rates respectively; finally, the four sub - features after multi - scale feature learning are concatenated, and feature fusion is performed through 1×1 convolution and depth - separable 3×3 convolution (DSConv) to obtain global semantic features.
[0159] S34: Establish a channel interaction module CIM to enhance the interaction between features of different sizes and highlight the features in the changing regions.
[0160] Input the obtained low - level difference features , high - level difference features and global semantic features into the channel interaction module CIM for processing; first, perform global average pooling (GAP) and 1×1 convolution operations on , , respectively to obtain corresponding vectors , and for calculating attention weights; then reshape , and so that performs matrix multiplication with , respectively to obtain two similarity matrices, add the two similarity matrices and generate attention weights through the sigmoid function; reshape and perform matrix multiplication with , then process it through 1×1 convolution, batch normalization (BatchNorm), and ReLU activation function, and establish a residual connection with to obtain enhanced difference features
[0161] S35: Construct a global decoder Decoder to restore the spatial resolution and generate a change detection map.
[0162] Input , and into the decoding unit DU for processing; first, pass and through 1×1 convolution respectively, and then perform bilinear upsampling to make and have the same resolution as , and combine , and After element-wise addition, the result is input into a depthwise separable convolution of DSConv3×3, and finally, the prediction result of the current layer is output through a Conv1×1 convolution.
[0163] Step S4, model training. The specific process is as follows:
[0164] Train the established dual-stream change detection network model, given a training set of dual-temporal remote sensing images and their corresponding ground truth label data;
[0165] Input the paired remote sensing images into RFAMNet. After being processed by RFAMNet, the prediction result of change detection is output ;
[0166] Adopt the binary cross-entropy loss function and the dice loss function to construct the total loss function as:
[0167]
[0168] where represents the layer of the decoder, corresponding to the prediction result after decoding the difference features from to ; , respectively represent the binary cross-entropy loss function and the dice loss function of the decoder layer;
[0169] The formula for the binary cross-entropy loss function is:
[0170]
[0171] where and are the width and height of the input remote sensing image respectively, represents the ground truth label of the pixel at coordinate , represents the predicted value of the pixel at coordinate ;
[0172] The formula for the dice loss function is:
[0173]
[0174] where represents the size of the intersection of the ground truth label and the prediction result , represents the size of , represent the size of;
[0175] Minimize the total loss function through the backpropagation algorithm , perform iterative optimization training on the model parameters, and train until the value converges and the model reaches the optimal;
[0176] During the iterative optimization training process, use the validation set to verify the model training accuracy in real time and save the model weights with the highest accuracy.
[0177] Step S5, model inference test, the specific process is: Input the paired image and label data in the test set into the trained model, predict the changed area in the test set, and calculate the intersection over union between the predicted changed area and the true changed area, and perform statistics according to all categories of the changed area to obtain the mean intersection over union to evaluate the prediction accuracy of change detection.
[0178]
[0179] where represents the number of categories of change detection, represents the predicted changed area, represents the changed area of the true label, The higher the value of, the higher the model accuracy.
[0180] As Figures 4 - 6 shown, input the comparison images of the same area in different periods, and after the trained RFAMNet, the predicted image of the comparison changed area is output, thus completing the operation of model test inference.
Claims
1. A dual-stream remote sensing image change detection method enhanced by fusing Mamba, characterized in that, It includes the following steps: S1, Image acquisition: Use a remote sensing image acquisition device to obtain remote sensing images of the same area at different times, that is, dual-temporal remote sensing images, to obtain a remote sensing image dataset; S2, Data preprocessing: Preprocess the obtained dual-temporal remote sensing images and divide them into a training set and a test set; S3, Model establishment: Design a dual-stream change detection network model RFAMNet that integrates Mamba enhancement. RFAMNet uses the lightweight MobileNet as the backbone network, constructs a dual-stream feature preprocessor LightWeight Encoder and a feature enhancement module FRM, and then generates a change detection map through a semantic segmentation aggregation module SSAM, a channel interaction module CIM, and a global decoder Decoder. Among them, the feature enhancement module FRM includes Mamba sequence modeling to enhance the global feature modeling and edge perception processing capabilities. The specific process of model establishment is as follows: S31: Construct a dual-stream feature preprocessor LightWeight Encoder, input the dual-temporal remote sensing images into the lightweight MobileNet, and extract multi-scale feature representations; S32: Construct a feature enhancement module FRM, integrate Mamba sequence modeling, dynamically capture long-range dependencies, and edge perception; S33: Establish a semantic segmentation aggregation module SSAM to integrate multi-scale differential features and mine global semantic information; S34: Establish a channel interaction module CIM to enhance the interaction between features of different sizes and highlight the features of the change area; S35: Construct a global decoder Decoder to restore the spatial resolution and generate a change detection map; S4, Model training: Use the training set data to train the constructed dual-stream change detection network model until the entire model converges, and save the optimal model; S5, Model Validation Inference: Input the test set data into the trained optimal model to predict the changed regions in the test set, calculate the intersection over union (IoU) between the predicted changed regions and the true changed regions, and conduct statistics based on all categories of the changed regions to obtain the average intersection over union (mIoU). Evaluate the prediction accuracy of change detection.
2. The dual-stream remote sensing image change detection method enhanced by fusing Mamba according to claim 1, wherein, In the step S1, the obtained remote sensing image dataset S is: ; Among them, represents the image data of the region during period and represents the image data of the region during period 3. The dual-stream remote sensing image change detection method enhanced by the fusion of Mamba according to claim 2, wherein The specific process of the step S2 is: S21: Image annotation; Target category annotation of the changed area of the remote sensing image dataset to obtain the annotated remote sensing image dataset It is as follows: ; Among them, represents the remote sensing image , and the category label map of the changed area for comparison; S22: Image segmentation; Pairwise The sliding window method is used for image segmentation, and the image and label datasets are segmented into images of size 1024×1024: ; ; Among them, denotes the th 1024×1024 sample image cut by the sliding window method, and is and the corresponding contrast change region label map; S23: Dataset division: Randomly divide the sample dataset into a training set, a validation set, and a test set according to the ratio of 8:1:1; S24: Data preprocessing: Perform normalization and data augmentation on the original dual-temporal remote sensing images.
4. The method for change detection of fused Mamba-enhanced dual-stream remote sensing images according to claim 3, characterized in that, In the step S22, the specific operation of the sliding window method is as follows: First, divide remote sensing images and the like into multiple unit grids of 1024×1024 pixels. If the remote sensing image cannot be divided proportionally, padding is performed. The step size of each unit grid is denoted as ; then sample the annotated area through a sliding window while retaining the relative position information of the window center; the obtained sliding window sampled image is : ; Among them 、 represent the abscissa and ordinate of the center position of the sliding window, represents the sliding step size; For the image edge area, use the mirror filling method to ensure that the edge area is completely sampled, and the size of the filled image remains 1024×1024.
5. The method for fusing Mamba-enhanced dual-stream remote sensing image change detection according to claim 4, characterized in that, The specific process of the step S31 is: Images of the period Images of the period and the corresponding label data of both After passing through the two-stream feature preprocessor Lightweight Encoder, preprocessing features are extracted , and then the preprocessing features extract multi-scale features through the lightweight MobileNet to obtain a global feature map , where is the global feature map obtained from the images of the period, and is the global feature map obtained from the images of the The dual-stream feature preprocessor, Lightweight Encoder, extracts features through a lightweight MobileNet. This network adopts a weights-sharing mechanism and consists of Conv3×3 convolution operations, BatchNorm normalization, and ReLU activation functions. MobileNet extracts multi-scale features through depthwise separable convolutions (DSConv) and outputs multi-scale feature representations , , are features of 5 different scales, where .
6. The fused Mamba-enhanced dual-stream remote sensing image change detection method according to claim 5, wherein The specific process of the step S32 is: Input the preprocessed multi-scale features into the Feature Enhancement Module (FRM), and use the feature at a specific level as the boundary. For the features lower than this specific level feature and first downsample them to the target resolution through max pooling, then align the channels through Conv1×1 convolution, and extract local features through DSConv3×3 depthwise separable convolution; Then, the features higher than the features at this specific level and are processed through Conv1×1 convolution to align channels, DSConv3×3 depthwise separable convolution, BatchNorm normalization, and ReLU activation function, and bilinearly upsampled to the target resolution; Then, the features at each level are respectively input into the MambaBlock module for sequence modeling enhancement; for serialization processing, the two-dimensional feature map is rearranged into a one-dimensional sequence , , , is the real number field, are all dimensions; then, state space modeling is performed to capture long-range dependencies through the parameterized state space model SSM. The formula is as follows: ; Among them , , , are learnable parameters, is the input signal, is the hidden state, and k represents the position in the feature sequence; Then, the output sequence is inversely mapped to a two-dimensional feature map , , ; By normalizing the enhanced features and connecting them with the input residuals, the enhanced feature of the th level is output. The enhanced features of each level are concatenated along the channel dimension, and then passed through a Conv1×1 convolution and a DSConv3×3 depthwise separable convolution in sequence to further fuse the cross-scale feature information. The fused features are connected with the original features of the target level by residuals, and finally the enhanced -th period -th level features are output , .
7. The method for detecting changes in dual-stream remote sensing images enhanced by integrating Mamba according to claim 6, wherein, The specific process of the step S33 is: Calculate the difference of the dual-temporal multi-scale features enhanced by the Feature Enhancement Module (FRM). Through the formula obtain the -level difference features , where represents the operation of taking the absolute value after element-wise subtraction, represents the -level features in the period, represents the -level features in the period. Take the 3rd-level difference feature as the low-level difference feature ; Input the 4th-level differential features and the 5th-level differential features into the semantic segmentation aggregation module SSAM. First, perform bilinear upsampling to make the resolution the same as , and then directly add the two element-wise to obtain the coarse global semantic features; then expand the number of channels through Conv3×3 convolution operation, and evenly divide them into four sub-features on the channel dimension. Each sub-feature performs multi-scale feature learning through dilated convolutions with different dilation rates; finally, the four sub-features after multi-scale feature learning are concatenated, and feature fusion is performed through Conv1×1 convolution and DSConv3×3 depthwise separable convolution to obtain the global semantic feature .
8. The method for change detection of fused Mamba-enhanced dual-stream remote sensing images according to claim 7, wherein, The specific process of the step S34 is: The obtained low-level differential features , high-level differential features and global semantic features are input into the channel interaction module CIM for processing; first, perform global average pooling GAP and Conv1×1 convolution operations on , , respectively to obtain corresponding vectors , and for calculating attention weights; then reshape , and so that performs matrix multiplication with , respectively to obtain two similarity matrices, add the two similarity matrices and generate attention weights through the sigmoid function; reshape and perform matrix multiplication with , then process through Conv1×1 convolution, batch BatchNorm normalization and ReLU activation function, and establish a residual connection with to obtain enhanced differential features .
9. The fusion Mamba-enhanced dual-stream remote sensing image change detection method according to claim 8, wherein The specific process of the step S35 is: Input , and into the decoding unit DU for processing; first, perform Conv1×1 convolution on and respectively, and then perform bilinear upsampling to make the resolutions of and the same as that of . Add , and element-wise and then input the result into the DSConv3×3 depthwise separable convolution. Finally, perform Conv1×1 convolution to output the prediction result of the current layer. 10. The method for dual-stream remote sensing image change detection enhanced by integrated Mamba according to claim 9, wherein, The specific process of the step S4 is: S41: Given the dual-temporal remote sensing images of the training set and their corresponding true label data; S42: Input paired remote sensing images into RFAMNet, and after the processing of RFAMNe, output the prediction results of change detection ; S43: Adopt the binary cross-entropy loss function and the dice loss function to construct the total loss function as follows: ; Among them represents the level of the decoder, corresponding to to the predicted results after decoding the differential features; and respectively represent the binary cross-entropy loss function and the dice loss function of the decoder level; Binary cross-entropy loss function formula It is as follows: ; where and are the width and height of the input remote sensing image respectively, represents the true label of the pixel at coordinate , represents the predicted value of the pixel at coordinate . Dice loss function formula is as follows: ; Among them represents the true label and the prediction result of the intersection size, represents the size of, represents the size of; Minimize the total loss function through the backpropagation algorithm , perform iterative optimization training on the model parameters, and train until the value converges and the model reaches the optimal state; During the iterative optimization training process, use the validation set to verify the model training accuracy in real time and save the model weights with the highest accuracy.
Citation Information
Patent Citations
Remote sensing image change detection system and method based on depth separable convolution module
CN116229283A
Remote sensing image change detection method combining U-shaped network and self-attention mechanism
CN116740527A
SAM-fused fine-grained high-resolution remote sensing image change detection method
CN117911879A
Remote sensing small target detection method based on multi-dimensional feature aggregation enhancement and distribution mechanism
CN118397465A
Attention mechanism and Mama combined farmland non-agrochemical detection method
CN119169478A
Cited By
Cross-color space information missing distinguishing method and system for structural damage detection
CN120747038A
Remote sensing image change detection method and device, electronic equipment and storage medium
CN120823515A
Coastal wetland mud flat ecological environment monitoring method, system, equipment and medium
CN120853041A
Cross-domain image change detection method based on style randomization and similarity difference
CN120976770A
Cross-domain image change detection method based on style randomization and similarity difference
CN120976770B