A remote sensing image change detection method based on a Mamba-CNN double-branch network model

CN122454417BActive Publication Date: 2026-09-11NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610837981.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-11
Estimated Expiration
2046-06-11

AI Technical Summary

Technical Problem

解决现有模型对图像高频特征相对不敏感,细粒度边缘、纹理及小尺度目标特征易丢失,导致小目标漏检率高、变化边界模糊的问题;在保证检测精度达到SOTA水平的同时,维持较低的模型计算复杂度与参数规模,从而提升模型在实际工程场景中的部署效率与应用价值

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454417B_ABST
    Figure CN122454417B_ABST
Patent Text Reader

Abstract

The application discloses a remote sensing image change detection method based on a Mamba-CNN double-branch network model, which comprises the following steps: collecting double-phase remote sensing image data of a target area at different shooting time sequences; inputting the data into a convolutional encoder to extract shallow texture detail features, middle-layer transition features and deep-layer global semantic features, outputting multi-scale coding features corresponding to the double phases, and then inputting a space-time reorganization modeling module to reconstruct three double-phase space-time dependence relationships of sequence, cross and parallel through a space-time feature reorganization mode, and generating multi-scale joint difference features after splicing; inputting the multi-scale joint difference features into a decoder for multi-scale feature fusion to generate a remote sensing image prediction binary graph; and training the network model in combination with a loss function. The application can guarantee that the detection precision reaches the SOTA level, maintain a low model calculation complexity and parameter size, and thus improve the deployment efficiency and application value of the model in actual engineering scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing image detection, and in particular relates to a method for detecting changes in remote sensing images based on a Mamba-CNN dual-branch network model. Background Technology

[0002] Remote sensing image change detection aims to identify and extract surface change information from dual-temporal remote sensing images. It is a key technology in the field of remote sensing and has been widely applied in tasks such as urban planning, disaster assessment, and land resource management. With the continuous development of high-resolution remote sensing technology, change patterns in complex urban scenes are becoming increasingly diverse, with problems such as the coexistence of multi-scale changing targets, dense distribution of small targets, and blurred change boundaries becoming more prominent. These factors not only increase the difficulty of identifying changed areas but also place higher demands on the detection accuracy and computational efficiency of models.

[0003] In existing change detection methods, traditional methods typically rely on manually designed features, heavily depending on prior experience. This not only incurs high time and labor costs but also makes them susceptible to factors such as illumination changes, noise interference, and spurious changes in complex scenes, resulting in relatively limited robustness. CNN-based methods, leveraging their convolutional structures, possess strong feature extraction capabilities within local receptive fields, effectively capturing details such as texture and edges. However, due to limitations in their local modeling mechanisms, they are insufficient for representing long-distance spatiotemporal dependencies. Transformer-based methods achieve global context modeling through self-attention mechanisms, demonstrating strong feature representation capabilities in complex scene change detection tasks. However, these methods are usually accompanied by high computational complexity and memory overhead, posing challenges when processing large-scale, high-resolution remote sensing images.

[0004] In recent years, methods based on the Mamba state-space model have gradually become a research hotspot in the field of remote sensing change detection due to their advantages such as lower linear time complexity, strong ability to model long-distance dependencies, and high inference efficiency. These methods alleviate, to some extent, the computational overhead of Transformers as the input size increases, providing a new approach for efficient processing of high-resolution remote sensing images. However, Mamba tends to focus more on low-frequency global semantic information during feature modeling, and its ability to perceive high-frequency detailed features such as edges, textures, and small-scale targets is relatively insufficient. Therefore, in complex scenes, it is prone to problems such as missed detection of small targets and unclear change boundaries. Summary of the Invention

[0005] Purpose of the Invention: The purpose of this invention is to provide a remote sensing image change detection method based on a Mamba-CNN dual-branch network model. It addresses the problems of existing models being relatively insensitive to high-frequency image features, easily losing fine-grained edge, texture, and small-scale target features, leading to high false negative rates for small targets and blurred change boundaries. While ensuring state-of-the-art (SOTA) detection accuracy, it maintains low model computational complexity and parameter size, thereby improving the model's deployment efficiency and application value in practical engineering scenarios.

[0006] Technical solution: The present invention provides a remote sensing image change detection method based on a Mamba-CNN dual-branch network model, comprising the following steps:

[0007] Step 1: Collect dual-temporal remote sensing image data of the target area at different shooting times, and preprocess the data;

[0008] Step 2: Input the preprocessed data into the Mamba convolutional encoder to extract shallow texture detail features, mid-level transition features and deep global semantic features, and output multi-scale encoded features corresponding to both time periods.

[0009] Step 3: Input the multi-scale encoded features into the spatiotemporal reconstruction modeling module, reconstruct the three types of bi-temporal spatiotemporal dependencies of sequential, cross, and parallel through spatiotemporal feature reconstruction, and then concatenate the three types of bi-temporal spatiotemporal dependencies to generate multi-scale joint difference features;

[0010] Step 4: Input the multi-scale joint difference features into the decoder to perform multi-scale feature fusion and generate a predicted binary map of the remote sensing image.

[0011] Step 5: Train the Mamba-CNN dual-branch network model by combining the cross-entropy (CE) loss function and the Dice loss function, and iteratively optimize the model through backpropagation.

[0012] Furthermore, in step 1, the preprocessing includes cropping the dual-temporal remote sensing image data into a uniform pixel standard size and normalizing noise reduction.

[0013] Furthermore, in step 2, the Mamba convolutional encoder includes four levels of stages, with channel dimensions of {64, 128, 384, 512} for each level and stacking number of modules of each level of {4, 4, 12, 6}. Each level of stage works in collaboration with a local module and a progressive fusion module PFHR that combines high-frequency residuals.

[0014] Furthermore, the local module employs a lightweight local feature extraction unit, constructed using deep convolution combined with the GELU activation function, ensuring that the dimensions of the input and output features remain consistent. This is used to lightweight capture local edge and high-frequency texture details in dual-temporal remote sensing image data.

[0015] Furthermore, the structure of the progressive fusion module PFHR, which incorporates high-frequency residuals, includes a convolution branch and an SS2D scanning branch.

[0016] The input feature channels are proportionally divided using a channel splitting coefficient α. Channels with an α-ratio enter the SS2D scanning branch, while those with a 1-α ratio enter the convolution branch. A progressive fusion strategy is used to dynamically adjust the coefficient α, with α values ​​set sequentially to {0.25, 0.5, 0.5, 0.75} across four stages, allowing the model to dynamically balance the contributions of global and local features at different levels. In the SS2D scanning branch, the feature first undergoes a depthwise convolution, then high-frequency residuals are used for detail enhancement, and finally, the feature enters the SS2D module to capture global contextual information. In the convolution branch, the residual terms... and features entering the convolution branch spliced ​​together along the channel dimension It undergoes a depthwise convolution to enhance local detail features;

[0017] Finally, the outputs of the SS2D scanning branch and the convolution branch are concatenated to output the corresponding multi-scale encoded features in both time periods.

[0018] Furthermore, the detailed enhancement using high-frequency residuals specifically involves: processing the input features... Features are obtained by performing average pooling. Then, its spatial resolution is restored through nearest neighbor upsampling, and compared with the input features. The residual term is obtained by subtracting pixels. This can be expressed as the following formula:

[0019]

[0020] Where Pool represents average pooling and Upsample represents upsampling.

[0021] Furthermore, in step 3, the three types of dual-temporal spatiotemporal dependencies—sequential, intersecting, and parallel—are reconstructed through spatiotemporal feature reorganization. Specifically, the sequential relationship involves unfolding and arranging the features of the two temporal phases in chronological order, i.e., scanning the features of the first temporal phase one by one in sequence, and then scanning the features of the second temporal phase one by one in the same order, thereby capturing the sequential dependency relationship across temporal phases. The intersecting relationship strengthens the interaction between spatiotemporal features by alternating the arrangement of the two temporal phases, enhancing the association learning between the same regions of the two temporal phases. The parallel relationship involves splicing the features of the two temporal phases in the channel dimension to perform joint spatiotemporal modeling, simultaneously expressing spatial and temporal features, and realizing the collaborative fusion of dual-temporal information.

[0022] Furthermore, step 4 specifically involves the following steps: the difference features from different stages are fed into the fusion module in the decoder. First, the upsampled features are passed through a smoothing layer consisting of a convolutional layer and an activation function to map the number of channels of the low-level features to the same dimension as the high-level features. Then, the high and low-level features are added and fused element by element. Next, the output is obtained through a residual smoothing layer and passed to the next stage for progressive decoding. The spatial resolution is gradually restored, and finally, a binary image of the remote sensing image is generated.

[0023] Furthermore, step 5 specifically involves calculating the cross-entropy (CE) loss function and the Dice loss function using the following formulas:

[0024]

[0025]

[0026] in, It is the first The actual value of each pixel. Indicates the first The probability of a pixel; This represents the number of pixels; the overall loss function is expressed as:

[0027]

[0028] in, and This represents the weighting coefficients in the overall loss function.

[0029] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method of the present invention.

[0030] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0031] This invention solves the problem that existing models are relatively insensitive to high-frequency features of images, and are prone to losing fine-grained edge, texture and small-scale target features, resulting in high false negative rates for small targets and blurred boundary changes.

[0032] This invention can better balance the collaborative modeling of global semantic information and local detailed features, and exhibits stronger change representation capabilities in complex multi-scale scenes.

[0033] This invention addresses the shortcomings of existing simple Mamba-CNN fusion methods, which lack a hierarchical collaborative dynamic adjustment mechanism and cannot effectively preserve and utilize multi-scale features at each level.

[0034] This invention combines the feature differences and correlations of dual-temporal images in change detection tasks, and performs joint modeling and collaborative analysis of dual-temporal features to make the feature interaction between different temporal phases more complete and effectively enhance the ability to distinguish differences in features.

[0035] This invention ensures that the detection accuracy reaches the state-of-the-art level while maintaining low model computational complexity and parameter scale, thereby improving the deployment efficiency and application value of the model in actual engineering scenarios, and ultimately achieving high-precision, high-efficiency, and high-robust detection of multi-scale changes in complex urban scenarios. Attached Figure Description

[0036] Figure 1 This is the overall network framework diagram of the present invention;

[0037] Figure 2 The following is a structural diagram of the encoder module of the present invention, wherein (a) is an architecture diagram of the Mamba convolution module, (b) is a structural diagram of the local module, and (c) is a structural diagram of the PFHR.

[0038] Figure 3 The diagram shows the structure of the differential feature module of the present invention, wherein (a) is the structure diagram of spatiotemporal reconstruction modeling; and (b) is the structure diagram of VSSBlcok.

[0039] Figure 4 This is a structural diagram of the fusion module of the present invention;

[0040] Figure 5 The images show the visualization results of different methods on the LEVIR-CD dataset, where (a) is Example 1 of the dataset; and (b) is Example 2 of the dataset.

[0041] Figure 6 These are visualization results of different methods on the LEVIR-CD+ dataset of this invention, where (a) is Example 1 in the dataset; and (b) is Example 2 in the dataset.

[0042] Figure 7 These are visualization results of different methods on the WHU-CD dataset of this invention, where (a) is Example 1 in the dataset; and (b) is Example 2 in the dataset.

[0043] Figure 8 This is a graph comparing the model complexity of different methods in terms of computational cost, number of parameters, and IoU on the LEVIR-CD dataset.

[0044] Figure 9 The following are visualization diagrams of the features of the present invention, wherein (a) is the diagram of Example 1; and (b) is the diagram of Example 2. Detailed Implementation

[0045] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0046] This invention provides a hierarchical collaborative Mamba-CNN dual-branch remote sensing image change detection network, the model architecture of which is as follows: Figure 1 As shown. The model mainly consists of three parts: a Mamba convolutional encoder, a spatiotemporal reconstruction modeling, and a multi-scale fusion decoder. The specific process is as follows: 1) The encoder consists of four stages: In each stage, the input features are first downsampled, and the Mamba convolutional module extracts multi-scale features from the dual-temporal remote sensing images before and after the change, and outputs the results. and 2) Subsequently, multi-scale features and The data is transmitted to a spatiotemporal reconstruction model for joint spatiotemporal feature analysis and modeling, enabling full interaction between features before and after the change, and generating multi-scale difference features. 3) Finally, the differences in characteristics The code then proceeds to the multi-scale fusion decoder. The decoder and encoder are structurally symmetrical, each consisting of four levels. Specifically, the decoder's stages are inversely aligned to their corresponding encoder stages, meaning shallower layers align with deeper layers, progressing step-by-step. The fusion module in the decoder will... and The fusion process is performed, and the fusion result is obtained by upsampling. The features serve as input values ​​for the next stage; the spatial resolution is gradually restored in the decoder, and finally a classification head is used to generate an accurate change detection binary map.

[0047] The encoder consists of four levels, each level consisting of, for example, ... Figure 2 The model consists of Mamba convolutional modules as shown in (a), with channel dimensions of {64, 128, 384, 512} at each level and stacking numbers of modules at each level of {4, 4, 12, 6}. Except for the third stage, the first, second, and fourth stages all adopt a structure of first stacking several local modules and then introducing a Progressive Fusion Module with High-Frequency Residuals (PFHR) at the end of the stage; the third stage adds an additional Progressive Fusion Module with High-Frequency Residuals (PFHR) in the middle to strengthen global semantic modeling.

[0048] like Figure 2As shown in (b), the local module is a lightweight local feature extraction module that effectively captures local spatial information in images while maintaining computational efficiency through deep convolution combined with the GELU activation function. This module maintains consistency in the input and output feature dimensions, facilitating multi-layer stacking within the network and creating a good complementarity of local information with the subsequent global modeling process.

[0049] The structure of the progressive fusion module PFHR, which incorporates high-frequency residuals, is as follows: Figure 2 As shown in (c), the input features are dynamically divided into two branches along the channel (C) dimension: a convolutional branch and an SS2D scanning branch. The input feature channels are proportionally divided using a channel partitioning coefficient α. Channels with a ratio of α enter the SS2D scanning branch, while those with a ratio of 1-α enter the convolutional branch. This invention employs a progressive fusion strategy to dynamically adjust the coefficient α, with the α values ​​set sequentially to {0.25, 0.5, 0.5, 0.75} across four stages. This allows the model to dynamically balance the contributions of global and local features at different levels, achieving hierarchical synergy.

[0050] In the SS2D scanning branch, the data first undergoes a depthwise convolution, followed by detail enhancement using high-frequency residuals, and finally enters the SS2D module to capture global context information. The high-frequency residual mechanism is implemented through Laplace decoupling, and its specific steps are as follows: For features... Features are obtained by performing average pooling. Then, its spatial resolution is restored through nearest neighbor upsampling, compared with the original input. The residual term is obtained by subtracting pixels. The above steps can be expressed as the following formula:

[0051] (1)

[0052] In the convolution branch, the residual term and the characteristics of entering this branch spliced ​​together along the channel dimension The SS2D scanning branch and the convolutional branch are then subjected to a depthwise convolution to enhance local detail features. Finally, the outputs of the SS2D scanning branch and the convolutional branch are concatenated, and the concatenated features are added to the original feature input to obtain the final output features of the PFHR. The output features are then processed through a feed-forward network (FFN) and residual connections to improve feature representation, thus obtaining the final output of the Mamba convolutional module.

[0053] Figure 3 (a) illustrates the structure of spatiotemporal reconstruction modeling. First, the bi-temporal features are recombined in the spatiotemporal feature reconstruction module to generate three types of spatiotemporal dependencies: sequential, intersecting, and parallel. Then, through methods such as... Figure 3 The VisualState Space Block (VSS Block) shown in (b) learns the corresponding spatiotemporal relationships and models them respectively. (1) Sequential relationship: By unfolding and arranging the feature tokens of the two time phases in chronological order, that is, the scanning will first scan the feature tokens of the first time phase one by one in sequence, and then scan the feature tokens of the second time phase one by one in the same order, thereby capturing the sequential dependency relationship across time phases. This chronological unfolding method helps the model perceive the continuity of time and then extract the temporal change pattern. (2) Cross relationship: By alternating the tokens of the two time phases, the interaction between spatiotemporal features is strengthened, which can enhance the association learning between the same areas of the two time phases and enable the model to better capture the fine spatiotemporal difference features. (3) Parallel relationship: By splicing the features of the two time phases in the channel dimension, joint spatiotemporal modeling is performed to express both spatial and temporal features, and the collaborative fusion of dual-time phase information is realized. Finally, the feature results of the three spatiotemporal relationship models are spliced ​​and merged, and the difference features are obtained through a convolution operation.

[0054] Differences in different stages The feature fusion module in the decoder is then processed. The structure of the multi-scale feature fusion module is as follows: Figure 4 As shown. First, the features are upsampled. A smoothing layer consisting of convolutional layers and activation functions maps the number of channels in low-level features to the number of channels in high-level features. Consistent dimensions are obtained; then, the high and low layer features are added and fused element by element, and then the output is obtained through a residual smoothing layer, which is passed to the next stage for step-by-step decoding. The spatial resolution is gradually restored, and finally a predicted binary map is generated.

[0055] To mitigate the negative impact of foreground-background imbalance in image data on training performance, a strategy combining the cross-entropy (CE) loss function and the Dice loss function was adopted, as shown in the following expression:

[0056] (2)

[0057] (3)

[0058] in, It is the first The actual value of each pixel. Indicates the first The probability of a pixel; This represents the number of pixels. The overall loss function is expressed as:

[0059] (4)

[0060] in, and These represent the weighting coefficients in the overall loss function, with values ​​of 0.75 and 0.25 respectively.

[0061] Example

[0062] To verify the effectiveness of the proposed method, experiments and analyses were conducted on three publicly available datasets: LEVIR-CD, LEVIR-CD+, and WHU-CD. The model performance was evaluated from both quantitative and qualitative perspectives. Simultaneously, comparative experiments were performed using current mainstream advanced change monitoring methods, including three typical architectures: CNN-based methods (FC-EF, FC-Siam-Diff, IFN, and SNUNet); Transformer-based methods (Changeformer and BIT); and Mamba-based methods (ChangeMamba, MCD, and CDMamba).

[0063] Table 1. Accuracy results of each model on the LEVIR-CD dataset.

[0064]

[0065] Table 2. Accuracy results of each model on the LEVIR+-CD dataset.

[0066]

[0067] Table 3. Accuracy results of each model on the WHU-CD dataset.

[0068]

[0069] Tables 1, 2, and 3 present the quantitative evaluation indicators for LEVIR-CD, LEVIR-CD+, and WHU-CD, respectively. The best results are indicated in bold.

[0070] Table 1 presents the experimental results on the LEVIR-CD dataset. Our model achieved optimal or near-optimal performance on several core metrics, with Pre, F1, IoU, and OA reaching 92.70%, 91.18%, 83.81%, and 99.12%, respectively, comprehensively outperforming other comparative methods. This demonstrates a significant advantage in the accuracy of distinguishing changed regions and the overall consistency of detection results. Although methods such as IFN, ChangeFormer, and ChangeMamba have higher Rec values, our model performs better on Pre and F1, indicating greater robustness in suppressing spurious changes and reducing false detection rates.

[0071] Table 2 presents the experimental results on the LEVIR-CD+ dataset. This dataset features richer scenes and more complex variations, significantly increasing the detection difficulty. Our model achieved state-of-the-art performance on several key metrics, with Pre, Rec, F1, and IoU reaching 86.50%, 84.60%, 85.54%, and 74.68%, respectively, all outperforming other comparative methods. The results demonstrate that even under complex backgrounds and diverse interference, our model maintains a good balance between accuracy and completeness in identifying changed regions, providing a more accurate and stable portrayal of real-world changes.

[0072] Table 3 shows the experimental results on the WHU-CD dataset. This model demonstrates strong competitiveness across various metrics, achieving Recall, F1, and IoU of 93.80%, 94.32%, and 89.24%, respectively, placing it at a near-optimal level; its OA reaches 99.55%, outperforming all compared methods. Although it did not achieve optimal results on all datasets, this model maintains a balance between Pre and Recall while maintaining the best OA, avoiding performance imbalance caused by an excessively high single metric. For example, ChangeFormer's Pre reaches 96.15%, 1.32% higher than this model, but its F1 and IoU are 0.99% and 0.44% lower than DMCNet, respectively.

[0073] The experimental results from the three datasets show that the model maintains optimal or near-optimal performance in core metrics such as F1, IoU, and OA, with minimal performance fluctuations across different datasets. This indicates that the model can not only accurately identify changing regions but also maintain stable output even with significant variations in scene complexity, demonstrating its good generalization ability and robustness in complex environments.

[0074] Figure 5 , Figure 6 and Figure 7 The visualization results for LEVIR-CD, LEVIR-CD+, and WHU-CD are presented respectively. The visualization legends are as follows: white represents the correctly changed area (TP), black represents the correctly unchanged area (TN), red represents the false positive area (FP), and green represents the false negative area (FN).

[0075] Figure 5 The visualization results for the LEVIR-CD dataset are presented. This model's predictions of change boundaries are smoother and more complete, with a significant reduction in missed detections and false positives. Figure 5 In the comparison results (a), for complex edge variation areas, other methods have both missed detections and false detections, while this model only has a small number of missed detections and no obvious false detections. Figure 5 In the comparison results (b) in the model, the method of this model missed fewer detections in the change region compared with other methods, showing better integrity.

[0076] Figure 6 The visualization results of the LEVIR-CD+ dataset are presented. Figure 6 In the comparison results (a), other methods have varying degrees of missed detection of changes in small target buildings, while this model can effectively capture their changes. Figure 6 In the comparison results (b) above, some other methods are prone to misjudgment under complex background conditions, incorrectly identifying background changes as building changes, resulting in false detections. In contrast, this model can effectively suppress background interference and is more refined in characterizing the boundary of change.

[0077] Figure 7 The visualization results of the WHU-CD dataset are presented. Figure 7 In (a) of the model, the proposed model demonstrates good robustness in complex background regions, with almost no false detections caused by background interference. Furthermore, the model maintains high continuity and integrity in representing changing regions, effectively avoiding breaks or omissions. Similarly, in Figure 7 In (b) of the model, there are fewer false positives in the background compared to other methods, and no large-scale missed detections occur within the changing region. Furthermore, this model demonstrates higher prediction accuracy in edge region characterization, more accurately reconstructing the structural contours of the changing target, and its overall performance surpasses other comparative methods.

[0078] Table 4 and Figure 8 This paper presents a comparison of the parameter size and computational complexity of different change detection models on the LEVIR-CD dataset. Thanks to the hybrid network architecture combining Mamba and CNN, this model effectively controls the computational cost to 26.65G while maintaining a moderate model size of 36.7M parameters. Its computational cost is the lowest among all compared Transformer and Mamba architectures, significantly lower than mainstream models such as BIT, ChangeFormer, and ChangeMamba, and also significantly better than some classic CNN methods such as IFN and SNUNet. Even compared to these models with much higher computational costs, this model demonstrates a significant advantage in detection accuracy, achieving the best IoU performance on the LEVIR-CD dataset.

[0079] Specifically, while achieving optimal IoU, this model's Flops are only 23.21% of ChangeMamba's, 90.09% of MCD's, 89.91% of CDMamba's, and 15.11% of SNUNet's. In summary, the experimental results fully verify that this model achieves state-of-the-art (SOTA) detection performance while maintaining computational efficiency, further demonstrating that it achieves a good balance between efficiency and accuracy, showcasing its comprehensive advantages in change detection tasks.

[0080] Table 4 Computational costs of different models

[0081]

[0082] To visually demonstrate the feature learning and evolution process of this model, this invention employs Grad-CAM for model visualization analysis. For example... Figure 9 As shown, in the shallow layers of the encoder, the model focuses on local details and texture features in the input image; as the network depth increases, the encoder gradually shifts to modeling global semantic information. Based on this, the spatiotemporal reconstruction modeling module significantly enhances the model's ability to respond to features in areas of building change. By fully interacting and jointly modeling bi-temporal features, it generates more discriminative differential feature representations. Ultimately, the feature response region output by the decoder closely matches the predicted change map.

[0083] The above embodiments are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make several improvements and equivalent substitutions without departing from the principle of the present invention. All such improvements and equivalent substitutions to the claims of the present invention fall within the protection scope of the present invention.

Claims

1. A method for detecting changes in remote sensing images based on a Mamba-CNN dual-branch network model, characterized in that, Includes the following steps: Step 1: Collect dual-temporal remote sensing image data of the target area at different shooting times, and preprocess the data; Step 2: Input the preprocessed data into the Mamba convolutional encoder to extract shallow texture detail features, mid-level transition features and deep global semantic features, and output multi-scale encoded features corresponding to both time periods. In step 2, the Mamba convolutional encoder includes four levels, with channel dimensions of {64, 128, 384, 512} for each level and a stacking number of modules of {4, 4, 12, 6} for each level. Each level works in collaboration with a local module and a progressive fusion module (PFHR) incorporating high-frequency residuals. The PFHR module includes a convolutional branch and an SS2D scanning branch. The input feature channels are proportionally divided by the channel division coefficient α. The channel features with the α ratio enter the SS2D scanning branch, and the channel features with the 1-α ratio enter the convolution branch. A progressive fusion strategy is employed to dynamically adjust the coefficient α, with the α values ​​set sequentially to {0.25, 0.5, 0.5, 0.75} across the four stages. This allows the model to dynamically balance the contributions of global and local features at different levels. In the SS2D scanning branch, the model first undergoes a depthwise convolution, then high-frequency residuals are used for detail enhancement, and finally, it enters the SS2D module to capture global contextual information. In the convolution branch, the residual term... and features entering the convolution branch spliced ​​together along the channel dimension It undergoes a depthwise convolution to enhance local detail features; Finally, the outputs of the SS2D scanning branch and the convolution branch are concatenated to output the corresponding multi-scale encoded features in both time periods. Step 3: Input the multi-scale encoded features into the spatiotemporal reconstruction modeling module, reconstruct the three types of bi-temporal spatiotemporal dependencies of sequential, cross, and parallel through spatiotemporal feature reconstruction, and then concatenate the three types of bi-temporal spatiotemporal dependencies to generate multi-scale joint difference features; In step 3, the three types of dual-temporal spatiotemporal dependencies—sequential, intersecting, and parallel—are reconstructed through spatiotemporal feature recombination. Specifically, the sequential dependency is captured by arranging the features of the two temporal phases in chronological order, i.e., scanning the features of the first temporal phase sequentially and then scanning the features of the second temporal phase sequentially in the same order. The intersecting dependency strengthens the interaction between spatiotemporal features by alternating the features of the two temporal phases, enhancing the association learning between the same regions of the two temporal phases. The parallel dependency splices the features of the two temporal phases along the channel dimension to perform joint spatiotemporal modeling, simultaneously expressing spatial and temporal features and achieving the collaborative fusion of dual-temporal information. Step 4: Input the multi-scale joint difference features into the decoder to perform multi-scale feature fusion and generate a predicted binary map of the remote sensing image. Step 5: Train the Mamba-CNN dual-branch network model by combining the cross-entropy (CE) loss function and the Dice loss function, and iteratively optimize the model through backpropagation.

2. The remote sensing image change detection method based on the Mamba-CNN dual-branch network model according to claim 1, characterized in that, In step 1, the preprocessing includes cropping the dual-temporal remote sensing image data into a uniform pixel standard size and normalizing noise reduction.

3. The remote sensing image change detection method based on the Mamba-CNN dual-branch network model according to claim 1, characterized in that, The local module employs a lightweight local feature extraction unit, constructed using deep convolution combined with the GELU activation function. The input and output feature dimensions are kept consistent, and it is used to capture local edge and high-frequency texture details in dual-temporal remote sensing image data.

4. The remote sensing image change detection method based on the Mamba-CNN dual-branch network model according to claim 1, characterized in that, The detailed enhancement using high-frequency residuals specifically involves: processing the input features... Features are obtained by performing average pooling. Then, its spatial resolution is restored through nearest neighbor upsampling, and compared with the input features. The residual term is obtained by subtracting pixels. This can be expressed as the following formula: ; Where Pool represents average pooling and Upsample represents upsampling.

5. The remote sensing image change detection method based on the Mamba-CNN dual-branch network model according to claim 1, characterized in that, Step 4 specifically involves the following steps: the difference features from different stages are fed into the fusion module in the decoder. First, the upsampled features are passed through a smoothing layer consisting of a convolutional layer and an activation function to map the number of channels of the low-level features to the same dimension as the high-level features. Then, the high and low-level features are fused element by element and then passed through a residual smoothing layer to obtain the output, which is then passed to the next stage for progressive decoding. The spatial resolution is gradually restored, and finally, a binary image of the remote sensing image is generated.

6. The remote sensing image change detection method based on the Mamba-CNN dual-branch network model according to claim 1, characterized in that, Step 5 specifically involves the calculation formulas for the cross-entropy (CE) loss function and the Dice loss function, as follows: ; ; in, It is the first The actual value of each pixel. Indicates the first The probability of a pixel; This represents the number of pixels; the overall loss function is expressed as: ; in, and This represents the weighting coefficients in the overall loss function.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method of claim 1.

Citation Information

Patent Citations

  • Remote sensing semantic change detection method based on local detail continuity keeping

    CN121121528A

  • Remote sensing image semantic segmentation method based on global dependence and local detail collaboration

    CN122135023A