Remote sensing image change detection method based on hybrid vision Mama network
Through the twin hybrid encoder and dual-branch fusion module of the hybrid vision Mamba network, the problem of local information ignorance in remote sensing change detection is solved, and a more refined detection effect is achieved.
Patent Information
- Application Number
- CN202510455092.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing remote sensing change detection method based on Mamba ignores local information in intensive prediction tasks, making it difficult to achieve fine detection.
The hybrid vision Mamba network is adopted to extract global and local features through a twin hybrid encoder, and the global differential decoder and local differential decoder are used to process feature differences respectively. The feature fusion is enhanced with the dual-branch fusion module to generate the final prediction result.
Effectively retaining the global context information and local features of the bi-time phase image alleviates the problem of Mamba lacking local clues in the change detection task and improves detection accuracy.
Smart Images

Figure CN120298902A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing image detection, and particularly relates to a remote sensing image change detection method based on a hybrid vision Mamba network. Background Art
[0002] The goal of the change detection task is to detect meaningful changes by inputting two remote sensing images acquired at the same location but at different times. Change detection plays an important role in various fields such as urban planning, land cover analysis, disaster assessment, ecosystem monitoring, and resource management.
[0003] The development of deep learning technology has brought a promising solution to the field of change detection. A common practice in existing models is to use CNN or Transformer as a siamese encoder to extract dual-temporal features. CNN is very effective in extracting complex and relevant features, especially good at local feature extraction, but often insufficient in capturing global dependencies. Transformer, through its self-attention mechanism, is good at dealing with complex spatial transformations and capturing long-range feature dependencies, which helps to form a comprehensive global representation. However, the complexity of using Transformer for image processing is quadratic with respect to the length of the image patch. This leads to huge computational costs and makes it unsuitable for dense prediction tasks such as change detection. Recently, Mamba introduced time-varying parameters into the state space model (SSM) to achieve data-dependent global modeling with linear complexity, achieving great success in the field of natural language processing. Subsequently, the Mamba architecture was extended to the field of computer vision and used in change detection, being considered an effective alternative to Transformer. However, in dense prediction tasks such as change detection, local information still plays a crucial role in accurate detection. Summary of the Invention
[0004] Aiming at the above deficiencies in the prior art, the remote sensing image change detection method based on a hybrid vision Mamba network provided by the present invention solves the problem that in the Mamba-based remote sensing change detection method, local information is ignored and it is difficult to achieve fine detection in dense prediction tasks.
[0005] To achieve the above invention purpose, the technical solution adopted by the present invention is: a remote sensing image change detection method based on a hybrid vision Mamba network, including the following steps:
[0006] S1. Obtain dual-temporal remote sensing images, and input the dual-temporal remote sensing images into a siamese hybrid encoder to obtain global features and local features;
[0007] S2. Input the global features and local features into a global difference decoder and a local difference decoder respectively to obtain global difference features and local difference features;
[0008] S3. Input the global difference features and local difference features into the dual-branch fusion module to obtain the fused features;
[0009] S4. Input the fused features into the fusion decoder to obtain the final prediction result.
[0010] Furthermore: In S1, the siamese hybrid encoder includes a backbone module, a first hybrid module, a second hybrid module, and a third hybrid module connected in sequence;
[0011] Among them, the backbone module includes a 3×3 convolutional layer connected to each other;
[0012] The first hybrid module, the second hybrid module, and the third hybrid module have the same structure, and each includes a parallel CNN branch and Mamba branch, and the CNN branch and the Mamba branch are connected by a dual-branch interaction module;
[0013] Among them, the Mamba branch is provided with a VSS sub-module, the CNN branch is provided with a first local sub-module and a second local sub-module connected to each other, and the dual-branch interaction module includes a first interaction sub-module and a second interaction sub-module;
[0014] The output end of the first local sub-module is connected to the input end of the VSS sub-module through the first interaction sub-module, and the output end of the VSS sub-module is connected to the input end of the second local sub-module through the second interaction sub-module.
[0015] Furthermore: The VSS sub-module includes a first normalization layer, a first linear layer, a depth convolutional layer, an SS2D unit, and a second normalization layer connected in sequence. The output end of the second normalization layer is connected to the first input end of the first element multiplication. The input end of the first normalization layer is used as the input end of the VSS sub-module. The output end of the first normalization layer is connected to the second input end of the first element multiplication through the second linear layer. The output end of the first element multiplication is connected to the first input end of the second element multiplication through the third linear layer. The output end of the first normalization layer is also connected to the second input end of the second element multiplication. The output end of the second element multiplication is used as the output end of the VSS sub-module;
[0016] The first local sub-module and the second local sub-module have the same structure, and each includes a reparameterized 3×3 convolution, a first 1×1 convolution, a GELU activation function, and a second 1×1 convolution connected in sequence. The second 1×1 convolution is also connected to the first element addition. The reparameterized 3×3 convolution is also connected to the first element addition.
[0017] Further: The dual-temporal remote sensing images in S1 include the remote sensing images before and after the change. The remote sensing images before and after the change are input into the twin hybrid encoder to obtain the global features and local features of the input images. Among them, the global features include the first to third Mamba features, and the local features include the first to third CNN features. The method for the twin hybrid encoder to obtain the output features according to the input images is specifically as follows:
[0018] S11. Input the input image into the backbone module, extract the main features in the input image, and downsample its size to 1 / 4 of the input image to obtain the first feature map;
[0019] S12. Input the first feature map into the first hybrid module to obtain the first CNN feature and the first Mamba feature;
[0020] S13. Input the first CNN feature and the first Mamba feature into the second hybrid module and the third hybrid module in sequence. Output the second CNN feature and the second Mamba feature through the second hybrid module, and output the third CNN feature and the third Mamba feature through the third hybrid module.
[0021] Further: In S13, the method for the second hybrid module to output the second CNN feature and the second Mamba feature is specifically as follows:
[0022] S131. Input the first CNN feature into the first local sub-module to obtain the second feature map;
[0023] S132. Perform 1×1 convolution, deformation, and normalization operations on the second feature map in sequence through the first interaction sub-module to obtain the third feature map;
[0024] S133. Add the first Mamba feature to the third feature map and input it into the VSS sub-module to obtain the second Mamba feature;
[0025] S134. Perform deformation, 1×1 convolution, and normalization operations on the second Mamba feature in sequence through the second interaction sub-module to obtain the fourth feature map;
[0026] S135. Add the fourth feature map to the second feature map and input it into the second local sub-module to obtain the second CNN feature.
[0027] Further: In S2, the global difference decoder and the local difference decoder have the same structure, both including the first to third DIFF modules and the linear fusion module. The first to third DIFF modules are all connected to the linear fusion module;
[0028] Among them, the first to third DIFF modules have the same structure, including a connected first 3×3 convolutional layer and a second 3×3 convolutional layer.
[0029] Furthermore: S2 includes the following steps:
[0030] S21. Input the global feature into the global difference decoder to obtain the global difference feature;
[0031] S22. Input the local feature into the local difference decoder to obtain the local difference feature;
[0032] Among them, the global difference decoder and the local difference decoder process the input features in the same way. S21 includes the following sub-steps:
[0033] S211. Input the first to third Mamba features into the first to third DIFF modules respectively;
[0034] S212. Merge the results generated by the first to third DIFF modules through the linear fusion module to obtain the first to third multi-scale difference features;
[0035] S213. Upsample the first to third multi-scale difference features to 1 / 4 of the size of the dual-temporal remote sensing image, connect the first to third multi-scale difference features along the channel dimension, and use 1×1 convolution to adjust the number of channels to obtain the global difference feature.
[0036] Furthermore: In S3, the dual-branch fusion module includes a first dynamic feature enhancement module, a second dynamic feature enhancement module, and a cross-branch fusion Mamba module. Both the first dynamic feature enhancement module and the second dynamic feature enhancement module are connected to the cross-branch fusion Mamba module;
[0037] S3 includes the following sub-steps:
[0038] S31. Fuse the global difference feature and the local difference feature to obtain the coarse-grained fusion feature;
[0039] S32. Input the global difference feature, the local difference feature, and the coarse-grained fusion feature into the first dynamic feature enhancement module to obtain the first enhanced feature;
[0040] S33. Input the global difference feature, the local difference feature, and the coarse-grained fusion feature into the second dynamic feature enhancement module to obtain the second enhanced feature;
[0041] S34. Input the first enhanced feature and the second enhanced feature into the cross-branch fusion Mamba module to obtain the fusion feature.
[0042] Furthermore: S32 includes the following sub-steps:
[0043] S321. Subtract the local difference feature from the global difference feature to obtain the first difference feature;
[0044] S322. After processing the first difference feature through a pooling operation, calculate the difference weight between the global difference feature and the local difference feature through dynamic difference-aware attention, and use it to amplify the feature difference of the coarse-grained fusion feature to obtain a second difference feature;
[0045] S323. Process the global difference feature through learnable descriptive convolution, and add the generated result, the second difference feature, and the global difference feature to obtain a first enhanced feature;
[0046] In S34, the expression for obtaining the fusion feature X is specifically:
[0047]
[0048] In the formula, ECA is the channel attention module, is the element-wise addition operation, D1 is the first enhanced feature, D2 is the second enhanced feature, LN is the normalization layer, H f is the result of adding the mixed features, and its expression is specifically:
[0049]
[0050] In the formula, H1 is the first mixed feature, H2 is the second mixed feature, and their expressions are specifically:
[0051]
[0052] In the formula, SS2D is the selective scanning 2D unit, is the element-wise multiplication operation, SiLU is a neural network activation function, Linear is the linear layer, H is the initial mixed feature, and its expression is specifically:
[0053]
[0054] In the formula, Dwc(·) is the depth convolution operation.
[0055] Furthermore: S4 is specifically:
[0056] Upsample the fusion feature by a factor of four through transposed convolution, input the upsampled result into the prediction head composed of convolutions, generate a change mask with the same size as the dual-temporal remote sensing image, and use it as the final prediction result.
[0057] The beneficial effects of the present invention are:
[0058] (1) The present invention provides a remote sensing image change detection method based on a hybrid vision Mamba network, constructs a siamese hybrid encoder based on multiple hybrid modules, where the hybrid module is composed of a CNN and a VMamba, and enables the features of the two branches of the encoder to interact and couple gradually through a double-branch interaction module, so as to maximize the retention of the global context information and local features of the bi-temporal images, and alleviate the problem that Mamba lacks local clues when dealing with change detection tasks.
[0059] (2) The present invention proposes a double-branch fusion module, which includes a dynamic feature enhancement module for enhancing fine texture features and perception difference features, and a cross-branch fusion Mamba module for effectively exploring the correlation of global and local difference features. Through the synergistic effect of these two sub-modules, more discriminative change features are extracted for the detection task. Description of the Drawings
[0060] Figure 1 It is a flowchart of the remote sensing image change detection method based on the hybrid vision Mamba network of the present invention.
[0061] Figure 2 It is a schematic structural diagram of the remote sensing image change detection method based on the hybrid vision Mamba network.
[0062] Figure 3 It is a schematic structural diagram of the hybrid module.
[0063] Figure 4 It is a schematic structural diagram of the VSS module.
[0064] Figure 5 It is a schematic structural diagram of the SS2D unit.
[0065] Figure 6 It is a schematic diagram of the double-branch fusion module.
[0066] Figure 7 It is a schematic diagram of the first dynamic feature enhancement module.
[0067] Figure 8 It is a schematic diagram of the LDC.
[0068] Figure 9 It is a schematic diagram of the cross-branch fusion Mamba module. Detailed Embodiments
[0069] The specific embodiments of the present invention will be described below to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0070] As Figure 1 shown, in an embodiment of the present invention, a remote sensing image change detection method based on a hybrid vision Mamba network includes the following steps:
[0071] S1. Obtain dual-temporal remote sensing images, and input the dual-temporal remote sensing images into a Siamese hybrid encoder to obtain global features and local features;
[0072] S2. Input the global features and local features into a global difference decoder and a local difference decoder respectively to obtain global difference features and local difference features;
[0073] S3. Input the global difference features and local difference features into a two-branch fusion module to obtain fusion features;
[0074] S4. Input the fusion features into a fusion decoder to obtain the final prediction result.
[0075] In this embodiment, the network structure of the present invention is as Figure 2 shown. The Siamese hybrid encoder takes the pre-change remote sensing image I pre and the post-change remote sensing image I post as inputs, calculates the corresponding CNN features pre and Mamba features of I and the corresponding CNN features post and Mamba features of I where i is the stage and i = 1, 2, 3.
[0076] In S1, the Siamese hybrid encoder includes a backbone module, a first hybrid module, a second hybrid module, and a third hybrid module connected in sequence;
[0077] Among them, the backbone module includes a 3×3 convolutional layer connected to each other;
[0078] As Figure 3 shown, the structures of the first hybrid module, the second hybrid module, and the third hybrid module are the same, and each includes a parallel CNN branch and Mamba branch, and the CNN branch and the Mamba branch are connected through a two-branch interaction module;
[0079] Among them, the Mamba branch is provided with a VSS sub-module, the CNN branch is provided with a first local sub-module and a second local sub-module connected to each other, and the interaction module of the dual branches includes a first interaction sub-module and a second interaction sub-module;
[0080] The output end of the first local sub-module is connected to the input end of the VSS sub-module through the first interaction sub-module, and the output end of the VSS sub-module is connected to the input end of the second local sub-module through the second interaction sub-module.
[0081] In this embodiment, in order to make full use of local features and global representations, the present invention designs the hybrid module as a concurrent structure, which includes a CNN branch, a Mamba branch, and an interaction module (FIM) of the dual branches.
[0082] As Figure 4 shown, the VSS sub-module includes a first normalization layer, a first linear layer, a depth convolution layer, an SS2D unit, and a second normalization layer connected in sequence. The output end of the second normalization layer is connected to the first input end of the first element multiplication. The input end of the first normalization layer is used as the input end of the VSS sub-module. The output end of the first normalization layer is connected to the second input end of the first element multiplication through the second linear layer. The output end of the first element multiplication is connected to the first input end of the second element multiplication through the third linear layer. The output end of the first normalization layer is also connected to the second input end of the second element multiplication. The output end of the second element multiplication is used as the output end of the VSS sub-module;
[0083] The structures of the first local sub-module and the second local sub-module are the same, and both include a reparameterized 3×3 convolution, a first 1×1 convolution, a GELU activation function, and a second 1×1 convolution connected in sequence. The second 1×1 convolution is also connected to the first element addition, and the reparameterized 3×3 convolution is also connected to the first element addition.
[0084] In this embodiment, the VSS sub-module is composed of a linear layer and a depth convolution layer, followed by an SS2D unit. The structure of the SS2D unit is as Figure 5 shown. The SS2D unit flattens the input and processes it in four different directions, namely from top left to bottom right, from bottom right to top left, from top right to bottom left, and from bottom left to top. Thus, remote dependencies can be extracted in multiple directions. For each direction, a selective scanning state model is learned by learning parameters A, B, C, D, Δ, and the results are combined to generate the final output of the SS2D unit.
[0085] The dual-temporal remote sensing images in S1 include the remote sensing images before and after the change. The remote sensing images before and after the change are input into the Siamese hybrid encoder to obtain the global features and local features of the input images. Among them, the global features include the first to third Mamba features, and the local features include the first to third CNN features. The method for the Siamese hybrid encoder to obtain the output features according to the input images is specifically as follows:
[0086] S11. Input the input image into the backbone module, extract the main features in the input image, and downsample its size to 1 / 4 of the input image to obtain the first feature map;
[0087] S12. Input the first feature map into the first hybrid module to obtain the first CNN feature and the first Mamba feature;
[0088] S13. Input the first CNN feature and the first Mamba feature into the second hybrid module and the third hybrid module in sequence. The second hybrid module outputs the second CNN feature and the second Mamba feature, and the third hybrid module outputs the third CNN feature and the third Mamba feature.
[0089] In S13, the method for the second hybrid module to output the second CNN feature and the second Mamba feature is specifically as follows:
[0090] S131. Input the first CNN feature into the first local sub-module to obtain the second feature map;
[0091] S132. Perform 1×1 convolution, deformation, and normalization operations on the second feature map in sequence through the first interaction sub-module to obtain the third feature map;
[0092] S133. Add the first Mamba feature to the third feature map and then input it into the VSS sub-module to obtain the second Mamba feature;
[0093] S134. Perform deformation, 1×1 convolution, and normalization operations on the second Mamba feature in sequence through the second interaction sub-module to obtain the fourth feature map;
[0094] S135. Add the fourth feature map to the second feature map and then input it into the second local sub-module to obtain the second CNN feature.
[0095] In this embodiment, the CNN branch uses 2 local sub-modules to provide a fulcrum for the two-way interaction between the two branches. Each local sub-module uses a re-parameterized 3×3 convolution to extract local features. The interaction sub-module includes a 1×1 convolution for adjusting the number of feature channels, and a series of deformation and normalization operations for solving the misalignment problem between the feature map of the CNN branch and the patch embedding of the Mamba branch. Considering the complementarity of global features and local features, the present invention uses the interaction sub-module to gradually integrate the local features of the CNN branch into the embedding block, enriching the Mamba branch with local details. Similarly, the global context from the Mamba branch is fed back to the feature map to enhance the global perception ability of the CNN branch. During the interaction between the two branches, the CNN features and the Mamba features are mutually embedded.
[0096] As Figure 2 shown, in S2, the global difference decoder and the local difference decoder have the same structure, both including the first to third DIFF modules and a linear fusion module. The first to third DIFF modules are all connected to the linear fusion module;
[0097] Among them, the first to third DIFF modules have the same structure, including a first 3×3 convolutional layer and a second 3×3 convolutional layer connected to each other.
[0098] S2 includes the following steps:
[0099] S21. Input the global feature into the global difference decoder to obtain the global difference feature;
[0100] S22. Input the local feature into the local difference decoder to obtain the local difference feature;
[0101] Among them, the global difference decoder and the local difference decoder process the input features in the same way, so the working process of the local difference decoder will not be elaborated here.
[0102] S21 includes the following sub-steps:
[0103] S211. Input the first to third Mamba features into the first to third DIFF modules respectively;
[0104] S212. Combine the results generated by the first to third DIFF modules through the linear fusion module to obtain the first to third multi-scale difference features;
[0105] S213. Upsample the first to third multi-scale difference features to 1 / 4 of the size of the dual-temporal remote sensing image, connect the first to third multi-scale difference features along the channel dimension, and use a 1×1 convolution to adjust the number of channels to obtain the global difference feature.
[0106] As Figure 2As shown, in this embodiment, the present invention also inputs the global difference features into the prediction head to generate a binary change map, representing the prediction result of the global decoder. Similarly, the local difference features generated by the local decoder are input into the prediction head to generate a binary change map, representing the prediction result of the local decoder. By generating a binary change map with a size of one-fourth of the input image for each of the two decoders and supervised by the change labels, the length of the gradient backpropagation path is effectively reduced, facilitating model training.
[0107] As Figure 6 shown, in S3, the dual-branch fusion module includes a first dynamic feature enhancement module, a second dynamic feature enhancement module, and a cross-branch fusion Mamba module. Both the first dynamic feature enhancement module and the second dynamic feature enhancement module are connected to the cross-branch fusion Mamba module;
[0108] S3 includes the following sub-steps:
[0109] S31. Fuse the global difference features and the local difference features to obtain coarse-grained fusion features;
[0110] S32. Input the global difference features, the local difference features, and the coarse-grained fusion features into the first dynamic feature enhancement module to obtain the first enhanced features;
[0111] S33. Input the global difference features, the local difference features, and the coarse-grained fusion features into the second dynamic feature enhancement module to obtain the second enhanced features;
[0112] S34. Input the first enhanced features and the second enhanced features into the cross-branch fusion Mamba module to obtain the fusion features.
[0113] In this embodiment, the dynamic feature enhancement module takes the global difference features and the local difference features as inputs and performs coarse-grained fusion in the module to obtain coarse-grained fusion features. By subtracting the features from different branches to obtain the difference features, the mapping of these difference features is enhanced. Subsequently, these difference features are merged with the original features, and the information from other branches is used to supplement and enrich the difference features. This process effectively extracts and amplifies the complementary features and texture details inherent in the image, thereby improving the overall fusion performance, helping the model to effectively capture the subtle differences between the input features, thus improving the resolution and perception between different features, and helping the model to enhance the fusion performance.
[0114] As Figure 7 shown, S32 includes the following sub-steps:
[0115] S321. Subtract the global difference features from the local difference features to obtain the first difference features;
[0116] S322. After processing the first difference feature through a pooling operation, calculate the difference weight between the global difference feature and the local difference feature through dynamic difference-aware attention, and use it to amplify the feature difference in the coarse-grained fusion feature to obtain the second difference feature; among them, the calculation of the difference weight can effectively amplify the feature difference;
[0117] S323. Process the global difference feature through learnable descriptive convolution, and add the generated result, the second difference feature, and the global difference feature to obtain the first enhanced feature;
[0118] In this embodiment, the difference between the methods of the first dynamic feature enhancement module and the second dynamic feature enhancement module for processing the input feature lies in the different inputs of the LDC. The input of the LDC in the first dynamic feature enhancement module is the global difference feature, and the input of the LDC in the second dynamic feature enhancement module is the local difference feature. LDC uses learnable mask parameters and convolution operations to enhance the texture processing of the input feature map. By adjusting the weights of the convolution kernels, it emphasizes the texture information and enhances the texture feature perception of the model.
[0119] As Figure 9 shown, in S34, the expression for obtaining the fusion feature X is specifically:
[0120]
[0121] In the formula, ECA is the channel attention module, is the element-wise addition operation, D1 is the first enhanced feature, D2 is the second enhanced feature, LN is the normalization layer, H f is the result of adding the mixed features, and its expression is specifically:
[0122]
[0123] In the formula, H1 is the first mixed feature, H2 is the second mixed feature, and their expressions are specifically:
[0124]
[0125] In the formula, SS2D is the selective scanning 2D unit, is the element-wise multiplication operation, SiLU is the neural network activation function, Linear is the linear layer, H is the initial mixed feature, and its expression is specifically:
[0126]
[0127] In the formula, Dwc(·) is the depth convolution operation.
[0128] In this embodiment, the cross-branch fusion Mamba module is as Figure 9As shown, the cross-branch fusion Mamba module takes D1 and D2 as inputs, performs fine-grained fusion, explores the information correlation between different branches, generates the initial mixed feature H, and then inputs the initial mixed feature into the SS2D layer to capture the long-range spatial dependence, generating H1 and H2. The channel attention module is used to reduce channel redundancy to obtain the final fused feature X.
[0129] S4 specifically is:
[0130] The fused feature is upsampled four times through transposed convolution, and the upsampled result is input into the prediction head composed of convolutions to generate a change mask with the same size as the bi-temporal remote sensing image, which is used as the final prediction result.
[0131] The present invention optimizes the predicted change maps obtained by the three decoders through the cross-entropy loss, and the cross-entropy loss is expressed as:
[0132]
[0133] In the formula, y represents the true value, represents the predicted value, and N represents the number of pixels.
[0134] The beneficial effects of the present invention are as follows: The present invention provides a remote sensing image change detection method based on a hybrid vision Mamba network, constructs a siamese hybrid encoder based on multiple hybrid modules, the hybrid module is composed of a CNN and a VMamba, and through the dual-branch interaction module, the features of the two branches of the encoder are gradually interactively coupled, so as to maximize the retention of the global context information and local features of the bi-temporal images, and alleviate the problem that Mamba lacks local clues when dealing with change detection tasks.
[0135] The present invention proposes a dual-branch fusion module, which includes a dynamic feature enhancement module for enhancing fine texture features and perceptual difference features, and a cross-branch fusion Mamba module for effectively exploring the correlation between global and local difference features. Through the synergistic effect of these two sub-modules, more discriminative change features are extracted for the detection task.
[0136] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of technical features. Therefore, the features defined by "first", "second", and "third" may explicitly or implicitly include one or more of such features.
Claims
1. A remote sensing image change detection method based on a hybrid vision Mamba network, characterized in that, Including the following steps: S1. Obtain dual-temporal remote sensing images, input the dual-temporal remote sensing images into the Siamese hybrid encoder, and obtain global features and local features; S2. Input the global features and local features into the global difference decoder and the local difference decoder respectively to obtain global difference features and local difference features; S3. Input the global difference features and local difference features into the dual-branch fusion module to obtain fusion features; S4. Input the fusion features into the fusion decoder to obtain the final prediction result.
2. The remote sensing image change detection method based on the hybrid vision Mamba network according to claim 1, wherein In S1, the Siamese hybrid encoder includes a backbone module, a first hybrid module, a second hybrid module, and a third hybrid module connected in sequence; Among them, the backbone module includes a 3×3 convolutional layer connected to each other; The first hybrid module, the second hybrid module, and the third hybrid module have the same structure, and both include a parallel CNN branch and a Mamba branch, and the CNN branch and the Mamba branch are connected through a dual-branch interaction module; Among them, the Mamba branch is provided with a VSS sub-module, the CNN branch is provided with a first local sub-module and a second local sub-module connected to each other, and the dual-branch interaction module includes a first interaction sub-module and a second interaction sub-module; The output end of the first local sub-module is connected to the input end of the VSS sub-module through the first interaction sub-module, and the output end of the VSS sub-module is connected to the input end of the second local sub-module through the second interaction sub-module.
3. The remote sensing image change detection method based on the hybrid vision Mamba network according to claim 2, wherein The VSS sub-module includes a first normalization layer, a first linear layer, a depth convolutional layer, an SS2D unit, and a second normalization layer connected in sequence. The output end of the second normalization layer is connected to the first input end of the first element multiplication. The input end of the first normalization layer is used as the input end of the VSS sub-module. The output end of the first normalization layer is connected to the second input end of the first element multiplication through the second linear layer. The output end of the first element multiplication is connected to the first input end of the second element multiplication through the third linear layer. The output end of the first normalization layer is also connected to the second input end of the second element multiplication. The output end of the second element multiplication is used as the output end of the VSS sub-module; The first local sub-module and the second local sub-module have the same structure, and both include a reparameterized 3×3 convolution, a first 1×1 convolution, a GELU activation function, and a second 1×1 convolution connected in sequence. The second 1×1 convolution is also connected to the first element addition, and the reparameterized 3×3 convolution is also connected to the first element addition.
4. The remote sensing image change detection method based on the hybrid vision Mamba network according to claim 3, wherein, In S1, the dual-temporal remote sensing images include the remote sensing images before and after the change. Input the remote sensing images before and after the change into the Siamese hybrid encoder to obtain the global features and local features of the input images. Among them, the global features include the first to third Mamba features, and the local features include the first to third CNN features. The method for the Siamese hybrid encoder to obtain the output features according to the input images is specifically as follows: S11. Input the input images into the backbone module, extract the main features in the input images, and downsample their sizes to 1 / 4 of the input images to obtain the first feature map; S12. Input the first feature map into the first hybrid module to obtain the first CNN feature and the first Mamba feature; S13. Input the first CNN feature and the first Mamba feature into the second hybrid module and the third hybrid module in sequence. Output the second CNN feature and the second Mamba feature through the second hybrid module, and output the third CNN feature and the third Mamba feature through the third hybrid module.
5. The remote sensing image change detection method based on the hybrid vision Mamba network according to claim 4, characterized in that, In S13, the method for the second hybrid module to output the second CNN feature and the second Mamba feature is specifically as follows: S131. Input the first CNN feature into the first local sub-module to obtain a second feature map. S132. Through the first interaction sub-module, perform 1×1 convolution, deformation, and normalization operations on the second feature map in sequence to obtain a third feature map. S133. Add the first Mamba feature to the third feature map and input it into the VSS sub-module to obtain the second Mamba feature. S134. Through the second interaction sub-module, perform deformation, 1×1 convolution, and normalization operations on the second Mamba feature in sequence to obtain a fourth feature map. S135. Add the fourth feature map to the second feature map and input it into the second local sub-module to obtain the second CNN feature.
6. The remote sensing image change detection method based on the hybrid vision Mamba network according to claim 4, wherein, In S2, the structures of the global difference decoder and the local difference decoder are the same, both including the first to third DIFF modules and a linear fusion module. The first to third DIFF modules are all connected to the linear fusion module. Among them, the structures of the first to third DIFF modules are the same, including a first 3×3 convolutional layer and a second 3×3 convolutional layer connected to each other.
7. The remote sensing image change detection method based on the hybrid vision Mamba network according to claim 6, wherein S2 includes the following steps: S21. Input the global feature into the global difference decoder to obtain a global difference feature. S22. Input the local feature into the local difference decoder to obtain a local difference feature. Among them, the methods for the global difference decoder and the local difference decoder to process the input features are the same. S21 includes the following sub-steps: S211. Input the first to third Mamba features into the first to third DIFF modules respectively. S212. Combine the results generated by the first to third DIFF modules through the linear fusion module to obtain the first to third multi-scale difference features. S213. Upsample the first to third multi-scale difference features to 1 / 4 of the size of the dual-temporal remote sensing image, connect the first to third multi-scale difference features along the channel dimension, and use 1×1 convolution to adjust the number of channels to obtain the global difference feature.
8. The remote sensing image change detection method based on the hybrid vision Mamba network according to claim 1, characterized in that, In S3, the dual-branch fusion module includes a first dynamic feature enhancement module, a second dynamic feature enhancement module, and a cross-branch fusion Mamba module. The first dynamic feature enhancement module and the second dynamic feature enhancement module are both connected to the cross-branch fusion Mamba module. S3 includes the following sub-steps: S31. Fuse the global difference feature and the local difference feature to obtain a coarse-grained fusion feature. S32. Input the global difference feature, the local difference feature, and the coarse-grained fusion feature into the first dynamic feature enhancement module to obtain a first enhanced feature. S33. Input the global difference feature, the local difference feature, and the coarse-grained fusion feature into the second dynamic feature enhancement module to obtain a second enhanced feature. S34. Input the first enhanced feature and the second enhanced feature into the cross-branch fusion Mamba module to obtain a fused feature.
9. The remote sensing image change detection method based on the hybrid vision Mamba network according to claim 8, wherein, S32 includes the following sub-steps: S321. Subtract the global difference feature from the local difference feature to obtain a first difference feature. S322. After processing the first difference feature through a pooling operation, calculate the difference weight between the global difference feature and the local difference feature through dynamic difference-aware attention, and use it to amplify the feature difference of the coarse-grained fused feature to obtain a second difference feature. S323. Process the global difference feature through learnable descriptive convolution, and add the generated result, the second difference feature, and the global difference feature to obtain a first enhanced feature. In S34, the expression for obtaining the fused feature X is specifically: where ECA is the channel attention module, is the element-wise addition operation, D1 is the first enhanced feature, D2 is the second enhanced feature, LN is the normalization layer, and H f is the result of adding the mixed features, and its expression is specifically: In the formula, H1 is the first mixed feature, and H2 is the second mixed feature. Their expressions are specifically: where SS2D is the selective scan 2D unit, is the element-wise multiplication operation, SiLU is the neural network activation function, Linear is the linear layer, and H is the initial mixed feature, and its expression is specifically: In the formula, Dwc(·) is a depth convolution operation.
10. The remote sensing image change detection method based on the hybrid vision Mamba network according to claim 1, characterized in that, S4 is specifically: Upsample the fused feature by a factor of four through transposed convolution, input the upsampled result into a prediction head composed of convolutions, generate a change mask with the same size as the dual-temporal remote sensing image, and use it as the final prediction result.
Citation Information
Patent Citations
Remote sensing image change detection network and detection method based on double twinborn branches
CN116524361A
High-resolution remote sensing image target detection method based on multi-scale network
CN118485927A
Infrared unmanned aerial vehicle group detection method based on state space model
CN118506222A
Remote sensing image building change detection method and system based on morphological constraint
CN118736425A
AU2020102569A4