Remote sensing image change detection method based on task-driven Mamba joint model

By decoupling the remote sensing image change detection task into multiple sub-tasks using a task-driven Mamba joint model, a U-shaped three-layer architecture is constructed and a joint loss function is used to train the network. This solves the problems of edge blurring and false change recognition in remote sensing image change detection, and achieves efficient change region recognition.

CN121010867APending Publication Date: 2025-11-25DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511139934.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing remote sensing image change detection technologies suffer from technical bottlenecks in high-resolution images, such as edge blurring, semantic distortion, and false change recognition. Furthermore, it is difficult to ensure computational efficiency while simultaneously taking into account multi-granularity feature collaboration mechanisms.

Method used

A task-driven Mamba joint model is adopted, which decouples the remote sensing image change detection task into three sub-tasks: dual temporal feature interaction, difference feature capture, and target detail reconstruction. A U-shaped three-layer remote sensing image change detection network is constructed, and the network is trained using a joint loss function of cross-entropy and Lovasz-softmax, including a Stem module, a temporal interactive Mamba module, an edge-focusing Mamba difference module, and a dual-attention Mamba reconstruction module.

Benefits of technology

It significantly improves the accuracy of changing region identification, simplifies the network structure, enhances the stability and analytical capability of the model, and improves the effect of remote sensing image change detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010867A_ABST
    Figure CN121010867A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image change detection method based on a task-driven Mama joint model. The method comprises the following steps: acquiring double-time remote sensing image groups of different time phases at the same place through a public data set; end-to-end change detection is realized by designing a task-decoupled remote sensing image change detection network model of a U-shaped three-layer network architecture, and a dual-time remote sensing image group can obtain a binary change graph with change characteristics in an image through the network model; the network model comprises three core modules: a time-phase interactive Mama module extracts double-time-phase feature maps of three scales through multi-stage down-sampling, and a time-phase interactive Mama module extracts double-time-phase feature maps of three scales through multi-stage down-sampling; an edge focusing Mamba difference module independently captures a time phase characteristic difference at each scale to generate a multi-scale difference characteristic graph; and the double-attention Mama reconstruction module fuses the time phase features and the feedback information of the AFF module, and reconstructs a three-scale double-detail change feature graph group. According to the method, the long sequence modeling capability of Mamba is utilized, so that the remote sensing image change detection precision and the edge detail retention effect are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of remote sensing image change detection, and in particular to a remote sensing image change detection method based on a task-driven Mamba combined model. BACKGROUND

[0002] Remote sensing image change detection is widely used in urban expansion monitoring, disaster assessment and environmental change analysis by analyzing the differences between images of the same geographical area at different times. Although remote sensing change detection technology based on deep learning has evolved from CNN local modeling, ViT global perception to hybrid architecture, it still faces three contradictions: CNN and ViT are limited by local representation limitations and computational complexity, hybrid models integrate multi-scale features but are accompanied by parameter inflation; the twin network and the difference module are difficult to balance complex dependency modeling and fine-grained feature preservation in cross-time interaction; the emerging Mamba architecture realizes linear complexity long-range modeling through a selective state mechanism, but its variants have deficiencies in scene-specific representation, multi-level feature fusion and false change suppression. It is currently necessary to build a dynamic interaction network to break through the technical bottlenecks of edge blur, semantic distortion and false change recognition in high-resolution images while ensuring computational efficiency, and to strengthen the task-driven multi-granularity feature collaboration mechanism. SUMMARY

[0003] The present application provides a remote sensing image change detection method based on a task-driven Mamba combined model to overcome the problem that existing network models cannot focus on the characteristics of the remote sensing image change detection task.

[0004] To achieve the above purpose, the technical scheme of the present application is:

[0005] A remote sensing image change detection method based on a task-driven Mamba combined model, comprising:

[0006] S1, obtaining a double-time remote sensing image group through a public data set, the double-time remote sensing image group being two remote sensing images of the same place at different times;

[0007] S2, based on the subtasks decoupled from the remote sensing image change detection task based on deep learning, designing a U-shaped three-layer architecture remote sensing image change detection network model; and training the remote sensing image change detection network model using a joint loss function of cross-entropy and Lovasz-softmax; the decoupled subtasks are double-time feature interaction tasks, difference feature capture tasks and target detail reconstruction tasks;

[0008] The remote sensing image change detection network model comprises a Stem module, a time-interaction Mamba module, an edge-focused Mamba difference module, a double-attention Mamba reconstruction module, an AFF fusion module and a classifier module.

[0009] The Stem module preprocesses the dual-time remote sensing image set to obtain a high-dimensional feature map set containing two preprocessed images. The temporal interactive Mamba module extracts the correlation features of the high-dimensional feature map set to obtain two initial temporal feature maps, and then performs two-level downsampling processing to obtain temporal feature maps at two secondary scale levels, ultimately resulting in six temporal feature maps at three scales. The edge-focusing Mamba difference module extracts the difference features of the two temporal feature maps at each scale, generating the difference features corresponding to each scale. The system first extracts six temporal feature maps, then obtains three scale difference feature maps. The dual-attention Mamba reconstruction module receives the output of the six temporal feature maps and the AFF fusion module, and fuses them to reconstruct a set of dual-detail feature maps at three scales. The AFF fusion module fuses the set of dual-detail feature maps at each scale with its corresponding difference feature map, and outputs the fusion results of the first two levels to the next level dual-attention Mamba reconstruction module, and outputs the fusion result of the third level to the classifier module. The classifier module compresses the channels of the dual-detail feature maps to finally obtain a binary change map.

[0010] S3. Input the dual-time remote sensing image to be detected into the trained remote sensing image change detection network model to obtain the binary change map of the dual-time remote sensing image.

[0011] Furthermore, the dual-time remote sensing image group is preprocessed, including increasing the number of channels of the dual-time remote sensing image group through two convolutional layers to obtain a high-dimensional feature map group.

[0012] Furthermore, the temporal interactive Mamba module is constructed based on the dual-temporal feature interaction task; the temporal interactive Mamba module includes a weight-shared residual submodule, a first Mamba submodule, a weight-independent MLP submodule, and multiple convolutional blocks; its specific execution process includes:

[0013] S211. The high-dimensional feature map group is received through the weight-sharing residual submodule, which is used to learn the correlation features between the high-dimensional feature map group and output two correlation feature maps.

[0014] S212. The associated features of the two associated feature maps are concatenated channel by channel through a convolutional block, and then normalized and activated using an activation function; the activated features are then merged into a single feature map.

[0015] S213. Receive the single feature map through the first Mamba submodule, and recombine the single feature map into a fused feature map according to the cross-scanning mechanism, selective state space mechanism and cross-merging mechanism of the first Mamba submodule.

[0016] S214. The fused feature map and the two associated feature maps in step 211 are fused by a convolutional block to generate two corresponding temporal feature maps;

[0017] S215. The two corresponding time-phase feature maps are channel-adjusted through the weighted independent MLP submodule, and the two time-phase feature maps are output to obtain the two initial time-phase feature maps.

[0018] S216. According to steps S211 to S215, perform two-level downsampling operations on each initial temporal feature map to obtain the corresponding temporal feature maps at two secondary scale levels. That is, combine the two initial temporal feature maps and the corresponding temporal feature maps at the two secondary scale levels to obtain six temporal feature maps.

[0019] Furthermore, based on the differential feature capture task, the edge-focusing Mamba differential module is constructed. The edge-focusing Mamba differential module includes a differential submodule, a double-derivative edge extraction submodule, a second Mamba submodule, and a convolutional block; its specific execution process includes:

[0020] S221. The edge-focusing Mamba difference module receives the temporal feature map output by the temporal interactive Mamba module, and the two temporal feature maps of the same scale are spliced ​​and fused channel by channel and the number of channels is adjusted by the convolution block to obtain the preliminary fused features.

[0021] S222. The preliminary fusion features are received through the double derivative edge extraction submodule. The preliminary fusion features are processed by first-order derivative edge constraints and second-order edge constraints to extract first-order edge features and second-order edge features respectively. The first-order derivative edge constraints consist of convolutional layers initialized by the Scharr operator. The second-order derivative edge constraints consist of Laplace convolutional layers constrained by the Gauss algorithm.

[0022] S223. The difference submodule performs difference operations and channel splicing on the preliminary fused features and the first-order edge features and second-order edge features to obtain two difference features.

[0023] S224. The two difference features are normalized by a convolutional block and activated using an activation function, and then fused along the channels to obtain fused difference features;

[0024] S225. The fused difference features are processed through the dual-branch mechanism of the second Mamba submodule. The result of the second branch is multiplied by the result of the first branch to obtain the Mamba processed features. The first branch of the dual-branch mechanism includes: a normalization layer, a linear embedding layer, a SiLU activation function, a depthwise separable convolutional layer, and an SS2D module. The second branch includes: a FADC frequency adaptive dilated convolution and a normalization layer.

[0025] S226. Add the fused difference features to the Mamba processing features to obtain the final difference feature map;

[0026] S227. According to steps S221 to S226, extract the corresponding difference feature maps from the dual-temporal feature maps of the other two scales, and finally obtain the difference feature maps of the three scales.

[0027] Furthermore, based on the target object reconstruction task, a dual-attention Mamba reconstruction module is constructed. This module includes a channel attention Mamba submodule, a spatial attention Mamba submodule, and multiple convolutional blocks. The specific process includes:

[0028] S231, The first-layer dual-attention Mamba reconstruction module receives two temporal feature maps at the first-level scale;

[0029] S232. Two temporal feature maps are concatenated channel by channel through a convolutional block, then normalized and activated using an activation function to obtain fused features;

[0030] S233. The fusion feature is input into the channel attention Mamba submodule. Through the processing of the Mamba branch and the channel attention branch, the processing results of the two branches are multiplied channel by channel, and then the multiplication result is added to the fusion feature to obtain the channel attention enhancement feature.

[0031] S234. Input the channel attention enhancement feature into the spatial attention Mamba submodule. Process it through the Mamba branch and the spatial attention branch, multiply the processing results of the two branches channel by channel, and then add the multiplication result to the channel attention enhancement feature to obtain the detail change feature map of the corresponding scale and output it to the AFF module.

[0032] S235, the second and third layer dual-attention Mamba reconstruction module receives temporal feature maps of two secondary scales and the output of the AFF module; through a convolutional block, the two temporal feature maps of the same scale and the output of the AFF module are concatenated channel by channel, normalized and activated using an activation function to generate fused features; by executing steps S233 to S234, the detailed change feature map of the corresponding scale is obtained and output to the AFF module.

[0033] Furthermore, the steps for establishing the joint loss function include:

[0034] S241. Establish the cross-entropy loss function, whose expression is:

[0035]

[0036] In the formula, y i This represents the true value in the i-th pixel; L represents the probability of the i-th pixel; N represents the number of pixels; L represents the probability of the i-th pixel. ce The cross-entropy loss function;

[0037] S242. Based on the aforementioned cross-entropy loss function, establish a joint loss function of cross-entropy and Lovasz-softmax, the expression of which is:

[0038] L total =L ce +λ*L lov (2)

[0039] In the formula, λ is the joint weight of the loss function; L lov The loss function is Lovasz-softmax; L total This is the joint loss function.

[0040] The present invention has the following beneficial effects:

[0041] This invention explicitly decouples the change detection task into three sub-tasks: dual-temporal feature interaction, differential feature capture, and target detail reconstruction. It then constructs a temporally interactive Mamba module, an edge-focusing Mamba differential module, and a dual-attention Mamba reconstruction module to specifically align the network learning process with the task logic, significantly improving the accuracy of changed region identification. A U-shaped three-layer architecture is used to form a joint network, making the network more concise and efficient. Joint loss is employed to train the network, thereby enhancing the model's stability and analytical capabilities. The use of Mamba model-based techniques has significant theoretical implications for remote sensing image change detection methods. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 This is a flowchart of the remote sensing image change detection method of the present invention;

[0044] Figure 2 This is a network structure diagram of the deep network model of the present invention;

[0045] Figure 3a This is the first phase pseudo-color image selected from the test dataset in this embodiment of the invention;

[0046] Figure 3b This is the second temporal pseudo-color image selected from the test dataset in this embodiment of the invention;

[0047] Figure 3c This is a graph showing the change detection results of the test dataset in this embodiment of the invention;

[0048] Figure 3d This is a ground truth map for change detection of the test dataset in this embodiment of the invention. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] This embodiment provides a remote sensing image change detection method based on a task-driven Mamba joint model, such as... Figure 1 As shown, it includes:

[0051] S1. Obtain a dual-time remote sensing image set through a public dataset. The dual-time remote sensing image set consists of two remote sensing images of the same location at different times.

[0052] S2. Based on the decoupled subtasks of the remote sensing image change detection task using deep learning, a U-shaped three-layer architecture remote sensing image change detection network model is designed; and the remote sensing image change detection network model is trained using a joint loss function of cross-entropy and Lovasz-softmax; the decoupled subtasks are a dual-time feature interaction task, a difference feature capture task, and a target detail reconstruction task.

[0053] The remote sensing image change detection network model includes a Stem module, a temporal interactive Mamba module, an edge-focusing Mamba difference module, a dual-attention Mamba reconstruction module, an AFF fusion module, and a classifier module;

[0054] The Stem module preprocesses the dual-time remote sensing image set to obtain a high-dimensional feature map set containing two preprocessed images. The temporal interactive Mamba module extracts the correlation features of the high-dimensional feature map set to obtain two initial temporal feature maps, and then performs two-level downsampling processing to obtain temporal feature maps at two secondary scale levels, ultimately resulting in six temporal feature maps at three scales. The edge-focusing Mamba difference module extracts the difference features of the two temporal feature maps at each scale, generating the difference features corresponding to each scale. The system first extracts six temporal feature maps, then obtains three scale difference feature maps. The dual-attention Mamba reconstruction module receives the output of the six temporal feature maps and the AFF fusion module, and fuses them to reconstruct a set of dual-detail feature maps at three scales. The AFF fusion module fuses the set of dual-detail feature maps at each scale with its corresponding difference feature map, and outputs the fusion results of the first two levels to the next level dual-attention Mamba reconstruction module, and outputs the fusion result of the third level to the classifier module. The classifier module compresses the channels of the dual-detail feature maps to finally obtain a binary change map.

[0055] S3. Input the dual-time remote sensing image to be detected into the trained remote sensing image change detection network model to obtain the binary change map of the dual-time remote sensing image.

[0056] Specifically, firstly, a set of dual-temporal remote sensing images is obtained from a public dataset as the input data source for the change detection task. Based on the idea of ​​task decoupling and combined with existing deep learning technology, the deep learning remote sensing image change detection task is decomposed into three key sub-tasks: dual-temporal feature interaction, difference feature capture, and target detail reconstruction. Based on this, a U-shaped three-layer detection network is designed. Since the sub-tasks are obtained by task decoupling based on deep learning, the details of this part will not be repeated: The Stem module enhances the spatial features and increases the dimensionality of the input image to generate a high-dimensional feature map set; the temporal interactive Mamba module extracts the multi-scale correlation of dual-temporal features through temporal-aware state space modeling and performs downsampling to capture the global context; the edge-focusing Mamba difference module strengthens the edge response of difference features to accurately locate the change area in order to address the problem of blurred change boundaries; the dual-attention Mamba reconstruction module fuses multi-scale temporal features and difference information and reconstructs target details through a channel-space dual attention mechanism; the AFF fusion module adaptively aggregates multi-scale detail features to generate a pixel-level binary change map.

[0057] To optimize training performance, a joint loss function of cross-entropy and Lovasz-softmax is adopted. The former constrains pixel classification error, while the latter directly optimizes the segmentation intersection-union ratio, thereby improving the detection sensitivity of small target change regions. Finally, the dual-temporal images are input into the fully trained network to output a high-precision binary change map.

[0058] In this embodiment, the dual-time remote sensing image group is preprocessed, including increasing the number of channels of the dual-time remote sensing image group through two convolutional layers to obtain a high-dimensional feature map group.

[0059] Specifically, the first convolutional layer increases the number of channels in the 256×256×3 dual-time remote sensing image group from 3 to 32; the second convolutional layer increases the number of channels in the image after the channel increase to 64; resulting in two high-dimensional feature map groups with a size of 256×256×64.

[0060] In this embodiment, the temporal interactive Mamba module is constructed based on the dual-temporal feature interaction task; the temporal interactive Mamba module includes a weight-shared residual submodule, a first Mamba submodule, a weight-independent MLP submodule, and multiple convolutional blocks;

[0061] The weight-sharing residual submodule includes a normalization layer, a linear layer, and a convolutional layer.

[0062] The first Mamba submodule includes a normalization layer, a linear embedding layer, a SiLU activation function, a depthwise separable convolutional layer, and an SS2D module;

[0063] The SS2D module includes a cross-scan block, a selective state space, and a cross-merge block.

[0064] The MLP submodule includes a normalization layer, a linear layer that expands the mapping space, and a linear layer that restores the mapping space.

[0065] The specific execution process of the time-phase interactive Mamba module includes:

[0066] S211. The high-dimensional feature map group is received through the weight-sharing residual submodule, which is used to learn the correlation features between the high-dimensional feature map group and output two correlation feature maps with a size of 256×256×64.

[0067] S212. The associated features of the two associated feature maps are concatenated channel by channel through a convolutional block, and then normalized and activated using an activation function; the activated features are merged into a single feature map of size 256×256×64.

[0068] S213. The single feature map is received through the first Mamba submodule. The cross-scanning mechanism of the first Mamba submodule decomposes the single feature map into 4 one-dimensional sequences. Each sequence is learned through the state space in the selective state space mechanism. The one-dimensional sequences are recombined into a fused feature map of size 256×256×64 through the cross-merging mechanism.

[0069] S214. The fused feature map and the two associated feature maps in step 211 are fused by a convolutional block to generate two corresponding temporal feature maps of size 256×256×64.

[0070] S215. The two corresponding time-phase feature maps are channel-adjusted through the weighted independent MLP submodule, and the two time-phase feature maps are output to obtain the two initial time-phase feature maps.

[0071] S216. According to steps S211 to S215, perform a two-level downsampling operation on each initial temporal feature map to obtain temporal feature maps corresponding to two secondary scale levels. That is, combine the two initial temporal feature maps and the temporal feature maps corresponding to the two secondary scale levels to obtain six temporal feature maps. The scales of the temporal feature maps are 256×256×64, 128×128×128 and 64×64×256.

[0072] In this embodiment, the edge-focusing Mamba difference module is constructed based on the difference feature capture task. The edge-focusing Mamba difference module includes a difference submodule, a double-derivative edge extraction submodule, a second Mamba submodule, and a convolutional block; its specific execution process includes:

[0073] S221. The temporal feature map output by the temporal interactive Mamba module is received through the edge-focusing Mamba difference module, and the two temporal feature maps of the same scale are spliced ​​and fused channel by channel and the number of channels is adjusted through a convolutional block to obtain preliminary fused features; in this embodiment, the size of the two temporal feature maps is selected as 64×64×256.

[0074] S222. The preliminary fusion features are received through the double derivative edge extraction submodule. The preliminary fusion features are processed by first-order derivative edge constraints and second-order edge constraints to extract first-order edge features and second-order edge features respectively. The first-order derivative edge constraints consist of convolutional layers initialized by the Scharr operator. The second-order derivative edge constraints consist of Laplace convolutional layers constrained by the Gauss algorithm.

[0075] Specifically, the Scharr operator calculates gradient maps in four directions (45°, 90°, 135°, and 180°) of the image to obtain gradient maps in these four directions. The gradient magnitude is calculated to obtain the gradient magnitude value of the gradient maps in these four directions. The larger the gradient magnitude value, the higher the probability that the pixel belongs to the edge; thus, the edge of the changing region is located.

[0076] The initialization process for a Laplace convolutional layer is as follows:

[0077] x,y=meshgrid([-1,0,1],[-1,0,1]) (1)

[0078]

[0079] Laplace=Gauss(σ1)-Gauss(σ2) (3)

[0080] In the formula, σ represents the weights of the Guass algorithm; meshgrid represents the function of assembling a one-dimensional matrix into a two-dimensional matrix.

[0081] S223. The difference submodule performs difference operations and channel splicing on the preliminary fused features and the first-order edge features and second-order edge features to obtain two difference features.

[0082] Specifically, the preliminary fusion features and the first-order edge features are residually connected to obtain two corresponding temporal features, namely the first-order A temporal feature and the first-order B temporal feature; at the same time, the preliminary fusion features and the second-order edge features are residually connected to obtain two corresponding temporal features, namely the second-order A temporal feature and the second-order B temporal feature; the features of the same order are subtracted and concatenated with the channels to obtain two new difference features.

[0083] S224. The two difference features are normalized by a convolutional block and activated using an activation function, and then fused along the channels to obtain fused difference features;

[0084] S225. The fused difference features are processed through the dual-branch mechanism of the second Mamba submodule. The result of the second branch is multiplied by the result of the first branch to obtain the Mamba processed features. The first branch of the dual-branch mechanism includes: a normalization layer, a linear embedding layer, a SiLU activation function, a depthwise separable convolutional layer, and an SS2D module. The second branch includes: a FADC frequency adaptive dilated convolution and a normalization layer.

[0085] The FADC frequency adaptive dilated convolution process is as follows:

[0086]

[0087] In the formula, Y(p) represents the pixel value at position p in the output feature map; K is the convolution kernel size; W i These are the weight parameters of the convolution kernel; X(p+Δp) i ) represents the offset Δp from position p in the input feature map. i The pixel value at that location; D is the porosity for expanding the receptive field;

[0088] S226. Add the fused difference features to the Mamba processing features to obtain the final difference feature map;

[0089] S227. According to steps S221 to S226, extract the corresponding difference feature maps from the dual-temporal feature maps of the other two scales, and finally obtain the difference feature maps of the three scales.

[0090] In this embodiment, the dual-attention Mamba reconstruction module is constructed based on the target object reconstruction task. The dual-attention Mamba reconstruction module includes a channel attention Mamba submodule, a spatial attention Mamba submodule, and multiple convolutional blocks; the specific process includes:

[0091] S231, The first-layer dual-attention Mamba reconstruction module receives two temporal feature maps at the first-level scale;

[0092] S232. Two temporal feature maps are concatenated channel by channel through a convolutional block, then normalized and activated using an activation function to obtain fused features;

[0093] S233. The fusion feature is input into the channel attention Mamba submodule. Through the processing of the Mamba branch and the channel attention branch, the processing results of the two branches are multiplied channel by channel, and then the multiplication result is added to the fusion feature to obtain the channel attention enhancement feature.

[0094] S234. Input the channel attention enhancement feature into the spatial attention Mamba submodule. Process it through the Mamba branch and the spatial attention branch, multiply the processing results of the two branches channel by channel, and then add the multiplication result to the channel attention enhancement feature to obtain the detail change feature map of the corresponding scale and output it to the AFF module.

[0095] S235, the second and third layer dual-attention Mamba reconstruction module receives temporal feature maps of two secondary scales and the output of the AFF module; through a convolutional block, the two temporal feature maps of the same scale and the output of the AFF module are concatenated channel by channel, normalized and activated using an activation function to generate fused features; by executing steps S233 to S234, the detailed change feature map of the corresponding scale is obtained and output to the AFF module.

[0096] In a specific embodiment, the steps for establishing the joint loss function include:

[0097] S241. Establish the cross-entropy loss function, whose expression is:

[0098]

[0099] In the formula, y i This represents the true value in the i-th pixel; L represents the probability of the i-th pixel; N represents the number of pixels; L represents the probability of the i-th pixel. ce The cross-entropy loss function;

[0100] S242. Based on the aforementioned cross-entropy loss function, establish a joint loss function of cross-entropy and Lovasz-softmax, the expression of which is:

[0101] L total =L ce +λ*L lov (6)

[0102] In the formula, λ is the joint weight of the loss function; L lov The loss function is Lovasz-softmax; L total This is the joint loss function.

[0103] In this embodiment, a remote sensing image change detection method based on a task-driven Mamba joint model is used to conduct experiments on the LEVIR-CD dataset. The experimental results are shown in Table 1:

[0104] Table 1. Detection scores (%) on the LEVIR-CD dataset.

[0105] Score LEVIR-CD Precision 92.34 Recall 90.70 F1 91.51 OA 99.14 Kappa 91.05 mIoU 91.72 IoU 84.35

[0106] To more objectively evaluate the roles of the main modules and training functions in the remote sensing image change detection method based on the task-driven Mamba joint model of this invention, existing ablation experiments were added for illustration. Single modules or combinations of different modules were added to the ordinary prototype network to compare the experimental results. Specific experimental results are shown in Table 2.

[0107] Table 2. Change detection scores (%) for different modules

[0108]

[0109] From the above results, we can conclude that:

[0110] 1. The experimental results in Table 1 show that the proposed remote sensing image change detection method based on the task-driven Mamba joint model has good change detection performance, which proves that the method has excellent performance in remote sensing image change detection.

[0111] 2. The ablation experiment data in Table 2 show that the change detection results using the temporal interactive Mamba encoder are significantly better than those using the twin Mamba encoder, proving that the encoder using the temporal interactive architecture exhibits more robust performance in modeling representation.

[0112] 3. The ablation experiment data in Table 2 show that the change detection results with the addition of the edge-focused difference Mamba module are significantly better than those without the enhancement module. At the same time, the change detection results with the edge-focused Mamba module to enhance the difference extraction are significantly better than those with the basic difference module. This proves that the edge-focused Mamba module plays an important role in the extraction of difference features from remote sensing images and is beneficial to improving the difference extraction effect of the network model.

[0113] 4. The ablation experiment data in Table 2 show that the use of the dual-attention Mamba decoder has a significant impact on the experimental results. The detection results using this module are significantly better than those using the Mamba prototype network alone, proving that the Mamba module, together with the dual-attention mechanism, plays an important role in the reconstruction of the target object.

[0114] 5. The ablation experiment data in Table 2 show that, based on the use of the temporal interactive Mamba encoder, the edge focus difference Mamba module, and the dual attention Mamba decoder, the detection score using the joint loss is better than the detection score using the single loss. This proves that the addition of the joint loss makes the network fitting more accurate and is more conducive to improving the detection effect.

[0115] The experimental input set of dual-time remote sensing images is as follows: Figure 3a and Figure 3b As shown in the figure, the final output of the experiment shows the change detection results. Figure 3c As shown, the change detection truth map is as follows: Figure 3d As shown.

[0116] The present invention has the following beneficial effects:

[0117] This invention explicitly decouples the change detection task into three sub-tasks: dual-temporal feature interaction, differential feature capture, and target detail reconstruction. It then constructs a temporally interactive Mamba module, an edge-focusing Mamba differential module, and a dual-attention Mamba reconstruction module to specifically align the network learning process with the task logic, significantly improving the accuracy of changed region identification. A U-shaped three-layer architecture is used to form a joint network, making the network more concise and efficient. Joint loss is employed to train the network, thereby enhancing the model's stability and analytical capabilities. The use of Mamba model-based techniques has significant theoretical implications for remote sensing image change detection methods.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting changes in remote sensing images based on a task-driven Mamba joint model, characterized in that, include: S1. Obtain a dual-time remote sensing image set through a public dataset. The dual-time remote sensing image set consists of two remote sensing images of the same location at different times. S2. Based on the decoupled subtasks of the remote sensing image change detection task using deep learning, a U-shaped three-layer architecture remote sensing image change detection network model is designed; and the remote sensing image change detection network model is trained using a joint loss function of cross-entropy and Lovasz-softmax; the decoupled subtasks are a dual-time feature interaction task, a difference feature capture task, and a target detail reconstruction task. The remote sensing image change detection network model includes a Stem module, a temporal interactive Mamba module, an edge-focusing Mamba difference module, a dual-attention Mamba reconstruction module, an AFF fusion module, and a classifier module; The Stem module is used to preprocess the dual-time remote sensing image set to obtain a high-dimensional feature map set containing two preprocessed images; the temporal interactive Mamba module is used to extract the associated features of the high-dimensional feature map set to obtain two initial temporal feature maps, and then perform two-level downsampling processing to obtain the corresponding temporal feature maps at two secondary scale levels, finally obtaining temporal feature maps at three scales, i.e., six temporal feature maps. The edge-focusing Mamba difference module extracts the difference features of the two temporal feature maps at each scale, generating the difference feature map corresponding to each scale, thus obtaining three scale difference feature maps. The dual-attention Mamba reconstruction module receives the output of the six temporal feature maps and the AFF fusion module, and fuses them to reconstruct a set of dual-detail feature maps at three scales. The AFF fusion module fuses the set of dual-detail feature maps at each scale with its corresponding difference feature map, and outputs the fusion results of the first two levels to the next level dual-attention Mamba reconstruction module, and outputs the fusion result of the third level to the classifier module. The classifier module compresses the channels of the dual detail variation feature map to obtain a binary variation map. S3. Input the dual-time remote sensing image to be detected into the trained remote sensing image change detection network model to obtain the binary change map of the dual-time remote sensing image.

2. The remote sensing image change detection method based on a task-driven Mamba joint model according to claim 1, characterized in that, The dual-time remote sensing image group is preprocessed, including increasing the number of channels of the dual-time remote sensing image group through two convolutional layers to obtain a high-dimensional feature map group.

3. The remote sensing image change detection method based on a task-driven Mamba joint model according to claim 1, characterized in that, The temporal interactive Mamba module is constructed based on the dual-temporal feature interaction task; the temporal interactive Mamba module includes a weight-shared residual submodule, a first Mamba submodule, a weight-independent MLP submodule, and multiple convolutional blocks; Its specific execution process includes: S211. The high-dimensional feature map group is received through the weight-sharing residual submodule, which is used to learn the correlation features between the high-dimensional feature map group and output two correlation feature maps. S212. The associated features of the two associated feature maps are concatenated channel by channel through a convolutional block, normalized, and activated using an activation function; the activated features are then merged into a single feature map. S213. Receive the single feature map through the first Mamba submodule, and recombine the single feature map into a fused feature map according to the cross-scanning mechanism, selective state space mechanism and cross-merging mechanism of the first Mamba submodule. S214. The fused feature map and the two associated feature maps in step 211 are fused by a convolutional block to generate two corresponding temporal feature maps; S215. The two corresponding time-phase feature maps are channel-adjusted through the weighted independent MLP submodule, and the two time-phase feature maps are output to obtain the two initial time-phase feature maps. S216. According to steps S211 to S215, perform two-level downsampling operations on each initial temporal feature map to obtain the corresponding temporal feature maps at two secondary scale levels. That is, combine the two initial temporal feature maps and the corresponding temporal feature maps at the two secondary scale levels to obtain six temporal feature maps.

4. The remote sensing image change detection method based on a task-driven Mamba joint model according to claim 1, characterized in that, The edge-focusing Mamba difference module is constructed based on the difference feature capture task. The edge-focusing Mamba difference module includes a difference sub-module, a double derivative edge extraction sub-module, a second Mamba sub-module, and a convolutional block. Its specific execution process includes: S221. The edge-focusing Mamba difference module receives the temporal feature map output by the temporal interactive Mamba module, and the two temporal feature maps of the same scale are spliced ​​and fused channel by channel and the number of channels is adjusted by the convolution block to obtain the preliminary fused features. S222. The preliminary fusion features are received through the double derivative edge extraction submodule. The preliminary fusion features are processed by first-order derivative edge constraints and second-order edge constraints to extract first-order edge features and second-order edge features respectively. The first-order derivative edge constraints consist of convolutional layers initialized by the Scharr operator. The second-order derivative edge constraints consist of Laplace convolutional layers constrained by the Gauss algorithm. S223. The difference submodule performs difference operations and channel splicing on the preliminary fused features and the first-order edge features and second-order edge features to obtain two difference features. S224. The two difference features are normalized by a convolutional block and activated using an activation function, and then fused along the channels to obtain fused difference features; S225. The fused difference features are processed through the dual-branch mechanism of the second Mamba submodule. The result of the second branch is multiplied by the result of the first branch to obtain the Mamba processed features. The first branch of the dual-branch mechanism includes: a normalization layer, a linear embedding layer, a SiLU activation function, a depthwise separable convolutional layer, and an SS2D module. The second branch includes: a FADC frequency adaptive dilated convolution and a normalization layer. S226. Add the fused difference features to the Mamba processing features to obtain the final difference feature map; S227. According to steps S221 to S226, extract the corresponding difference feature maps from the dual-temporal feature maps of the other two scales, and finally obtain the difference feature maps of the three scales.

5. The remote sensing image change detection method based on a task-driven Mamba joint model according to claim 1, characterized in that, The dual-attention Mamba reconstruction module is constructed based on the target object reconstruction task. The dual-attention Mamba reconstruction module includes a channel attention Mamba submodule, a spatial attention Mamba submodule, and multiple convolutional blocks. Its specific process includes: S231, The first-layer dual-attention Mamba reconstruction module receives two temporal feature maps at the first-level scale; S232. Two temporal feature maps are concatenated channel by channel through a convolutional block, then normalized and activated using an activation function to obtain fused features; S233. The fusion feature is input into the channel attention Mamba submodule. Through the processing of the Mamba branch and the channel attention branch, the processing results of the two branches are multiplied channel by channel, and then the multiplication result is added to the fusion feature to obtain the channel attention enhancement feature. S234. Input the channel attention enhancement feature into the spatial attention Mamba submodule. Process it through the Mamba branch and the spatial attention branch, multiply the processing results of the two branches channel by channel, and then add the multiplication result to the channel attention enhancement feature to obtain the detail change feature map of the corresponding scale and output it to the AFF module. S235, the second and third layer dual-attention Mamba reconstruction module receives temporal feature maps of two secondary scales and the output of the AFF module; through a convolutional block, the two temporal feature maps of the same scale and the output of the AFF module are concatenated channel by channel, normalized and activated using an activation function to generate fused features; by executing steps S233 to S234, the detailed change feature map of the corresponding scale is obtained and output to the AFF module.

6. The remote sensing image change detection method based on a task-driven Mamba joint model according to claim 1, characterized in that, The steps for establishing the joint loss function include: S241. Establish the cross-entropy loss function, whose expression is: In the formula, y i This represents the true value in the i-th pixel; L represents the probability of the i-th pixel; N represents the number of pixels; L represents the probability of the i-th pixel. ce The cross-entropy loss function; S242. Based on the aforementioned cross-entropy loss function, establish a joint loss function of cross-entropy and Lovasz-softmax, the expression of which is: L total =L ce +λ*L lov (2) In the formula, λ is the joint weight of the loss function; L lov The loss function is Lovasz-softmax; L total This is the joint loss function.

Citation Information

Cited By

  • Ultrasonic image segmentation method based on edge perception and space channel Mama

    CN121305093A

  • Multi-task unsupervised change detection method fusing image domain alignment and segmentation network

    CN121725373A