Feature decoupling and fusion deep network for remote sensing change interpretation

By introducing feature decoupling and fusion deep networks into remote sensing change detection technology, the coupling problem of changing characteristics and non-change characteristics is solved, and higher detection accuracy and generalization capabilities are achieved.

CN120182822APending Publication Date: 2025-06-20ZHONGBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510264845.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing remote sensing change detection technology has the problem of coupling between changing characteristics and non-change characteristics, resulting in reduced detection accuracy and poor generalization.

Method used

A deep network of feature decoupling and fusion for remote sensing change interpretation is proposed. Mamba Out network is used for multi-scale feature extraction. The feature decoupling module designs a feature decoupling module to distinguish between phase changes and invariant features, and avoids information loss through the reconstruction module. Finally, PIPP loss regularization is introduced to enhance attention to positive example samples.

Benefits of technology

Effectively decoupling of changing and unchange characteristics improves the accuracy and generalization ability of change detection, and significantly improves the performance of the model on different data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182822A_ABST
    Figure CN120182822A_ABST
Patent Text Reader

Abstract

The invention relates to the field of remote sensing image change detection, in particular to a remote sensing change interpretation-oriented feature decoupling and fusion deep network, which comprises an encoder, a decoupling module, a reconstruction module and a change detection head, inputting the dual-time-phase remote sensing images into an encoder to extract multi-scale features; inputting the obtained multi-scale features into a decoupling module so as to obtain specific features and shared features of each scale of each input image; finally, inputting the specific features and the shared features of each scale into a change detection head to generate a final change graph; in the middle, a reconstruction module is used for respectively combining specific features and shared features of each scale and reconstructing the specific features and the shared features into a reconstruction graph, in a training stage, reconstruction loss is calculated for an obtained reconstruction graph and an input graph, change graph loss is calculated for a generated final change graph and a label GT, and then the weight of each parameter in the network is adjusted by using a loss value obtained by calculation. Experimental results show that the network disclosed by the invention has good effects in indexes such as F1 and mIoU and qualitative comparison.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image change detection, and specifically to a feature decoupling and fusion deep network for remote sensing change interpretation. Background Technique

[0002] The main purpose of remote sensing image change detection is to find out the changed areas from remote sensing images of different time phases. With the continuous development of remote sensing technology, change detection using remote sensing images as data sources has become one of the research hotspots and has now been widely applied in environmental monitoring, economic development, national defense construction, etc.

[0003] Traditional remote sensing image change detection technologies rely on methods such as algebraic calculation, post-classification comparison, and image transformation. Their processes are relatively unified and generally divided into three steps: The first step is to preprocess the data to make the data more comprehensive; the second step is to use a model to perform image difference or image comparison operations to generate a difference map; the third step is to classify it to obtain a change map. In the past few years, researchers have developed many representative change detection algorithms, such as change vector analysis, slow feature analysis, and Fourier transform. Although the above methods have greatly improved the detection accuracy, they lack adaptability. Changes in data distribution or scenarios may lead to performance degradation, severely limiting the generalization of traditional methods and their applications in actual scenarios, making it difficult for them to become the mainstream.

[0004] Due to the powerful non-linear representation ability and feature extraction ability of deep neural networks, deep learning has made breakthrough progress in the applications of multiple industries and fields. It shows natural advantages in high-resolution remote sensing image change detection with its effective generalization ability. Currently, the research on using deep learning for high-resolution remote sensing image change detection has become a hot topic in the field of remote sensing applications.

[0005] In the field of deep learning, the mainstream networks for change detection cover convolutional neural networks, recurrent neural networks, and generative adversarial networks. First of all, the convolutional neural network CNN, as a classic deep learning network, has a powerful deep feature extraction ability and is commonly used for feature extraction and classification tasks in change detection. Many deep learning networks for change detection are built based on CNN. According to the data input form and network structure, it can be divided into two categories: single-stream structure and two-stream structure. The single-stream mode usually uses a single neural network to merge and input dual-temporal images to perform feature extraction and change map generation tasks, mainly covering two types: direct classification and mapping transformation-based. The two-stream structure, as the mainstream framework for change detection, consists of two parallel neural networks. The dual-temporal images are respectively input into the networks for feature extraction, and then the generated paired features are used to generate a change map, including the siamese structure and the post-classification structure.

[0006] Secondly, the Recurrent Neural Network (RNN) is a recursive neural network that takes sequential data as input. The input data can be regarded as sequential data, and then the transformation features between images can be directly obtained through the network. Furthermore, the Generative Adversarial Network (GAN) belongs to an unsupervised deep learning model, which consists of a generator and a discriminator. Through continuous generative adversarial training, an optimal generation model can be obtained. In the field of change detection, it can not only be used to generate training samples to expand the scale of the dataset, but also generate change maps based on the given reference image and query image.

[0007] Although CNN, RNN, and GAN have achieved remarkable results in the field of remote sensing image change detection, there are still some problems. For example, Figure 1 as shown in Fig. a, the coupling of change features and non-change features makes it easy to be interfered by non-change factors when detecting actual changes in change detection, resulting in a decrease in accuracy. Therefore, in-depth research in this field is of great significance. Summary of the Invention

[0008] Aiming at the problem that the current remote sensing change interpretation deep network using hybrid feature extraction and fusion easily leads to blurred prediction class boundaries, the present invention proposes a feature decoupling and fusion deep network for remote sensing change interpretation, as Figure 1 shown in Fig. b. First, use the Mamba Out network for multi-scale feature extraction; secondly, design a feature decoupling module to decouple the multi-scale features to distinguish the change and non-change features of each time phase; then, to ensure the completeness of the decoupling result, a reconstruction module is proposed to reconstruct the decoupled features and compare them with the original image; finally, aiming at the problem of data imbalance in current remote sensing image change detection, a new regularization loss is introduced to achieve accurate detection of remote sensing image changes.

[0009] The present invention is implemented by the following technical solutions: A feature decoupling and fusion deep network for remote sensing change interpretation, including an encoder, a decoupling module, a reconstruction module, and a change detection head; First, input the dual-temporal remote sensing images into the encoder to extract multi-scale features; then, input the obtained multi-scale features into the decoupling module to obtain the unique features and shared features of each scale of each input image; finally, input the unique features and shared features of each scale into the change detection head, and the unique features and shared features of each scale are concatenated at the channel level and convolution operations are performed to obtain the features of each scale. These scale features are processed by convolution and upsampling to generate the final change map; in the middle, the reconstruction module merges the unique features and shared features of each scale respectively, and then reconstructs the merged result into a reconstructed map. In the training stage, calculate the reconstruction loss between the obtained reconstructed map and the input map; and calculate the change map loss between the generated final change map and the label GT, and then use the calculated loss value to adjust the weight of each parameter in the network.

[0010] The above-mentioned feature decoupling and fusion deep network for remote sensing change interpretation uses the MambaOut network as the encoder.

[0011] For the above-mentioned feature decoupling and fusion deep network for remote sensing change interpretation, the reconstruction process of the reconstruction module is as follows: In the formula, T i P represents the unique feature obtained after the decoupling module processes the i-th image, and T i S represents the shared feature obtained after the decoupling module processes the i-th image. Here, i, j ∈ {1, 2}, i ≠ j, and θ represents the network parameters in the reconstruction module G.

[0012] For the above-mentioned feature decoupling and fusion deep network for remote sensing change interpretation, the specific process of the reconstruction module is as follows: perform residual processing and convolution operations on the unique feature and the shared feature respectively. Then, cascade the two and perform convolution operations again to obtain the overall feature at each scale. Transpose-convolve the overall features at each scale starting from the small scale and splice them with the features of the previous scale. Finally, process the obtained features through convolution and the Tanh activation function to obtain the reconstructed image.

[0013] For the above-mentioned feature decoupling and fusion deep network for remote sensing change interpretation, calculate the loss by using the input image to which the unique feature included in the reconstructed image belongs and the reconstructed image. The reconstruction loss

[0014]

[0015] For the above-mentioned feature decoupling and fusion deep network for remote sensing change interpretation, when calculating the loss between the final change map and the label GT, use the binary cross-entropy loss function and the PIPP loss function. The binary cross-entropy loss The PIPP loss Among them, N nc represents the number of unchanged pixels, N c represents the number of pixels that have not changed, Y hw represents the label pixel value at positions h and w, P hw represents the probability that the pixel at positions h and w is predicted to change. τ represents a parameter. The change map loss L chan and the total loss L total are: L change = L PIPP + λL BCE L total = L change + βL recon .

[0016] This invention focuses on the problem of blurred prediction class boundaries caused by the use of hybrid feature extraction and fusion in the dual-temporal remote sensing image change detection task, and proposes a neural network for feature decoupling and fusion. First, Mamba Out is used as the backbone network for multi-scale feature extraction; subsequently, a feature decoupling module is used to distinguish the unique features and shared features at each scale; then, the decoupled features are reconstructed through a reconstruction module to avoid information loss during the decoupling process; finally, PIPP loss regularization is introduced to enhance the attention to positive example samples. Description of the Drawings

[0017] Figure 1 This is a comparison chart of the method of the present invention and the traditional method. In the figure, a represents the traditional method, and b represents the method of the present invention.

[0018] Figure 2 This is the structural diagram of the network of the present invention.

[0019] Figure 3 This is the chart of the qualitative comparison experimental results on the LEVIR-CD dataset.

[0020] Figure 4 This is the chart of the qualitative comparison experimental results on the CDD dataset.

[0021] Figure 5 This is the chart of the qualitative comparison experimental results on the SYSUCD dataset.

[0022] Figure 6 This is the network visualization chart using the LEVIR CD dataset. Detailed Implementation Manner

[0023] A feature decoupling and fusion deep network FDFCD for remote sensing change interpretation, the network structure is as Figure 2 shown, including an encoder, a decoupling module, a reconstruction module, and a change detection head. First, to more effectively retain local details and abstract semantic information simultaneously and adapt to different scale change targets. The dual-temporal remote sensing images are respectively input into the encoder to extract multi-scale features After that, the obtained multi-scale features are input into the decoupling module, so as to obtain the unique features (i.e., the changed areas) and shared features (i.e., the unchanged areas) at each scale of each input image. Subsequently, the unique features and shared features at each scale are respectively merged through the reconstruction module, and then the results after merging the four scales are reconstructed into a reconstructed image. Immediately afterwards, the unique features and shared features at each scale are concatenated at the channel level and convolution operations are performed to obtain the features at each scale. Finally, these four-scale features are processed through convolution and upsampling to generate the final change map. In the training stage, the four obtained reconstructed images need to calculate the loss with the input images; and the generated change map needs to calculate the loss with the label GT. The loss function uses binary cross-entropy and PIPP loss, and then the calculated loss values are used to adjust the weight of each parameter in the network.

[0024] Encoder

[0025] The Mamba structure takes SSM as the core, has the ability to effectively process sequence information and the advantage of parallel training, and was initially widely used in natural language processing, speech processing and time series prediction. Due to the deficiencies of Mamba in visual tasks, the Mamba Out network was further proposed, and experiments have proved the feasibility of the Mamba Out network in visual tasks. The present invention uses the Mamba Out network as the encoder.

[0026] Assuming that the input feature is x and the output feature is y, the calculation formula for the Mamba structure is as follows:

[0027]

[0028] For the improved Mamba Out network, the calculation formula is as follows:

[0029]

[0030] Among them, Linear represents the linear layer, Conv represents the convolutional layer, represents the Cartesian product, σ represents the activation function, and the activation function used here is the GELU activation function.

[0031] Decoupling module

[0032] As Figure 2 shown in a, two groups of multi-scale features belong to two different time domains. To obtain the change map, the present invention first performs a decoupling operation on them. Among them, BN refers to Batch Normalization (batch normalization operation), and the multi-scale features F1 of the dual-time image m and are respectively decomposed into unique features and shared features through the decoupling module. Specifically, the shared features and unique features of the dual-time image are extracted by the decoupling modules E p and E s :

[0033] T i P = E p (F i m ; θ) (3)

[0034] T i S = E s (F i m ; θ) (4)

[0035] Among them, i = {1, 2} represents images of different time phases, and θ represents the network parameters in the decoupling module. In a deep neural network, as the number of network layers increases, the gradient may disappear or explode during backpropagation. Therefore, a resblock is introduced in the decoupling module to solve this problem. The skip connection of the resblock helps the gradient to propagate more stably, reducing the risk of gradient disappearance or explosion, and thus improving the training effect.

[0036] Reconstruction module

[0037] To improve the extraction quality of specific features and shared features, instead of directly sending the features to the change detection head as usual, the present invention designs a reconstruction module to reconstruct the obtained specific features (i.e., change features) and shared features (i.e., non-change features) back to the original image, and then calculates the loss between the reconstruction result and the input image respectively, as shown in the following formula:

[0038]

[0039] In the formula, T i P represents the specific features obtained after the i-th image passes through the decoupling module, and T i S represents the shared features obtained after the i-th image passes through the decoupling module, i, j ∈ {1, 2}, i ≠ j, and θ represents the network parameters in the reconstruction module G. Formulas 5 and 6 respectively represent the self-consistency reconstruction and cross-reconstruction of heterogeneous features. By putting a set of changing and unchanging features into the reconstruction module, a reconstructed image can be obtained.

[0040] The reconstruction module network is as shown in Figure 2 Figure b. In the figure, Upsampling is upsampling, and the transposed convolution method is used for upsampling here. First, residual processing and convolution operations are respectively performed on the specific features and shared features, and then the two are concatenated and convolution operation is performed again to obtain the overall features of each scale. Then, the overall features of each scale are sequentially subjected to transposed convolution starting from the small scale and spliced with the features of the previous scale. Finally, the obtained features are processed by convolution and the Tanh activation function to obtain the reconstructed image. In the training stage, for the link of calculating the loss between the reconstructed image and the input image, the selection of the input image follows a specific basis. Specifically, the input image to which the specific features included in the reconstructed image belong is selected to calculate the loss using the reconstructed image. The internal logic is that the goal of training is to minimize the intra-class difference in the unchanged area. Therefore, the selection method of the present invention conforms to the training goal, helps to more accurately measure the accuracy of changing and unchanging features, and further improves the performance of the entire model.

[0041] The specific reconstruction loss is as follows:

[0042]

[0043] Change detection head

[0044] In this module, first, for each group of input features, the features of the same scale are sequentially processed through modules such as convolution and residual to obtain the corresponding features for each scale. Then, the features of each scale are sequentially upsampled, convolved, and cascaded in ascending order to obtain the final change map. Finally, the obtained change map is used to calculate the loss with the label GT, and the loss function uses cross-entropy and PIPP loss, as Figure 2 shown in c.

[0045] Loss function

[0046] For remote sensing change detection tasks, there is often a situation where the number of positive and negative samples is extremely unbalanced, and the number of negative samples is much larger than that of positive samples. To make the accuracy index reach a higher value, the model tends to predict all samples as negative samples during training, resulting in poor detection ability for positive samples and a large number of false negatives.

[0047] As a traditional loss function, binary cross-entropy loss has certain limitations when facing this unbalanced data, although it performs well in many classification tasks. It treats all samples equally and does not consider the difference in the number of positive and negative samples. The formula is as follows:

[0048]

[0049] To solve the above problems, a positive example push-pull loss regular term is proposed to improve the detection ability for unbalanced samples by paying more attention to positive example samples, as Figure 2 shown in d, and the formula is as follows:

[0050]

[0051] Among them, N nc represents the number of unchanged pixels, N c represents the number of pixels that have not changed, Y hw is the label pixel value at the position (h, w), and P hw represents the probability that the pixel at this position is predicted to change. The PIPP loss function introduces the parameter τ, and during training, it makes the model assign higher weights to changing pixels, thereby balancing the influence of changing pixels and unchanged pixels in the learning process to a certain extent, effectively alleviating the data imbalance problem, and then significantly improving the model performance.

[0052] Therefore, the present invention selects to use positive example push-pull loss to handle the data imbalance problem, and uses binary cross-entropy to assist the model in more comprehensive learning of sample features. The change map loss and the total loss are as follows:

[0053] L chang = L PIPP + λL BCE (10)

[0054] L total = L chang + βL recon (11)

[0055] Experiment and Result Analysis

[0056] The present invention verifies the effectiveness of the network of the present invention through some experiments. First, three data sets used in the experiments are introduced. Second, the parameter settings used in the experiments are introduced. Then, the evaluation metrics used in the quantitative experiments are introduced. Finally, the experimental results are analyzed and the effectiveness of the PIPP loss function is verified through ablation experiments.

[0057] Introduction to the Data Sets

[0058] To evaluate the effectiveness of the network of the present invention, experiments are carried out based on three benchmark remote sensing image change detection data sets (LEVIR-CD, CDD, SYSUCD). The detailed information of these three data sets is as follows:

[0059] LEVIR-CD consists of 637 pairs of very high-resolution (0.5 m / pixel) Google Earth image patches, each patch with a size of 1024×1024 pixels, covering various types of buildings such as villas, high-rise apartments, small garages, and large warehouses. The rich building types included in this data set provide diverse samples for accurately detecting building-related changes, which helps to comprehensively evaluate the change detection ability of the model in complex building scenarios.

[0060] CDD is a public data set constructed based on satellite images captured in different seasons. The sizes of its changed areas vary, covering elements such as buildings, roads, and vehicles. Since its images are from different seasons, it can reflect the changes caused by various factors such as seasonal changes and human activities, increasing data diversity and complexity, and posing higher requirements for the generalization ability of the model. The present invention cuts the images into 256×256 pixel blocks and divides them into 10000, 2998, and 3000 pairs for training, validation, and testing respectively.

[0061] The SYSUCD data set mainly comes from scenes such as campuses and contains a large number of pedestrian images under different postures, different perspectives, and different lighting conditions, providing rich and diverse samples for pedestrian detection algorithms to adapt to pedestrian detection tasks in various practical application scenarios.

[0062] Parameter Settings

[0063] This network is implemented in the Pytorch framework and trained on a single NVIDIA RTX 4090 GPU. The SGD optimizer is used to optimize the network, with its momentum set to 0.9 and the weight decay coefficient to 5×10 -4 , and the initial learning rate lr is set to 0.01. The learning rate of the optimizer adopts a dynamic adjustment strategy, decaying by 0.7 times every 20 training epochs. The parameter τ in the loss function is set to 1, and λ and β are set to 10 and 0.5. In addition, each dataset is trained 150 times, and the batch size is set to 24. After each round of training, the validation set is used for validation, the F1 metric and other relevant evaluation metrics are calculated, the performance of the model on the validation set is recorded in detail, and the model parameters at the highest F1 metric are selected as the final prediction weights.

[0064] Evaluation Metrics

[0065] To comprehensively evaluate the model, the present invention uses five evaluation metrics to quantitatively evaluate the change detection results, namely Precision (abbreviated as P), Recall (abbreviated as R), Overall Accuracy (abbreviated as OA), F1-score (abbreviated as F1), and mean Intersection over Union (abbreviated as mIoU). These metrics measure the performance of the model in the change detection task from different perspectives. Precision reflects the proportion of samples predicted as positive and actually being positive by the model; Recall represents the proportion of samples that are actually positive and correctly predicted by the model; Overall Accuracy indicates the proportion of correctly predicted samples in the total number of samples; The F1-score comprehensively considers Precision and Recall and can more comprehensively evaluate the performance of the model; The mean Intersection over Union is used to measure the overlap degree between the prediction result and the true label. Their calculation formulas are as follows:

[0066]

[0067]

[0068] Among them, TN represents the number of correctly predicted negative samples; TP represents the number of correctly predicted positive samples; FP represents the number of incorrectly predicted positive samples; FN represents the number of incorrectly predicted negative samples.

[0069] Comparative Experiments

[0070] To verify the effectiveness of the proposed network, it is compared with a variety of advanced change detection methods, including FC-EF, FC-Siam-Diff, FC-Siam-Conc, DTCDSCN, BIT, Change Former, ICIFNe t, SRCDNet, and STANet. Experiments are carried out on a unified platform using the publicly available codes and default parameters of these methods.

[0071] The quantitative comparison results of various methods on LEVIR-CD, CDD and SYSUCD datasets are listed in Tables 1, 2 and 3 (the bold black part in the table indicates the best effect). It can be seen from the experimental results in the table that FDFCD shows excellent results on both datasets, among which the indicators F1, OA, and mIoU all reach the optimal values. These indicators comprehensively consider the accuracy of the model in change detection and the ability to handle samples of different categories. The network performs best in these indicators on different datasets, indicating that it can effectively capture the image change characteristics in various scenarios, and is not affected by the differences in datasets and the size of the changing target, thus proving that the network has good generalization ability and adaptability.

[0072] Table 1 Comparison of performance of various methods on the LEVIR-CD dataset

[0073]

[0074] Table 2 Comparison of performance of various methods on CDD dataset

[0075]

[0076] Table 3 Comparison of performance of various methods on the SYSUCD dataset

[0077]

[0078] Figure 3 , Figure 4 and Figure 5Qualitative comparison of experimental results on the LEVIR-CD, CDD, and SYSUCD datasets. In the experiment, white represents TP, black represents TN, red represents FP, and green represents FN. Among them, (a) image at time T1, (b) image at time T2, (c) GT (Ground Truth, true label), (d) FC-EF, (e) FC-Siam-Diff, (f) FC-Siam-Conc, (g) DTCDSCN, (h) BIT, (i) ChangeFormer, (j) ICIFNet, (k) SRCDNet, (I) Our FDFCD. Obviously, for small targets, models such as FC-EF and BIT have missed detection phenomena; in terms of large target areas, models such as SEIFNet have inaccurate edge detection, while FDFCD performs excellently in terms of detection accuracy, edge integrity, and adaptability to different scale targets.

[0079] Ablation experiment

[0080] Since this model is not limited to a specific backbone network, experiments are carried out on the LEVIR-CD dataset for different backbone networks.

[0081] As shown in Table 4: The experimental results are for using Mamba Out, Mamba, and ResNet as backbone networks respectively. The experimental results show that among the backbone networks used, Mamba Out has the best performance, ResNet is the second, and Mamba ranks behind. This indicates that using Mamba Out as the backbone network, its synergy with subsequent convolutional operations effectively improves the accuracy and integrity of feature extraction, and thus is superior to other backbone network combinations in overall performance.

[0082] Table 4 Comparison of experimental results of different backbone networks

[0083]

[0084] Based on the PIPP loss function constructed according to the present invention, ablation experiments are carried out on the LEVIR-CD dataset. The details are shown in Table 5:

[0085] Table 5 Comparison of performance indicators of different loss functions

[0086]

[0087] The experimental results clearly show that after using the PIPP loss function, indicators such as F1 and mIoU have extremely significant improvements, which fully proves that the PIPP loss function is feasible and can effectively optimize the model performance.

[0088] To verify the influence of different loss function coefficients on the model performance, corresponding experiments were conducted on the LEVIR-CD dataset. The experimental results are shown in Table IV, where λ and β represent the coefficients of the binary cross-entropy loss function and the reconstruction loss term, respectively. The experimental results show that the method performs best when λ = 10 and β = 0.5.

[0089] TABLE IV

[0090] INFLUENCE OF REGULARIZATION COEFFICIENT ON LEVIR-CD

[0091]

[0092] *Color Convention: best.

[0093] Network Visualization

[0094] To understand FDFCD more intuitively, a representative sample was selected from LEVIR-CD to visualize the features generated at different stages, as Figure 6 shown. Given the dual-temporal images and labels ( Figure 6 a), after passing through the encoder with Mamba Out as the backbone, the decoupling module is used to decouple the features, generating unique features and shared features ( Figure 6 b and Figure 6 c), which indicates that the model effectively decouples the changing and unchanging regions. Then, the decoupled features are respectively passed through the reconstruction module to generate the reconstructed image ( Figure 6 d) and the heatmap ( Figure 6 e). They demonstrate the accuracy obtained by the decoupling module, which can thus be used for subsequent change detection.

[0095] The present invention focuses on the problem that the use of hybrid feature extraction and fusion in the dual-temporal remote sensing image change detection task leads to blurred prediction class boundaries, and proposes a neural network for feature decoupling and fusion. First, Mamba Out is used as the backbone network for multi-scale feature extraction; subsequently, a feature decoupling module is used to distinguish unique features and shared features at each scale; then, the decoupled features are reconstructed through a reconstruction module to avoid information loss during the decoupling process; finally, PIPP loss regularization is introduced to enhance the attention to positive example samples. The experimental results show that FDFCD has achieved good results in terms of indicators such as F1 and mIoU and qualitative comparison.

Claims

1. A feature decoupling and fusion deep network for remote sensing change interpretation, characterized by: It includes an encoder, a decoupling module, a reconstruction module and a change detection head. First, the dual-phase remote sensing images are respectively input into the encoder to extract multi-scale features. Then, the obtained multi-scale features are input into the decoupling module to obtain the unique features and shared features of each scale of each input image. Finally, each scale-specific feature and shared feature are input into the change detection head. Each scale-specific feature and shared feature are cascaded and convolved on the channel to obtain each scale feature. These scale features are convolved and upsampled to generate the final change map. In the middle, the scale-specific features and shared features are merged by the reconstruction module, and then the merged results are reconstructed into a reconstructed map. In the training stage, the reconstruction loss is calculated by comparing the obtained reconstructed map with the input map. The generated final change map and the label GT are used to calculate the change map loss, and then the calculated loss value is used to adjust the weights of each parameter in the network.

2. A feature decoupling and fusion deep network for remote sensing change interpretation according to claim 1, characterized in that: The encoder uses the Mamba Out network.

3. A feature decoupling and fusion deep network for remote sensing change interpretation according to claim 1, characterized in that: The reconstruction process of the reconstruction module is: Where, T i P represents the unique features of the i-th image after passing through the decoupling module, represents the shared features obtained after the i-th image passes through the decoupling module, i,j∈{1,2},i≠j, and θ represents the network parameters in the reconstruction module G.

4. A feature decoupling and fusion deep network for remote sensing change interpretation according to claim 3, characterized in that: The specific process of the reconstruction module is as follows: residual processing and convolution operations are performed on the unique features and shared features respectively, and then the two are cascaded and convolution operations are performed again to obtain the overall features of each scale, and the overall features of each scale are transposed and convolved in sequence from the small scale and concatenated with the features of the previous scale. Finally, the obtained features are processed by convolution and Tanh activation function to obtain the reconstructed image.

5. A feature decoupling and fusion deep network for remote sensing change interpretation according to claim 3 or 4, characterized in that: Select the input image and the reconstructed image to which the unique features contained in the reconstructed image belong to calculate the loss, the reconstruction loss 6. A feature decoupling and fusion deep network for remote sensing change interpretation according to claim 5, characterized in that: The binary cross entropy loss function and PIPP loss function are used to calculate the loss between the final change map and the label GT. PIPP Losses Among them, N nc Represents the number of unchanged pixels, N c Indicates the number of unchanged pixels, Y hw Indicates the label pixel value at position h, w, P hw represents the probability that the pixel at position h, w is predicted to be changed, τ represents the parameter, Change map loss L change And the total loss L total For: L change =L PIPP +λL BCE , L total =L chang +βL recon .