Remote Sensing Image Change Detection Method Based on Dual-Stream Temporal Feature Adapter
Through the dual-stream time-phase feature adapter method, combined with the multi-receptive field feature pyramid and dynamic timing perception enhancement module, the problem of taking into account local details and global context in remote sensing image change detection is solved, which improves detection accuracy and consistency, and reduces the computational complexity.
Patent Information
- Application Number
- CN202510534851.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When processing remote sensing images, existing remote sensing image change detection methods are difficult to take into account local details and global context, cross-architecture deep interaction and domain adaptation, and have high computational complexity. The existing methods have insufficient model generalization capabilities and catastrophic forgetting in cross-domain tasks.
Using a method based on the dual-stream time-phase feature adapter, a bidirectional dynamic fusion feature extraction is performed by sharing the pre-trained basic visual Transformer model backbone network and two time-phase feature adapter paths, combining the multi-receptive field feature pyramid module, dynamic timing perception enhancement module and bidirectional interaction module to realize multi-scale feature extraction and timing feature perception, reducing the impact of non-semantic timing variation.
It improves the accuracy and consistency of remote sensing image change detection, reduces the computational overhead, enhances the robustness and cross-domain adaptability of the model, and generates a high-resolution change detection mask.
Smart Images

Figure CN120047450B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and remote sensing technology, and particularly to a remote sensing image change detection method and system based on a two-stream temporal feature adapter. Background Art
[0002] Remote sensing image change detection (CD) plays an important role in fields such as urban planning, environmental monitoring, and disaster assessment. By comparing remote sensing images from different periods, changes in surface cover can be identified, such as the increase or decrease of buildings, changes in vegetation, etc. With the continuous improvement of satellite technology and sensor resolution, it has become easier to obtain high-resolution remote sensing data, which also poses higher requirements for change detection technology.
[0003] Convolutional Neural Network (CNN) is a deep learning model, especially suitable for processing data with a grid structure, such as images. CNN automatically extracts local features in images, such as edges, textures, etc., through convolutional layers, and reduces the feature dimension through pooling layers to improve computational efficiency. Although CNN performs well in many computer vision tasks, when dealing with large-scale scenes, its receptive field is limited, making it difficult to fully model global context information. For example, FC-EF (Fully Convolutional Early Fusion) is a CNN-based change detection method that relies on convolutional operations to capture local texture patterns, but due to its limited receptive field, it is difficult to fully model global context information when dealing with large-scale scenes. In addition, CNN models usually require a large amount of training data to obtain good generalization ability, which is a challenge for remote sensing images with high annotation costs.
[0004] In traditional methods, many CNN-based architectures (such as FC-EF) and subsequent improved methods generally adopt a siamese network architecture. This architecture processes remote sensing images before and after change through two independent branch networks respectively, and then realizes change detection through feature fusion or difference calculation. However, although this design can capture local features to a certain extent, it has significant limitations: First, the siamese network needs to perform independent calculations on two sets of images, resulting in a multiple increase in computational resource consumption; Second, since the front and rear images are independently processed through the same network branch, it is easy to have feature misalignment problems. Especially in complex scenes or when the image has geometric distortion, it is difficult to maintain the spatial correspondence relationship of features. In addition, the siamese network lacks an effective adaptation mechanism for the domain difference between natural images and remote sensing images, resulting in insufficient cross-scene generalization ability.
[0005] In recent years, the Transformer model has received extensive attention due to its powerful ability to capture long-range dependencies. Transformer is a deep learning model based on the self-attention mechanism, initially used in the field of natural language processing, but has also been widely applied to computer vision tasks in recent years. Different from CNN, Transformer can capture long-range dependencies between arbitrary positions in the input sequence through the self-attention mechanism, thus enhancing the understanding of global context. For example, ChangeFormer can better capture long-range dependencies through the self-attention mechanism, thereby enhancing the understanding of global context. However, this method has the problem of high computational complexity, especially when performing pixel-level localization on high-resolution images, the performance is not good. In addition, when the Transformer model processes high-resolution images, it consumes a large amount of computing resources, which limits its popularization in practical applications.
[0006] The Vision Transformer (ViT) is an image recognition model developed based on the Transformer architecture in the field of natural language processing. Taking the Segment Anything Model (SAM) as an example, this is an advanced segmentation model that can perform high-quality segmentation on any object in the open world. SAM uses a powerful pre-trained ViT model to extract hierarchical semantic features of images and achieves generality for various tasks through fine-tuned adapters. However, when directly applying ViT to remote sensing image change detection, two main dilemmas will be encountered: First, the distribution difference between natural images and remote sensing images seriously affects the segmentation prior effect of ViT. Remote sensing images usually have higher spatial resolution and more complex scene structures, which makes it challenging for ViT to process remote sensing images; Second, existing methods are difficult to coordinate the hierarchical semantic abstraction of ViT with the per-pixel localization accuracy required for change detection. How to make ViT adapt to fine-grained change reasoning in different fields while maintaining its powerful segmentation ability is an open challenge that has not been solved.
[0007] Existing change detection methods still have many deficiencies when dealing with remote sensing images, especially in terms of balancing local details and global context, cross-architecture deep interactions, and domain adaptation. Traditional methods are difficult to effectively model global context information in large-scale scenarios, while Transformer-based methods face the problem of high computational complexity. Although hybrid architectures attempt to combine the advantages of both, they often introduce redundant parameters and implicit fusion strategies, affecting the efficiency and stability of the model. For example, TransUNetCD embeds Transformer modules in the U-Net structure, which improves the processing ability of optical images but still has a large computational overhead. In addition, pre-trained base models are prone to catastrophic forgetting in cross-domain tasks, and existing lightweight adapters cannot effectively retain the hierarchical segmentation priors of pre-trained models, resulting in poor performance in cross-domain tasks.
[0008] In summary, although existing deep learning methods have made some progress in remote sensing image change detection, there are still many challenges to be solved. How to effectively support remote sensing image change detection based on making full use of the segmentation priors of pre-trained vision Transformer models remains an important research direction. Future work needs to pay more attention to the design of domain adaptation strategies to improve the robustness and generalization ability of models in different application scenarios. Summary of the Invention
[0009] The main purpose of the present invention is to provide a remote sensing image change detection method based on a two-stream temporal feature adapter, which can overcome the deficiencies of the prior art in remote sensing image change detection, obtain multi-scale feature differences more efficiently, and improve the accuracy and consistency of change detection.
[0010] Provide a remote sensing image change detection method based on a two-stream temporal feature adapter, including the following steps:
[0011] S1. Obtain the remote sensing image of the previous time phase and the remote sensing image of the subsequent time phase;
[0012] S2. Use the remote sensing image of the previous phase and the remote sensing image of the subsequent phase as the inputs of the encoder. Through the backbone network of the pre-trained basic Vision Transformer model of the encoder and two temporal feature adapter paths, perform bidirectional dynamic fusion feature extraction processing in n stages to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase in n stages. Among them, the backbone network of the pre-trained basic Vision Transformer model includes a patch embedding layer and n sequentially connected block processing blocks. Each temporal feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules. Each block processing block corresponds to a bidirectional gating module on each temporal feature adapter path to perform bidirectional dynamic fusion feature extraction processing in one stage, and obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase in this stage. n is a positive integer;
[0013] S3. Input the global context feature maps of the previous phase and the global context feature maps of the subsequent phase in n stages into the decoder for decoding respectively to obtain the change detection masks of the remote sensing images of the previous and subsequent phases.
[0014] In one embodiment, the step of using the remote sensing image of the previous phase and the remote sensing image of the subsequent phase as the inputs of the encoder, and through the backbone network of the pre-trained basic Vision Transformer model of the encoder and two temporal feature adapter paths to perform bidirectional dynamic fusion feature extraction processing in n stages to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase in n stages includes:
[0015] Input the remote sensing image of the previous phase into multiple convolutional layers in the first temporal feature adapter path of the encoder for processing to obtain previous-phase feature maps with different resolutions;
[0016] Input the remote sensing image of the subsequent phase into multiple convolutional layers in the second temporal feature adapter path of the encoder for processing to obtain subsequent-phase feature maps with different resolutions;
[0017] Input the remote sensing image of the previous phase and the remote sensing image of the subsequent phase into the patch embedding layer of the backbone network of the pre-trained basic Vision Transformer model of the encoder for processing to obtain a previous-phase high-dimensional vector sequence and a subsequent-phase high-dimensional vector sequence;
[0018] Based on the pre-temporal feature maps with different resolutions, the post-temporal feature maps with different resolutions, the pre-temporal high-dimensional vector sequence, and the post-temporal high-dimensional vector sequence, perform bidirectional dynamic fusion feature extraction processing in n stages through n block processing blocks and n bidirectional gating modules on two temporal feature adapter paths to obtain the pre-temporal global context feature maps and post-temporal global context feature maps in n stages.
[0019] In one embodiment, the processing method for each stage is the same, and the output of the previous stage is the input of the next stage;
[0020] The processing process of the first stage is as follows:
[0021] Based on the pre-temporal feature maps with different resolutions, the pre-temporal high-dimensional vector sequence, and the pre-temporal global context feature map of the first stage output by the first block processing block, perform processing through the first bidirectional gating module in the first temporal feature adapter path to obtain the pre-temporal fusion feature map of the first stage and the added pre-temporal fusion feature map. Among them, the pre-temporal fusion feature map of the first stage is input into the first block processing block, and the added pre-temporal fusion feature map of the first stage is input into the second bidirectional gating module in the first temporal feature adapter path;
[0022] Based on the post-temporal feature maps with different resolutions, the post-temporal high-dimensional vector sequence, and the post-temporal global context feature map of the first stage output by the first block processing block, perform processing through the first bidirectional gating module in the second temporal feature adapter path to obtain the post-temporal fusion feature map of the first stage and the added post-temporal fusion feature map. Among them, the post-temporal fusion feature map of the first stage is input into the first block processing block, and the added post-temporal fusion feature map of the first stage is input into the second bidirectional gating module in the second temporal feature adapter path;
[0023] Based on the pre-temporal high-dimensional vector sequence, the post-temporal high-dimensional vector sequence, the pre-temporal fusion feature map of the first stage, and the post-temporal fusion feature map of the first stage, perform processing through the first block processing block to obtain the pre-temporal global context feature map and post-temporal global context feature map of the first stage. Among them, the pre-temporal global context feature map of the first stage is input into the second block processing block and the first bidirectional gating module in the first temporal feature adapter path, and the post-temporal global context feature map of the first stage is input into the second block processing block and the first bidirectional gating module in the second temporal feature adapter path.
[0024] In one embodiment, the structures of each two-way gating module are the same. The two-way gating module includes a multi-receptive field feature pyramid module, a dynamic temporal perception enhancement module, a first two-way interaction module, and a second two-way interaction module;
[0025] Among them, the multi-receptive field feature pyramid module is used to process the input feature map based on receptive fields of different sizes to form a multi-scale feature map; the dynamic temporal perception enhancement module is used to process the multi-scale feature map output by the multi-receptive field feature pyramid module, construct temporal features within a single phase, and generate an enhanced multi-scale feature map; the first two-way interaction module is used to interact and fuse the output of the dynamic temporal perception enhancement module in the current stage with the feature map output by the block processing block in the previous stage and output the first fusion feature map in the current stage; the second two-way interaction module is used to interact and fuse the output of the dynamic temporal perception enhancement module in the current stage with the feature map output by the block processing block in the current stage and output the second fusion feature map in the current stage; perform an addition operation on the output of the dynamic temporal perception enhancement module in the current stage and the second fusion feature map in the current stage to obtain the added and fused feature map in the current stage.
[0026] In one embodiment, the multi-receptive field feature pyramid module reduces the dimension of the received feature map through the first linear projection layer, divides the dimension-reduced feature map along the channel dimension into multiple groups, each group corresponding to a receptive field of a different size. After each group passes through a depthwise separable convolution operation, it then passes through the second linear projection layer to restore the original dimension, forming a multi-scale feature map.
[0027] In one embodiment, the dynamic temporal perception enhancement module receives the multi-scale feature map output from the multi-receptive field feature pyramid module. Among them, the feature map of each scale in the multi-scale feature map corresponds to a different spatial resolution;
[0028] Extract global statistical information from the feature map of each scale respectively through global average pooling operation to generate statistical vectors at the channel level, input each of the statistical vectors into the first multi-layer perceptron, generate each modulation vector through non-linear transformation, and input each of the statistical vectors into the second multi-layer perceptron to generate each channel weight;
[0029] The feature map of each scale is respectively input into the corresponding depthwise separable convolution for processing to obtain each processed feature map;
[0030] Use each of the modulation vectors to perform per-channel scaling on the corresponding processed feature map to obtain each modulated feature map;
[0031] Perform channel re-weighting on each of the modulated feature maps according to each of the channel weights to obtain each re-weighted feature map;
[0032] Add the received multi-scale feature maps to the corresponding re-weighted feature maps through residual connections, and apply layer normalization operations to generate enhanced multi-scale feature maps.
[0033] In one embodiment, based on the previous-phase high-dimensional vector sequence and the subsequent-phase high-dimensional vector sequence, the previous-phase fusion feature map of the first stage and the subsequent-phase fusion feature map of the first stage, process through the first block processing block to obtain the previous-phase global context feature map and the subsequent-phase global context feature map of the first stage, including:
[0034] Perform an addition operation on the previous-phase high-dimensional vector sequence and the previous-phase fusion feature map of the first stage to obtain an added previous-phase feature map;
[0035] Perform an addition operation on the subsequent-phase high-dimensional vector sequence and the subsequent-phase fusion feature map of the first stage to obtain an added subsequent-phase feature map;
[0036] Process the added previous-phase feature map through the first block processing block to obtain the previous-phase global context feature map of the first stage;
[0037] Process the added subsequent-phase feature map through the first block processing block to obtain the subsequent-phase global context feature map of the first stage.
[0038] In one embodiment, input the previous-phase global context feature maps and the subsequent-phase global context feature maps of n stages into the decoder for decoding to obtain the change detection mask of the remote sensing images of the previous and subsequent phases, including:
[0039] Input the previous-phase global context feature maps and the subsequent-phase global context feature maps of n stages into the decoder respectively, perform pixel-level difference operations on the previous-phase global context feature maps and the subsequent-phase global context feature maps of the corresponding stages to generate respective difference feature maps; then aggregate the respective difference feature maps based on a lightweight multi-layer perceptron decoder to obtain the change detection mask of the remote sensing images of the previous and subsequent phases.
[0040] In one embodiment, n≥4.
[0041] A remote sensing image change detection system based on a two-stream temporal feature adapter, including an image acquisition module, an encoder, and a decoder, wherein:
[0042] The image acquisition module is used to acquire remote sensing images of the previous phase and remote sensing images of the subsequent phase;
[0043] The encoder is used to perform bidirectional dynamic fusion feature extraction processing on the input remote sensing image of the previous phase and the remote sensing image of the subsequent phase through the pre-trained basic vision Transformer model backbone network of the encoder and two temporal feature adapter paths for n stages, to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase for n stages. Among them, the pre-trained basic vision Transformer model backbone network includes a patch embedding layer and n sequentially connected block processing blocks. Each temporal feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules. Each block processing block corresponds to a bidirectional gating module on each temporal feature adapter path to perform bidirectional dynamic fusion feature extraction processing for one stage, to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase for this stage, where n is a positive integer;
[0044] The decoder is used to decode the input global context feature maps of the previous phase and the global context feature maps of the subsequent phase for n stages respectively, to obtain the change detection masks of the remote sensing images of the previous and subsequent phases.
[0045] The present invention also provides a computer storage medium, which stores a computer program executable by a processor, and this computer program executes the remote sensing image change detection method based on a dual-stream temporal feature adapter described in the above technical solution.
[0046] The beneficial effects produced by the present invention are as follows: The present invention processes the remote sensing images of the previous and subsequent phases respectively through a shared pre-trained basic vision Transformer model backbone network (ViT) and two independent temporal feature adapter paths for bidirectional interaction and fusion. And bidirectional gating modules corresponding to different stages of the backbone network are set in the temporal feature adapter, which can sense the internal temporal features of the phase (such as changes in lighting conditions and shooting conditions, etc.) while performing multi-scale feature extraction, so as to reduce the influence of non-semantic temporal variations on the detection results and further improve the robustness of feature representation; and the present invention effectively bridges the domain differences between natural images and remote sensing images through the architecture of the shared ViT combined with the dual-stream temporal feature adapter, more efficiently obtains multi-scale feature differences, improves the change detection accuracy and consistency of remote sensing images, and at the same time significantly reduces the computational overhead, and finally generates a change detection mask of high-resolution remote sensing images of the previous and subsequent phases.
[0047] Furthermore, the multi-receptive field feature pyramid module (MRFP) in the bidirectional gating module of the present invention has a series of linear projection layers and depthwise separable convolutions, with different receptive fields, which is used to enhance the feature extraction ability of the CNN branch and ensure that each path can be optimized according to the characteristics of its respective phase.
[0048] Furthermore, the bidirectional gating adapter of the present invention activates the CNN-Transformer bidirectional interaction module (CTI), which can achieve dynamic fusion between the CNN features sensitive to local details and the Transformer features with global semantic perception.
[0049] Furthermore, the lightweight multi-layer perceptron (MLP) decoder aggregates the difference feature maps to generate a high-resolution change detection mask, further improving the robustness and practicality of the system.
[0050] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 is a flowchart of the remote sensing image change detection method based on the dual-stream temporal feature adapter according to the embodiment of the present invention;
[0053] Figure 2 is a schematic structural diagram of the encoder and decoder of the remote sensing image change detection method system based on the dual-stream pyramid feature adapter according to the embodiment of the present invention;
[0054] Figure 3 is a schematic structural diagram of the multi-receptive field feature pyramid module according to the embodiment of the present invention;
[0055] Figure 4 is a schematic structural diagram of the dynamic temporal perception enhancement module according to the embodiment of the present invention;
[0056] Figure 5 is a schematic structural diagram of the bidirectional interaction module according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further elaborates on the present invention with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0058] It should be noted that the illustrations provided in the embodiments of the present invention only schematically illustrate the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0059] In the present invention, it should also be noted that when terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. appear, the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present application. In addition, when terms such as "first" and "second" appear, they are only used for descriptive and distinguishing purposes, and cannot be construed as indicating or implying relative importance.
[0060] In addition, it should also be noted that the features of various embodiments of the present invention can be partially or wholly combined or integrated, and as can be understood by those skilled in the art, they can interact and operate in different ways. Each embodiment can be implemented independently of each other or in an associated relationship.
[0061] In one embodiment, as Figure 1 and Figure 2 shown, a remote sensing image change detection method based on a two-stream temporal feature adapter is provided, including the following steps:
[0062] Step S1, obtain the remote sensing image of the previous time phase and the remote sensing image of the subsequent time phase.
[0063] Step S2, take the remote sensing image of the previous time phase and the remote sensing image of the subsequent time phase as the input of the encoder, and perform n-stage bidirectional dynamic fusion feature extraction processing through the pre-trained basic vision Transformer model backbone network of the encoder and two temporal feature adapter paths to obtain the global context feature maps of the previous time phase and the subsequent time phase in n stages. Among them, the pre-trained basic vision Transformer model backbone network includes a patch embedding layer and n sequentially connected block processing blocks. Each temporal feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules. Each block processing block corresponds to a bidirectional gating module on each temporal feature adapter path to perform one-stage bidirectional dynamic fusion feature extraction processing to obtain the global context feature maps of the previous time phase and the subsequent time phase in this stage. n is a positive integer.
[0064] Among them, the remote sensing images of the previous time phase and the remote sensing images of the later time phase are processed using the same pre-trained basic vision Transformer model backbone network.
[0065] Among them, at the beginning and end of each block processing stage, each bidirectional gating module in each path performs bidirectional dynamic feature fusion with the corresponding block processing block, and generates corresponding fused feature maps at each stage.
[0066] Among them, the bidirectional gating module is used to first perform pyramid multi-scale feature image processing on the input image, and then through dynamic modulation and channel re-weighting driven by global statistical information, construct temporal features within a single time phase, generate enhanced feature maps, and perform interactive fusion with the block processing blocks at the corresponding stages. Finally, the feature representations are unified through a multi-scale self-attention mechanism to generate the fused feature maps at the corresponding stages.
[0067] Among them, the multiple convolutional layers on the two time-phase feature adapter paths are independent of each other. The multiple convolutional layers are used to process the input remote sensing images to obtain feature maps with different resolutions. The feature maps with different resolutions overlap in a pyramid shape. The resolutions of these feature maps with different resolutions can be 1 / 8, 1 / 16, and 1 / 32 of the original image. Each feature map contains a D-dimensional feature representation (that is, the data points in each feature map are represented by vectors composed of D numerical values).
[0068] Step S3: Input the global context feature maps of the previous time phase and the global context feature maps of the later time phase of n stages into the decoder for decoding respectively to obtain the change detection masks of the remote sensing images of the previous and later time phases.
[0069] As Figure 2 shown, in one embodiment, the remote sensing images of the previous time phase and the remote sensing images of the later time phase are used as the input of the encoder, and through the pre-trained basic vision Transformer model backbone network of the encoder and two time-phase feature adapter paths, n-stage bidirectional dynamic fusion feature extraction processing is performed to obtain the global context feature maps of the previous time phase and the global context feature maps of the later time phase of n stages, including:
[0070] The remote sensing images of the previous phase are processed through multiple convolutional layers in the first temporal feature adapter path of the encoder to obtain previous-phase feature maps of different resolutions; the remote sensing images of the subsequent phase are processed through multiple convolutional layers in the second temporal feature adapter path of the encoder to obtain subsequent-phase feature maps of different resolutions; the remote sensing images of the previous phase and the remote sensing images of the subsequent phase are processed through the patch embedding layer of the pre-trained basic vision Transformer model backbone network of the encoder to obtain a previous-phase high-dimensional vector sequence and a subsequent-phase high-dimensional vector sequence; based on the previous-phase feature maps of different resolutions, the subsequent-phase feature maps of different resolutions, the previous-phase high-dimensional vector sequence, and the subsequent-phase high-dimensional vector sequence, n-stage bidirectional dynamic fusion feature extraction processing is performed through n block processing blocks and n bidirectional gating modules on two temporal feature adapter paths to obtain previous-phase global context feature maps and subsequent-phase global context feature maps of n stages.
[0071] Among them, the network structures in the first temporal feature adapter path and the second temporal feature adapter path are the same. The difference is that each path processes the remote sensing images of one temporal phase, and then outputs the feature maps corresponding to the temporal phase.
[0072] As Figure 2 shown, in one embodiment, the processing method for each stage is the same, and the output of the previous stage is the input of the next stage; the processing process of the first stage is:
[0073] Based on the pre-phase feature maps with different resolutions, the pre-phase high-dimensional vector sequences, and the pre-phase global context feature maps of the first stage output by the first block processing block, they are processed through the first bidirectional gating module in the first phase feature adapter path to obtain the pre-phase fusion feature maps of the first stage and the pre-phase fusion feature maps after addition. Among them, the pre-phase fusion feature maps of the first stage are input into the first block processing block, and the pre-phase fusion feature maps after addition of the first stage are input into the second bidirectional gating module in the first phase feature adapter path; Based on the post-phase feature maps with different resolutions, the post-phase high-dimensional vector sequences, and the post-phase global context feature maps of the first stage output by the first block processing block, they are processed through the first bidirectional gating module in the second phase feature adapter path to obtain the post-phase fusion feature maps of the first stage and the post-phase fusion feature maps after addition. Among them, the post-phase fusion feature maps of the first stage are input into the first block processing block, and the post-phase fusion feature maps after addition of the first stage are input into the second bidirectional gating module in the second phase feature adapter path; Based on the pre-phase high-dimensional vector sequences and the post-phase high-dimensional vector sequences, the pre-phase fusion feature maps of the first stage and the post-phase fusion feature maps of the first stage, they are processed through the first block processing block to obtain the pre-phase global context feature maps and the post-phase global context feature maps of the first stage. Among them, the pre-phase global context feature maps of the first stage are input into the second block processing block and the first bidirectional gating module in the first phase feature adapter path, and the post-phase global context feature maps of the first stage are input into the second block processing block and the first bidirectional gating module in the second phase feature adapter path.
[0074] It should be understood that the processing method of each stage is the same, only the input of the first stage is different from that of other stages. The input of the first stage is the output of the patch embedding layer and multiple convolutional layers in each phase feature adapter path.
[0075] Such as Figure 2As shown, in one embodiment, the structures of each bidirectional gating module are the same. The bidirectional gating module includes a multi-receptive field feature pyramid module, a dynamic temporal perception enhancement module, a first bidirectional interaction module, and a second bidirectional interaction module. Among them, the multi-receptive field feature pyramid module is used to process the input feature map based on receptive fields of different sizes to form a multi-scale feature map. The dynamic temporal perception enhancement module is used to process the multi-scale feature map output by the multi-receptive field feature pyramid module, construct the temporal features within a single phase, and generate an enhanced multi-scale feature map. The first bidirectional interaction module is used to interactively fuse the output of the dynamic temporal perception enhancement module at the current stage with the feature map output by the block processing block at the previous stage and output the first fusion feature map at the current stage. The second bidirectional interaction module is used to interactively fuse the output of the dynamic temporal perception enhancement module at the current stage with the feature map output by the block processing block at the current stage and output the second fusion feature map at the current stage. An addition operation is performed on the output of the dynamic temporal perception enhancement module at the current stage and the second fusion feature map at the current stage to obtain the fused feature map after addition at the current stage.
[0076] Among them, if the current stage is the first stage, the first bidirectional interaction module is used to interactively fuse the output of the dynamic temporal perception enhancement module at the current stage with the feature maps output by multiple convolutional layers on this path and output the first fusion feature map at the current stage.
[0077] Among them, if the path where the bidirectional gating module is located is to process the remote sensing image of the previous phase, the first fusion feature map at the current stage is the fusion feature map of the previous phase input into the block processing block at the current stage. If the path where the bidirectional gating module is located is to process the remote sensing image of the later phase, the first fusion feature map at the current stage is the fusion feature map of the later phase input into the block processing block at the current stage.
[0078] It should be understood that if the path where the bidirectional gating module is located is to process the remote sensing image of the previous phase, then all types of feature maps processed or output in the bidirectional gating module are of the previous phase. Correspondingly, if the path where the bidirectional gating module is located is to process the remote sensing image of the later phase, then all types of feature maps processed or output in the bidirectional gating module are of the later phase.
[0079] As Figure 3 shown, in one embodiment, the multi-receptive field feature pyramid module reduces the dimension of the received feature map through the first linear projection layer, divides the feature map after dimension reduction along the channel dimension into multiple groups, each group corresponding to a receptive field of a different size. After each group passes through a depthwise separable convolution operation, it then restores the original dimension through the second linear projection layer to form a multi-scale feature map.
[0080] It should be understood that if the path where the bidirectional gating module is located is the remotely sensed image of the pre-processing phase, all types of feature maps processed or output by the multi-receptive field feature pyramid module in the bidirectional gating module are of the pre-phase; correspondingly, if the path where the bidirectional gating module is located is the remotely sensed image of the post-processing phase, all types of feature maps processed or output by the multi-receptive field feature pyramid module in the bidirectional gating module are of the post-phase.
[0081] As Figure 4 shown, in one embodiment, the dynamic temporal perception enhancement module receives the multi-scale feature maps output from the multi-receptive field feature pyramid module, where each scale of the feature maps of the multi-scale feature maps corresponds to a different spatial resolution; global statistical information is extracted from each scale of the feature maps respectively through global average pooling operations to generate respective statistical vectors at the channel level, the respective statistical vectors are input into the first multi-layer perceptron, and respective modulation vectors are generated through non-linear transformation, the respective statistical vectors are input into the second multi-layer perceptron to generate respective channel weights; each scale of the feature maps is respectively input into the corresponding depthwise separable convolution for processing to obtain respective processed feature maps; the respective processed feature maps are scaled channel by channel by using the respective modulation vectors to obtain respective modulated feature maps; the respective modulated feature maps are re-weighted at the channel level according to the respective channel weights to obtain respective re-weighted feature maps; the received multi-scale feature maps are added to the corresponding re-weighted feature maps through residual connections, and layer normalization operations are applied to generate enhanced multi-scale feature maps.
[0082] It should be understood that if the path where the bidirectional gating module is located is the remotely sensed image of the pre-processing phase, all types of feature maps processed or output by the dynamic temporal perception enhancement module in the bidirectional gating module are of the pre-phase; correspondingly, if the path where the bidirectional gating module is located is the remotely sensed image of the post-processing phase, all types of feature maps processed or output by the dynamic temporal perception enhancement module in the bidirectional gating module are of the post-phase.
[0083] In one embodiment, the first bidirectional interaction module is specifically configured to, at the beginning of each stage, fuse the output of the dynamic temporal perception enhancement module with the output of the block processing block of the previous stage, and then apply the multi-scale deformable attention mechanism to uniformly represent the fused features, align the feature sizes of each scale and transfer them to the block processing block of the current stage through interpolation operations.
[0084] In one embodiment, the second bidirectional interaction module is specifically configured to, at the end of each stage, fuse the output of the dynamic temporal perception enhancement module with the output of the block processing block of the current stage, then apply the multi-scale deformable attention mechanism to the unified fused feature representation, align the feature sizes of each scale, and after interpolation operations, generate the fused feature maps of different scales corresponding to the current stage, add them to the output of the dynamic temporal perception enhancement module of the current stage, and then transfer the result to the bidirectional gating module of the next stage.
[0085] In one embodiment, based on the high-dimensional vector sequence of the previous time phase and the high-dimensional vector sequence of the next time phase, the fused feature map of the previous time phase of the first stage and the fused feature map of the next time phase of the first stage, they are processed by the first block processing block to obtain the global context feature map of the previous time phase and the global context feature map of the next time phase of the first stage, including:
[0086] Perform an addition operation on the high-dimensional vector sequence of the previous time phase and the fused feature map of the previous time phase of the first stage to obtain the added previous time phase feature map; perform an addition operation on the high-dimensional vector sequence of the next time phase and the fused feature map of the next time phase of the first stage to obtain the added next time phase feature map; process the added previous time phase feature map by the first block processing block to obtain the global context feature map of the previous time phase of the first stage; process the added next time phase feature map by the first block processing block to obtain the global context feature map of the next time phase of the first stage.
[0087] As Figure 2 shown, in one embodiment, the global context feature maps of the previous time phase and the global context feature maps of the next time phase of n stages are respectively input into the decoder for decoding to obtain the change detection mask of the remote sensing images of the previous and next time phases, including:
[0088] The global context feature maps of the previous time phase and the global context feature maps of the next time phase of n stages are respectively input into the decoder, perform pixel-level difference operations on the global context feature maps of the previous time phase and the global context feature maps of the next time phase of the corresponding stages to generate the respective difference feature maps; then aggregate the respective difference feature maps based on the lightweight multi-layer perceptron decoder to obtain the change detection mask of the remote sensing images of the previous and next time phases.
[0089] The above remote sensing image change detection method based on a dual-stream temporal feature adapter processes the remote sensing images of the front and back phases respectively through a shared pre-trained basic vision Transformer model backbone network (ViT) and two independent temporal feature adapter paths, conducts bidirectional interactive fusion, and a bidirectional gating module corresponding to different stages of the backbone network is set in the temporal feature adapter, while performing multi-scale feature extraction, it also perceives the internal temporal features of the phase (such as changes in lighting conditions and shooting conditions, etc.), thereby reducing the impact of non-semantic temporal variations on the detection results and further enhancing the robustness of feature representation; and the present invention effectively bridges the domain differences between natural images and remote sensing images through the architecture of a shared ViT combined with a dual-stream temporal feature adapter, more efficiently obtains multi-scale feature differences, improves the change detection accuracy and consistency of remote sensing images, while significantly reducing the computational overhead, and finally generates a change detection mask for high-resolution remote sensing images of the front and back phases.
[0090] In one embodiment, as Figure 2 shown, each independent temporal feature adapter path of the encoder interacts with the intermediate shared basic vision Transformer model backbone network at each stage. The parameters of the shared pre-trained basic vision Transformer model backbone network (Vision Transformer, abbreviated as ViT) remain frozen during use, avoiding performance degradation caused by re-training or fine-tuning, and using its powerful semantic prior to enhance the generalization ability of the model. The basic vision Transformer model backbone network includes n block processing blocks, and the size of n can be selected according to actual needs. Generally, n≥4. In this embodiment, n = 4 is selected, that is, the basic vision Transformer model backbone network contains four block processing blocks, and the backbone network is divided into four processing stages.
[0091] Correspondingly, each temporal feature adapter of each path also contains four bidirectional gating modules, which interact and fuse with the block processing blocks at the beginning and end of each stage, continuously updating the input and output of the block processing blocks at each stage, and also continuously updating the input and output of the bidirectional gating modules at each stage. And at the end of each stage, fusion feature maps of different scales at the corresponding level are generated, which fuse the different scale features of the remote sensing images of the front and back phases. The Siamese adapter architecture designed by the present invention provides independent feature extraction and fusion processes for the remote sensing images of the front and back phases, ensuring that each path can be optimized according to the characteristics of its respective phase, thereby improving the overall performance of the system.
[0092] In addition, as Figure 2The Patch Embedding layer in it is a step in ViT to convert the input image into sequence data suitable for Transformer processing. Specifically, the original image (i.e., the remote sensing image of the previous time phase or the remote sensing image of the later time phase) is first segmented into multiple small patches of a fixed size; then, each patch is flattened into a one-dimensional vector and mapped to the hidden dimension of the encoder through a linear projection layer to generate the feature representation of each patch; in order to retain the position information of these patches in the original image, position encoding is also added to each patch. After Patch Embedding, the image is thus converted into a sequence composed of multiple high-dimensional vectors with spatial information encoding, that is, the high-dimensional vector sequence of the previous time phase or the high-dimensional vector sequence of the later time phase, so that the subsequent Transformer architecture can effectively process this information.
[0093] Among them, the bidirectional gating module includes a multi-receptive field feature pyramid module MRFP, a dynamic temporal perception enhancement module DTAEM, and two cross-temporal interaction modules CTI. Among them, MRFP is used to form multi-scale feature maps based on receptive fields of different sizes; DTAEM is used to process the multi-scale feature maps output by MRFP, construct the temporal features within a single time phase, and generate enhanced multi-scale feature maps; CTI is used to interact and fuse the output of DTAEM with the block processing blocks of the corresponding stage and output the fused feature maps of different stages.
[0094] During specific processing, the remote sensing images of the previous and later time phases are simultaneously input into the shared backbone network of the basic vision Transformer model and their respective independent time-phase feature adapter paths. In the time-phase feature adapter path of each time-phase remote sensing image, it first passes through a series of independent convolutional layers for processing to generate initial feature maps (i.e., the previous time-phase feature maps of different resolutions or the later time-phase feature maps of different resolutions); the initial feature maps then pass through the multi-receptive field feature pyramid module (MRFP) to generate feature maps of different scales, such as Figure 3As shown in the figure, the multi-receptive field feature pyramid module includes a series of linear projection layers and depthwise separable convolutions, with different receptive fields. The input feature map is dimensionally reduced through the linear projection layer, and the dimensionally reduced feature map is divided into multiple groups along the channel dimension, with each group corresponding to a receptive field of a different size. Each group undergoes a depthwise separable convolution operation to expand the receptive field and enhance the long-range modeling ability. The processed feature map passes through the linear projection layer again to restore the original dimension, forming a multi-scale feature map. Among them, by using depthwise separable convolution (Depthwise Separable Convolution) and the multi-receptive field feature pyramid module (MRFP), the receptive field of each convolutional layer can be expanded, enabling the encoder to consider a wider range of context information. Depthwise separable convolution is an efficient convolution method that first independently applies spatial convolution to each channel and then integrates cross-channel information through 1x1 convolution. This method not only reduces the computational amount and the number of parameters but also allows the model to capture more complex patterns and long-range dependencies by combining multiple convolutional kernels with different receptive fields.
[0095] As Figure 4 shown, the dynamic temporal awareness enhancement module (DTAEM) models the temporal features within a single phase through a dynamic modulation and channel reweighting mechanism driven by global statistical information, enhancing the robustness of feature representation. Among them, by constructing the temporal features within a single phase (such as changes in lighting conditions), the robustness of the multi-scale feature map to phase features can be enhanced, thereby reducing the interference of non-semantic temporal variations on change detection and improving the adaptability and discriminability of feature representation. The specific process includes: 1) DTAEM receives the multi-scale feature maps from the multi-receptive field feature pyramid module (MRFP), where each feature map corresponds to a different spatial resolution; 2) Global statistical data extraction: For each scale of the feature map, first, global statistical information is extracted through global average pooling operation to generate a channel-level statistical vector to characterize the overall distribution characteristics of the feature map; subsequently, the statistical vector is input into a multi-layer perceptron (MLP), and a modulation vector is generated through non-linear transformation. Based on the same statistical vector, a channel weight is generated through another multi-layer perceptron; 3) Using this modulation vector to scale the output of the depthwise separable convolution channel by channel, dynamically modulating the convolution response to adapt to temporal characteristics; 4) Channel recalibration: Using the channel weight to perform channel reweighting on the modulated feature map, highlighting the feature channels related to change detection and suppressing the influence of non-semantic temporal variations; 5) Finally, the multi-scale feature maps from the multi-receptive field feature pyramid module are added to the reweighted feature map through a residual connection, and a layer normalization operation is applied to generate an enhanced multi-scale feature map, providing a robust feature representation for subsequent feature fusion.
[0096] As Figure 2 、 5As shown, the CNN-Transformer based Bidirectional Interaction Module (CTI) is mainly used to achieve the dynamic fusion between the CNN features sensitive to local details and the Transformer features with global semantic perception. It can dynamically fuse local and global features at different scales, improve the expression ability and detection accuracy of the system, and unify the representation differences between different modalities through direct addition and multi-scale self-attention mechanism, enhancing the model's performance in complex scenarios. The specific process includes: 1) At the beginning of each stage, directly add and fuse the different-scale features from the temporal feature adapter and the ViT branch features (i.e., the output features of each block processing block); 2) Apply the multi-scale deformable attention mechanism to further unify the feature representation, align the feature sizes of each scale and transfer them to the next layer of the ViT branch (the ViT branch refers to the block processing block. When n = 4, the four block processing blocks correspond to Figure 2 ViT-S1, ViT-S2, ViT-S3, and ViT-S4 in
[0097] ); 3) Generate feature maps of different scales at the corresponding levels and transfer them to the next stage for further processing.
[0098] Furthermore, the decoder can aggregate the difference feature maps through a lightweight multi-layer perceptron (MLP) decoder to generate a high-resolution change detection mask, further improving the robustness and practicality of the system. The decoding process of the decoder mainly includes the following steps:
[0099] 1) Perform pixel-level difference operations on the pre-temporal global context feature map and the post-temporal global context feature map of the corresponding stage to generate the difference feature map of the corresponding stage;
[0100] In one embodiment, it is based on SAM-ViT and n = 4. SAM-ViT refers to the backbone network of a pre-trained Vision Transformer (ViT) model based on the Segment Anything Model (SAM). In SAM, ViT is used as its backbone network, which means that ViT undertakes the main task of processing and understanding the input image in SAM. Specifically, SAM uses the pre-trained ViT to extract the feature representation of the image. These features are efficient and information-rich image representations, which are crucial for performing subsequent segmentation tasks. The ViT backbone network of SAM has been pre-trained on a massive dataset including SA-1B, which contains more than 1 billion mask annotations. This large-scale data enables the model to learn diverse segmentation patterns. In addition, combined with the contrastive learning method, ViT can further enhance the sensitivity and discrimination ability to object boundaries. Based on this SAM-ViT, this embodiment uses the SAM-CNN encoder for encoding, that is, to implement step S2. The SAM-CNN encoder is not an existing independent model, but is used to describe the encoder structure formed by combining SAM-ViT mentioned in this embodiment with a CNN (Convolutional Neural Network) branch. The global features are extracted through the shared SAM-ViT, and at the same time, two independent temporal feature adapter paths are used to process the pre-temporal remote sensing image and the post-temporal remote sensing image respectively, generating multi-scale local features and fusing them with the global features. The SAM-CNN encoder consists of two parts: (a) a shared SAM-ViT backbone network with frozen parameters, (b) a temporal feature adapter. First, for the SAM-ViT branch, the pre-temporal remote sensing image and the post-temporal remote sensing image with the shape of H×W×3 are respectively input into the patch embedding layer to obtain a feature representation with the resolution reduced to 1 / 16 of the original image. At the same time, for the temporal feature adapter branch, the pre-temporal remote sensing image and the post-temporal remote sensing image are processed through a series of independent convolutional layers on the corresponding paths, generating feature pyramids with resolutions of 1 / 8, 1 / 16, and 1 / 32 respectively (i.e., Figure 3Among C3, C4, and C5), each feature map of the feature pyramid contains a D-dimensional feature representation. Next, the features on the two paths undergo 4 stages of feature interaction. In each stage, first, the MRFP enhances the feature pyramid to obtain a multi-scale feature map to capture richer spatial information. Subsequently, the DTAEM performs global statistics-driven dynamic modulation and channel re-weighting on the multi-scale feature map output by the MRFP, models the temporal features within a single phase, and generates an enhanced feature map. Then, two CTIs interact bidirectionally with the features of SAM-ViT to obtain multi-scale features with rich semantic information. It should be noted that one CTI runs at the beginning of each stage, and the other CTI runs at the end of each stage to ensure that the features can be fully and effectively interacted. After 4 stages of feature interaction, the features of the two phases in each stage are subsequently input into the decoder for feature difference calculation.
[0101] Among them, to enhance the encoder's ability to capture spatial information at different scales, the MRFP is introduced. As Figure 3 shown, the MRFP consists of a series of linear projection layers and depthwise separable convolutions with different receptive fields. Specifically, the input includes feature maps C3, C4, C5 with different resolutions. First, dimensionality reduction is performed through the linear projection layer, and then a reduced-dimensional feature representation C is obtained in , where R represents real numbers, and H and W are used to represent the spatial dimensions. Subsequently, these features are divided into M groups along the channel dimension, and each group corresponds to a convolutional layer with a different kernel size, thereby expanding the receptive field and enhancing the model's long-distance modeling ability for CNN features. Then, the processed features are connected through another linear projection layer and restored to the original dimension, and a multi-scale feature map is output. This process can be mathematically expressed as:
[0102] ;
[0103] where FC(•) represents the linear projection operation, DWConv(•) represents a group of depth convolutions with different kernel sizes, and represents the output of the MRFP. The MRFP not only provides rich multi-scale information but also enhances the model's ability to capture local details and global backgrounds of images by combining different receptive fields, thereby significantly improving the model's performance in dense prediction tasks.
[0104] In remote sensing change detection tasks, due to factors such as lighting conditions, seasonal changes, or atmospheric effects, images taken at different time stages often exhibit significant feature changes. Although these attributes of specific time phases do not directly show semantic changes on the ground surface, they introduce interference during the feature extraction process, thus affecting the accuracy of change detection. To solve this problem, the present invention specifically designs DTAEM, which can dynamically adjust the feature representations corresponding to each time phase processed by each adapter to ensure their consistency with specific time backgrounds (such as changes in lighting intensity or vegetation conditions in different seasons).
[0105] As Figure 4 shown, DTAEM processes the multi-scale feature maps generated by MRFP. Among them, the multi-scale feature maps , i = each feature map in 3, 4, 5 corresponds to a different spatial resolution. Among them, B represents the batch size, and H i represents the height of i represents the width of. D is used to represent the channel dimension. This DTAEM processes these multi-scale feature maps through a series of operations, which can dynamically adjust the convolutional response and recalibrate the channel importance according to the global feature statistics to ensure consistency with the specific time features of each image. The workflow first independently analyzes the feature maps of each scale, using depthwise separable convolution, temporal modulation, and channel reweighting to generate enhanced feature representations that are robust to temporal artifacts.
[0106] At the beginning of the processing, first perform depthwise separable convolution on each feature map to effectively capture local spatial patterns. To make this operation temporally adaptive, global statistical information is extracted through global average pooling operation The mathematical formula is as follows:
[0107] ;
[0108] Among them, i = 3, 4, 5, represents the output of the global average pooling operation, which encompasses the overall feature distribution across the spatial dimensions. Then the statistical vector is input into a multi-layer perceptron (MLP) to generate a modulation vector:
[0109] ;
[0110] Among them, i = 3, 4, 5, and σ represents the sigmoid activation function. The modulation vector scales the output of the depthwise separable convolution DWConv. According to Dynamically adjust the feature response of the encoded temporal context. The mathematical expression for this step is:
[0111] ,
[0112] ;
[0113] where, i = 3, 4, 5, expand(•) means broadcasting M i to match the spatial dimension, represents the result after per-channel scaling of the output of the depthwise separable convolution using the modulation vector, DWConv(•) represents the depthwise separable convolution using a 3×3 kernel, represents the output of the depthwise separable convolution, ⊙ represents element-wise multiplication with broadcast to match the spatial dimension. This modulation mechanism enables the model to selectively enhance or suppress feature channels based on the correlation between feature channels and temporal changes, thus effectively distinguishing real changes from temporal artifacts.
[0114] After modulation, DTAEM uses an adaptive reweighting mechanism to recalibrate channel importance, thereby further refining the feature map. Using the same global statistics , the second MLP calculates the per-channel scaling weights . These weights are applied to the modulated features to emphasize the channels most relevant to change detection while suppressing the channels dominated by noise or irrelevant temporal changes:
[0115] ;
[0116] where, i = 3, 4, 5, represents the reweighted feature map.
[0117] This reweighting step ensures the effective allocation of computational resources and enhances the model's robustness to non-semantic temporal noise. To maintain the integrity of the original features and stabilize training, a residual connection combines the input feature map with the reweighted output . Then the resulting features are layer-normalized to maintain the consistency of the feature scale throughout the network:
[0118] ;
[0119] where, i = 3, 4, 5, LayerNorm(•) represents layer normalization, Represents the enhanced multi-scale feature map output by DTAEM.
[0120] This workflow is repeatedly applied at all scales to ensure that the dynamic time enhancement is consistent across different spatial resolutions. As Figure 5 shown, the enhanced multi-scale feature map Subsequently, it is sent to the CTI module for two-way integration with the global context feature map from SAM-ViT. This method is based on globally feature statistic-driven modulation and reweighting to ensure that the resulting feature representation is both robust and semantically focused, thus supporting accurate change detection in complex remote sensing environments.
[0121] To effectively integrate the features of the ViT and CNN branches (i.e., the temporal feature adapter path) at each stage, CTI is also introduced. At the beginning of each stage, the multi-scale features of the temporal feature enhancement obtained from DTAEM are fused with the features from the ViT branch directly add F4 to X to generate a new set of features F′={F3,F4 ,F5}, and this process can be expressed as: Subsequently, self-attention calculation is performed on F′ to unify the representational differences of different modalities:
[0122] , where norm(•) represents layer normalization, Attention(•) represents multi-scale deformable attention, FFN(•) represents the feed-forward network, and O represents the output of CTI. Then, the bilinear interpolation method is used to align the sizes of the feature map O3 and the feature map O5 output by CTI with the size of the feature map O4, and the results are added to the input of the corresponding ViT branch to form the input of this ViT branch: ;
[0123] where, is the updated feature of the ViT branch, α is a learnable variable, represents the O3 feature map aligned with the size of O4, represents the O5 feature map aligned with the size of O4.
[0124] At the end of each stage, features F′={F3,F4 ,F5} and O={O3,O4,O5} are obtained through the same steps, and the features of the corresponding scales are directly added to enhance the feature representation ability of the CNN branch at each scale. Finally, these fused features are passed to the CNN branch of the next stage to ensure that the model can more accurately capture the multi-level information in the image.
[0125] The decoder used in this embodiment is an MLP decoder. The task of the MLP decoder is to aggregate multi-level difference feature maps from difference calculation to generate a high-resolution change detection mask. Using a sequence of multi-layer perceptron (MLP) layers and strategic upsampling operations, this lightweight decoder can effectively synthesize multi-scale difference features into a coherent representation, thereby accurately identifying the changed regions in the remote sensing image pair.
[0126] The decoder receives the pre-temporal global context feature maps of n stages output by the encoder and the post-temporal global context feature maps , where i = 1, 2, 3, 4. The pre-temporal global context feature maps and the post-temporal global context feature maps of the same stage perform pixel-level difference operations to generate the difference feature maps of this stage as input to the MLP decoder.
[0127] The MLP decoder receives a set of four difference feature maps as input. When i = 1, 2, 3, 4, the difference feature maps are denoted as , and each difference feature map is obtained through difference calculation. The spatial resolution of these difference feature maps is , and the channel dimension of the difference feature map is D i . Due to the differences in scale and dimension, a structured process is required to coordinate these features for effective fusion and final prediction.
[0128] To initiate this process, each difference feature map is transformed through an MLP layer to standardize its channel dimension to a unified embedding size C ebd . This channel unification step ensures the compatibility of each layer, and its mathematical definition is as follows:
[0129] ;
[0130] where, represents the channel dimension of the difference feature map of the corresponding stage, and represents the i-th layer difference feature map after being processed by the MLP layer.
[0131] Subsequently, the spatial resolution of these feature maps is adjusted to a common scale ( ) through bilinear interpolation. This upsampling operation prepares for connecting the features, and its expression is:
[0132] ;
[0133] where, Denote the i-th layer of differential feature map after upsampling, Upsample(•) represents upsampling, and bilinear indicates using the bilinear interpolation method for upsampling.
[0134] After normalizing the channel and spatial dimensions, the feature map after upsampling is concatenated along the channel axis to obtain a composite feature map with a channel dimension of 4C ebd Then, an additional MLP layer is used to process this combined representation, fusing the multi-level differential information into a unified fused feature map F fused , which is expressed as follows:
[0135] ;
[0136] where Cat(•) represents concatenation.
[0137] To generate the final change detection mask at the original image resolution, the fused feature map F fused will be upsampled through a 2D transposed convolutional layer with a stride of 4 and a kernel size of 3 to restore the spatial dimension to H×W. These operations are defined as follows:
[0138] ,
[0139] ;
[0140] where, denotes the feature map after the transposed convolution operation, ConvTranspose2D represents the 2D transposed convolutional layer, denotes the class parameter, i.e., the number of classes of the final output, S represents the stride, which is 4 here, K represents the size of the convolutional kernel, which is 3×3 here, and CM represents the predicted change mask, with a dimension of H×W×N cls .
[0141] The streamlined architecture of the MLP decoder balances computational efficiency and the ability to effectively integrate multi-scale differential features. By using MLP layers for feature transformation and fusion, and supplemented with precise upsampling, this module can ensure the generation of accurate and spatially consistent change maps.
[0142] In one embodiment, a remote sensing image change detection system based on a two-stream temporal feature adapter is provided, including an image acquisition module, an encoder, and a decoder, where:
[0143] The image acquisition module is used to acquire the remote sensing image of the previous time phase and the remote sensing image of the later time phase;
[0144] The encoder is used to perform bidirectional dynamic fusion feature extraction processing on the input remote sensing images of the previous phase and the remote sensing images of the subsequent phase through the pre-trained basic vision Transformer model backbone network of the encoder and two temporal feature adapter paths for n stages, to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase for n stages. Among them, the pre-trained basic vision Transformer model backbone network includes a patch embedding layer and n sequentially connected block processing blocks. Each temporal feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules. Each block processing block corresponds to a bidirectional gating module on each temporal feature adapter path to perform bidirectional dynamic fusion feature extraction processing for one stage, to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase for this stage. n is a positive integer;
[0145] The decoder is used to decode the input global context feature maps of the previous phase and the global context feature maps of the subsequent phase for n stages respectively, to obtain the change detection masks of the remote sensing images of the previous and subsequent phases.
[0146] Each module is mainly used to implement each step of the above method embodiment, which will not be elaborated here one by one.
[0147] This application also provides a computer-readable storage medium, such as flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, server, App application mall, etc., on which a computer program is stored. When the program is executed by a processor, the corresponding functions are implemented. When the computer-readable storage medium of this embodiment is executed by a processor, it implements the remote sensing image change detection method based on a dual-stream temporal feature adapter in the method embodiment.
[0148] In summary, unlike the traditional twin network architecture that relies on high computational cost and low feature alignment efficiency, the present invention proposes a shared pre-trained basic visual Transformer (ViT) model backbone network combined with a dual-stream pyramid feature adapter architecture, which effectively bridges the domain differences between natural images and remote sensing images, while significantly reducing computational overhead. Through two independent lightweight pyramid feature adapter paths, the remote sensing images of the previous and next phases are processed respectively, and the dynamic fusion between the CNN features sensitive to local details and the Transformer features of global semantic perception is realized, thereby enhancing the cross-modal feature enhancement effect. In addition, a dynamic temporal awareness enhancement module (DTAEM) is specially designed, which models the temporal features within a single phase through dynamic modulation and channel reweighting driven by global statistical information, thereby reducing the impact of non-semantic temporal variation on the detection results and further improving the robustness of feature representation. It not only provides a resource-efficient solution, but also significantly improves the accuracy and interpretability of remote sensing image change detection in complex scenes.
[0149] It should be pointed out that, according to the needs of implementation, the various steps / components described in this application can be split into more steps / components, and two or more steps / components or partial operations of steps / components can be combined into new steps / components to achieve the purpose of the present invention.
[0150] The order of execution of each step in the above embodiment does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0151] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all these improvements and changes should fall within the scope of protection of the appended claims of the present invention.
Claims
1. A remote sensing image change detection method based on a dual-stream temporal feature adapter, characterized in that The remote sensing image change detection method based on the two-stream temporal feature adapter includes: S1. Obtain the remote sensing image of the previous phase and the remote sensing image of the subsequent phase; S2. Take the remote sensing image of the previous phase and the remote sensing image of the subsequent phase as the input of the encoder, and perform bidirectional dynamic fusion feature extraction processing in n stages through the pre-trained basic vision Transformer model backbone network of the encoder and two temporal feature adapter paths to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase in n stages. Among them, the pre-trained basic vision Transformer model backbone network includes a patch embedding layer and n sequentially connected block processing blocks. Each temporal feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules. Each block processing block corresponds to a bidirectional gating module on each temporal feature adapter path to perform bidirectional dynamic fusion feature extraction processing in one stage to obtain the global context feature map of the previous phase and the global context feature map of the subsequent phase in this stage. n is a positive integer; S3. Input the global context feature maps of the previous phase and the global context feature maps of the subsequent phase in n stages into the decoder for decoding respectively to obtain the change detection masks of the remote sensing images of the previous and subsequent phases; The step of taking the remote sensing image of the previous phase and the remote sensing image of the subsequent phase as the input of the encoder, and performing bidirectional dynamic fusion feature extraction processing in n stages through the pre-trained basic vision Transformer model backbone network of the encoder and two temporal feature adapter paths to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase in n stages includes: The remote sensing image of the previous phase is input into multiple convolutional layers in the first temporal feature adapter path of the encoder for processing to obtain previous-phase feature maps with different resolutions; The remote sensing image of the subsequent phase is input into multiple convolutional layers in the second temporal feature adapter path of the encoder for processing to obtain subsequent-phase feature maps with different resolutions; The remote sensing image of the previous phase and the remote sensing image of the subsequent phase are input into the patch embedding layer of the pre-trained basic vision Transformer model backbone network of the encoder for processing to obtain a previous-phase high-dimensional vector sequence and a subsequent-phase high-dimensional vector sequence; Based on the previous-phase feature maps with different resolutions, the subsequent-phase feature maps with different resolutions, the previous-phase high-dimensional vector sequence, and the subsequent-phase high-dimensional vector sequence, perform bidirectional dynamic fusion feature extraction processing in n stages through n block processing blocks and n bidirectional gating modules on two temporal feature adapter paths to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase in n stages.
2. The remote sensing image change detection method based on the two-stream temporal feature adapter according to claim 1, wherein, The processing method of each stage is the same, and the output of the previous stage is the input of the next stage; The One processing procedure of the first stage is as follows: Based on the pre-phase feature maps of different resolutions, the pre-phase high-dimensional vector sequence, and the pre-phase global context feature map of the first stage output by the first block processing block, it is processed through the first bidirectional gating module in the first phase feature adapter path to obtain the pre-phase fusion feature map of the first stage and the pre-phase fusion feature map after addition, wherein the pre-phase fusion feature map of the first stage is input into the first block processing block, and the pre-phase fusion feature map after addition of the first stage is input into the second bidirectional gating module in the first phase feature adapter path; Based on the post-phase feature maps of different resolutions, the post-phase high-dimensional vector sequence, and the post-phase global context feature map of the first stage output by the first block processing block, it is processed through the first bidirectional gating module in the second phase feature adapter path to obtain the post-phase fusion feature map of the first stage and the post-phase fusion feature map after addition, wherein the post-phase fusion feature map of the first stage is input into the first block processing block, and the post-phase fusion feature map after addition of the first stage is input into the second bidirectional gating module in the second phase feature adapter path; Based on the pre-phase high-dimensional vector sequence and the post-phase high-dimensional vector sequence, the pre-phase fusion feature map of the first stage and the post-phase fusion feature map of the first stage, it is processed through the first block processing block to obtain the pre-phase global context feature map and the post-phase global context feature map of the first stage, wherein the pre-phase global context feature map of the first stage is input into the second block processing block and the first bidirectional gating module in the first phase feature adapter path, and the post-phase global context feature map of the first stage is input into the second block processing block and the first bidirectional gating module in the second phase feature adapter path.
3. The remote sensing image change detection method based on a two-stream temporal feature adapter according to claim 2, wherein The structures of each bidirectional gating module are the same, and the bidirectional gating module includes a multi-receptive field feature pyramid module, a dynamic temporal perception enhancement module, a first bidirectional interaction module, and a second bidirectional interaction module; Among them, the multi-receptive field feature pyramid module is used to process the input feature map based on receptive fields of different sizes to form a multi-scale feature map; the dynamic temporal perception enhancement module is used to process the multi-scale feature map output by the multi-receptive field feature pyramid module to construct the temporal features within a single phase and generate an enhanced multi-scale feature map; the first bidirectional interaction module is used to interact and fuse the output of the dynamic temporal perception enhancement module in the current stage with the feature map output by the block processing block in the previous stage and output the first fusion feature map of the current stage; the second bidirectional interaction module is used to interact and fuse the output of the dynamic temporal perception enhancement module in the current stage with the feature map output by the block processing block in the current stage and output the second fusion feature map of the current stage; an addition operation is performed on the output of the dynamic temporal perception enhancement module in the current stage and the second fusion feature map of the current stage to obtain the fusion feature map after addition of the current stage.
4. The remote sensing image change detection method based on the dual-stream temporal feature adapter according to claim 3, wherein The multi-receptive field feature pyramid module reduces the dimension of the received feature map through the first linear projection layer, divides the dimension-reduced feature map into multiple groups along the channel dimension, each group corresponding to a receptive field of a different size. After each group passes through a depthwise separable convolution operation, it then passes through the second linear projection layer to restore the original dimension, forming a multi-scale feature map.
5. The remote sensing image change detection method based on a dual-stream temporal feature adapter according to claim 3, wherein The dynamic temporal perception enhancement module receives the multi-scale feature map output from the multi-receptive field feature pyramid module. Among them, the feature map of each scale in the multi-scale feature map corresponds to a different spatial resolution. Global average pooling operations are used to extract global statistical information from the feature map of each scale respectively, generating statistical vectors at the channel level. These statistical vectors are input into the first multi-layer perceptron, and modulation vectors are generated through non-linear transformation. These statistical vectors are input into the second multi-layer perceptron to generate channel weights. The feature map of each scale is respectively input into the corresponding depthwise separable convolution for processing to obtain processed feature maps. Each processed feature map is scaled channel by channel using the corresponding modulation vector to obtain modulated feature maps. Channel re-weighting is performed on each modulated feature map according to the channel weights to obtain re-weighted feature maps. The received multi-scale feature map is added to the corresponding re-weighted feature map through residual connection, and layer normalization operation is applied to generate enhanced multi-scale feature maps.
6. The remote sensing image change detection method based on a dual-stream temporal feature adapter according to claim 2, wherein Based on the previous-phase high-dimensional vector sequence, the post-phase high-dimensional vector sequence, the previous-phase fusion feature map of the first stage, and the post-phase fusion feature map of the first stage, processing is performed through the first block processing block to obtain the previous-phase global context feature map and the post-phase global context feature map of the first stage, including: An addition operation is performed on the previous-phase high-dimensional vector sequence and the previous-phase fusion feature map of the first stage to obtain an added previous-phase feature map. An addition operation is performed on the post-phase high-dimensional vector sequence and the post-phase fusion feature map of the first stage to obtain an added post-phase feature map. The added previous-phase feature map is processed through the first block processing block to obtain the previous-phase global context feature map of the first stage. The added post-phase feature map is processed through the first block processing block to obtain the post-phase global context feature map of the first stage.
7. The remote sensing image change detection method based on a dual-stream temporal feature adapter according to any one of claims 1-6, characterized in that The previous-phase global context feature maps and the post-phase global context feature maps of n stages are respectively input into the decoder for decoding to obtain a change detection mask for the remote sensing images of the previous and post phases, including: The previous-phase global context feature maps and the post-phase global context feature maps of n stages are respectively input into the decoder. Pixel-level difference operations are performed on the previous-phase global context feature maps and the post-phase global context feature maps of the corresponding stages to generate difference feature maps. Then, based on the lightweight multi-layer perceptron decoder, the difference feature maps are aggregated to obtain a change detection mask for the remote sensing images of the previous and post phases.
8. The remote sensing image change detection method based on the dual-stream temporal feature adapter according to any one of claims 1-6, characterized in that n≥4。 9. A remote sensing image change detection system based on a dual-stream temporal feature adapter, characterized in that, It includes an image acquisition module, an encoder, and a decoder, where: The image acquisition module is used to acquire remote sensing images of the previous phase and remote sensing images of the post phase. The encoder is used to perform bidirectional dynamic fusion feature extraction processing on the input remote sensing image of the previous phase and the remote sensing image of the subsequent phase through the pre-trained basic vision Transformer model backbone network of the encoder and two temporal feature adapter paths for n stages, to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase for n stages. Among them, the pre-trained basic vision Transformer model backbone network includes a patch embedding layer and n sequentially connected block processing blocks. Each temporal feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules. Each block processing block corresponds to a bidirectional gating module on each temporal feature adapter path to perform bidirectional dynamic fusion feature extraction processing for one stage, to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase for this stage. n is a positive integer; The decoder is used to decode the input global context feature maps of the previous phase and the global context feature maps of the subsequent phase for n stages respectively, to obtain the change detection mask of the remote sensing images of the previous and subsequent phases; The step of performing bidirectional dynamic fusion feature extraction processing on the input remote sensing image of the previous phase and the remote sensing image of the subsequent phase through the pre-trained basic vision Transformer model backbone network of the encoder and two temporal feature adapter paths for n stages, to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase for n stages, includes: The remote sensing image of the previous phase is input into multiple convolutional layers in the first temporal feature adapter path of the encoder for processing, to obtain the previous-phase feature maps with different resolutions; The remote sensing image of the subsequent phase is input into multiple convolutional layers in the second temporal feature adapter path of the encoder for processing, to obtain the subsequent-phase feature maps with different resolutions; The remote sensing image of the previous phase and the remote sensing image of the subsequent phase are input into the patch embedding layer of the pre-trained basic vision Transformer model backbone network of the encoder for processing, to obtain the previous-phase high-dimensional vector sequence and the subsequent-phase high-dimensional vector sequence; Based on the previous-phase feature maps with different resolutions, the subsequent-phase feature maps with different resolutions, the previous-phase high-dimensional vector sequence, and the subsequent-phase high-dimensional vector sequence, bidirectional dynamic fusion feature extraction processing is performed through n block processing blocks and n bidirectional gating modules on two temporal feature adapter paths for n stages, to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase for n stages.
Citation Information
Patent Citations
Self-attention feature fused high-resolution remote sensing image semantic change detection method
CN116486255A
Remote sensing image building change detection method based on multi-scale attention
CN118298305A
Cited By
Remote sensing image change detection method and system and electronic equipment
CN121121491A