Remote sensing image change detection method based on double-current time-phase feature adapter

By adopting the architecture of a dual-stream time-phase feature adapter in remote sensing image change detection, and using bidirectional dynamic fusion feature extraction processing and lightweight decoder, the problems of high computing resource consumption and feature misalignment in the prior art are solved, and efficient change detection and stronger robustness are achieved.

CN120047450AActive Publication Date: 2025-05-27NANJING UNIV OF INFORMATION SCI & TECH
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510534851.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-27
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The prior art has problems such as high computing resource consumption, misalignment of features, and insufficient cross-scene generalization capabilities in remote sensing image change detection, making it difficult to effectively model global context information in large-scale scenarios.

Method used

Using a dual-stream time-phase feature adapter architecture, a bidirectional dynamic fusion feature extraction process is performed through a shared pre-trained basic visual Transformer model backbone network and two time-phase feature adapter paths, to obtain the global context feature map of the front and back time-phase phases, and a change detection mask is generated through a lightweight multi-layer perceptron decoder.

Benefits of technology

It improves the accuracy and consistency of remote sensing image change detection, significantly reduces the computational overhead, enhances the robustness of feature representation and cross-scene adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047450A_ABST
    Figure CN120047450A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image change detection method based on a double-current time-phase feature adapter. The method comprises the following steps: acquiring remote sensing images of front and back time phases; the method comprises the following steps: taking remote sensing images of front and back time phases as the input of an encoder, and carrying out n stages of bidirectional dynamic fusion feature extraction processing through a pre-training basic visual Transform model backbone network of the encoder and two time phase feature adapter paths to obtain n stages of front time phase global context feature maps and back time phase global context feature maps, and respectively inputting the front-time-phase global context feature maps and the rear-time-phase global context feature maps of the n stages into a decoder for decoding to obtain change detection masks of the remote sensing images of the front and rear time phases. According to the method, the domain difference between the natural image and the remote sensing image is effectively bridged, the multi-scale feature difference is efficiently obtained, and the change detection precision and consistency of the remote sensing image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and remote sensing technology, and particularly to a remote sensing image change detection method and system based on a two-stream temporal feature adapter. Background Art

[0002] Remote sensing image change detection (CD) plays an important role in fields such as urban planning, environmental monitoring, and disaster assessment. By comparing remote sensing images from different periods, changes in land cover can be identified, such as the increase or decrease of buildings, changes in vegetation, etc. With the continuous improvement of satellite technology and sensor resolution, it has become easier to obtain high-resolution remote sensing data, which also poses higher requirements for change detection technology.

[0003] A convolutional neural network (CNN) is a deep learning model, especially suitable for processing data with a grid structure, such as images. The CNN automatically extracts local features in the image, such as edges, textures, etc., through convolutional layers, and reduces the feature dimension through pooling layers to improve computational efficiency. Although the CNN performs well in many computer vision tasks, when dealing with large-scale scenes, its receptive field is limited, making it difficult to fully model global context information. For example, FC-EF (Fully Convolutional Early Fusion) is a CNN-based change detection method that relies on convolutional operations to capture local texture patterns, but due to its limited receptive field, it is difficult to fully model global context information when dealing with large-scale scenes. In addition, CNN models usually require a large amount of training data to obtain good generalization ability, which is a challenge for remote sensing images with high annotation costs.

[0004] In traditional methods, many CNN-based architectures (such as FC-EF) and subsequent improved methods generally adopt a siamese network architecture. This architecture processes remote sensing images before and after change through two independent branch networks respectively, and then realizes change detection through feature fusion or difference calculation. However, although this design can capture local features to a certain extent, it has significant limitations: First, the siamese network needs to perform independent calculations on two sets of images, resulting in a multiple increase in computational resource consumption; Second, since the front and back images are independently processed through the same network branch, it is easy to have a feature misalignment problem. Especially in complex scenes or when the image has geometric distortion, it is difficult to maintain the spatial correspondence relationship of features. In addition, the siamese network lacks an effective adaptation mechanism for the domain difference between natural images and remote sensing images, resulting in insufficient cross-scene generalization ability.

[0005] In recent years, the Transformer model has received extensive attention due to its powerful ability to capture long-range dependencies. Transformer is a deep learning model based on the self-attention mechanism, initially used in the field of natural language processing, but has also been widely applied to computer vision tasks in recent years. Different from CNN, Transformer can capture long-range dependencies between any positions in the input sequence through the self-attention mechanism, thus enhancing the understanding of global context. For example, ChangeFormer can better capture long-range dependencies through the self-attention mechanism, thereby enhancing the understanding of global context. However, this method has the problem of high computational complexity, especially when performing pixel-level localization on high-resolution images, the performance is not good. In addition, when the Transformer model processes high-resolution images, it consumes a large amount of computing resources, which limits its promotion in practical applications.

[0006] The Vision Transformer (ViT) is an image recognition model developed based on the Transformer architecture in the field of natural language processing. Taking the Segment Anything Model (SAM) as an example, this is an advanced segmentation model that can perform high-quality segmentation on any object in the open world. SAM uses a powerful pre-trained ViT model to extract hierarchical semantic features of images and achieves generality for various tasks through fine-tuned adapters. However, when directly applying ViT to remote sensing image change detection, two main dilemmas will be encountered: First, the distribution difference between natural images and remote sensing images seriously affects the segmentation prior effect of ViT. Remote sensing images usually have higher spatial resolution and more complex scene structures, which makes it a great challenge for ViT to process remote sensing images; Second, existing methods are difficult to coordinate the hierarchical semantic abstraction of ViT with the per-pixel localization accuracy required for change detection. How to make ViT adapt to fine-grained change reasoning in different fields while maintaining its powerful segmentation ability is an open challenge that has not been solved yet.

[0007] Existing change detection methods still have many deficiencies when dealing with remote sensing images, especially in terms of balancing local details and global context, cross-architecture deep interactions, and domain adaptation. Traditional methods are difficult to effectively model global context information in large-scale scenarios, while Transformer-based methods face the problem of high computational complexity. Although hybrid architectures attempt to combine the advantages of both, they often introduce redundant parameters and implicit fusion strategies, affecting the efficiency and stability of the model. For example, TransUNetCD embeds Transformer modules in the U-Net structure, which improves the processing ability of optical images but still has a large computational overhead. In addition, pre-trained base models are prone to catastrophic forgetting in cross-domain tasks, and existing lightweight adapters cannot effectively retain the hierarchical segmentation priors of pre-trained models, resulting in poor performance in cross-domain tasks.

[0008] In summary, although existing deep learning methods have made some progress in remote sensing image change detection, there are still many challenges to be solved. How to effectively support remote sensing image change detection based on making full use of the segmentation priors of pre-trained vision Transformer models remains an important research direction. Future work needs to pay more attention to the design of domain adaptation strategies to improve the robustness and generalization ability of the model in different application scenarios. Summary of the Invention

[0009] The main purpose of the present invention is to provide a remote sensing image change detection method based on a two-stream temporal feature adapter, which can overcome the deficiencies of the prior art in remote sensing image change detection, obtain multi-scale feature differences more efficiently, and improve the accuracy and consistency of change detection.

[0010] Provide a remote sensing image change detection method based on a two-stream temporal feature adapter, including the following steps: S1. Obtain the remote sensing image of the previous time phase and the remote sensing image of the subsequent time phase; S2. Use the remote sensing image of the previous time phase and the remote sensing image of the later time phase as the inputs of the encoder. Through the pre-trained basic Vision Transformer model backbone network of the encoder and two time-phase feature adapter paths, perform bidirectional dynamic fusion feature extraction processing for n stages to obtain the global context feature maps of the previous time phase and the global context feature maps of the later time phase at n stages. Among them, the pre-trained basic Vision Transformer model backbone network includes a patch embedding layer and n sequentially connected block processing blocks. Each time-phase feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules. Each block processing block corresponds to a bidirectional gating module on each time-phase feature adapter path to perform bidirectional dynamic fusion feature extraction processing for one stage, obtaining the global context feature maps of the previous time phase and the global context feature maps of the later time phase at this stage. n is a positive integer; S3. Input the global context feature maps of the previous time phase and the global context feature maps of the later time phase at n stages into the decoder for decoding respectively to obtain the change detection masks of the remote sensing images of the previous and later time phases.

[0011] In one embodiment, the step of using the remote sensing image of the previous time phase and the remote sensing image of the later time phase as the inputs of the encoder, and through the pre-trained basic Vision Transformer model backbone network of the encoder and two time-phase feature adapter paths, performing bidirectional dynamic fusion feature extraction processing for n stages to obtain the global context feature maps of the previous time phase and the global context feature maps of the later time phase at n stages includes: Input the remote sensing image of the previous time phase into multiple convolutional layers in the first time-phase feature adapter path of the encoder for processing to obtain the previous time-phase feature maps with different resolutions; Input the remote sensing image of the later time phase into multiple convolutional layers in the second time-phase feature adapter path of the encoder for processing to obtain the later time-phase feature maps with different resolutions; Input the remote sensing image of the previous time phase and the remote sensing image of the later time phase into the patch embedding layer of the pre-trained basic Vision Transformer model backbone network of the encoder for processing to obtain the previous time-phase high-dimensional vector sequence and the later time-phase high-dimensional vector sequence; Based on the previous time-phase feature maps with different resolutions, the later time-phase feature maps with different resolutions, the previous time-phase high-dimensional vector sequence, and the later time-phase high-dimensional vector sequence, perform bidirectional dynamic fusion feature extraction processing for n stages through n block processing blocks and n bidirectional gating modules on two time-phase feature adapter paths to obtain the global context feature maps of the previous time phase and the global context feature maps of the later time phase at n stages.

[0012] In one embodiment, the processing method for each stage is the same, and the output of the previous stage is the input of the next stage; The processing process of the first stage is as follows: Based on the previous-phase feature maps of different resolutions, the previous-phase high-dimensional vector sequence, and the previous-phase global context feature map of the first stage output by the first block processing block, it is processed through the first bidirectional gating module in the first temporal feature adapter path to obtain the previous-phase fusion feature map of the first stage and the added previous-phase fusion feature map. Among them, the previous-phase fusion feature map of the first stage is input into the first block processing block, and the added previous-phase fusion feature map of the first stage is input into the second bidirectional gating module in the first temporal feature adapter path; Based on the post-phase feature maps of different resolutions, the post-phase high-dimensional vector sequence, and the post-phase global context feature map of the first stage output by the first block processing block, it is processed through the first bidirectional gating module in the second temporal feature adapter path to obtain the post-phase fusion feature map of the first stage and the added post-phase fusion feature map. Among them, the post-phase fusion feature map of the first stage is input into the first block processing block, and the added post-phase fusion feature map of the first stage is input into the second bidirectional gating module in the second temporal feature adapter path; Based on the previous-phase high-dimensional vector sequence and the post-phase high-dimensional vector sequence, the previous-phase fusion feature map of the first stage and the post-phase fusion feature map of the first stage, it is processed through the first block processing block to obtain the previous-phase global context feature map and the post-phase global context feature map of the first stage. Among them, the previous-phase global context feature map of the first stage is input into the second block processing block and the first bidirectional gating module in the first temporal feature adapter path, and the post-phase global context feature map of the first stage is input into the second block processing block and the first bidirectional gating module in the second temporal feature adapter path.

[0013] In one embodiment, the structure of each bidirectional gating module is the same. The bidirectional gating module includes a multi-receptive field feature pyramid module, a dynamic temporal perception enhancement module, a first bidirectional interaction module, and a second bidirectional interaction module; Among them, the multi-receptive field feature pyramid module is used to process the input feature map based on receptive fields of different sizes to form a multi-scale feature map; the dynamic temporal perception enhancement module is used to process the multi-scale feature map output by the multi-receptive field feature pyramid module, construct temporal features within a single phase, and generate an enhanced multi-scale feature map; the first bidirectional interaction module is used to interact and fuse the output of the dynamic temporal perception enhancement module at the current stage with the feature map output by the block processing block at the previous stage and output the first fusion feature map at the current stage; the second bidirectional interaction module is used to interact and fuse the output of the dynamic temporal perception enhancement module at the current stage with the feature map output by the block processing block at the current stage and output the second fusion feature map at the current stage; perform an addition operation on the output of the dynamic temporal perception enhancement module at the current stage and the second fusion feature map at the current stage to obtain the fused feature map after addition at the current stage.

[0014] In one embodiment, the multi-receptive field feature pyramid module reduces the dimension of the received feature map through the first linear projection layer, divides the dimension-reduced feature map along the channel dimension into multiple groups, each group corresponding to a receptive field of a different size. After each group passes through a depthwise separable convolution operation, it then passes through the second linear projection layer to restore the original dimension, forming a multi-scale feature map.

[0015] In one embodiment, the dynamic temporal perception enhancement module receives the multi-scale feature map output from the multi-receptive field feature pyramid module. Among them, the feature map of each scale in the multi-scale feature map corresponds to a different spatial resolution; Extract global statistical information from the feature map of each scale through global average pooling operations to generate statistical vectors at the channel level, input the statistical vectors into the first multi-layer perceptron, generate modulation vectors through non-linear transformation, and input the statistical vectors into the second multi-layer perceptron to generate channel weights; The feature map of each scale is respectively input into the corresponding depthwise separable convolution for processing to obtain processed feature maps; Use the modulation vectors to scale the corresponding processed feature maps channel by channel to obtain modulated feature maps; Perform channel re-weighting on the modulated feature maps according to the channel weights to obtain re-weighted feature maps; Add the received multi-scale feature map and the corresponding re-weighted feature map through residual connection and apply layer normalization operations to generate an enhanced multi-scale feature map.

[0016] In one embodiment, based on the pre-phase high-dimensional vector sequence, the post-phase high-dimensional vector sequence, the pre-phase fusion feature map of the first stage, and the post-phase fusion feature map of the first stage, they are processed by the first block processing block to obtain the pre-phase global context feature map and the post-phase global context feature map of the first stage, including: Perform an addition operation on the pre-phase high-dimensional vector sequence and the pre-phase fusion feature map of the first stage to obtain an added pre-phase feature map; Perform an addition operation on the post-phase high-dimensional vector sequence and the post-phase fusion feature map of the first stage to obtain an added post-phase feature map; Process the added pre-phase feature map through the first block processing block to obtain the pre-phase global context feature map of the first stage; Process the added post-phase feature map through the first block processing block to obtain the post-phase global context feature map of the first stage.

[0017] In one embodiment, the pre-phase global context feature maps and the post-phase global context feature maps of n stages are respectively input into a decoder for decoding to obtain a change detection mask of the remote sensing images in the pre-phase and post-phase, including: Respectively input the pre-phase global context feature maps and the post-phase global context feature maps of n stages into the decoder, perform pixel-level difference operations on the pre-phase global context feature maps and the post-phase global context feature maps of the corresponding stages to generate respective difference feature maps; then aggregate the respective difference feature maps based on a lightweight multi-layer perceptron decoder to obtain a change detection mask of the remote sensing images in the pre-phase and post-phase.

[0018] In one embodiment, n≥4.

[0019] A remote sensing image change detection system based on a two-stream temporal feature adapter, including an image acquisition module, an encoder, and a decoder, where: The image acquisition module is used to acquire a pre-phase remote sensing image and a post-phase remote sensing image; The encoder is used to perform bidirectional dynamic fusion feature extraction processing on the input remote sensing images of the previous phase and the remote sensing images of the subsequent phase through the pre-trained basic vision Transformer model backbone network of the encoder and two temporal feature adapter paths for n stages, to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase at n stages. Among them, the pre-trained basic vision Transformer model backbone network includes a patch embedding layer and n sequentially connected block processing blocks. Each temporal feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules. Each block processing block corresponds to a bidirectional gating module on each temporal feature adapter path to perform bidirectional dynamic fusion feature extraction processing for one stage, to obtain the global context feature maps of the previous phase and the global context feature maps of the subsequent phase at this stage, where n is a positive integer; The decoder is used to decode the input global context feature maps of the previous phase and the global context feature maps of the subsequent phase at n stages respectively, to obtain the change detection masks of the remote sensing images of the previous and subsequent phases.

[0020] The present invention also provides a computer storage medium, which stores a computer program executable by a processor, and this computer program executes the method for remote sensing image change detection based on a dual-stream temporal feature adapter described in the above technical solution.

[0021] The beneficial effects produced by the present invention are as follows: The present invention processes the remote sensing images of the previous and subsequent phases respectively through a shared pre-trained basic vision Transformer model backbone network (ViT) and two independent temporal feature adapter paths for bidirectional interaction and fusion. And bidirectional gating modules corresponding to different stages of the backbone network are set in the temporal feature adapter, which can perceive the internal temporal features of each phase (such as changes in lighting conditions and shooting conditions, etc.) while performing multi-scale feature extraction, so as to reduce the influence of non-semantic temporal variations on the detection results and further improve the robustness of feature representation; and the present invention effectively bridges the domain differences between natural images and remote sensing images through the architecture of the shared ViT combined with the dual-stream temporal feature adapter, more efficiently obtains multi-scale feature differences, improves the change detection accuracy and consistency of remote sensing images, and at the same time significantly reduces the computational overhead, and finally generates a high-resolution change detection mask for the remote sensing images of the previous and subsequent phases.

[0022] Furthermore, the multi-receptive field feature pyramid module (MRFP) in the bidirectional gating module of the present invention has a series of linear projection layers and depthwise separable convolutions, with different receptive fields, which are used to enhance the feature extraction ability of the CNN branch and ensure that each path can be optimized according to the characteristics of its respective phase.

[0023] Furthermore, the bidirectional gating adapter of the present invention activates the CNN-Transformer bidirectional interaction module (CTI), which can achieve dynamic fusion between local detail-sensitive CNN features and global semantic-aware Transformer features.

[0024] Furthermore, the lightweight multi-layer perceptron (MLP) decoder aggregates the differential feature maps to generate a high-resolution change detection mask, further improving the robustness and practicality of the system.

[0025] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 is a flowchart of the remote sensing image change detection method based on the dual-stream temporal feature adapter in the embodiments of the present invention; Figure 2 is a schematic structural diagram of the encoder and decoder of the remote sensing image change detection method system based on the dual-stream pyramid feature adapter in the embodiments of the present invention; Figure 3 is a schematic structural diagram of the multi-receptive field feature pyramid module in the embodiments of the present invention; Figure 4 is a schematic structural diagram of the dynamic temporal perception enhancement module in the embodiments of the present invention; Figure 5 is a schematic structural diagram of the bidirectional interaction module in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the present invention in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0029] It should be noted that the diagrams provided in the embodiments of the present invention only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape and size of the components in actual implementation. The type, quantity and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0030] In the present invention, it should also be noted that when terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. appear, the orientation or positional relationship indicated thereby is based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application. In addition, when terms such as "first" and "second" appear, they are only used for descriptive and distinguishing purposes, and should not be construed as indicating or implying relative importance.

[0031] In addition, it should also be noted that the features of various embodiments of the present invention can be partially or wholly combined or integrated, and as can be understood by those skilled in the art, they can interact and operate in different ways. Each embodiment can be implemented independently of each other or in an associated relationship.

[0032] In one embodiment, as Figure 1 and Figure 2 shown, a remote sensing image change detection method based on a two-stream temporal feature adapter is provided, including the following steps:

[0033] Step S1: Obtain the remote sensing image of the previous time phase and the remote sensing image of the subsequent time phase.

[0034] Step S2: Take the remote sensing image of the previous time phase and the remote sensing image of the subsequent time phase as the inputs of the encoder, and perform bidirectional dynamic fusion feature extraction processing in n stages through the pre-trained basic vision Transformer model backbone network of the encoder and two temporal feature adapter paths to obtain the global context feature maps of the previous time phase and the global context feature maps of the subsequent time phase in n stages. Among them, the pre-trained basic vision Transformer model backbone network includes a patch embedding layer and n sequentially connected block processing blocks. Each temporal feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules. Each block processing block corresponds to a bidirectional gating module on each temporal feature adapter path to perform bidirectional dynamic fusion feature extraction processing in one stage to obtain the global context feature map of the previous time phase and the global context feature map of the subsequent time phase in this stage. n is a positive integer.

[0035] Among them, the remote sensing image of the previous time phase and the remote sensing image of the subsequent time phase share the same pre-trained basic vision Transformer model backbone network for processing.

[0036] Among them, at the beginning and end of each block processing block processing stage, each bidirectional gating module in each path performs bidirectional dynamic fusion of features with the corresponding block processing block, and generates corresponding fused feature maps in each stage.

[0037] Among them, the bidirectional gating module is used to perform pyramid multi-scale feature image processing on the input image, and then through dynamic modulation and channel re-weighting driven by global statistical information, construct temporal features within a single phase, generate an enhanced feature map, interact and fuse with the block processing blocks at the corresponding stage, and finally unify the feature representation through a multi-scale self-attention mechanism to generate the fusion feature map at the corresponding stage.

[0038] Among them, the multiple convolutional layers on the two temporal feature adapter paths are independent of each other. The multiple convolutional layers are used to process the input remote sensing image to obtain feature maps with different resolutions. The feature maps with different resolutions overlap in a pyramid shape. The resolutions of these feature maps with different resolutions can be 1 / 8, 1 / 16, and 1 / 32 of the original image. Each feature map contains a D-dimensional feature representation (that is, the data points in each feature map are represented by a vector composed of D numerical values).

[0039] Step S3: Input the pre-temporal global context feature maps and post-temporal global context feature maps of n stages into the decoder for decoding to obtain the change detection masks of the pre-temporal and post-temporal remote sensing images.

[0040] As Figure 2 shown, in one embodiment, the pre-temporal remote sensing image and the post-temporal remote sensing image are used as the input of the encoder. Through the pre-trained basic vision Transformer model backbone network of the encoder and two temporal feature adapter paths, n-stage bidirectional dynamic fusion feature extraction processing is performed to obtain the pre-temporal global context feature maps and post-temporal global context feature maps of n stages, including: The pre-temporal remote sensing image is input into the multiple convolutional layers in the first temporal feature adapter path of the encoder for processing to obtain pre-temporal feature maps with different resolutions; the post-temporal remote sensing image is input into the multiple convolutional layers in the second temporal feature adapter path of the encoder for processing to obtain post-temporal feature maps with different resolutions; the pre-temporal remote sensing image and the post-temporal remote sensing image are input into the patch embedding layer of the pre-trained basic vision Transformer model backbone network of the encoder for processing to obtain a pre-temporal high-dimensional vector sequence and a post-temporal high-dimensional vector sequence; based on the pre-temporal feature maps with different resolutions, the post-temporal feature maps with different resolutions, the pre-temporal high-dimensional vector sequence, and the post-temporal high-dimensional vector sequence, through n block processing blocks and n bidirectional gating modules on the two temporal feature adapter paths, n-stage bidirectional dynamic fusion feature extraction processing is performed to obtain the pre-temporal global context feature maps and post-temporal global context feature maps of n stages.

[0041] Among them, the network structures in the first temporal feature adapter path and the second temporal feature adapter path are the same. The difference is that each path processes the remote sensing images of one temporal phase, and then outputs the feature maps corresponding to the temporal phase.

[0042] As Figure 2 shown, in one embodiment, the processing method in each stage is the same, and the output of the previous stage is the input of the next stage; the processing process of the first stage is as follows: Based on the previous temporal feature maps with different resolutions, the previous temporal high-dimensional vector sequences, and the previous temporal global context feature maps of the first stage output by the first block processing block, they are processed through the first bidirectional gating module in the first temporal feature adapter path to obtain the previous temporal fusion feature maps of the first stage and the previous temporal fusion feature maps after addition. Among them, the previous temporal fusion feature maps of the first stage are input into the first block processing block, and the previous temporal fusion feature maps after addition of the first stage are input into the second bidirectional gating module in the first temporal feature adapter path; based on the subsequent temporal feature maps with different resolutions, the subsequent temporal high-dimensional vector sequences, and the subsequent temporal global context feature maps of the first stage output by the first block processing block, they are processed through the first bidirectional gating module in the second temporal feature adapter path to obtain the subsequent temporal fusion feature maps of the first stage and the subsequent temporal fusion feature maps after addition. Among them, the subsequent temporal fusion feature maps of the first stage are input into the first block processing block, and the subsequent temporal fusion feature maps after addition of the first stage are input into the second bidirectional gating module in the second temporal feature adapter path; based on the previous temporal high-dimensional vector sequences and the subsequent temporal high-dimensional vector sequences, the previous temporal fusion feature maps of the first stage and the subsequent temporal fusion feature maps of the first stage, they are processed through the first block processing block to obtain the previous temporal global context feature maps and the subsequent temporal global context feature maps of the first stage. Among them, the previous temporal global context feature maps of the first stage are input into the second block processing block and the first bidirectional gating module in the first temporal feature adapter path, and the subsequent temporal global context feature maps of the first stage are input into the second block processing block and the first bidirectional gating module in the second temporal feature adapter path.

[0043] It should be understood that the processing method in each stage is the same, only the input of the first stage is different from that of other stages, and the input of the first stage is the output of the patch embedding layer and multiple convolutional layers in each temporal feature adapter path.

[0044] As Figure 2As shown, in one embodiment, the structures of each bidirectional gating module are the same. The bidirectional gating module includes a multi-receptive field feature pyramid module, a dynamic temporal perception enhancement module, a first bidirectional interaction module, and a second bidirectional interaction module. Among them, the multi-receptive field feature pyramid module is used to process the input feature map based on receptive fields of different sizes to form a multi-scale feature map. The dynamic temporal perception enhancement module is used to process the multi-scale feature map output by the multi-receptive field feature pyramid module, construct the temporal features within a single phase, and generate an enhanced multi-scale feature map. The first bidirectional interaction module is used to interact and fuse the output of the dynamic temporal perception enhancement module in the current stage with the feature map output by the block processing block in the previous stage and output the first fused feature map in the current stage. The second bidirectional interaction module is used to interact and fuse the output of the dynamic temporal perception enhancement module in the current stage with the feature map output by the block processing block in the current stage and output the second fused feature map in the current stage. An addition operation is performed on the output of the dynamic temporal perception enhancement module in the current stage and the second fused feature map in the current stage to obtain the fused feature map after addition in the current stage.

[0045] Among them, if the current stage is the first stage, the first bidirectional interaction module is used to interact and fuse the output of the dynamic temporal perception enhancement module in the current stage with the feature maps output by multiple convolutional layers on this path and output the first fused feature map in the current stage.

[0046] Among them, if the path where the bidirectional gating module is located is to process the remote sensing image of the previous phase, the first fused feature map in the current stage is the fused feature map of the previous phase input into the block processing block in the current stage. If the path where the bidirectional gating module is located is to process the remote sensing image of the later phase, the first fused feature map in the current stage is the fused feature map of the later phase input into the block processing block in the current stage.

[0047] It should be understood that if the path where the bidirectional gating module is located is to process the remote sensing image of the previous phase, then all types of feature maps processed or output in the bidirectional gating module are of the previous phase. Correspondingly, if the path where the bidirectional gating module is located is to process the remote sensing image of the later phase, then all types of feature maps processed or output in the bidirectional gating module are of the later phase.

[0048] As Figure 3 shown, in one embodiment, the multi-receptive field feature pyramid module reduces the dimension of the received feature map through the first linear projection layer, divides the dimension-reduced feature map along the channel dimension into multiple groups, each group corresponding to a receptive field of a different size. After each group passes through the depthwise separable convolution operation, it then restores the original dimension through the second linear projection layer to form a multi-scale feature map.

[0049] It should be understood that if the path where the bidirectional gating module is located is the remotely sensed image of the pre-processing phase, all types of feature maps processed or output by the multi-receptive field feature pyramid module in the bidirectional gating module are of the pre-phase; correspondingly, if the path where the bidirectional gating module is located is the remotely sensed image of the post-processing phase, all types of feature maps processed or output by the multi-receptive field feature pyramid module in the bidirectional gating module are of the post-phase.

[0050] As Figure 4 shown, in one embodiment, the dynamic temporal perception enhancement module receives the multi-scale feature maps output from the multi-receptive field feature pyramid module, where each scale of the feature maps in the multi-scale feature maps corresponds to a different spatial resolution; global statistical information is extracted from each scale of the feature maps through global average pooling operations to generate respective statistical vectors at the channel level, the respective statistical vectors are input into the first multi-layer perceptron, and respective modulation vectors are generated through non-linear transformation, the respective statistical vectors are input into the second multi-layer perceptron to generate respective channel weights; each scale of the feature maps is respectively input into the corresponding depthwise separable convolution for processing to obtain respective processed feature maps; the respective processed feature maps are scaled channel by channel using the respective modulation vectors to obtain respective modulated feature maps; the respective modulated feature maps are re-weighted in channels according to the respective channel weights to obtain respective re-weighted feature maps; the received multi-scale feature maps are added to the corresponding re-weighted feature maps through residual connections, and layer normalization operations are applied to generate enhanced multi-scale feature maps.

[0051] It should be understood that if the path where the bidirectional gating module is located is the remotely sensed image of the pre-processing phase, all types of feature maps processed or output by the dynamic temporal perception enhancement module in the bidirectional gating module are of the pre-phase; correspondingly, if the path where the bidirectional gating module is located is the remotely sensed image of the post-processing phase, all types of feature maps processed or output by the dynamic temporal perception enhancement module in the bidirectional gating module are of the post-phase.

[0052] In one embodiment, the first bidirectional interaction module is specifically configured to, at the beginning of each stage, fuse the output of the dynamic temporal perception enhancement module with the output of the block processing block of the previous stage, and then apply the multi-scale deformable attention mechanism to the unified fused feature representation to align the feature sizes of each scale and transfer them to the block processing block of the current stage through interpolation operations.

[0053] In one embodiment, the second bidirectional interaction module is specifically configured to, at the end of each stage, fuse the output of the dynamic temporal perception enhancement module with the output of the block processing block of the current stage, then apply a multi-scale deformable attention mechanism to the unified fused feature representation, align the feature sizes of each scale, and after interpolation operations, generate fused feature maps of different scales corresponding to the current stage, add them to the output of the dynamic temporal perception enhancement module of the current stage, and then pass the result to the bidirectional gating module of the next stage.

[0054] In one embodiment, based on the high-dimensional vector sequence of the previous time phase, the high-dimensional vector sequence of the subsequent time phase, the fused feature map of the previous time phase of the first stage, and the fused feature map of the subsequent time phase of the first stage, processing is performed through the first block processing block to obtain the global context feature map of the previous time phase and the global context feature map of the subsequent time phase of the first stage, including: Perform an addition operation on the high-dimensional vector sequence of the previous time phase and the fused feature map of the previous time phase of the first stage to obtain the added previous time phase feature map; perform an addition operation on the high-dimensional vector sequence of the subsequent time phase and the fused feature map of the subsequent time phase of the first stage to obtain the added subsequent time phase feature map; process the added previous time phase feature map through the first block processing block to obtain the global context feature map of the previous time phase of the first stage; process the added subsequent time phase feature map through the first block processing block to obtain the global context feature map of the subsequent time phase of the first stage.

[0055] As Figure 2 shown, in one embodiment, the global context feature maps of the previous time phase and the global context feature maps of the subsequent time phase of n stages are respectively input into the decoder for decoding to obtain the change detection mask of the remote sensing images of the previous and subsequent time phases, including: The global context feature maps of the previous time phase and the global context feature maps of the subsequent time phase of n stages are respectively input into the decoder, and pixel-level difference operations are performed on the global context feature maps of the previous time phase and the global context feature maps of the subsequent time phase of the corresponding stages to generate respective difference feature maps; then, based on the lightweight multi-layer perceptron decoder, the respective difference feature maps are aggregated to obtain the change detection mask of the remote sensing images of the previous and subsequent time phases.

[0056] The above remote sensing image change detection method based on a dual-stream temporal feature adapter processes the remote sensing images of the front and back phases respectively through a shared pre-trained basic vision Transformer model backbone network (ViT) and two independent temporal feature adapter paths, conducts two-way interactive fusion, and a two-way gating module corresponding to different stages of the backbone network is set in the temporal feature adapter. While performing multi-scale feature extraction, it also perceives the internal temporal features of the phase (such as changes in lighting conditions and shooting conditions, etc.), thereby reducing the impact of non-semantic temporal variations on the detection results and further enhancing the robustness of feature representation. Moreover, the present invention effectively bridges the domain differences between natural images and remote sensing images through the architecture of a shared ViT combined with a dual-stream temporal feature adapter, more efficiently obtains multi-scale feature differences, improves the change detection accuracy and consistency of remote sensing images, and at the same time significantly reduces the computational overhead, and finally generates a change detection mask for high-resolution remote sensing images of the front and back phases.

[0057] In one embodiment, as Figure 2 shown, each independent temporal feature adapter path of the encoder interacts with the intermediate shared basic vision Transformer model backbone network at each stage. The parameters of the shared pre-trained basic vision Transformer model backbone network (Vision Transformer, abbreviated as ViT) remain frozen during use, avoiding performance degradation caused by re-training or fine-tuning, and leveraging its powerful semantic prior to enhance the generalization ability of the model. The basic vision Transformer model backbone network includes n block processing blocks, and the size of n can be selected according to actual needs. Generally, n≥4. In this embodiment, n = 4 is selected, that is, the basic vision Transformer model backbone network contains four block processing blocks, and the backbone network is divided into four processing stages.

[0058] Correspondingly, each temporal feature adapter in each path also contains four two-way gating modules, which interact and fuse with the block processing blocks at the beginning and end of each stage, continuously updating the input and output of the block processing blocks at each stage, and also continuously updating the input and output of the two-way gating modules at each stage. And at the end of each stage, fused feature maps of different scales at the corresponding level are generated, which fuse the different scale features of the remote sensing images of the front and back phases. The siamese adapter architecture designed by the present invention provides independent feature extraction and fusion processes for the remote sensing images of the front and back phases, ensuring that each path can be optimized according to the characteristics of its respective phase, thereby improving the overall performance of the system.

[0059] In addition, as Figure 2The Patch Embedding layer in it is a step in ViT to convert the input image into sequence data suitable for Transformer processing. The specific process is as follows: First, the original image (i.e., the remote sensing image of the previous time phase or the remote sensing image of the later time phase) is segmented into multiple small patches of a fixed size; then, each patch is flattened into a one-dimensional vector and mapped to the hidden dimension of the encoder through a linear projection layer, thereby generating the feature representation of each patch; in order to retain the position information of these patches in the original image, position encoding is also added to each patch. After Patch Embedding, the image is thus converted into a sequence composed of multiple high-dimensional vectors with spatial information encoding, that is, the high-dimensional vector sequence of the previous time phase or the high-dimensional vector sequence of the later time phase, so that the subsequent Transformer architecture can effectively process this information.

[0060] Among them, the bidirectional gating module includes a multi-receptive field feature pyramid module MRFP, a dynamic temporal perception enhancement module DTAEM, and two cross-temporal interaction modules CTI. Among them, MRFP is used to form multi-scale feature maps based on receptive fields of different sizes; DTAEM is used to process the multi-scale feature maps output by MRFP, construct the temporal features within a single time phase, and generate enhanced multi-scale feature maps; CTI is used to interact and fuse the output of DTAEM with the block processing blocks at the corresponding stages and output the fused feature maps at different stages.

[0061] During specific processing, the remote sensing images of the previous and later time phases are simultaneously input into the shared backbone network of the basic vision Transformer model and their respective independent temporal feature adapter paths. In their respective temporal feature adapter paths, the remote sensing images of each time phase first pass through a series of independent convolutional layers for processing to generate initial feature maps (i.e., the previous-time-phase feature maps of different resolutions or the later-time-phase feature maps of different resolutions); the initial feature maps then pass through the multi-receptive field feature pyramid module (MRFP) to generate feature maps of different scales, such as Figure 3As shown, the multi-receptive field feature pyramid module includes a series of linear projection layers and depthwise separable convolutions with different receptive fields. The input feature map is dimensionally reduced through the linear projection layer, and the dimensionally reduced feature map is divided into multiple groups along the channel dimension, with each group corresponding to a receptive field of a different size. Each group undergoes depthwise separable convolution operations to expand the receptive field and enhance the long-range modeling ability. The processed feature map passes through the linear projection layer again to restore the original dimension, forming a multi-scale feature map. Among them, by using depthwise separable convolution (Depthwise Separable Convolution) and the multi-receptive field feature pyramid module (MRFP), the receptive field of each convolutional layer can be expanded, enabling the encoder to consider a wider range of context information. Depthwise separable convolution is an efficient convolution method that first independently applies spatial convolution to each channel and then integrates cross-channel information through 1x1 convolution. This method not only reduces the computational amount and the number of parameters but also allows the model to capture more complex patterns and long-range dependencies by combining multiple convolutional kernels with different receptive fields.

[0062] As Figure 4 shown, the dynamic temporal awareness enhancement module (DTAEM) models the temporal features within a single phase through a dynamic modulation and channel reweighting mechanism driven by global statistical information, enhancing the robustness of feature representation. Among them, by constructing the temporal features within a single phase (such as changes in lighting conditions), the robustness of the multi-scale feature map to phase features can be enhanced, thereby reducing the interference of non-semantic temporal variations on change detection and improving the adaptability and discriminability of feature representation. The specific process includes: 1) DTAEM receives the multi-scale feature maps from the multi-receptive field feature pyramid module (MRFP), where each feature map corresponds to a different spatial resolution; 2) Global statistical data extraction: For each scale of the feature map, first, global statistical information is extracted through global average pooling operations to generate a channel-level statistical vector to characterize the overall distribution characteristics of the feature map; subsequently, the statistical vector is input into a multi-layer perceptron (MLP), and a modulation vector is generated through non-linear transformation. Based on the same statistical vector, a channel weight is generated through another multi-layer perceptron; 3) Using this modulation vector to scale the output of the depthwise separable convolution channel by channel, dynamically modulating the convolution response to adapt to temporal characteristics; 4) Channel recalibration: Using the channel weight to reweight the channels of the modulated feature map, highlighting the feature channels related to change detection and suppressing the influence of non-semantic temporal variations; 5) Finally, the multi-scale feature maps from the multi-receptive field feature pyramid module are added to the reweighted feature map through a residual connection, and a layer normalization operation is applied to generate an enhanced multi-scale feature map, providing a robust feature representation for subsequent feature fusion.

[0063] As Figure 2 、 5As shown, the CNN-Transformer based Bidirectional Interaction Module (CTI) is mainly used to achieve the dynamic fusion between the locally detailed CNN features and the globally semantic-aware Transformer features. It can dynamically fuse local and global features at different scales, enhance the expression ability and detection accuracy of the system, and unify the representation differences between different modalities through direct addition and multi-scale self-attention mechanism, enhancing the model's performance in complex scenarios. The specific process includes: 1) At the beginning of each stage, directly add and fuse the different-scale features from the temporal feature adapter with the ViT branch features (i.e., the output features of each block processing block); 2) Apply the multi-scale deformable attention mechanism to further unify the feature representation, align the feature sizes of each scale and transfer them to the next layer of the ViT branch through interpolation operations (the ViT branch refers to the block processing block. When n = 4, the four block processing blocks correspond to Figure 2 ViT-S1, ViT-S2, ViT-S3, and ViT-S4 in

[0064] ); 3) Generate feature maps of different scales at the corresponding levels and transfer them to the next stage for further processing. 1) Perform pixel-level difference operations on the pre-temporal global context feature map and the post-temporal global context feature map of the corresponding stage to generate the difference feature map of the corresponding stage; 2) The multi-layer perceptron (MLP) decoder receives the four different-level difference feature maps from the difference features as inputs, processes them through a series of MLP layers, gradually reduces the number of channels and increases the spatial resolution, and uses the upsampling operation to restore the feature map to the original image resolution, and finally outputs a high-resolution change detection mask indicating the specific location of the change area.

[0065] In one embodiment, it is based on SAM-ViT, and n = 4. SAM-ViT refers to the backbone network of a pre-trained Vision Transformer (ViT) model based on the Segment Anything Model (SAM). In SAM, ViT is used as its backbone network, which means that ViT undertakes the main task of processing and understanding the input image in SAM. Specifically, SAM uses the pre-trained ViT to extract the feature representation of the image. These features are efficient and information-rich image representations, which are crucial for performing subsequent segmentation tasks. The ViT backbone network of SAM is pre-trained on a massive dataset including SA-1B, which contains more than 1 billion mask annotations. This large-scale data enables the model to learn diverse segmentation patterns. In addition, combined with the contrastive learning method, ViT can further enhance the sensitivity and discrimination ability to object boundaries. Based on this SAM-ViT, this embodiment adopts a SAM-CNN encoder for encoding, that is, step S2 is implemented. The SAM-CNN encoder is not an existing independent model, but is used to describe the encoder structure formed by combining SAM-ViT with a CNN (Convolutional Neural Network) branch in this embodiment. The global features are extracted through the shared SAM-ViT, and at the same time, two independent temporal feature adapter paths are used to process the pre-temporal remote sensing image and the post-temporal remote sensing image respectively, generating multi-scale local features and fusing them with the global features. The SAM-CNN encoder consists of two parts: (a) a shared parameter-frozen SAM-ViT backbone network, (b) a temporal feature adapter. First, for the SAM-ViT branch, the pre-temporal remote sensing image and the post-temporal remote sensing image with the shape of H×W×3 are respectively input into the patch embedding layer to obtain a feature representation with a resolution reduced to 1 / 16 of the original image. At the same time, for the temporal feature adapter branch, the pre-temporal remote sensing image and the post-temporal remote sensing image are processed through a series of independent convolutional layers on the corresponding paths, generating feature pyramids with resolutions of 1 / 8, 1 / 16, and 1 / 32 respectively (i.e., Figure 3 C in 3 、C 4 and C 5), each feature map of the feature pyramid contains a D-dimensional feature representation. Next, the features on the two paths undergo 4 stages of feature interaction. At each stage, first, the MRFP enhances the feature pyramid to obtain multi-scale feature maps to capture richer spatial information. Subsequently, the DTAEM performs global statistic-driven dynamic modulation and channel re-weighting on the multi-scale feature maps output by the MRFP, models the temporal features within a single phase, and generates enhanced feature maps. Then, two CTIs interact bidirectionally with the features of SAM-ViT to obtain multi-scale features with rich semantic information. It should be noted that one CTI runs at the beginning of each stage, and the other CTI runs at the end of each stage to ensure that the features can be fully and effectively interacted. After 4 stages of feature interaction, the features of the two phases at each stage are subsequently input into the decoder for feature difference calculation.

[0066] Among them, to enhance the encoder's ability to capture spatial information at different scales, the MRFP is introduced. As Figure 3 shown, the MRFP consists of a series of linear projection layers and depthwise separable convolutions with different receptive fields. Specifically, the input includes feature maps C 3 、C 4 、C 5 with different resolutions. First, dimensionality reduction is performed through a linear projection layer, and then a dimensionality-reduced feature representation C is obtained in , where R represents real numbers, and H and W are used to represent spatial dimensions. Subsequently, these features are divided into M groups along the channel dimension, and each group corresponds to a convolutional layer with a different kernel size, thereby expanding the receptive field and enhancing the model's long-distance modeling ability for CNN features. Then, the processed features are connected through another linear projection layer and restored to the original dimension, and the multi-scale feature maps are output. This process can be mathematically expressed as: ; where FC(•) represents the linear projection operation, DWConv(•) represents a group of depth convolutions with different kernel sizes, and represents the output of the MRFP. The MRFP not only provides rich multi-scale information but also enhances the model's ability to capture local details and global backgrounds of images by combining different receptive fields, thus significantly improving the model's performance in dense prediction tasks.

[0067] In remote sensing change detection tasks, due to factors such as lighting conditions, seasonal changes, or atmospheric effects, images taken at different time stages often exhibit significant feature changes. Although these attributes of specific time phases do not directly show semantic changes on the ground surface, they introduce interference during the feature extraction process, thus affecting the accuracy of change detection. To solve this problem, the present invention specifically designs DTAEM, which can dynamically adjust the feature representations corresponding to each time phase processed by each adapter to ensure their consistency with specific time backgrounds (such as changes in lighting intensity or vegetation conditions in different seasons).

[0068] As Figure 4 shown, DTAEM processes the multi-scale feature maps generated by MRFP. Among them, the multi-scale feature maps , i = 3, 4, 5 for each feature map correspond to different spatial resolutions. Among them, B represents the batch size, H i represents height, W i represents width, and D is used to represent the channel dimension. This DTAEM processes these multi-scale feature maps through a series of operations, which can dynamically adjust the convolution response and recalibrate the channel importance according to the global feature statistics to ensure consistency with the specific time features of each image. The workflow first independently analyzes the feature maps of each scale, using depthwise separable convolution, temporal modulation, and channel reweighting to generate enhanced feature representations that are robust to temporal artifacts.

[0069] At the beginning of the processing, first perform depthwise separable convolution on each feature map to effectively capture local spatial patterns. To make this operation temporally adaptive, global statistical information is extracted through global average pooling operation , and the mathematical formula is as follows: ; where, i = 3, 4, 5, represents the output of the global average pooling operation, which encompasses the overall feature distribution across the spatial dimension. Then the statistical vector is input into a multi-layer perceptron (MLP) to generate a modulation vector: ; where, i = 3, 4, 5, and σ represents the sigmoid activation function. The modulation vector scales the output of the depthwise separable convolution DWConv, dynamically adjusting the feature response according to the temporal background encoded in . The mathematical expression for this step is: , ; Among them, i = 3, 4, 5, expand(•) means broadcasting M i to the spatial dimension of the match . represents the result after per-channel scaling of the output of the depthwise separable convolution using the modulation vector, DWConv(•) represents the depthwise separable convolution using a 3×3 kernel, represents the output of the depthwise separable convolution, ⊙ represents element-wise multiplication with the broadcast elements to match the spatial dimension. This modulation mechanism enables the model to selectively enhance or suppress feature channels according to the correlation between feature channels and temporal changes, thereby effectively distinguishing true changes and temporal artifacts.

[0070] After modulation, DTAEM uses an adaptive reweighting mechanism to recalibrate channel importance, thereby further refining the feature map. Using the same global statistics , the second MLP calculates the per-channel scaling weights . These weights are applied to the modulated features to emphasize the channels most relevant to change detection while suppressing the channels dominated by noise or irrelevant temporal changes: ; Among them, i = 3, 4, 5, represents the reweighted feature map.

[0071] This reweighting step ensures the effective allocation of computing resources and enhances the robustness of the model to non-semantic temporal noise. To maintain the integrity of the original features and stabilize training, a residual connection combines the input feature map with the reweighted output . Then the resulting features are layer-normalized to maintain the consistency of the feature scale throughout the network: ; Among them, i = 3, 4, 5, LayerNorm(•) represents layer normalization, represents the enhanced multi-scale feature map output by DTAEM.

[0072] This workflow is repeatedly applied at all scales to ensure that dynamic temporal enhancement is effective at different spatial resolutions. As Figure 5 shown, the enhanced multi-scale feature map Subsequently, it is sent to the CTI module for two-way integration with the global context feature map from SAM-ViT. This method is based on globally feature-statistics-driven modulation and reweighting, ensuring that the resulting feature representation is both robust and semantics-focused, thus supporting precise change detection in complex remote sensing environments.

[0073] To effectively integrate the features of the ViT and CNN branches (i.e., the temporal feature adapter path) at each stage, CTI is also introduced. At the beginning of each stage, the temporally feature-enhanced multi-scale features obtained from DTAEM are fused with the features from the ViT branch by directly adding F 4 to X to generate a new set of features F′ = {F 3 , F 4 , F 5}, and this process can be expressed as: Subsequently, self-attention calculation is performed on F′ to unify the representational differences of different modalities: , where norm(•) represents layer normalization, Attention(•) represents multi-scale deformable attention, FFN(•) represents the feed-forward network, and O represents the output of CTI. Then, the feature map O output by CTI 3 and the feature map O 5 are aligned in size with the feature map O 4 , and the results are added to the input of the corresponding ViT branch to form the input of this ViT branch: ; where, is the updated feature of the ViT branch, α is a learnable variable, represents the O 4 feature map aligned in size with O 3 , represents the O 4 feature map aligned in size with O 5 feature map.

[0074] At the end of each stage, the features F′ = {F 3 , F 4 , F 5} and O = {O 3 , O 4 , O 5}, directly add the features of the corresponding scales to enhance the feature representation ability of the CNN branch at each scale, and finally pass these fused features to the CNN branch of the next stage to ensure that the model can more accurately capture the multi-level information in the image.

[0075] The decoder used in this embodiment is an MLP decoder. The task of the MLP decoder is to aggregate the multi-level difference feature maps from the difference calculation to generate a high-resolution change detection mask. Using a sequence of multi-layer perceptron (MLP) layers and strategic upsampling operations, this lightweight decoder can effectively synthesize the multi-scale difference features into a coherent representation, thereby accurately identifying the changed regions in the remote sensing image pair.

[0076] The decoder receives the pre-phase global context feature maps of n stages output by the encoder and the post-phase global context feature maps , where i = 1, 2, 3, 4. The pre-phase global context feature maps and the post-phase global context feature maps of the same stage perform pixel-level difference operations to generate the difference feature maps of this stage and input them into the MLP decoder.

[0077] The MLP decoder receives a set of four difference feature maps as input. When i = 1, 2, 3, 4, the difference feature maps are denoted as , and each difference feature map is obtained through difference calculation. The spatial resolution of these difference feature maps is , and the channel dimension of the difference feature map is D i . Due to the differences in scale and dimension, a structured process is required to coordinate these features to achieve effective fusion and final prediction.

[0078] To initiate this process, each difference feature map is transformed through an MLP layer to standardize its channel dimension to a unified embedding size C ebd . This channel unification step ensures the compatibility of each layer, and its mathematical definition is as follows: ; where, represents the channel dimension of the difference feature map of the corresponding stage, and represents the i-th layer difference feature map after being processed by the MLP layer.

[0079] Subsequently, the spatial resolution of these feature maps is adjusted to a common scale ( ) through bilinear interpolation. This upsampling operation prepares for connecting the features, and its expression is: ; Among them, represents the i-th layer of differential feature map after upsampling processing, Upsample(•) represents upsampling, and bilinear represents upsampling using the bilinear interpolation method.

[0080] After normalizing the channel and spatial dimensions, the feature map after upsampling processing is concatenated along the channel axis to obtain a composite feature map with a channel dimension of 4C ebd Then, an additional MLP layer is used to process this combined representation, fusing multi-level differential information into a unified fused feature map F fused , which is expressed as follows: ; where Cat(•) represents concatenation.

[0081] To generate the final change detection mask at the original image resolution, the fused feature map F fused will be upsampled through a two-dimensional transposed convolutional layer with a stride of 4 and a kernel size of 3 to restore the spatial dimension to H×W. These operations are defined as follows: , ; where, represents the feature map after the transposed convolution operation, ConvTranspose2D represents the two-dimensional transposed convolutional layer, represents the class parameter, that is, the number of classes of the final output, S represents the stride, which is 4 here, K represents the size of the convolutional kernel, which is 3×3 here, CM represents the predicted change mask, with a dimension of H×W×N cls .

[0082] The streamlined architecture of the MLP decoder takes into account both computational efficiency and the ability to effectively integrate multi-scale differential features. By using the MLP layer for feature transformation and fusion, and supplemented by precise upsampling, this module can ensure the generation of accurate and spatially consistent change maps.

[0083] In one embodiment, a remote sensing image change detection system based on a two-stream temporal feature adapter is provided, including an image acquisition module, an encoder, and a decoder, where: The image acquisition module is used to acquire the remote sensing image of the previous time phase and the remote sensing image of the later time phase; The encoder is used to perform bidirectional dynamic fusion feature extraction processing on the input pre-temporal remote sensing image and post-temporal remote sensing image through the pre-trained basic vision Transformer model backbone network of the encoder and two temporal feature adapter paths for n stages, to obtain the pre-temporal global context feature maps and post-temporal global context feature maps of n stages. Among them, the pre-trained basic vision Transformer model backbone network includes a patch embedding layer and n sequentially connected block processing blocks. Each temporal feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules. Each block processing block corresponds to a bidirectional gating module on each temporal feature adapter path to perform bidirectional dynamic fusion feature extraction processing for one stage, to obtain the pre-temporal global context feature map and post-temporal global context feature map of this stage, where n is a positive integer; The decoder is used to decode the input pre-temporal global context feature maps and post-temporal global context feature maps of n stages respectively, to obtain the change detection masks of the pre-temporal and post-temporal remote sensing images.

[0084] Each module is mainly used to implement each step of the above method embodiment, which will not be elaborated here one by one.

[0085] The present application also provides a computer-readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, a server, an App application mall, etc., on which a computer program is stored, and when the program is executed by a processor, the corresponding functions are realized. When the computer-readable storage medium of this embodiment is executed by a processor, it realizes the remote sensing image change detection method based on a dual-stream temporal feature adapter in the method embodiment.

[0086] In summary, different from the traditional Siamese network architecture that relies on high computational costs and low feature alignment efficiency, the present invention proposes an architecture that combines a shared pre-trained basic Vision Transformer (ViT) model backbone network with a dual-stream pyramid feature adapter, effectively bridging the domain differences between natural images and remote sensing images while significantly reducing the computational overhead. By processing the remote sensing images of the front and back time phases through two independent lightweight pyramid feature adapter paths respectively, the dynamic fusion between the local detail-sensitive CNN features and the global semantic-aware Transformer features is achieved, enhancing the cross-modal feature enhancement effect. In addition, a dynamic temporal awareness enhancement module (DTAEM) is specially designed to model the temporal features within a single time phase through dynamic modulation and channel re-weighting driven by global statistical information, thereby reducing the impact of non-semantic temporal variations on the detection results and further improving the robustness of feature representation. It not only provides a resource-efficient solution but also significantly improves the accuracy and interpretability of remote sensing image change detection in complex scenarios.

[0087] It should be noted that, according to the needs of implementation, each step / component described in this application can be split into more steps / components, or two or more steps / components or partial operations of steps / components can be combined into new steps / components to achieve the purpose of the present invention.

[0088] The sequence numbers of the steps in the above embodiments do not indicate the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0089] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A remote sensing image change detection method based on a dual-stream temporal feature adapter, characterized in that: The remote sensing image change detection method based on the dual-stream temporal feature adapter includes: S1, acquiring a remote sensing image of a front phase and a remote sensing image of a back phase; S2. The remote sensing image of the front phase and the remote sensing image of the rear phase are used as the input of the encoder, and the pre-trained basic visual Transformer model backbone network of the encoder and the two phase feature adapter paths perform n-stage bidirectional dynamic fusion feature extraction processing to obtain the n-stage front phase global context feature map and the rear phase global context feature map, wherein the pre-trained basic visual Transformer model backbone network includes a patch embedding layer and n sequentially connected block processing blocks, each phase feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules, each block processing block corresponds to a bidirectional gating module on each phase feature adapter path to perform one stage of bidirectional dynamic fusion feature extraction processing to obtain the front phase global context feature map and the rear phase global context feature map of the stage, and n is a positive integer; S3. Input the global context feature map of the previous phase and the global context feature map of the subsequent phase of n stages into the decoder for decoding respectively, and obtain the change detection mask of the remote sensing image of the previous and subsequent phases.

2. The remote sensing image change detection method based on the dual-stream temporal feature adapter according to claim 1 is characterized in that: The remote sensing image of the front phase and the remote sensing image of the back phase are used as inputs of the encoder, and bidirectional dynamic fusion feature extraction processing of n stages is performed through the pre-trained basic visual Transformer model backbone network of the encoder and the two phase feature adapter paths to obtain the global context feature map of the front phase and the global context feature map of the back phase of n stages, including: The remote sensing image of the front phase is input into a plurality of convolutional layers in the first phase feature adapter path of the encoder for processing to obtain front phase feature maps of different resolutions; The remote sensing image of the posterior phase is input into a plurality of convolutional layers in the second phase feature adapter path of the encoder for processing to obtain posterior phase feature maps of different resolutions; The remote sensing image of the front phase and the remote sensing image of the back phase are input into the patch embedding layer of the pre-trained basic visual Transformer model backbone network of the encoder for processing to obtain a high-dimensional vector sequence of the front phase and a high-dimensional vector sequence of the back phase; Based on the pre-phase feature maps of different resolutions, the post-phase feature maps of different resolutions, the pre-phase high-dimensional vector sequence and the post-phase high-dimensional vector sequence, n stages of bidirectional dynamic fusion feature extraction processing are performed through n block processing blocks and n bidirectional gating modules on two phase feature adapter paths to obtain n stages of pre-phase global context feature maps and post-phase global context feature maps.

3. The remote sensing image change detection method based on the dual-stream temporal feature adapter according to claim 2 is characterized in that: Each stage is processed in the same way, and the output of the previous stage is the input of the next stage; No. one The processing steps of each stage are: Based on the pre-phase feature maps of different resolutions, the pre-phase high-dimensional vector sequence and the pre-phase global context feature map of the first stage output by the first block processing block, the pre-phase fusion feature map of the first stage and the pre-phase fusion feature map after addition are processed by the first bidirectional gating module in the first phase feature adapter path to obtain the pre-phase fusion feature map of the first stage, wherein the pre-phase fusion feature map of the first stage is input into the first block processing block, and the pre-phase fusion feature map after addition of the first stage is input into the second bidirectional gating module in the first phase feature adapter path; Based on the post-phase feature maps of different resolutions, the post-phase high-dimensional vector sequence and the post-phase global context feature map of the first stage output by the first block processing block, the post-phase fusion feature map of the first stage and the added post-phase fusion feature map are processed by the first bidirectional gating module in the second phase feature adapter path to obtain the post-phase fusion feature map of the first stage, wherein the post-phase fusion feature map of the first stage is input into the first block processing block, and the added post-phase fusion feature map of the first stage is input into the second bidirectional gating module in the second phase feature adapter path; Based on the previous phase high-dimensional vector sequence and the later phase high-dimensional vector sequence, the previous phase fusion feature map of the first stage and the later phase fusion feature map of the first stage, the first block processing block is used to process the previous phase global context feature map and the later phase global context feature map of the first stage, wherein the previous phase global context feature map of the first stage is input into the second block processing block and the first bidirectional gating module in the first phase feature adapter path, and the later phase global context feature map of the first stage is input into the second block processing block and the first bidirectional gating module in the second phase feature adapter path.

4. The remote sensing image change detection method based on the dual-stream temporal feature adapter according to claim 3 is characterized in that: The structure of each bidirectional gating module is the same, and the bidirectional gating module includes a multi-receptive field feature pyramid module, a dynamic temporal perception enhancement module, a first bidirectional interaction module and a second bidirectional interaction module; Among them, the multi-receptive field feature pyramid module is used to process the input feature map based on receptive fields of different sizes to form a multi-scale feature map; the dynamic temporal perception enhancement module is used to process the multi-scale feature map output by the multi-receptive field feature pyramid module, construct the temporal features within a single phase, and generate an enhanced multi-scale feature map; the first two-way interactive module is used to interactively fuse the output of the dynamic temporal perception enhancement module at the current stage with the feature map output by the block processing block at the previous stage and output the first fused feature map of the current stage; the second two-way interactive module is used to interactively fuse the output of the dynamic temporal perception enhancement module at the current stage with the feature map output by the block processing block at the current stage and output the second fused feature map of the current stage; the output of the dynamic temporal perception enhancement module at the current stage is added to the second fused feature map of the current stage to obtain the added fused feature map of the current stage.

5. The remote sensing image change detection method based on the dual-stream temporal feature adapter according to claim 4 is characterized in that: The multi-receptive field feature pyramid module reduces the dimension of the received feature map through the first linear projection layer, divides the reduced feature map into multiple groups along the channel dimension, each group corresponds to a receptive field of a different size, and each group is further restored to its original dimension through a second linear projection layer after a depth-wise separable convolution operation to form a multi-scale feature map.

6. The remote sensing image change detection method based on a dual-stream temporal feature adapter according to claim 4 is characterized in that: The dynamic temporal perception enhancement module receives a multi-scale feature map outputted from a multi-receptive field feature pyramid module, wherein each scale feature map of the multi-scale feature map corresponds to a different spatial resolution; Extracting global statistical information from feature maps of each scale through a global average pooling operation, generating channel-level statistical vectors, inputting the statistical vectors into a first multi-layer perceptron, generating modulation vectors through a nonlinear transformation, and inputting the statistical vectors into a second multi-layer perceptron to generate channel weights; The feature maps of each scale are input into the corresponding depth-wise separable convolution for processing to obtain the processed feature maps; Using each of the modulation vectors to scale the corresponding processed feature maps channel by channel to obtain each modulated feature map; Performing channel reweighting on each of the modulated feature maps according to each of the channel weights to obtain each reweighted feature map; The received multi-scale feature map is added to the corresponding reweighted feature map through residual connection, and layer normalization operation is applied to generate the enhanced multi-scale feature map.

7. The remote sensing image change detection method based on a dual-stream temporal feature adapter according to claim 3 is characterized in that: The method of obtaining the global context feature map of the first phase and the global context feature map of the second phase based on the high-dimensional vector sequence of the previous phase and the high-dimensional vector sequence of the next phase, the fusion feature map of the previous phase in the first stage and the fusion feature map of the next phase in the first stage by processing through the first block processing block includes: Performing an addition operation on the previous phase high-dimensional vector sequence and the previous phase fusion feature map of the first stage to obtain an added previous phase feature map; Performing an addition operation on the post-phase high-dimensional vector sequence and the post-phase fusion feature map of the first stage to obtain an added post-phase feature map; The first block is used to process the pre-addition phase feature map to obtain the pre-phase global context feature map of the first stage; The post-addition phase feature map is processed by the first block processing block to obtain the post-phase global context feature map of the first stage.

8. The remote sensing image change detection method based on a dual-stream temporal feature adapter according to any one of claims 1 to 7, characterized in that: The method of inputting the global context feature map of the previous phase and the global context feature map of the next phase of n stages into the decoder for decoding to obtain the change detection mask of the remote sensing image of the previous and next phases includes: The global context feature maps of the previous phase and the global context feature maps of the subsequent phase of n stages are respectively input into the decoder, and the pixel-level difference operation is performed on the global context feature maps of the previous phase and the global context feature maps of the subsequent phase of the corresponding stage to generate each difference feature map; then, the difference feature maps are aggregated based on a lightweight multi-layer perceptron decoder to obtain the change detection mask of the remote sensing image of the previous and next phases.

9. The remote sensing image change detection method based on a dual-stream temporal feature adapter according to any one of claims 1 to 7, characterized in that: n≥4。 10. A remote sensing image change detection system based on a dual-stream temporal feature adapter, characterized in that: It includes an image acquisition module, an encoder and a decoder, wherein: The image acquisition module is used to acquire the remote sensing image of the front phase and the remote sensing image of the back phase; The encoder is used to perform n-stage bidirectional dynamic fusion feature extraction processing on the input remote sensing image of the front phase and the remote sensing image of the back phase through the pre-trained basic visual Transformer model backbone network of the encoder and two phase feature adapter paths to obtain the n-stage front phase global context feature map and the back phase global context feature map, wherein the pre-trained basic visual Transformer model backbone network includes a patch embedding layer and n sequentially connected block processing blocks, each phase feature adapter path includes multiple convolutional layers and n sequentially connected bidirectional gating modules, each block processing block corresponds to a bidirectional gating module on each phase feature adapter path to perform one stage of bidirectional dynamic fusion feature extraction processing to obtain the front phase global context feature map and the back phase global context feature map of the stage, and n is a positive integer; The decoder is used to respectively decode the input n-stage front-phase global context feature map and the back-phase global context feature map to obtain the change detection mask of the remote sensing image of the front- and back-phases.

Citation Information

Patent Citations

  • Remote sensing image change detection method combining convolutional neural network and Transform

    CN116402766A

  • Self-attention feature fused high-resolution remote sensing image semantic change detection method

    CN116486255A

  • Instance constraint change detection method and device for broken image spots of remote sensing image

    CN117253157A

  • Remote sensing image building change detection method based on multi-scale attention

    CN118298305A

  • Dual-time remote sensing image semantic change detection method based on twin residual network

    CN118429819A