Method and device for detecting change of remote sensing image, and electronic equipment

By combining the advantages of CNN and Transformer, a deep learning network with a hybrid encoder and a multi-scale progressive fusion decoder is used to solve the problems of low detection accuracy and high computational complexity in remote sensing image change detection, achieving efficient feature extraction and accurate detection.

CN121810657APending Publication Date: 2026-04-07CHONGQING UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing remote sensing image change detection methods suffer from low accuracy when faced with complex ground features and variable environmental conditions in high-resolution remote sensing images. They also struggle to effectively balance local feature extraction with global context modeling and have high computational complexity.

Method used

A deep learning network employing a hybrid encoder and a multi-scale progressive fusion decoder leverages the local feature extraction capabilities of CNNs and the global modeling advantages of Transformers, combined with a variation feature extraction module and a multi-level progressive attention module, to achieve the fusion and enhancement of multi-level variation features.

Benefits of technology

It improves the accuracy and robustness of remote sensing image change detection, significantly enhances the model's practicality and deployment efficiency, and reduces computational complexity while maintaining high detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121810657A_ABST
    Figure CN121810657A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing image processing, and discloses a remote sensing image change detection method, which comprises the following steps: acquiring a dual-time-phase to-be-detected remote sensing image; inputting the dual-time-phase to-be-detected remote sensing image into a preset remote sensing image change detection model to obtain a detection result of the target area; the remote sensing image change detection model comprises a hybrid encoder and a multi-scale progressive fusion decoder, and the hybrid encoder is used for extracting a multi-scale feature map of a dual-time-phase remote sensing image to be detected; a change feature extraction module of the multi-scale progressive fusion decoder is used for extracting a change feature map of a corresponding scale from each scale feature map; and a multi-stage progressive attention module of the multi-scale progressive fusion decoder is used for carrying out feature enhancement on the change feature patterns of each scale, and carrying out channel splicing and fusion on each enhanced change feature pattern to obtain a detection result. The method can improve the accuracy of remote sensing image change detection. The invention further discloses a remote sensing image change detection device and electronic equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of remote sensing image processing technology, such as a method, apparatus, and electronic device for detecting changes in remote sensing images. Background Technology

[0002] Remote sensing change detection, as one of the core technologies in the field of remote sensing image processing, has always been committed to accurately identifying surface change information from registered remote sensing images of the same region at different time phases. This technology plays an irreplaceable role in many important areas of the national economy, such as urban and rural development monitoring, environmental change assessment, disaster emergency response, and agricultural resource surveys. With the continuous improvement of high-resolution Earth observation systems, the acquisition of massive amounts of multi-temporal remote sensing data has provided a rich data foundation for surface change monitoring.

[0003] Traditional change detection methods have evolved from algebraic computation-based to image transformation-based and then to classifier-based methods. Early algebraic methods such as image interpolation and change vector analysis, while computationally simple, were extremely sensitive to image registration accuracy and radiometric correction quality, easily producing numerous false detections in complex scenarios. Image transformation-based methods such as principal component analysis and wavelet transform improved feature representation capabilities to some extent, but remained limited by manually designed feature extraction strategies. Machine learning classifier methods such as support vector machines and random forests improved detection accuracy by introducing contextual information; however, their feature learning capabilities were limited, making them difficult to adapt to diverse development scenarios. These traditional methods often exhibit significant limitations when facing the complex ground features and variable environmental conditions presented by high-resolution remote sensing imagery, resulting in low accuracy in remote sensing image change detection.

[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.

[0006] This disclosure provides a method, apparatus, and electronic device for detecting changes in remote sensing images, which can improve the accuracy of remote sensing image change detection.

[0007] In some embodiments, the method includes: acquiring dual-temporal remote sensing images to be measured, wherein the dual-temporal remote sensing images to be measured are remote sensing images of a target area at two different times; inputting the dual-temporal remote sensing images to be measured into a preset remote sensing image change detection model to obtain a detection result of the target area; wherein the remote sensing image change detection model includes: a hybrid encoder and a multi-scale progressive fusion decoder, the hybrid encoder including multiple feature extraction modules, each feature extraction module being composed of a preset residual module and a preset attention module, the hybrid encoder being used to extract multi-scale feature maps of the dual-temporal remote sensing images to be measured; the multi-scale progressive fusion decoder including a change feature extraction module and a multi-level progressive attention module; the change feature extraction module being used to extract change feature maps of corresponding scales from the feature maps of each scale; the multi-level progressive attention module being used to perform feature enhancement on the change feature maps of each scale to obtain enhanced change feature maps of the corresponding scale, and to perform channel stitching and fusion on the enhanced change feature maps to obtain a detection result.

[0008] In some embodiments, the apparatus includes: an acquisition module configured to acquire dual-temporal remote sensing images to be measured, wherein the dual-temporal remote sensing images to be measured are remote sensing images of a target area at two different times; and a detection module configured to input the dual-temporal remote sensing images to be measured into a preset remote sensing image change detection model to obtain a detection result of the target area; wherein the remote sensing image change detection model includes: a hybrid encoder and a multi-scale progressive fusion decoder, the hybrid encoder including multiple feature extraction modules, each feature extraction module consisting of a preset residual module and a preset attention module, the hybrid encoder being used to extract multi-scale feature maps of the dual-temporal remote sensing images to be measured; the multi-scale progressive fusion decoder including a change feature extraction module and a multi-level progressive attention module; the change feature extraction module being used to extract change feature maps of corresponding scales from the feature maps of each scale; the multi-level progressive attention module being used to perform feature enhancement on the change feature maps of each scale to obtain enhanced change feature maps of the corresponding scale, and to perform channel stitching and fusion on the enhanced change feature maps to obtain a detection result.

[0009] In some embodiments, the electronic device includes a processor and a memory storing program instructions, the processor being configured to execute the method described above for detecting changes in remote sensing images when the program instructions are executed.

[0010] The method, apparatus, and electronic equipment for detecting changes in remote sensing images provided in this disclosure can achieve the following technical effects: By inputting dual-temporal remote sensing images into a pre-defined remote sensing image change detection model, the model is tested. This model comprises a hybrid encoder and a multi-scale progressive fusion decoder. The hybrid encoder's multiple feature extraction modules consist of pre-defined residual modules and pre-defined Transformer modules, while the multi-scale progressive fusion decoder includes a change feature extraction module and a multi-level progressive attention module. This leverages the local feature extraction capabilities of CNNs (Convolutional Neural Networks) to effectively enhance the model's feature representation ability and training stability. The global modeling advantages of Transformers enable efficient global information aggregation across multiple scales. Furthermore, the collaborative work of the change feature extraction module and the multi-level progressive attention module achieves full fusion and semantic enhancement of multi-level change features, thereby effectively improving the accuracy and robustness of remote sensing image change detection. Therefore, using this remote sensing image change detection model to detect dual-temporal remote sensing images can improve the accuracy of remote sensing image change detection.

[0011] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description

[0012] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein: Figure 1 This is a schematic diagram of a method for detecting changes in remote sensing images provided in an embodiment of this disclosure; Figure 2 This is a schematic diagram of the structure of a change detection model provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram of the structure of a preset residual module provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of the structure of a sliding window attention module provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of the structure of a change feature extraction submodule provided in an embodiment of this disclosure; Figure 6 This is a schematic diagram of the structure of an attention submodule provided in an embodiment of this disclosure; Figure 7 This is a visualization comparison of the remote sensing image change detection method provided in this embodiment with other traditional change detection methods in a remote sensing scene; Figure 8This is a schematic diagram of an apparatus for detecting changes in remote sensing images provided in an embodiment of this disclosure; Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0013] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.

[0014] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0015] Unless otherwise stated, the term "multiple" means two or more.

[0016] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0017] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0018] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.

[0019] The method for detecting changes in remote sensing images provided in this disclosure is applied to electronic devices, including servers or computers.

[0020] When processing multi-temporal and multi-scale remote sensing data, related technologies often struggle to effectively balance local feature extraction with global context modeling, resulting in limited detection accuracy in complex terrain backgrounds. Furthermore, these technologies lack sufficient ability to suppress spurious changes and suffer from high model computational complexity, limiting their application in practical engineering. The remote sensing image change detection method proposed in this disclosure first ensures the quality and consistency of input data through a preprocessing workflow, including radiometric correction, atmospheric correction, image registration, and cropping. Then, a deep learning network based on a hybrid encoder and a multi-scale progressive fusion decoder is constructed. The hybrid encoder achieves effective fusion of local features and global context by alternately stacking improved residual blocks and Transformer modules. The multi-scale progressive fusion decoder achieves progressive fusion and enhancement of multi-level features through the collaborative work of a change feature extraction module and a multi-level progressive attention module. The remote sensing image change detection method disclosed in this embodiment constructs an efficient feature extraction system by leveraging the complementary advantages of CNN and Transformer. It achieves accurate extraction of change features through a dual-branch feature interaction mechanism in the change feature extraction module, and ensures the integrity of feature representation through a multi-scale progressive fusion strategy in a multi-level progressive attention module. Thus, it significantly improves the practicality and deployment efficiency of the model while maintaining high detection accuracy.

[0021] Combination Figure 1 As shown, this disclosure provides a method for detecting changes in remote sensing images, including: Step S101: Acquire dual-temporal remote sensing images of the target area at two different times. The dual-temporal remote sensing images include a first-temporal image acquired earlier and a second-temporal image acquired later. The interval between the acquisition time of the first-temporal image and the acquisition time of the second-temporal image is a preset duration threshold.

[0022] In some embodiments, initial remote sensing images of the target area in a first time phase and initial remote sensing images in a second time phase are acquired via a preset Earth observation satellite. A preset time interval threshold is set between the acquisition time of the initial remote sensing images in the first time phase and the acquisition time of the initial remote sensing images in the second time phase. Preprocessing operations, including radiometric correction, atmospheric correction, image registration, and cropping, are sequentially performed on the initial remote sensing images of the first and second time phases. The preprocessed initial remote sensing images of the target area in the first and second time phases are then identified as dual-temporal remote sensing images to be measured.

[0023] Step S102: Input the dual-temporal remote sensing images to be tested into a preset remote sensing image change detection model to obtain the detection results of the target area. The remote sensing image change detection model includes a hybrid encoder and a multi-scale progressive fusion decoder. The hybrid encoder includes multiple feature extraction modules, each composed of a preset residual module and a preset Transformer module. The hybrid encoder is used to extract multi-scale feature maps from the dual-temporal remote sensing images to be tested. The multi-scale progressive fusion decoder includes a change feature extraction module and a multi-level progressive attention module. The change feature extraction module is used to extract change feature maps of corresponding scales from the feature maps at each scale. The multi-level progressive attention module is used to enhance the change feature maps at each scale, obtaining enhanced change feature maps of the corresponding scale, and then performing channel stitching and fusion on the enhanced change feature maps to obtain the detection results.

[0024] The method for remote sensing image change detection provided in this disclosure involves inputting dual-temporal remote sensing images to a preset remote sensing image change detection model. This model includes a hybrid encoder and a multi-scale progressive fusion decoder. The hybrid encoder's multiple feature extraction modules are composed of preset residual modules and preset Transformer modules, respectively. The multi-scale progressive fusion decoder includes a change feature extraction module and a multi-level progressive attention module. This leverages the local feature extraction capabilities of CNNs (Convolutional Neural Networks) to effectively enhance the model's feature representation ability and training stability. The global modeling advantages of Transformers enable efficient global information aggregation across multiple scales. Furthermore, the collaborative work of the change feature extraction module and the multi-level progressive attention module achieves full fusion and semantic enhancement of multi-level change features, thereby effectively improving the accuracy and robustness of remote sensing image change detection. Therefore, using this remote sensing image change detection model to detect dual-temporal remote sensing images can improve the accuracy of remote sensing image change detection.

[0025] Optionally, the preset remote sensing image change detection model is obtained by acquiring multiple pairs of dual-temporal sample remote sensing images, where the dual-temporal sample remote sensing images are remote sensing images of a preset area at two different times. Each pair of dual-temporal sample remote sensing images is sequentially input into the preset change detection model for iterative training until a preset stopping condition is met. The trained change detection model is then determined as the remote sensing image change detection model.

[0026] In some embodiments, a preset number of pairs of dual-temporal sample remote sensing images are acquired. The preset number is greater than 200. Each pair of dual-temporal sample remote sensing images includes a first-temporal sample remote sensing image acquired first and a second-temporal sample remote sensing image acquired later. The interval between the acquisition time of the first-temporal sample remote sensing image and the acquisition time of the second-temporal sample remote sensing image is a preset duration threshold.

[0027] High-resolution remote sensing images of a preset region are acquired using multi-source Earth observation satellites at preset time intervals, resulting in multiple pairs of dual-temporal raw remote sensing images for the preset region. Remote sensing images with a resolution between 0.7m and 2m are considered high-resolution. Each pair of dual-temporal raw remote sensing images for the preset region includes a first-temporal raw remote sensing image acquired earlier and a second-temporal raw remote sensing image acquired later. The interval between the acquisition times of the first and second-temporal raw remote sensing images is the preset time threshold. Each pair of dual-temporal raw remote sensing images for the preset region undergoes preprocessing operations including radiometric correction, atmospheric correction, image registration, and cropping. The preprocessed dual-temporal raw remote sensing images of the preset region are then identified as dual-temporal sample remote sensing images. Preprocessing multiple pairs of dual-temporal raw remote sensing images for the preset region eliminates various noises and geometric distortions in the raw remote sensing images, transforming them into high-quality data with unified standards that can be directly used for model analysis. Furthermore, it can eliminate spatial deviations caused by factors such as sensor perspective and terrain undulations, thereby providing a reliable data foundation for subsequent change detection algorithms with pixel-level spatial alignment, laying a key prerequisite for model training and application.

[0028] In some embodiments, radiometric calibration and radiometric correction are performed on each original remote sensing image of a preset area to convert the raw digital quantization values ​​recorded by the sensor into apparent surface reflectance with clear physical meaning, eliminating the influence of differences in sensor response. Subsequently, a preset atmospheric correction model is used to further remove interference from atmospheric molecules, aerosol scattering and absorption, and to invert the true surface reflectance of the land cover, thereby reducing spectral distortion caused by different atmospheric conditions at the time of imaging in the original remote sensing images of different time phases. Based on this, image registration is performed on the dual-time phase original remote sensing images that have undergone radiometric and atmospheric correction. Specifically, by selecting a preset number of control points, a polynomial or triangular mesh algorithm is used to achieve sub-pixel-level spatial alignment to ensure that the position of the same land cover is completely consistent in the remote sensing images of different time phases. Finally, the registered remote sensing images are cropped according to the vector boundary of the preset area to obtain remote sensing image pairs with consistent range and pixel alignment. In this way, by preprocessing each pair of dual-temporal original remote sensing images of the preset area, the consistency of the input data in radiometric, spectral, and spatial dimensions is ensured to the greatest extent, providing reliable and high-quality input for the subsequent change detection model, which is a key prerequisite for improving model performance and result accuracy.

[0029] In one embodiment, Figure 2 This is a schematic diagram of the change detection model. (Combined with...) Figure 2 As shown, optionally, the change detection model includes a hybrid encoder and a multi-scale progressive fusion decoder to form an end-to-end change detection network. The multi-scale progressive fusion decoder includes a change feature extraction module and a multi-level progressive attention module. The following operations are performed during each training iteration of the change detection model: The dual-temporal sample remote sensing images are input into a hybrid encoder, which extracts multi-scale feature maps from the images. A change feature extraction module extracts change feature maps for each scale of the feature maps. A multi-level progressive attention module enhances the change feature maps at each scale, resulting in enhanced change feature maps for each scale. Finally, channel stitching and fusion of these enhanced change feature maps outputs the change image.

[0030] Optionally, the hybrid encoder includes multiple feature extraction modules, each of which consists of a preset residual module and a preset attention module. The hybrid encoder includes: a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, a fifth feature extraction module, a sixth feature extraction module, a seventh feature extraction module, and an eighth feature extraction module. The preset residual modules include: a first residual module, a second residual module, a third residual module, a fourth residual module, a fifth residual module, a sixth residual module, a seventh residual module, an eighth residual module, a ninth residual module, and a tenth residual module. The preset attention modules include: a first attention module, a second attention module, a third attention module, a fourth attention module, a fifth attention module, a sixth attention module, a seventh attention module, and an eighth attention module.

[0031] like Figure 2As shown, the first feature extraction module includes a first residual module 201, a second residual module 202, and a first attention module 203; the second feature extraction module includes a third residual module 204, a fourth residual module 205, and a second attention module 206; the third feature extraction module includes a fifth residual module 207 and a third attention module 208; the fourth feature extraction module includes a sixth residual module 209 and a fourth attention module 210; the fifth feature extraction module includes a seventh residual module 211 and a fifth attention module 212; the sixth feature extraction module includes an eighth residual module 213 and a sixth attention module 214; the seventh feature extraction module includes a ninth residual module 215 and a seventh attention module 216; and the eighth feature extraction module includes a tenth residual module 217 and an eighth attention module 218. Thus, the hybrid encoder, by alternating residual modules and attention modules, achieves multi-scale feature extraction and long-range dependency modeling of remote sensing images, thereby obtaining high-level feature representations with rich semantic information while maintaining computational efficiency.

[0032] Extracting multi-scale feature maps from the dual-temporal sample remote sensing images using the hybrid encoder includes: inputting a first-temporal sample remote sensing image into a first feature extraction module to obtain a first-scale feature map of the first-temporal sample remote sensing image; inputting a second-temporal sample remote sensing image into a second feature extraction module to obtain a first-scale feature map of the second-temporal sample remote sensing image; inputting the first-scale feature map of the first-temporal sample remote sensing image into a third feature extraction module to obtain a second-scale feature map of the first-temporal sample remote sensing image; inputting the first-scale feature map of the second-temporal sample remote sensing image into a fourth feature extraction module to obtain a second-scale feature map of the second-temporal sample remote sensing image; inputting the second-scale feature map of the first-temporal sample remote sensing image into a fifth feature extraction module to obtain a third-scale feature map of the first-temporal sample remote sensing image; inputting the second-scale feature map of the second-temporal sample remote sensing image into a sixth feature extraction module to obtain a third-scale feature map of the second-temporal sample remote sensing image; and inputting the third-scale feature map of the first-temporal sample remote sensing image into a seventh feature extraction module to obtain a fourth-scale feature map of the first-temporal sample remote sensing image. The third-scale feature map of the second-temporal sample remote sensing image is input into the eighth feature extraction module to obtain the fourth-scale feature map of the second-temporal sample remote sensing image.

[0033] like Figure 2 As shown, the first temporal sample remote sensing image Image-T1 has a size of H×W×3. Here, H is the height of the sample remote sensing image, W is the width of the sample remote sensing image, and 3 is the number of channels. Inputting the first temporal sample remote sensing image Image-T1 into the first residual module 201 of the first feature extraction module yields a result of size... The feature map is obtained from the first residual module and then passed through the second residual module 202 and the first attention module 203 to obtain a feature map of size . The feature map, i.e., the first-scale feature map of the first-phase sample remote sensing image. , The first-scale feature map of the first-phase sample remote sensing image is input into the third feature extraction module, and then sequentially passes through the fifth residual module 207 and the third attention module 208 to obtain a feature map of size [size missing]. The feature map, i.e., the second-scale feature map of the first-phase sample remote sensing image. , The second-scale feature map of the first temporal sample remote sensing image is input into the fifth feature extraction module, and then sequentially passes through the seventh residual module 211 and the fifth attention module 212 to obtain a feature map of size [value missing]. The feature map, namely the third-scale feature map of the first-phase sample remote sensing image. , The third-scale feature map of the first time-phase sample remote sensing image is input into the seventh feature extraction module, and then sequentially passes through the ninth residual module 215 and the seventh attention module 216 to obtain a feature map of size [value missing]. The feature map, namely the fourth-scale feature map of the first-phase sample remote sensing image. , .

[0034] like Figure 2 As shown, the second temporal sample remote sensing image Image-T2 has a size of H×W×3. Inputting the second temporal sample remote sensing image Image-T2 into the third residual module 204 of the second feature extraction module yields a result of size H×W×3. The feature map obtained from the third residual module is then input into the fourth residual module 205 and the second attention module 206 to obtain a feature map of size 205. The feature map, namely the first-scale feature map of the second-phase sample remote sensing image. , The first-scale feature map of the second-phase sample remote sensing image is input into the fourth feature extraction module, and then sequentially passes through the sixth residual module 209 and the fourth attention module 210 to obtain a feature map of size [size missing]. The feature map, namely the second-scale feature map of the second-phase sample remote sensing image. , The second-scale feature map of the second-phase sample remote sensing image is input into the sixth feature extraction module, and then sequentially passes through the eighth residual module 213 and the sixth attention module 214 to obtain a feature map of size [value missing]. The feature map, namely the third-scale feature map of the second-phase sample remote sensing image. , The third-scale feature map of the second-phase sample remote sensing image is input into the eighth feature extraction module, and then sequentially passes through the tenth residual module 217 and the eighth attention module 218 to obtain a feature map of size [size missing]. The feature map, namely the fourth-scale feature map of the second-phase sample remote sensing image. , .

[0035] For the first temporal sample remote sensing image, the hybrid encoder performs multi-level downsampling operations through the first and second residual modules of the first feature extraction module. For the second temporal sample remote sensing image, the hybrid encoder performs multi-level downsampling operations through the third and fourth residual modules of the second feature extraction module, achieving a four-fold downsampling of the original remote sensing images of both temporal phases to extract more discriminative local features. Then, for the first-scale feature map of the first temporal sample remote sensing image, the third, fifth, and seventh feature extraction modules are sequentially passed through the third, fourth, sixth, and eighth feature extraction modules, respectively, to extract local texture features. These features are then encoded at the corresponding attention modules at the four scales, thereby obtaining deep semantic features with rich contextual information at four resolutions: 1 / 4, 1 / 8, 1 / 16, and 1 / 32.

[0036] In some embodiments, Figure 3 This is a schematic diagram of the pre-defined residual block structure, combined with... Figure 3As shown, the preset residual module includes one downsampled residual block 301 and N standard residual blocks 302. The downsampled residual block 301 contains a main branch and a skip connection branch. The feature map of the input residual module first passes through the first convolutional layer on the main branch of the downsampled residual block 301. The first convolutional layer is a 3×3 convolution (Conv3×3) with a stride of 2 to achieve a 2x downsampling. It then sequentially undergoes a first batch normalization (Bn) process, a first ReLU activation function, and then a second convolutional layer. The second convolutional layer is a 3×3 convolution (Conv3×3) with a stride of 1. Finally, a second batch normalization (Bn) process is performed to obtain the output feature map of the main branch. The skip connection branch passes through a third convolutional layer, which is a 5×5 convolution with a stride of 2 (Conv5×5) for double downsampling and a third batch normalization (Bn) process. The output feature maps of the two branches are perfectly matched in both spatial and channel dimensions. The output feature maps of the main branch and the skip connection branch are fused element-wise, then passed through a second ReLU activation function to obtain the output feature map of the downsampled residual block. This downsampled residual block output feature map is then input into the first standard residual block. Each standard residual block has the same structure. In the standard residual block, the data first passes through a fourth convolutional layer (a 3×3 convolution with a stride of 1, Conv3×3), then through a fourth batch normalization (Bn) process, a third ReLU activation function, and a fifth convolutional layer (also a 3×3 convolution with a stride of 1), followed by a fifth batch normalization (Bn) process. The resulting intermediate feature map is then element-wise added to the input feature map of the standard residual block via skip connections. After passing through the fourth ReLU activation function, the output feature map of the first standard residual block is output. This output feature map is then input to the next standard residual block, and so on, until the output feature map of the Nth standard residual block is obtained. The output feature map of the Nth standard residual block is the output feature map of the corresponding residual module.

[0037] The first and second residual modules have the same structure, each consisting of one downsampled residual block and two standard residual blocks. The third and fourth residual modules also have the same structure, each consisting of one downsampled residual block and three standard residual blocks. The fifth and sixth residual modules have the same structure, each consisting of one downsampled residual block and four standard residual blocks. The seventh and eighth residual modules have the same structure, each consisting of one downsampled residual block and six standard residual blocks. The ninth and tenth residual modules also have the same structure, each consisting of one downsampled residual block and three standard residual blocks. In this way, by stacking multiple standard residual blocks within each residual module, deep semantic features are extracted progressively while maximizing the preservation of spatial details helpful for change detection. This provides a high-quality feature foundation for subsequent Transformer global modeling and change feature extraction and fusion.

[0038] Each residual module first achieves feature compression and spatial downsampling through downsampled residual blocks, and then deepens the network and enhances semantic expressive power through stacked standard residual blocks. The convolutional operations in the skip connection branches of the downsampled residual blocks in each residual module use 5×5 convolutional kernels, providing a larger receptive field that better matches the combined receptive field of the two 3×3 convolutional layers in the main branch. Furthermore, the larger kernels capture richer local contextual information, effectively preserving key spatial features such as edges and textures. Simultaneously, by maintaining consistency with the receptive field of the main branch, semantic alignment during feature fusion is ensured. This improves the continuity and stability of feature propagation.

[0039] In some embodiments, the first attention module includes two shift window attention blocks, the second attention module includes two shift window attention blocks, the third attention module includes two shift window attention blocks, the fourth attention module includes two shift window attention blocks, the fifth attention module includes 18 shift window attention blocks, the sixth attention module includes 18 shift window attention blocks, the seventh attention module includes two shift window attention blocks, and the eighth attention module includes two shift window attention blocks. The shift window attention blocks mainly consist of a multi-head self-attention mechanism, a shift window multi-head self-attention mechanism, a multilayer perceptron, a normalization layer, and residual connections. The core idea is to perform attention calculations within locally partitioned windows and achieve cross-window feature interaction through a shift window strategy, thereby balancing local feature sensitivity and global contextual relevance. Thus, the attention blocks in each feature extraction module adopt a structure design based on a shift window self-attention mechanism to effectively reduce computational complexity while ensuring global feature modeling capabilities, thereby achieving efficient spatiotemporal feature interaction and representation enhancement. Furthermore, each feature extraction module stacks a different number of shift window attention blocks. This hierarchical design enables the network to progressively enhance its global feature modeling capabilities at different scales. The shallow stages primarily capture local spatial features, the mid-level stages establish moderate dependencies, and the deep stages focus on integrating global semantic information. In particular, the stacking of 18 sliding window attention modules in the third stage significantly enhances the model's ability to model complex spatial contexts, providing rich global information support for accurately identifying changing regions. This improves the feature representation and change characterization capabilities of remote sensing imagery across multiple spatial scales. Through this progressive feature enhancement strategy, the attention modules and residual modules effectively complement each other, jointly constructing a highly efficient feature extraction system that combines local detail preservation with global context modeling capabilities.

[0040] In one embodiment, Figure 4 This is a schematic diagram of the structure of a Shift Window AttentionBlock, combined with... Figure 4 As shown, for example, The feature map of the input sliding window attention module is used. l For positive integers greater than or equal to 1, this feature map First, the feature map undergoes a first-layer normalization (LayerNorm, LN) to obtain the feature matrix T. This matrix is ​​then input into a window-based multi-head self-attention (W-MHSA) module for local relation modeling. Finally, the input feature map is processed through a first residual connection. The first intermediate feature is obtained by adding it to the processing result of the W-MHSA module. Next, the first intermediate feature The first intermediate features are then processed sequentially through a second normalized LN layer and a first multilayer perceptron (MLP) for nonlinear transformation, and then connected through a second residual connection. The second intermediate feature is obtained by adding it to the processing result of the first MLP. Subsequently, the second intermediate feature Moving to the next stage, the data first passes through a third normalized LN layer, then is input into a Shifted Window-based Multi-Head Self-Attention (SW-MHSA) module for processing to achieve cross-window information interaction. Similarly, the processing result of the SW-MHSA module is connected to the second intermediate feature through a third residual connection. Adding them together yields the third intermediate feature. Finally, the third intermediate feature The third intermediate feature is then processed sequentially through a fourth normalized LN layer and a second multilayer perceptron (MLP), and then connected via a fourth residual connection. The output feature map of the sliding window attention module is obtained by adding the result of the second MLP. The output feature map Input the next sliding window attention module. The feature map output by the last sliding window attention module is the output result of the corresponding attention module.

[0041] In its implementation, the sliding window attention module alternates between window multi-head self-attention and shifted window multi-head self-attention mechanisms to achieve feature interaction and global information fusion between adjacent windows. Feature updates are achieved by combining a multilayer perceptron and layer normalization. The shifted window mechanism achieves cross-window information interaction by cyclically shifting the window position while maintaining computational efficiency. The overall computation process of the sliding window attention module is as follows:

[0042]

[0043]

[0044]

[0045] in, Indicates the first The output feature map of the multi-head self-attention W-MHSA module in the layer window. Indicates the first The output feature map of the first MLP layer. Indicates the first Output feature map of the sliding window multi-head self-attention SW-MHSA module with +1 layer Indicates the first The output feature map of the second MLP layer, each Each module includes a two-layer fully connected network and an activation function. The residual connection structure accelerates network convergence while maintaining stable feature propagation.

[0046] The self-attention mechanism is used to capture feature dependencies within a local window. Attention weights are obtained by calculating the similarity between the query matrix and the key matrix, and then the value matrices are weighted and summed. The sliding window attention module in this embodiment calculates attention within the divided local window, effectively reducing computational complexity. The feature matrix T obtained after the first layer of normalized LN is input into the W-MHSA module, and is first mapped to the query matrix through a linear transformation. ,key Sum The calculation process of windowed multi-head self-attention (W-MHSA) with three feature matrices is as follows:

[0047]

[0048]

[0049]

[0050] in, This is the output feature map of the W-MHSA module. For the feature matrix of the input W-MHSA module, These represent the weight matrices of the linear projection layer. These represent the query, key, and value matrices, respectively. M For window size, For channel dimension, It is a learnable position offset matrix used to supplement position information to enhance spatial modeling capabilities.

[0051] Combination Figure 2 As shown, optionally, the Change Feature Extraction Module (CFEM) includes: a first change feature extraction submodule 101, a second change feature extraction submodule 102, a third change feature extraction submodule 103, and a fourth change feature extraction submodule 104.

[0052] The change feature extraction module extracts change feature maps for each scale of feature maps, including: inputting the dual-temporal feature map of the first scale into the first change feature extraction submodule 101 to obtain the change feature map of the first scale. The first-scale dual-temporal feature map includes: the first-scale feature map of the first-temporal sample remote sensing image. First-scale feature map of second-phase sample remote sensing images The dual-temporal feature map at the second scale is input into the second change feature extraction submodule 102 to obtain the change feature map at the second scale. The second-scale dual-temporal feature map includes: the second-scale feature map of the first-temporal sample remote sensing image. Second-scale feature map of second-phase sample remote sensing images The dual-temporal feature map at the third scale is input into the third change feature extraction submodule 103 to obtain the change feature map at the third scale. The third-scale dual-temporal feature map includes: the third-scale feature map of the first-temporal sample remote sensing image. Third-scale feature map of second-phase sample remote sensing images The bi-temporal feature map at the fourth scale is input into the fourth change feature extraction submodule 104 to obtain the change feature map at the fourth scale. The fourth-scale dual-temporal feature map includes: the fourth-scale feature map of the first-temporal sample remote sensing image. Fourth-scale feature map of second-phase sample remote sensing images .

[0053] The change feature extraction module employs a channel attention weighting mechanism. By fusing differential feature extraction with channel weighting, it extracts significantly changing regions from dual-temporal feature maps at various scales and suppresses spurious change interference caused by seasonal, illumination, or vegetation changes. Specifically, the change feature extraction module combines global average pooling and max pooling strategies to generate a semantic weight map and uses a multilayer perceptron to adaptively allocate channel weights, outputting a highly discriminative change feature map.

[0054] In some embodiments, the first change feature extraction submodule, the second change feature extraction submodule, the third change feature extraction submodule, and the fourth change feature extraction submodule have the same structure. Figure 5 This is a schematic diagram of the variation feature extraction submodule, which uses a dual-branch structure for feature extraction and fusion processing. Figure 5 As shown, The first phase of remote sensing image i Scale feature map The second phase sample remote sensing image i Scale feature mapi Indicates the first i Each scale. The variation feature extraction submodule receives the input dual-temporal feature map. and The first branch takes the input biphase feature map... and Concatenation is performed along the channel dimension to obtain concatenated features. Then, a first convolutional layer (a 1×1 convolution, Conv1×1) is used for channel compression and feature fusion, followed by batch normalization and ReLU activation. Next, global average pooling (GAP) is performed along the spatial dimension to extract global semantic information. Finally, the pooling result is input into a first multilayer perceptron (MLP) for non-linear encoding to generate a first sequence containing global correlations. , for i The first sequence of scales is calculated using the following formula:

[0055] The second branch focuses on capturing significant differences between the two temporal feature maps. The second branch processes the input two temporal feature maps... and Element-wise subtraction and absolute value (Abstraction & Subtraction) is performed to highlight pixel-level changes. Then, a second convolutional layer (a 1×1 convolution) is used for channel enhancement. Next, global max pooling (GMP) is performed along the spatial dimension to extract responses from regions of significant change. Finally, the pooling result is input into a first multilayer perceptron (MLP) for encoding, resulting in the second sequence. , for i The second sequence of the scale is calculated using the following formula:

[0056] The first sequence With the second sequence Element-wise addition yields a joint feature representation that includes both global correlation and local difference information. This joint feature representation is then reconstructed using a second multilayer perceptron and a sigmoid activation function. Subsequently, the reconstructed vector is divided into two weight sequences through a segmentation operation. and , for i The first weight sequence of the scale corresponds to the channel weighting factor of the first phase sample remote sensing image. for iThe second weight sequence at the scale corresponds to the channel weighting factor of the second temporal sample remote sensing image. The calculation formula is as follows:

[0057] Will i First weight sequence of scale Temporal feature map of the input Element-wise product is performed along the channel dimension to obtain i The first weighted feature of the scale will i The second weight sequence of the scale Temporal feature map of the input Element-wise multiplication along the channel dimension yields... i The second weighted feature is then used to calculate the scale; the first weighted feature is then added to the second weighted feature (element-wise addition), and feature fusion and channel integration are performed through a 1×1 convolution to output the first weighted feature. i Scale variation feature map The calculation formula is as follows:

[0058] in, This indicates a splicing operation. This indicates the inclusion of batch normalization and ReLU activation functions. Convolutional layer, GAP GMP represents global average pooling. This represents global max pooling, MLP. This represents a multilayer perceptron. This indicates the operation of splitting the sequence in half.

[0059] The change feature extraction module achieves a synergistic effect of global semantic constraints and local significant change capture through a dual-branch structure. The first branch ensures that global contextual information is preserved and suppresses false change detection caused by large-scale environmental changes. The second branch emphasizes significantly changing pixels, improving sensitivity to small targets or edge changes. The fusion reconstruction unit achieves effective fusion of dual-temporal features through adaptively generated channel weights, so that the output feature map retains important change information while reducing background interference.

[0060] The change feature extraction module can extract highly discriminative change features in a multi-scale feature space, providing reliable input for the subsequent multi-scale progressive attention module, thereby significantly improving the accuracy and robustness of the overall change detection model in complex remote sensing scenarios.

[0061] Combination Figure 2As shown, optionally, the Multilevel Progressive Attention Module (MPAM) includes: a first attention submodule 105, a second attention submodule 106, a third attention submodule 107, and a fourth attention submodule 108.

[0062] Feature enhancement is performed on the feature maps at each scale using a multi-level progressive attention module to obtain enhanced feature maps at the corresponding scales, including: enhancing the feature map at the fourth scale. Input the fourth attention submodule 108 to obtain the enhanced change feature map at the fourth scale. Combine the enhanced change feature map at the fourth scale with the change feature map at the third scale. Input the third attention submodule to obtain the enhanced change feature map at the third scale. Combine the enhanced change feature map at the third scale with the change feature map at the second scale. Input the second attention submodule to obtain the enhanced change feature map at the second scale. Combine the enhanced change feature map at the second scale with the change feature map at the first scale. The first attention submodule is input to obtain the enhanced feature map at the first scale. The multi-level progressive attention module employs a pyramid-shaped feature fusion strategy to effectively fuse and enhance features at adjacent scales. This module enhances semantic expressiveness while preserving detailed information by combining top-down feature propagation and attention weighting mechanisms.

[0063] In some embodiments, the first attention submodule, the second attention submodule, and the third attention submodule have the same structure. Each of the first, second, and third attention submodules contains two branches. Figure 6 This is a structural diagram of the first, second, and third attention submodules, combined with... Figure 6 As shown, the first i- 1. Feature map of changes in the input received by the attention submodule and , i ∈{2, 3, 4}, For the first i Scale-enhanced variation feature map, For the first i- A feature map showing changes at scale 1. Through this... i- 1. Attention submodule, obtain the first i- 1-scale enhanced feature map. First branch, for the input... i Scale-enhanced variation feature map Perform bilinear upsampling to make its spatial resolution comparable to that of the first... i -1 scale variation feature map To maintain consistency, a 3×3 convolutional layer (Conv3×3) is then used to adjust the number of channels, thereby enhancing the feature maps in the lower layers. Capable of displaying the characteristics of changes at higher levels Alignment is performed along the channel dimension. This operation ensures the fusionability of multi-scale features while preserving as much fine-grained texture information as possible in low-level features, thus providing a reliable foundation for subsequent fusion. Then, the enhanced feature map after adjusting the number of channels is... With change feature map Concatenate along the channel dimension, then fuse using a 1×1 convolution (Conv1×1) to obtain preliminary fused features. The calculation formula is as follows:

[0064] The second branch will adjust the enhanced feature map after adjusting the number of channels. With change feature map Element-wise addition performs global average pooling and global max pooling on the added features along the channel dimension, concatenates the two pooling results, and generates a spatial attention weight map through a 1×1 convolution (Conv1×1). This spatial attention weight map is used to highlight important spatial regions and suppress irrelevant background. The calculation formula is as follows:

[0065] Subsequently, the generated spatial attention weight map Change characteristics diagram of high-level Element-wise product is performed to enhance space.

[0066] Then, the multiplied features are combined with the initial fused features. The features are added together and then integrated using a 1×1 convolution (Conv1×1) to obtain the first... i- 1-Scale Enhanced Change Feature Map. This operation ensures the full fusion of spatial details from low-level features and semantic information from high-level features, while dynamically emphasizing regions of significant change, thereby improving the accuracy and robustness of change detection. The calculation formula is as follows:

[0067] in, This indicates a bilinear interpolation upsampling operation. This indicates a splicing operation. This represents a 1×1 convolutional layer containing batch normalization and ReLU activation functions. This represents a 3×3 convolutional layer containing batch normalization and ReLU activation functions. This indicates an operation that performs average pooling and max pooling along the channel dimension and then concatenates the results.

[0068] In some embodiments, the fourth-scale change feature map is enhanced using only a spatial attention weight generation mechanism to optimize its semantic expressive power. Specifically, the fourth-scale change feature map... The fourth attention submodule is input, and global average pooling and global max pooling are performed along the channel dimension. The results of the two pooling operations are concatenated and then generated into a spatial attention weight map through a 1×1 convolution (Conv1×1). Subsequently, the generated spatial attention weight map will be... Change feature map with fourth scale Element-wise product is performed to obtain an enhanced change feature map at the fourth scale, achieving spatial augmentation. This provides a reliable high-level feature base for subsequent decoding or output modules, thereby improving the accuracy and robustness of the entire change detection network.

[0069] The multi-scale progressive attention module employs a pyramid-shaped multi-branch structure to achieve step-by-step fusion of multi-scale change features and restoration of spatial details. The first branch performs bilinear interpolation upsampling on low-scale features and aligns them with high-scale features via channels before stitching and fusing them. The multi-channel attention module then enhances the features in the changed areas, restoring the image to one-quarter the size of the original image. The second branch adjusts the multi-scale features to the same spatial size using bilinear interpolation before fusing them, and then gradually restores the image to its original size through continuous upsampling, ultimately generating high-precision change detection results.

[0070] The multi-scale progressive attention module achieves complementary enhancement of low-level and high-level features through the aforementioned multi-step operations. It also weights spatially important regions, ensuring accurate localization of changing areas even in complex backgrounds, under varying lighting conditions, or with vegetation interference. Furthermore, the module supports top-down pyramid fusion, effectively transferring high-level semantic information to low-level features while preserving texture details in the lower layers.

[0071] Optionally, channel stitching and fusion are performed on the enhanced change feature maps, including: performing a fourth upsampling operation on the fourth-scale enhanced change feature map to obtain an upsampled fourth-scale enhanced change feature map; performing a third upsampling operation on the third-scale enhanced change feature map to obtain an upsampled third-scale enhanced change feature map; performing a second upsampling operation on the second-scale enhanced change feature map to obtain an upsampled second-scale enhanced change feature map; concatenating the upsampled fourth-scale, third-scale, and second-scale enhanced change feature maps with the first-scale enhanced change feature map to obtain a concatenated feature; performing channel fusion on the concatenated feature using a preset convolutional block to obtain a fused feature; and performing a first upsampling operation on the fused feature to obtain a changed image.

[0072] The pre-defined convolutional block (Convblock) consists of a 1×1 convolution (Conv1×1), batch normalization (Bn) processing, and a ReLU activation function. The fourth upsampling operation consists of a bilinear interpolation upsampling operation (8x) and a 3x3 convolution (Bilinear8X&Conv3x3). The third upsampling operation consists of a bilinear interpolation upsampling operation (4x) and a 3x3 convolution (Bilinear 4X&Conv3x3). The second upsampling operation consists of a bilinear interpolation upsampling operation (2x) and a 3x3 convolution (Bilinear 2X&Conv3x3). The first upsampling operation consists of a bilinear interpolation upsampling operation (4x) and a sigmoid activation function (Bilinear 4X&sigmoid).

[0073] In some embodiments, combined with Figure 2As shown, an 8x upsampling operation is performed on the enhanced change feature map at the fourth scale to obtain an upsampled enhanced change feature map at the fourth scale. A 4x upsampling operation is performed on the enhanced change feature map at the third scale to obtain an upsampled enhanced change feature map at the third scale. A 2x upsampling operation is performed on the enhanced change feature map at the second scale to obtain an upsampled enhanced change feature map at the second scale. The upsampled enhanced change feature maps at the fourth, third, and second scales are then concatenated with the enhanced change feature map at the first scale to obtain a concatenated feature. The concatenated feature is then fused using a pre-defined convolutional block (Convblock) to obtain a fused feature. A 4x upsampling operation using bilinear interpolation and a sigmoid activation function (Bilinear 4X & sigmoid) are then performed on the fused feature to obtain a change probability map. This change probability map is then binarized to obtain a binary map of surface change, i.e., a change map. This change map is then output. The process of binarizing the probability map includes setting pixels with a probability greater than a set threshold to 1, which represents pixels that have changed, and setting pixels with a probability less than the set threshold to 0, which represents background that has not changed.

[0074] This remote sensing image change detection model, based on multi-scale spatiotemporal differential recognition, comprehensively utilizes the local feature extraction capabilities of residual convolutional structures and the global modeling advantages of Transformer structures to achieve differential enhancement and hierarchical fusion of change information in a multi-scale feature space. Compared with traditional single-path convolutional models, the remote sensing image change detection model provided in this disclosure can capture spatiotemporal change features more comprehensively, effectively improving the accuracy and robustness of remote sensing change detection, especially exhibiting superior change recognition performance under conditions of illumination changes, surface disturbances, or complex backgrounds.

[0075] In some embodiments, the training process of a remote sensing image change detection model based on a hybrid encoder and a multi-scale progressive fusion decoder is optimized using a joint loss function, thereby improving the model's accuracy in identifying areas of surface change and the stability of training convergence.

[0076] Optionally, the preset stopping conditions include: stopping the training of the change detection model when the joint loss function reaches a preset threshold, and determining the trained change detection model as the remote sensing image change detection model.

[0077] Optionally, the preset stopping condition includes: iterating the change detection model a preset number of times, and determining the change detection model corresponding to the minimum joint loss function value as the remote sensing image change detection model. The preset number of iterations is 200.

[0078] The joint loss function consists of binary cross-entropy loss. With Dice loss The loss function is constructed by linear addition, which balances pixel-level classification accuracy with the ability to identify small target change regions. During the training of the change detection model, the binary cross-entropy loss and the Dice loss are added to form a joint loss function, which is used for end-to-end optimization of the model parameters. Its calculation formula is as follows: in, For binary cross-entropy loss, This is a loss for Dice.

[0079] Binary cross-entropy loss measures the difference at each pixel between the predicted change image output by the change detection model during each training iteration and the actual change image from the bi-temporal sample remote sensing image. Both the predicted and actual change images are binarized images. The formula for calculating the binary cross-entropy loss is as follows: ;in, , Here are the pixel coordinates of the image, where h is the row, w is the column, H is the height of the image, and W is the width of the image. To predict the location in a changed image The predicted value at that location, Location in the actual change image The actual label value at the location, function This represents the cross-entropy loss function.

[0080] Through calculation get .

[0081] Dice loss measures the similarity between the predicted changed image and the actual changed image in terms of region overlap, thereby enhancing the model's ability to identify small targets or unevenly changed regions. Its calculation formula is as follows: ; The joint loss function consists of binary cross-entropy loss. and Dice loss The linear addition structure aims to synergistically address the class imbalance problem commonly found in change detection tasks and improve the convergence speed and stability of model training. Binary cross-entropy loss measures the difference between the predicted result and the true label pixel-by-pixel, ensuring the model's classification accuracy for each pixel. Meanwhile, Dice loss evaluates the consistency between the predicted result and the actual changed region from the perspective of region overlap, effectively enhancing the model's overall ability to identify changed regions.

[0082] By employing a joint loss function, the model not only maintains classification accuracy at the pixel level but also improves its ability to perceive changed regions at the region level, exhibiting particularly good recognition performance for small-scale changed regions in remote sensing images. During training, this joint loss function serves as the optimization objective, guiding the parameter updates of the hybrid encoder and multi-scale progressive fusion decoder through backpropagation. This achieves deep supervision and collaborative optimization of feature learning at each layer of the model, significantly improving the accuracy and robustness of the change detection model.

[0083] In practical applications, dual-temporal remote sensing images are input into a trained remote sensing image change detection model to obtain detection results for target areas. Specifically, the remote sensing image change detection model receives the input dual-temporal remote sensing images and generates a binarized change image. This facilitates the identification of changed areas in the remote sensing image through the change image, thereby providing effective information for land surface change monitoring, environmental management, and resource regulation. Through the output results after multi-scale fusion, the remote sensing image change detection model can generate a refined change detection map and, combined with data from the entire image range, achieve spatial distribution extraction and structural feature analysis of changed areas.

[0084] The remote sensing image change detection method provided in this disclosure first acquires multi-temporal, high-resolution remote sensing image data covering a preset area, and performs preprocessing such as atmospheric correction, radiometric calibration, color correction, image registration, and cropping to form a high-quality, temporally consistent remote sensing image dataset. Subsequently, a remote sensing image change detection model based on a hybrid encoder and a multi-scale progressive fusion decoder is constructed and trained using a deep learning algorithm. The hybrid encoder includes a convolutional residual module and a Transformer module based on a shift-window self-attention mechanism, while the multi-scale progressive fusion decoder includes a change feature extraction module and a multi-level progressive attention module. This structure can achieve differential enhancement and fusion of change features at different scales and semantic levels.

[0085] During the model inference phase, the model models the features of dual-temporal images from two perspectives: local texture features and global contextual information. It extracts significantly changed regions through modules, and uses an adaptive feature fusion module to aggregate features at different scales hierarchically. Then, a module performs pyramid-style fusion and spatial enhancement on change features at adjacent scales, thereby improving detection accuracy and model generalization ability. The final generated land change area detection results accurately reflect the extent and structural characteristics of the changed areas and effectively suppress spurious change interference caused by seasonal, illumination, or vegetation changes. This model has low parameter count and high computational efficiency, making it suitable for land change monitoring and intelligent identification of large-scale, multi-temporal remote sensing images, balancing detection accuracy and deployment practicality.

[0086] Figure 7This is a visualization comparison of the remote sensing image change detection method provided in this disclosure with other traditional change detection methods in a remote sensing scene. Combined with... Figure 7 As shown, there are five pairs of dual-temporal remote sensing images to be detected: a, b, c, d, and e. T1 Image is the first temporal image to be detected, acquired earlier, and T2 Image is the second temporal image to be detected, acquired later. GT represents the actual changed image label. FC-EF, FC-Siam-Conc, FC-Siam-Diff, STANet, BIT, ChangeFormer, SNUNet, SwinSUNet, CSI-Net, and EGPNet all represent the detection results of other traditional change detection methods. Ours represents the detection result of the remote sensing image change detection method provided in this embodiment. From Figure 7 As can be seen, traditional change detection methods often suffer from blurred boundaries, missed detections, or false detections of change areas when faced with complex surface environments. Red and green represent false detection areas. In contrast, the remote sensing image change detection method provided in this disclosure can more accurately capture the spatial and semantic features of change areas, significantly improving the clarity and integrity of change area boundaries.

[0087] Specifically, in scenarios with significant changes in building areas, other traditional change detection methods, due to their limited ability to model long-range dependencies, are prone to semantic mismatches between images from different time periods, resulting in numerous false detections. The remote sensing image change detection method provided in this disclosure introduces multi-scale feature fusion and a global self-attention mechanism, achieving efficient correlation modeling between features across time phases, effectively suppressing background interference, and maintaining stable detection performance even under complex textures and scale changes.

[0088] Furthermore, from a visual perspective, the remote sensing image change detection method provided in this disclosure is more accurate in extracting the contours of changed areas, can completely identify the true boundaries of changed targets, and maintains a smooth and consistent background area with almost no obvious noise artifacts. This indicates that the remote sensing image change detection method provided in this disclosure balances detail accuracy and global consistency in change detection tasks, and is superior to traditional change detection methods in terms of robustness and generalization ability.

[0089] Combination Figure 8As shown, this embodiment of the present disclosure provides an apparatus 800 for detecting changes in remote sensing images, including: an acquisition module 801 and a detection module 802. The acquisition module 801 is configured to acquire dual-temporal remote sensing images to be measured, wherein the dual-temporal remote sensing images to be measured are remote sensing images of a target area at two different times. The detection module 802 is configured to input the dual-temporal remote sensing image to be tested into a preset remote sensing image change detection model to obtain the detection result of the target area. The remote sensing image change detection model includes a hybrid encoder and a multi-scale progressive fusion decoder. The hybrid encoder includes multiple feature extraction modules, each composed of a preset residual module and a preset attention module. The hybrid encoder is used to extract multi-scale feature maps from the dual-temporal remote sensing image to be tested. The multi-scale progressive fusion decoder includes a change feature extraction module and a multi-level progressive attention module. The change feature extraction module is used to extract change feature maps of corresponding scales from the feature maps at each scale. The multi-level progressive attention module is used to enhance the change feature maps at each scale to obtain enhanced change feature maps of the corresponding scale, and then performs channel stitching and fusion on the enhanced change feature maps to obtain the detection result.

[0090] The apparatus for remote sensing image change detection provided in this disclosure detects changes by inputting dual-temporal remote sensing images into a preset remote sensing image change detection model. This preset model includes a hybrid encoder and a multi-scale progressive fusion decoder. The hybrid encoder's multiple feature extraction modules are composed of preset residual modules and preset Transformer modules, respectively. The multi-scale progressive fusion decoder includes a change feature extraction module and a multi-level progressive attention module. This effectively enhances the model's feature representation ability and training stability by utilizing the local feature extraction capabilities of CNNs (Convolutional Neural Networks). It leverages the global modeling advantages of Transformers to achieve efficient global information aggregation at multiple scales. Furthermore, the collaborative work of the change feature extraction module and the multi-level progressive attention module enables the full fusion and semantic enhancement of multi-level change features, thereby effectively improving the accuracy and robustness of remote sensing image change detection. Therefore, using the remote sensing image change detection model to detect dual-temporal remote sensing images can improve the accuracy of remote sensing image change detection.

[0091] Optionally, the device for detecting changes in remote sensing images further includes a training module; the training module is configured to acquire multiple pairs of dual-temporal sample remote sensing images, wherein the dual-temporal sample remote sensing images are remote sensing images of a preset area at two different times. Each pair of dual-temporal sample remote sensing images is sequentially input into a preset change detection model for iterative training until a preset stopping condition is met, and the trained change detection model is determined as the remote sensing image change detection model.

[0092] In this embodiment, high-resolution dual-temporal remote sensing image data registered in the same geographical area are acquired and preprocessed, including radiometric correction, atmospheric correction, image registration, and cropping, to obtain standardized sample remote sensing image pairs. This effectively eliminates interference caused by differences in imaging conditions and provides a high-quality data foundation for subsequent change detection. A change detection model based on a hybrid encoder and a multi-scale progressive fusion decoder is constructed using deep learning algorithms. End-to-end training is performed on standardized sample remote sensing image pairs to obtain a high-performance remote sensing image change detection model. The hybrid encoder, by alternately stacking improved residual blocks and Swing Transformer blocks, fully leverages the complementary advantages of CNNs in local feature extraction and Transformers in global context modeling. The multi-scale progressive fusion decoder, through the collaborative work of a change feature extraction module and a multi-level progressive attention module, achieves full fusion and semantic enhancement of multi-level change features.

[0093] The network architecture design of the remote sensing image change detection model provided in this disclosure effectively solves the technical challenges of traditional methods in global information modeling, multi-scale feature utilization, and pseudo-change suppression. The hybrid encoder can simultaneously capture local detail features and global semantic information, providing rich feature representations for change detection. The change feature extraction module achieves refined extraction of differential features through a dual-branch structure, significantly improving the recognition accuracy of truly changed areas. The multi-level progressive attention module effectively maintains the integrity of change boundaries through a pyramid-style fusion strategy. The organic combination of these technologies enables the remote sensing image change detection method provided in this disclosure to maintain excellent detection performance even in complex scenarios, significantly improving the accuracy and robustness of change detection.

[0094] Furthermore, the remote sensing image change detection method provided in this disclosure ensures detection accuracy while also considering the model's practicality and deployment efficiency. Through a shifted window attention mechanism and optimized residual structure design, computational complexity is reduced while maintaining the model's expressive power; a progressive fusion strategy of multi-scale features achieves reasonable allocation of computational resources. This allows it to adapt to different application scenarios, providing a reliable technical solution for practical engineering applications of remote sensing change detection. Experimental results show that the remote sensing image change detection method provided in this disclosure achieves leading performance on multiple public datasets, demonstrating significant advantages in detection accuracy, robustness, and computational efficiency.

[0095] Combination Figure 9As shown, this disclosure provides an electronic device including a processor 900 and a memory 901 storing program instructions. Optionally, the electronic device may further include a communication interface 902 and a bus 903. The processor 900, communication interface 902, and memory 901 can communicate with each other via the bus 903. The communication interface 902 can be used for information transmission. The processor 904 can call the program instructions in the memory 901 to execute the method for remote sensing image change detection described in the above embodiment.

[0096] Furthermore, the logic instructions in the aforementioned memory 901 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0097] The memory 901, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 904 executes functional applications and data processing by running the program instructions / modules stored in the memory 901, thereby implementing the method for remote sensing image change detection in the above embodiments.

[0098] The memory 901 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 901 may include high-speed random access memory and may also include non-volatile memory.

[0099] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and other media capable of storing program code; it can also be a transient storage medium.

[0100] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0101] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0102] The methods and products disclosed in the embodiments herein (including but not limited to devices and equipment) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A method for detecting changes in remote sensing images, characterized in that, include: Acquire dual-temporal remote sensing images of the target area at two different times; The dual-temporal remote sensing images to be tested are input into a preset remote sensing image change detection model to obtain the detection result of the target area. The remote sensing image change detection model includes a hybrid encoder and a multi-scale progressive fusion decoder. The hybrid encoder includes multiple feature extraction modules, each composed of a preset residual module and a preset attention module. The hybrid encoder is used to extract multi-scale feature maps from the dual-temporal remote sensing images to be tested. The multi-scale progressive fusion decoder includes a change feature extraction module and a multi-level progressive attention module. The change feature extraction module is used to extract change feature maps of corresponding scales from the feature maps at each scale. The multi-level progressive attention module is used to enhance the change feature maps at each scale to obtain enhanced change feature maps of the corresponding scale, and then performs channel stitching and fusion on the enhanced change feature maps to obtain the detection result.

2. The method according to claim 1, characterized in that, The preset remote sensing image change detection model is obtained through the following methods: Acquire multiple pairs of dual-temporal sample remote sensing images, wherein the dual-temporal sample remote sensing images are remote sensing images of a preset area at two different times; Each pair of dual-temporal sample remote sensing images is sequentially input into a preset change detection model for iterative training until a preset stopping condition is met. The trained change detection model is then determined as the remote sensing image change detection model.

3. The method according to claim 2, characterized in that, The change detection model includes a hybrid encoder and a multi-scale progressive fusion decoder, wherein the multi-scale progressive fusion decoder includes a change feature extraction module and a multi-level progressive attention module; the following operations are performed during each training iteration of the change detection model: The dual-temporal sample remote sensing image is input into the hybrid encoder, and the multi-scale feature map of the dual-temporal sample remote sensing image is extracted by the hybrid encoder; The change feature extraction module extracts the corresponding scale change feature maps from the feature maps at each scale. The multi-level progressive attention module enhances the feature maps at each scale to obtain enhanced feature maps at the corresponding scale. Then, the enhanced feature maps are spliced ​​and fused to output the change image.

4. The method according to claim 3, characterized in that, The hybrid encoder includes: a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, a fifth feature extraction module, a sixth feature extraction module, a seventh feature extraction module, and an eighth feature extraction module; the hybrid encoder extracts multi-scale feature maps of the dual-temporal sample remote sensing images, including: The first time-phase sample remote sensing image is input into the first feature extraction module to obtain the first scale feature map of the first time-phase sample remote sensing image; the second time-phase sample remote sensing image is input into the second feature extraction module to obtain the first scale feature map of the second time-phase sample remote sensing image. The first-scale feature map of the first-temporal sample remote sensing image is input into the third feature extraction module to obtain the second-scale feature map of the first-temporal sample remote sensing image; the first-scale feature map of the second-temporal sample remote sensing image is input into the fourth feature extraction module to obtain the second-scale feature map of the second-temporal sample remote sensing image. The second-scale feature map of the first time-phase sample remote sensing image is input into the fifth feature extraction module to obtain the third-scale feature map of the first time-phase sample remote sensing image; the second-scale feature map of the second time-phase sample remote sensing image is input into the sixth feature extraction module to obtain the third-scale feature map of the second time-phase sample remote sensing image. The third-scale feature map of the first time-phase sample remote sensing image is input into the seventh feature extraction module to obtain the fourth-scale feature map of the first time-phase sample remote sensing image; the third-scale feature map of the second time-phase sample remote sensing image is input into the eighth feature extraction module to obtain the fourth-scale feature map of the second time-phase sample remote sensing image.

5. The method according to claim 4, characterized in that, The change feature extraction module includes: a first change feature extraction submodule, a second change feature extraction submodule, a third change feature extraction submodule, and a fourth change feature extraction submodule; the change feature extraction module extracts change feature maps of corresponding scales from feature maps of each scale, including: The first-scale dual-temporal feature map is input into the first change feature extraction submodule to obtain the first-scale change feature map; The second-scale dual-temporal feature map is input into the second change feature extraction submodule to obtain the second-scale change feature map; The dual-temporal feature map at the third scale is input into the third change feature extraction submodule to obtain the change feature map at the third scale; The bi-temporal feature map at the fourth scale is input into the fourth change feature extraction submodule to obtain the change feature map at the fourth scale.

6. The method according to claim 5, characterized in that, The multi-level progressive attention module includes: a first attention sub-module, a second attention sub-module, a third attention sub-module, and a fourth attention sub-module; the multi-level progressive attention module enhances the feature maps at each scale to obtain enhanced feature maps at the corresponding scale, including: The fourth-scale change feature map is input into the fourth attention submodule to obtain the enhanced fourth-scale change feature map; The enhanced change feature map at the fourth scale and the change feature map at the third scale are input into the third attention submodule to obtain the enhanced change feature map at the third scale. The enhanced change feature map at the third scale and the change feature map at the second scale are input into the second attention submodule to obtain the enhanced change feature map at the second scale. The enhanced change feature map at the second scale and the change feature map at the first scale are input into the first attention submodule to obtain the enhanced change feature map at the first scale.

7. The method according to claim 6, characterized in that, Channel stitching and fusion are performed on each enhanced feature map, including: A fourth upsampling operation is performed on the enhanced change feature map at the fourth scale to obtain an upsampled enhanced change feature map at the fourth scale; a third upsampling operation is performed on the enhanced change feature map at the third scale to obtain an upsampled enhanced change feature map at the third scale; and a second upsampling operation is performed on the enhanced change feature map at the second scale to obtain an upsampled enhanced change feature map at the second scale. The enhanced change feature maps at the fourth scale, the third scale, the second scale, and the first scale after upsampling are concatenated to obtain the concatenated features. The spliced ​​features are fused using a preset convolutional block to obtain fused features; A first upsampling operation is performed on the fused features to obtain the changed image.

8. An apparatus for detecting changes in remote sensing images, characterized in that, include: The acquisition module is configured to acquire dual-temporal remote sensing images to be measured, wherein the dual-temporal remote sensing images to be measured are remote sensing images of the target area at two different times. The detection module is configured to input the dual-temporal remote sensing image to be tested into a preset remote sensing image change detection model to obtain the detection result of the target area. The remote sensing image change detection model includes a hybrid encoder and a multi-scale progressive fusion decoder. The hybrid encoder includes multiple feature extraction modules, each consisting of a preset residual module and a preset attention module. The hybrid encoder is used to extract multi-scale feature maps from the dual-temporal remote sensing image to be tested. The multi-scale progressive fusion decoder includes a change feature extraction module and a multi-level progressive attention module. The change feature extraction module is used to extract change feature maps of corresponding scales from the feature maps at each scale. The multi-level progressive attention module is used to enhance the change feature maps at each scale to obtain enhanced change feature maps of the corresponding scale, and then performs channel stitching and fusion on the enhanced change feature maps to obtain the detection result.

9. The apparatus according to claim 8, characterized in that, Also includes: Training module; The training module is configured to acquire multiple pairs of dual-temporal sample remote sensing images, wherein the dual-temporal sample remote sensing images are remote sensing images of a preset area at two different times. Each pair of dual-temporal sample remote sensing images is sequentially input into a preset change detection model for iterative training until a preset stopping condition is met. The trained change detection model is then determined as the remote sensing image change detection model.

10. An electronic device comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute, when running the program instructions, the method for detecting changes in remote sensing images as described in any one of claims 1 to 7.