Dual-time-phase image change detection method, system, equipment and medium

The dual-temporal image change detection method, which integrates multi-scale feature fusion and attention mechanisms, solves the accuracy and efficiency problems of remote sensing image change detection in complex scenarios. It achieves efficient and accurate multi-temporal remote sensing image change region detection and is suitable for various application scenarios.

CN120877133APending Publication Date: 2025-10-31GANTRY LAB +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510953079.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing remote sensing image change detection technologies have significant shortcomings in terms of accuracy, efficiency, and adaptability. They are particularly difficult to achieve high-precision real-time detection in complex scenarios, have limited ability to detect small-sized changing targets, consume large amounts of computing resources, and are difficult to deploy on edge devices.

Method used

A dual-temporal image change detection method based on multi-scale feature fusion and attention mechanism is adopted. Through the dual-branch feature extraction network of ResNet architecture, combined with cross-temporal feature fusion module and multi-level feature decoding module, multi-scale feature extraction and fusion are realized. The self-attention mechanism is used to enhance feature representation ability and reduce computational complexity.

Benefits of technology

It improves detection accuracy, enhances adaptability to complex scenarios, significantly reduces computational complexity, supports rapid change detection in large-scale remote sensing images, and is applicable to fields such as natural disaster monitoring, urban development assessment, agricultural monitoring, environmental protection, and land use analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877133A_ABST
    Figure CN120877133A_ABST
Patent Text Reader

Abstract

The invention provides a dual-temporal image change detection method, system and device based on multi-scale feature fusion and an attention mechanism, and a medium, and relates to the field of computer vision and remote sensing image processing, and the method comprises the steps: taking a dual-temporal remote sensing image semantic change detection data set as original remote sensing image data; preprocessing the original remote sensing image data; extracting multi-scale features of the processed remote sensing image through a double-branch feature extraction network, wherein the double-branch feature extraction network adopts a ResNet architecture as a backbone network; a cross-time-phase feature fusion module is adopted to fuse multi-scale features of different time phases, a multi-stage feature decoding module is adopted to perform up-sampling and scale fusion on the fused features, and a final remote sensing image change detection probability graph is output after decoding operation. According to the invention, by introducing technical means such as deep learning, multi-scale feature extraction and a self-attention mechanism, efficient detection of a multi-time remote sensing image change area is realized, and the detection precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and remote sensing image processing technology, specifically to a method, system, device, and medium for detecting changes in dual-temporal images based on multi-scale feature fusion and attention mechanisms. Background Technology

[0002] The development of current remote sensing image change detection technology faces numerous technical bottlenecks and practical application challenges. From a technological evolution perspective, this field has undergone a transformation from traditional pixel-level methods to object-level analysis, and then to deep learning methods. However, existing solutions still have significant shortcomings in terms of accuracy, efficiency, and adaptability. At the traditional method level, pixel-level change detection techniques (such as image difference analysis, change vector analysis, and principal component analysis) mainly rely on spectral feature differences for change identification. While these methods are simple in algorithm and computationally efficient, their performance is constrained by many factors: First, the methods require extremely high image registration accuracy; even a registration error of 0.5 pixels can lead to a false alarm rate of over 20%. Second, relying solely on spectral features makes it difficult to distinguish between real surface changes and interference factors such as lighting conditions and seasonal variations; in crop rotation areas, the false detection rate can reach 40-50%. Third, these methods completely ignore the spatial context information of the image, resulting in very limited detection capabilities for structural changes such as partial changes in buildings and road expansions.

[0003] Object-oriented change detection methods (such as multi-scale segmentation methods based on the eCognition platform) improve the efficiency of spatial information utilization to some extent by introducing image segmentation technology. However, these methods have three key drawbacks: First, the selection of segmentation scale requires a lot of manual intervention, and different segmentation parameters need to be set for different land cover types, resulting in low automation. Second, the segmentation process loses detailed information, and the detection capability for small-scale changes (such as newly added temporary buildings, vehicles, etc.) is insufficient, with a false negative rate of over 35% for targets smaller than 10×10 pixels. Third, the algorithm complexity is high, and processing remote sensing images of a 1 square kilometer area usually requires more than 30 minutes of computation time, making it difficult to meet the needs of real-time monitoring.

[0004] In recent years, deep learning methods have shown great potential in remote sensing change detection, but several technical challenges remain to be addressed. Regarding feature extraction, existing dual-branch convolutional networks (such as FC-EF and DSIFN) typically employ simple feature concatenation or arithmetic operations to achieve temporal fusion. This approach struggles to effectively eliminate radiometric differences caused by factors such as lighting conditions and shooting angles. Experimental data shows that in farmland areas with significant seasonal changes, the false detection rate of these methods remains as high as 30-40%. In terms of feature fusion, while the introduction of attention mechanisms (such as the spatial attention module in STANet) has improved the targeting of feature selection, it has also led to a significant increase in computational complexity, with typical 256×256 pixel image processing speeds generally below 10 FPS. More importantly, existing attention mechanisms have limited ability to suppress interference factors such as cloud cover and shadows, potentially leading to a 15-20% increase in false detection rates in certain scenarios. Regarding multi-scale processing, current mainstream architectures (such as UNet++ and HRNet) primarily employ fixed dilated convolution strategies or simple feature pyramid structures, making it difficult to adapt to the detection needs of targets with varying sizes. In addition, existing methods generally consume a lot of computing resources, with the number of parameters typically exceeding 50M, posing a serious challenge to deployment on edge devices.

[0005] Industry assessment data (IEEE GRSS2023) shows that, while maintaining real-time processing speed (>15 FPS), the theoretical upper limit of detection accuracy for existing technologies is approximately 82% (IoU), which is significantly lower than the practical application requirements (>90%). From a practical engineering application perspective, the bottlenecks of existing technologies are mainly reflected in three aspects: First, insufficient adaptability to complex scenarios, especially poor detection stability under conditions such as cloud cover, seasonal changes, and shadow interference; second, limited detection capability for small-sized changing targets, significantly reducing its practical value in scenarios such as smart city management and illegal building inspection; and third, a prominent contradiction between algorithm efficiency and accuracy, making it difficult to achieve high-precision real-time detection on resource-constrained edge devices. These technical deficiencies severely restrict the large-scale application of remote sensing change detection technology in key areas such as land monitoring, disaster assessment, and smart cities, and urgently require breakthroughs through technological innovation. Summary of the Invention

[0006] In view of this, embodiments of this application provide a method, system, device and medium for detecting changes in dual-temporal images. By introducing deep learning, multi-scale feature extraction and self-attention mechanism and other technical means, it achieves efficient detection of change areas in multi-temporal remote sensing images, and can be widely used in multiple fields such as disaster monitoring, urban development assessment and land use change analysis.

[0007] This application provides the following technical solution: a dual-temporal image change detection method based on multi-scale feature fusion and attention mechanism, comprising:

[0008] The semantic change detection dataset of dual-temporal remote sensing images is used as the original remote sensing image data, and the original remote sensing image data is preprocessed.

[0009] Multi-scale features of the preprocessed remote sensing images are extracted by a dual-branch feature extraction network, wherein the dual-branch feature extraction network adopts the ResNet architecture as the backbone network.

[0010] A cross-temporal feature fusion module is used to perform cross-temporal fusion on the multi-scale features respectively; a multi-level feature decoding module is used to upsample and scale-fuse the fused features, and after decoding, the final remote sensing image change detection probability map is output.

[0011] According to one embodiment of this application, preprocessing the raw remote sensing image data includes:

[0012] The remote sensing image data and labels in the random part of the training sample are flipped horizontally, randomly cropped and rotated. At the same time, the remote sensing images are color enhanced and salt and pepper noise is added to the labels.

[0013] The remote sensing image data, excluding the random portion mentioned above, is stitched together with three random batches of remote sensing image data and labels. Then, the new remote sensing image data and corresponding labels are flipped left and right and rotated randomly. At the same time, the remote sensing images are color-enhanced and salt-and-pepper noise is added to the labels.

[0014] According to one embodiment of this application, the dual-branch feature extraction network adopts the ResNet 18 architecture as the backbone network; the backbone network structure of the ResNet 18 architecture includes multiple residual modules, each of which contains two or three convolutional layers, and the input and output are added layer by layer through skip connections.

[0015] According to one embodiment of this application, the fully connected layers are removed from the backbone network structure of the ResNet 18 architecture, and four feature output layers are retained. Each feature output layer achieves feature extraction of different receptive fields by adjusting the convolution kernel size and downsampling stride.

[0016] According to one embodiment of this application, a cross-temporal feature fusion module is used to perform cross-temporal fusion on the multi-scale features, including:

[0017] Feature representations under different receptive fields are extracted simultaneously by multi-scale convolutional units, and the input multi-scale features are concatenated along the channel dimension; the attention weight of each channel is measured by a channel attention mechanism to obtain the channel attention weight coefficient, and the channel attention weight coefficient is multiplied by the concatenated features to obtain the enhanced features;

[0018] Local features are extracted from the input features through convolution operations. These local features are then fused with the enhanced features, and batch normalization and ReLU activation operations are performed to obtain the fused features.

[0019] Calculate the channel mean and variance of the fused features, generate adaptive enhancement weights based on the channel mean and variance, and multiply the adaptive enhancement weights with the input fused features to generate the final fused features.

[0020] According to one embodiment of this application, a multi-level feature decoding module upsamples and scales the fused features, and outputs a final remote sensing image change detection probability map after decoding, including:

[0021] The high-level feature map is upsampled using bilinear interpolation to generate a feature map with double resolution. The double-resolution feature map is then concatenated with the lower-level feature map along the channel dimension to form an intermediate feature map.

[0022] The intermediate feature map is input into the decoding unit of the dual-branch architecture for decoding. The dual-branch architecture includes a local branch and a global branch. The decoding unit performs channel rearrangement on the input intermediate feature map, and inputs the rearranged data into the local branch and the global branch respectively to extract local features and global features. The local features and the global features are then fused to output the decoded feature map.

[0023] According to one embodiment of this application, the local branch extracts local features of the feature map through a convolutional layer-normalization-activation function layer, the global branch extracts global features using a channel attention and spatial attention concatenation mechanism, and fuses global and local features through weighted summation. Finally, the final remote sensing image change detection probability map is generated through convolutional layers and interpolation operations.

[0024] This application also provides a dual-temporal image change detection system based on multi-scale feature fusion and attention mechanism, including:

[0025] The preprocessing module is used to preprocess the original remote sensing image data by using the semantic change detection dataset of dual-temporal remote sensing images as the original remote sensing image data.

[0026] The feature extraction module is used to extract multi-scale features of the preprocessed remote sensing image through a dual-branch feature extraction network, wherein the dual-branch feature extraction network adopts the ResNet architecture as the backbone network.

[0027] The feature fusion and decoding module is used to perform cross-temporal fusion of the multi-scale features using the cross-temporal feature fusion module; and to perform upsampling and scale fusion of the fused features through the multi-level feature decoding module, and output the final remote sensing image change detection probability map after decoding.

[0028] This application also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described dual-temporal image change detection method.

[0029] This application also provides a computer-readable storage medium storing a computer program that performs the above-described dual-temporal image change detection method.

[0030] Compared with existing technologies, the beneficial effects achieved by at least one of the above-mentioned technical solutions adopted in the embodiments of this specification include at least the following: First, the multi-scale feature extraction module can simultaneously capture local detail changes and global structural changes in images, thereby improving detection accuracy. Second, the dynamic feature enhancement module significantly improves the model's ability to represent features in changing regions through a self-attention mechanism, and can suppress interference from background information in complex scenes. Finally, the modular decoding architecture and efficient implementation significantly reduce computational complexity, supporting rapid change detection of large-scale remote sensing images.

[0031] The method of this invention is applicable to various dual-temporal image change detection scenarios, including natural disaster monitoring (such as surface changes caused by floods, earthquakes, and landslides), urban development assessment (such as building construction and demolition, road expansion, etc.), agricultural monitoring (such as changes in farmland area and crop growth status assessment), environmental protection (such as deforestation, wetland changes, and lake area changes), and land use analysis (such as changes in land use classification and land erosion detection). By flexibly adjusting the model architecture and parameters, this method can adapt to the change detection needs of remote sensing images with various resolutions, multispectral dimensions, and time spans. In summary, this invention, by designing a deep learning framework based on multi-scale feature enhancement and attention guidance, solves the problems of insufficient accuracy and low model efficiency in change detection under complex scenarios by traditional methods, achieving efficient and accurate detection of change areas in multi-temporal remote sensing images, and has broad application value and potential for technology promotion. Attached Figure Description

[0032] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is a schematic flowchart of the dual-temporal image change detection method according to an embodiment of the present invention;

[0034] Figure 2 This is a block diagram of a dual-temporal image change detection system according to an embodiment of the present invention;

[0035] Figure 3 This is a schematic diagram of the structure of the computer device of the present invention. Detailed Implementation

[0036] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0037] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0038] like Figure 1 As shown, this embodiment of the invention provides a dual-temporal image change detection method based on multi-scale feature fusion and attention mechanism, including:

[0039] S101. Use the semantic change detection dataset of dual-temporal remote sensing images as the original remote sensing image data, and preprocess the original remote sensing image data;

[0040] S102. Multi-scale features of the preprocessed remote sensing images are extracted using a dual-branch feature extraction network, wherein the dual-branch feature extraction network uses a ResNet architecture as the backbone network;

[0041] S103. The multi-scale features are fused across time using a cross-temporal feature fusion module; the fused features are upsampled and scale-fused using a multi-level feature decoding module, and the final remote sensing image change detection probability map is output after decoding.

[0042] In one embodiment of the present invention, the original remote sensing image data is preprocessed: data preprocessing is a crucial step in the dual-temporal image change detection (remote sensing image change detection) task, aiming to provide standardized input data for subsequent feature extraction and fusion modules. The preprocessing process in this embodiment primarily employs two different data augmentation methods, the specific implementation methods of which are as follows:

[0043] A1. Perform left and right flipping, random cropping, and rotation operations on a random portion of remote sensing image data and labels in the training samples. At the same time, enhance the brightness, contrast, and other colors of the remote sensing images, and add salt and pepper noise to the labels.

[0044] A2. The remaining remote sensing image data from the training samples are stitched together with three random batches of data and labels. Then, the new remote sensing image data and its labels are flipped horizontally and rotated randomly. Simultaneously, color enhancement is applied to the remote sensing images, and salt-and-pepper noise is added to the labels. Different data augmentation methods are used to generate diverse samples to improve the model's robustness.

[0045] In one embodiment of the present invention, the dual-branch feature extraction network uses a ResNet 18 architecture as its backbone network. The ResNet 18 backbone network structure includes multiple residual modules, each containing two or three convolutional layers, and uses skip connections to achieve layer-by-layer addition of input and output. Specifically, the ResNet 18 backbone network structure removes fully connected layers and retains four feature output layers. Each feature output layer extracts features from different receptive fields by adjusting the convolutional kernel size and downsampling stride.

[0046] In practical implementation, the feature extraction network is used to extract multi-level, multi-scale depth features from the two preprocessed remote sensing images, providing a foundation for subsequent feature fusion and change detection. This embodiment of the invention employs a dual-branch feature extraction network with ResNet as the backbone during the feature extraction process, combined with an input channel adjustment mechanism and multi-scale feature hierarchy output, as specifically implemented below:

[0047] Deep Dual-Branch Feature Extraction Network: This embodiment of the invention uses ResNet as the backbone network for feature extraction. Based on task complexity and hardware resource limitations, ResNet18 is selected as the default architecture. The residual modules of ResNet can effectively alleviate the vanishing gradient problem in deep networks and improve feature representation capabilities. The specific structure of the backbone network is as follows:

[0048] Input layer: First, the input image is subjected to preliminary convolution and downsampling operations. Low-level features are extracted through a 7×7 convolution kernel, and batch normalization and ReLU activation function are used to enhance non-linear expressive power.

[0049] x1 = ReLU(BN(Conv2d(x)))

[0050] The spatial resolution of the feature map is then further compressed using a 3×3 max-pooling layer.

[0051] Residual Modules: The residual structure is implemented using residual blocks (such as BasicBlock or Bottleneck) in the ResNet backbone, while custom modules (such as module1) do not use residual design. Each residual block contains the main path (2-3 convolutional layers + BN + ReLU) and skip connections, which are ultimately fused element-wise, formally as follows:

[0052] x l+1 =ReLU(x l +F(x l ))

[0053] Where F represents a combination of convolution, batch normalization, and ReLU.

[0054] Multi-scale feature output: The backbone network contains four feature layers (layer1, layer2, layer3, layer4). Each layer extracts features from different receptive fields by adjusting the convolutional kernel size and downsampling stride, corresponding to different spatial scales.

[0055] layer1: Output channels are 64, mainly capturing low-level texture features;

[0056] layer2: Output channels are 128, extracting intermediate edge and local structural features;

[0057] layer3: Output channels are 256, describing higher-order geometry;

[0058] layer4: Output channels are 512, generating global semantic features.

[0059] Input Channel Adjustment: To accommodate different numbers of channels in the input image (e.g., 3-channel RGB image or multi-channel remote sensing image), this embodiment of the invention designs an adjustable input layer. If the number of channels in the input image is not 3, the first convolutional kernel of the ResNet is replaced.

[0060] Conv1=Conv2d(in_ch,64,kernel_size=7,stride=2,padding=3)

[0061] Feature hierarchy separation: This invention separates and outputs features at different levels during the feature extraction stage to facilitate multi-scale processing in subsequent fusion and decoding modules. Assuming input images A and B, their multi-scale features are extracted separately through the backbone network:

[0062]

[0063] in: Low-level texture features; Mid-level edge features; High-level geometric features; Global semantic features. The outputs of these feature levels provide a sufficient information foundation for subsequent feature fusion modules.

[0064] Modular Implementation: To improve the flexibility and scalability of the model, in one embodiment of the present invention, the feature extraction network is encapsulated as a modular component, which facilitates the replacement of the backbone network or the adjustment of the feature hierarchy configuration. For example, by specifying parameters, a deeper ResNet50 or ResNet101 can be selected as the backbone network to adapt to remote sensing image change detection tasks of varying complexity.

[0065] In one embodiment of the present invention, a cross-temporal feature fusion module is used to fuse the multi-scale features, including:

[0066] Feature representations under different receptive fields are extracted simultaneously by multi-scale convolutional units, and the input multi-scale features are concatenated along the channel dimension; the attention weight of each channel is measured by a channel attention mechanism to obtain the channel attention weight coefficient, and the channel attention weight coefficient is multiplied by the concatenated features to obtain the enhanced features;

[0067] Local features are extracted from the input features through convolution operations. These local features are then fused with the enhanced features, and batch normalization and ReLU activation operations are performed to obtain the fused features.

[0068] Calculate the channel mean and variance of the fused features, generate adaptive enhancement weights based on the channel mean and variance, and multiply the adaptive enhancement weights with the input fused features to generate the final fused features.

[0069] In this embodiment, the cross-temporal feature fusion module aims to effectively fuse features extracted from two remote sensing images through interactive processing of multi-temporal features, providing a more accurate feature representation for subsequent decoding and change detection. This embodiment proposes a multi-module collaborative fusion strategy, including a multi-scale convolution module, a local feature extraction unit, and a feature enhancement unit, specifically implemented as follows:

[0070] Multi-scale Convolutional Unit (PASR): This embodiment incorporates a multi-scale convolutional unit (PASR) during feature fusion to simultaneously extract feature representations from different receptive fields, thereby enhancing global variation information. For the input features x and y, they are first concatenated along the channel dimension: f cat =Concat(x,y), where x and y are features extracted from the two images, f cat ∈R C×H×W Subsequently, features are extracted through multi-scale convolution operations: f ms =MConv(f cat The multi-scale convolution module contains four convolutional kernels of different sizes (3×3, 5×5, 7×7, 9×9). Each branch extracts features within a specific receptive region, and finally, these features are concatenated and fused, as shown below:

[0071] f ms =Concat(Conv 3×3 (f cat ),Conv 5×5 (f cat ),Conv 7×7 (f cat ),Conv 9×9 (f cat ))

[0072] To more effectively distinguish change information in images from different time phases, this embodiment introduces a channel attention mechanism in the multi-scale convolutional unit. The channel attention mechanism enhances salient features by measuring the contribution weight of each channel, while suppressing irrelevant or redundant features. The specific steps are as follows:

[0073] Compression: Compression of the spliced ​​feature f cat Perform global average pooling to compress the feature map of each channel into a single value, and calculate the global information for each channel: q = AvgPool(f cat ), q∈R C×1×1 Activation: A convolution-ReLU-convolution layer is then used to extract inter-channel features. Simultaneously, the ReLU activation function is used to introduce non-linearity, enhancing the model's expressive power, ultimately generating a weight vector. Reweighting: The generated weight vector is normalized to the [0,1] range using the Sigmoid function, making it usable as weight coefficients w. se Feature enhancement: The weighting coefficients w se Multiply by the spliced ​​and merged features: f Mconv =f ms ·w se .

[0074] Finally, the enhanced features are added to the original image to obtain a feature map with global feature enhancement.

[0075] Local Feature Extraction Unit: This embodiment also uses a local feature extraction unit during the feature fusion process: convolutional layer-batch normalization-activation function to extract local features from the input image, expressed by the formula: f CBR =ReLU(BN(Conv) 1×1 (f cat Local features are extracted from the input feature map through convolution operations; batch normalization is used to normalize the extracted features to reduce internal covariate bias; and the ReLU activation function is used to introduce nonlinearity to enhance the model's expressive power.

[0076] Feature fusion is performed on the above feature maps, followed by batch normalization and ReLU activation operations to further enhance feature representation capabilities. A feature enhancement module further strengthens the feature representation capabilities of salient regions. The feature enhancement module is implemented as follows:

[0077] Feature Enhancement Unit (SSFC): To further improve the discriminative ability of fused features, this embodiment proposes a Feature Attention Enhancement Unit (SSFC). This module dynamically adjusts feature weights by analyzing the statistical characteristics of each channel, thereby enhancing the feature representation ability of salient regions. Specifically, it is implemented as follows:

[0078] Mean calculation: Calculate the fusion feature f att Channel mean:

[0079] Variance calculation: Calculate the variance for each channel:

[0080] Generate weights: Generate adaptive enhancement weights based on feature differences and channel mean. ∈ is a very small value to prevent division by zero errors or numerical instability in numerical calculations.

[0081] Feature enhancement: The enhancement weights are multiplied by the input features to generate the final fused features: f enhanced =f att ·w c ;

[0082] Multi-layer feature fusion: The feature fusion module operates independently at each scale feature layer (layer1, layer2, layer3, layer4), fusing features from two images layer by layer. For example, for layer l:

[0083] f fused,l =Conv l (f enhanced,l )

[0084] The fused features are then passed as input to the decoding module to generate the final change detection results.

[0085] In summary, the feature fusion module of this embodiment achieves efficient information interaction and feature enhancement at the spatial and semantic levels through multi-scale convolution, local feature extraction units, and feature enhancement units (SSFC), significantly improving the accuracy and robustness of change detection.

[0086] In one embodiment of the present invention, a multi-level feature decoding module upsamples and fuses the fused features, and after decoding, outputs the final remote sensing image change detection probability map, including:

[0087] The high-level feature map is upsampled using bilinear interpolation to generate a feature map with double resolution. The double-resolution feature map is then concatenated with the lower-level feature map along the channel dimension to form an intermediate feature map.

[0088] The intermediate feature map is input into the decoding unit of the dual-branch architecture for decoding. The dual-branch architecture includes a local branch and a global branch. The decoding unit performs channel rearrangement on the input intermediate feature map, and inputs the rearranged data into the local branch and the global branch respectively to extract local features and global features. The local features and the global features are then fused to output the decoded feature map.

[0089] This embodiment designs an efficient multi-level decoding structure to fully utilize the multi-level features generated during feature extraction and feature fusion, gradually restoring spatial resolution and extracting information about changing regions. The decoding module consists of multi-level sub-modules, each responsible for upsampling and fusing features at a specific level, thereby preserving a balance between high-level semantic information and low-level spatial detail information while restoring spatial resolution. The decoding module includes an upsampling and splicing unit and a decoding unit, specifically implemented as follows:

[0090] Upsampling and stitching unit. This embodiment designs an upsampling and stitching unit to fuse high-level features with low-level features and use the result as input to the decoding unit. The upsampling and stitching unit upsamples the high-level feature map using bilinear interpolation to generate a feature map with twice the resolution (consistent with the resolution of the low-level feature map), and then stitches it with the low-level feature map along the channel dimension to form an intermediate feature.

[0091] Decoding Unit. This embodiment designs a dual-branch architecture decoding unit to decode the feature map. The decoding unit takes the intermediate features obtained after fusing the features at different levels as input, performs channel rearrangement on the input image to improve the model's expressive power, and inputs the rearranged data into the local branch and global branch respectively to extract local and global features, performs feature fusion, and outputs the decoded features. The specific implementation of the decoding unit is as follows:

[0092] Local Branch: The local branch mainly analyzes the local features of the feature map. First, a convolutional layer-normalization-activation function layer is used to extract local features; then, the features are fed into convolutional layers with kernels of (1,3), (3,1), and 3×3 respectively to extract the horizontal, vertical, and local features of the feature map, calculate the mean of the three features, and add batch normalization and ReLU activation operations to further improve the feature representation capability.

[0093] Global Branch: The global branch primarily analyzes the global features of the feature map. First, it extracts local features through a convolutional layer-normalization-activation function layer. Then, it uses a self-attention mechanism to extract global features from the extracted feature map and performs a weighted sum with the original image to further enhance the global features of the image.

[0094] Feature fusion: Feature fusion is achieved by adding the feature maps output by the local branch and the global branch point by point, and the nonlinear expressive power is enhanced by the ReLU activation function.

[0095] The decoding process in this embodiment of the invention is specifically implemented as follows:

[0096] The feature map decoding process begins with the highest-level features (i.e., the fused layer 4 features). Let's assume the fused layer 4 features are represented as follows: Where C4, H4, and W4 represent the number of channels, spatial height, and width of the feature, respectively. First, F4 is fused with the resulting layer 3 feature. The input is fed into the upsampling and stitching unit. Specifically, F4 generates a feature representation with twice the resolution. The formula is expressed as: Subsequently, the upsampled features Concatenate F3 with F3 along the channel dimension to form an intermediate feature representation: To reduce the dimensionality of the concatenated features and further extract feature relationships, a decoding unit is used. This process ensures efficient feature fusion while avoiding an explosive increase in the number of channels. Similar operations are repeated in subsequent feature processing. The processed features are then processed... Upsampling is performed and combined with the fused layer 2 features. To perform the splicing, the formula is: The concatenated features are then processed by the Decoder2 unit to generate features. Next, for Upsampling is performed, and the fused layer 1 features are then combined. The data is concatenated and then processed by the decoding unit Decoder1 to generate the final decoded features. :

[0097]

[0098] Finally, a lightweight convolutional network is used to... Further processing is performed to generate the final change detection result O∈R. 1×H×W This process is completed by the final decoding module: During the decoding process, in order to match the resolution of the input image, the final detection result O is adjusted to the input size using bilinear interpolation:

[0099] O final =Upsample(O,size=(H,W),mode='bilinear')

[0100] Through the multi-level decoding operations described above, this invention can fully utilize the spatial details and semantic information of features at different levels, thereby generating high-quality change detection results. This design not only improves the model's adaptability to multi-scale changes but also ensures the accuracy and robustness of the output results, providing an efficient and accurate solution for change detection in dual-temporal images.

[0101] like Figure 2 As shown, this application also provides a dual-temporal image change detection system 200 based on multi-scale feature fusion and attention mechanism, comprising:

[0102] Preprocessing module 201 is used to preprocess the original remote sensing image data by using the dual-temporal remote sensing image semantic change detection dataset as the original remote sensing image data.

[0103] The feature extraction module 202 is used to extract multi-scale features of the preprocessed remote sensing image through a dual-branch feature extraction network, wherein the dual-branch feature extraction network adopts the ResNet architecture as the backbone network.

[0104] The feature fusion and decoding module 203 is used to perform cross-temporal fusion on the multi-scale features using the cross-temporal feature fusion module; and to perform upsampling and scale fusion on the fused features through the multi-level feature decoding module, and output the final remote sensing image change detection probability map after decoding.

[0105] The deep neural network architecture of this invention includes a preprocessing module, a feature extraction module, and a feature fusion and decoding module; the feature fusion and decoding module specifically includes a feature fusion and enhancement module, a decoding module, and an output module. Through the efficient collaboration of these modules, the model can extract multi-level, multi-scale change information from input multi-temporal remote sensing images and output change detection results consistent with the size of the input images.

[0106] The feature extraction module is based on an improved version of deep convolutional neural networks (such as ResNet), employing a hierarchical network structure to extract multi-level features from images, including edge information, texture information, and semantic information. In the network design, the fully connected layers and global average pooling layers of ResNet are removed to adapt to the specific needs of change detection tasks, while also supporting flexible adjustments for non-standard three-channel inputs (such as multispectral images).

[0107] The feature fusion and enhancement module achieves efficient feature fusion through a multi-scale convolutional module (PASR) and a self-attention feature enhancement module (SSFC). The PASR module employs multi-scale convolutional kernels (such as 3×3, 5×5, 7×7, and 9×9) to capture variation information at different scales, from local details to global structure. The SSFC module utilizes the difference between global mean features and local features to generate dynamic attention weights. By weighting and adjusting the feature representation, it highlights the salient features of changing regions while effectively suppressing the interference of background noise on the detection results. The fused features are then normalized and processed using activation functions to generate enhanced multi-scale feature representations.

[0108] The decoding module progressively restores high-dimensional features to high-resolution change detection results through layer-by-layer upsampling and low-level feature concatenation. During decoding, skip connections are used to fuse high-level semantic information and low-level detail information, ensuring the accuracy and resolution of the detection results. Specifically, the model progressively decodes the input feature map through multiple decoder modules (Decoder1, Decoder2, Decoder3). Each decoder module includes feature concatenation, convolution processing, and upsampling operations, thereby achieving progressive optimization from high-dimensional features to change detection results.

[0109] The output module generates a final change detection probability map through convolutional layers and interpolation operations. This probability map represents the probability value of each pixel belonging to a changed region. Bilinear interpolation is used to adjust the resolution of the output result to match the size of the input image, facilitating subsequent visualization analysis and practical applications.

[0110] This invention first preprocesses the remote sensing image data, then extracts four levels of features (64, 128, 256, and 512 channels) from the dual-temporal remote sensing images using a dual-branch ResNet backbone network. Next, an innovative multi-scale feature fusion module is employed to achieve cross-temporal feature interaction. This module includes a multi-scale convolutional unit (PASR) and a self-attention feature enhancement unit (SSFC). PASR captures multi-scale contextual information through parallel convolutions with different dilation rates, while SSFC automatically learns attention weights for important regions based on a feature variance normalization mechanism. Finally, a three-level progressive decoder is used to progressively upsample and fuse features at each level, ultimately outputting a pixel-level change probability map.

[0111] Experiments show that the embodiments of the present invention achieve an F1 score of 91.52% and an IoU of 84.36% on the LEVIR-CD dataset, which is an improvement of 4.2-6.8 percentage points compared with existing mainstream methods. This demonstrates that the embodiments of the present invention can be effectively applied to remote sensing monitoring scenarios such as urban planning and disaster assessment. Through the synergistic design of multi-scale feature fusion and attention mechanisms, the embodiments of the present invention significantly improve the accuracy and robustness of change detection in complex scenarios.

[0112] In one embodiment, a computer device is provided, such as Figure 3 As shown, it includes a memory 301, a processor 302, and a computer program stored in the memory 301 and executable on the processor 302. When the processor 302 executes the computer program, it implements the above-described dual-phase image change detection method.

[0113] Specifically, the computer device can be a computer terminal, a server, or a similar computing device.

[0114] In this embodiment, a computer-readable storage medium is provided, which stores a computer program that performs the above-described dual-temporal image change detection method.

[0115] Specifically, computer-readable storage media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media does not include transient media, such as modulated data signals and carrier waves.

[0116] Obviously, those skilled in the art should understand that the modules or steps of the above-described embodiments of the present invention can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the embodiments of the present invention are not limited to any particular hardware and software combination.

[0117] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for detecting changes in dual-temporal images based on multi-scale feature fusion and attention mechanism, characterized in that, include: The semantic change detection dataset of dual-temporal remote sensing images is used as the original remote sensing image data, and the original remote sensing image data is preprocessed. Multi-scale features of the preprocessed remote sensing images are extracted by a dual-branch feature extraction network, wherein the dual-branch feature extraction network adopts the ResNet architecture as the backbone network. A cross-temporal feature fusion module is used to perform cross-temporal fusion on the multi-scale features respectively; a multi-level feature decoding module is used to upsample and scale-fuse the fused features, and after decoding, the final remote sensing image change detection probability map is output.

2. The dual-temporal image change detection method based on multi-scale feature fusion and attention mechanism according to claim 1, characterized in that, Preprocessing the raw remote sensing image data includes: The remote sensing image data and labels in the random part of the training sample are flipped horizontally, randomly cropped and rotated. At the same time, the remote sensing images are color enhanced and salt and pepper noise is added to the labels. The remote sensing image data, excluding the random portion mentioned above, is stitched together with three random batches of remote sensing image data and labels. Then, the new remote sensing image data and corresponding labels are flipped left and right and rotated randomly. At the same time, the remote sensing images are color-enhanced and salt-and-pepper noise is added to the labels.

3. The dual-temporal image change detection method based on multi-scale feature fusion and attention mechanism according to claim 1, characterized in that, The dual-branch feature extraction network uses the ResNet 18 architecture as its backbone network. The backbone network structure of the ResNet 18 architecture includes multiple residual modules, each of which contains two or three convolutional layers and achieves layer-by-layer addition of input and output through skip connections.

4. The dual-temporal image change detection method based on multi-scale feature fusion and attention mechanism according to claim 3, characterized in that, The ResNet 18 architecture removes fully connected layers from its backbone network structure and retains four feature output layers. Each feature output layer extracts features from different receptive fields by adjusting the convolution kernel size and downsampling stride.

5. The dual-temporal image change detection method based on multi-scale feature fusion and attention mechanism according to claim 1, characterized in that, The multi-scale features are fused across time using a cross-temporal feature fusion module, including: Feature representations under different receptive fields are extracted simultaneously by multi-scale convolutional units, and the input multi-scale features are concatenated along the channel dimension; the attention weight of each channel is measured by a channel attention mechanism to obtain the channel attention weight coefficient, and the channel attention weight coefficient is multiplied by the concatenated features to obtain the enhanced features; Local features are extracted from the input features through convolution operations. These local features are then fused with the enhanced features, and batch normalization and ReLU activation operations are performed to obtain the fused features. Calculate the channel mean and variance of the fused features, generate adaptive enhancement weights based on the channel mean and variance, and multiply the adaptive enhancement weights with the input fused features to generate the final fused features.

6. The dual-temporal image change detection method based on multi-scale feature fusion and attention mechanism according to claim 1, characterized in that, The multi-level feature decoding module upsamples and scales the fused high-level features, and after decoding, outputs the final remote sensing image change detection probability map, including: The high-level feature map is upsampled using bilinear interpolation to generate a feature map with double resolution. The double-resolution feature map is then concatenated with the lower-level feature map along the channel dimension to form an intermediate feature map. The intermediate feature map is input into the decoding unit of the dual-branch architecture for decoding. The dual-branch architecture includes a local branch and a global branch. The decoding unit performs channel rearrangement on the input intermediate feature map, and inputs the rearranged data into the local branch and the global branch respectively to extract local features and global features. The local features and the global features are then fused to output the decoded feature map.

7. The dual-temporal image change detection method based on multi-scale feature fusion and attention mechanism according to claim 6, characterized in that, The local branch extracts local features from the feature map through a convolutional layer-normalization-activation function layer, while the global branch extracts global features using a channel attention and spatial attention concatenation mechanism. The global and local features are then fused through weighted summation. Finally, the final remote sensing image change detection probability map is generated through convolutional layers and interpolation operations.

8. A dual-temporal image change detection system based on multi-scale feature fusion and attention mechanism, characterized in that, include: The preprocessing module is used to preprocess the original remote sensing image data by using the semantic change detection dataset of dual-temporal remote sensing images as the original remote sensing image data. The feature extraction module is used to extract multi-scale features of the preprocessed remote sensing image through a dual-branch feature extraction network, wherein the dual-branch feature extraction network adopts the ResNet architecture as the backbone network. The feature fusion and decoding module is used to perform cross-temporal fusion of the multi-scale features using the cross-temporal feature fusion module; and to perform upsampling and scale fusion of the fused features through the multi-level feature decoding module, and output the final remote sensing image change detection probability map after decoding.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the dual-temporal image change detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that performs the dual-temporal image change detection method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Image change detection method and device, electronic equipment and storage medium

    CN121074038A