A High-Resolution Remote Sensing Image Change Detection Method and Device Combining Transformer and CNN

By combining Transformer and CNN, the global and local features of remote sensing images are extracted and fused, the problem that the existing technology is difficult to understand the change law from a global perspective is solved, and the accurate detection of the changes of remote sensing images is achieved.

CN116310828BActive Publication Date: 2025-06-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310289429.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-23
Publication Date
2025-06-27
Estimated Expiration
2043-03-23

AI Technical Summary

Technical Problem

Existing change detection methods based on deep learning are difficult to understand the change laws from a global perspective, and it is difficult to distinguish between real changes and pseudo-changes, especially under interference under different imaging conditions.

Method used

Combining Transformer and CNN, the global features of the image are extracted through the Transformer branch, the local features are extracted through the convolutional neural network branch, and fused through the adaptive feature fusion module, and finally the change detection is performed through the Decoder branch.

Benefits of technology

Accurate detection of bi-time phase image changes is achieved, image information can be better expressed and real changes and pseudo-change are distinguished.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310828B_ABST
    Figure CN116310828B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and particularly to a high-resolution remote sensing image change detection method and device combining Transformer and CNN. The method includes using a Transformer branch to extract global features of dual-temporal remote sensing images from preprocessed dual-temporal remote sensing images, and using a convolutional neural network branch to extract local features of the dual-temporal remote sensing images; inputting the feature maps at each depth of the Transformer branch and the feature maps at each depth of the convolutional neural network branch into an adaptive feature fusion module for feature fusion to obtain global-local features; inputting the global-local features into a Decoder branch for layer-by-layer decoding, and using a classifier to output the result of change detection for the decoded feature maps; the present invention can accurately detect the areas where changes occur in the dual-temporal images, and the extracted features can better express the image information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and computer vision, and particularly relates to a high-resolution remote sensing image change detection method and device combining Transformer and CNN. Background Art

[0002] Change detection aims to detect changes in land cover types in a pair of registered images acquired at different times. Changes in man-made facilities (such as buildings, vehicles, etc.), changes in vegetation, and changes in the environment (such as polar ice sheet melting, deforestation, damage caused by disasters) are generally considered as changes related to land cover types. Change detection hopes to identify these related changes while avoiding other complex and irrelevant changes caused by seasonal changes, building shadows, atmospheric changes, and lighting condition changes. Change detection in optical remote sensing images has received extensive attention due to its important role in earth observation and environmental monitoring. With the development of spaceborne / airborne optical imaging technology, change detection in optical remote sensing images has become one of the important tasks in fields such as disaster assessment and urban land renewal surveys.

[0003] In recent years, with the development of deep learning technology and the increase in optical remote sensing data, many deep learning-based change detection methods have been proposed. Due to their powerful automatic feature extraction ability, these methods are superior to traditional change detection methods based on handcrafted features. Since convolutional operations have strong local modeling capabilities, in existing deep learning-based change detection methods, convolutional neural networks are widely used for the extraction of remote sensing image features. However, due to the inherent local feature extraction ability of convolutional operations and the limited receptive field, methods based on pure convolutional structures are difficult to understand the change law from a global perspective and are difficult to distinguish between real changes and pseudo-changes in bi-temporal images under the interference of different imaging conditions (such as illumination changes and atmospheric environment). Summary of the Invention

[0004] In order to achieve high-resolution remote sensing image change detection, the present invention provides a high-resolution remote sensing image change detection method combining Transformer and CNN, which specifically includes the following steps:

[0005] Crop paired bi-temporal remote sensing images to a fixed size and preprocess the cropped images;

[0006] Construct a convolutional neural network branch, a Transformer branch, an adaptive feature fusion module, and a Decoder branch;

[0007] Use the Transformer branch to extract the global features of the bi-temporal remote sensing images and the convolutional neural network branch to extract the local features of the bi-temporal remote sensing images from the preprocessed bi-temporal remote sensing images;

[0008] Input the feature maps of each depth of the Transformer branch and the feature maps of each depth of the convolutional neural network branch into the adaptive feature fusion module for feature fusion to obtain global-local features;

[0009] Input the global-local features into the Decoder branch for layer-by-layer decoding, and use a classifier to output the results of change detection for the decoded feature maps.

[0010] Furthermore, when preprocessing the cropped images, data augmentation is performed using random flipping, random blurring, and random color adjustment.

[0011] Furthermore, the convolutional neural network branch uses a ResNet50 backbone network, which includes five cascaded units. Among them: the first unit includes cascaded convolutional layers, BN layers, ReLU activation functions, and MaxPooling layers, and this segment downsamples the input feature maps by a factor of 4; the other four units include 3, 4, 6, and 3 residual blocks respectively, and each residual block includes a convolutional layer with a 1×1 convolutional kernel, a convolutional layer with a 3×3 convolutional kernel, and a convolutional layer with a 1×1 convolutional kernel, and the convolutional layer with a 3×3 convolutional kernel is downsampled through a convolutional operation with a stride of 2.

[0012] Furthermore, the Transformer branch includes four cascaded units. Among them: the first unit divides the picture into patches through PatchEmbedding; the other three units pass through 2, 2, and 6 VIT blocks respectively, and each VIT block includes an Encoder in the Transformer, and each stage also includes a downsampling with a factor of 2, so that the output feature map of each time is 1 / 2 of the input.

[0013] Furthermore, the VIT block in the Transformer branch adopts a two-stream cross-attention mechanism, that is, the two images of the two-temporal remote sensing images are used as the input of the VIT block, and the processing of the two images includes the following steps:

[0014] Obtain the first query vector Q1, the first key vector K1, and the first value vector V1 of the first image in the two-temporal remote sensing images, and the second query vector Q2, the second key vector K2, and the second value vector V2 of the second image;

[0015] Multiply the first query vector Q1 by the second key vector K2, perform a softmax operation, and then multiply by the second value vector V2 to obtain the first differential feature of the two images;

[0016] Subtract the first differential feature from the first value vector Q1 to obtain the differential feature dominated by the first image, and use this differential feature as the input of the next level;

[0017] After multiplying the second query vector Q2 by the first key vector K1, performing a softmax operation, and then multiplying by the first value vector V1, the second differential features of the two images are obtained;

[0018] Subtract the second differential features from the second value vector Q1 to obtain the difference features dominated by the second image, and use the difference features as the input for the next level.

[0019] Preferably, the first VIT block takes the dual-temporal remote sensing image of the current unit as input, and its subsequent cascaded VIT blocks take the two feature images output by the previous level as input.

[0020] Furthermore, when the adaptive feature fusion module performs fusion, it specifically includes the following steps:

[0021] Integrate the features from the local branch and the global branch by element-wise summation, and then embed the integrated features into the channel space through global average pooling operation;

[0022] Use two fully connected layers to generate two adaptive attention weights for the local branch and the global branch respectively with the help of the channel-embedded features;

[0023] Fuse the global and local features by using the attention weights.

[0024] Furthermore, the Decoder branch includes five units. Each unit includes an upsampling operation, a convolutional layer that halves the number of feature channels, and two 3×3 convolutional layers, each followed by a ReLU; a 1×1 convolution is used in the last unit to map the feature vector to 2.

[0025] The present invention also provides a high-resolution remote sensing image change detection device combining Transformer and CNN for implementing a high-resolution remote sensing image change detection method combining Transformer and CNN, including a preprocessing module, a convolutional neural network branch, a Transformer branch, an adaptive feature fusion module, and a Decoder branch, where:

[0026] The preprocessing module is used to crop the input dual-temporal remote sensing image to a fixed size and perform image enhancement on the cropped image;

[0027] The convolutional neural network branch is used to extract the global features in the dual-temporal remote sensing image;

[0028] The Transformer branch is used to extract the local features in the dual-temporal remote sensing image;

[0029] The adaptive feature fusion module is used to fuse the global features and the local features to obtain global-local features;

[0030] The Decoder branch is used to determine whether there are changes in the dual - time remote sensing images based on the global - local features.

[0031] The beneficial effects of the present invention are as follows:

[0032] 1) The present invention proposes a high - resolution remote sensing image change detection method combining Transformer and CNN, which can accurately detect the changed areas in the dual - time images.

[0033] 2) This method uses a convolutional neural network to extract local features of the image, uses Transformer to extract global features of the image, and uses an adaptive feature fusion module to fuse the global features and local features, so that the extracted features can better represent the image information. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 FIG. is a schematic diagram of the overall process of a high - resolution remote sensing image change detection method combining Transformer and CNN;

[0035] Figure 2 FIG. is a schematic diagram of the Transformer - CNN model structure;

[0036] Figure 3 FIG. is a schematic diagram of the structure of the adaptive feature fusion module;

[0037] Figure 4 FIG. is a schematic diagram of the two - stream cross attention in the Transformer branch. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] The present invention provides a high - resolution remote sensing image change detection method combining Transformer and CNN, which specifically includes the following steps:

[0040] Crop the paired dual - time remote sensing images to a fixed size and pre - process the cropped images;

[0041] Construct a convolutional neural network branch, a Transformer branch, an adaptive feature fusion module, and a Decoder branch;

[0042] The pre - processed dual - time remote sensing images are used to extract the global features of the dual - time remote sensing images through the Transformer branch, and the local features of the dual - time remote sensing images are extracted through the convolutional neural network branch;

[0043] The feature maps of each depth of the Transformer branch and the feature maps of each depth of the convolutional neural network branch are input into the adaptive feature fusion module for feature fusion to obtain global - local features;

[0044] The global - local features are input into the Decoder branch for layer - by - layer decoding, and the decoded feature maps are output by the classifier to obtain the results of change detection.

[0045] In the present invention, the LEVIR dataset is adopted. First, the LEVIR dataset is cropped. The size of the original dataset images is 1024×1024, and the images are cropped into 224×224 to form images of a fixed size. Then, the parameters of the entire network are initialized, and the network is trained with labeled data. The loss function uses the classification loss function, and the parameters are continuously adjusted to train the network model.

[0046] The LEVIR dataset used this time is collected through the Geogle Earth API, collecting 637 pairs of very high - resolution Geogle Earth (GE) image patches of size 1024×1024 pixels. These dual - temporal images come from 20 different regions in several cities in Texas, USA, and the collection time ranges from 2002 to 2018.

[0047] Figure 1 This is the overall flow diagram of the high - resolution remote sensing image change detection method based on Transformer and CNN in the present invention, as Figure 1 shown. The method of the present invention specifically includes the following steps:

[0048] S1: Pre - processing of data. The paired dual - time remote sensing images I1 and I2 are cropped to a fixed size and divided into a training set, a validation set, and a test set;

[0049] The original dataset contains 637 pairs of dual - time high - resolution remote sensing images of 1024×1024. The dataset is cropped into images of size 224×224, a total of 15925 pairs of 224×224 images, and the data is divided into a training set, a validation set, and a test set.

[0050] S2: Construct a convolutional neural network branch, a Transformer branch, an adaptive feature fusion module, and a Decoder branch;

[0051] Among them, the convolutional neural network branch uses the ResNet50 backbone network, and adjusts the convolutional channel dimensions in several stages of ResNet50 to [8, 16, 32, 64, 256] to adapt to the channel dimensions in the Transformer. The Transformer branch contains 4 stages, the Transformer depth is [2, 2, 6, 2], and the Embedding dimension is selected as 48. The structural schematic diagram of the adaptive module is as shown in Figure 3 shown. The adaptive module uses the convolutional features and Transformer features to generate two weight matrices respectively. After multiplying the convolutional features and Transformer features with the two weight matrices respectively, the feature matrices are added to generate the finally fused global-local features. The decoder uses the decoder of Unet. Each Decoder block contains two convolutional layers, and then an upsampling module is used to double the size.

[0052] S3: Use the Transformer branch and the convolutional neural network branch to extract the global features and local features of the I1 and I2 phases from the processed dual-temporal remote sensing images respectively.

[0053] S4: Send the feature maps of different depths in the five stages of the ResNet50 backbone network and the different scale feature maps in the four stages of the Transformer branch into the adaptive fusion module respectively. Make full use of the global-local features from the Transformer and the convolutional neural network, and send the fused features into the Decoder branch for layer-by-layer decoding.

[0054] S5: Use the classifier to generate the final binary change map from the finally obtained decoded feature map.

[0055] As Figure 2 , this embodiment includes a convolutional neural network branch (CNN branch) and a Transformer branch. The two branches process the dual-temporal remote sensing images respectively.

[0056] This embodiment also provides another specific implementation. A high-resolution remote sensing image change detection method combining Transformer and convolutional neural network in this implementation includes the following steps:

[0057] S1: Crop the paired dual-temporal remote sensing images I1 and I2 to a fixed size, divide them into a training set, a validation set and a test set, and at the same time perform data preprocessing; specifically, crop the dual-temporal high-resolution remote sensing images to a fixed size of 224×224. For the images in the training set, data augmentation is performed by random flipping, random blurring, and random color adjustment.

[0058] S2: Construct a convolutional neural network branch, a Transformer branch, an adaptive feature fusion module, and a Decoder branch. Among them, the convolutional neural network branch uses a ResNet50 backbone network, which includes five stages (each stage is Figure 2 a CNN BLOCK in the middle); the Transformer branch uses a VIT network with 4 stages (each VIT network is Figure 2 a VIT BLOCK in the middle); the Decoder branch uses the Decoder branch in the UNet network (the Decoder branch includes multiple DeBLOCKs and Classifiers as shown in Figure 2 . In this embodiment, the Decoder branch uses skip connections, that is, the first four DeBLOCKs are respectively skip-connected to the first four CNN BLOCKs in the convolutional neural network branch and the VIT BLOCKs in the Transformer branch, and the Classifier layer is skip-connected to the last layer of the convolutional neural network branch), which contains 5 stages, and each stage uses bilinear interpolation for upsampling. Specifically, the specific structures of each branch are as follows:

[0059] 1) For the convolutional neural network branch, use a ResNet50 backbone network, which contains five stages. The starting stage includes: a convolutional layer, a BN layer, a ReLU activation function, and a MaxPooling layer, which downsamples the input by a factor of 4. The next four stages each go through 3, 4, 6, and 3 residual blocks respectively. Each residual block includes: a convolutional layer with a 1×1 convolutional kernel, a convolutional layer with a 3×3 convolutional kernel, and a convolutional layer with a 1×1 convolutional kernel. Among them, the convolutional layer with a 3×3 convolutional kernel downsamples through a convolutional operation with a stride of 2.

[0060] 2) For the Transformer branch, it altogether contains four stages. First, the image is divided into patches through Patch Embedding. The next three stages each go through 2, 2, and 6 VIT blocks respectively. Each VIT block only contains the Encoder in the Transformer. Each stage also includes a downsampling with a factor of 2, so that the output feature map is 1 / 2 of the input each time.

[0061] 3) The adaptive feature fusion module. The adaptive feature fusion module contains three steps: feature integration, attention calculation, and feature selection. First, the adaptive fusion module integrates the features from the local branch and the global branch through element-wise summation, and then embeds the integrated features into the channel space through global average pooling operation. Next, two fully connected layers are used to generate two adaptive attention weights for the local branch and the global branch respectively with the help of the channel-embedded features. Finally, the global and local features are fused by using the attention weights.

[0062] 4) For the Decoder branch, the Decoder branch consists of a total of five stages. Each stage in the Decoder branch includes upsampling of the feature map, followed by a convolutional layer that halves the number of feature channels, two 3x3 convolutions, each followed by a ReLU. In the last layer, a 1x1 convolution is used to map the feature vector to 2, thereby generating a binary change detection map.

[0063] The VIT block in the Transformer branch adopts a two-stream cross-attention mechanism, that is, two images of the two-temporal remote sensing images are used as the input of the VIT block, as Figure 4 , and the processing of the two images includes the following steps:

[0064] Obtain the first query vector Q1, the first key vector K1, the first value vector V1 of the first image in the two-temporal remote sensing image, and the second query vector Q2, the second key vector K2, and the second value vector V2 of the second image;

[0065] After multiplying the first query vector Q1 by the second key vector K2 and performing a softmax operation, then multiplying by the second value vector V2, the first differential feature of the two images is obtained;

[0066] Subtract the first differential feature from the first value vector Q1 to obtain the difference feature dominated by the first image, and use this difference feature as the input of the next level;

[0067] After multiplying the second query vector Q2 by the first key vector K1 and performing a softmax operation, then multiplying by the first value vector V1, the second differential feature of the two images is obtained;

[0068] Subtract the second differential feature from the second value vector Q1 to obtain the difference feature dominated by the second image, and use this difference feature as the input of the next level.

[0069] S3: Use the Transformer branch and the convolutional neural network branch to extract the global features and local features of the I1 and I2 phases from the processed two-temporal remote sensing images respectively. The specific steps are as follows:

[0070] 1) Extract the global features of the two-temporal images I1 and I2. Divide the two-temporal images cropped to 224×224 into 4×4 patches using PatchEmbedding, and then pass through 2, 2, and 6 VIT blocks respectively to extract multiple feature maps of different scales. In the global feature extraction branch, a total of 4 feature maps of different scales are extracted.

[0071] 2) Extract the local features of the dual-temporal images I1 and I2. The dual-temporal images with a size of 224×224 are directly input into the ResNet50 backbone network without loading the pre-trained model. To facilitate the fusion with the features of the global feature branch, in the present invention, the channel dimension of feature extraction in the ResNet50 backbone network is reduced to 8, 16, 32, 64, 128, and 256. In the local feature extraction branch, five stages of feature maps with different scales are extracted.

[0072] S4: Send the feature maps with different depths in the five stages of the ResNet50 backbone network and the feature maps with different scales in the four stages of the Transformer branch into the adaptive fusion module respectively, make full use of the global-local features from the Transformer and the convolutional neural network, and send the fused features into the Decoder branch for layer-by-layer decoding. The specific steps are as follows:

[0073] 1) Send the global features and local features extracted by the Transformer branch and the convolutional branch into the adaptive feature fusion module respectively. The adaptive feature module is shown in the attached figure. The adaptive feature fusion module first concatenates the global features and local features of the dual-temporal images respectively, then performs an addition operation on the concatenated global feature matrix and local feature matrix, and then after the global average pooling operation, two learnable weight matrices are generated by two fully connected layers respectively and multiplied with the concatenated global feature matrix and local feature matrix, and the final results are added to obtain the final fused feature.

[0074] 2) The decoder adopts the decoder of Unet. Each Decoder block contains two convolutional layers, and then an upsampling module is used to double the size.

[0075] S5: Use a classifier to generate the final binary change map for the finally obtained decoded feature map. Specifically, for the feature map finally generated in step 4, a classifier is constructed to perform binary classification in the channel dimension to generate the final binary change map. The classifier is specifically implemented as: using a 2D convolution to perform channel dimension transformation to generate the final binary change map.

[0076] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A high-resolution remote sensing image change detection method combining Transformer and CNN, characterized in that, Specifically, it includes the following steps: Crop the paired dual-temporal remote sensing images to a fixed size and preprocess the cropped images; Construct a convolutional neural network branch, a Transformer branch, an adaptive feature fusion module, and a Decoder branch. The Transformer branch includes four cascaded units, where: the first unit divides the picture into patches through Patch Embedding; the other three units go through 2, 2, and 6 VIT blocks respectively. Each VIT block includes the Encoder in Transformer, and each stage also includes a downsampling with a coefficient of 2, so that the output feature map is 1 / 2 of the input each time. The VIT blocks in the Transformer branch adopt a two-stream cross-attention mechanism, that is, the two images of the dual-temporal remote sensing images are used as the input of the VIT blocks. The processing of the two images includes the following steps: Obtain the first query vector Q1, the first key vector K1, and the first value vector V1 of the first image in the dual-temporal remote sensing images, and the second query vector Q2, the second key vector K2, and the second value vector V2 of the second image; Multiply the first query vector Q1 by the second key vector K2, perform a softmax operation, and then multiply by the second value vector V2 to obtain the first differential feature of the two images; Subtract the first differential feature from the first value vector Q1 to obtain the differential feature dominated by the first image, and use this differential feature as the input of the next level; Multiply the second query vector Q2 by the first key vector K1, perform a softmax operation, and then multiply by the first value vector V1 to obtain the second differential feature of the two images; Subtract the second differential feature from the second value vector Q1 to obtain the differential feature dominated by the second image, and use this differential feature as the input of the next level; When the adaptive feature fusion module performs fusion, it specifically includes the following steps: Integrate the features from the local branch and the global branch by element-wise summation, and then embed the integrated features into the channel space through global average pooling operation; Use two fully connected layers to generate two adaptive attention weights for the local branch and the global branch respectively with the help of the channel-embedded features; Fuse the global and local features by using the attention weights; Use the Transformer branch to extract the global features of the dual-temporal remote sensing images from the preprocessed dual-temporal remote sensing images, and use the convolutional neural network branch to extract the local features of the dual-temporal remote sensing images; Input the feature maps of each depth of the Transformer branch and the feature maps of each depth of the convolutional neural network branch into the adaptive feature fusion module for feature fusion to obtain global-local features; Input the global-local features into the Decoder branch for layer-by-layer decoding, and use the classifier to output the result of change detection for the decoded feature map.

2. A high-resolution remote sensing image change detection method combining Transformer and CNN according to claim 1, characterized in that When preprocessing the cropped images, data augmentation is performed by random flipping, random blurring, and random color adjustment.

3. A high-resolution remote sensing image change detection method combining Transformer and CNN according to claim 1, characterized in that, The convolutional neural network branch adopts the ResNet50 backbone network, which includes five cascaded units. Among them: the first unit includes a cascaded convolutional layer, a BN layer, a ReLU activation function, and a MaxPooling layer, and this unit downsamples the input feature map with a coefficient of 4; the other four units respectively include 3, 4, 6, and 3 residual blocks, and each residual block includes a convolutional layer with a 1×1 convolutional kernel, a convolutional layer with a 3×3 convolutional kernel, and a convolutional layer with a 1×1 convolutional kernel, where the convolutional layer with a 3×3 convolutional kernel downsamples through a convolutional operation with a stride of 2.

4. A high-resolution remote sensing image change detection method combining Transformer and CNN according to claim 1, characterized in that, The Decoder branch includes five units, each unit includes an upsampling operation, a convolutional layer that halves the number of feature channels, and two 3×3 convolutional layers, and there is a ReLU after each convolution; a 1×1 convolution is used in the last unit to map the feature vector to 2.

5. A high-resolution remote sensing image change detection device combining Transformer and CNN, characterized in that, A high-resolution remote sensing image change detection method combining Transformer and CNN for implementing the claim 1, including a preprocessing module, a convolutional neural network branch, a Transformer branch, an adaptive feature fusion module, and a Decoder branch, where: The preprocessing module is used to crop the input dual-temporal remote sensing image to a fixed size and perform image enhancement on the cropped image; The convolutional neural network branch is used to extract the global features in the dual-temporal remote sensing image; The Transformer branch is used to extract the local features in the dual-temporal remote sensing image; The adaptive feature fusion module is used to fuse the global features and the local features to obtain global-local features; The Decoder branch is used to determine whether there is a change in the dual-temporal remote sensing image according to the global-local features.

Citation Information

Patent Citations

  • Dual-time remote sensing change detection method combining local representation and global modeling

    CN114821303A

  • Remote sensing image semantic segmentation method based on double-branch feature fusion

    CN115797931A