A Remote Sensing Image Change Detection Method and System

By combining trans-time phase Transformer and convolutional neural network in remote sensing image change detection, a new attention mechanism and Transformer decoder are constructed, which solves the problem that existing methods fail to make full use of Transformer characteristics, and achieves accurate detection of image changes and higher detection effects.

CN115049922BActive Publication Date: 2025-07-01SHANDONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210540288.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2025-07-01
Estimated Expiration
2042-05-18

AI Technical Summary

Technical Problem

The existing change detection methods fail to fully utilize the characteristics of Transformer, resulting in the failure to effectively capture long-term dependencies and sequence information in image change detection.

Method used

A remote sensing image change detection method based on transtime phase Transformer and convolutional neural network is proposed. By constructing a new attention mechanism and Transformer decoder, global change information is captured, and local information is extracted in combination with CNN, and the change detection results are finally obtained through feature stacking.

Benefits of technology

Accurate capture and detection of image changes is achieved, the accuracy and effectiveness of change detection is improved, and the objective evaluation indexes are better than the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115049922B_ABST
    Figure CN115049922B_ABST
Patent Text Reader

Abstract

The present disclosure provides a remote sensing image change detection method and system. Feature extraction is performed on the image before the change and the image after the change is collected, and local information is extracted through CNN. Based on constructing a cross-temporal Transformer and a convolutional neural network structure, the characteristics of the Transformer are fully utilized, so that the images before and after the change respectively enter the cross-temporal Transformer with a new attention mechanism and the Transformer decoder to extract global change information, and the final change detection result is obtained through feature stacking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of image processing, and particularly relates to a remote sensing image change detection method and system. Background Art

[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] Change detection technology usually performs change detection on images of the same area at different time periods. Currently, change detection technology has been applied in many fields, such as land use, urban expansion, farmland change, forest protection, etc. With the rapid development of deep learning, change detection methods based on deep learning have been continuously proposed, including convolutional neural networks, deep neural networks, etc.

[0004] The core of change detection technology is the extraction of image change information. The extraction of image change information mainly includes the following three categories: 1) Pixel-based change information extraction. This type of algorithm takes pixels as processing units and calculates change information pixel by pixel. This method has a good effect on change detection of large areas, but is prone to noise. 2) Time series analysis methods. This type of method is for long-term impact analysis, but has relatively high requirements for time resolution. 3) Deep learning methods. This type of method has an end-to-end network structure and is currently widely used in change detection technology, including neural networks, deep neural networks, and recurrent neural networks, etc. And in recent years, Transformer has also been applied to computer vision. This method can better mine sequence information and capture long-term dependencies. However, the current change detection methods do not fully utilize the characteristics of Transformer and only use Transformer as a tool for feature learning. Summary of the Invention

[0005] In order to solve the above problems, the present disclosure proposes a remote sensing image change detection method and system. Based on constructing a cross-temporal Transformer and a convolutional neural network structure, the characteristics of Transformer are fully utilized, so that the images before and after change respectively enter the cross-temporal Transformer with a new attention mechanism and the Transformer decoder to capture global change information, combine with CNN to extract local information, and obtain the final change detection result through feature stacking.

[0006] According to some embodiments, the present disclosure adopts the following technical solutions:

[0007] A remote sensing image change detection method, and the specific training steps include:

[0008] Collect the images before and after the change and perform preprocessing to obtain paired training data;

[0009] Based on the cross-temporal Transformer deep neural network structure, use the convolutional layer to extract features from the image data before and after the change to obtain the original feature map;

[0010] Based on the cross-temporal Transformer structure of the cross-temporal Transformer deep neural network, capture the changed regions of the feature map after feature extraction to obtain the changed features;

[0011] Stack the obtained changed features with the original feature map, perform convolutional feature extraction, and obtain the final change detection map.

[0012] According to some other embodiments, the present disclosure also adopts the following technical solutions:

[0013] An image acquisition module for collecting the images before and after the change and performing preprocessing to obtain paired training data;

[0014] A feature extraction module for using the convolutional layer to extract features from the image data before and after the change to obtain the original feature map;

[0015] A feature block module for capturing the changed regions of the feature map after feature extraction based on the cross-temporal Transformer structure of the cross-temporal Transformer deep neural network to obtain the changed features;

[0016] A feature stacking module for stacking the obtained changed features with the original feature map;

[0017] A feature training module for performing convolutional feature extraction and outputting the final change detection map.

[0018] Compared with the prior art, the beneficial effects of the present disclosure are:

[0019] The present disclosure constructs a cross-temporal Transformer and a convolutional neural network structure, fully captures the features of image changes, utilizes the characteristics of the Transformer, and proposes a new attention mechanism to capture the changed regions, and combines with CNN to achieve accurate change detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure.

[0021] Figure 1Flowchart for implementing the method of the present disclosure;

[0022] Figure 2 Result diagram of change detection of the present disclosure;

[0023] Figure 3 Schematic structural diagram of two cross-temporal Transformers of the present disclosure;

[0024] Figure 4 Schematic structural diagram of the corresponding Multi-Head Cross-Temporal Attention in the cross-temporal Transformer of the present disclosure. Detailed implementation manners

[0025] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.

[0026] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.

[0027] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0028] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as methods, systems, or computer program products. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0029] Embodiment 1

[0030] The present disclosure provides a remote sensing image change detection method, including the following steps:

[0031] S101: Collect the images before and after the change and perform preprocessing to obtain paired training data;

[0032] S102: Based on the deep neural network structure of cross-temporal Transformer, and use the convolutional layer to extract features from the image data before and after the change to obtain the original feature map;

[0033] S103: Use the cross-temporal Transformer structure of the deep neural network based on cross-temporal Transformer to capture the changed regions of the feature map after feature extraction to obtain the changed features;

[0034] S104: Stack the obtained changed features with the original feature map, perform convolutional feature extraction, and obtain the final change detection map.

[0035] In step S101, collect the image before the change and the image after the change of the image to be detected, denoted as I1 and I2 respectively. During training, input the image I1 before the change and the image I2 after the change respectively. Specifically, the data in the public large-scale building change detection dataset can be used, including 637 pairs of high-resolution (0.5m) remote sensing images, and the image size is 1024×1024.

[0036] Preprocess the images, mainly to specify the size of the images, delimit the specifications for the sizes of the images collected before and after the change, divide the images into small blocks of size 3×256×256, and set overlaps between the small blocks to obtain paired training data.

[0037] Construct an improved deep neural network, that is, set a cross-temporal Transformer module in the deep neural network. By constructing a deep neural network containing multiple convolutional modules and cross-temporal Transformer modules, extract features from the images before and after the change and capture the changed regions. In the proposed deep neural network, the first convolutional module uses a residual network (Resnet 18) to downsample the collected and processed images, that is, the input images, to obtain the extracted features.

[0038] In step S102, based on the deep neural network structure of cross-temporal Transformer, and use the convolutional layer to extract features from the image data before and after the change to obtain the original feature map.

[0039] Specifically, after delimiting the specification sizes of the images collected before and after the change, respectively perform feature extraction on the images I1 and I2 before and after the change through two identical residual network structures, and respectively output the original feature maps F1 and F2, with a size of 32×128×128. The image before the change passes through the residual network module to obtain the feature map F1, and the image after the change passes through the residual network module to obtain the feature map F2.

[0040] In step S103, the cross-temporal Transformer structure of the deep neural network based on the cross-temporal Transformer captures the changing regions of the feature maps after feature extraction to obtain the changing features;

[0041] The feature maps F1 and F2 are divided into tokens token1 and token2 through feature chunking. Feature chunking is implemented by a convolution with a kernel size equal to the chunk size and a stride equal to the chunk size, combined with the Flatten and Transpose functions. The chunk size is selected as 16×16, and the sizes of the obtained token1 and token2 are 8×64×32.

[0042] The feature maps F1 and F2 after feature extraction of the images before and after the change are respectively transformed through feature chunking to obtain token1 and token2. Then, query and value are obtained from token1, and key is obtained from token2 and input into the cross-temporal Transformer. The changing regions are captured through the attention mechanism.

[0043] Specifically, feature chunking performs a convolution on the input F1 and F2 with a kernel size equal to the chunk size and a stride equal to the chunk size to obtain a four-dimensional tensor. The last two dimensions of the four-dimensional tensor are the square root of the number of chunks. Then, the four-dimensional tensor is reduced to three dimensions by the flatten function, and the last dimension is the number of chunks. The transpose function swaps the last and the second-to-last dimensions to obtain token1 and token2.

[0044] Query and value are obtained from token2, and key is obtained from token1 and input into another cross-temporal Transformer. The changing regions are captured through the attention mechanism to obtain the new tokens token3 and token4.

[0045] Specifically, token1 undergoes a Linear linear mapping to obtain query and value; token2 undergoes a Linear linear mapping to obtain key; the obtained query, value, and key are simultaneously input into the cross-temporal Transformer, and the changing regions are captured through the newly proposed attention mechanism; token2 undergoes a Linear linear mapping to obtain query and value; token1 undergoes a Linear linear mapping to obtain key; the obtained query, value, and key are input into another cross-temporal Transformer, and the changing regions are obtained through the newly proposed attention mechanism. The new tokens token3 and token4 are obtained.

[0046] The corresponding Multi-Head Cross-Temporal Attention structure in the cross-temporal Transformer is as follows, and its specific expression is given by the formula:

[0047]

[0048]

[0049] Where Q represents query, K represents key, and V represents value; Q1, K1, and V1 are linearly mapped from token1, Q2, K2, and V2 are linearly mapped from token2, and d is the number of columns of Q1 and K1; abs() represents the absolute value operation. Softmax is the normalized exponential function.

[0050] In step S104, the obtained changed features are stacked with the original feature map, and convolutional feature extraction is performed to obtain the final change detection map.

[0051] Among them, the obtained feature maps F1 and F2 are downsampled, token3 and the feature map F1 are input into the Transformer decoder, and token4 and the feature map F2 are input into another Transformer decoder to obtain the change feature maps T1 and T2.

[0052] The obtained change feature map T1 and the feature map F1, and the change feature map T2 and the feature map F2 are respectively stacked to obtain new feature change maps T1' and T2'.

[0053] The obtained feature change maps T1' and T2' are stacked to obtain a feature map T, and then the feature map T is input into the convolutional feature extraction module to output the final change detection map.

[0054] Embodiment 2

[0055] Construct a network structure including a cross-temporal Transformer and a convolutional network, make full use of the characteristics of the Transformer, and let the images before and after the change enter the cross-temporal Transformer and the Transformer decoder with a new attention mechanism respectively, and finally obtain the change detection result through feature stacking. The specific implementation steps are as follows:

[0056] (1) Input images:

[0057] Input the image I1 before the change and the image I2 after the change respectively, and perform feature extraction on the image before the change and the image after the change through a convolutional layer to obtain paired training data.

[0058] (2) Construct a deep neural network based on the cross-temporal Transformer:

[0059] A deep neural network containing multiple convolutional modules and cross-temporal Transformer modules is constructed to extract features from the images before and after the change and capture the changed regions. In the deep neural network proposed by the present invention, the first convolutional module uses a residual network (Resnet18) to extract features from the input image.

[0060] (2a) The image before the change passes through the residual network module to obtain the feature map F1, and the image after the change passes through the residual network module to obtain the feature map F2.

[0061] (2b) The feature maps F1 and F2 are respectively transformed into tokens1 and tokens2 through the feature block module. The feature block is implemented by a convolution with a kernel size of the block size and a stride of the block size in combination with the Flatten and Transpose functions. A convolution with a kernel size of the block size and a stride of the block size is performed on the input F1 and F2 to obtain a four-dimensional tensor. The last two dimensions of the four-dimensional tensor are the square root of the number of blocks. Then, the four-dimensional tensor is reduced to three dimensions by the flatten function, and the last dimension is the number of blocks. The transpose function swaps the last and the second-to-last dimensions to obtain tokens1 and tokens2.

[0062] (2c) The query and value are obtained from token1, and the key is obtained from token2 and input into the cross-temporal Transformer to capture the changed region through the newly proposed attention mechanism; the query and value are obtained from token2, and the key is obtained from token1 and input into another cross-temporal Transformer to obtain the changed region through the newly proposed attention mechanism. New tokens3 and tokens4 are obtained. Token1 undergoes a Linear linear mapping to obtain the query and value; token2 undergoes a Linear linear mapping to obtain the key; the obtained query, value, and key are simultaneously input into the cross-temporal Transformer to capture the changed region through the newly proposed attention mechanism; token2 undergoes a Linear linear mapping to obtain the query and value; token1 undergoes a Linear linear mapping to obtain the key; the obtained query, value, and key are input into another cross-temporal Transformer to obtain the changed region through the newly proposed attention mechanism. New tokens3 and tokens4 are obtained.

[0063] The structures of two cross-temporal Transformers are as Figure 3 shown:

[0064] As Figure 3 shown in Figure (a), Q1, K1, and V1 are three matrices linearly mapped from token1, and K2 is a matrix linearly mapped from token2. K2, Q1, K1, and V1 enter the cross-temporal Transformer, which consists of a multi-head cross-temporal Attention and a feed-forward network layer, and adopts residual connections and normalization modules.

[0065] As Figure 3 shown in Figure (b), Q2, K2, and V2 are three matrices linearly mapped from token2, and K1 is a matrix linearly mapped from token2. K1, Q2, K2, and V2 enter the cross-temporal Transformer, which consists of a multi-head cross-temporal Attention and a feed-forward network layer, and adopts residual connections and normalization modules.

[0066] The corresponding Multi-Head Cross-Temporal Attention structure in the cross-temporal Transformer is as Figure 4 shown:

[0067] As Figure 4 shown in Figure (a), Q1, K1, V1, and K2 are three-dimensional matrices linearly mapped from token1 and token2, and a four-dimensional matrix is obtained through the rearrange function. Then, the transpose of K1 multiplied by Q1 is subtracted from the transpose of K2 multiplied by Q1, and the absolute value is taken. The resulting value is divided by the square root of d, where d is the dimension of Q. Finally, it is multiplied by V1 and normalized.

[0068] As Figure 4 shown in Figure (b), Q2, K2, V2, and K1 are three-dimensional matrices linearly mapped from token2 and token1, and a four-dimensional matrix is obtained through the rearrange function. Then, the transpose of K2 multiplied by Q2 is subtracted from the transpose of K1 multiplied by Q2, and the absolute value is taken. The resulting value is divided by the square root of d, where d is the number of columns of Q. Finally, it is multiplied by V2 and normalized.

[0069] The specific expression of this structure is given by the formula:

[0070]

[0071]

[0072] Among them, Q1, K1, and V1 are matrices linearly mapped from token1, Q2, K2, and V2 are matrices linearly mapped from token2, d is the number of columns of Q and K, abs() represents the absolute value operation, and Softmax is the normalized exponential function.

[0073] (2d), Downsample F1 and F2. Input the token1 obtained in (2c) and F1 obtained in (2b) into the Transformer decoder, and input the token2 obtained in (2c) and F2 obtained in (2b) into another Transformer decoder to obtain the change features T1 and T2.

[0074] (2e), Stack the feature maps T1 obtained in (2d) and F1 in (2b), and stack the feature maps T2 obtained in (2d) and F2 in (2b). Obtain the new feature maps T1' and T2'.

[0075] (2f), Stack the feature maps T1' and T2' obtained in (2e) to obtain the feature map T.

[0076] (2g), Input the feature map T in (2f) into the convolutional feature extraction module to obtain the final change detection map.

[0077] (3) Use the training samples generated in step 1 and the stochastic gradient descent algorithm to train the network. The loss function is to construct the objective equation: minimize the cross-entropy loss to optimize the network parameters.

[0078]

[0079] Among them, l(P hw , y) = -log(P hwy ) is the cross-entropy loss, and Y hw is the pixel value at the (h, w) position.

[0080] (4) Training and testing:

[0081] Use the training samples obtained in step (1), and adopt the stochastic gradient descent algorithm to train the deep neural network to obtain the trained deep neural network. Input the images before and after the change to be detected into the trained deep neural network to obtain the image of the changed part after change detection.

[0082] Embodiment 3

[0083] The present disclosure also provides a remote sensing image change detection system, specifically including:

[0084] An image acquisition module, configured to acquire images before and after changes and perform preprocessing to obtain paired training data;

[0085] A feature extraction module, configured to use a convolutional layer to extract features from the image data before and after changes to obtain an original feature map;

[0086] A feature chunking module, configured to capture the changed regions of the feature map after feature extraction based on the cross-temporal Transformer structure of a deep neural network with cross-temporal Transformer to obtain changed features;

[0087] A feature stacking module, configured to stack the obtained changed features and the original feature map;

[0088] A feature training module, configured to perform convolutional feature extraction and output a final change detection map.

[0089] The above modules implement the following method steps:

[0090] S101: Acquire images before and after changes and perform preprocessing to obtain paired training data;

[0091] S102: Based on the deep neural network structure with cross-temporal Transformer, use a convolutional layer to extract features from the image data before and after changes to obtain an original feature map;

[0092] S103: Capture the changed regions of the feature map after feature extraction based on the cross-temporal Transformer structure of a deep neural network with cross-temporal Transformer to obtain changed features;

[0093] S104: Stack the obtained changed features and the original feature map, perform convolutional feature extraction, and obtain a final change detection map.

[0094] The effects of the present disclosure can be further illustrated by the following simulations:

[0095] 1. Simulation environment:

[0096] PyCharm Community Edition 2021.02 x64, NVIDIA 2080Ti GPU, Ubuntu 16.04.

[0097] 2. Simulation content:

[0098] Simulation 1. In this disclosure, data from the public large building change detection dataset is used, which includes 637 pairs of high-resolution (0.5m) remote sensing images with a size of 1024×1024 and a time span of 5 to 14 years of bit images. The detection results are as Figure 2 shown, where:

[0099] Figure 2 (a) is the remote sensing image before the change, with a size of 3×256×256.

[0100] Figure 2 (b) is the remote sensing image after the change, with a size of 3×256×256.

[0101] Figure 2 (c) is the change detection image obtained by performing change detection on Figure 2 (a) and Figure 2 (b) by the present invention, with a size of 3×256×256.

[0102] Figure 2 (d) is Figure 2 (a) and Figure 2 (b) The true value of the change detection is 3×256×256.

[0103] It can be Figure 2 seen that the detection results of the method proposed in this disclosure are basically consistent with the true values, achieving a very accurate change detection effect.

[0104] Simulation 2. To prove the effect of the present invention, the method of the present invention is used to perform change detection on the images of Figure 2 (a) and Figure 2 (b), and objective index evaluation is performed on the detection results. The evaluation indexes are all in %, and the results are shown in Table 1.

[0105] Table 1. Objective evaluation of change detection results of various methods

[0106]

[0107]

[0108] The evaluation indexes are as follows: Pre is the precision, also known as the recall rate. That is, the proportion of correctly predicted positive samples among all detected predicted positive samples. The value range is [0,1], and the larger the better.

[0109] Recall is the recall rate, also known as the completeness rate. That is, the proportion of correctly predicted samples among all positive samples. The value range is [0,1], and the larger the better.

[0110] F1 is the harmonic mean of the precision and the accuracy rate. The value range is [0,1], and the larger the better.

[0111] IoU is the intersection over union. The value ranges from [0, 1], and the larger the value, the better.

[0112] OA is the overall accuracy rate. That is, the proportion of correctly detected samples in all samples. The value ranges from [0, 1], and the larger the value, the better.

[0113] As can be seen from Table 1, the F1, IoU, and Recall of the present disclosure are all greater than the evaluation indicators of the prior art. It can be seen from this that most of the objective evaluation indicators of the present invention are superior to those of the prior art.

[0114] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0115] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0117] The above are only the preferred embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, the present disclosure can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

[0118] Although the specific embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, they are not intended to limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications or variations that can be made without creative efforts on the basis of the technical solutions of the present disclosure are still within the scope of protection of the present disclosure.

Claims

1. A remote sensing image change detection method, characterized in that, The specific training steps include: Collect the images before and after the change and perform preprocessing to obtain paired training data; Based on the deep neural network structure of the cross-temporal Transformer, and use the convolutional layer to extract features from the image data before and after the change to obtain the original feature map; Based on the cross-temporal Transformer structure of the deep neural network of the cross-temporal Transformer, capture the changing regions of the feature map after feature extraction to obtain the changing features; Stack the obtained changing features with the original feature map, perform convolutional feature extraction, and obtain the final change detection map; Based on the cross-temporal Transformer structure of the deep neural network of the cross-temporal Transformer, capture the changing regions of the feature map. Specifically: the feature maps F1 and F2 after feature extraction of the images before and after the change are respectively subjected to feature block transformation to obtain token1 and token2, and then query and value are obtained from the token1, and key is obtained from the token2 and input into the cross-temporal Transformer, and the changing regions are captured through the attention mechanism; The token2 obtains query and value, and the token1 obtains key and inputs them into another cross-temporal Transformer to capture the changing regions through the attention mechanism to obtain new token3 and token4.

2. The remote sensing image change detection method according to claim 1, characterized in that, Specify the specifications for the sizes of the images collected before and after the change, with the size being 3×256×256. Then, the images before and after the change respectively pass through two identical residual network structures for feature extraction, and the original feature maps F1 and F2 are respectively output, with the size being 32×128×128.

3. The remote sensing image change detection method according to claim 1, characterized in that The feature block is realized by a convolution with a convolution kernel size of the block size and a step size of the block size in combination with the Flatten and Transpose functions.

4. The remote sensing image change detection method according to claim 1, characterized in that The specific expression of the attention mechanism structure is: where Q1, K1, and V1 are matrices linearly mapped from token1, Q2, K2, and V2 are matrices linearly mapped from token2, d is the number of columns of Q and K, abs() represents the absolute value operation, and Softmax is the normalized exponential function.

5. A remote sensing image change detection method according to claim 1, characterized in that, Specifically, downsample the obtained feature maps F1 and F2, input token3 and the feature map F1 into the Transformer decoder, input token4 and the feature map into another Transformer decoder, and obtain the change feature maps T1 and T2.

6. The remote sensing image change detection method according to claim 5, wherein Stack the obtained change feature maps T1 with the feature map F1 and the change feature maps T2 with the feature map F2 respectively to obtain new feature change maps T1' and T2'.

7. The remote sensing image change detection method according to claim 6, wherein, Stack the obtained feature change maps T1' and T2' to obtain the feature map T, and then input the feature map T into the convolutional feature extraction module to output the final change detection map.

8. A remote sensing image change detection system, characterized in that, Including: An image acquisition module for collecting the images before and after the change and performing preprocessing to obtain paired training data; A feature extraction module, which is used to extract features from the image data before and after the change by using a convolutional layer to obtain an original feature map; A feature block module, which is used to capture the changed area of the feature map after feature extraction based on the cross-temporal Transformer structure of a deep neural network based on the cross-temporal Transformer to obtain changed features; A feature stacking module, which is used to stack the obtained changed features with the original feature map; A feature training module, which is used to perform convolutional feature extraction and output a final change detection map; Capturing the changed area of the feature map based on the cross-temporal Transformer structure of a deep neural network based on the cross-temporal Transformer specifically includes: respectively performing feature block transformation on the feature maps F1 and F2 obtained by feature extraction of the images before and after the change to obtain token1 and token2, then obtaining query and value from the token1, and obtaining key from the token2 and inputting them into the cross-temporal Transformer, and capturing the changed area through the attention mechanism; The token2 obtains query and value, and the token1 obtains key and inputs them into another cross-temporal Transformer to capture the changed area through the attention mechanism to obtain new token3 and token4.

Citation Information

Patent Citations

  • Remote sensing image change detection method and device, computer equipment and storage medium

    CN114022788A

  • Transform-based remote sensing VHR image change detection method

    CN114170154A