A method for change detection of optical remote sensing images based on deep learning
By combining Resnet18 and feature pyramid structure for preliminary feature extraction and information fusion, and then using Vision Transformer for further feature extraction, the existing optical remote sensing image change detection methods are solved in terms of detection accuracy and universality, and more accurate and efficient change detection is achieved.
Patent Information
- Application Number
- CN202311080420.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-08-25
AI Technical Summary
The existing optical remote sensing image change detection methods have insufficient detection accuracy, and the probability of false detection and missed detection is relatively high, and the universality is poor on different data sets.
The two-step progressive method is used to extract the changing features. First, the lightweight network Resnet18 is used for preliminary feature extraction, and then the feature pyramid structure is used to fuse low-level details and high-level semantic information. The fused features are sent to Vision Transformer for further feature extraction, and finally the differential feature extraction module is used to obtain the changing content.
It improves the accuracy of changing feature modeling and extraction, reduces the probability of false detection and missed detection, and shows high universality on different data sets.
Smart Images

Figure CN117152630B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of optical remote sensing image change detection, and relates to an optical remote sensing image change detection method based on deep learning. Background Art
[0002] In recent years, with the rapid development of satellite remote sensing technology, China has continuously launched multi-sensor remote sensing satellites. Remote sensing satellite sensors can detect ground conditions and obtain a large amount of remote sensing images in real time. Remote Sensing Change Detection (RSCD) generally refers to detecting the change content of two or more remote sensing images obtained at different times, including temporal change information and spatial change information. The detection results can be the position change of ground objects or the morphological change of ground objects, and it has important application value and significance in many fields such as land resource planning, urban expansion management, natural disaster assessment, ecological environment monitoring, military activity monitoring, and ocean resource exploration.
[0003] Remote sensing images can be divided into Synthetic Aperture Radar (SAR) images and optical remote sensing images, and optical remote sensing images exhibit the following characteristics: high spatial resolution, high temporal resolution, and high spectral resolution. Therefore, they have received much attention in the field of change detection.
[0004] The change detection methods for remote sensing images include classical machine learning algorithms and deep learning-based algorithms. Classical machine learning includes mathematical analysis methods, spatial transformation methods, and feature classification methods, etc. Its advantages are that the results are highly interpretable and have low requirements for data volume. However, it also has some disadvantages, such as low detection accuracy, more manual intervention, low automation, weak big data processing ability, and large time consumption. For optical remote sensing images, due to the significant improvement of its spectral resolution, and at the same time, the spatial resolution and temporal resolution of the images have also been greatly improved. Therefore, traditional machine learning algorithms cannot meet the requirements for the change detection accuracy of optical remote sensing images. Change detection methods related to neural networks and deep learning have emerged as the times require. Deep learning-based methods mainly detect the change information of remote sensing images by designing various network models. There are mainly two major directions. The first direction is to send two or more remote sensing images into the network model for feature extraction and recognition respectively, and then compare the extracted features to obtain the change detection result map. The second direction is to first fuse two or more remote sensing images, such as image subtraction, image stitching, etc., and then send them into the change detection network to obtain the final detection result. Currently, domestic and foreign remote sensing image change detection methods based on deep learning include using the idea of transfer learning and the Long Short-Term Memory (LSTM) mechanism to obtain the change information of long-sequence remote sensing images; there are also methods that design a U-Net network with dense skip connections for multi-scale feature fusion to reduce the detection error. With the introduction of the Transformer mechanism from the field of natural language processing to the visual field, many change detection methods based on the Transformer mechanism have emerged one after another, including the BIT model proposed by researchers to solve the problems of high detection difficulty for high-resolution remote sensing images with complex scenes and large difficulty in capturing the context relationship of spatio-temporal information; there are also models based on the sliding window Swin Transformer to solve the problem that information cannot be exchanged between windows during the feature extraction process in the basic Transformer model. Summary of the Invention
[0005] The object of the present invention is to provide a method for change detection of optical remote sensing images based on deep learning, which uses a two-step progressive method to extract change features. First, the lightweight network Resnet18 is used for preliminary feature extraction, then the feature pyramid structure is used to fuse the detailed information contained in the low level and the semantic information contained in the high level, and then the fused features are fed into the Vision Transformer for further feature extraction. For the features before and after the change obtained by the extraction, the differential feature extraction module is used to obtain the corresponding changed content. The present invention improves the accuracy of change feature modeling and extraction, and solves the problems of low detection accuracy, large probability of false detection and missed detection of the existing change detection methods for optical remote sensing images, as well as poor universality of the change detection algorithm on different data sets.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A method for change detection of optical remote sensing images based on deep learning, comprising the following steps:
[0008] Step 1: Obtain multi-temporal optical remote sensing image data and perform preprocessing on it;
[0009] Step 2: Input the image image1 before the change and the image image2 after the change read into the Resnet18 network respectively, and extract the feature maps of different levels from stage1 to stage5 of each image;
[0010] Step 3: For the two images image1 and image2 before and after the change, input the feature maps of different levels from stage1 to stage5 into the feature pyramid structure respectively, fuse the detailed information contained in the low level and the semantics contained in the high level, and obtain the fused feature maps;
[0011] Step 4: Input the fused feature maps into the Vision Transformer model. First, use the token extractor to convert the feature maps into the semantic token space, then use the encoder module to perform context modeling and feature extraction on the tokens, and finally use the decoder module to map the features from the token space to the pixel space, thereby obtaining the feature maps containing rich information;
[0012] Step 4.1: For the features of the two images before and after the change, use the Token extractor to extract the semantic tokens of the fused features respectively, convert the features into the space based on semantic tokens for representation, and obtain the semantic sets Token1 and Token2 respectively;
[0013] Step 4.2: Concatenate the semantic sets Token1 and Token2 corresponding to the two images before the change respectively to obtain the concatenated Tokens;
[0014] Step 4.3: For the concatenated Tokens, perform trainable position embedding. The specific operation is as follows: Add a trainable vector PE with the same size as Tokens to the tokens. Position embedding can perform relative position encoding on the elements in the semantic token space, so it is beneficial for context modeling and feature extraction of features;
[0015] Step 4.4: Send the Tokens with position encoding information to the encoder to perform context information modeling in the semantic token-based space to obtain the semantic set Token_new with rich information. The encoder module is stacked by Ne basic modules, and each module contains a multi-head self-attention module and a multi-layer perceptron;
[0016] Step 4.5: Split Token_new to obtain the corresponding semantic sets Token_new1 and Token_new2 before and after the change respectively;
[0017] Step 4.6: For the semantic sets Token_new1 and Token_new2, use the decoder to project them back into the pixel space to obtain the enhanced feature maps corresponding to before and after the change respectively. The decoder module is stacked by Nd basic modules, and each module contains a multi-head cross-attention module and a multi-layer perceptron;
[0018] Step 5: Input the extracted feature map with rich information into the channel attention module. The channel attention module can calculate the importance of each channel, give different attention to each channel, and thus obtain the enhanced feature map;
[0019] Step 6: Send the enhanced feature maps before and after the change to the differential feature extraction module to obtain the difference feature map of the two remote sensing images;
[0020] Step 7: Send the difference feature map to the classifier to obtain the final change detection binary image, and compare it with the change reference map.
[0021] Furthermore, the specific steps of Step 1 are as follows:
[0022] Step 1.1: After obtaining the multi-temporal optical remote sensing image data, segment it into non-overlapping image blocks with a resolution of 256*256, and divide it into three data sets: training set, validation set, and test set according to the original ratio of the data bureau;
[0023] Step 1.2: Read the segmented remote sensing image;
[0024] Step 1.3: Perform data augmentation operations on the read training set, including random horizontal flipping, random vertical flipping, random image cropping, and Gaussian blurring of the image. The validation set and test set do not require data augmentation operations. The purpose of data augmentation is to increase the complexity during the training process and enhance the robustness of the model.
[0025] Furthermore, the specific steps of step 4.1 are as follows:
[0026] Step 4.1.1: For the input feature map X i , obtain L feature maps by performing two-dimensional convolution on it and perform a softmax operation to obtain the feature map A with spatial attention information i , and the calculation formula is as follows:
[0027] A i = softmax(Conv(X i ))
[0028] Step 4.1.2: Use the feature map A with spatial attention information i to perform pixel-level weighted calculation on the original input feature map X i to obtain the semantic token set T of length L i , and the calculation formula is as follows:
[0029]
[0030] Furthermore, the specific steps of step 4.4 are as follows:
[0031] Step 4.4.1: The concatenated tokens with position encoding information are input into the encoder. Here, the network is divided into two branches. The main branch first passes through layer normalization (Layer Norm), and then through the multi-head self-attention module. At the same time, a residual branch is introduced to fuse with the main branch; the calculation formula of the multi-head attention module is as follows:
[0032] MHAtten(Q, K, V) = Concat(h1,…,h i ,…,h n )W O
[0033] h i = Atten(QW i Q , KW i K , VW iV )
[0034] where h i is the single-head attention mechanism, and its calculation formula is:
[0035]
[0036] where Q, K, and V respectively represent the Query query vector, Key key vector, and Value value vector, and W i Q , W i K , W i V respectively represent the mapping matrices corresponding to the three vectors Q, K, and V in each head, and W O represents the output mapping matrix, and d k represents the vector dimension size. It can be seen from the formula that the single-head attention mechanism calculates how much attention should be given to the V vector through the Q vector and the K vector, and multiplies it by the V vector, so that the V vector with attention information can be obtained; while the multi-head attention mechanism simultaneously executes multiple single-head attentions in parallel, concatenates their results, and finally performs output mapping; for the multi-head self-attention module, its three vectors Q, K, and V all come from the same object, that is, the tokens with position encoding information obtained in step 4.3;
[0037] Step 4.4.2: The data from the above step 4.2.1 also passes through two branches respectively. The main branch sequentially passes through layer normalization Layer Norm and a multi-layer perceptron, and at the same time introduces a residual branch to be added and fused with the main branch.
[0038] Furthermore, the specific steps of step 4.6 are as follows:
[0039] Step 4.6.1: Input the semantic sets Token_new1 and Token_new2 from step 4.5 into the decoder respectively. Here, the network is also divided into two branches. The main branch first passes through layer normalization Layer Norm, and then passes through the multi-head cross-attention module, and at the same time introduces a residual branch to be fused with the main branch; the multi-head cross-attention MCA also contains three vectors Q, K, and V. Different from the multi-head self-attention MSA, the three vectors Q, K, and V of the multi-head cross-attention MCA come from different objects, where the Q vector comes from the feature X fused by the feature pyramid in step 3.1 i , while the K and V vectors come from the features Token_new encoded by the encoder module segmented in step 4.5 i ;
[0040] Step 4.6.2: The data from the above Step 4.6.1 also passes through two branches respectively. The main branch sequentially passes through Layer Normalization (Layer Norm) and a multi-layer perceptron, and at the same time, a residual branch is introduced to be added and fused with the main branch.
[0041] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:
[0042] 1. An optical remote sensing image change detection method based on deep learning, which adopts a two-place progressive network structure, combines a lightweight Resnet18 network structure with a Vision Transformer model based on a multi-head attention mechanism. First, the Resnet18 network is used to extract preliminary feature information, and then the encoder and decoder modules in the Vision Transformer model are used to model and capture the context information of the features, so as to obtain enhanced feature information.
[0043] 2. In step 3 of the present invention, a feature pyramid structure is used to fuse the information from different levels in the Resnet18 network, and the detailed information contained in the low level and the semantic information contained in the high level are well fused together, so that the information contained in the obtained feature map is more abundant.
[0044] 3. In step 4.4 of the present invention, in the Transformer encoder module, multi-head self-attention (MSA) is used to extract context information from the semantic token space. The Q, K, and V vectors all come from the Tokens with position encoding information. In the Transformer decoder module in step 4.6, multi-head cross-attention (MCA) is adopted. The Q, K, and V vectors come from different objects. The Q vector comes from the features obtained by feature pyramid fusion, that is, the original pixel space, while the K and V vectors come from the features encoded by the Transformer encoder module, that is, the semantic token space. In this way, more accurate features can be obtained when projecting the features back to the pixel space, improving the accuracy and integrity of feature extraction.
[0045] 4. In step 6 of the present invention, a difference feature extraction module is introduced. By constructing a network, the module can learn the difference information between the two feature maps before and after the change through training by itself, rather than specifying two feature maps to perform subtraction, reducing manual intervention.
[0046] 5. The present invention is trained and tested on the LEVIR-CD dataset and the DSIFN-CD dataset respectively. The experimental results show that the proposed method effectively optimizes and improves the problems of low detection accuracy, large false detection and missed detection probabilities of existing change detection methods on optical remote sensing images with high resolution, and poor universality of change detection algorithms on different datasets. Brief Description of the Drawings
[0047] The present invention will be described by way of examples with reference to the accompanying drawings, where:
[0048] Figure 1 is a flowchart of an optical remote sensing image change detection method based on deep learning;
[0049] Figure 2 is the distribution of the dataset used in Embodiment 1 of the present invention, where (a) is the division of the LEVIR-CD dataset and (b) is the division of the DSIFN-CD dataset;
[0050] Figure 3 is the structure diagram of the feature pyramid network used in Embodiment 1 of the present invention;
[0051] Figure 4 is the structure diagram of the Vision Transformer model used in Embodiment 1 of the present invention;
[0052] Figure 5 is the structure diagram of the Transformer encoder module used in Embodiment 1 of the present invention;
[0053] Figure 6 is the structure diagram of the Transformer decoder module used in Embodiment 1 of the present invention;
[0054] Figure 7 is the structure diagram of the differential feature extraction module used in Embodiment 1 of the present invention;
[0055] Figure 8 is the display of the change detection results of Embodiment 1 of the present invention on the LEVIR-CD test set, where (a) is the pre-change image, (b) is the post-change image, (c) is the change reference image, and (d) is the obtained detection result image;
[0056] Figure 9 is the display of the change detection results of Embodiment 1 of the present invention on the DSIFN-CD test set, where (a) is the pre-change image, (b) is the post-change image, (c) is the change reference image, and (d) is the obtained detection result image. Detailed Description of the Preferred Embodiments
[0057] All features disclosed in this specification, or all steps in the disclosed methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.
[0058] The following is a detailed description of the present invention in conjunction with Figures 1 to 9 Make a detailed description of the present invention.
[0059] Embodiment 1:
[0060] Step 1: Obtain multi-temporal optical remote sensing image data and preprocess it;
[0061] Step 1.1: After obtaining the multi-temporal optical remote sensing image data, segment it into non-overlapping image patches with a resolution of 256*256, and divide the dataset into three parts: train / val / test according to the original ratio. The dataset division is as Figure 2 shown, where (a) is the dataset division of LEVIR-CD, and (b) is the dataset division of DSIFN-CD;
[0062] Step 1.2: Read the segmented remote sensing images, including the pre-change image, the post-change image, and the corresponding binary change image;
[0063] Step 1.3: Perform data augmentation operations on the read training set, including random horizontal flipping, random vertical flipping, random image cropping, Gaussian blur of the image, etc. The validation set and the test set do not require data augmentation operations. The purpose of data augmentation is to increase the complexity during the training process and enhance the robustness of the model;
[0064] Step 2: Input the pre-change image image1 and the post-change image image2 into the Resnet18 network respectively, and extract the feature maps of different levels from stage1 to stage5 of each image;
[0065] Step 3: For the two images image1 and image2 before and after the change, input the feature maps of different levels from stage1 to stage5 into the feature pyramid structure respectively, so as to well fuse the detailed information contained in the low level and the semantics contained in the high level, and obtain the fused feature maps. The feature pyramid network structure is as Figure 3 shown;
[0066] Step 4: Input the fused feature maps into the Vision Transformer model. The Vision Transformer model is as Figure 4 shown. First, use the token extractor to convert the feature maps into the semantic token space, then use the Transformer encoder module to perform context modeling and feature extraction on the Tokens, and finally use the Transformer decoder module to map the features from the token space to the pixel space, thereby obtaining feature maps containing rich information;
[0067] Step 4.1: For the features of the two images before and after the change, use the token extractor to extract semantic tokens from the features fused at different levels respectively, convert the features to be represented in the space based on semantic tokens, and obtain semantic sets Token1 and Token2 respectively;
[0068] Step 4.1.1: For the input feature map X i , perform two-dimensional convolution on it to obtain L feature maps, and perform a softmax operation to obtain a feature map A with spatial attention information i , and the calculation formula is as follows:
[0069] A i = softmax(Conv(X i ))
[0070] Step 4.1.2: Use the feature map A with spatial attention information i to perform pixel-level weighted calculation on the original input feature map X i to obtain a semantic token set T of length L i , and the calculation formula is as follows:
[0071]
[0072] Step 4.2: Concatenate the semantic sets Token1 and Token2 corresponding to the two images before the change respectively to obtain the concatenated Tokens;
[0073] Step 4.3: For the concatenated Tokens, perform trainable position embedding Position Embedding. The specific operation is: add a trainable vector PE with parameters on the basis of Tokens, and this vector has the same size as Tokens; Position embedding can perform relative position encoding on the elements in the space based on semantic tokens, so it is beneficial to model and extract the context of features;
[0074] Step 4.4: Send the Tokens with position encoding information to the Transformer encoder to model the context information in the space based on semantic tokens, and obtain a semantic set Token_new containing rich information. The Transformer encoder module is stacked by N e basic modules, as Figure 5 shown, and each basic module contains multi-head self-attention Multi-Head Self-Attention (MSA) and multi-layer perceptron Multilayer Perceptron (MLP);
[0075] Step 4.4.1: The concatenated Tokens with positional encoding information are input into the Transformer encoder. Here, the network is divided into two branches. The main branch first passes through Layer Normalization (Layer Norm), and then through Multi-Head Self-Attention (MSA). Meanwhile, a residual branch is introduced to fuse with the main branch. The calculation formula of Multi-Head Attention (MHA) is as follows:
[0076] MHAtten(Q, K, V) = Concat(h1, … h n )W O
[0077] h i = Atten(QW i Q , KW i K , VW i V )
[0078] where h i is the single-head attention mechanism, and its calculation formula is:
[0079]
[0080] where Q, K, and V represent the Query query vector, Key key vector, and Value value vector respectively. W i Q , W i K , W i V represent the mapping matrices corresponding to the three vectors Q, K, and V in each head respectively. W O represents the output mapping matrix, d k represents the vector dimension size. It can be seen from the formula that the single-head attention mechanism calculates how much attention should be given to the V vector by the Q vector and the K vector, and multiplies it with the V vector, so as to obtain the V vector with attention information. The multi-head attention mechanism simultaneously executes multiple single-head attentions in parallel, concatenates their results, and finally performs output mapping. For Multi-Head Self-Attention (MSA), the three vectors Q, K, and V all come from the same object, that is, the Tokens with positional encoding information obtained in Step 4.3;
[0081] Step 4.4.2: The data from the above Step 4.2.1 also passes through two branches respectively. The main branch sequentially passes through Layer Normalization (Layer Norm) and Multilayer Perceptron (MLP), and at the same time, a residual branch is introduced to be added and fused with the main branch;
[0082] Step 4.5: Split Token_new to obtain the corresponding semantic sets Token_new1 and Token_new2 before and after the change respectively;
[0083] Step 4.6: For the semantic sets Token_new1 and Token_new2, use the Transformer decoder to project them back into the pixel space to obtain the enhanced feature maps corresponding to before and after the change respectively; The Transformer decoder module is stacked by N d basic modules, as Figure 6 shown. Each basic module contains Multi-Head Cross-Attention (MCA) and Multilayer Perceptron (MLP);
[0084] Step 4.6.1: Input the semantic sets Token_new1 and Token_new2 from Step 4.5 into the Transformer decoder respectively. Here, the network is also divided into two branches. The main branch first passes through Layer Normalization (Layer Norm), and then through Multi-Head Cross-Attention (MCA). At the same time, a residual branch is introduced to be fused with the main branch; The Multi-Head Cross-Attention (MCA) also contains three vectors Q, K, and V. Different from the Multi-Head Self-Attention (MSA), the three vectors Q, K, and V of the Multi-Head Cross-Attention (MCA) come from different objects. Among them, the Q vector comes from the feature X i obtained by fusing through the Feature Pyramid in Step 3.1, while the K and V vectors come from the features Token_new encoded by the Transformer encoder module obtained by splitting in Step 4.5 i ;
[0085] Step 4.6.2: The data from the above Step 4.6.1 also passes through two branches respectively. The main branch sequentially passes through Layer Normalization (Layer Norm) and Multilayer Perceptron (MLP), and at the same time, a residual branch is introduced to be added and fused with the main branch;
[0086] Step 5: Input the feature map containing rich information extracted into the channel attention module. The channel attention module can calculate the importance of each channel, give different attentions to each channel, and thus obtain the enhanced feature map;
[0087] Step 6: Send the enhanced feature maps before and after the change to the difference information extraction module to obtain the difference feature map of the two remote sensing images. The difference information extraction module is as Figure 7 shown, including a two-dimensional convolutional layer, an activation function RELU, and batch normalization;
[0088] Step 7: Send the difference feature map into the classifier to obtain the final change detection binary image, and compare it with the actual change reference map for display output. The results on the LEVIR-CD test set are as Figure 8 shown, where (a) is the image before change, (b) is the image after change, (c) is the actual change reference image, (d) is the obtained detection result image. In figures (c) and (d), white represents the changed area, and the black part represents the unchanged area; the results on the DSIFN-CD test set are as Figure 9 shown. Similarly, where (a) is the image before change, (b) is the image after change, (c) is the change reference image, (d) is the obtained detection result image. In figures (c) and (d), white represents the changed area, and the black part represents the unchanged area.
[0089] The present invention combines the lightweight Resnet18 network structure with the VisionTransformer model based on the multi-head attention mechanism. First, the Resnet18 network is used to preliminarily extract the feature information of the images before and after the change. Then, the feature pyramid structure is used to fuse the detailed information contained in the low level and the semantics contained in the high level. Furthermore, the encoder and decoder modules in the Vision Transformer model are used to further model and capture the context information of the fused features, obtaining enhanced feature information. Through such a two-step incremental feature extraction network, it is possible to better capture the rich information contained in the image. In addition, the present invention introduces a difference information extraction module. By constructing a network structure, the model can learn the difference information between the two feature maps before and after the change through training by itself, rather than specifying two feature maps for subtraction, reducing manual intervention. Through the above method, the present invention effectively improves the problems of low detection accuracy and large false detection and missed detection probabilities of the existing change detection methods on optical remote sensing images with high resolution. Moreover, the present invention conducts the training and testing processes on two different datasets, LEVIR-CD and DSIFN-CD, respectively, verifying the universality of the proposed method on different datasets.
[0090] As described above, it is only the preferred embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be thought of by those skilled in the art within the technical scope disclosed by the present invention without creative labor should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope defined by the claims.
Claims
1. A method for optical remote sensing image change detection based on deep learning, comprising the following steps: Step 1: Obtain multi-temporal optical remote sensing image data and preprocess it; Step 2: Input the image image1 before change and the image image2 after change read into the Resnet18 network respectively, and extract the feature maps of different levels from stage1 to stage5 of each image; Step 3: For the two images image1 and image2 before and after change, input the feature maps of different levels from stage1 to stage5 into the feature pyramid structure respectively, fuse the detailed information contained in the low level and the semantics contained in the high level, and obtain the fused feature maps; Step 4: Input the fused feature maps into the Vision Transformer model. First, use the token extractor to transform the feature maps into the semantic token space, then use the encoder module to perform context modeling and feature extraction on the tokens, and finally use the decoder module to map the features from the token space to the pixel space, thereby obtaining feature maps containing rich information; Step 4.1: For the features of the two images before and after change, use the Token extractor to extract semantic tokens from the fused features respectively, transform the features into a space represented based on semantic tokens, and obtain semantic sets Token1 and Token2 respectively; Step 4.2: Concatenate the semantic sets Token1 and Token2 corresponding to the two images before change respectively to obtain the concatenated Tokens; Step 4.3: For the concatenated Tokens, perform trainable position embedding Position Embedding. The specific operation is as follows: add a parameter-trainable vector PE on the basis of the tokens, and this vector has the same size as the Tokens; position embedding can perform relative position encoding on the elements in the semantic token space, so it is beneficial to perform context modeling and extraction on the features; Step 4.4: Send the Tokens with position encoding information to the encoder to perform context information modeling in the semantic token-based space, and obtain a semantic set Token_new containing rich information. The encoder module is stacked by Ne basic modules, and each module contains a multi-head self-attention module and a multi-layer perceptron; Step 4.5: Split Token_new to obtain the semantic sets Token_new1 and Token_new2 corresponding to before and after change respectively; Step 4.6: For the semantic sets Token_new1 and Token_new2, use the decoder to project them back into the pixel space to obtain the enhanced feature maps corresponding to before and after change respectively; The decoder module is stacked by Nd basic modules, and each module contains a multi-head cross-attention module and a multi-layer perceptron; Step 5: Input the feature map containing rich information extracted into the channel attention module. The channel attention module can calculate the importance of each channel, give different attentions to each channel, and thus obtain the enhanced feature map; Step 6: Feed the enhanced feature maps before and after the change into the differential feature extraction module to obtain the difference feature map of the two remote sensing images; Step 7: Feed the difference feature map into the classifier to obtain the final change detection binary image, and compare it with the change reference map.
2. The method for optical remote sensing image change detection based on deep learning according to claim 1, wherein, The specific steps of step 1 are as follows: Step 1.1: After obtaining the multi-temporal optical remote sensing image data, segment it into non-overlapping image patches with a resolution of 256*256, and divide it into three data sets: training set, validation set, and test set according to the original ratio of the data bureau; Step 1.2: Read the segmented remote sensing images; Step 1.3: Perform data augmentation operations on the read training set, including random flipping in the horizontal direction, random flipping in the vertical direction, random cropping of the image, and Gaussian blurring of the image. The validation set and test set do not require data augmentation operations. The purpose of data augmentation is to increase the complexity during the training process and enhance the robustness of the model.
3. The method for optical remote sensing image change detection based on deep learning according to claim 1, wherein, The specific steps of step 4.1 are as follows: Step 4.1.1: For the input feature map X i , perform two-dimensional convolution on it to obtain L feature maps, and perform softmax operation to obtain the feature map A with spatial attention information i , and the calculation formula is as follows: A i = softmax(Conv(X i )) Step 4.1.2: Use the feature map A with spatial attention information i Perform pixel-level weighted calculation on the original input feature map X i To obtain a semantic token set T of length L i , and the calculation formula is as follows:
4. The method for optical remote sensing image change detection based on deep learning according to claim 1, wherein, The specific steps of step 4.4 are as follows: Step 4.4.1: The concatenated tokens with position encoding information are input into the encoder. Here, the network is divided into two branches. The main branch first passes through Layer Norm, and then through the multi-head self-attention module. At the same time, a residual branch is introduced to be fused with the main branch; the calculation formula of the multi-head attention module is as follows: MHAtten(Q, K, V) = Concat(h1, …, h i , …, h n )W O where h i is the single-head attention mechanism, and its calculation formula is: Among them, Q, K, and V represent the Query query vector, Key key vector, and Value value vector respectively, and W i Q , represent the mapping matrices corresponding to the three vectors Q, K, and V in each head respectively, and W O represents the output mapping matrix, and d k represents the vector dimension size. It can be seen from the formula that the single-head attention mechanism calculates how much attention should be given to the V vector through the Q vector and the K vector, and multiplies it by the V vector, so that the V vector with attention information can be obtained; while the multi-head attention mechanism simultaneously executes multiple single-head attentions in parallel, concatenates their results, and finally performs output mapping; for the multi-head self-attention module, its three vectors Q, K, and V all come from the same object, that is, the tokens with position encoding information obtained in step 4.3; Step 4.4.2: The data from the above step 4.4.1 also passes through two branches respectively. The main branch passes through Layer Norm and the multi-layer perceptron in sequence. At the same time, a residual branch is introduced to be added and fused with the main branch.
5. The method for optical remote sensing image change detection based on deep learning according to claim 1, wherein,The specific steps of step 4.6 are as follows: Step 4.6.1: Input the semantic sets Token_new1 and Token_new2 from Step 4.5 into the decoder respectively. Here, the network is also divided into two branches. The main branch first passes through layer normalization (Layer Norm), and then through the multi-head cross-attention module. Meanwhile, a residual branch is introduced to fuse with the main branch. The multi-head cross-attention (MCA) also contains three vectors, Q, K, and V. Different from the multi-head self-attention (MSA), the three vectors Q, K, and V of the multi-head cross-attention (MCA) come from different objects. Among them, the Q vector comes from the feature X obtained by fusing through the feature pyramid in Step 3 i , while the K and V vectors come from the features Token_new encoded by the encoder module segmented in Step 4.5 i ; Step 4.6.2: The data from the above step 4.6.1 also passes through two branches respectively. The main branch passes through Layer Norm and the multi-layer perceptron in sequence. At the same time, a residual branch is introduced to be added and fused with the main branch.
Citation Information
Cited By
Remote sensing image detection method based on dual-time-phase interaction enhancement CNN-Transform
CN118135392A
Remote sensing image detection method based on dual temporal interactive enhancement CNN-Transformer
CN118135392B