A Remote Sensing Image Change Detection Method Based on Difference Enhancement and Attention Module

By introducing image difference enhancement module and attention module in UNet3+ network, the problem of the failure of the existing technology to fully utilize remote sensing image difference information is solved, and higher change detection accuracy and efficiency are achieved.

CN116824359BActive Publication Date: 2025-06-27DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310483693.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-04
Publication Date
2025-06-27
Estimated Expiration
2043-05-04

AI Technical Summary

Technical Problem

The prior art fails to fully utilize image pair differential information in remote sensing image change detection, resulting in detection error detection and missed detection, and the model structure cannot be effectively improved to adapt to the characteristics of remote sensing images.

Method used

Based on the UNet3+ network, an image difference enhancement module is built, and differential features are generated through additional fill convolution (APC) units, and an attention module is introduced into the encoder and decoder to optimize the feature extraction and decoding process and reduce information loss.

Benefits of technology

By making full use of the difference information of the image pair and the feature extraction optimized by the attention module, the accuracy and efficiency of remote sensing image change detection are significantly improved, and missed detection and missed detection are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116824359B_ABST
    Figure CN116824359B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of remote sensing image processing, and provides a remote sensing image change detection method based on difference enhancement and attention module. The steps are as follows: (1) Preprocess the dual-temporal remote sensing images; (2) Comprise an image difference enhancement module by APC units with shared weights and generate image difference features at all levels; (3) The encoder receives the auxiliary features from the image difference enhancement module and generates encoder features at all levels; (4) The decoder decodes the features level by level; (5) The last-layer decoder features generate the final change detection result through 3×3 convolution. The present invention can make full use of the difference information of remote sensing images, mine the deep features in the features, effectively improve the change detection accuracy, and has broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image processing, and relates to a remote sensing image change detection method based on difference enhancement and attention module. Background Technique

[0002] The task of remote sensing image change detection refers to analyzing two or more remote sensing data of the same location to detect changes on the earth's surface over time. It has now been widely applied in fields such as urban planning, agricultural and forestry monitoring, marine and inland water body monitoring, and natural disaster monitoring, and is of great significance for environmental protection, sustainable development, and natural disaster resolution. With the progress of remote sensing technology in the field of earth observation, there are now many sensors that can provide ultra-high-resolution remote sensing images. Due to the improvement of their spatial resolution, these images can provide a large amount of details and shape features of ground objects, enabling us to perform more refined ground object recognition. Although high-resolution remote sensing images bring more application possibilities for change detection, they also bring a series of new technical problems and challenges, such as the need for higher storage capacity and processing power to handle large amounts of data, atmospheric interference may lead to data instability, large differences in ground objects may cause change detection errors and missed detections, and color differences may cause false changes.

[0003] The fully convolutional network is the first end-to-end network model for pixel-level prediction. It replaces the last layer in the convolutional neural network with a convolutional layer, avoiding the problems of repeated storage and convolution calculation caused by using pixel blocks, and becoming the basic framework for semantic segmentation. Therefore, many studies at home and abroad have designed various fully convolutional networks to complete the above-mentioned remote sensing image change detection tasks. Daudt et al. proposed three fully convolutional (FC) network architectures in the paper "Daudt R, Bertrand LS, Alexandre B. Fully Convolutional Siamese Networks for Change Detection[C] / / IEEE International Conference on Image Processing, ICIP 2018, Athens, Greece, October 7-10, 2018: 4063–4067.", namely the fully convolutional network with early fusion (FC-EF), the fully convolutional network with siamese encoders and feature fusion by concatenation (FC-Siam-Conc), and the fully convolutional network with siamese encoders and feature fusion by absolute difference (FC-Siam-Diff). Further, Daudt et al. tried to combine residual blocks with FC-EF in the paper "Daudt R, Bertrand LS, Alexandre B, et al. Multitask learning for large-scale semantic change detection[J]. Computer Vision and Image Understanding, 2019, 187: 102783." and proposed FC-EF-Res. Zhou et al. proposed the Unet++ network in the paper "ZHOU Z, SIDDIQUEE M, TAJBAKHSH N, et al. UNet++: Redesigning skip connections to exploit multiscale features in image segmentation[J]. IEEE Transactions on Medical Imaging, 2020, 39(6): 1856–1867". This model is a densely connected network that improves skip connections to integrate features at different depths and can be regarded as an effective collection of UNets with different depths.Fang et al. proposed the SNUNet-CD network in the paper "FANG S, LI K, SHAO J, et al. SNUNet-CD: A densely connected Siamese network for change detection of VHR images[J]. IEEE Geoscience and Remote Sensing Letters, 2021, Art no. 8007805." This model is an effective combination of Unet++, Siamese network and attention module. However, the above method does not fully utilize the difference information of the bi-temporal image pair, nor does it improve the model structure according to the characteristics of remote sensing images themselves. As we all know, remote sensing images contain extremely rich information, such as texture information, color information, location information, etc. If the model feature extraction method can be improved, the deep features of remote sensing images can be mined, and the difference information of the image pair can be fully utilized, which can not only improve the detection accuracy but also improve the detection efficiency. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a remote sensing image change detection method based on difference enhancement and attention module for the deficiency that the above-mentioned existing technology pays more attention to the local detail information of remote sensing images and does not fully utilize the difference information of image pairs for change detection. Based on the UNet3+ network, the present invention constructs an image difference enhancement module to provide auxiliary features for the encoder, and adds different attention modules in the encoder and decoder to refine the features, considers the internal relationship between the changing objects, reduces the phenomena of missed detection and false detection, and realizes the improvement of detection accuracy.

[0005] In order to improve the change detection accuracy and efficiency, the technical solution adopted by the present invention is as follows:

[0006] A remote sensing image change detection method based on difference enhancement and attention module ( Figure 2 ), comprising the following steps:

[0007] Step 1, input the preprocessed bi-temporal images into the additional padding convolution (APC) units with shared weights respectively, and obtain two first-layer difference features ( Figure 2 in and ). After splicing the bi-temporal images, input them into the encoder, and obtain the first-layer encoder features through convolution operations (Figure 3(b) is a schematic diagram of the encoder convolution operation) ( Figure 2 in ). Where C, H, and W respectively represent the number of channels, height, and width of the image, and m and n in APC(m,n) and conv(m,n) represent the number of channels of the features before and after convolution.

[0008] Step 2: Perform M×M max pooling on both and , where M is an even number, preferably 2, and then input them into the APC units with shared weights respectively to obtain two second-layer difference features ( Figure 2 in and ). Concatenate the absolute value of the difference between and with F EN1 , and then perform M×M max pooling, convolution operation, and convolutional block attention module (CBAM) feature optimization in sequence to obtain the second-layer encoder feature ( Figure 2 F in EN2 ).

[0009] Step 3: Continuously repeat the operations similar to those in Step 2 to achieve feature extraction. Specifically, perform M×M max pooling on both and , and then input them into the APC units with shared weights respectively to obtain two (i + 1)-layer difference features ( Figure 2 in and ). Concatenate the absolute value of the difference between and with F ENi , and then perform M×M max pooling, convolution operation, and CBAM feature optimization in sequence to obtain the (i + 1)-layer encoder feature ( Figure 2 F in ENi+1 ). The F EN3 and F EN4 are generated successively through the above process, that is, i in the above process is equal to 2 and 3 respectively.

[0010] Step 4: Concatenate the absolute value of the difference between and with F EN4 , and then perform M×M max pooling, convolution operation, and CBAM feature optimization in sequence to obtain the fifth-layer encoder feature ( Figure 2 F in EN5 ). So far, the feature extraction task of the encoder has been completed. Since there are 4 M×M max poolings in the encoder, the height and width of F EN5 should be

[0011] Step 5: Use the fifth-layer encoder feature F EN5Decode is performed. The specific process is to perform full-scale skip connection (FsSC), convolution operation (Figure 3(c) is a schematic diagram of the decoder convolution operation), and channel optimization of the improved efficient channel attention module (imECA) in sequence to obtain the fourth-layer decoder feature F. DE4 The multi-level features of the encoder are aggregated through the full-scale skip connection, reducing information loss. Subsequently, the convolution operation extracts variable information from the multi-level features, and finally the imECA dynamically adjusts the weights of each channel to efficiently mine deep features.

[0012] Step 6: Repeat the operation similar to Step 5 for feature decoding. Specifically, perform FsSC, convolution operation, and imECA channel optimization on the j-th layer decoder feature F DEj in sequence to obtain the (j - 1)-th layer decoder feature F DEj-1 . That is, F DE4 generates F DE3 , then F DE3 generates F DE2 , and finally F DE2 generates F DE1 . Therefore, in the above process, j is equal to 4, 3, and 2 in sequence.

[0013] Step 7: Perform a conventional 3×3 convolution on the first-layer decoder feature F DE1 to obtain a binary change detection map.

[0014] The present invention first generates a difference feature through the APC unit, and then the encoder receives the difference feature for feature extraction. The height and width of the obtained encoder feature continuously decrease. After the encoder performs 4 times of feature extraction, the decoder decodes and restores the size of the feature, and finally generates a change detection map with the same height and width as the input.

[0015] Furthermore, the preprocessing process of the bi-temporal images in Step 1 is as follows:

[0016] (1) Random flipping: Randomly flip the image with a probability of 50% to achieve data augmentation and prevent the model from overfitting.

[0017] (2) Random rotation: Rotate the images in the dataset 0 to 2 times by 90° to achieve data augmentation and improve the generalization ability of the model.

[0018] In steps 1-3 described above, the structure of the APC unit is shown in Fig. 3(a). Assuming that C1 and C2 represent the number of channels of the input features and output features respectively, the specific process of the APC unit is as follows: First, use a 1×1 convolution to modify the number of channels of the input features, obtaining an intermediate feature with a channel number three times that of the output features, and then divide it into three parts. Padding is added to two of the feature blocks in the horizontal and vertical directions respectively, specifically by setting the values from the start to 1 / 8 of the corresponding direction of the feature block to zero. Then, 3×3 convolutions are performed on all three feature blocks to achieve feature extraction. Finally, the three feature blocks are added together, and after batch normalization (BN) and ReLU non-linear activation, the final output is obtained. The APC unit used in the present invention is a feature extraction unit more applicable to remote sensing images. It can mine more effective information, highlight the difference information of the image pair, provide auxiliary features for the encoder, and achieve more precise change detection.

[0019] The mathematical description of CBAM in step 2 is shown in formula (1):

[0020]

[0021]

[0022] Where, In represents the input, Out′ represents the intermediate output, Out″ represents the final output, E c (·) represents channel attention, E s (·) represents spatial attention. The mathematical descriptions of channel attention and spatial attention are shown in formulas (2) and (3) respectively.

[0023] E c (In) = σ(MLP(AvgPool(In)) + MLP(MaxPool(In)))

[0024] = σ(W1(W0(In avg )) + W1(W0(In max ))) (2)

[0025] E s (Out′) = σ(f 7×7 ([AvgPool(Out′); MaxPool(Out′)]))

[0026] = σ(f 7×7 ([Out′ avg ; Out′ max )) (3)

[0027] Where, MLP represents a multi-layer perceptron, W0 and W1 are the weight matrices of the MLP, and σ represents the sigmoid function.

[0028] The schematic diagram of the convolution operation of the encoder in steps 1-4 is shown in Fig. 3(b). The input features are successively subjected to 3×3 convolution, BN layer, ReLU non-linear activation, 3×3 convolution, BN layer and ReLU non-linear activation, and the output features are obtained through the above operations.

[0029] Furthermore, the mathematical description of the full-scale skip connection in step 5 is shown in formula (4):

[0030]

[0031] where, is the output of the i-th level of the decoder, is the feature map of the same scale from the encoder. H(·) is followed by convolution, BN layer and ReLU non-linear activation. C(·) represents the convolution operation. D(·) and U(·) respectively represent the downsampling and upsampling operations. [·] represents the concatenation operation. N represents the number of layers included in the encoder. In the present invention, N = 5. Scales represents the scale range of this operation.

[0032] Furthermore, the structure of the improved efficient channel attention module imECA in step 5 is as Figure 4 shown. It is improved from the efficient channel attention network ECA. ECA is an attention module. This module first undergoes global average pooling, so that the feature block becomes a vector of C×1×1 from a matrix of C×H×W. Subsequently, a one-dimensional convolution with a convolution kernel size of k and sigmoid processing are performed on this vector to obtain the weight of each channel in the feature block. Finally, the weight is multiplied element-wise with the original input feature map to obtain the output result. The mathematical expression of ECA when k = 7 can be summarized as shown in formula (5)

[0033]

[0034] where, In and Out are respectively the input and output of ECA. GAP represents global average pooling, and f1 i×i represents a one-dimensional convolution operation with a convolution kernel of i×i.

[0035] The imECA proposed by the present invention adds multi-scale fusion of features on the basis of the original ECA. Its specific process is as follows: First, global average pooling is performed on the input features, then one-dimensional convolution operations of 5×5, 7×7 and 9×9 are respectively performed on them. Subsequently, the obtained multi-scale features are concatenated and a 1×1 convolution operation is performed. Finally, sigmoid processing is performed to obtain the channel weight matrix. The mathematical expression of imECA is shown in formula (6).

[0036]

[0037] where In and Out are the input features and output features of imECA respectively, GAP represents global average pooling, and f1 i×i represents a one-dimensional convolution operation with a convolution kernel of i×i, and f 1×1 represents a conventional 1×1 convolution. Through imECA, more feature information can be mined to optimize the weights of each channel.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] (1) From the perspectives of the practicability and operability of the model, the present invention constructs an efficient remote sensing image change detection network. This network can not only obtain local information through the encoder but also obtain the difference information of the image pair through the APC unit, making full use of the local information and difference information to perform change detection on the image, overcoming the problems of misdetection and missed detection of change objects caused by the existing fully convolutional neural network not using difference information, and making the present invention have the advantage of high detection accuracy.

[0040] (2) The present invention makes full use of the image features, emphasizes the difference features of the image pair, and assists the encoder in mining deep features, making the present invention have excellent visual effects for change detection.

[0041] (3) The detection network constructed by the present invention uses the decoded features rather than the encoded features to calculate the change information, and first uses FsSC to aggregate the full-scale feature information during the decoding process, and then uses imECA to dynamically adjust the weights of each channel, reducing the information loss during the decoding process, making the present invention more accurate in identifying change objects. Description of the Drawings

[0042] Figure 1 is a flowchart of the remote sensing image change detection method based on the difference enhancement and attention module;

[0043] Figure 2 is a structural diagram of the remote sensing image change detection method based on the difference enhancement and attention module;

[0044] Figure 3(a) is a structural diagram of the APC unit;

[0045] Figure 3(b) is a schematic diagram of the convolutional operation of the encoder;

[0046] Figure 3(c) is a schematic diagram of the convolutional operation of the decoder;

[0047] Figure 4 is a structural diagram of imECA;

[0048] Figure 5 is the change detection result of the present invention and the prior art. Detailed Embodiments

[0049] To make the method problems solved by the present invention, the adopted method solutions, and the achieved method effects clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the convenience of description, only the parts related to the present invention are shown in the accompanying drawings, rather than all the content.

[0050] Optical remote sensing images are remote sensing images collected by optical sensors. Their characteristic is that they can provide high-resolution surface information. However, a large amount of information also brings many challenges, such as data quality, occlusion problems, illumination and noise problems, the complexity of object changes, etc. Detecting changes in complex optical remote sensing images is still a challenging task. The present invention aims at the problem that most methods do not make full use of the difference information of dual-temporal image pairs, and improves the model structure according to the characteristics of remote sensing images themselves, and proposes an effective change detection method. The present invention can be applied to change detection tasks in many scenarios. For example, after obtaining remote sensing images through satellites or drones, using appropriate data to train the model can meet the actual needs such as vegetation change detection, building change detection, and coastline change detection.

[0051] 1. Data and operating environment

[0052] In this example, the CDD dataset is used for simulation experiments. The hardware platform for the simulation experiments of the present invention is: the CPU is Intel(R) Core(TM) i7-10700, the main frequency is 2.90 GHz, the memory is 32 GB, the GPU is NVIDIA GeForce RTX 2060, and the memory is 6G. The software platform for the simulation experiments of the present invention is: Python 3.8.

[0053] 2. Implementation steps

[0054] Step 1: Dataset preparation.

[0055] In this example, the publicly available CDD change detection data is used, which contains a total of 16,000 image pairs. The size of all images is 256×256 pixels, and the spatial resolution ranges from 0.03 to 1 meter. The collection locations of this dataset include roads, snowfields, housing buildings, vegetation, rivers, etc., and include many images with large-scale changes. In this example, all the images of the CDD dataset are used for model training, and they are divided into a training set, a validation set, and a test set according to the ratio of 10:3:3.

[0056] Step 2: Data preprocessing.

[0057] First, randomly flip the dual-temporal images with a probability of 50%, and then randomly rotate them 0 to 2 times by 90° to achieve data augmentation and prevent model overfitting.

[0058] Step 3: Construct the loss function.

[0059] The loss function is a combination of the WCE loss and the IoU loss. The WCE loss aims to solve the problem of class imbalance, and its expression is shown in Equation (7).

[0060]

[0061] Among them, c is the number of label categories, w c represents the weight of class c, g(i) represents the label value, and p(i, c) represents the predicted probability that the pixel point i belongs to class c.

[0062] The IoU loss is a common loss function used in the fields of image segmentation and object detection, and its expression is shown in Equation (8).

[0063]

[0064] Finally, the loss function of the model is: L(i) = L wce (i) + L IoU (i). (9)

[0065] Step 4: Parameter setting.

[0066] The number of channels of the four-layer features generated by the image difference enhancement module in the network are 8, 16, 32, and 64 respectively. The number of channels of the five-layer features extracted by the encoder are 16, 32, 64, 128, and 256 respectively. The number of channels of each layer of features in the decoder is 80. The epoch is set to 50 in the network training process, the initial learning rate lr = 0.001, and the Adam optimizer is used to calculate the gradient, with its parameters betas = (0.9, 0.999). The Batchsize for training, validation, and testing processes is all 4.

[0067] Step 5: Network training, validation, and testing.

[0068] Step 5.1, input the training set into the network in batches, calculate the loss between the binary change detection result output by the network and the true label, and use the Adam optimization algorithm to iteratively update the network parameters until all the training sets have participated in the training. Input the validation set into the current trained model, calculate 4 evaluation metrics, and save the current model.

[0069] Step 5.2, based on the previous trained model, repeat Step 5.1 until the training of all epochs is completed.

[0070] Step 5.3: Observe the performance of the trained model obtained in each epoch on the validation set, and select the trained model obtained in the optimal epoch for testing.

[0071] Step 5.4: Input the test set into the optimal trained model in Step 5.3 and calculate the evaluation metrics.

[0072] Among them, the four evaluation metrics are Overall Accuracy (OA), precision (p), recall (r), and F1-score.

[0073] Overall Accuracy measures the closeness between the predicted value and the true value, that is, the proportion of correctly predicted (un)changed categories, and the expression is as follows.

[0074]

[0075] Among them, TP, TN, FP, and FN are true positive, true negative, false positive, and false negative examples respectively.

[0076] Precision measures the proportion of correctly predicted changed pixels to all pixels predicted as changed, and the expression is as follows.

[0077]

[0078] Recall measures the proportion of correctly predicted changed pixels to the truly changed pixels, and the expression is as follows.

[0079]

[0080] The F1-score can be regarded as a harmonic mean of the model's precision and recall, and the expression is as follows.

[0081]

[0082] Step 6: Model comparison and analysis.

[0083] Six existing technologies are used for comparative analysis with the present invention. The performance of all technologies on the CDD test set is shown in Table 1. Among them, "the present invention" refers to the proposed remote sensing image change detection method based on difference enhancement and attention module, "FC-EF", "FC-Siam-Conc", "FC-Siam-Diff", and "FC-EF-Res" refer to four fully convolutional change detection networks proposed by Daudt et al., "Unet++" refers to the dense connection network proposed by ZHOU et al., and "SNUNet-CD" refers to the efficient network proposed by FANG et al.

[0084] Table 1 Performance evaluation table of the present invention and existing remote sensing image change detection models

[0085] Method OA(%) p(%) r(%) F1(%) FC-EF 91.9 67.0 61.9 64.3 FC-Siam-Conc 92.3 67.7 66.1 66.9 FC-Siam-Diff 92.3 70.9 58.3 64.0 FC-EF-Res 96.4 85.6 83.3 84.5 Unet++ 95.5 81.9 79.0 80.4 SNUNet-CD 96.8 88.1 84.1 86.0 The present invention 97.2 87.4 89.1 88.3

[0086] As can be seen from Table 1, the present invention has obtained satisfactory results in all four evaluation indicators, and has a relatively high improvement in the comprehensive indicators of "OA" and "F1 score" compared with other technologies, proving that the present invention can achieve higher change detection accuracy. In addition, Figure 5 The visualization results of the present invention and the prior art on one of the test samples are shown. It can be seen that the present invention can obtain a cleaner visual effect with almost no noise points, and at the same time has a better detection effect on boundary pixels. In summary, it can be considered that higher change detection accuracy can be obtained through the present invention, more accurate change detection maps can be generated, and practical tasks such as coastline change detection and building change detection can be completed.

[0087] Finally, it should be noted that the above examples are only used to illustrate the implementation mode of the present invention. It should be understood that the examples are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can be made. Modifications to various equivalent forms of the present invention all fall within the scope defined by the appended claims of this application, and all belong to the protection scope of the present invention.

Claims

1. A remote sensing image change detection method based on difference enhancement and attention module, characterized in that Including the following steps: Step 1, input the preprocessed dual-temporal images into the weight-sharing additional padding convolutional APC units respectively, to obtain two first-layer difference features and Concatenate the dual-temporal images and input them into the encoder, and obtain the first-layer encoder features through convolutional operations Step 2, perform max pooling on both and , and then input them into the APC units with shared weights respectively to obtain two second-layer differential features and Concatenate the absolute value of the difference between and with F EN1 . Subsequently, perform max pooling, convolution operation, and CBAM feature optimization of the convolutional block in sequence to obtain the second-layer encoder feature F EN2 ; Step 3, continuously repeat the operation similar to Step 2 to achieve feature extraction Max-pool both and , then input them into the APC units with shared weights respectively to obtain two difference features of the (i + 1)-th layer and Take the absolute value of the difference between and , concatenate it with F ENi , then perform max-pooling, convolution operation, and CBAM feature optimization in sequence to obtain the encoder feature F ENi+1 of the (i + 1)-th layer. The above process is performed with i equal to 2 and 3 respectively; generate F EN3 and F EN4 in sequence; Step 4: and The absolute value of the difference and F EN4 Splicing, followed by maximum pooling, convolution operation, and CBAM feature optimization, to obtain the 5th layer encoder feature F EN5 ; At this point, the encoder's feature extraction task has been completed; Step 5, decode the encoder features F of the fifth layer EN5 ​ Perform full-scale skip connection FsSC, convolution operation, and improved efficient channel attention module imECA channel optimization in sequence to obtain the fourth-layer decoder feature F DE4 ; Step 6, repeat the operation in Step 5 for feature decoding The decoder feature F of the j-th layer DEj is successively subjected to FsSC, convolution operation, and imECA channel optimization to obtain the decoder feature F of the (j - 1)-th layer DEj-1 ; that is, F DE4 is used to generate F DE3 , and then F DE3 is used to generate F DE2 , and finally F DE2 is used to generate F DE1 . In the above process, j is successively equal to 4, 3, and 2; Step 7, perform a conventional 3×3 convolution on the first-layer decoder feature F DE1 to obtain a binary change detection map 2. The remote sensing image change detection method based on difference enhancement and attention module according to claim 1, wherein The preprocessing process of the dual-temporal image in Step 1 is as follows: (1) Random flipping: Randomly flip the image with a probability of 50% to achieve data augmentation and prevent model overfitting; (2) Random rotation: Rotate the images in the dataset by 0 to 2 times of 90° to achieve data augmentation and improve the generalization ability of the model.

3. A remote sensing image change detection method based on difference enhancement and attention module according to claim 1, characterized in that, In Steps 1 to 3, assuming that C1 and C2 represent the number of channels of the input feature and the output feature respectively, the specific process of the APC unit is as follows: First, use a 1×1 convolution to modify the number of channels of the input feature to obtain an intermediate feature with a channel number three times that of the output feature. Then divide it into three parts, and perform additional padding on two feature blocks in the horizontal and vertical directions respectively. Then perform 3×3 convolution on all three feature blocks to achieve feature extraction. Finally, add the three feature blocks, and after batch normalization BN and ReLU non-linear activation, the final output is obtained.

4. A remote sensing image change detection method based on difference enhancement and attention module according to claim 1, characterized in that The mathematical description of CBAM in Step 2 is shown in Formula (1): Among them, In represents the input, Out′ represents the intermediate output, Out″ represents the final output, and E c (·) represents channel attention, and E s (·) represents spatial attention; the mathematical descriptions of channel attention and spatial attention are shown in formulas (2) and (3) respectively; E c (In) = σ(MLP(AvgPool(In)) + MLP(MaxPool(In))) E s (Out′) = σ(f 7×7 ([AvgPool(Out′); MaxPool(Out′)])) Among them, MLP represents a multi-layer perceptron, W0 and W1 are the weight matrices of the MLP, and σ represents the sigmoid function.

5. A remote sensing image change detection method based on difference enhancement and attention module according to claim 1, characterized in that The convolution operation of the encoder in Steps 1 to 4: Perform 3×3 convolution, BN layer, ReLU non-linear activation, 3×3 convolution, BN layer, and ReLU non-linear activation on the input feature in sequence, and obtain the output feature through the above operations.

6. A remote sensing image change detection method based on difference enhancement and attention module according to claim 1, characterized in that The full-scale skip connection in Step 5 is shown in Formula (4): Among them, is the output of the i-th stage of the decoder, is the feature map of the same scale from the encoder. H(·) is followed by a convolutional layer, a BN layer, and a ReLU non-linear activation. C(·) represents the convolutional operation. D(·) and U(·) represent the downsampling and upsampling operations respectively. [·] represents the concatenation operation. N represents the number of layers in the encoder. Scales represents the scale range of this operation.

7. A remote sensing image change detection method based on difference enhancement and attention module according to claim 1, characterized in that The improved efficient channel attention module imECA in Step 5 is to add multi-scale fusion of features on the basis of the original efficient channel attention network ECA. Its specific process is as follows: First, perform global average pooling on the input feature, then perform one-dimensional convolution operations of 5×5, 7×7, and 9×9 on it respectively. Subsequently, splice the obtained multi-scale features and perform 1×1 convolution operation, and finally perform sigmoid processing to obtain the channel weight matrix; the mathematical expression of imECA is shown in Formula (5); where In and Out are the input features and output features of imECA respectively, GAP represents global average pooling, and f1 i×i represents a one-dimensional convolution operation with a convolution kernel of i×i, and f 1×1 represents a conventional 1×1 convolution; more feature information can be mined through imECA to optimize the weights of each channel.

Citation Information

Patent Citations

  • Unsupervised change detection method and system for homologous or heterologous remote sensing image

    CN113901900A

  • Remote sensing image change detection method and device based on multi-scale CNN-Transform

    CN115861703A