Contrastive learning based remote sensing image anti-cloud and fog interference change detection method

By constructing a sample set containing remote sensing images and foggy images, a feature extractor is trained and global and local features are fused. Spatial and channel attention modules are used to improve the efficiency and accuracy of remote sensing image change detection, solving the problem of false changes caused by thin cloud interference and achieving fast and accurate change detection.

CN120014446BActive Publication Date: 2026-05-08BEIJING INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING INST OF TECH
Filing Date
2025-01-17
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing remote sensing image change detection methods struggle to achieve rapid and accurate change detection when faced with thin cloud cover interference, especially since the problem of spurious changes remains unresolved.

Method used

A sample set is constructed, including remote sensing images and foggy images, and a feature extractor is trained. A change detection network is built using a feature processing module, a spatial attention module, a converter module, and a channel attention module. Global and local features are fused to improve detection efficiency and accuracy.

Benefits of technology

It enables rapid and accurate detection of changes in remote sensing images with cloud interference, thus improving the robustness of remote sensing image change detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014446B_ABST
    Figure CN120014446B_ABST
Patent Text Reader

Abstract

The application discloses a remote sensing image anti-cloud and fog interference change detection method based on contrast learning, comprising the following steps: constructing a sample set; training a feature extractor by using the sample set to obtain a trained feature extractor; constructing a change detection network by using the trained feature extractor, a feature processing module, a spatial attention module, a converter module, a channel attention module, a connection module and a classifier; inputting two detection images into the change detection network, the two detection images being different in shooting time and same in shooting area, and the change detection network outputting a predicted change image. The change detection network constructed by the trained feature extractor fuses global features and local features of images, and uses the spatial attention module and the channel attention module to improve efficiency and accuracy, so that quick and accurate detection can be realized even in a remote sensing image change detection task with cloud interference, and the robustness of remote sensing image change detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image change detection, and more specifically, to a method for detecting changes in remote sensing images that are resistant to cloud and fog interference based on contrastive learning. Background Technology

[0002] Bi-temporal change detection is a crucial task in remote sensing, involving the comparison and identification of changes in registered remote sensing images of the same area across different temporal phases. It has wide applications in disaster assessment, urban planning, agricultural surveys, resource management, and environmental monitoring. Due to factors such as complex textures, seasonal variations, climate change, and evolving needs, pseudo-change is a long-standing and significant problem in change detection. In remote sensing images from different temporal phases, the phenomenon where unchanged areas are incorrectly identified as changed areas due to other differences is called pseudo-change. Typical causes of pseudo-change can be categorized as color changes, temporary objects, shadow shapes, and cloud interference. These pseudo-change problems lead to the incorrect detection of shadow and projection differences in unchanged areas as changed areas. Among these, pseudo-change caused by thin cloud interference is a particularly noteworthy issue. Existing remote sensing image change detection datasets generally do not incorporate the influence of thin cloud interference, resulting in most change detection methods based on these datasets failing to consider the impact of thin cloud interference.

[0003] Traditional variation methods based on handcrafted features can achieve good results in simple scenarios, but they usually perform poorly in complex scenarios.

[0004] Algorithms based on deep convolutional neural networks (CNNs) perform better. Deep convolutional neural networks are widely used in change detection to extract discriminative local features, including classic convolutional neural networks and their extended architectures, such as ResNet (residual network) and UNet (image segmentation network).

[0005] Compared to pure convolutional neural networks, algorithms based on the Transformer architecture have achieved impressive results in change detection tasks by effectively modeling global contextual information through an encoder-decoder architecture. However, they are limited by the inherent limitations of the Transformer itself, resulting in limited utilization of local features.

[0006] Algorithms that fuse CNNs and Transformers have been developed, such as BIT-CD (a change detection model), which has achieved good performance in change detection tasks. However, there is still much room for improvement in the performance of change detection networks based on this approach, especially in dealing with the problem of false changes caused by cloud interference.

[0007] Chinese patent application (application number: 201410441207.2, application date: 2014.09.01) discloses a method for detecting changes in remote sensing images. After inputting remote sensing images before and after the change, it is necessary to determine whether the remote sensing image contains fog. Foggy images are then defogged before identification. This method requires extensive processing and computation during image judgment and defogging, resulting in slow processing efficiency.

[0008] Therefore, there is an urgent need to provide a method that can resist cloud and fog interference and achieve accurate and rapid change detection. Summary of the Invention

[0009] In view of this, the present invention provides a method for detecting changes in remote sensing images that are resistant to cloud and fog interference based on contrastive learning.

[0010] This invention provides a method for detecting changes in remote sensing images against cloud and fog interference based on contrastive learning, comprising:

[0011] Construct a sample set, which includes remote sensing images and foggy images. The number of remote sensing images is at least two, and any two remote sensing images are different. Each remote sensing image corresponds one-to-one with a foggy image, and each foggy image is obtained from its corresponding remote sensing image.

[0012] The feature extractor is trained using the sample set to obtain the trained feature extractor;

[0013] Constructing a change detection network includes:

[0014] Constructing a sub-network includes: providing a trained feature extractor for receiving an input image, extracting 32×32×32, 32×64×64, and 32×128×128 feature maps from the input image, performing upsampling on the 32×32×32 feature map to obtain upsampled features, and performing downsampling on the 32×128×128 feature map to obtain downsampled features; connecting the input of a feature processing module to the output of the trained feature extractor to fuse the 32×64×64 feature map, the upsampled features, and the downsampled features to obtain a fused feature map; connecting the input of a spatial attention module to the output of the feature processing module, the input of a converter module to the output of the spatial attention module, and the input of a channel attention module to the output of the converter module to process the fused feature map sequentially through the spatial attention module, the converter module, and the channel attention module to obtain a final feature map;

[0015] Provide two of the aforementioned sub-networks;

[0016] The input of the connection module is connected to the output of the channel attention module of the two sub-networks respectively to obtain two final feature maps. The two final feature maps are then connected in pairs along the channel dimension to obtain the overall image.

[0017] The input of the classifier is connected to the output of the connection module to output a predicted change map based on the overall map.

[0018] Two detection images are provided, which are captured at different times but in the same area. The two detection images are used as input images to the change detection network. The input of the feature extractor trained in the change detection network corresponds one-to-one with the input images. The change detection network outputs the corresponding predicted change map.

[0019] Optionally, constructing the sample set includes:

[0020] Simulated cloud images were obtained by combining Berlin noise with fractal Brownian motion.

[0021] Provide at least two of the aforementioned remote sensing images;

[0022] Each of the remote sensing images is fused with the cloud simulation image to obtain the foggy image corresponding to the remote sensing image;

[0023] All of the remote sensing images and all of the foggy images constitute the sample set.

[0024] Optionally, the remote sensing image is fused with the cloud simulation image to obtain the foggy image corresponding to the remote sensing image, calculated in the following manner:

[0025] I(x)=J(x)t(x)+A(1-t(x))

[0026] Wherein, J(x) is the remote sensing image, t(x) is a parameter representing the light that does not scatter and reaches the camera, A is the cloud simulation image, and I(x) is the foggy image corresponding to the remote sensing image.

[0027] Optionally, the feature extractor is trained using the sample set to obtain the trained feature extractor, including:

[0028] Two of the aforementioned feature extractors are provided;

[0029] Two images are randomly selected from the sample set. The two selected images are either any two remote sensing images, or any two foggy images, or one of the two selected images is any remote sensing image and the other is any foggy image.

[0030] The two selected images are input into two feature extractors, each corresponding to one of the images. The feature extractors output feature vectors based on the input images, thus obtaining two feature vectors.

[0031] Determine whether there is a correspondence between the two selected images. If there is a correspondence between the two images, set the label value to 0. If there is no correspondence between the two images, set the label value to 1.

[0032] Calculate the loss function based on the two feature vectors and the label values;

[0033] The feature extractor is trained based on the loss function.

[0034] Optionally, the loss function is calculated based on the two feature vectors and the label value, in the following manner:

[0035] loss=(1-label)euclidean(f1,f2) 2 +label(1-euclidean(f1, f2)) 2

[0036] Where loss is the loss value, label is the label value, euclidean(f1,f2) is the normalized Euclidean distance between the two feature vectors, f1 is the feature vector corresponding to one of the selected images, and f2 is the feature vector corresponding to the other selected image.

[0037] Optionally, the fused feature map is processed sequentially through the spatial attention module, the converter module, and the channel attention module to obtain the final feature map, including:

[0038] The spatial attention module processes the fused feature map to obtain a variable feature map;

[0039] The converter module and the channel attention module process the variable feature map to obtain the final feature map.

[0040] Optionally, the spatial attention module processes the fused feature map to obtain the variable feature map, and calculates it in the following manner:

[0041]

[0042] in, The fused feature map is defined as follows: σ is the activation function, f is the convolution kernel operation, Maxpool is the max pooling operation, and Avgpool is the average pooling operation. For two-dimensional spatial attention, For element-wise multiplication, This is the feature map of the variable.

[0043] Optionally, the converter module and the channel attention module process the variable feature map to obtain the final feature map, calculated in the following manner:

[0044]

[0045] M c =MLP(M i )

[0046]

[0047] Maxpool represents the max pooling operation. The variable feature map is defined as T(), which is processed by the converter module. For element-wise addition, MLP is a multilayer perceptron processing method. For element-wise multiplication, This is the final feature map.

[0048] Compared with existing technologies, the remote sensing image anti-cloud and fog interference change detection method based on contrastive learning provided by this invention achieves at least the following beneficial effects:

[0049] This invention provides a remote sensing image change detection method based on contrastive learning, which is robust against cloud and fog interference. The method includes: constructing a sample set; training a feature extractor using the sample set to obtain a trained feature extractor; constructing a change detection network using the trained feature extractor, a feature processing module, a spatial attention module, a converter module, a channel attention module, a connection module, and a classifier; inputting two detection images (captured at different times but in the same area) into the change detection network, and having the network output a predicted change map. The feature extractor is trained using a sample set containing both remote sensing and foggy images, resulting in a trained feature extractor with good resistance to cloud interference. The change detection network constructed using the trained feature extractor fuses global and local features of the image, and employs spatial and channel attention modules to improve efficiency and accuracy. Even in remote sensing image change detection tasks with cloud interference, this method achieves fast and accurate detection, improving the robustness of remote sensing image change detection.

[0050] Of course, any product implementing this invention does not necessarily need to achieve all of the technical effects described above at the same time.

[0051] Other features and advantages of the invention will become clear from the following detailed description of exemplary embodiments of the invention with reference to the accompanying drawings. Attached Figure Description

[0052] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.

[0053] Figure 1 This is a flowchart illustrating a remote sensing image anti-cloud and fog interference change detection method based on contrastive learning provided by the present invention.

[0054] Figure 2 This is a schematic diagram of a change detection network.

[0055] Figure 3 This is another flowchart illustrating the remote sensing image anti-cloud and fog interference change detection method based on contrastive learning provided by the present invention.

[0056] Figure 4 This is a rendering of a cloud layer based on a remote sensing image.

[0057] Figure 5 This is a flowchart illustrating a comparative learning process.

[0058] Figure 6 This is a schematic diagram of a spatial attention module.

[0059] Figure 7 This is a schematic diagram of a channel attention module.

[0060] In the diagram: 1. Subnetwork; 2. Feature extractor after training; 3. Feature processing module; 4. Spatial attention module; 5. Transformer module; 6. Channel attention module; 7. Connection module; 8. Classifier. Detailed Implementation

[0061] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the invention.

[0062] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.

[0063] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0064] In all the examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0065] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0066] Example 1

[0067] Combination Figure 1 and Figure 2 , Figure 1 This is a flowchart illustrating a remote sensing image anti-cloud and fog interference change detection method based on contrastive learning provided by the present invention. Figure 2 This is a schematic diagram of a change detection network, illustrating a specific embodiment of the remote sensing image change detection method based on contrastive learning to resist cloud and fog interference provided by the present invention, including:

[0068] S101: Construct a sample set, which includes remote sensing images and foggy images. The number of remote sensing images is at least two, and no two remote sensing images are the same. There is a one-to-one correspondence between the remote sensing images and the foggy images. The foggy images are obtained from their corresponding remote sensing images.

[0069] S102: Train the feature extractor using the sample set to obtain the trained feature extractor;

[0070] S103: Construct the change detection network (ACCDNet), including:

[0071] S1031: Constructing a sub-network, including: providing a trained feature extractor for receiving an input image, extracting 32×32×32, 32×64×64, and 32×128×128 feature maps from the input image, performing upsampling on the 32×32×32 feature map to obtain upsampled features, and performing downsampling on the 32×128×128 feature map to obtain downsampled features; connecting the input of the feature processing module to the output of the trained feature extractor to fuse the 32×64×64 feature map, upsampled features, and downsampled features to obtain a fused feature map; connecting the input of the spatial attention module to the output of the feature processing module, the input of the converter module to the output of the spatial attention module, and the input of the channel attention module to the output of the converter module to process the fused feature map sequentially through the spatial attention module, the converter module, and the channel attention module to obtain the final feature map;

[0072] S1032: Provides two sub-networks;

[0073] S1033: Connect the input of the connection module to the output of the channel attention module of the two sub-networks respectively to obtain two final feature maps. Connect the two final feature maps in pairs along the channel dimension to obtain the overall map.

[0074] S1034: Connect the input of the classifier to the output of the connection module to output a predicted change map based on the overall map;

[0075] S104: Provide two detection images. The two detection images were captured at different times but in the same area. The two detection images are used as two input images to the change detection network. The input of the trained feature extractor in the change detection network corresponds one-to-one with the input images. The change detection network outputs the corresponding predicted change map.

[0076] It should be noted that in this embodiment, the feature extractor used is ResNet-50. ResNet-50 can perform multi-scale feature extraction on the input image. Of course, it is not limited to this; the feature extractor can be selected according to actual needs. (Refer to...) Figure 2 In this embodiment, the two detected images are denoted as I1 and I2, respectively. The trained feature extractor is a pre-trained ResNet-50, the spatial attention module is SAM, the transformer module is a Transformer (including an encoder and a decoder), the channel attention module is CAM, and the classifier is a Classifier. In step S1032, two sub-networks are provided to construct a Siamese network structure. The two sub-networks share weights, giving them the same weights.

[0077] It is understandable that for the input image I i (i∈{1,2})∈R 3×H×W We used a pre-trained ResNet-50 backbone to extract feature maps at three different scales, denoted as follows: and It is a 32×32×32 feature map. It is a 32×64×64 feature map. The feature map is 32×128×128. (For...) Upsampling yields upsampled features, for Downsampling yields downsampled features, which are then combined with upsampled features and downsampled features. By fusion To fuse feature maps. The input spatial attention module SAM is obtained For variable feature maps. The inputs are sequentially fed into the Transformer module and the Channel Attention module (CAM) to obtain... This is the final feature map. The final feature maps of the same scale from the two sub-networks are paired and concatenated along the channel dimension, and then fed into the classifier to obtain the predicted change map.

[0078] This embodiment provides a remote sensing image change detection method based on contrastive learning, which is robust against cloud and fog interference. The method includes: constructing a sample set; training a feature extractor using the sample set; constructing a change detection network using the trained feature extractor, a feature processing module, a spatial attention module, a converter module, a channel attention module, a connection module, and a classifier; inputting two detection images (captured at different times but in the same area) into the change detection network, and outputting a predicted change map. The feature extractor is trained using a sample set containing both remote sensing and foggy images, resulting in a trained feature extractor with good resistance to cloud interference. The change detection network constructed using the trained feature extractor fuses global and local features of the image, and uses spatial and channel attention modules to improve efficiency and accuracy. Even in remote sensing image change detection tasks with cloud interference, it can achieve fast and accurate detection, improving the robustness of remote sensing image change detection.

[0079] Example 2

[0080] Combination Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 7 , Figure 3 This is another flowchart illustrating the remote sensing image anti-cloud and fog interference change detection method based on contrastive learning provided by the present invention. Figure 4 It is a rendering of cloud formations based on remote sensing images. Figure 5 This is a flowchart illustrating a comparative learning process. Figure 6 This is a schematic diagram of a spatial attention module. Figure 7 This is a schematic diagram of a channel attention module, illustrating another specific embodiment of the remote sensing image anti-cloud and fog interference change detection method based on contrastive learning provided by the present invention, including:

[0081] S201: Construct a sample set, which includes remote sensing images and foggy images. The number of remote sensing images is at least two, and no two remote sensing images are the same. There is a one-to-one correspondence between the remote sensing images and the foggy images, and the foggy images are obtained from their corresponding remote sensing images.

[0082] S202: The feature extractor is trained using a sample set to obtain the trained feature extractor, including:

[0083] S2021: Provides two feature extractors;

[0084] S2022: Select any two images from the sample set. The two selected images are either any two remote sensing images, or any two foggy images, or one of the selected images is any remote sensing image and the other is any foggy image.

[0085] S2023: Input the two selected images into two feature extractors. Each feature extractor corresponds to one image. The feature extractor outputs a feature vector based on the input image, resulting in two feature vectors.

[0086] S2024: Determine whether there is a correspondence between the two selected images. If there is a correspondence between the two images, set the label value to 0. If there is no correspondence between the two images, set the label value to 1.

[0087] S2025: Calculate the loss function based on the two feature vectors and the label values, as follows:

[0088] loss=(1-label)euclidean(f1,f2) 2 +label(1-euclidean(f1,f2)) 2

[0089] Where loss is the loss value, label is the label value, euclidean(f1,f2) is the normalized Euclidean distance between two feature vectors, f1 is the feature vector corresponding to one selected image, and f2 is the feature vector corresponding to another selected image.

[0090] S2026: Train the feature extractor based on the loss function.

[0091] S203: Construct a change detection network, including:

[0092] S2031: Construct a sub-network, including: providing a trained feature extractor to receive the input image, extracting 32×32×32, 32×64×64, and 32×128×128 feature maps from the input image, upsampling the 32×32×32 feature map to obtain upsampled features, and downsampling the 32×128×128 feature map to obtain downsampled features; connecting the input of the feature processing module to the output of the trained feature extractor to fuse the 32×64×64 feature map, upsampled features, and downsampled features to obtain a fused feature map; connecting the input of the spatial attention module to the output of the feature processing module, the input of the converter module to the output of the spatial attention module, and the input of the channel attention module to the output of the converter module to process the fused feature map sequentially through the spatial attention module, the converter module, and the channel attention module to obtain the final feature map;

[0093] S2032: Provides two sub-networks;

[0094] S2033: Connect the input of the connection module to the output of the channel attention module of the two sub-networks respectively to obtain two final feature maps. Connect the two final feature maps in pairs along the channel dimension to obtain the overall map.

[0095] S2034: Connect the input of the classifier to the output of the connection module to output a predicted change map based on the overall map.

[0096] S204: Provide two detection images. The two detection images were captured at different times but in the same area. The two detection images are used as two input images to the change detection network. The input of the trained feature extractor in the change detection network corresponds one-to-one with the input images. The change detection network outputs the corresponding predicted change map.

[0097] It should be noted that in step S201, constructing the sample set includes: obtaining a simulated cloud image using Perlin noise combined with fractal Brownian motion; providing at least two remote sensing images; fusing each remote sensing image with the simulated cloud image to obtain a foggy image corresponding to the remote sensing image; and all remote sensing images and all foggy images constitute the sample set. Perlin noise is a type of noise that can generate continuous and smooth random values, suitable for simulating various phenomena in nature, such as the undulation of mountains, the shape of clouds, and the dynamic changes of flames, and has the characteristics of smoothness, multidimensionality, and predictability. In this embodiment, Perlin noise of different frequencies combined with fractal Brownian motion is used to simulate the interference effect of thin clouds in nature, referring to... Figure 4 , Figure 4The LEVIR-CD dataset is a dataset composed of remote sensing images, while the LEVIR-CD-cloud dataset is a dataset composed of foggy images corresponding to the remote sensing images in the LEVIR-CD dataset. The size of the foggy images is 1024×1024 pixels.

[0098] Specifically, the remote sensing image is fused with the cloud simulation image to obtain the foggy image corresponding to the remote sensing image, and the calculation is performed as follows:

[0099] I(x)=J(x)t(x)+A(1-t(x))

[0100] Where J(x) is the remote sensing image, t(x) is a parameter representing the non-scattering light reaching the camera, A is the ambient light, representing the cloud simulation image, I(x) is the foggy image corresponding to the remote sensing image, and x = (x, y) are the image coordinates.

[0101] For each remote sensing image in the LEVIR-CD dataset, a cloud simulation with random initial values ​​was performed to generate a corresponding foggy image. Four remote sensing images and their corresponding foggy images were randomly selected for reference. Figure 4 ,Depend on Figure 4 It can be seen that the perlin noise combined with fractal Brownian motion effectively simulates the effect of thin cloud interference.

[0102] Reference Figure 5 In step S202, a feature extractor is trained using a sample set. The trained feature extractor is based on contrastive learning. Remote sensing images and their corresponding foggy images are used as positive sample pairs, while other combinations are used as negative sample pairs. When the input image is a positive sample pair, label = 0 (label value 0); when the input image is a negative sample pair, label = 1 (label value 1). Thus, the constructed loss function achieves the following: when a positive sample pair is input, the smaller the difference between the extracted feature vectors f1 and f2, the smaller the total loss function; when a negative sample pair is input, the larger the difference between the extracted feature vectors f1 and f2, the smaller the total loss function. This maximizes the similarity between positive sample pairs and the difference between negative sample pairs.

[0103] Understandably, in step S2031, the feature extractor removes the initial fully connected layer as the backbone and extracts multi-scale features from the input images I1 and I2. The ResNet backbone of the feature extractor consists of five main blocks, including one 7×7 convolutional layer and four residual blocks. For ease of understanding, these five blocks are referred to as Conv1, Res2, Res3, Res4, and Res5, respectively. Res3 and Res4 perform downsampling with a stride of 2. For the input bitemporal image I... i(i∈{1,2})∈R 3×H×w Three feature maps of different scales were extracted from Res2, Res3, and Res5, respectively. and ).

[0104] In step S2031, the fused feature map is processed sequentially through a spatial attention module, a converter module, and a channel attention module to obtain the final feature map, including:

[0105] The spatial attention module processes the fused feature map to obtain the variable feature map;

[0106] The converter module and the channel attention module process the variable feature map to obtain the final feature map.

[0107] The spatial attention module processes the fused feature map to obtain the variable feature map, which is calculated as follows:

[0108]

[0109] in, To fuse feature maps, σ is the activation function, f is the convolution kernel operation, Maxpool is the max pooling operation, and Avgpool is the average pooling operation. For two-dimensional spatial attention, For element-wise multiplication, For variable feature maps.

[0110] The converter module and the channel attention module process the variable feature map to obtain the final feature map, which is calculated as follows:

[0111]

[0112] M c =MLP(M i )

[0113]

[0114] Maxpool represents the max pooling operation. Here, T() represents the variable feature map, T() represents the converter module processing, ⊕ represents element-wise addition, and MLP represents multilayer perceptron processing. For element-wise multiplication, For the final feature map, For a given variable feature map of a detection image, This is a variable feature map of another detection image provided.

[0115] The feature map is obtained by mixing the extracted features. The subsequent Spatial Attention Module (SAM) automatically emphasizes important information relevant to the feature map at location. (See reference...) Figure 6 SAM utilizes two-dimensional spatial attention In each Element-wise multiplication is implemented. Meaningful features related to positional changes are given greater weight. In this way, the Spatial Attention Module (SAM) can effectively highlight changing regions in bi-temporal images and suppress features in irrelevant regions. To obtain two-dimensional spatial attention... Perform average pooling and max pooling operations along the channel dimension, and then concatenate the results of the pooling operations to generate...

[0116] Reference Figure 7 The variable feature map is obtained after passing through the Spatial Attention Module (SAM). The final feature map is then generated using the Transformer module and the Channel Attention module (CAM). The Transformer module utilizes encoder and decoder blocks and is plug-and-play in the change detection network provided in this embodiment. In this embodiment, the Spatial Attention Module (SAM) and the Transformer module are used to model spatial context information and global context information, respectively. The Channel Attention Module (CAM) models channel context information by highlighting channels relevant to changes. Figure 7 As shown, for image I i Multiple features share the same channel attention M c To compute channel attention, feature maps of the same scale from two sub-branches are first fused using element-wise summation. Then, max-pooling is applied along the spatial dimension of the fused result. Next, element-wise summation is used again to merge the multi-scale results of the max-pooling operation, and the fused result is passed through a multilayer perceptron (MLP) to obtain channel attention. This MLP consists of a fully convolutional layer, a ReLU activation function, a fully convolutional layer, and a sigmoid activation function. Figure 7 In this context, Maxpool stands for max-pooling.

[0117] The change detection network provided in this embodiment fuses global and local image features and employs spatial and channel attention to improve efficiency and accuracy. It utilizes Perlin noise to simulate natural thin cloud interference and adds it to the classic change dataset to generate a new change detection dataset with thin cloud interference. This enables the change detection network to have superior performance in change detection tasks of remote sensing images with cloud interference, especially thin cloud interference.

[0118] Example 3

[0119] Reference Figure 2 Based on the same inventive concept, this invention also provides a remote sensing image anti-cloud and fog interference change detection network based on contrastive learning, comprising:

[0120] Two identical subnetworks 1, each including a trained feature extractor 2, a feature processing module 3, a spatial attention module 4, a converter module 5, and a channel attention module 6. The input of the feature processing module 3 is connected to the output of the trained feature extractor 2, the input of the spatial attention module 4 is connected to the output of the feature processing module 3, the input of the converter module 5 is connected to the output of the spatial attention module 4, and the input of the channel attention module 6 is connected to the output of the converter module 5.

[0121] The input of the connection module 7 is connected to the output of the channel attention module 6 of the two sub-networks 1 respectively;

[0122] Classifier 8, the input of classifier 8 is connected to the output of connection module 7.

[0123] It should be noted that the trained feature extractor 2 is used to receive the input image and extract 32×32×32, 32×64×64, and 32×128×128 feature maps from the input image. The 32×32×32 feature map is upsampled to obtain upsampled features, and the 32×128×128 feature map is downsampled to obtain downsampled features. The feature processing module 3 is used to fuse the 32×64×64 feature map, upsampled features, and downsampled features to obtain a fused feature map. The spatial attention module 4, converter module 5, and channel attention module 6 are used to process the fused feature map sequentially through the spatial attention module 4, converter module 5, and channel attention module 6 to obtain the final feature map. The connection module 7 is used to obtain two final feature maps and connect the two final feature maps in pairs along the channel dimension to obtain the overall map. The classifier 8 is used to output a predicted change map based on the overall map.

[0124] Understandably, the remote sensing image anti-cloud and fog interference change detection network constructed using the trained feature extractor 2 fuses global and local features of the image, and uses spatial attention module 4 and channel attention module 6 to improve efficiency and accuracy. Even in remote sensing image change detection tasks with cloud interference, it can achieve fast and accurate detection, thus improving the robustness of remote sensing image change detection.

[0125] As can be seen from the above embodiments, the remote sensing image anti-cloud and fog interference change detection method based on contrastive learning provided by the present invention achieves at least the following beneficial effects:

[0126] This invention provides a remote sensing image change detection method based on contrastive learning, which is robust against cloud and fog interference. The method includes: constructing a sample set; training a feature extractor using the sample set to obtain a trained feature extractor; constructing a change detection network using the trained feature extractor, a feature processing module, a spatial attention module, a converter module, a channel attention module, a connection module, and a classifier; inputting two detection images (captured at different times but in the same area) into the change detection network, and having the network output a predicted change map. The feature extractor is trained using a sample set containing both remote sensing and foggy images, resulting in a trained feature extractor with good resistance to cloud interference. The change detection network constructed using the trained feature extractor fuses global and local features of the image, and employs spatial and channel attention modules to improve efficiency and accuracy. Even in remote sensing image change detection tasks with cloud interference, this method achieves fast and accurate detection, improving the robustness of remote sensing image change detection.

[0127] While specific embodiments of the invention have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of the invention. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of the invention. The scope of the invention is defined by the appended claims.

Claims

1. A method for detecting changes in remote sensing images against cloud and fog interference based on contrastive learning, characterized in that, include: Construct a sample set, which includes remote sensing images and foggy images. The number of remote sensing images is at least two, and any two remote sensing images are different. Each remote sensing image corresponds one-to-one with a foggy image, and each foggy image is obtained from its corresponding remote sensing image. The feature extractor is trained using the sample set to obtain the trained feature extractor; Constructing a change detection network includes: Constructing a sub-network includes: providing a trained feature extractor for receiving an input image, extracting 32×32×32, 32×64×64, and 32×128×128 feature maps from the input image, performing upsampling on the 32×32×32 feature map to obtain upsampled features, and performing downsampling on the 32×128×128 feature map to obtain downsampled features; connecting the input of a feature processing module to the output of the trained feature extractor to fuse the 32×64×64 feature map, the upsampled features, and the downsampled features to obtain a fused feature map; connecting the input of a spatial attention module to the output of the feature processing module, the input of a converter module to the output of the spatial attention module, and the input of a channel attention module to the output of the converter module to process the fused feature map sequentially through the spatial attention module, the converter module, and the channel attention module to obtain a final feature map; Provide two of the aforementioned sub-networks; The input of the connection module is connected to the output of the channel attention module of the two sub-networks respectively to obtain two final feature maps. The two final feature maps are then connected in pairs along the channel dimension to obtain the overall image. The input of the classifier is connected to the output of the connection module to output a predicted change map based on the overall map. Two detection images are provided, which are captured at different times but in the same area. The two detection images are used as input images to the change detection network. The input of the feature extractor trained in the change detection network corresponds one-to-one with the input images. The change detection network outputs the corresponding predicted change map.

2. The method for detecting changes in remote sensing images against cloud and fog interference based on contrastive learning according to claim 1, characterized in that, Constructing the sample set includes: Simulated cloud images were obtained by combining Berlin noise with fractal Brownian motion. Provide at least two of the aforementioned remote sensing images; Each of the remote sensing images is fused with the cloud simulation image to obtain the foggy image corresponding to the remote sensing image; All of the remote sensing images and all of the foggy images constitute the sample set.

3. The remote sensing image anti-cloud and fog interference change detection method based on contrastive learning according to claim 2, characterized in that, The remote sensing image is fused with the cloud simulation image to obtain the foggy image corresponding to the remote sensing image, and the calculation is performed as follows: I(x)=J(x)t(x)+A(1-t(x)) Wherein, J(x) is the remote sensing image, t(x) is a parameter representing the light that does not scatter and reaches the camera, A is the cloud simulation image, and I(x) is the foggy image corresponding to the remote sensing image.

4. The method for detecting changes in remote sensing images based on contrastive learning to resist cloud and fog interference according to claim 1, characterized in that, The feature extractor is trained using the sample set to obtain the trained feature extractor, including: Two of the aforementioned feature extractors are provided; Two images are randomly selected from the sample set. The two selected images are either any two remote sensing images, or any two foggy images, or one of the two selected images is any remote sensing image and the other is any foggy image. The two selected images are input into two feature extractors, each corresponding to one of the images. The feature extractors output feature vectors based on the input images, thus obtaining two feature vectors. Determine whether there is a correspondence between the two selected images. If there is a correspondence between the two images, set the label value to 0. If there is no correspondence between the two images, set the label value to 1. Calculate the loss function based on the two feature vectors and the label values; The feature extractor is trained based on the loss function.

5. The method for detecting changes in remote sensing images based on contrastive learning to resist cloud and fog interference according to claim 4, characterized in that, The loss function is calculated based on the two feature vectors and the label value, as follows: loss=(1-label)euclidean(f1,f2) 2 +label(1-euclidean(f1,f2)) 2 Wherein, loss is the loss value, label is the label value, euclidean(f1, f2) is the normalized Euclidean distance between the two feature vectors, f1 is the feature vector corresponding to one of the selected images, and f2 is the feature vector corresponding to the other selected image.

6. The method for detecting changes in remote sensing images against cloud and fog interference based on contrastive learning according to claim 1, characterized in that, The fused feature map is processed sequentially through the spatial attention module, the converter module, and the channel attention module to obtain the final feature map, including: The spatial attention module processes the fused feature map to obtain a variable feature map; The converter module and the channel attention module process the variable feature map to obtain the final feature map.

7. The method for detecting changes in remote sensing images based on contrastive learning to resist cloud and fog interference according to claim 6, characterized in that, The spatial attention module processes the fused feature map to obtain the variable feature map, which is calculated in the following manner: in, The fused feature map is defined as follows: σ is the activation function, f is the convolution kernel operation, Maxpool is the max pooling operation, and Avgpool is the average pooling operation. For two-dimensional spatial attention, For element-wise multiplication, This is the feature map of the variable.

8. The method for detecting changes in remote sensing images based on contrastive learning to resist cloud and fog interference according to claim 6, characterized in that, The converter module and the channel attention module process the variable feature map to obtain the final feature map, which is calculated in the following manner: M c =MLP(M i ) Maxpool represents the max pooling operation. Here, T() represents the variable feature map, T() represents the processing by the converter module, ⊕ represents element-wise addition, and MLP represents multilayer perceptron processing. For element-wise multiplication, For the final feature map, The variable feature map of the provided detection image. The variable feature map is provided for another of the detected images.

Citation Information

Patent Citations

  • Change Detection Method of Remote Sensing Image

    CN104182985B