Remote sensing image change detection method and system based on depth feature alignment

By constructing a deep feature alignment model, the problem of insufficient feature alignment in remote sensing image change detection is solved, and high-precision and robust change detection is achieved, which is suitable for remote sensing image change detection in complex scenarios.

CN120451791APending Publication Date: 2025-08-08JINAN GUIHUA DESIGN RES YUAN
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510575706.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

When the existing remote sensing image change detection methods deal with high-resolution images, large-scale areas and complex scenes, there are problems such as insufficient alignment of depth features and poor robustness, resulting in low detection accuracy and susceptibility to noise interference.

Method used

Using a remote sensing image change detection method based on depth feature alignment, a model including the first reference encoder, the second reference encoder, a differential feature compensator, a depth feature extractor and a feature alignment function is constructed, and the depth feature extractor is extracted, compensated and aligned. The model is optimized by using the loss function and the optimizer to output an accurate change image.

Benefits of technology

It significantly improves the accuracy and robustness of change detection, can accurately detect changing areas in complex scenarios, reduce background noise interference, improve the reliability and adaptability of detection results, and is suitable for change detection tasks of multiple resolutions and multi-source data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451791A_ABST
    Figure CN120451791A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing image detection, in particular to a remote sensing image change detection method and system based on depth feature alignment, and the method specifically comprises the following steps: collecting remote sensing images of sub-regions at different moments, constructing a data set, and carrying out the preprocessing, thereby obtaining complete remote sensing images at different moments; manually marking change areas of the remote sensing graphs at any two moments to obtain a reference change image; constructing a remote sensing image change detection model which comprises a first reference encoder, a second reference encoder, a difference feature compensator, a first depth feature extractor, a second depth feature extractor, a feature alignment function and a depth decoder; inputting the remote sensing images at two different moments into the model, and outputting a predicted change image through the processing of the model; and the model is optimized through a # imgabs0 # loss function and a # imgabs1 # optimizer. The method can be used for application scenes such as city change monitoring and environment change analysis to improve the precision and robustness of remote sensing image change detection in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image detection, and in particular to a remote sensing image change detection method and system based on depth feature alignment. Background Art

[0002] Urban construction land, as the cornerstone of urban development, faces significant challenges in its efficient utilization and scientific management, particularly in relation to sustainable urban development, improved quality of life for residents, and ecological conservation. With the acceleration of urbanization, land resources are becoming increasingly scarce. Scientifically assessing the performance potential of urban construction land and implementing dynamic monitoring have become key issues in urban renewal and management. Remote sensing image change detection is a key technology in remote sensing image analysis, widely used in fields such as land use change monitoring and urban expansion monitoring. The fundamental task of change detection is to compare remote sensing images from different time periods to identify areas of change, analyze the nature of these changes, and assess the performance potential of construction land, thereby enabling change monitoring and optimized management of the detected areas. However, remote sensing image change detection faces numerous challenges, particularly when processing high-resolution imagery, large-scale regions, and complex scenes. Traditional change detection methods primarily rely on pixel-level difference detection or spectral feature comparison, such as image differencing, ratioing, and principal component analysis (PCA). While these methods have achieved promising results in some applications, they also suffer from numerous limitations. First, due to differences in remote sensing image acquisition conditions (such as sensor type, shooting angle, and time difference), the spectral characteristics or spatial structure of the image may vary significantly, making direct comparison of image pixel values inaccurate or unreliable. Second, traditional methods often ignore the contextual information of the image, resulting in detection results that are affected by factors such as noise, shadows, and object occlusion, resulting in low accuracy and poor robustness.

[0003] With the rapid development of deep learning technology, remote sensing image change detection methods based on deep learning models such as convolutional neural networks (CNNs) have gradually gained popularity. Deep learning can automatically learn useful feature representations from large amounts of data, thereby improving the accuracy and robustness of change detection. In particular, in the complex task of remote sensing image change detection, deep learning models can effectively capture semantic information in images through hierarchical feature learning, compensating for the shortcomings of traditional methods in feature extraction. However, existing remote sensing image change detection methods based on deep learning still face several challenges, particularly the alignment of deep features between images. Due to the different capture times and viewpoints of remote sensing images, even images from the same region can exhibit significant differences in shape, scale, texture, and other aspects. These differences can cause spatial misalignment or misalignment of the image's deep features, thus affecting the accuracy and reliability of change detection.

[0004] However, existing deep feature alignment methods still have some shortcomings, such as insufficient adaptability to complex scenes and poor robustness to noise, and it is still difficult to achieve ideal results in practical applications. The present invention proposes a remote sensing image change detection method based on deep feature alignment. In response to the shortcomings of the existing technology, a new method is proposed. By effectively aligning the deep features of remote sensing images, the accuracy and robustness of change detection are significantly improved. By introducing a deep feature alignment module, this method solves the image alignment problem of remote sensing images at different times, different sensors or different perspectives, thereby being able to accurately detect the changed areas in the image, especially in high-complexity and large-scale remote sensing data processing tasks, and shows excellent performance.

[0005] Therefore, the present invention proposes a remote sensing image change detection method and system based on deep feature alignment to solve the above problems. Summary of the Invention

[0006] In response to the shortcomings of the existing technology, the present invention develops a remote sensing image change detection method and system based on depth feature alignment. By effectively aligning the depth features of remote sensing images, the present invention significantly improves the accuracy and robustness of change detection. By introducing a deep feature alignment module, the changed areas in the image can be accurately detected.

[0007] On the one hand, the technical solution to the technical problem of the present invention is a remote sensing image change detection method based on deep feature alignment, comprising the following steps: S1. Collect remote sensing images of sub-regions at different times to construct Dataset, for Preprocess the remote sensing images in the dataset to obtain complete remote sensing images at different times. Manually mark the change areas of the remote sensing images at any two times to obtain the corresponding benchmark change images. ,Finally, the preprocessed dataset is divided into training set and test set; S2. Constructing a remote sensing image change detection model, which includes a first reference encoder, a second reference encoder, a difference feature compensator, a first deep feature extractor, a second deep feature extractor, a feature alignment function, and a deep decoder; S3: Input any two remote sensing images at different times in the training set into the remote sensing image change detection model, and output the predicted change image after being processed by the model. ; S4, through Loss function calculates predicted change image and baseline change images Loss, reuse The optimizer optimizes the remote sensing image change detection model to obtain a trained remote sensing image change detection model; S5. Input the remote sensing images in the test set into the trained remote sensing image change detection model and output the final predicted change image. .

[0008] S1 is as follows: S1.1. Collect remote sensing images to build a dataset: Collect remote sensing images of a certain area by using multiple drones equipped with cameras to take pictures simultaneously. Remote sensing images of each sub-area in the area are collected. The remote sensing images of each sub-area are aggregated to cover the entire area. The remote sensing images of the sub-areas collected at the same time are grouped as a group of images. Remote sensing images at multiple different times are collected. S1.2. Preprocess the remote sensing images in the dataset: Use ENVI software to perform image radiation correction, and then stitch the remote sensing images of multiple sub-areas into a complete remote sensing image to obtain the preprocessed data set , , Represents the preprocessed dataset The CCP uses A complete remote sensing image at a moment, Indicates the Complete remote sensing images at all times, ; S1.3 Image Annotation: Select remote sensing images at any two moments, use QGIS semi-automatic annotation tools combined with manual review, and manually annotate the changed areas in the remote sensing images at the two moments to obtain the baseline change image. ; S1.4. Divide the data set: The preprocessed dataset is divided into training and testing sets in proportion.

[0009] S2 is as follows: (1) The first benchmark encoder consists of the first convolutional pooling layer, the second convolutional pooling layer, the depthwise separable convolutional pooling layer, and the fully connected layer. The convolution kernel size of the first convolutional pooling layer is 3×3, the number of convolution kernels is 64, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, the step size is 2; the second convolution pooling layer has a convolution kernel size of 3×3, the number of convolution kernels is 128, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, the step size is 2; the depth-separable convolution pooling layer has a convolution kernel size of 7×7, the number of convolution kernels is 512, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects average pooling, the pooling size is 3×3, and the stride is 1; the fully connected layer uses one layer of full connection, and the output is a 2048-dimensional vector; (2) The second benchmark encoder sequentially includes the first convolution pooling layer, the second convolution pooling layer, the dilated convolution pooling layer, and the fully connected layer. The convolution kernel size of the first convolution pooling layer is 1×1, the number of convolution kernels is 64, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, the step size is 2; the second convolution pooling layer has a convolution kernel size of 3×3, the number of convolution kernels is 128, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, and the step size is 2; the convolution kernel size of the void convolution pooling layer is 5×5, the number of convolution kernels is 256, the void rate is set to 2, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects average pooling, the pooling size is 3×3, and the stride is 1; the fully connected layer uses one layer of full connection, and the output is a 2048-dimensional vector; (3) The difference feature compensator sequentially includes an input layer, a feature difference calculation module, a compensator convolution module, and a compensation feature generation module; Among them, the feature difference calculation module is used to calculate the feature difference after the shape is changed; The compensator convolution module includes three convolution layers for feature extraction and transformation of feature differences. The first convolution layer has 256 input channels, 128 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. The second convolution layer has 128 input channels, 64 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. The third convolution layer has 64 input channels, 256 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. Each convolution layer is followed by batch normalization and Activation function; The compensation feature generation module generates compensation features by fusing features; (4) The first deep feature extractor includes the first convolutional layer, the second convolutional layer, the deep feature enhancer and the flattening layer. The first convolutional layer has 256 input channels, 128 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. It contains a maximum pooling layer with a window size of 2×2 and a stride of 2. The output spatial resolution becomes 32×32. The second convolutional layer has 128 input channels, 256 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. It is followed by batch normalization and Activation function; the deep feature enhancer contains a global average pooling layer, which reduces the spatial resolution to 1×1 and retains only the global features of each channel; (5) The second deep feature extractor is the same as the first deep feature extractor and receives the compensation features, but the weights are different from those of the first deep model; (6) The deep decoder consists of the first deconvolution layer, the second deconvolution layer, the third deconvolution layer, Activation function and thresholding function, where the convolution kernel size of the first deconvolution layer is 4×4, the input channel is 128, the output channel is 64, the step size is 2, and the padding is 1; the convolution kernel size of the second deconvolution layer is 4×4, the input channel is 64, the output channel is 32, the step size is 2, and the padding is 1; the convolution kernel size of the third deconvolution layer is 4×4, the input channel is 32, the output channel is 1, the step size is 2, and the padding is 1.

[0010] S3 is as follows: The preprocessed dataset The complete remote sensing image at any two moments and Input to the remote sensing image change detection model, Indicates the A complete remote sensing image at a moment, Indicates the A complete remote sensing image at a moment, , , ; Input to the first reference encoder, Input into the second benchmark encoder, the two images are sequentially passed through the first convolution pooling layer, the second convolution pooling layer, the depthwise separable convolution pooling layer, and the fully connected layer in each benchmark encoder to obtain the first The benchmark encoding features at each moment Hedi The benchmark encoding features at each moment ; The features and Input to the difference feature compensator, and pass through the input layer of the difference feature compensator Function changes characteristics and The shapes of and features , and then calculate the feature through the feature difference calculation module and features The difference between the characteristics , , and then the difference features The input is sent to the compensation convolution module through three convolution layers, and the compensation feature is finally output. , the calculation formula is as follows: , in, and represents two different learnable parameters, ; Then, the compensation feature Input to the first deep feature extractor and pass through the first convolution layer to obtain ,Will Input to the second convolutional layer to get ,Will Input into the deep feature enhancer and flattening layer to obtain the first deep feature , similarly, the compensation feature Input to the second deep feature extractor and pass through the first convolution layer to obtain ,Will Input to the second convolutional layer to get ,Will Input to the deep feature enhancer and flattening layer to obtain the second deep feature ; Then the features and the first depth feature Add up the features , the features and the second deep feature Add up the features , and then the features and Input into the feature alignment function and calculate the alignment features , the specific calculation is as follows: , in, Used to limit the output range to , Represents the hyperparameter used to adjust the weight of the difference feature, setting , Represents the hyperparameter used to adjust the weight of collaborative features, setting , Indicates the smoothing factor, the default value is , represents element-by-element addition, represents element-wise multiplication, represents the two-norm normalization operation; Finally, the features ,feature and align features After adding, it is input into the depth decoder and passes through the first deconvolution layer to obtain ,Will Input to the second deconvolution layer to get ,Will Input to the third deconvolution layer and In the activation function, the change probability map is obtained , change probability map Each pixel of is between [0,1], and then the change probability map is thresholded. Each pixel value is thresholded to obtain the predicted change image , the calculation process is as follows: , in, Indicates the threshold value, set , represents the pixel position, express exist The pixel value at express exist The pixel value at , .

[0011] S4 is as follows: pass Loss Function Calculate the predicted change image and baseline change images The loss is calculated as follows: , in, Indicates the total number of pixels, express The index of Indicates the first pixels, Represents the predicted change in the image pixels.

[0012] On the other hand, the present invention also provides a remote sensing image change detection system based on deep feature alignment, which implements a remote sensing image change detection method based on deep feature alignment.

[0013] The effects provided in the summary of the invention are only the effects of the embodiments, rather than all the effects of the invention. The above technical solution has the following advantages or beneficial effects: The present invention effectively solves the technical problems of low detection accuracy and large noise interference caused by insufficient alignment of remote sensing image features in traditional change detection methods by constructing a remote sensing image change detection method based on deep feature alignment. By extracting, compensating and aligning deep features of remote sensing images at different times, this method achieves efficient capture and accurate identification of change information, and has significant advantages in terms of technical details and overall application effects. From a technical point of view, the present invention compensates for the feature deviation problem of time-series remote sensing images caused by environmental changes or radiation differences by innovatively introducing a difference feature compensator and a feature alignment function, making the remote sensing images at any two times more consistent in the feature space. At the same time, the design of the dual-depth feature extractor can extract deeper change information from the compensated features, effectively improving the expressiveness of detailed features. In the feature alignment phase, the baseline encoding features at two moments are combined with the deep features and then processed by the feature alignment function. This greatly enhances the significance of the change features and reduces the interference of background noise. Combined with the efficient reconstruction capability of the deep decoder, the final output change image is more accurate in restoring details and global features. In addition, the iterative optimization using the BCE loss function and the Adam optimizer effectively ensures the rapid convergence and accuracy of the model during training.

[0014] From the overall effect, the change detection method proposed in the present invention has high robustness and adaptability. Through the radiation correction and feature compensation mechanism of ENVI software, it can process remote sensing images in various complex scenes and is not significantly affected by factors such as lighting and weather conditions. In the change detection task, the generated change image can not only accurately depict the changed area, but also has the advantages of clear boundaries and complete textures, which greatly improves the reliability of the detection results. In addition, the method has strong versatility and is suitable for change detection tasks of various resolutions and multi-source data, providing an efficient and reliable technical means for dynamic monitoring, environmental assessment and disaster response in the remote sensing field. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0016] Figure 1 Schematic diagram of the method of the present invention.

[0017] Figure 2 This is a comparative example of the method of the present invention and other methods. Figure 2 (a) is the complete remote sensing image at time t1. Figure 2 (b) is the complete remote sensing image at time t2. Figure 2 (c) is the baseline change image, Figure 2 (d) is the image predicted by the method of the present invention. Figure 2 Middle (e) is the image predicted by the FC-EF method. DETAILED DESCRIPTION

[0018] To clearly illustrate the technical features of this solution, the present invention is described in detail below through specific embodiments and in conjunction with the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and configurations of specific examples are described below.

[0019] Example 1 A remote sensing image change detection method based on deep feature alignment includes the following steps: S1. Collect remote sensing images of sub-regions at different times to construct Dataset, for Preprocess the remote sensing images in the dataset to obtain complete remote sensing images at different times. Manually mark the change areas of the remote sensing images at any two times to obtain the corresponding benchmark change images. ,Finally, the preprocessed dataset is divided into training set and test set; S2. Constructing a remote sensing image change detection model, which includes a first reference encoder, a second reference encoder, a difference feature compensator, a first deep feature extractor, a second deep feature extractor, a feature alignment function, and a deep decoder; S3: Input any two remote sensing images at different times in the training set into the remote sensing image change detection model, and output the predicted change image after being processed by the model. ; S4, through Loss function calculates predicted change image and baseline change images Loss, reuse The optimizer optimizes the remote sensing image change detection model to obtain a trained remote sensing image change detection model; S5. Input the remote sensing images in the test set into the trained remote sensing image change detection model and output the final predicted change image. .

[0020] S1 is as follows: S1.1. Collect remote sensing images to build a dataset: Collect remote sensing images of a certain area by using multiple drones equipped with cameras to take pictures simultaneously. Remote sensing images of each sub-area in the area are collected. The remote sensing images of each sub-area are aggregated to cover the entire area. The remote sensing images of the sub-areas collected at the same time are grouped as a group of images. Remote sensing images at multiple different times are collected. S1.2. Preprocess the remote sensing images in the dataset: Use ENVI software to perform image radiation correction, and then stitch the remote sensing images of multiple sub-areas into a complete remote sensing image to obtain the preprocessed data set , , Represents the preprocessed dataset The CCP uses A complete remote sensing image at a moment, Indicates the Complete remote sensing images at all times, ; S1.3 Image Annotation: Select remote sensing images at any two moments, use QGIS semi-automatic annotation tools combined with manual review, and manually annotate the changed areas in the remote sensing images at the two moments to obtain the baseline change image. ; S1.4. Divide the data set: The preprocessed dataset is divided into training and testing sets in proportion.

[0021] S2 is as follows: (1) The first benchmark encoder consists of the first convolutional pooling layer, the second convolutional pooling layer, the depthwise separable convolutional pooling layer, and the fully connected layer. The convolution kernel size of the first convolutional pooling layer is 3×3, the number of convolution kernels is 64, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, the step size is 2; the second convolution pooling layer has a convolution kernel size of 3×3, the number of convolution kernels is 128, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, the step size is 2; the depth-separable convolution pooling layer has a convolution kernel size of 7×7, the number of convolution kernels is 512, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects average pooling, the pooling size is 3×3, and the stride is 1; the fully connected layer uses one layer of full connection, and the output is a 2048-dimensional vector; (2) The second benchmark encoder sequentially includes the first convolution pooling layer, the second convolution pooling layer, the dilated convolution pooling layer, and the fully connected layer. The convolution kernel size of the first convolution pooling layer is 1×1, the number of convolution kernels is 64, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, the step size is 2; the second convolution pooling layer has a convolution kernel size of 3×3, the number of convolution kernels is 128, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, and the step size is 2; the convolution kernel size of the void convolution pooling layer is 5×5, the number of convolution kernels is 256, the void rate is set to 2, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects average pooling, the pooling size is 3×3, and the stride is 1; the fully connected layer uses one layer of full connection, and the output is a 2048-dimensional vector; (3) The difference feature compensator sequentially includes an input layer, a feature difference calculation module, a compensator convolution module, and a compensation feature generation module; Among them, the feature difference calculation module is used to calculate the feature difference after the shape is changed; The compensator convolution module includes three convolution layers for feature extraction and transformation of feature differences. The first convolution layer has 256 input channels, 128 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. The second convolution layer has 128 input channels, 64 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. The third convolution layer has 64 input channels, 256 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. Each convolution layer is followed by batch normalization and Activation function; The compensation feature generation module generates compensation features by fusing features; (4) The first deep feature extractor includes the first convolutional layer, the second convolutional layer, the deep feature enhancer and the flattening layer. The first convolutional layer has 256 input channels, 128 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. It contains a maximum pooling layer with a window size of 2×2 and a stride of 2. The output spatial resolution becomes 32×32. The second convolutional layer has 128 input channels, 256 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. It is followed by batch normalization and Activation function; the deep feature enhancer contains a global average pooling layer, which reduces the spatial resolution to 1×1 and retains only the global features of each channel; (5) The second deep feature extractor is the same as the first deep feature extractor and receives the compensation features, but the weights are different from those of the first deep model; (6) The deep decoder consists of the first deconvolution layer, the second deconvolution layer, the third deconvolution layer, Activation function and thresholding function, where the convolution kernel size of the first deconvolution layer is 4×4, the input channel is 128, the output channel is 64, the step size is 2, and the padding is 1; the convolution kernel size of the second deconvolution layer is 4×4, the input channel is 64, the output channel is 32, the step size is 2, and the padding is 1; the convolution kernel size of the third deconvolution layer is 4×4, the input channel is 32, the output channel is 1, the step size is 2, and the padding is 1.

[0022] S3 is as follows: The preprocessed dataset The complete remote sensing image at any two moments and Input to the remote sensing image change detection model, Indicates the A complete remote sensing image at a moment, Indicates the A complete remote sensing image at a moment, , , ; Input to the first reference encoder, Input into the second benchmark encoder, the two images are sequentially passed through the first convolution pooling layer, the second convolution pooling layer, the depthwise separable convolution pooling layer, and the fully connected layer in each benchmark encoder to obtain the first The benchmark encoding features at each moment Hedi The benchmark encoding features at each moment ; The features and Input to the difference feature compensator, and pass through the input layer of the difference feature compensator Function changes characteristics and The shapes of and features , and then calculate the feature through the feature difference calculation module and features The difference between the characteristics , , and then the difference features The input is sent to the compensation convolution module through three convolution layers, and the compensation feature is finally output. , the calculation formula is as follows: , in, and represents two different learnable parameters, ; Then, the compensation feature Input to the first deep feature extractor and pass through the first convolution layer to obtain ,Will Input to the second convolutional layer to get ,Will Input into the deep feature enhancer and flattening layer to obtain the first deep feature , similarly, the compensation feature Input to the second deep feature extractor and pass through the first convolution layer to obtain ,Will Input to the second convolutional layer to get ,Will Input to the deep feature enhancer and flattening layer to obtain the second deep feature ; Then the features and the first depth feature Add up the features , the features and the second deep feature Add up the features , and then the features and Input into the feature alignment function and calculate the alignment features , the specific calculation is as follows: , in, Used to limit the output range to , Represents the hyperparameter used to adjust the weight of the difference feature, setting , Represents the hyperparameter used to adjust the weight of collaborative features, setting , Indicates the smoothing factor, the default value is , represents element-by-element addition, represents element-wise multiplication, represents the two-norm normalization operation; Finally, the features ,feature and align features After adding, it is input into the depth decoder and passes through the first deconvolution layer to obtain ,Will Input to the second deconvolution layer to get ,Will Input to the third deconvolution layer and In the activation function, the change probability map is obtained , change probability map Each pixel of is between [0,1], and then the change probability map is thresholded. Each pixel value is thresholded to obtain the predicted change image , the calculation process is as follows: , in, Indicates the threshold value, set , represents the pixel position, express exist The pixel value at express exist The pixel value at , .

[0023] S4 is as follows: pass Loss Function Calculate the predicted change image and baseline change images The loss is calculated as follows: , in, Indicates the total number of pixels, express The index of Indicates the first pixels, Represents the predicted change in the image pixels.

[0024] Example 2 A remote sensing image change detection system based on deep feature alignment implements a remote sensing image change detection method based on deep feature alignment.

[0025] Example 3 In order to better demonstrate the technical effect of the present invention, the method of the present invention is compared with the existing methods. As shown in Table 1, the method of the present invention is compared with the existing methods FC-EF (Fully Convolutional Encoder-Feature, a fully convolutional neural network FCN structure for change detection) and DiffMatch (an open source project developed by Google engineer Kevin Decker to solve the problems of text difference comparison, matching and patch application). The evaluation indicators are detection accuracy, recall rate and precision. Detection accuracy (Accuracy) indicates the correctness of the overall prediction of the detection model; recall rate (Recall) indicates the proportion of all real change areas successfully detected by the model; precision (Precision) indicates the proportion of real change areas among the change areas detected by the model; Table 1 Comparison between the method of the present invention and the existing method As shown in Table 1, the FC-EF method achieved a detection accuracy of 83.2%, the worst performance among the three methods. This may be due to its limited feature extraction and change region detection capabilities, resulting in the model's relatively weak overall prediction. The DiffMatch method achieved a detection accuracy of 85.6%, a 2.4% increase over FC-EF, indicating some improvement in feature alignment and change region identification. The proposed method achieved a detection accuracy of 94.8%, a 9.2% increase over DiffMatch. This significant advantage is attributed to the proposed differential feature compensator and deep feature alignment mechanism, which effectively reduce false detections and missed detections. The FC-EF method had a recall rate of 81.5%, the worst performance, indicating that it missed many changed regions. The DiffMatch method had a recall rate of 84.2%, a 2.7% improvement over FC-EF, indicating that it was better at detecting more real changed regions. The proposed method achieved a recall rate of 92.3%, significantly higher than other methods. This excellent performance is attributed to its innovative deep feature extractor and feature alignment mechanism, which can capture more change information and reduce missed regions. The FC-EF method achieved an accuracy of 80.8%, also the worst performance, indicating that it detected many false positives in the changed regions. The DiffMatch method achieved an accuracy of 83.9%, a 3.1% improvement over FC-EF, reducing the false positive rate to a certain extent. The proposed method achieved an accuracy of 93.7%, a 9.8% improvement over DiffMatch, demonstrating significant false positive suppression capabilities. This is due to the deep decoder and feature alignment function of the proposed method, resulting in more accurate detection of the changed regions.

[0026] In summary, the present invention effectively solves the problems of missed detection and false detection caused by insufficient feature alignment in traditional methods by introducing a difference feature compensator and a deep feature alignment mechanism, and achieves a comprehensive improvement in detection performance. Whether it is overall detection accuracy, missed detection rate control, or false detection area suppression, it is superior to the other two methods.

[0027] Example 4 In order to better demonstrate the technical effect of the present invention, the processing results of the actually collected remote sensing images are taken as an example. Figure 2 As shown, remote sensing images at two moments are collected and manually annotated to obtain corresponding baseline change images. The remote sensing images collected at the two moments are processed by the method of the present invention and the existing method FC-EF respectively to obtain prediction results. The prediction results of the two methods are compared with the baseline change images to verify the superiority of the method of the present invention in actual scenarios. It can be seen from the baseline change area image that the change area is mainly concentrated in the location of the building (such as the addition of a new building or the demolition of a building). The baseline image clearly annotates these change areas, providing an objective standard for evaluating the detection method. The detection results of the method of the present invention are highly consistent with the baseline change area, and can accurately identify the change area where a building is added or disappeared. The detected change area has sharp edges and is almost error-free with the baseline change image. There is no obvious false detection or redundant noise area, indicating that the method of the present invention has excellent noise suppression ability. In the darker areas or areas with unclear details in the image, the method of the present invention can still accurately identify the change area; However, the detection results of the FC-EF method have certain defects. The boundaries of the changed areas are not clear enough, there is a certain deviation from the reference image, and there are some obvious false detections in the background areas where there is no actual change, which affects the reliability of the results. Some real changed areas are not successfully detected, especially small or detailed areas. From the detection effect point of view, the method of the present invention is superior to the traditional method (FC-EF) in terms of accurate identification of changed areas and noise suppression capabilities. Whether in boundary clarity, detection integrity, or background noise suppression, the method of the present invention shows significant advantages. Through this experimental result, it can be proved that the method of the present invention is applicable and efficient in remote sensing image change detection, providing a reliable solution for dynamic monitoring in complex scenes.

[0028] Although the above describes the specific implementation methods of the invention in conjunction with the accompanying drawings, it does not limit the scope of protection of the invention. Based on the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present invention.

Claims

1. A remote sensing image change detection method based on deep feature alignment, characterized in that: The following steps are involved: S1. Collect remote sensing images of sub-regions at different times to construct Dataset, for Preprocess the remote sensing images in the dataset to obtain complete remote sensing images at different times. Manually mark the change areas of the remote sensing images at any two times to obtain the corresponding benchmark change images. ,Finally, the preprocessed dataset is divided into training set and test set; S2. Constructing a remote sensing image change detection model, which includes a first reference encoder, a second reference encoder, a difference feature compensator, a first deep feature extractor, a second deep feature extractor, a feature alignment function, and a deep decoder; S3: Input any two remote sensing images at different times in the training set into the remote sensing image change detection model, and output the predicted change image after being processed by the model. ; S4, through Loss function calculates predicted change image and baseline change images Loss, reuse The optimizer optimizes the remote sensing image change detection model to obtain a trained remote sensing image change detection model; S5. Input the remote sensing images in the test set into the trained remote sensing image change detection model and output the final predicted change image. .

2. The remote sensing image change detection method based on deep feature alignment according to claim 1 is characterized in that: S1 is as follows: S1.

1. Collect remote sensing images to build a dataset: Collect remote sensing images of a certain area by using multiple drones equipped with cameras to take pictures simultaneously. Remote sensing images of each sub-area in the area are collected. The remote sensing images of each sub-area are aggregated to cover the entire area. The remote sensing images of the sub-areas collected at the same time are grouped as a group of images. Remote sensing images at multiple different times are collected. S1.

2. Preprocess the remote sensing images in the dataset: Use ENVI software to perform image radiation correction, and then stitch the remote sensing images of multiple sub-areas into a complete remote sensing image to obtain the preprocessed data set , , Represents the preprocessed dataset The CCP uses A complete remote sensing image at a moment, Indicates the Complete remote sensing images at all times, ; S1.3 Image Annotation: Select remote sensing images at any two moments, use QGIS semi-automatic annotation tools combined with manual review, and manually annotate the changed areas in the remote sensing images at the two moments to obtain the baseline change image. ; S1.

4. Divide the data set: The preprocessed dataset is divided into training and testing sets in proportion.

3. The remote sensing image change detection method based on deep feature alignment according to claim 2, wherein S2 The details are as follows: (1) The first benchmark encoder consists of the first convolutional pooling layer, the second convolutional pooling layer, the depthwise separable convolutional pooling layer, and the fully connected layer. The convolution kernel size of the first convolutional pooling layer is 3×3, the number of convolution kernels is 64, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, the step size is 2; the second convolution pooling layer has a convolution kernel size of 3×3, the number of convolution kernels is 128, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, the step size is 2; the depth-separable convolution pooling layer has a convolution kernel size of 7×7, the number of convolution kernels is 512, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects average pooling, the pooling size is 3×3, and the stride is 1; the fully connected layer uses one layer of full connection, and the output is a 2048-dimensional vector; (2) The second benchmark encoder sequentially includes the first convolution pooling layer, the second convolution pooling layer, the dilated convolution pooling layer, and the fully connected layer. The convolution kernel size of the first convolution pooling layer is 1×1, the number of convolution kernels is 64, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, the step size is 2; the second convolution pooling layer has a convolution kernel size of 3×3, the number of convolution kernels is 128, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects the maximum pooling, the pooling size is 2×2, and the step size is 2; the convolution kernel size of the void convolution pooling layer is 5×5, the number of convolution kernels is 256, the void rate is set to 2, the step size is 1, and the activation function uses , using batch normalization, the pooling layer selects average pooling, the pooling size is 3×3, and the stride is 1; the fully connected layer uses one layer of full connection, and the output is a 2048-dimensional vector; (3) The difference feature compensator sequentially includes an input layer, a feature difference calculation module, a compensator convolution module, and a compensation feature generation module; Among them, the feature difference calculation module is used to calculate the feature difference after the shape is changed; The compensator convolution module includes three convolution layers for feature extraction and transformation of feature differences. The first convolution layer has 256 input channels, 128 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. The second convolution layer has 128 input channels, 64 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. The third convolution layer has 64 input channels, 256 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. Each convolution layer is followed by batch normalization and Activation function; The compensation feature generation module generates compensation features by fusing features; (4) The first deep feature extractor includes the first convolutional layer, the second convolutional layer, the deep feature enhancer and the flattening layer. The first convolutional layer has 256 input channels, 128 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. It contains a maximum pooling layer with a window size of 2×2 and a stride of 2. The output spatial resolution becomes 32×32. The second convolutional layer has 128 input channels, 256 output channels, a convolution kernel size of 3×3, a stride of 1, and a padding of 1. It is followed by batch normalization and Activation function; the deep feature enhancer contains a global average pooling layer, which reduces the spatial resolution to 1×1 and retains only the global features of each channel; (5) The second deep feature extractor is the same as the first deep feature extractor and receives the compensation features, but the weights are different from those of the first deep model; (6) The deep decoder consists of the first deconvolution layer, the second deconvolution layer, the third deconvolution layer, Activation function and thresholding function, where the convolution kernel size of the first deconvolution layer is 4×4, the input channel is 128, the output channel is 64, the step size is 2, and the padding is 1; the convolution kernel size of the second deconvolution layer is 4×4, the input channel is 64, the output channel is 32, the step size is 2, and the padding is 1; the convolution kernel size of the third deconvolution layer is 4×4, the input channel is 32, the output channel is 1, the step size is 2, and the padding is 1.

4. The remote sensing image change detection method based on deep feature alignment according to claim 3 is characterized in that: S3 is as follows: The preprocessed dataset The complete remote sensing image at any two moments and Input to the remote sensing image change detection model, Indicates the A complete remote sensing image at a moment, Indicates the A complete remote sensing image at a moment, , , ; Input to the first reference encoder, Input into the second benchmark encoder, the two images are sequentially passed through the first convolution pooling layer, the second convolution pooling layer, the depthwise separable convolution pooling layer, and the fully connected layer in each benchmark encoder to obtain the first The benchmark encoding features at each moment Hedi Benchmark coding features at each moment ; The features and Input to the difference feature compensator, and pass through the input layer of the difference feature compensator Function changes characteristics and The shapes of and features , and then calculate the feature through the feature difference calculation module and features The difference between the characteristics , , and then the difference features The input is sent to the compensation convolution module through three convolution layers, and the compensation features are finally output. , the calculation formula is as follows: , in, and represents two different learnable parameters, ; Then, the compensation feature Input to the first deep feature extractor and pass through the first convolution layer to obtain ,Will Input to the second convolutional layer to get ,Will Input into the deep feature enhancer and flattening layer to obtain the first deep feature , similarly, the compensation feature Input to the second deep feature extractor and pass through the first convolution layer to obtain ,Will Input to the second convolutional layer to get ,Will Input to the deep feature enhancer and flattening layer to obtain the second deep feature ; Then the features and the first depth feature Add up the features , the features and the second deep feature Add up the features , and then the features and Input into the feature alignment function and calculate the alignment features , the specific calculation is as follows: , in, Used to limit the output range to , Represents the hyperparameter used to adjust the weight of the difference feature, setting , Represents the hyperparameter used to adjust the weight of collaborative features, setting , Indicates the smoothing factor, the default value is , represents element-by-element addition, represents element-wise multiplication, represents the two-norm normalization operation; Finally, the features ,feature and align features After adding, it is input into the depth decoder and passes through the first deconvolution layer to obtain ,Will Input to the second deconvolution layer to get ,Will Input to the third deconvolution layer and In the activation function, the change probability map is obtained , change probability map Each pixel of is between [0,1], and then the change probability map is thresholded. Each pixel value is thresholded to obtain the predicted change image , the calculation process is as follows: , in, Indicates the threshold value, set , represents the pixel position, express exist The pixel value at express exist The pixel value at , .

5. The remote sensing image change detection method based on deep feature alignment according to claim 4 is characterized in that: S4 is as follows: pass Loss Function Calculate the predicted change image and baseline change images The loss is calculated as follows: , in, Indicates the total number of pixels, express The index of Indicates the first pixels, Represents the predicted change in the image pixels.

6. A remote sensing image change detection system based on deep feature alignment, characterized by: Execute a remote sensing image change detection method based on deep feature alignment as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Remote sensing image change detection system and method based on depth separable convolution module

    CN116229283A

  • Remote sensing image change detection method based on hierarchical attention residual UNet + +

    CN116958800A

  • Dual-time remote sensing image semantic change detection method based on twin residual network

    CN118429819A