A Multi-Source High-Resolution Remote Sensing Image Change Detection Method and System

The modal conversion of remote sensing images is performed through the CycleGAN model, and features are extracted using weighted fusion and spatiotemporal attention mechanisms, which solves the problems of information loss and insufficient feature fusion in the prior art, and improves the efficiency and accuracy of remote sensing image change detection.

CN119360230BActive Publication Date: 2025-05-27CHINA COMM SERVICE APPL & SOLUTION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411907127.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-27
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

The existing multi-source high-resolution remote sensing image change detection technology has problems such as information loss, insufficient feature fusion, and high computing resource consumption, resulting in low detection efficiency and low accuracy.

Method used

Modal transformation is performed through the CycleGAN model, synthetic images with different modal styles are generated, and global features are extracted using weighted fusion and spatiotemporal attention mechanisms to form the final fusion features for change detection.

Benefits of technology

The effectiveness and accuracy of multimodal feature fusion is improved, the spatio-temporal modeling capabilities of change detection are enhanced, the detection efficiency and accuracy are improved, and the cost of computing resources is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360230B_ABST
    Figure CN119360230B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image data processing, and discloses a multi-source high-resolution remote sensing image change detection method and system, including obtaining remote sensing images; converting the remote sensing images into synthetic images; extracting the global features of the first image, the first synthetic image, the second image, and the second synthetic image; performing weighted fusion on the extracted global features to form the fusion features at the first moment and the fusion features at the second moment; and performing correlation modeling on the fusion features at the first moment and the fusion features at the second moment to extract the spatio-temporal relationship between the images of the ground objects at different time phases, generating the final fusion features at the first moment and the final fusion features at the second moment; detecting the image changes between the first moment and the second moment according to the final fusion features at the first moment and the final fusion features at the second moment. The present invention solves the problems of low efficiency and low accuracy in the existing remote sensing image change detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image data processing, and particularly relates to a method and system for multi-source high-resolution remote sensing image change detection. Background Technique

[0002] Remote sensing image change detection is an important task in remote sensing image processing research, aiming to identify and locate ground object changes through the comparison of multi-temporal images. Currently, the sources of remote sensing images provided on the Internet are very extensive, covering a variety of remote sensing platforms, such as optical satellites, synthetic aperture radar satellites, unmanned aerial vehicles, Internet map platforms Google Earth and Internet map platform OpenStreetMap, etc. The data information types of remote sensing images obtained by these platforms are diverse, including optical images, radar images, thermal infrared images and other modalities. The diversity of remote sensing image data information provides a richer and more comprehensive data source for change detection, making it possible to observe the target area from different perspectives, so as to more comprehensively and accurately reflect the changes of ground objects.

[0003] The existing techniques for multi-source high-resolution remote sensing image change detection mainly perform modality conversion on data information of different modalities based on a cyclic generative adversarial network model, converting different modalities into the same modality, and performing feature extraction and change detection in the same modality. However, the existing detection methods have problems such as information loss in modality conversion, insufficient feature fusion, and high consumption of computing resources, resulting in low detection efficiency and low accuracy. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for multi-source high-resolution remote sensing image change detection, so as to solve the problems of low efficiency and low accuracy in existing remote sensing image change detection.

[0005] The present invention is achieved by the following technical solutions:

[0006] A method for multi-source high-resolution remote sensing image change detection includes the following steps:

[0007] S1. Obtain remote sensing images, where the remote sensing images include a first image at a first moment and a second image at a second moment, the first image is an image of a first modality, the second image is an image of a second modality, and the first modality and the second modality are different modalities;

[0008] S2. Convert the remote sensing images into synthetic images, including: converting the first image into a first synthetic image with the style of the second modality, and converting the second image into a second synthetic image with the style of the first modality;

[0009] S3. Extract the global features of the first image, the first synthetic image, the second image, and the second synthetic image;

[0010] S4. Perform weighted fusion on the extracted global features to form the fused features at the first moment and the fused features at the second moment; and perform correlation modeling on the fused features at the first moment and the fused features at the second moment to extract the spatio-temporal relationship between the images of the ground objects at different time phases, and generate the final fused features at the first moment and the final fused features at the second moment;

[0011] S5. Detect the image changes between the first moment and the second moment according to the final fused features at the first moment and the final fused features at the second moment.

[0012] In some embodiments, the step of generating the final fused features at the first moment in step S4 includes:

[0013] S401. Respectively perform different linear transformations on the fused features at the first moment and the fused features at the second moment to obtain Q1 corresponding to the fused features at the first moment, K1 corresponding to the fused features at the second moment, and V1. The linear transformation formula is:

[0014] Q1 = W Q F combined,t1 ,K1 = W K F combined,t2 ,V1 = W V F combined,t2 ,

[0015] wherein, W Q 、W K 、W V all represent learnable linear transformation matrices for adjusting the dimensions of the features; F combined,t1 represents the fused features at the first moment, and F combined,t2 represents the fused features at the second moment; Q1 represents the query vector at each feature position in the first moment; K1 represents the key vector at each feature position in the second moment; V1 represents the weight value at each feature position in the second moment;

[0016] S402. Calculate the similarity at each feature position according to the query vector Q1 at each feature position in the first moment and the key vector K1 at each feature position in the second moment to obtain the first attention weight C1. The calculation formula of the first attention weight C1 is:

[0017] ,

[0018] wherein, d k represents the dimension of the feature for scaling the similarity score; K1 TDenote the transpose of K1, which represents the transposed form of the key vectors at each feature position in the second moment and is used to calculate the similarity between Q1 and K1; C1 represents the correlation between each feature position in the first moment and the second moment;

[0019] S403. Weighted sum the weight values V1 at each feature position in the second moment according to the first attention weight C1 to generate the weighted feature representation F1 of the first moment attended ;

[0020] S404. Fuse the fused feature F of the first moment combined,t1 and the weighted feature representation F1 of the first moment attended to generate the final fused feature F1 of the first moment final ; The fusion formula is:

[0021] F1 final =F combined,t1 +F1 attended 。

[0022] In some embodiments, the steps of generating the final fused feature of the second moment in step S4 include:

[0023] S411. Respectively perform different linear transformations on the fused feature of the first moment and the fused feature of the second moment to obtain Q2 corresponding to the fused feature of the second moment, K2 corresponding to the fused feature of the first moment, and V2. The linear transformation formula is:

[0024] Q2 = W Q F combined,t2 ,K2 = W K F combined,t1 ,V2 = W V F combined,t1 ,

[0025] where Q2 represents the query vector at each feature position in the second moment; K2 represents the key vector at each feature position in the first moment; V2 represents the weight value at each feature position in the first moment;

[0026] S412. Calculate the similarity at each feature position according to the query vector Q2 at each feature position in the second moment and the key vector K2 at each feature position in the first moment to obtain the second attention weight C2. The calculation formula of the second attention weight C2 is:

[0027] ,

[0028] where K2 TDenote the transpose of K2, which represents the transposed form of the key vector at each feature position in the first moment and is used to calculate the similarity between Q2 and K2; C2 represents the correlation between each feature position at the second moment and the first moment.

[0029] S413. Weighted sum the weight values V2 at each feature position in the first moment according to the second attention weight C2 to generate the weighted feature representation F2 of the second moment. attended ;

[0030] S414. Fuse the fused feature F of the second moment combined,t2 and the weighted feature representation F2 of the second moment attended to generate the final fused feature F2 of the second moment final ; The fusion formula is:

[0031] F2 final = F combined,t2 + F2 attended .

[0032] In some embodiments, step S2 includes:

[0033] Input the first image and the second image into the CycleGAN model, and through the generator G of the CycleGAN model 1 convert the first image into the first synthetic image with the style of the second modality, and at the same time through the generator G of the CycleGAN model 2 convert the second image into the second synthetic image with the style of the first modality;

[0034] Through the discriminator D of the CycleGAN model 1 judge the difference in the modality features between the first synthetic image and the second image, and feedback it to the generator G 1 , and the generator G 1 improves according to the difference information fed back by the discriminator D 1 so that the generator G 1 can output the first synthetic image closer to the style of the second modality; At the same time, through the discriminator D of the CycleGAN model 2 judge the difference in the modality features between the second synthetic image and the first image, and feedback it to the generator G 2 , and the generator G 2 improves according to the difference information fed back by the discriminator D 2 so that the generator G 2 can output the second synthetic image closer to the style of the first modality.

[0035] In some embodiments, step S2 further includes a step of performing edge-preserving loss processing on the first synthetic image and the second synthetic image, including:

[0036] Calculate the edge differences between the first image, the second image, the first synthesized image, and the second synthesized image respectively by the Sobel operator, and feedback them to the generator G of the CycleGAN model 1 and the generator G 2 , the generator G 1 and the generator G 2 Improve according to the feedback edge difference information, so that the generator G 1 can output the first synthesized image retaining the edge features of the first image, and the generator G 2 can output the second synthesized image retaining the edge features of the second image. Use the improved generator G 1 to output the first synthesized image retaining the edge features of the first image according to the first image, and use the improved generator G 2 to output the second synthesized image retaining the edge features of the second image according to the second image; the Sobel operator is:

[0037] L edge =E A [||Edge(A)-Edge(G 1 (A))||1]+E B [||Edge(B)-Edge(G 2 (B))||1],

[0038] wherein, L edge represents the edge difference information between the first image and the second image and the first synthesized image and the second synthesized image, and E A represents the edge loss value of the first image, and E B represents the edge loss value of the second image. A represents the first image, B represents the second image, and G 1 (A) represents the first synthesized image, and G 2 (B) represents the second synthesized image.

[0039] In some embodiments, step S2 further includes a step of performing spatial consistency loss processing on the remote sensing image and the corresponding synthesized image, including:

[0040] Calculate the spatial difference between the remote sensing image and the corresponding synthesized image by the spatial consistency loss formula, and feedback it to the corresponding generator in the CycleGAN model. The generator improves according to the feedback spatial difference information, so that the generator can output the synthesized image retaining the spatial structure of the remote sensing image. Use the improved generator to output the synthesized image retaining the spatial structure of the remote sensing image according to the remote sensing image; the spatial consistency loss formula is:

[0041] L spatial =E x[||A spatial -Gi(X) spatial || 2 ,

[0042] Among them, L spatial represents the spatial difference information between the remote sensing image and the corresponding synthetic image, E x represents the edge loss value of the remote sensing image, A spatial represents the spatial feature representation of the remote sensing image, and Gi(X) spatial represents the spatial feature representation of the synthetic image corresponding to the remote sensing image.

[0043] In some embodiments, step S3 includes:

[0044] Using convolution kernels of different sizes of a multi-scale convolutional neural network to extract the global features of the first image, the first synthetic image, the second image, and the second synthetic image.

[0045] In some embodiments, step S3 further includes performing channel attention mechanism processing on the global feature maps generated when extracting global features, and using a multi-scale convolutional neural network to perform feature extraction on the processed global feature maps to generate global features;

[0046] The step of performing channel attention mechanism processing on the global feature maps includes:

[0047] When extracting the global features of the first image, the first synthetic image, the second image, and the second synthetic image, global feature maps of different sizes are generated. Global average pooling is performed on the global feature maps to obtain the global statistical information of each channel in the global feature maps. A fully connected network is used to calculate the weights of the importance of each channel according to the global statistical information of each channel, and the weights are applied to each channel of the global feature maps.

[0048] In some embodiments, step S5 includes:

[0049] According to the final fusion feature F1 at the first moment final and the final fusion feature F2 at the second moment final perform element-wise subtraction and then take the absolute value to obtain the differential feature representation D. The calculation formula of the differential feature representation D is:

[0050] D = |F1 final - F2 final |;

[0051] Input the differential feature representation D into a classifier to detect the image changes between the first moment and the second moment, and generate a change map.

[0052] The present invention also relates to a multi-source high-resolution remote sensing image change detection system based on the above multi-source high-resolution remote sensing image change detection method, including:

[0053] A modality conversion module, configured to convert the first image into a first synthesized image with the style of the second modality, and convert the second image into a second synthesized image with the style of the first modality. The first image is a first modality image corresponding to the first moment, the second image is a second modality image corresponding to the second moment, and the first modality and the second modality are different modalities;

[0054] A feature extraction module, configured to extract the global features of the first image, the first synthesized image, the second image, and the second synthesized image;

[0055] A feature fusion module, configured to perform weighted fusion on the extracted global features to form the fusion features of the first moment and the fusion features of the second moment; and perform correlation modeling on the fusion features of the first moment and the fusion features of the second moment to extract the spatio-temporal relationship between the images of the ground objects at different time phases, and generate the final fusion features of the first moment and the final fusion features of the second moment;

[0056] A change detection module, configured to detect the image change between the first moment and the second moment according to the final fusion features of the first moment and the final fusion features of the second moment.

[0057] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0058] 1. Aiming at the problems that the existing fusion methods cannot fully consider the information differences between different modalities, ignore the importance of weights between different modality features, and cannot make full use of complementary information from different data sources, resulting in the lack of sufficient discriminative power of the fused features in change detection, making the influence of some modalities on the final decision too large, and thus reducing the accuracy of change detection, the present invention adopts weighted fusion, which can dynamically adjust the fusion weights according to the feature importance of each modality, thereby improving the effectiveness and accuracy of multi-modal feature fusion; for remote sensing images of different time phases and different modalities, the existing fusion methods fail to fully consider the spatio-temporal correlation between different time phases and ignore the influence of the time dimension, i.e., the phase change, treating remote sensing images at different time points as independent inputs, thus reducing the accuracy of change detection. The present invention adopts associative modeling of the fused features at the first moment and the fused features at the second moment to extract the spatio-temporal relationship between the images of the same ground object at different time phases, which can comprehensively consider the change information in the time dimension, enhance the spatio-temporal modeling ability of change detection, better understand the relationship between the first moment and the second moment in change detection, effectively capture the spatio-temporal relationship between different time phases, more comprehensively capture the change information, enable the full fusion of multi-modal features, and thus improve the efficiency and accuracy of change detection and reduce the computational resource cost.

[0059] 2. The edge-preserving loss method of the CycleGAN model improves the generator through the Sobel operator, which can effectively retain the edge information and detail information of the image during the modality conversion, making the output synthetic image retain the edge features of the remote sensing image, ensuring that the edge features of the remote sensing image and the synthetic image are as consistent as possible, preventing the details from being blurred after the conversion of the high-resolution image, and thus improving the accuracy of change detection.

[0060] 3. The spatial consistency loss method of the CycleGAN model is used to improve the generator, making the output synthetic image retain the spatial structure of the corresponding remote sensing image, solving the problem of geometric distortion of the remote sensing image during the modality conversion of the high-resolution remote sensing image, ensuring that the object positions and geometric shapes in the remote sensing image remain unchanged before and after the conversion, guaranteeing the consistency of the spatial structure between the converted synthetic image and the remote sensing image, reducing the geometric distortion, and ensuring the accuracy of the object positions and structures in change detection, thereby improving the accuracy of change detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0062] Figure 1 This is the flowchart of the existing remote sensing image change detection method in the implementation of the present invention.

[0063] Figure 2 This is the flowchart of the remote sensing image change detection method in the implementation of the present invention.

[0064] Figure 3 This is the block diagram of the remote sensing image change detection system in the implementation of the present invention. Detailed implementation manners

[0065] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention.

[0066] In the current change detection task for multi-source high-resolution remote sensing images, the change detection method is based on the CycleGAN model. CycleGAN can perform modality conversion on data information of different modalities to convert them into the same modality. For example, if it is necessary to perform change detection between optical images obtained from multi-source Internet platforms and SAR images of Synthetic Aperture Radar, it is very difficult to directly perform change detection on optical images and SAR images from different remote sensing platforms because the modality characteristics of these images are quite different. However, through CycleGAN, SAR images can be converted into optical-style images, or optical images can be converted into SAR-style images, so as to perform feature extraction and change detection in the same modality. This modality conversion greatly reduces the impact of modality differences on the change detection task, making the feature extraction and detection results more consistent and reliable.

[0067] Taking the optical images and SAR images in the high-resolution image data information of multi-source Internet platforms as an example, referring to Figure 1 , the flow of the existing change detection method includes:

[0068] 1) Multi-modal image preprocessing: First, the optical images and SAR images are georegistered and geometrically corrected, and then the optical images and SAR images are cropped into image blocks of a preset size, such as 256×256.

[0069] 2) Modal conversion: The preprocessed optical images and SAR images are input into the CycleGAN modal conversion module. CycleGAN is used to perform modal conversion on data information of different modalities to convert optical images and SAR images into the same modality. Specifically, CycleGAN contains two generators and two discriminators. The generators are responsible for converting images of one modality into images of another modality, that is, converting optical images into SAR-style images or converting SAR images into optical-style images. The discriminators are used to determine whether the converted images are consistent with the target modality, thereby continuously improving the conversion effect of the generators to reduce the differences between modalities.

[0070] 3) Feature extraction: The images after modal conversion are input into a convolutional neural network to extract representative spatial features. The convolutional neural network can capture and extract features such as edges, textures, and local structures in the images through multiple layers of convolution, providing a basis for subsequent change detection.

[0071] 4) Feature fusion: For the features after different modal conversions, they are fused by simple concatenation or superposition to form a joint feature representation. This fusion method aims to combine information of different modalities so that subsequent change detection can utilize these features.

[0072] 5) Change detection: The method of feature difference is used to generate a change map, and the change areas of ground objects are identified by comparing the image features at different times.

[0073] 6) Generation of change detection map: Calculate the difference between the joint feature representations of two time phases, such as the feature difference between time points t1 and t2. Usually, the differential feature is obtained by element-wise subtraction and then taking the absolute value. Then, based on the differential result of the features, the threshold method or classifier is used to determine which areas have undergone significant changes, and finally, a change map is generated.

[0074] Existing technical solutions use the CycleGAN (Cyclic Generative Adversarial Network) model for modal conversion to eliminate differences between different modalities, such as the differences between optical images and SAR images. However, due to the conversion process of CycleGAN relying on adversarial generation, the converted images often have the problem of information loss. Especially in high-resolution images, the detailed features after conversion cannot be fully retained. This information loss will directly affect the accuracy of subsequent change detection and reduce the accuracy of distinguishing between changed areas and unchanged areas.

[0075] In the prior art, the fusion of multi-modal features is mainly achieved by simple splicing or superposition. This simple fusion method ignores the importance of weights between different modal features, fails to fully utilize the complementary information from different data sources, and results in the fused features lacking sufficient discriminative power in the change detection task. This simple fusion method also fails to fully consider the spatio-temporal correlation between different time phases, limiting the performance of multi-modal features in change detection.

[0076] The above method also has the problem of high computational resource consumption.

[0077] Embodiment 1

[0078] A method for change detection of multi-source high-resolution remote sensing images, referring to Figure 2 , includes the following steps:

[0079] S1. Obtain remote sensing images, where the remote sensing images include a first image at a first moment and a second image at a second moment. The first image is an image of a first modality, the second image is an image of a second modality, and the first modality and the second modality are different modalities;

[0080] Obtain high-resolution remote sensing images from a multi-source Internet platform, and perform preprocessing on the obtained remote sensing images, including: performing georegistration and geometric correction processing on the first image and the second image, and then cropping the first image and the second image into images with a size of 256×256.

[0081] S2. Convert the preprocessed remote sensing images into composite images, including: converting the first image into a first composite image with the style of the second modality, and converting the second image into a second composite image with the style of the first modality;

[0082] S21. Input the first image and the second image into the CycleGAN model, and through the generator G of the CycleGAN model 1 convert the first image into a first composite image with the style of the second modality, and at the same time, through the generator G of the CycleGAN model 2 convert the second image into a second composite image with the style of the first modality.

[0083] S22. Through the discriminator D of the CycleGAN model 1 judge the difference in modal features between the first composite image and the second image, and feedback it to the generator G 1 , and the generator G 1 improves according to the difference information fed back by the discriminator D 1 so that the generator G 1 can output a first composite image closer to the style of the second modality; at the same time, through the discriminator D of the CycleGAN model 2Judge the differences in modal features between the second synthetic image and the first image, and feed them back to the generator G 2 , the generator G 2 Improve according to the difference information fed back by the discriminator D 2 so that the generator G 2 can output a second synthetic image that is closer to the style of the first modality.

[0084] S23. Perform edge-preserving loss processing on the first synthetic image and the second synthetic image.

[0085] Including: Calculate the edge differences between the first image, the second image, the first synthetic image, and the second synthetic image respectively through the Sobel operator, and feed them back to the generator G of the CycleGAN model 1 and the generator G 2 , the generator G 1 and the generator G 2 Improve according to the fed-back edge difference information, so that the generator G 1 can output a first synthetic image that retains the edge features of the first image, and the generator G 2 can output a second synthetic image that retains the edge features of the second image. Use the improved generator G 1 to output a first synthetic image that retains the edge features of the first image according to the first image, and use the improved generator G 2 to output a second synthetic image that retains the edge features of the second image according to the second image; the Sobel operator is:

[0086] L edge =E A [||Edge(A)-Edge(G 1 (A))||1]+E B [||Edge(B)-Edge(G 2 (B))||1],

[0087] where L edge represents the edge difference information between the first image and the second image and the first synthetic image and the second synthetic image, and E A represents the edge loss value of the first image, and E B represents the edge loss value of the second image. A represents the first image, B represents the second image, and G 1 (A) represents the first synthetic image, and G 2 (B) represents the second synthetic image.

[0088] The edge-preserving loss method using the CycleGAN model improves the generator through the Sobel operator, which can effectively retain the edge information and detail information of the image during the modality conversion process, enabling the output synthetic image to retain the edge features of the remote sensing image, ensuring that the edge features of the remote sensing image and the synthetic image are as consistent as possible, preventing the details of the high-resolution image from being blurred after conversion, and thus improving the accuracy of change detection.

[0089] S24. Perform spatial consistency loss processing on the remote sensing image and the corresponding synthetic image.

[0090] This includes: calculating the spatial difference between the remote sensing image and the corresponding synthetic image through the spatial consistency loss formula and feeding it back to the corresponding generator in the CycleGAN model. The generator is improved according to the fed-back spatial difference information, enabling the generator to output a synthetic image that retains the spatial structure of the remote sensing image. The improved generator is used to output a synthetic image that retains the spatial structure of the remote sensing image according to the remote sensing image; the spatial consistency loss formula is:

[0091] L spatial =E x [||A spatial -Gi(X) spatial || 2 ,

[0092] where L spatial represents the spatial difference information between the remote sensing image and the corresponding synthetic image, E x represents the edge loss value of the remote sensing image, A spatial represents the spatial feature representation of the remote sensing image, and Gi(X) spatial represents the spatial feature representation of the synthetic image corresponding to the remote sensing image.

[0093] The spatial consistency loss method using the CycleGAN model improves the generator, enabling the output synthetic image to retain the spatial structure of the remote sensing image corresponding to the synthetic image, solving the problem of geometric distortion of the remote sensing image during the modality conversion process of high-resolution remote sensing images, ensuring that the object positions and geometric shapes in the remote sensing image remain unchanged before and after conversion, guaranteeing the consistency of the spatial structure between the converted synthetic image and the remote sensing image, reducing the geometric shape distortion, ensuring the accuracy of the object positions and structures in change detection, and thus improving the accuracy of change detection.

[0094] S3. Extract the global features of the first image, the first synthetic image, the second image, and the second synthetic image;

[0095] S31. Use convolutional kernels of different sizes in the multi-scale convolutional neural network to extract the global features of the first image, the first synthetic image, the second image, and the second synthetic image.

[0096] The convolutional kernels include 3x3, 5x5, 7x7, etc.

[0097] S32. Perform channel attention mechanism processing on the global feature map generated when extracting global features, and use a multi-scale convolutional neural network to extract features from the processed global feature map to generate global features;

[0098] The steps for performing channel attention mechanism processing on the global feature map include:

[0099] When extracting the global features of the first image, the first synthetic image, the second image, and the second synthetic image, use convolutional kernels of different sizes in the multi-scale convolutional neural network for convolutional processing to capture information at different scales in the images, generate global feature maps of different sizes, perform global average pooling on the global feature maps to obtain the global statistical information of each channel in the global feature maps, use a fully connected network to calculate the weights of the importance of each channel according to the global statistical information of each channel, and apply the weights to each channel of the global feature map, so that important features in the global feature map can be amplified and unimportant features in the global feature map can be suppressed.

[0100] Through the channel attention mechanism, the multi-scale convolutional neural network can learn that certain global features are the most important for change detection, thereby enhancing these global features and improving the expression ability of the global features and the distinguishability of the change regions.

[0101] Use convolutional kernels of different sizes in the multi-scale convolutional neural network to extract global features at different scales, and introduce a channel attention mechanism to improve the expression ability of the global features and the distinguishability of the change regions, so that when the multi-scale convolutional neural network faces multi-source images, the accuracy and robustness of change detection are improved.

[0102] S4. Perform weighted fusion on the extracted global features to form the fusion features at the first moment and the fusion features at the second moment; and perform correlation modeling on the fusion features at the first moment and the fusion features at the second moment to extract the spatio-temporal relationship between the images of the ground objects at different time phases, and generate the final fusion features at the first moment and the final fusion features at the second moment;

[0103] S41. Perform weight assignment according to the global features of the first image, the first synthetic image, the second image, and the second synthetic image, and perform weighted summation according to the assigned weights to generate the fusion features at the first moment and the fusion features at the second moment, realizing multi-modal feature fusion.

[0104] Aiming at the problem that the existing fusion methods cannot fully consider the information differences between different modalities, ignore the importance of weights between different modality features, and cannot make full use of complementary information from different data sources, resulting in the fused features lacking sufficient discriminative power in change detection, making the influence of certain modalities on the final decision too large, and thus reducing the accuracy of change detection, the present invention adopts weighted fusion, which can dynamically adjust the fusion weights according to the feature importance of each modality, thereby improving the effectiveness and accuracy of multi-modal feature fusion.

[0105] S42. Use a spatio-temporal attention mechanism to perform correlation modeling on the fused features at the first moment and the fused features at the second moment, extract the spatio-temporal relationship between the images of the ground objects at different time phases, and generate the final fused features at the first moment and the final fused features at the second moment.

[0106] Among them, the steps of generating the final fused features at the first moment include:

[0107] S401. Respectively perform different linear transformations on the fused features at the first moment and the fused features at the second moment to obtain Q1 corresponding to the fused features at the first moment, K1 corresponding to the fused features at the second moment, and V1. The linear transformation formula is:

[0108] Q1 = W Q F combined,t1 ,K1 = W K F combined,t2 ,V1 = W V F combined,t2 ,

[0109] Among them, W Q 、W K 、W V all represent learnable linear transformation matrices for adjusting the dimensions of features; F combined,t1 represents the fused features at the first moment, and F combined,t2 represents the fused features at the second moment; Q1 represents the query vector at each feature position in the first moment; K1 represents the key vector at each feature position in the second moment; V1 represents the weight value at each feature position in the second moment;

[0110] S402. Calculate the similarity at each feature position according to the query vector Q1 at each feature position in the first moment and the key vector K1 at each feature position in the second moment to obtain the first attention weight C1. The calculation formula of the first attention weight C1 is:

[0111] ,

[0112] Among them, d kRepresents the dimension of the feature, used to scale the similarity score to ensure numerical stability; K1 T Represents the transpose of K1, which represents the transposed form of the key vector at each feature position in the second moment and is used to calculate the similarity between Q1 and K1; C1 represents the correlation between each feature position in the first moment and the second moment;

[0113] S403. Weightedly sum the weight values V1 at each feature position in the second moment according to the first attention weight C1 to generate the weighted feature representation F1 of the first moment attended ;

[0114] S404. Fuse the fused feature F of the first moment combined,t1 and the weighted feature representation F1 of the first moment attended to generate the final fused feature F1 of the first moment final ; The fusion formula is:

[0115] F1 final =F combined,t1 +F1 attended .

[0116] The steps to generate the final fused feature of the second moment include:

[0117] S411. Respectively perform different linear transformations on the fused feature of the first moment and the fused feature of the second moment to obtain Q2 corresponding to the fused feature of the second moment, K2 corresponding to the fused feature of the first moment, and V2, and the linear transformation formula is:

[0118] Q2 = W Q F combined,t2 , K2 = W K F combined,t1 , V2 = W V F combined,t1 ,

[0119] where, Q2 represents the query vector at each feature position in the second moment; K2 represents the key vector at each feature position in the first moment; V2 represents the weight value at each feature position in the first moment;

[0120] S412. Calculate the similarity at each feature position according to the query vector Q2 at each feature position in the second moment and the key vector K2 at each feature position in the first moment to obtain the second attention weight C2, and the calculation formula of the second attention weight C2 is:

[0121] ,

[0122] where, K2 TDenotes the transpose of K2, representing the transposed form of the key vectors at each feature position in the first moment, which is used to calculate the similarity between Q2 and K2; C2 represents the correlation between each feature position between the second moment and the first moment;

[0123] S413. Weighted sum the weight values V2 at each feature position in the first moment according to the second attention weight C2 to generate the weighted feature representation F2 of the second moment attended ;

[0124] S414. Fuse the fused feature F of the second moment combined,t2 and the weighted feature representation F2 of the second moment attended to perform fusion processing to generate the final fused feature F2 of the second moment final ; The fusion formula is:

[0125] F2 final = F combined,t2 + F2 attended 。

[0126] For remote sensing images of different time phases and different modalities, the existing fusion methods fail to fully consider the spatio-temporal correlation between different time phases, ignore the influence of the time dimension, i.e., the phase change, and regard the remote sensing images at different time points as independent inputs, thus reducing the accuracy of change detection. The present invention uses a spatio-temporal attention mechanism to perform correlation modeling on the fused feature of the first moment and the fused feature of the second moment, extracts the spatio-temporal relationship between the images of ground objects at different time phases, can comprehensively consider the change information in the time dimension, enhances the spatio-temporal modeling ability of change detection, can better understand the relationship between the first moment and the second moment in change detection, effectively capture the spatio-temporal relationship between different time phases, more comprehensively capture the change information, enable the multi-modal features to be fully fused, and thus improve the efficiency and accuracy of change detection and reduce the computational resource cost.

[0127] S5. Detect the image change between the first moment and the second moment according to the final fused feature of the first moment and the final fused feature of the second moment.

[0128] According to the final fused feature F1 of the first moment final and the final fused feature F2 of the second moment final perform element-wise subtraction and then take the absolute value to obtain the differential feature representation D. The calculation formula of the differential feature representation D is:

[0129] D = |F1 final - F2 final |;

[0130] Input the differential feature representation D into a classifier. The classifier detects the image changes between the first moment and the second moment, and generates a change map by marking the change regions obtained through change determination.

[0131] The present invention also relates to a multi-source high-resolution remote sensing image change detection system based on the above multi-source high-resolution remote sensing image change detection method. Refer to Figure 3 , including:

[0132] A modality conversion module for converting the first image into a first synthesized image with the style of the second modality, and converting the second image into a second synthesized image with the style of the first modality. The first image is the first modality image corresponding to the first moment, the second image is the second modality image corresponding to the second moment, and the first modality and the second modality are different modalities;

[0133] A feature extraction module for extracting the global features of the first image, the first synthesized image, the second image, and the second synthesized image;

[0134] A feature fusion module for performing weighted fusion on the extracted global features to form the fusion features of the first moment and the fusion features of the second moment; and performing correlation modeling on the fusion features of the first moment and the fusion features of the second moment to extract the spatio-temporal relationship between the images of the ground objects at different time phases, and generating the final fusion features of the first moment and the final fusion features of the second moment;

[0135] A change detection module for detecting the image changes between the first moment and the second moment according to the final fusion features of the first moment and the final fusion features of the second moment.

[0136] As described above, it is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Any simple modification or equivalent change made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the present invention.

Claims

1. A multi-source high-resolution remote sensing image change detection method, characterized in that: The following steps are involved: S1. Acquire a remote sensing image, wherein the remote sensing image includes a first image at a first moment and a second image at a second moment, wherein the first image is an image of a first modality, and the second image is an image of a second modality, and the first modality and the second modality are different modalities; S2, converting the remote sensing image into a synthetic image, including: converting the first image into a first synthetic image having a second modality style, and converting the second image into a second synthetic image having the first modality style; S3, extracting global features of the first image, the first synthesized image, the second image, and the second synthesized image; S4, weighting the extracted global features of the first image, the first synthetic image, the second image, and the second synthetic image, and performing weighted summation according to the assigned weights to generate fusion features at the first moment and fusion features at the second moment; and correlating the fusion features at the first moment and the fusion features at the second moment to model, extract the spatiotemporal relationship between the images of the ground object at different time phases, generate the final fusion features at the first moment and the final fusion features at the second moment, and realize multimodal feature fusion; S5. Detect image changes between the first moment and the second moment according to the final fusion features at the first moment and the final fusion features at the second moment.

2. The method for detecting changes in multi-source high-resolution remote sensing images according to claim 1, characterized in that: The step of generating the final fusion feature at the first moment in step S4 includes: S401, subjecting the fused features at the first moment and the fused features at the second moment to different linear transformations, respectively, to obtain Q1 corresponding to the fused features at the first moment, and K1 and V1 corresponding to the fused features at the second moment. The linear transformation formula is: Q1=W Q F combined,t1 ,K1=W K F combined,t2 ,V1=W V F combined,t2 , Among them, W Q , W K , W V Both represent learnable linear transformation matrices, which are used to adjust the dimension of features; F combined,t1 represents the fusion feature at the first moment, F combined,t2 represents the fused features at the second moment; Q1 represents the query vector of each feature position at the first moment; K1 represents the key vector of each feature position at the second moment; V1 represents the weight value of each feature position at the second moment; S402, calculating the similarity of each feature position according to the query vector Q1 of each feature position at the first moment and the key vector K1 of each feature position at the second moment to obtain a first attention weight C1, and the calculation formula of the first attention weight C1 is: , Among them, d k Represents the dimension of the feature, used to scale the similarity score; K1 T represents the transpose of K1, which represents the transposed form of the key vector of each feature position at the second moment, and is used to calculate the similarity between Q1 and K1; C1 represents the correlation between each feature position at the first moment and the second moment; S403: Perform weighted summation on the weight value V1 of each feature position at the second moment according to the first attention weight C1 to generate a weighted feature representation F1 at the first moment. attended ; S404, the fusion feature F at the first moment combined,t1 And the weighted feature representation F1 at the first moment attended Perform fusion processing to generate the final fusion feature F1 at the first moment final ; The fusion formula is: F1 final =F combined,t1 +F1 attended 。 3. The method for detecting changes in multi-source high-resolution remote sensing images according to claim 2, characterized in that: The step of generating the final fusion feature at the second moment in step S4 includes: S411, subjecting the fused features at the first moment and the fused features at the second moment to different linear transformations, respectively, to obtain Q2 corresponding to the fused features at the second moment, and K2 and V2 corresponding to the fused features at the first moment. The linear transformation formula is: <h2 style=";text-align:left;direction:ltr">Q2=W<h2 style=";text-align:left;direction:ltr"> Q <h2 style=";text-align:left;direction:ltr"> F<h2 style=";text-align:left;direction:ltr"> combined,t2 <h2 style=";text-align:left;direction:ltr"> K2=W<h2 style=";text-align:left;direction:ltr"> K <h2 style=";text-align:left;direction:ltr"> F<h2 style=";text-align:left;direction:ltr"> combined,t1 <h2 style=";text-align:left;direction:ltr"> V2=W<h2 style=";text-align:left;direction:ltr"> V <h2 style=";text-align:left;direction:ltr"> F<h2 style=";text-align:left;direction:ltr"> combined,t1 <h2 style=";text-align:left;direction:ltr"> , Wherein, Q2 represents the query vector of each feature position at the second moment; K2 represents the key vector of each feature position at the first moment; V2 represents the weight value of each feature position at the first moment; S412: Calculate the similarity of each feature position according to the query vector Q2 of each feature position at the second moment and the key vector K2 of each feature position at the first moment to obtain a second attention weight C2. The calculation formula of the second attention weight C2 is: , Among them, K2 T represents the transpose of K2, which represents the transposed form of the key vector of each feature position at the first moment, and is used to calculate the similarity between Q2 and K2; C2 represents the correlation between each feature position at the second moment and the first moment; S413: Perform weighted summation on the weight value V2 of each feature position at the first moment according to the second attention weight C2 to generate a weighted feature representation F2 at the second moment. attended ; S414, the fusion feature F at the second moment combined,t2 And the weighted feature representation F2 at the second moment attended Perform fusion processing to generate the final fusion feature F2 at the second moment final ; The fusion formula is: <h2 style=";text-align:left;direction:ltr">F2<h2 style=";text-align:left;direction:ltr"> final <h2 style=";text-align:left;direction:ltr"> =F<h2 style=";text-align:left;direction:ltr"> combined,t2 <h2 style=";text-align:left;direction:ltr"> +F2<h2 style=";text-align:left;direction:ltr"> attended <h2 style=";text-align:left;direction:ltr"> 。 4. The method for detecting changes in multi-source high-resolution remote sensing images according to claim 1, characterized in that: Step S2 includes: Inputting the first image and the second image into the CycleGAN model, converting the first image into a first synthetic image having the style of the second modality through the generator G1 of the CycleGAN model, and converting the second image into a second synthetic image having the style of the first modality through the generator G2 of the CycleGAN model; The discriminator D1 of the CycleGAN model determines the difference in modal features between the first synthetic image and the second image, and feeds back to the generator G1. The generator G1 is improved according to the difference information fed back by the discriminator D1, so that the generator G1 can output the first synthetic image that is closer to the second modal style. At the same time, the discriminator D2 of the CycleGAN model determines the difference in modal features between the second synthetic image and the first image, and feeds back to the generator G2. The generator G2 is improved according to the difference information fed back by the discriminator D2, so that the generator G2 can output the second synthetic image that is closer to the first modal style.

5. The method for detecting changes in multi-source high-resolution remote sensing images according to claim 4, characterized in that: Step S2 also includes the step of performing edge preservation loss processing on the first synthetic image and the second synthetic image, including: The edge differences between the first image, the second image and the first synthetic image and the second synthetic image are calculated respectively by the Sobel operator and fed back to the generator G1 and the generator G2 of the CycleGAN model. The generator G1 and the generator G2 are improved according to the fed-back edge difference information so that the generator G1 can output the first synthetic image retaining the edge features of the first image, and the generator G2 can output the second synthetic image retaining the edge features of the second image. The improved generator G1 outputs the first synthetic image retaining the edge features of the first image according to the first image, and the improved generator G2 outputs the second synthetic image retaining the edge features of the second image according to the second image.

6. The method for detecting changes in multi-source high-resolution remote sensing images according to claim 4, characterized in that: Step S2 also includes the step of performing spatial consistency loss processing on the remote sensing image and the corresponding synthetic image, including: The spatial difference between the remote sensing image and the corresponding synthetic image is calculated by the spatial consistency loss formula and fed back to the corresponding generator in the CycleGAN model. The generator is improved according to the fed-back spatial difference information so that the generator can output a synthetic image that retains the spatial structure of the remote sensing image. The improved generator is used to output a synthetic image that retains the spatial structure of the remote sensing image based on the remote sensing image. The spatial consistency loss formula is: L spatial =E x [||A spatial -Gi(X) spatial ||2], Among them, L spatial Represents the spatial difference information between the remote sensing image and the corresponding synthetic image, E x Represents the edge loss value of the remote sensing image, A spatial Represents the spatial feature representation of remote sensing images, Gi(X) spatial Representation of spatial features of synthetic images corresponding to remote sensing images.

7. The method for detecting changes in multi-source high-resolution remote sensing images according to claim 1, characterized in that: Step S3 includes: Convolution kernels of different sizes of a multi-scale convolutional neural network are used to extract global features of the first image, the first synthesized image, the second image, and the second synthesized image.

8. The method for detecting changes in multi-source high-resolution remote sensing images according to claim 7, characterized in that: Step S3 also includes performing a channel attention mechanism on the global feature map generated when extracting the global feature, and using a multi-scale convolutional neural network to extract features from the processed global feature map to generate global features; The steps of processing the channel attention mechanism on the global feature map include: When extracting the global features of the first image, the first synthetic image, the second image, and the second synthetic image, global feature maps of different sizes are generated, and global average pooling is performed on the global feature maps to obtain the global statistical information of each channel in the global feature maps. A fully connected network is used to calculate the weight of the importance of each channel according to the global statistical information of each channel, and the weight is applied to each channel of the global feature map.

9. The method for detecting changes in multi-source high-resolution remote sensing images according to claim 1, characterized in that: Step S5 includes: According to the final fusion feature F1 at the first moment final And the final fusion feature F2 at the second moment final After performing element-by-element subtraction, the absolute value is taken to obtain the differential feature representation D. The calculation formula of the differential feature representation D is: D=|F1 final - F2 final |; The differential feature representation D is input into the classifier to detect the image changes between the first moment and the second moment and generate a change map.

10. A multi-source high-resolution remote sensing image change detection system based on the multi-source high-resolution remote sensing image change detection method according to any one of claims 1 to 9, characterized in that: include: a modality conversion module, configured to convert a first image into a first synthetic image having a second modality style, and convert a second image into a second synthetic image having a first modality style, wherein the first image is a first modality image corresponding to a first moment, the second image is a second modality image corresponding to a second moment, and the first modality and the second modality are different modalities; A feature extraction module, used for extracting global features of the first image, the first synthetic image, the second image, and the second synthetic image; A feature fusion module, used to assign weights to the global features of the extracted first image, the first synthetic image, the second image, and the second synthetic image, and to perform weighted summation according to the assigned weights to generate a fusion feature at the first moment and a fusion feature at the second moment; The fusion features at the first moment and the fusion features at the second moment are associated and modeled, the spatiotemporal relationship between the images of the objects at different time phases is extracted, the final fusion features at the first moment and the final fusion features at the second moment are generated, and multi-modal feature fusion is realized; The change detection module is used to detect the image change between the first moment and the second moment based on the final fusion feature of the first moment and the final fusion feature of the second moment.

Citation Information

Patent Citations

  • Multi-modal remote sensing image change detection method, model generation method and terminal equipment

    CN113298056A

  • Ground feature change detection method combining optical and polarimetric SAR (Synthetic Aperture Radar) data

    CN117437556A