A multi-modal remote sensing data change detection method

CN117523401BActive Publication Date: 2026-09-25FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311596046.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2026-09-25
Estimated Expiration
2043-11-28

AI Technical Summary

Technical Problem

[0004]有鉴于此,本发明的目的在于提供一种多模态遥感数据变化检测方法,克服了现有方法上不同模态影像之间特征交流不充分,双时态域之间差异大,无法充分学习到不同模态影像特征之间的联系,导致变化信息不突出易出现误检和漏检的问题,为灾害发生前后无法获取同类型影像的变化信息快速识别提供技术支撑

Benefits of technology

[0040]与现有技术相比,本发明具有以下有益效果:(1)运用通道特征交换和注意力机制,通过相互学习感知不同模态特征的上下文信息,让两个模态特征之间的分布更加相似,实现双时态域之间的自动域适应,解决了传统方法特征交流不充分、特征之间无法比较的问题;(2)变化信息增强模块和多尺度特征融合解码器,对不同层级交互后的影像特征进行变化区域的增强识别,并将低层的变化信息特征与高层的变化信息特征融合在一起,克服传统方法在识别变化区域时易出现漏检和误检的问题,为利用多模态影像进行变化区域的识别提供技术支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117523401B_ABST
    Figure CN117523401B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of multi-modal remote sensing data change detection method.The present application constructs multi-modal feature interaction module, and let model automatically learn the information between different modal images.The present application faces optical and SAR (Synthetic Aperture Radar) two modal remote sensing data, utilizes deep learning method, adopts pseudo twin convolutional neural network CNN to extract the deep feature of different modal images respectively, establishes cross-modal feature interaction module, change information enhancement module and multi-scale fusion decoder, constructs a kind of change detection model for multi-modal remote sensing image, realizes the automatic detection of the change region of two different modal images, solves the problem that same type image in different periods is difficult to obtain under special environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of change detection technology for dual-temporal multimodal remote sensing images, and in particular to a method for detecting changes in multimodal remote sensing data. Background Technology

[0002] Change detection (CD) using remote sensing imagery is widely applied in various fields, such as urban sprawl detection, geological hazard monitoring, and urban damage assessment. Currently, most CD methods focus on a single data source, primarily optical imagery. However, optical imagery is affected by weather conditions, especially during disasters such as typhoons and floods, making it difficult to acquire similar images from different periods. Synthetic Aperture Radar (SAR) imagery offers all-weather, all-time coverage, compensating for the shortcomings of optical imagery. Utilizing multimodal remote sensing imagery for change detection has practical application significance.

[0003] Because remote sensing images of different modalities have inconsistent imaging mechanisms, radiometric characteristics, and geometric features, existing change detection models based on homogeneous images are not suitable for multimodal images. Current multimodal remote sensing data change detection methods mainly focus on homogenization methods for different modalities, that is, converting different spatial images to the same feature space and performing supervised or unsupervised image transformation (CD). However, these methods are prone to losing image details and introducing unnecessary image noise after conversion between different modalities. Furthermore, stylistic differences exist between image domains at different times, making the model's understanding of changed areas unclear and prone to false positives and false negatives. Therefore, for practical applications such as disaster occurrence, how to quickly and accurately identify changed areas based on multimodal remote sensing images is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a method for detecting changes in multimodal remote sensing data, which overcomes the problems of insufficient feature exchange between different modal images, large differences between the two temporal domains, and inability to fully learn the relationship between features of different modal images in existing methods, resulting in unclear change information and easy false detection and missed detection. This provides technical support for the rapid identification of change information before and after disasters when it is impossible to obtain images of the same type.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for detecting changes in multimodal remote sensing data, comprising the following steps:

[0006] Step S1: Acquire remote sensing images of the study area at two different times and in different modalities, where the earlier T1 is an optical image and the later T2 is a synthetic aperture radar (SAR) image.

[0007] Step S2: Perform preprocessing operations on the T1 temporal optical image obtained in step S1, including radiometric calibration, orthorectification, atmospheric correction, and fusion.

[0008] Step S3: Perform preprocessing operations on the T2 temporal SAR image obtained in step S1, including orbit correction, noise reduction, radiometric calibration, and orthorectification.

[0009] Step S4: Perform image registration on the optical image and SAR image after preprocessing in steps S2 and S3. Based on the image with higher resolution, resample the image with lower resolution.

[0010] Step S5: Crop and augment the two preprocessed images, manually screen and correct them, establish a multimodal tile sample pair dataset of optical and SAR images, and divide it into training sample set, validation sample set and test sample set according to a certain ratio.

[0011] Step S6: Based on the multimodal training sample dataset and validation sample dataset established in step S5, input the corresponding T1 optical image and T2 SAR image sample pairs into two pseudo-twin feature coding layers respectively to extract the image semantic features of different levels of dual-temporal different modal images.

[0012] Step S7: Construct a cross-modal feature interaction module. Through the cross-modal feature interaction module, establish the connection between the features of different modal images at different levels extracted in step S6, eliminate the problem of feature differences between different modal images, and obtain the image features after interaction at different levels.

[0013] Step S8: Apply the change information enhancement module to the image features after different levels of interaction obtained in step S7 to obtain enhanced change information features at different levels.

[0014] Step S9: Based on the enhanced change information features at different levels obtained in step S8, the low-level change information features and the high-level change information features are fused through a multi-scale feature fusion decoder to obtain change detection results with strong semantic information.

[0015] Step S10: Train the model using the training sample set obtained in step S5. Use the binary classification cross-entropy loss function to calculate the difference between the model prediction result and the label to fine-tune the model parameters. Use the validation sample set to adjust the hyperparameters, monitor whether overfitting occurs, and output the model weights.

[0016] Step S11: Based on the model weights obtained in step S10, use the weights in the test sample set constructed in step S5 to evaluate the generalization ability of the model, and use the weights to complete the identification of the change area of ​​the entire region.

[0017] In a preferred embodiment, in step S6, a pseudo-Siamese convolutional neural network is used as a dual-branch architecture feature extractor for feature extraction of different modal images, so as to capture different levels of non-interfering and robust image semantic features of different modal images; the dual-branch feature extractor has the same structure with non-shared weights, and each branch consists of four parts, each part including a convolutional layer with a kernel size of 3×3, a normalization layer BatchNorm, a ReLU activation function layer and a max pooling layer.

[0018] In a preferred embodiment, step S7 is specifically implemented as follows:

[0019] Step S71: For the multi-level features of the dual-temporal images obtained in step S6, design a cross-modal feature interaction module. This module uses group convolution combined with multi-scale convolution kernels, where the convolution kernel sizes are set to 3×3, 5×5, 7×7, and 9×9, respectively. The corresponding group sizes are 1, 4, 8, and 16. The output channel of each convolution kernel is one-quarter of the total number of output channels. Each quarter of the channel feature map contains spatial information corresponding to the scale of the convolution kernel.

[0020] Step S72: Using the dual-temporal channel feature maps extracted in step S71 via group convolution, the channel feature maps of the two modalities are swapped pairwise. Channel feature maps of other modalities are interspersed within the channel feature maps of each modality. Then, the channel squeezing and activation module SEWeight is used to extract the cross-channel attention weight information of each modality after the channel feature maps are swapped, resulting in attention weight vectors of different scales for the cross-modal information. The definition of the channel squeezing and activation module SEWeight is as follows:

[0021] w c =σ(W1δ(W0(g) c )))

[0022] In the formula, δ represents the ReLU activation function. and Represents two fully connected layers, g c This represents the global draw pooling operation, w c This is the calculated attention weight vector;

[0023] Step S73: Recalibrate the attention weight vectors of different scales of the dual-temporal cross-modal information obtained in step S72 using the Softmax activation function so that the weights contain all positional and channel information of the cross-modal; Multiply the attention weight vectors of the two cross-modal information obtained by the step S72 with the multi-scale channel feature maps extracted from optical and SAR images respectively to obtain the feature maps after the interaction of different modal information.

[0024] Step S74: Insert the cross-modal feature interaction module after each convolutional basic module in the feature extractor. A total of 4 cross-modal feature interaction modules are inserted in the network to strengthen the connection and interaction between different modal image features extracted by the pseudo-twin convolutional network, make the feature distribution between the two branches more similar, and realize automatic domain adaptation between the two temporal domains.

[0025] Step S75: The feature map obtained after the modal information interaction in step S74 is further fed into the next convolutional coding layer to extract the semantic features of the image again, so that the two branches of the network can learn the contextual information between different modal features.

[0026] In a preferred embodiment, step S8 is specifically implemented as follows:

[0027] Step S81: Construct a change information enhancement module. This module takes image feature maps resulting from the interaction of features from two different modalities, performs subtraction and addition operations, and then applies spatial and channel-level squeezing and activation (scSE) to the results. This enhances the relevant features of the change region spatially and channel-level, resulting in a more accurate attention weight vector focusing on the change region. The spatial and channel-level squeezing and activation (scSE) combines channel attention and spatial attention mechanisms in parallel. Channel attention is achieved by performing global flat pooling on the feature maps and using 1×1×1 convolutions. The first step involves processing the line information, then normalizing it using the Sigmoid function, and finally multiplying it with the original feature map along the channel to obtain a channel-calibrated feature map. Spatial attention is achieved by first applying a 1×1×1 convolution to the feature map, transforming it from [C,H,W] to [1,H,W], then applying the Sigmoid activation function and multiplying it with the original feature map to obtain a spatially calibrated feature map. Finally, the channel-calibrated and spatially calibrated feature maps are added together to form a more accurate feature map resulting from spatial and channel compression and activation.

[0028] Step S82: Add the attention weight vectors obtained in step S81 and subtract them respectively, and recalibrate the weights using the Sigmoid activation function. The weights contain information about the enhanced change region.

[0029] Step S83: The image feature maps at two different times are fed into a 3×3 convolution to reduce the number of channels to half of the original number and stack them in terms of channel dimension. The stacked image feature map is multiplied with the weights obtained in S82 to obtain the difference feature map after the change information is enhanced.

[0030] Step S84: Insert the change information enhancement module after the cross-modal feature interaction module to extract change information from the dual-temporal image features after interaction. A total of four change information enhancement modules are added to the model. Based on the dual-branch feature map obtained after each cross-modal feature interaction in step S7, extract four different levels of enhanced difference feature maps.

[0031] In a preferred embodiment, step S9 is specifically implemented as follows:

[0032] Step S91: Based on the four different levels of enhanced differential feature maps obtained in step S8, the enhanced differential feature maps of different levels are fused through a multi-scale feature fusion decoder.

[0033] Step S92: For the multi-scale feature fusion decoder, the difference feature map after the change information enhancement of each layer is first subjected to a 3×3 convolution operation to reduce the number of channels to 64. Then, starting from the difference feature map with the lowest spatial resolution, a 2x upsampling operation is performed and added to the difference feature map of the previous layer to fuse the change information features of the lower layer with the change information features of the higher layer. The original spatial scale is restored layer by layer from bottom to top. Finally, two convolutional layers are used in the last layer to output the recognition result of the change region, thus completing the change information detection of the multimodal image.

[0034] In a preferred embodiment, step S10 is specifically implemented as follows:

[0035] Step S101: Using the multimodal tile sample dataset constructed in step S5, adjust the network hyperparameters appropriately, and set the initial learning rate, optimizer type, training batch of the model, and number of model iterations.

[0036] Step S102: The binary cross-entropy loss function (BCE Loss) is used to calculate and evaluate the degree of difference between the model's predicted results and the true labels. The BCE Loss is defined as follows:

[0037] L BCE =-p(x)·log((q(x))+(1-p(x))·log(1-q(x))

[0038] In the formula, p(x) represents the predicted probability distribution, and q(x) represents the true probability distribution;

[0039] By minimizing the loss function, the model can adjust its parameters to improve prediction accuracy. The goal of training is to continuously optimize the loss function to its minimum value, and to verify the model's generalization performance using a validation set while preserving the optimal model weights.

[0040] Compared with the prior art, the present invention has the following beneficial effects: (1) By using channel feature exchange and attention mechanism, the context information of different modal features is perceived through mutual learning, making the distribution between the two modal features more similar, realizing automatic domain adaptation between the two temporal domains, and solving the problems of insufficient feature exchange and incomparability between features in the traditional method; (2) The change information enhancement module and the multi-scale feature fusion decoder enhance the recognition of change areas of the image features after interaction at different levels, and fuse the change information features of the lower level with the change information features of the higher level, overcoming the problem of missed detection and false detection in the traditional method when recognizing change areas, and providing technical support for recognizing change areas using multimodal images. Attached Figure Description

[0041] Figure 1 This is a schematic diagram of the method flow of a preferred embodiment of the present invention.

[0042] Figure 2 This is a structural diagram of a pseudo-twin convolutional neural network according to a preferred embodiment of the present invention.

[0043] Figure 3 This is a structural diagram of the cross-modal feature interaction module in a preferred embodiment of the present invention.

[0044] Figure 4 This is a structural diagram of the change information enhancement module in a preferred embodiment of the present invention.

[0045] Figure 5 This is a diagram showing the extraction results of a dataset from a certain region, as a preferred embodiment of the present invention. Detailed Implementation

[0046] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0047] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0048] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0049] like Figure 1 As shown in the figure, this embodiment provides a method for detecting changes in multimodal remote sensing data, including the following steps:

[0050] Step S1: Acquire remote sensing images of the study area at two different times and in different modalities, where the earlier (T1) image is an optical image and the later (T2) image is a Synthetic Aperture Radar (SAR) image.

[0051] Step S2: Perform preprocessing operations on the T1 temporal optical image obtained in step S1, including radiometric calibration, orthorectification, atmospheric correction, and fusion.

[0052] Step S3: Perform preprocessing operations on the T2 time-phase SAR image obtained in step S1, including orbit correction, noise reduction, radiometric calibration, orthorectification, etc.

[0053] Step S4: Perform image registration on the optical image and SAR image after preprocessing in steps S2 and S3. Based on the image with higher resolution, resample the image with lower resolution.

[0054] Step S5: Crop and augment the two preprocessed images, manually screen and correct them, establish a multimodal tile sample pair dataset of optical and SAR images, and divide the training sample set, validation sample set and test sample set into an 8:1:1 ratio.

[0055] Step S6: Based on the multimodal training sample dataset and validation sample dataset established in step S5, input the corresponding T1 optical image and T2 SAR image sample pairs into two pseudo-twin feature coding layers respectively to extract the image semantic features of different levels of dual-temporal different modal images.

[0056] Step S7: Construct a cross-modal feature interaction module. Through the cross-modal feature interaction module, establish the connection between the features of different modal images at different levels extracted in step S6, eliminate the problem of feature differences between different modal images, and obtain the image features after interaction at different levels.

[0057] Step S8: Apply the change information enhancement module to the image features obtained in step S7 after different levels of interaction to obtain enhanced change information features at different levels.

[0058] Step S9: Based on the enhanced change information features at different levels obtained in step S8, the low-level change information features and the high-level change information features are fused through a multi-scale feature fusion decoder to obtain change detection results with strong semantic information.

[0059] Step S10: Train the model using the training sample set obtained in step S5. Use the binary classification cross-entropy loss function to calculate the difference between the model prediction result and the label to fine-tune the model parameters. Use the validation sample set to adjust the hyperparameters, monitor whether overfitting occurs, and output the model weights.

[0060] Step S11: Based on the model weights obtained in step S10, use the weights in the test sample set constructed in step S5 to evaluate the generalization ability of the model, and use the weights to complete the identification of the change area of ​​the entire region.

[0061] In this embodiment, step S6 specifically includes the following steps:

[0062] In one embodiment of the present invention, in step S6, a pseudo-Siamese convolutional neural network is used as a feature extractor with a dual-branch architecture to capture robust image semantic features at different levels of images of different modalities. The feature extractor has the same structure, consisting of four parts, each containing a convolutional layer, a pooling layer, and an activation function.

[0063] In this embodiment, step S7 specifically includes the following steps:

[0064] Step S71: Construct a cross-modal feature interaction module for the multi-level features of the dual-temporal images obtained in step S6. The cross-modal feature interaction module uses group convolution combined with multi-scale convolution kernels to effectively extract spatial information at different scales on the feature maps of each channel of different modal image features.

[0065] Step S72: Using the dual-temporal channel feature maps extracted in step S71 through group convolution, the feature channel maps of the two modalities are exchanged. The feature channel maps of other modalities are interspersed in the feature channel maps of each modality. The channel squeezing and excitation module SEWeight is used to extract the cross-channel attention weight information of each modality, and the attention weight vectors of different scales of the cross-modal information are obtained.

[0066] Step S73: Recalibrate the attention weight vectors at different scales of the dual-temporal cross-modal information obtained in step S72 using the Softmax activation function, so that the weights include all positional and channel information of the cross-modal. Multiply the obtained attention weight vectors of the two cross-modalities with the feature maps extracted from optical and SAR images respectively to obtain the feature maps after the interaction of different modal information.

[0067] Step S74: Insert the cross-modal feature interaction module after each convolutional basic module in the feature extractor. A total of 4 cross-modal feature interaction modules are inserted in the network to strengthen the connection and interaction between image features of different modalities.

[0068] Step S75: The feature map obtained from the interaction of different modal information in step S74 is further fed into the next convolutional coding layer to extract the semantic features of the image again.

[0069] In this embodiment, step S8 specifically includes the following steps:

[0070] Step S81: For the bi-branch feature map obtained after each cross-modal feature interaction in step S7, select to construct a change information enhancement module. The change information enhancement module performs subtraction and addition operations on the two image feature maps from different modalities after feature interaction, and then performs spatial and channel compression and excitation on the results of the operations to enhance the relevant features of the change region and obtain the attention weight vector of the change region.

[0071] Step S82: Add the attention weight vectors obtained in step S81 and subtract them respectively, and recalibrate the weights using the Sigmoid activation function. The weights contain information about the enhanced change region.

[0072] Step S83: The image feature maps at two different times are fed into a 3×3 convolution to reduce the number of channels to half of the original number and stack them in terms of channel dimension. The stacked image feature map is multiplied with the weights obtained in S82 to obtain the difference feature map after the change information is enhanced.

[0073] Step S84: Insert the change information enhancement module after the cross-modal feature interaction module to extract change information from the dual-temporal image features after interaction. A total of four change information enhancement modules are added to the model to extract four different levels of enhanced difference feature maps.

[0074] In this embodiment, step S9 specifically includes the following steps:

[0075] Step S91: Based on the four different levels of difference feature maps obtained in step S8, design a multi-scale feature fusion decoder to make full use of the difference feature maps with different spatial resolutions.

[0076] Step S92: For the multi-scale feature fusion decoder, starting from the difference feature map with the lowest spatial resolution, a 2x upsampling operation is first performed, and then added to the difference feature map of the previous layer. The change information features of the lower layer are fused with the change information features of the higher layer. The original spatial scale is restored layer by layer from bottom to top, and the recognition results of the change region are output by two convolutional layers in the last layer to complete the change information detection of the multimodal image.

[0077] In this embodiment, step S10 specifically includes the following steps:

[0078] Step S101: Using the multimodal tile sample dataset constructed in step S5, set reasonable network hyperparameters, initial learning rate, optimizer type, model training batch, and model iteration number;

[0079] Step S102: The binary cross-entropy loss function (BCE Loss) is used to calculate and evaluate the degree of difference between the model's predicted results and the true labels. The BCE Loss is defined as follows:

[0080] L BCE =-p(x)·log((q(x))+(1-p(x))·log(1-q(x))

[0081] In the formula, p(x) represents the predicted probability distribution, and q(x) represents the true probability distribution;

[0082] By minimizing the loss function, the model can adjust its parameters to improve prediction accuracy. The goal of training is to continuously optimize the loss function to its minimum value, and to verify the model's generalization performance using a validation set while preserving the optimal model weights.

[0083] In this embodiment, a multimodal change detection dataset of optical and SAR images of a certain region, after preprocessing and cropping, is used. The spatial resolution is 3m, the optical images are in red, green, and blue bands, and the SAR images are single-band HH images. The original image size is 11216×13693. In this embodiment, 2768 tile images with a size of 256×256 after tile cropping and sample augmentation, along with corresponding labels, are used for model training, validation, and testing.

[0084] like Figure 2This is a structural diagram of the pseudo-twin convolutional neural network constructed in this embodiment. As shown in the diagram, the model mainly consists of a pseudo-twin convolutional feature extractor, a cross-modal feature interaction module, and a change information enhancement module. First, a pseudo-twin convolutional coding feature extractor with non-shared weights is used to extract the semantic features of optical and SAR images respectively. Each branch has four basic convolutional modules. After each feature extraction, the number of channels doubles, and the image spatial size is halved, obtaining four different levels of image features. The extraction of image features from different modalities does not interfere with each other. Next, the cross-modal feature interaction module is used to strengthen the connection between the features of two different modalities, making the feature distributions of the two branches more similar. This module then achieves automatic domain adaptation between the two temporal domains to a certain extent. Finally, the change information enhancement module is used to further extract the change regions between the features of the two temporal images after interaction, extracting change information at four different levels in total. Finally, in order to make full use of the change information at different levels, a multi-scale feature fusion decoder is adopted. Starting from the change information at higher levels, a 1×1 convolution is first performed to reduce the number of channels, and then a 2x upsampling is performed and added to the change information of the previous layer. In this way, the original image size is gradually restored to obtain the final change area information. The change information features at lower levels are fused with the change information features at higher levels to reduce the phenomenon of missed detection and false detection.

[0085] like Figure 3 This is the cross-modal feature interaction module used in this embodiment. First, the Squeeze and Stack module (SPC) is applied to the dual-temporal images. This module mainly utilizes group convolution combined with multi-scale convolution kernels to extract channel feature maps at different spatial scales. Here, we use four convolution kernels of different sizes combined with four types of group convolution, and stack the channel feature maps after group convolution along the channel dimension, so that each quarter of the channel feature map contains spatial information corresponding to the convolution kernel scale. The channel feature maps of the two modalities are then swapped pairwise. Next, the Channel Squeeze and Activation module (SEWeight) is used to extract the cross-channel attention weight information of each modality image after the channel feature maps are swapped. The weights are recalibrated using the Softmax activation function. Then, the obtained attention weight vectors of the two cross-modalities are multiplied with the multi-scale channel feature maps extracted from the optical and SAR images, respectively, to obtain the feature maps after the interaction of different modal information. Through the attention weight vectors of the cross-modalities, the two different branches learn to perceive the contextual information of the two temporalities, making the distribution of features more similar and enabling the comparability of image features from different modalities.

[0086] like Figure 4This is the change information enhancement module used in this example. Specifically, this module is used to extract change information from the interactive dual-temporal features. First, it performs subtraction and addition operations on the two image feature maps from different modalities after feature interaction. Then, it performs spatial and channel compression and scSE activation on the results to enhance the relevant features of the change region spatially and channelally, obtaining a more accurate attention weight vector for the change region. Then, it adds the obtained addition and subtraction attention weight vectors and recalibrates the weights using the Sigmoid activation function. These weights contain the enhanced change region information. At the same time, it performs a 3×3 convolution on the image feature maps at two different times, reducing the number of channels to half of the original and stacking them in the channel dimension. The stacked image feature map is multiplied with the obtained weights to obtain the difference feature map after change information enhancement.

[0087] like Figure 5 This figure shows partial experimental results of a multimodal change detection dataset in a processed area according to this embodiment. As can be seen from the figure, the change detection prediction results obtained using the method of this invention have a high degree of consistency with the true labels; the accuracy and refinement of the change shape are high in detecting change areas between consecutive temporal images, with few missed or false detections; the experimental results demonstrate the effectiveness of the method in change detection using multimodal remote sensing image data.

[0088] Compared with existing methods, the pseudo-twin convolutional multimodal change detection model proposed in this invention utilizes channel feature exchange and attention mechanisms. By learning from each other, it perceives the contextual information of different modal features, making the distributions of two modal features more similar and achieving automatic domain adaptation between the two temporal domains. This solves the problems of insufficient feature exchange and incomparability between features in traditional methods. At the same time, it uses a change information enhancement module and a multi-scale feature fusion decoder to enhance the identification of change regions in image features after interaction at different levels, and fuses the change information features of lower levels with those of higher levels. This overcomes the problems of missed detections and false detections that traditional methods are prone to when identifying change regions, providing technical support for the identification of change regions using multimodal images.

[0089] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for detecting changes in multimodal remote sensing data, characterized in that, Includes the following steps: Step S1: Acquire remote sensing images of the study area at two different times and in different modalities, where the earlier T1 is an optical image and the later T2 is a synthetic aperture radar (SAR) image. Step S2: Perform preprocessing operations on the T1 temporal optical image obtained in step S1, including radiometric calibration, orthorectification, atmospheric correction, and fusion. Step S3: Perform preprocessing operations on the T2 temporal SAR image obtained in step S1, including orbit correction, noise reduction, radiometric calibration, and orthorectification. Step S4: Perform image registration on the optical image and SAR image after preprocessing in steps S2 and S3. Based on the high-resolution image, resample the low-resolution image. Step S5: Crop and augment the two preprocessed images, manually screen and correct them, establish a multimodal tile sample pair dataset of optical and SAR images, and divide it into training sample set, validation sample set and test sample set according to a certain ratio. Step S6: Based on the multimodal training sample dataset and validation sample dataset established in step S5, input the corresponding T1 optical image and T2 SAR image sample pairs into two pseudo-twin feature coding layers respectively to extract the image semantic features of different levels of dual-temporal different modal images. Step S7: Construct a cross-modal feature interaction module. Through the cross-modal feature interaction module, establish the connection between the features of different modal images at different levels extracted in step S6, eliminate the problem of feature differences between different modal images, and obtain the image features after interaction at different levels. Step S8: Apply the change information enhancement module to the image features obtained in step S7 after different levels of interaction to obtain enhanced change information features at different levels. Step S9: Based on the enhanced change information features at different levels obtained in step S8, the low-level change information features and the high-level change information features are fused through a multi-scale feature fusion decoder to obtain change detection results with strong semantic information. Step S10: Train the model using the training sample set obtained in step S5. Use the binary classification cross-entropy loss function to calculate the difference between the model prediction result and the label to fine-tune the model parameters. Use the validation sample set to adjust the hyperparameters, monitor whether overfitting occurs, and output the model weights. Step S11: Based on the model weights obtained in step S10, use the weights in the test sample set constructed in step S5 to evaluate the generalization ability of the model, and use the weights to complete the identification of the change area of ​​the entire region. Step S7 is implemented as follows: Step S71: For the multi-level features of the dual-temporal images obtained in step S6, design a cross-modal feature interaction module. This module uses group convolution combined with multi-scale convolution kernels, where the convolution kernel sizes are set to 3×3, 5×5, 7×7, and 9×9, respectively. The corresponding group sizes are 1, 4, 8, and 16. The output channel of each convolution kernel is one-quarter of the total number of output channels. Each quarter of the channel feature map contains spatial information corresponding to the scale of the convolution kernel. Step S72: Using the dual-temporal channel feature maps extracted in step S71 via group convolution, the channel feature maps of the two modalities are swapped pairwise. Channel feature maps of other modalities are interspersed within the channel feature maps of each modality. Then, the channel squeezing and activation module SEWeight is used to extract the cross-channel attention weight information of each modality after the channel feature maps are swapped, resulting in attention weight vectors of different scales for the cross-modal information. The definition of the channel squeezing and activation module SEWeight is as follows: In the formula Represents the ReLU activation function. and This represents two fully connected layers. This indicates a global tie pooling operation. This is the calculated attention weight vector; Step S73: Recalibrate the attention weight vectors of different scales of the dual-temporal cross-modal information obtained in step S72 using the Softmax activation function so that the weights contain all positional and channel information of the cross-modal; Multiply the attention weight vectors of the two cross-modal information obtained by the step S72 with the multi-scale channel feature maps extracted from optical and SAR images respectively to obtain the feature maps after the interaction of different modal information. Step S74: Insert the cross-modal feature interaction module after each convolutional basic module in the feature extractor. A total of 4 cross-modal feature interaction modules are inserted in the network to strengthen the connection and interaction between different modal image features extracted by the pseudo-twin convolutional network, make the feature distribution between the two branches more similar, and realize automatic domain adaptation between the two temporal domains. Step S75: The feature map obtained after the modal information interaction in step S74 is further fed into the next convolutional coding layer to extract the semantic features of the image again, so that the two branches of the network can learn the contextual information between different modal features.

2. The method for detecting changes in multimodal remote sensing data according to claim 1, characterized in that, In step S6, a pseudo-Siamese convolutional neural network is used as a dual-branch architecture feature extractor for feature extraction of different modal images, so as to capture different levels of robust image semantic features that do not interfere with each other. The dual-branch feature extractor has the same structure with non-shared weights. Each branch consists of four parts, each of which includes a convolutional layer with a kernel size of 3×3, a normalization layer BatchNorm, a ReLU activation function layer, and a max pooling layer.

3. The method for detecting changes in multimodal remote sensing data according to claim 1, characterized in that, Step S8 is implemented as follows: Step S81: Construct a change information enhancement module. This module takes image feature maps resulting from the interaction of features from two different modalities, performs subtraction and addition operations, and then applies spatial and channel-level squeezing and activation (scSE) to the results. This enhances the relevant features of the change region spatially and channel-level, resulting in a more accurate attention weight vector focusing on the change region. The spatial and channel-level squeezing and activation (scSE) combines channel attention and spatial attention mechanisms in parallel. Channel attention is achieved by performing global flat pooling on the feature maps and using 1×1×1 convolutions. The first step involves processing the line information, then normalizing it using the Sigmoid function, and finally multiplying it with the original feature map along the channel to obtain a channel-calibrated feature map. Spatial attention is achieved by first applying a 1×1×1 convolution to the feature map, transforming it from [C,H,W] to [1,H,W], then applying the Sigmoid activation function and multiplying it with the original feature map to obtain a spatially calibrated feature map. Finally, the channel-calibrated and spatially calibrated feature maps are added together to form a more accurate feature map resulting from spatial and channel compression and activation. Step S82: Add the attention weight vectors obtained in step S81 and subtract them respectively, and recalibrate the weights using the Sigmoid activation function. The weights contain information about the enhanced change region. Step S83: The image feature maps at two different times are fed into a 3×3 convolution to reduce the number of channels to half of the original number and stack them in terms of channel dimension. The stacked image feature map is multiplied with the weights obtained in S82 to obtain the difference feature map after the change information is enhanced. Step S84: Insert the change information enhancement module after the cross-modal feature interaction module to extract change information from the dual-temporal image features after interaction. A total of four change information enhancement modules are added to the model. Based on the dual-branch feature map obtained after each cross-modal feature interaction in step S7, extract four different levels of enhanced difference feature maps.

4. The method for detecting changes in multimodal remote sensing data according to claim 1, characterized in that, Step S9 is implemented as follows: Step S91: Based on the four different levels of enhanced differential feature maps obtained in step S8, the enhanced differential feature maps of different levels are fused through a multi-scale feature fusion decoder. Step S92: For the multi-scale feature fusion decoder, the difference feature map after the change information enhancement of each layer is first subjected to a 3×3 convolution operation to reduce the number of channels to 64. Then, starting from the difference feature map with the lowest spatial resolution, a 2x upsampling operation is performed and added to the difference feature map of the previous layer to fuse the change information features of the lower layer with the change information features of the higher layer. The original spatial scale is restored layer by layer from bottom to top. Finally, two convolutional layers are used in the last layer to output the recognition result of the change region, thus completing the change information detection of the multimodal image.

5. The method for detecting changes in multimodal remote sensing data according to claim 1, characterized in that, Step S10 is implemented as follows: Step S101: Using the multimodal tile sample dataset constructed in step S5, adjust the network hyperparameters appropriately, and set the initial learning rate, optimizer type, training batch of the model, and number of model iterations. Step S102: The binary cross-entropy loss function (BCE Loss) is used to calculate and evaluate the degree of difference between the model's predicted results and the true labels. The BCE Loss is defined as follows: In the formula, Represents the predicted probability distribution. Represents the true probability distribution; By minimizing the loss function, the model adjusts its parameters to improve prediction accuracy. The goal of training is to continuously optimize the loss function to its minimum value, and to verify the model's generalization performance using a validation set while preserving the optimal model weights.

Citation Information

Patent Citations

  • Self-attention feature fused high-resolution remote sensing image semantic change detection method

    CN116486255A

  • RGB-D cross-modal interactive fusion mechanical arm grabbing detection method based on Transform-CNN hybrid architecture

    CN116912608A