A Multimodal Remote Sensing Image Fusion Classification Method Based on a Dual-Branch Dynamic Modulation Network

Through the dynamic gradient optimization and feature enhancement strategy of the dual-branch dynamic modulation network, the problem of modal uneven convergence in the multimodal remote sensing image fusion network is solved, and higher classification accuracy and feature consistency are achieved.

CN116912646BActive Publication Date: 2025-08-01NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310816242.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-05
Publication Date
2025-08-01
Estimated Expiration
2043-07-05

AI Technical Summary

Technical Problem

In the existing multimodal remote sensing image fusion network, each mode may converge at different rates, causing one mode to suppress the expression of another mode feature, affecting the classification accuracy.

Method used

The dual-branch dynamic modulation network is adopted to adjust the optimization process of each branch through a dynamic gradient optimization strategy, and the balanced convergence and complementary enhancement of modal features are achieved through the multimodal bidirectional enhancement module and feature distribution consistency loss function.

Benefits of technology

The coordinated classification accuracy of multimodal remote sensing data is improved, feature extraction capabilities are enhanced, and the consistency and accuracy of classification results are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116912646B_ABST
    Figure CN116912646B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal remote sensing image fusion and classification method based on a dual-branch dynamic modulation network, comprising the following steps: Step 1: Obtain a multi-modal remote sensing image data set and perform preprocessing; Step 2: Propose a dynamic multi-modal gradient optimization strategy to dynamically adjust the optimization process of each branch during the backpropagation process through gradient modulation; Step 3: Construct a multi-modal bidirectional enhancement module to achieve complementary enhancement of multi-modal remote sensing features; Step 4: Design a feature distribution consistency loss function to quantify the similarity between the integrated features and the dominant features; Step 5: Optimize using the loss function to obtain a well-trained network; Step 6: Use the trained model to predict the data of the test set to obtain the classification result. The present invention proposes a dual-branch dynamic modulation network, which solves the problem of unbalanced convergence of the dual-branch heterogeneous data extraction network, can dynamically correct the feature expression process, realizes the balanced convergence of the multi-modal model, and thus achieves better collaborative classification accuracy of multi-modal remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent interpretation of remote sensing images, and specifically relates to a multi-modal remote sensing image fusion classification method based on a dual-branch dynamic modulation network. Background Art

[0002] As a non-contact technology, remote sensing images have been widely used in multiple application scenarios such as land cover classification and urban development monitoring. With the development of advanced remote sensing technologies, it is now possible to collect descriptive information of potential surface cover objects from multiple sensors simultaneously (such as hyperspectral (HS), light detection and ranging (LiDAR), synthetic aperture radar (SAR), multispectral (MS), optical detectors, etc.). This has special significance in the intelligent interpretation of remote sensing data because different sensors can capture various attributes of the same geographical area, which helps to understand the entire scene. Therefore, it is of great value to make full use of the complementarity of different modal data information and study relevant information collaborative processing technologies to improve the accuracy and reliability of classification.

[0003] Inspired by the success of deep learning models in computer vision, many multi-modal feature fusion networks for remote sensing land cover classification have been proposed. In multi-modal feature fusion networks, a multi-branch architecture is usually adopted, with each branch corresponding to a feature extraction block, and then the features are fused through some fusion strategies (such as concatenation, alignment, etc.) to form a joint representation for further classification. Although these multi-branch deep learning architectures have achieved impressive performance, they have an important limitation. Specifically, due to the backpropagation of parameters, the feature representation of one modality may affect the feature representation of other modalities. This potential drawback should not be ignored when designing architectures for these types of multi-modal fusion tasks. Recent research has shown that multi-modal models using a joint training strategy to optimize a unified learning objective for all modalities may sometimes perform worse than single-modal models.

[0004] Therefore, the present application proposes a dual-branch dynamic modulation network, which solves the problem of unbalanced convergence of the dual-branch heterogeneous data extraction network, can dynamically correct the feature expression process, realizes the balanced convergence of the multi-modal model, and thus achieves better collaborative classification accuracy for multi-modal remote sensing data. Summary of the Invention

[0005] Aiming at the above technical problems, the present invention provides a multi-modal remote sensing image fusion classification method based on a dual-branch dynamic modulation network. This method solves the problem that each modality in multi-modal learning may converge at different rates, resulting in one modality suppressing the feature expression of another modality. It can dynamically correct the feature expression process, realize the balanced convergence of the multi-modal model, and thus achieve better collaborative classification accuracy for multi-modal remote sensing images.

[0006] The technical method adopted by the present invention is: a multimodal remote sensing image fusion classification method based on a dual-branch dynamic modulation network, which is characterized by including the following steps:

[0007] Step 1: Obtain the modal 1 image and the modal 2 image datasets and perform preprocessing;

[0008] Step 101: Obtain the modal 1 image and the modal 2 image 0 where the modal 1 image has C1 channels, the number of pixels of the modal 1 image is a1×b1, the modal 2 image has C2 channels, and the number of pixels of the modal 2 image is a2×b2;

[0009] Step 102: Perform preprocessing operations of registration, cropping, and annotation on the modal 1 and modal 2 images obtained in Step 101 to obtain a local modal 1 image a local modal 2 image and the corresponding label where N = a×b;

[0010] Step 103: Divide the modal 1 and modal 2 data obtained in Step 102 into a training set and a test set;

[0011] Step 2: Propose a dynamic multimodal gradient optimization strategy to dynamically adjust the optimization process of each branch during the backpropagation process through gradient modulation;

[0012] Step 201: Use the encoders Φ h (θ h ,·) and Φ l (θ l ,·) to extract the features of the modal 1 and modal 2 image patches;

[0013] Step 202: Calculate the approximate prediction values of the modal 1 feature extraction branch and the modal 2 feature extraction branch in the dual-branch network:

[0014] Step 203: Calculate the contribution ratio of the modal 1 branch:

[0015] Step 204: Calculate the contribution ratio of the modal 2 branch The contribution ratio of the modal 2 branch is defined as the reciprocal of;

[0016] Step 205: Dynamically select the dominant modality F d and the auxiliary modality. If then the corresponding modal 1 is defined as the dominant modality at the current moment, and the modal 2 is defined as the auxiliary modality. If Mode 1 is the auxiliary mode at the current moment, and mode 2 is the dominant mode;

[0017] Step 206: Use to limit the optimization of the dominant mode without affecting the auxiliary mode and minimize the suppressed optimization: where α is a hyperparameter used to adjust the degree of optimization;

[0018] Step 207: Incorporate into the Adaptive Moment Estimation optimization method Adam, and update the backpropagation gradient at iteration step k according to where represents the first moment estimate of the gradient, represents the second moment estimate of the gradient, and ε is a constant used to maintain numerical stability;

[0019] Step 3: Construct a multimodal bidirectional enhancement module to achieve complementary enhancement of the features of mode 1 and mode 2;

[0020] Step 301: Input the local feature map of mode 1 data where C is the number of channels, and H×W is the number of pixels, into three convolutional layers g1, θ1, φ, respectively, to generate three new feature maps: {F h,a ,F h,b ,F h,c}, where transpose the dimensions of the feature maps F h,a , F h,b , F h,c from C / 8×H×W, C / 8×H×W, C×H×W to (H×W)×C / 8, C / 8×(H×W), C×H×W;

[0021] Step 302: Calculate the spatial attention matrix of mode 1 data:

[0022] Step 303: Generate three other feature maps from the local feature map of mode 2 data which are: {F l,a ,F l,b ,F l,c}, where transpose the dimensions of the feature maps F h,a , F h,b , F h,c from C / 8×H×W, C / 8×H×W, C×H×W to (H×W)×C / 8, C / 8×(H×W), C×H×W;

[0023] Step 304: Pass through the softmax layer and F l,a and F l,bCalculate the spatial attention matrix of modality 2 data through matrix multiplication between

[0024] Step 305: Input the local feature map of modality 1 data and the local feature map of modality 2 data into the multi-modal bidirectional enhancement module to obtain the multi-modal integrated feature:

[0025]

[0026] Step Four: Design a feature distribution consistency loss function to quantify the similarity between the integrated feature and the dominant feature;

[0027] Step 401: Quantify the distance between the dominant feature map F d and the integrated feature map F m through KL divergence, and calculate the feature distribution consistency loss function accordingly: L FDC = KL(avg(M·F m ), F d );

[0028] Step Five: Optimize using the loss function to obtain a well-trained network;

[0029] Step 501: Input the training set data, and adjust the network parameters according to the predicted values and labels of the training set to optimize the loss function Loss. The calculation method of the loss function is: Loss = L CE + γL FDC , where L CE is the cross-entropy loss between the predicted value and the true label Y:

[0030]

[0031] Step 502: Iteratively execute Step Two to Step Five. After each iteration, increment the iteration count by one until the iteration count is equal to the maximum iteration count. The iteration ends, and a trained dual-branch dynamic modulation network is obtained;

[0032] Step Six: Use the trained model to predict the test set data to obtain the classification result.

[0033] The present invention has the following beneficial effects compared with the prior art:

[0034] 1. The steps of the present invention are simple, reasonably designed, and convenient to implement and use.

[0035] 2. The present invention adopts a dynamic multi-modal gradient optimization strategy, which realizes independently adjusting the optimization of each branch during the backpropagation process and dynamically controls gradient modulation, thus contributing to the fusion classification of multi-modal data;

[0036] 3. Through the multi-modal bidirectional enhancement module, the present invention uses the self-attention mask of the features from modality 1 to generate interaction features to focus on the spectral features of the data in modality 1, and uses the cross-attention mask of the features from modality 2 to enhance the spatial representation of the features in the other modality, thereby enhancing the feature extraction ability of the multi-modal data and promoting the fusion of information at different levels;

[0037] 4. The present invention proposes a feature distribution consistency loss function to measure the similarity between the integrated features and the dominant modality features, making the representations of each category as consistent as possible, and thus improving the classification accuracy.

[0038] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings

[0039] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments

[0040] The method of the present invention will be further described in detail below with reference to the drawings and the embodiments of the present invention.

[0041] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and embodiments.

[0042] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0044] For ease of description, spatial relative terms, such as "above", "over", "on the upper surface", "upper", etc., may be used herein to describe the spatial positional relationship of one device or feature to other devices or features as shown in the figures. It should be understood that the spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is inverted, a device described as "above" or "over" other devices or structures will then be positioned "below" or "under" the other devices or structures. Thus, the exemplary term "above" can include both orientations of "above" and "below". The device may also be positioned in other different ways (rotated 90 degrees or in other orientations), and corresponding interpretations of the spatial relative descriptions used herein will be made accordingly.

[0045] As Figure 1 shown, the present invention includes the following steps:

[0046] Step 1: Obtain the modal 1 image and modal 2 image datasets and perform preprocessing;

[0047] Step 101: Obtain the modal 1 image and modal 2 image covering the same geographical area, where the modal 1 image has C1 number of channels, the number of pixels of the modal 1 image is a1×b1, the modal 2 image has C2 channels, and the number of pixels of the modal 2 image is a2×b2;

[0048] Step 102: Perform preprocessing operations of registration, cropping, and annotation on the modal 1 and modal 2 images obtained in Step 101 to obtain a modal 1 local image modal 2 local image and the corresponding label where N = a×b;

[0049] Step 103: Divide the modal 1 and modal 2 data obtained in Step 102 into a training set and a test set;

[0050] Step 2: Propose a dynamic multi-modal gradient optimization strategy to dynamically adjust the optimization process of each branch during the backpropagation process through gradient modulation;

[0051] Step 201: Use the encoders Φ h (θ h ,·) and Φ l (θ l ,·) to extract the features of the modal 1 and modal 2 image patches;

[0052] Step 202: Calculate the approximate prediction values of the modality 1 feature extraction branch and the modality 2 feature extraction branch in the dual-branch network:

[0053] Step 203: Calculate the contribution ratio of the modality 1 branch:

[0054] Step 204: Calculate the contribution ratio of the modality 2 branch The contribution ratio of the modality 2 branch is defined as the reciprocal of;

[0055] Step 205: Dynamically select the dominant modality F d and the auxiliary modality. If then the corresponding modality 1 is defined as the dominant modality at the current moment, and modality 2 is defined as the auxiliary modality. If modality 1 is the auxiliary modality at the current moment, and modality 2 is the dominant modality;

[0056] Step 206: Use to restrict the optimization of the dominant modality without affecting the auxiliary modality and minimize the suppressed optimization: where α is a hyperparameter used to adjust the degree of optimization;

[0057] Step 207: Incorporate into the Adaptive Moment Estimation optimization method Adam and update the backpropagation gradient at iteration step k according to where represents the first moment estimate of the gradient, represents the second moment estimate of the gradient, and ε is a constant used to maintain numerical stability;

[0058] Step 3: Construct a multi-modal bidirectional enhancement module to achieve complementary enhancement of modality 1 and modality 2 features;

[0059] Step 301: Input the local feature map of the modality 1 data where C is the number of channels, and H×W is the number of pixels, into three convolutional layers g1, θ1, φ, respectively, to generate three new feature maps: {F h,a ,F h,b ,F h,c}, where Transpose the dimensions of the feature maps F h,a , F h,b , F h,c from C / 8×H×W, C / 8×H×W, C×H×W to (H×W)×C / 8, C / 8×(H×W), C×H×W;

[0060] Step 302: Calculate the spatial attention matrix of the modality 1 data:

[0061] Step 303: From the local feature map of the modality 2 data Generate another three feature maps, namely: {F l,a , F l,b , F l,c}, where Transpose the dimensions of the feature maps F h,a , F h,b , F h,c from C / 8×H×W, C / 8×H×W, C×H×W to (H×W)×C / 8, C / 8×(H×W), C×H×W; <F

[0062] Step 304: Calculate the spatial attention matrix of the modality 2 data through the matrix multiplication between the softmax layer and F l,a and F l,b

[0063] Step 305: Input the local feature map of the modality 1 data and the local feature map of the modality 2 data into the multi-modal bidirectional enhancement module to obtain the multi-modal integrated feature:

[0064]

[0065] Step 4: Design a feature distribution consistency loss function to quantify the similarity between the integrated feature and the dominant feature;

[0066] Step 401: Quantify the distance between the dominant feature map F d and the integrated feature map F m through the KL divergence, and calculate the feature distribution consistency loss function accordingly: L FDC = KL(avg(M·F m ), F d );

[0067] Step 5: Optimize using the loss function to obtain a well-trained network;

[0068] Step 501: Input the training set data, and adjust the network parameters according to the predicted values and labels of the training set to optimize the loss function Loss. The calculation method of the loss function is: Loss = L CE + γL FDC , where L CE is the cross-entropy loss between the predicted value and the true label Y:

[0069]

[0070] ​Step 502: Iteratively execute Steps Two to Five. After each iteration, increment the iteration count by one until the iteration count is equal to the maximum iteration count. Then, the iteration ends, and a trained dual-branch dynamic modulation network is obtained.

[0071] Step Six: Use the trained model to predict the test set data to obtain classification results.

[0072] As described above, it is only an embodiment of the present invention and does not impose any limitation on the present invention. Any simple modification, change, and equivalent structural change made to the above embodiments according to the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A multi-modal remote sensing image fusion and classification method based on a dual-branch dynamic modulation network, characterized in that It includes the following steps: Step 1: Obtain the modality 1 image and modality 2 image datasets and perform preprocessing; Step 101: Obtain a modality 1 image covering the same geographical area and a modality 2 image wherein, the modality 1 image has C1 channels, the number of pixels of the modality 1 image is a1×b1, the modality 2 image has C2 channels, and the number of pixels of the modality 2 image is a2×b2; Step 102: Perform preprocessing operations of registration, cropping, and annotation on the modality 1 and modality 2 images obtained in Step 101 to obtain a local modality 1 image with N pixel points Local modality 2 image and the corresponding label where N = a × b; Step 103: Divide the modality 1 and modality 2 data obtained in Step 102 into a training set and a test set; Step 2: Propose a dynamic multi-modal gradient optimization strategy to dynamically adjust the optimization process of each branch during backpropagation through gradient modulation; Step 201: Use the encoders Φ h (θ h , ·) and Φ l (θ l , ·) to extract the features of the image patches of modality 1 and modality 2; Step 202: Calculate the approximate predicted values of the modality 1 feature extraction branch and the modality 2 feature extraction branch in the dual-branch network: Step 203: Calculate the modal 1 branch contribution ratio: Step 204: Calculate the modal 2 branch contribution ratio The modal 2 branch contribution ratio is defined as the reciprocal of; Step 205: Dynamically select the dominant mode F d and the auxiliary mode. If then the corresponding mode 1 is defined as the dominant mode at the current moment, and mode 2 is defined as the auxiliary mode. If mode 1 is the auxiliary mode at the current moment and mode 2 is the dominant mode; Step 206: Utilize to optimize the dominant mode with constraints without affecting the auxiliary mode and minimize the suppressed optimization: where α is a hyperparameter for adjusting the degree of optimization; Step 207: Incorporate into the Adaptive Moment Estimation optimization method Adam, and update the backpropagation gradient at iteration step k according to , where represents the first moment estimate of the gradient, represents the second moment estimate of the gradient, and ε is a constant used to maintain numerical stability; Step 3: Construct a multi-modal bidirectional enhancement module to achieve complementary enhancement of modality 1 and modality 2 features; Step 301: The local feature map of the modality 1 data where C is the number of channels, H×W is the number of pixels, and input three convolutional layers g1, θ1, φ to generate three new feature maps: {F h,a , F h,b , F h,c}, where transpose the dimensions of the feature maps F h,a , F h,b , F h,c from C / 8×H×W, C / 8×H×W, C×H×W to (H×W)×C / 8, C / 8×(H×W), C×H×W; Step 302: Calculate the spatial attention matrix of the mode 1 data: Step 303: From the local feature map of the modality 2 data Generate another three feature maps, namely: {F l,a , F l,b , F l,c}, where transpose the dimensions of the feature maps F h,a , F h,b , F h,c from C / 8×H×W, C / 8×H×W, C×H×W to (H×W)×C / 8, C / 8×(H×W), C×H×W; Step 304: Calculate the spatial attention matrix of the modality 2 data through the matrix multiplication between the softmax layer and F l,a and F l,b ​ Step 305: Input the local feature maps of the modality 1 data and the local feature maps of the modality 2 data into the multimodal bidirectional enhancement module to obtain multimodal integrated features: Step 4: Design a feature distribution consistency loss function to quantify the similarity between the integrated features and the dominant features; Step 401: Quantify the distance between the dominant feature map F d and the integrated feature map F m and calculate the feature distribution consistency loss function accordingly: L FDC = KL(avg(M·F m ), F d ); Step 5: Adopt loss function optimization to obtain a well-trained network; Step 501: Input the training set data. Adjust the network parameters according to the predicted values and labels of the training set to optimize the loss function Loss. The calculation method of the loss function is: Loss = L CE + γL FDC , where L CE is the cross-entropy loss between the predicted value and the true label Y: Step 502: Iteratively execute Step 2 to Step 5. After each iteration, increment the iteration count by one until the iteration count is equal to the maximum iteration count. Then, end the iteration to obtain a trained dual-branch dynamic modulation network; Step 6: Use the trained model to predict the test set data to obtain classification results.

Citation Information

Patent Citations

  • Coding and decoding structure-based multi-modal remote sensing image semantic segmentation method

    CN115984701A

  • Multi-modal sentiment analysis method based on dynamic gradient and multi-view collaborative attention

    CN116204850A