A multimodal remote sensing image fusion classification method under modality missing conditions

By using specific encoders and shared encoders in the multimodal remote sensing image classification method, and combining residual fusion modules and auxiliary tasks, the problem of multimodal feature fusion under modal loss is solved, significantly improving the robustness and prediction accuracy of the model.

CN119580101BActive Publication Date: 2025-05-09WUXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510130910.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-09
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

In the absence of modality, the robustness and prediction accuracy of the multimodal remote sensing image fusion classification method are affected, making it difficult to effectively extract and fuse multimodal features.

Method used

A specific encoder is used to extract specific features of each modality, and a shared encoder is used to extract shared features between modalities. A specific and shared features are learned through residual fusion modules and auxiliary tasks (comparative learning and domain classification tasks), ensuring that the model can still effectively fuse features in the absence of modality.

Benefits of technology

In the absence of complete modes and modalities, the robustness and prediction accuracy of the multimodal remote sensing image classification model are significantly improved, ensuring the efficient performance of the model in multimodal data fusion analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580101B_ABST
    Figure CN119580101B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal remote sensing image fusion classification method under modality missing conditions. By constructing a multimodal remote sensing image fusion classification model, the data of three remote sensing image modalities are used as input, each modality data is first extracted through its own specific feature encoder and a shared feature encoder, and then the specific features and shared features are fused, and a jump link is used to generate a single modality fusion feature. Finally, the fusion features of the three modalities are fused again to generate the final fusion feature for model prediction. The present invention introduces two auxiliary tasks in the model, one is a domain classification task based on specific features, and the other is a comparative learning task for shared features. The model of the present invention is not only suitable for full modality training, but also effectively extracts and fuses multimodal features in the case of partial modality missing, thereby improving the classification performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of remote sensing image classification, and in particular relates to a multi-modal remote sensing image fusion classification method under modality missing conditions. Background Art

[0002] Using remote sensing image data for object classification is an important research field with broad application prospects. It is of great significance in urban planning, mineral exploration, environmental monitoring and other fields. With the development of remote sensing technology, descriptions of potential land cover targets can be collected simultaneously through a series of sensors, such as hyperspectral (HS), laser radar (LiDAR), synthetic aperture radar (SAR), multispectral (MS), etc. By analyzing remote sensing image data, different types of objects can be identified and classified, such as buildings, roads, water bodies, vegetation, etc. At first, researchers tried to use the specific advantages of single-source sensor data to classify objects. However, the feature dimensions and information provided by single-source sensor data are limited, and it is difficult to fully capture the complex properties of objects. This limitation gradually prompted researchers to turn to the fusion of multi-source remote sensing data. Multi-source data fusion can integrate the advantages of various sensors, thereby more comprehensively characterizing objects in terms of spatial, spectral and texture features, and effectively improving classification accuracy.

[0003] Although multimodal methods show better performance than single-modal methods, there are still certain challenges in practical applications. Among them, modality loss is a relatively common problem, especially when imaging a specific area. The sampling frequency of some sensors is low or affected by factors such as weather and sensor failure, making it difficult to synchronously obtain multimodal data. This problem significantly increases the uncertainty of multimodal methods in practical scenarios, and effective response strategies are urgently needed to deal with the problem of modality loss, so as to ensure the robustness of the classification model under the missing modality. Summary of the invention

[0004] In view of the problems existing in the prior art, the present invention provides a multimodal remote sensing image fusion classification method under modality missing conditions, which can effectively extract and fuse multimodal features and improve the classification performance of the model not only in the case of complete modality but also in the case of partial modality missing.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: A multimodal remote sensing image fusion classification method under modality missing conditions, using a specific encoder to extract specific features of each modality and using a shared encoder to extract shared features between modalities, and using two auxiliary tasks to learn specific features and shared features respectively, including the following steps: A multimodal remote sensing image fusion classification method under modality missing conditions, characterized in that a specific encoder is used to extract specific features of each modality and a shared encoder is used to extract shared features between modalities, and two auxiliary tasks are used to learn specific features and shared features respectively, including the following steps:

[0006] S1. Acquire multimodal data covering the same area to construct a multimodal dataset and perform preprocessing. The multimodal dataset includes: three types of image data, namely, hyperspectral image data, DSM image data, and SAR image data, or a number of real pixels extracted from one or two of the image data, and corresponding classification labels;

[0007] S2, taking the three image data, the hyperspectral image data, the DSM image data and the SAR image data, or a number of real pixels extracted from one or two of the image data as input, and corresponding specific features as output, to construct a specific feature encoder;

[0008] The specific feature encoder extracts specific features from a number of real pixels extracted from three types of image data, namely, the hyperspectral image data, the DSM image data, and the SAR image data, or from one or two of the image data, as shown in the following formula:

[0009]

[0010] Among them, X i represents the pixel of hyperspectral image data, DSM image data or SAR image data, i∈{1, 2, 3}, f speciflc represents a specific feature encoder;

[0011] The three image data of hyperspectral image data, DSM image data and SAR image data or one or two of them are connected as input, and the shared features corresponding to the respective modalities are obtained as output to construct a shared feature encoder; the shared feature encoder extracts the shared features corresponding to the respective modalities of the three image data of hyperspectral image data, DSM image data and SAR image data or one or two of them, as shown in the following formula:

[0012]

[0013] Among them, X i represents the pixel of hyperspectral image data, DSM image data or SAR image data, i∈{1, 2, 3}, f shared represents the shared feature encoder;

[0014] S3, using specific features and shared features as input and final features as output, constructing a residual fusion module; specifically including the following sub-steps:

[0015] S3.1. Residual fusion module for specific features and shared features After concatenation in the channel dimension and projection operation, its output is added to the shared features as a residual to form a more semantically rich modality embedding f i The following expression:

[0016]

[0017] Among them, f proj represents the projection operation. If one modality is missing, the embedding of the missing modality is obtained by averaging the shared features of the remaining existing modalities.

[0018] S3.2. Embed the obtained modes of each mode into f i Concatenate to form the final feature representation f;

[0019] The construction parts of the specific encoder and the shared encoder are the same. They first extract local features by three 3×3 convolution blocks. Each convolution contains a convolution layer, a batch normalization layer, and a ReLU activation function. A variety of convolution operations of different sizes are used, including 1×1 convolution, 3×3 convolution, and global average pooling to further extract multi-scale features and capture multi-level spatial information. Finally, the features are reduced in dimension through 1×1 convolution and pooling operations, and the final representation capability is enhanced by splicing the features of different convolution branches.

[0020] S4, construct a prediction module with the final features as input and the corresponding classification labels as output;

[0021] S5, constructing a multimodal remote sensing image fusion classification network based on the specific feature encoder, shared feature encoder, residual fusion module, and prediction module of step S1 to step S4, training the multimodal remote sensing image fusion classification network with the multimodal data set of step S1, and obtaining a multimodal remote sensing image fusion classification model; when training the multimodal remote sensing image fusion classification network, performing total loss optimization, including minimizing the distance between similar features at the same position of different modalities by using a contrast loss function, and maximizing the difference between features at different positions; using a domain classification loss function to assist in the learning of specific features; and using a main classification task loss function to guide the prediction results;

[0022] S6. Input three or one or two of the collected hyperspectral image data, DSM image data, and SAR image data into a multimodal remote sensing image fusion classification model to obtain a classification label, that is, a classification result of the terrain.

[0023] Furthermore, the aforementioned step S1 includes the following sub-steps:

[0024] S1.1. Obtain multimodal data covering the same area, including hyperspectral image data, DSM image data and SAR image data, denoted as X1∈R C×H×W ,X2∈R H×W and X3∈R 4×H×W , where W and H represent the width and height, and C represents the number of spectral bands of the hyperspectral image;

[0025] S1.2. Data preprocessing: Perform principal component analysis and dimensionality reduction on the hyperspectral image, extract the first d principal components, and the generated hyperspectral data after dimensionality reduction is represented as X′1∈R d×H×W .

[0026] Furthermore, in the aforementioned step S4, the prediction module inputs the final feature f into the fully connected layer and generates the prediction result y using the weight of the fully connected layer, as shown in the following expression:

[0027] y=Soft max(f fc (cat(f 1 , f 2 , f 3 ),

[0028] Among them, f fc (·) represents the fully connected layer operation, and W is the connection weight corresponding to the fully connected layer operation.

[0029] Furthermore, in the aforementioned step S4,

[0030] The main classification task loss function is as follows: Where H(·,·) is the cross entropy loss function, is the true label and y is the predicted result.

[0031] Furthermore, in the aforementioned step S4, the contrast loss includes: the loss of the same position pair of different modalities l + , the loss of different position pairs of the same modality l A- , l B- and the loss l for different position pairs across modalities AB- As follows:

[0032]

[0033]

[0034] in, represents the cosine similarity between modality A and modality B at different positions. Margin is the threshold. The max function ensures that only when the similarity is lower than margin, there will be a positive loss and optimization. At the same time, the mask method is used to reduce the interference of missing modality data. When modality A is missing, the mask m A is 0 when the mode exists, and 1 when the mode does not exist;

[0035] The contrast loss function is:

[0036] l c =m A ·m B ·l + +m A ·l A- +M B ·l B- +m A ·m B ·l AB- .

[0037] Furthermore, in the aforementioned step S4, the domain classification loss function is:

[0038]

[0039] in, is a remote sensing image dataset of N modalities, is the i-th modal data of the j-th sample, y j For its classification label, t i ∈{0, 1} N is a one-hot modal label, which is 1 at the i-th position and 0 at the rest of the positions, i = 1, 2, ..., N; represents the specific features of the i-th mode of the j-th sample; f d :S→Δ N-1 is a function that maps a specific feature to an N-1-dimensional probability simplex, where S represents the specific feature space and Δ N-1 represents a probability simplex with N classes.

[0040] Furthermore, in the aforementioned step S4, the total loss function is expressed as:

[0041] l=μ cls ·l cls +μ c ·l c +μ d ·l d ,

[0042] Among them, μ cls , μ c and μ d They are the loss weights for the main classification task, contrastive learning task, and domain classification task, respectively. The model parameters are optimized through back propagation to minimize the total loss function.

[0043] Compared with the prior art, the beneficial technical effects of the present invention using the above technical solution are as follows:

[0044] The present invention provides a multimodal data processing method that has both the ability to extract specific modal features and the ability to extract cross-modal shared features, which can accurately perform classification tasks in the case of complete modalities and random missing modalities. Through the auxiliary optimization of contrastive learning and domain classification tasks, the robustness and prediction accuracy of the model in multimodal data fusion analysis are effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is the complete modal structure diagram of the model.

[0046] Figure 2 It is the model missing modal structure diagram.

[0047] Figure 3 Are the true label map and color map; in the figure, (a) is the true label map, and (b) is the color map.

[0048] Figure 4 It is a comparison chart of model prediction results under complete modal training and complete modal testing conditions; in the figure, (a) is the prediction chart of the MMANet network, (b) is the prediction chart of the mmFormer network, and (c) is the prediction chart of the network of the present invention.

[0049] Figure 5 It is a comparison diagram of model prediction results under random missing modal training and complete modal testing. In the figure, (a) is the prediction diagram of MMANet network, (b) is the prediction diagram of mmFormer network, and (c) is the prediction diagram of the network of the present invention.

[0050] Figure 6 It is a schematic diagram of the prediction accuracy of different models under each complete modality training.

[0051] Figure 7 It is a schematic diagram of the prediction accuracy of different models under random missing mode training. DETAILED DESCRIPTION

[0052] In order to better understand the technical content of the present invention, specific embodiments are described below in conjunction with the accompanying drawings.

[0053] Various aspects of the invention are described herein with reference to the accompanying drawings, in which many illustrative embodiments are shown. The embodiments of the invention are not limited to those described in the accompanying drawings. It should be understood that the invention is implemented by any of the various concepts and embodiments described above, as well as the concepts and embodiments described in detail below, because the concepts and embodiments disclosed in the invention are not limited to any implementation. In addition, some aspects disclosed in the invention may be used alone or in any appropriate combination with other aspects disclosed in the invention.

[0054] The present invention takes data of three remote sensing image modalities as input. Each modality data is first subjected to feature extraction through its own specific modality encoder and a shared encoder. Then the obtained specific features and shared features are fused, and a jump link is used to generate a single modality fusion feature. Finally, the fused features of the three modalities are fused again to generate the final fusion feature for model prediction.

[0055] In addition, the present invention introduces two auxiliary tasks into the model to promote the learning of specific features and shared features. One is a domain classification task for learning specific features, and the other is a comparative learning task for shared features. This model is not only suitable for full-modal training, but also can effectively extract and fuse multimodal features when some modalities are missing. This model is a model for the modality missing problem of trimodal remote sensing data, and it aims to effectively extract and fuse multimodal features even when some modalities are missing, thereby improving the classification performance of the model.

[0056] refer to Figure 1 The present invention provides a multi-modal remote sensing image fusion classification method under modality missing conditions, comprising the following steps:

[0057] S1. Acquire multimodal data covering the same area to construct a multimodal dataset and perform preprocessing. The multimodal dataset includes: three types of image data, namely, hyperspectral image data, DSM image data, and SAR image data, or a number of real pixels extracted from one or two of the image data, and corresponding classification labels;

[0058] S2, taking a number of real pixels extracted from three types of image data, namely, hyperspectral image data, DSM image data and SAR image data, or one or two of the image data, as input, and corresponding specific features as output, to construct a specific feature encoder;

[0059] Taking the three image data of hyperspectral image data, DSM image data and SAR image data or connecting one or two of them as input, obtaining the shared features corresponding to each modality as output, and constructing a shared feature encoder;

[0060] S3, using specific features and shared features as input and final features as output to construct a residual fusion module; S4, using the final features as input and the corresponding classification labels as output to construct a prediction module;

[0061] S5, constructing a multimodal remote sensing image fusion classification network based on the specific feature encoder, shared feature encoder, residual fusion module, and prediction module of step S1 to step S4, training the multimodal remote sensing image fusion classification network with the multimodal data set of step S1, and obtaining a multimodal remote sensing image fusion classification model; when training the multimodal remote sensing image fusion classification network, performing total loss optimization, including minimizing the distance between similar features of the same position of different modalities by using the contrast loss function, maximizing the difference between features of different position pairs of different modalities and different position pairs of the same modality; using the domain classification loss function to assist the learning of specific features; and using the main classification task loss function to guide the prediction results;

[0062] S6. Input three or one or two of the collected hyperspectral image data, DSM image data, and SAR image data into a multimodal remote sensing image fusion classification model to obtain a classification label, that is, a classification result of the terrain.

[0063] To verify the effectiveness of this research method, we use the Augsburg dataset for experiments. This dataset consists of three different data sources, including hyperspectral images (HSI), synthetic aperture radar images (SAR), and digital surface model images (DSM). Specifically, the hyperspectral image contains 180 bands, the SAR image consists of four feature channels, and the DSM image is a grayscale image. The spatial resolution of each image is 332×485 pixels. Figure 3 The visualization dataset is divided into 7 categories and contains 78,294 ground truth pixels, of which 761 pixels are used for training and 77,533 pixels are used for testing. In the figure, (a) shows the true label and (b) shows the color map.

[0064] As a preferred embodiment of the present invention, step S1, obtain multimodal data covering the same area, including hyperspectral data, DSM data and SAR data, respectively denoted as X1∈R C×H×W ,X2∈R H×W and X3∈R 4×H×W , where and represent the width and height of the three, and represents the number of spectral bands of the hyperspectral image.

[0065] The hyperspectral image is subjected to principal component analysis (PCA) dimensionality reduction processing, and the first d principal components are extracted. The generated hyperspectral data after dimensionality reduction is represented as X′1∈R d×H×W , to reduce the high-dimensional redundancy of hyperspectral images.

[0066] As a preferred embodiment of the present invention, the pre-processed hyperspectral data, DSM data and SAR data are respectively input into their respective specific encoders to extract specific features represented as follows: and At the same time, the shared features extracted by the shared encoder are expressed as and Specifically, the construction parts of the specific encoder and the shared encoder are the same, and both first extract local features by three 3×3 convolution blocks. Each convolution contains a convolution layer, a batch normalization layer, and a ReLU activation function. These modules help to extract low-level features of a specific modality from the data. On this basis, a variety of convolution operations of different sizes (including 1×1 convolution, 3×3 convolution, and global average pooling) are used to further extract multi-scale features and capture multi-level spatial information. Finally, the features are reduced in dimension through 1×1 convolution and pooling operations, and the final representation capability is enhanced by splicing the features of different convolution branches. The shared encoder ensures that cross-modal data can share the same set of convolution weights, thereby extracting consistent shared features.

[0067] In step S2, the specific feature encoder extracts specific features from a number of real pixels extracted from three types of image data, namely, the hyperspectral image data, the DSM image data, and the SAR image data, or two of the image data, as shown in the following formula:

[0068]

[0069] Among them, X i represents the pixel of hyperspectral image data, DSM image data or SAR image data, i∈{1, 2, 3}, f specific Represents a specific feature encoder.

[0070] The shared feature encoder extracts the shared features corresponding to the respective modalities of the three image data, namely, the hyperspectral image data, the DSM image data and the SAR image data, or two of the image data, as shown in the following formula:

[0071]

[0072] Among them, X i represents the pixel of hyperspectral image data, DSM image data or SAR image data, i∈{1, 2, 3}, f shared Represents shared feature encoder.

[0073] As a preferred embodiment of the present invention, step S3 is to construct a residual fusion module to fuse the specific features with the shared features. and shared features They are input into the residual fusion module together.

[0074] Step S3 includes the following sub-steps:

[0075] S3.1. Residual fusion module for specific features and shared features After concatenation in the channel dimension and projection operation, its output is added to the shared features as a residual to form a more semantically rich modality embedding f i The following expression:

[0076]

[0077] Among them, f proj represents the projection operation. If one modality is missing, the embedding of the missing modality is obtained by averaging the shared features of the remaining existing modalities.

[0078] S3.2. Embed the obtained modes of each mode into f i Concatenate to form the final feature representation f.

[0079] As a preferred embodiment of the present invention, in step S4, the prediction module inputs the final feature f into the fully connected layer, and generates the prediction result y using the weight of the fully connected layer, as shown in the following expression:

[0080] y=Soft max(f fc (cat(f 1 , f 2 , f 3 );W))

[0081] Among them, f fc (·) represents the fully connected layer operation, and W is the connection weight corresponding to the fully connected layer operation.

[0082] When training the multimodal remote sensing image fusion classification network, the total loss is optimized, including using the contrast loss function to minimize the distance between similar features at the same position in different modalities and maximize the difference between features at different positions; using the domain classification loss function to assist in the learning of specific features; and using the main classification task loss function to guide the prediction results, as follows:

[0083] The loss function of the main classification task is:

[0084]

[0085] Where H(·,·) is the cross entropy loss function, is the true label and y is the predicted result.

[0086] To assist the learning of multimodal shared features, the present invention designs a contrast loss function based on the cosine similarity metric, which aims to minimize the distance between similar features at the same position in different modalities and maximize the difference between features at different positions.

[0087] Specifically, by calculating the cosine similarity between the same-modal and cross-modal feature pairs of modality A and modality B, the losses of positive and negative samples are obtained respectively.

[0088] The loss function consists of two parts: first, the positive sample loss l + The similarity of shared features between modalities is enhanced by minimizing the feature similarity of the same position pairs in different modalities. Secondly, the negative sample loss is used to maximize the distance between different position pairs (including the loss of different position pairs in the same modality). A- , l B- and the loss l of features at different positions across modalities AB- ) to reduce the interference of irrelevant features.

[0089]

[0090]

[0091] in, represents the cosine similarity between modality A and modality B at different positions. Margin is the threshold. The max function ensures that only when the similarity is lower than margin, there will be a positive loss and optimization. At the same time, the mask method is used to reduce the interference of missing modality data. When modality A is missing, the mask m A is 0 when the mode exists, and 1 when the mode does not exist;

[0092] The contrast loss function is:

[0093] l c =m A ·m B ·l + +m A ·l A- +m B ·l B- +m A ·m B ·l AB- .

[0094] Similarly, domain classification is performed to assist in the learning of specific features. In order to further explore the specific features of each modality, the present invention adopts domain classification technology to assist in the learning of specific features.

[0095] Specifically, the domain classification task is used for all existing modal data, which can be understood as predicting the classification accuracy of the modality to which it belongs by learning specific features and using it as the training loss. Its loss function is expressed as:

[0096]

[0097] in, is a remote sensing image dataset of N modalities,

[0098] is the i-th modal data of the j-th sample, y j For its classification label, t i ∈{0,1} N is a one-hot modal label, which is 1 at the i-th position and 0 at the rest of the positions, i = 1, 2, ..., N; represents the specific features of the i-th mode of the j-th sample; f d :S→Δ N-1 is a function that maps a specific feature to an N-1-dimensional probability simplex, where S represents the specific feature space and Δ N-1 represents a probability simplex with N classes.

[0099] Finally, the model performs total loss optimization. The total loss l of the model is composed of three parts: the main classification task loss, the contrastive learning task loss, and the domain classification task loss. Therefore, the formula for the total loss l is expressed as:

[0100] l=μ cls ·l cls +μ c ·l c +μ d ·l d

[0101] Among them, μ cls , μ c and μ d They are the loss weights for the main classification task, contrastive learning task, and domain classification task, respectively. The model parameters are optimized through back propagation to minimize the total loss function.

[0102] like Figure 4 As shown in the figure, (a) shows the MMANet network training test, (b) shows the mmFormer network training test, and (c) shows the training test of the network of the present invention. In the case of full modality training, the model prediction accuracy of using only single modality, dual modality and tri-modality testing respectively. Figure 6As shown in the figure, it can be seen that the prediction accuracy of single modality training is significantly lower than that of dual-modality training, while the prediction accuracy of tri-modal data training is higher than that of dual-modality training. This shows that each modality has unique data characteristics, and multi-modal fusion can make more comprehensive use of these characteristics, thereby improving the prediction performance of the model.

[0103] like Figure 5 As shown, (a) shows the training and testing results of the MMANet network, (b) shows the training and testing results of the mmFormer network, and (c) shows the training and testing results of the network of the present invention. Figure 7 It shows that the model can accurately predict the model using data from each modality when the modality is randomly missing during training. Figure 7 The comparison results show that, in the case of random modality loss, although the model performance is reduced, the model of the present invention can still maintain a high classification accuracy. This shows that the model of the present invention has good robustness when dealing with the modality loss problem.

[0104] In addition, the comparative experiments with the existing methods show that the model of the present invention has obvious advantages in the multimodal remote sensing image classification task. The MMANet model and the mmFormer model only use the specific features of each modality, but fail to fully explore the correlation between the shared features of each modality, so the classification effect is not as good as our model. In contrast, the model of the present invention improves the classification performance by deeply exploring the shared features and specific features of each modality.

[0105] In summary, no matter in the complete modality case or in the missing modality case, the model of the present invention demonstrates extremely high classification accuracy and robustness.

[0106] In summary, no matter in the complete modality case or in the missing modality case, the model of the present invention demonstrates extremely high classification accuracy and robustness.

[0107] Although the present invention has been described above with preferred embodiments, it is not intended to limit the present invention. A person skilled in the art of the present invention may make various modifications and improvements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the claims.

Claims

1. A multimodal remote sensing image fusion classification method under modality missing conditions, characterized in that: The specific features of each modality are extracted using a specific encoder and the shared features between modalities are extracted using a shared encoder. At the same time, two auxiliary tasks are used to learn the specific features and shared features respectively, including the following steps: S1. Acquire multimodal data covering the same area to construct a multimodal dataset and perform preprocessing. The multimodal dataset includes: three types of image data, namely, hyperspectral image data, DSM image data, and SAR image data, or a number of real pixels extracted from one or two of the image data, and corresponding classification labels; S2, taking the three image data, the hyperspectral image data, the DSM image data and the SAR image data, or a number of real pixels extracted from one or two of the image data as input, and corresponding specific features as output, to construct a specific feature encoder; The specific feature encoder extracts specific features from a number of real pixels extracted from three types of image data, namely, the hyperspectral image data, the DSM image data, and the SAR image data, or from one or two of the image data, as shown in the following formula: Among them, X i represents the pixel of hyperspectral image data, DSM image data or SAR image data, i∈{1,2,3}, f specific represents a specific feature encoder; The three image data of hyperspectral image data, DSM image data and SAR image data or one or two of them are connected as input, and the shared features corresponding to the respective modalities are obtained as output to construct a shared feature encoder; the shared feature encoder extracts the shared features corresponding to the respective modalities of the three image data of hyperspectral image data, DSM image data and SAR image data or one or two of them, as shown in the following formula: Among them, X i represents the pixel of hyperspectral image data, DSM image data or SAR image data, i∈{1,2,3}, f shared represents the shared feature encoder; S3, using specific features and shared features as input and final features as output, constructing a residual fusion module; specifically including the following sub-steps: S3.

1. Residual fusion module for specific features and shared features After concatenation in the channel dimension and projection operation, its output is added to the shared features as a residual to form a more semantically rich modality embedding f i The following expression: Among them, f proj represents the projection operation. If one modality is missing, the embedding of the missing modality is obtained by averaging the shared features of the remaining existing modalities. S3.

2. Embed the obtained modes of each mode into f i Concatenate to form the final feature representation f; The construction parts of the specific encoder and the shared encoder are the same. They first extract local features by three 3×3 convolution blocks. Each convolution contains a convolution layer, a batch normalization layer, and a ReLU activation function. A variety of convolution operations of different sizes are used, including 1×1 convolution, 3×3 convolution, and global average pooling to further extract multi-scale features and capture multi-level spatial information. Finally, the features are reduced in dimension through 1×1 convolution and pooling operations, and the final representation capability is enhanced by splicing the features of different convolution branches. S4, construct a prediction module with the final features as input and the corresponding classification labels as output; S5, constructing a multimodal remote sensing image fusion classification network based on the specific feature encoder, shared feature encoder, residual fusion module, and prediction module of step S1 to step S4, training the multimodal remote sensing image fusion classification network with the multimodal data set of step S1, and obtaining a multimodal remote sensing image fusion classification model; when training the multimodal remote sensing image fusion classification network, performing total loss optimization, including minimizing the distance between similar features at the same position of different modalities by using a contrast loss function, and maximizing the difference between features at different positions; using a domain classification loss function to assist in the learning of specific features; and using a main classification task loss function to guide the prediction results; S6. Input three or one or two of the collected hyperspectral image data, DSM image data, and SAR image data into a multimodal remote sensing image fusion classification model to obtain a classification label, that is, a classification result of the terrain.

2. According to the multi-modal remote sensing image fusion classification method under modality loss conditions of claim 1, it is characterized in that: Step S1 includes the following sub-steps: S1.

1. Obtain multimodal data covering the same area, including hyperspectral image data, DSM image data and SAR image data, denoted as X1∈R C×H×W ,X2∈R H×W and X3∈R 4×H×W , where W and H represent the width and height, and C represents the number of spectral bands of the hyperspectral image; S1.

2. Data preprocessing: Perform principal component analysis and dimensionality reduction on the hyperspectral image, extract the first d principal components, and the generated hyperspectral data after dimensionality reduction is represented as X1′∈R d×H×W .

3. The multimodal remote sensing image fusion classification method under modality missing conditions according to claim 2 is characterized in that: In step S4, the prediction module inputs the final feature f into the fully connected layer and generates the prediction result y using the weight of the fully connected layer, as shown in the following expression: y=Softmax(f fc (cat(f 1 ,f 2 ,f 3 );W)), Among them, f fc (·) represents the fully connected layer operation, and W is the connection weight corresponding to the fully connected layer operation.

4. The multimodal remote sensing image fusion classification method under modality loss conditions according to claim 1 is characterized in that: In step S4, The main classification task loss function is as follows: Where H(·,·) is the cross entropy loss function, is the true label and y is the predicted result.

5. The multi-modal remote sensing image fusion classification method under modality loss condition according to claim 4 is characterized in that: In step S4, the contrast loss includes: the loss of the same position pair of different modalities l + , the loss of different position pairs of the same modality l A- , l B- and the loss l for different position pairs across modalities AB- As follows: in, represents the cosine similarity between modality A and modality B at different positions. Margin is the threshold. The max function ensures that only when the similarity is lower than margin, there will be a positive loss and optimization. At the same time, the mask method is used to reduce the interference of missing modality data. When modality A is missing, the mask m A is 0 when the mode exists, and 1 when the mode does not exist; The contrast loss function is: l c =m A ·m B ·l + +m A ·l A- +m B ·l B- +m A ·m B ·l AB- 。 6. The multi-modal remote sensing image fusion classification method under modality loss conditions according to claim 5 is characterized in that: In step S4, the domain classification loss function is: in, is a remote sensing image dataset of N modalities, is the i-th modal data of the j-th sample, y j For its classification label, t i ∈{0, 1} N is a one-hot modal label, which is 1 at the i-th position and 0 at the rest of the positions, i = 1, 2, ..., N; represents the specific features of the i-th mode of the j-th sample; f d :S→△ N-1 is a function that maps a specific feature to an N-1-dimensional probability simplex, where S represents the specific feature space, △ N-1 represents a probability simplex with N classes.

7. The multimodal remote sensing image fusion classification method under modality missing conditions according to claim 6 is characterized in that: In step S4, the total loss function is expressed as: l=μ cls ·l cls +m c ·l c +m d ·l d , Among them, μ cls , μ c and μ d They are the loss weights for the main classification task, contrastive learning task, and domain classification task, respectively. The model parameters are optimized through back propagation to minimize the total loss function.

Citation Information

Patent Citations

  • Multi-mode cross-subject emotion recognition method and system, electronic equipment and medium

    CN117272227A

  • Robustness ground feature classification method and system based on multi-modal relation modeling and fusion network

    CN119339226A