Ground feature classification method and system based on uncertainty weighted multi-modal fusion network
By constructing an uncertainty-weighted multimodal fusion network, the problem of uncertainty estimation and interpretability of modal information in hyperspectral images and arbitrary complementary modal fusion is solved, and efficient multimodal feature fusion and robust geographic classification are achieved.
Patent Information
- Application Number
- CN202510262958.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art lacks uncertainty estimation of modal information in the fusion of hyperspectral images and any supplementary modality, ignores the uneven utilization of modal information and the lack of interpretability, and fails to effectively deal with the correlation noise caused by inter-class similarity and intra-class differences in remote sensing images, affecting the universality and robustness of multi-modal fusion.
A multimodal fusion network based on uncertainty weighted is constructed, and features are extracted by expressing enhanced modules, and features are extracted by channel dimension and spatial dimension feature extraction modules. Combining the intermodal feature fusion module and decision-level fusion module, the Dirichlet framework is used to estimate uncertainty to achieve uncertainty-weighted decision-level fusion.
It realizes effective utilization of modal information, filters between classes and differential noises within classes, improves the universality and robustness of multimodal fusion, and obtains more reliable classification decisions.
Smart Images

Figure CN120298875A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and deep learning, and particularly relates to a method and system for ground object classification based on an uncertainty-weighted multi-modal fusion network. Background Art
[0002] Hyperspectral (HS) images contain extensive information about surface cover and land use, and play an important role in various tasks based on geospatial analysis. However, in complex remote sensing scenarios, the low spatial resolution of HS images and the inherent limitations of imaging technologies pose certain challenges to achieving high-quality, fine-grained analysis. To address these issues, deep learning-based multi-modal fusion methods have been continuously developed, aiming to combine hyperspectral images with complementary modalities to enhance the accuracy of surface cover and land use analysis. Commonly used complementary modalities include multi-spectral (MS) data, light detection and ranging (LiDAR) data, and synthetic aperture radar (SAR) data. However, these methods are usually designed for the characteristics of a certain complementary modality and are difficult to be extended to the fusion tasks of other complementary modalities and hyperspectral images. Therefore, under the application requirements of fusing hyperspectral images with any complementary modality, a large amount of computing resources and engineering efforts are required. In addition, due to the singularity of the applicable objects, these methods ignore the potential knowledge conversion relationships between different complementary modalities. Therefore, how to implement a general ground object classification method based on multi-modal remote sensing data fusion to meet the fusion requirements of hyperspectral and any complementary modality is worthy of further exploration.
[0003] For the high-precision land cover analysis task of hyperspectral images, Roy et al. designed a Transformer network for multimodal fusion, using the feature information from other complementary modalities in the encoder to achieve better generalization. Wang et al. proposed a multilevel information complementary fusion network based on a flexible hybrid strategy for land cover classification analysis of HS image and complementary modality fusion. In addition, Zhang et al. proposed a fusion segmentation framework based on cross-modal attention perception from local to global, adopting two parallel semantic segmentation backbone architectures to learn information from HS and complementary modalities. Although the above methods have made significant contributions to the research of general frameworks for hyperspectral and arbitrary complementary modality fusion, they still face the following problems: (1) Lack of uncertainty estimation for modal information. Traditional multimodal fusion methods usually assume that information from different modalities is equally important, or simply set a fixed weight factor for each modal feature. This easily leads to inappropriate and unbalanced utilization of modal information, restricting the effective utilization of all available data. In addition, due to the lack of uncertainty estimation, the obtained prediction results are often unreliable in the case of modal missing or data quality. (2) Lack of interpretability. The current methods mainly perform multimodal feature extraction and fusion through implicit deep networks, resulting in these methods mainly relying on training data and lacking explicit guidance, which restricts the generality and robustness of the methods. More importantly, in a unified framework, through interpretable multimodal feature processing, knowledge from a certain modality can be systematically identified, learned, and effectively retained, thereby further promoting the understanding of another modality. For example, a unified fusion framework trained based on HS and SAR images is also helpful for land cover classification based on the fusion of HS and MS images. Therefore, it is of great significance to explore a unified multimodal learning framework for HS and arbitrary complementary modalities under explicit guidance. (3) Ignoring the feature correlation relationship. Different from natural images, one of the challenges in remote sensing image analysis is the high similarity between land cover classes and large intra-class differences. In addition, in multimodal learning, there are large inherent gaps between various modalities, and the feature embedding space of each modality contains potential noise, which will disrupt the multimodal interaction process.
[0004] Therefore, it is necessary to design a land cover classification method and system based on an uncertainty-weighted multimodal fusion network for the above problems. Summary of the Invention
[0005] The object of the present invention is to address the problems existing in the prior art, and provide a method and system for land cover classification based on an uncertainty-weighted multi-modal fusion network. Aiming at the problems of lack of uncertainty estimation of modal information and lack of interpretability, by constructing an expression enhancement module to filter the correlation noise caused by inter-class similarity and intra-class difference, and using a channel dimension feature extraction module and a spatial dimension feature extraction module to extract corresponding features; under the explicit guidance of a fusion-guided loss function, a fusion feature that simultaneously preserves spectral features and spatial structure features is obtained to achieve general and robust multi-modal fusion; through an uncertainty-weighted decision-level fusion strategy, the uncertainty of single-modal features and fusion feature evidence is estimated based on the Dirichlet framework and used as a weight to control the importance of different components in the final decision, so as to achieve a more reliable classification decision.
[0006] According to one aspect of this specification, a method for land cover classification based on an uncertainty-weighted multi-modal fusion network is provided, including:
[0007] Obtain multi-modal remote sensing image data;
[0008] Input the obtained multi-modal remote sensing image data into a trained land cover classification model to obtain the land cover classification result of the multi-modal remote sensing image; wherein, the training of the land cover classification model includes:
[0009] Construct a multi-modal remote sensing image dataset of the same detection target or scene;
[0010] Construct a land cover classification model based on an uncertainty-weighted multi-modal fusion network, including: a feature embedding module for generating a feature embedding space for the input multi-modal data to obtain embedding features; an expression enhancement module for enhancing the feature expression of the embedding features; a channel dimension feature extraction module for extracting channel dimension features in the channel dimension for the enhanced features of hyperspectral; a spatial dimension feature extraction module for extracting spatial dimension features in the spatial dimension for the enhanced features of supplementary modalities; an inter-modal feature fusion module for fusing the channel dimension features and the spatial dimension features; a decision-level fusion module for estimating the uncertainty of the channel dimension features, the spatial dimension features and the fusion feature evidence to obtain the land cover classification result;
[0011] Perform model training based on the constructed dataset to obtain a trained land cover classification model.
[0012] Furthermore, enhancing the feature expression of the embedding features includes:
[0013] A feature expansion process: by modeling the high-order correlation between features, enhancing the non-linear expression ability of the features;
[0014] Internal feature fusion process: By explicitly utilizing the correlation information between features, the discriminative ability of features within a modality is enhanced.
[0015] Furthermore, the construction of the inter-modal feature fusion module includes:
[0016] Calculate the cosine similarity matrix for the input channel dimension features and spatial dimension features;
[0017] Multiply the cosine similarity matrix row by row and column by column with the channel dimension features and spatial dimension features respectively to obtain channel fusion features and spatial fusion features, and then concatenate them;
[0018] Input the concatenated fusion features into a convolutional layer, output the fusion features, and complete the construction of the inter-modal feature fusion module.
[0019] Furthermore, the construction of the decision-level fusion module includes:
[0020] Based on the Dirichlet distribution framework, estimate the uncertainties of the evidence of channel dimension features, spatial dimension features, and fusion features;
[0021] Use the estimated uncertainty results as weights to achieve evidence fusion, obtain the ground object classification result, and complete the construction of the uncertainty-weighted decision-level fusion module.
[0022] Furthermore, the ground object classification model also includes a loss function, which is divided into two parts: the first part is the fusion guidance loss function, and the second part is the mean square loss function.
[0023] According to one aspect of this specification, a ground object classification system based on an uncertainty-weighted multi-modal fusion network is provided, including:
[0024] A data acquisition module for acquiring multi-modal remote sensing image data;
[0025] A ground object classification module for inputting the acquired multi-modal remote sensing image data into the trained ground object classification model to obtain the ground object classification result of the multi-modal remote sensing image; wherein, the training of the ground object classification model includes:
[0026] Construct a multi-modal remote sensing image dataset of the same detection target or scene;
[0027] Construct a ground object classification model based on an uncertainty-weighted multimodal fusion network, including: a feature embedding module for generating a feature embedding space for the input multimodal remote sensing image data to obtain embedded features; an expression enhancement module for enhancing the feature expression of the embedded features; a channel dimension feature extraction module for extracting channel dimension features in the channel dimension for the enhanced features of hyperspectral images; a spatial dimension feature extraction module for extracting spatial dimension features in the spatial dimension for the enhanced features of supplementary modalities; an inter-modal feature fusion module for fusing the channel dimension features and the spatial dimension features; a decision-level fusion module for estimating the uncertainties of the channel dimension features, the spatial dimension features, and the fusion feature evidence to obtain a ground object classification result;
[0028] Train the model based on the constructed dataset to obtain a trained ground object classification model.
[0029] According to one aspect of this specification, there is provided an electronic device including a memory and a processor, the memory storing a computer program, characterized in that when the processor executes the computer program, the steps of the ground object classification method based on the uncertainty-weighted multimodal fusion network are implemented.
[0030] According to one aspect of this specification, there is provided a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the ground object classification method based on the uncertainty-weighted multimodal fusion network are implemented.
[0031] According to one aspect of the specification of the present invention, there is provided a computer program product containing instructions, which when run on a computer, causes the computer to execute the steps of the ground object classification method based on the uncertainty-weighted multimodal fusion network.
[0032] Compared with the prior art, the beneficial effects of the present invention are:
[0033] 1. Aiming at the problem of lack of uncertainty estimation of modal information, the present invention uses the feature correlation information through the expression enhancement module to enhance the feature expression, thereby filtering out the correlation noise caused by inter-class similarity and intra-class difference, and avoiding the damage to the multimodal interaction process caused by the noise in the single-modal feature embedding space.
[0034] 2. In view of the problem of lack of interpretability, the present invention uses a channel - dimension feature extraction module for hyperspectral features to extract the channel - dimension features of hyperspectral data, and uses a spatial - dimension feature extraction module for other supplementary modality features to extract the spatial - dimension features of other supplementary modalities; under the explicit guidance of a fusion - guiding loss function, the channel - dimension features and spatial - dimension features are fused based on an inter - modality feature fusion module to obtain a fused feature that retains both spectral features and spatial - structure features, realizing multi - modality fusion with generality and robustness.
[0035] 3. The present invention proposes an uncertainty - weighted decision - level fusion strategy. Based on the Dirichlet framework, the uncertainty of single - modality features and fused - feature evidence is estimated and used as a weight to control the importance of different components in the final decision, realizing a more reliable classification decision. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0037] Figure 1 It is the overall process framework diagram of the embodiment of the present invention;
[0038] Figure 2 It is the structural schematic diagram of the expression enhancement module of the embodiment of the present invention;
[0039] Figure 3 It is the structural schematic diagram of the channel - dimension feature extraction module of the embodiment of the present invention;
[0040] Figure 4 It is the structural schematic diagram of the spatial - dimension feature extraction module of the embodiment of the present invention;
[0041] Figure 5 It is the structural schematic diagram of the inter - modality feature fusion module of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0043] Such as Figure 1As shown in the figure, the embodiment of the present invention also provides an overall process framework diagram of a ground object classification method based on an uncertainty-weighted multimodal fusion network, which specifically includes: obtaining multimodal remote sensing image data of the same detection target or scene, and preprocessing the remote sensing image data of each modality; building a ground object classification model, including: a feature embedding module, which is used to generate a feature embedding space for the input multimodal remote sensing image data to obtain embedded features; an expression enhancement module, which is used to enhance the feature expression of the embedded features to obtain enhanced features; a channel dimension feature extraction module, which is used to extract channel dimension features in the channel dimension for the enhanced features of hyperspectral images; a spatial dimension feature extraction module, which is used to extract spatial dimension features in the spatial dimension for the enhanced features of supplementary modalities; an inter-modal feature fusion module, which is used to fuse the channel dimension features and the spatial dimension features under the explicit guidance of a loss function; an uncertainty-weighted decision-level fusion module, which is used to estimate the uncertainties of the channel dimension features, the spatial dimension features, and the evidence of the fusion features based on the Dirichlet distribution framework, and use them as weights to control the importance of different components in the final decision to obtain the ground object classification result; inputting the preprocessed remote sensing data of each modality into the built ground object classification model for training to obtain a trained ground object classification model.
[0044] Specifically, the embodiment of the present invention also provides step 1, obtaining multimodal remote sensing data of the same detection target or scene, and performing normalization processing on the remote sensing data of each modality. The normalization processing refers to linear normalization processing, which is used to scale the data to within the interval, which can be expressed as:
[0045] (1)
[0046] where is the original remote sensing data of modality input into the network, are the minimum and maximum values in the original remote sensing data of modality respectively, is the remote sensing data of modality after normalization. Then, the normalized remote sensing data is cropped into a data block centered on one pixel with a size of .
[0047] Specifically, the embodiment of the present invention also provides step 2, constructing a feature embedding module. The feature embedding module is composed of 4 convolutional modules connected in sequence, each of which consists of a convolutional layer, a BatchNorm normalization layer, and a ReLU activation function layer. Among them, the ReLU activation function is:
[0048] (2)
[0049] where Represents the input features of the ReLU activation function.
[0050] Specifically, in the feature embedding module, the normalized and clipped data block is used as the input, and after passing through the convolutional module operation, the embedded features are obtained. The implementation process can be expressed as:
[0051] (3)
[0052] Where is the modality remote sensing data after clipping, and the data block has a size of , is the feature embedding module of modality , is the modality embedding feature of modality obtained after the operation.
[0053] Specifically, the embodiment of the present invention also provides step 3, constructing an expression enhancement module, as Figure 2 shown. The implementation process of the expression enhancement module includes a feature expansion process and an internal feature fusion process. The feature expansion process enhances the non-linear expression ability of features by modeling the high-order correlations between features. The internal feature fusion process improves the discriminative ability of features within the modality by explicitly using the correlation information between features.
[0054] Specifically, in the expression enhancement module, the modality embedding features of modality are used as the input, and after passing through the feature expansion process, the features with expanded dimensions are obtained. The implementation process can be expressed as:
[0055] (4)
[0056] (5)
[0057] (6)
[0058] Where represents the training data set, which contains modalities, and each modality consists of samples. is the th sample of modality , is the ground truth corresponding to this sample, that is, the corresponding land cover class. is the mapping function used to implement the expansion of the feature dimension. , is the th The modal embedding features of a sample, with a feature dimension of , that is, the feature contains sub - features . Modal The expanded features of the th sample have a feature dimension of , that is, the feature contains sub - features . Modal The expanded features of all samples form an expanded feature set . The mapping function used to implement the expansion of the feature dimension The specific implementation process can be expressed as:
[0059] (7)
[0060] (8)
[0061] Among them, is the expansion coefficient of the feature dimension, is the th sub - feature in the expanded features of the th sample of modal . is the th sub - feature in the modal embedding features of the th sample of modal, and r is the calculation parameter in the mapping function
[0062] Specifically, in the expression enhancement module, taking the expanded features as input, through the internal feature fusion process, enhanced features are obtained, and the implementation process can be expressed as:
[0063] (9)
[0064] (10)
[0065] Among them, is the mapping function used to implement internal feature fusion, and the feature dimension of the enhanced feature of the th sample of modal is , that is, the feature contains sub - features . The enhanced features of all samples of modal form a feature set . The mapping function used to implement internal feature fusion can be specifically expressed as:
[0066] (9)
[0067] (10)
[0068] Among them, is the feature correlation matrix calculated based on the features extended by the input expression enhancement module, and its size is . The matrix element represents the Euclidean distance between the sub-features and in the feature, and . is the j-th sub-feature of the enhanced feature of the th sample in the modality , is the j-th sub-feature of the extended feature of the th sample in the modality ; represents the th column of the feature correlation matrix ,
[0069]
[0070] (11) Among them, is the length of the sub-feature, and are respectively the th sub-feature and the th sub-feature of the th eigenvalue in the extended feature. is the length of the sub-feature in the extended feature, that is, the total number of eigenvalues.
[0071] Specifically, the embodiment of the present invention further provides step 4, constructing a channel dimension feature extraction module, as Figure 3 shown. The enhanced feature of the hyperspectral modality is used as the input, and is simultaneously calculated through an average pooling layer and a maximum pooling layer, and the calculation results are added and then sent into a module composed of a fully connected layer, a ReLU activation function layer, a fully connected layer, and a Sigmoid activation function layer in sequence. After the calculation result is multiplied by the result after the input enhanced feature passes through the residual branch, it is concatenated with the input enhanced feature, and after passing through a convolutional layer, the channel dimension feature is output. The said residual branch includes Convolution layer, ReLU activation function layer, and BatchNorm normalization layer. The Sigmoid activation function is as follows:
[0072] (12)
[0073] where represents the input feature of the Sigmoid activation function.
[0074] Specifically, the embodiment of the present invention further provides step 5 of constructing a spatial dimension feature extraction module, as Figure 4 shown. The enhanced features of the supplementary modality are used as the input, and are simultaneously calculated through an average pooling layer and a max pooling layer, and the calculation results are added and then sent into a module successively composed of a convolution layer, a ReLU activation function layer, a convolution layer, and a Sigmoid activation function layer. After the calculated result is multiplied by the result of the input enhanced feature passing through the residual branch, it is concatenated with the input enhanced feature, and after passing through a convolution layer, the spatial dimension feature is output. The residual branch includes a convolution layer, a ReLU activation function layer, and a BatchNorm normalization layer.
[0075] Specifically, the embodiment of the present invention further provides step 6 of constructing an inter-modal feature fusion module, as Figure 5 shown. Under the explicit guidance of the fusion guidance loss function, the channel dimension feature and the spatial dimension feature are fused to obtain a fusion feature that simultaneously retains spectral features and spatial structure features.
[0076] Specifically, the cosine similarity matrix is calculated for the channel dimension feature and the spatial dimension feature of the input module. The calculation process of the cosine similarity matrix can be expressed as:
[0077] (13)
[0078] (14)
[0079] where is the calculated cosine similarity matrix, and its size is . is the cosine similarity calculation for the sub-feature in the channel dimension feature and the sub-feature in the spatial dimension feature , and .
[0080] Specifically, the cosine similarity matrix is multiplied row by row and column by column with the channel dimension features and the spatial dimension features respectively to obtain the channel fusion features and the spatial fusion features. The calculation process can be expressed as:
[0081] (15)
[0082] (16)
[0083] Among them, and are the th sub-features of the channel fusion features and the spatial fusion features respectively. and represent the th row and the th column of the cosine similarity matrix is 's transpose matrix.
[0084] Specifically, the concatenated channel fusion features and spatial fusion features are fed into a convolutional layer, and then the fused features are output.
[0085] Specifically, it is continuously optimized under the explicit guidance of the fusion guidance loss function. The fusion guidance loss function consists of two parts, which are used to guide the preservation of spectral features and the preservation of spatial structure features respectively. The calculation process can be expressed as:
[0086] (17)
[0087] (18)
[0088] Among them, is the fused feature output by the inter-modal feature fusion module, and its size is . represents all the features in the th channel dimension of , represents all the features in the th spatial dimension of . is all the features in the th channel dimension of the channel dimension feature , is all the features in the th spatial dimension of the spatial dimension feature ; is the correlation coefficient between the input features , and ; Represents the input features calculated based on the correlation coefficient The feature distance between them. That is The All features in the The Feature distance between all features in the channel dimension of and That is The All features in the The Feature distance between all features in the channel dimension of and Represents the L2 norm calculation, Is the coefficient used to balance the convergence degree of the two parts.
[0089] Specifically, the embodiment of the present invention also provides step 7, constructing an uncertainty-weighted decision-level fusion module. Based on the Dirichlet framework, taking the fusion feature, channel dimension feature, and spatial dimension feature as inputs, estimating the uncertainties of the channel dimension feature, spatial dimension feature, and fusion feature evidence, and using them as weights to achieve evidence fusion, obtaining the ground object classification result.
[0090] Specifically, build the Dirichlet framework based on the Dirichlet distribution, which can be expressed as:
[0091] (19)
[0092] (20)
[0093] Among them, , is the distribution parameter. Is a K-dimensional polynomial function, and K is the total number of classification categories. The implementation process of the uncertainty estimation of the channel dimension feature, spatial dimension feature, and fusion feature evidence based on the Dirichlet framework can be expressed as:
[0094] (21)
[0095] (22)
[0096] (23)
[0097] (24)
[0098] Among them, Is the uncertainty, Is the category The confidence level, , is the confidence vector. Represents evidence, which is a feature quantity indicating that a sample is classified into a certain class. , is the Dirichlet strength. Denotes the predicted probability of the th class; is the
[0099] th item of the distribution parameter.
[0100] (25)
[0101] (26)
[0102] (27)
[0103] (28)
[0104] Among them, , and respectively represent the confidence vectors of the channel dimension feature, the spatial dimension feature, and the fusion feature evidence. , and are the uncertainties of the channel dimension feature, the spatial dimension feature, and the fusion feature evidence respectively. and are the distribution parameter, the confidence vector, and the uncertainty after fusion respectively, is the normalization coefficient. is the predicted probability vector, that is, the ground object classification result.
[0105] Specifically, the embodiment of the present invention also provides step 8, constructing an uncertainty-weighted multi-modal fusion network for high-precision and robust ground object classification, as Figure 1 shown. The specific steps are as follows: Using the feature embedding module, the expression enhancement module, the channel dimension feature extraction module, the spatial dimension feature extraction module, and the inter-modal feature fusion module described in step 2, step 3, step 4, step 5, and step 6, adding the uncertainty-weighted decision-level fusion module described in step 7, to obtain the ground object classification result.
[0106] Specifically, the multi-modal normalized and trimmed data blocks are respectively sent into the feature embedding modules of the hyperspectral branch and the supplementary modality branch to obtain the embedding features of each modality, that is, Figure 2 in .
[0107] Specifically, the embedded features obtained by the feature embedding module for each input modality are sent to the expression enhancement module described in step 3 to obtain the enhanced features of each modality, that is, Figure 2 middle .
[0108] Specifically, the enhanced features of the hyperspectral branch and the enhanced features of the complementary modal branch, i.e. Figure 3 In and Figure 4 In , respectively sent to the channel dimension feature extraction module and the space dimension feature extraction module described in step 4 and step 5 to obtain the channel dimension feature and the space dimension feature, that is, Figure 3 In and Figure 4 In .
[0109] Specifically, the obtained channel dimension features and spatial dimension features are sent to the inter-modal feature fusion module described in step 6. Under the explicit guidance of the fusion guidance loss function, the channel dimension features and spatial dimension features are fused to obtain a fusion feature that retains both spectral features and spatial structure features, that is, Figure 5 In .
[0110] Specifically, the fused features, channel dimension features and space dimension features are sent to the uncertainty weighted decision-level fusion module described in step 7 to estimate the uncertainty of the channel dimension features, space dimension features and fused feature evidence, and use them as weights to achieve evidence fusion and obtain the object classification result.
[0111] Specifically, the embodiment of the present invention further provides step 9, designing a loss function suitable for the network. It includes two parts: one is the fusion guided loss function , and the second is to calculate the accuracy of the network output ground object classification prediction results based on the mean square loss function, which can be expressed as:
[0112] (29)
[0113] in, The corresponding sample The ground-truth classification value encoded by one-hot vector, express The value corresponding to the jth category in . It is a sample The corresponding K-dimensional polynomial function, Representation sample The K-dimensional probability vector, yes The differential of denotes the variance of the probability corresponding to the j-th category in is the sample The predicted probability of the j-th class corresponding to. Therefore, the final network loss function can be expressed as:
[0114] (30)
[0115] Among them, is a constant used to define the weight of the fusion guidance loss function, which is set to 0.01 in this embodiment.
[0116] Specifically, for various data situations of complete modalities and missing modalities, training is respectively carried out on the multimodal remote sensing datasets Houston2013 composed of hyperspectral and multispectral data, the multimodal remote sensing dataset Berlin composed of hyperspectral and synthetic aperture radar data, and the multimodal remote sensing dataset MUUFL composed of hyperspectral and lidar data, and the land cover classification of remote sensing images is carried out according to the obtained uncertainty weighted multimodal fusion unified model.
[0117] Specifically, based on the land cover classification results obtained in steps 1-step 9 on the multimodal remote sensing datasets Houston2013 composed of hyperspectral and multispectral data, the multimodal remote sensing dataset Berlin composed of hyperspectral and synthetic aperture radar data, and the multimodal remote sensing dataset MUUFL composed of hyperspectral and lidar data, in order to compare with other methods, we use six multimodal remote sensing image land cover classification methods: FusAtNet, S 2 ENet, MFT, Flex-MCFNet, Cross-HL, DSymFuser are compared with the method of the present invention on the above datasets.
[0118] Specifically, in order to quantitatively evaluate the land cover classification results, we introduce the Overall Accuracy (OA), Average Accuracy (AA) and Kappa coefficient to evaluate the land cover classification accuracy of remote sensing images. The larger the index value, the higher the classification accuracy. In the case of complete modalities, the quantitative comparison results on the Houston2013 dataset are shown in Table 1.
[0119] Table 1 Quantitative analysis of different methods on the Houston2013 dataset
[0120]
[0121] Specifically, the quantitative comparison results on the Berlin dataset are shown in Table 2.
[0122] Table 2 Quantitative analysis of different methods on the Berlin dataset
[0123]
[0124] Specifically, the quantitative comparison results on the MUUFL dataset are shown in Table 3.
[0125] Table 3 Quantitative analysis of different methods on the MUUFL dataset
[0126] Specifically, in the case of missing modalities, the quantitative comparison results on the three datasets of Houston2013, Berlin, and MUUFL are shown in Table 4.
[0127] Table 4 Quantitative analysis of different methods on the three datasets
[0128]
[0129] Specifically, in order to verify that in the unified framework, through interpretable multi-modal feature processing, knowledge from a certain modality can be systematically identified, learned, and effectively retained, thereby further promoting the understanding of another modality. We used the pre-trained parameters obtained by training on the hyperspectral and other complementary modality datasets as the initialization parameters, and the prediction results on the three datasets of Houston2013, Berlin, and MUUFL are shown in Table 5. Pre on H means using the pre-trained parameters obtained by training on the Houston2013 dataset as the initialization parameters, Pre on B means using the pre-trained parameters obtained by training on the Berlin dataset as the initialization parameters, and Pre on M means using the pre-trained parameters obtained by training on the MUUFL dataset as the initialization parameters. w / o pre means not using pre-trained parameter initialization. Bold indicates that the ground object classification accuracy is higher than w / o pre.
[0130] Table 5 Mutual promotion analysis of the method of the present invention on the three datasets
[0131]
[0132] Specifically, the embodiments of the present invention also provide a representation of quantitative index results. For various data situations of complete modalities and missing modalities, the ground object classification results obtained by the method proposed by the present invention are generally superior to existing methods on the multi-modal remote sensing datasets Houston2013 composed of hyperspectral and multispectral data, the multi-modal remote sensing dataset Berlin composed of hyperspectral and synthetic aperture radar data, and the multi-modal remote sensing dataset MUUFL composed of hyperspectral and lidar data, and can meet the application requirements of high-precision and robust ground object classification for the fusion of hyperspectral and any supplementary modality. Moreover, under a unified framework, through interpretable multi-modal feature processing, the method proposed by the present invention can systematically identify, learn, and effectively retain the knowledge from a certain modality, thereby further promoting the understanding of another modality.
[0133] The implementation basis of each embodiment of the present invention is achieved through programmed processing by a device with processor functions. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this actual situation, on the basis of the above embodiments, the embodiments of the present invention provide a ground object classification system based on an uncertainty-weighted multi-modal fusion network, and this system is used to execute the ground object classification method based on the uncertainty-weighted multi-modal fusion network in the above method embodiments.
[0134] The system includes: a data acquisition module, which is used to acquire multi-modal remote sensing image data; a ground object classification module, which is used to input the acquired multi-modal remote sensing image data into a trained ground object classification model to obtain the ground object classification results of the multi-modal remote sensing image; wherein, the training of the ground object classification model includes: constructing a multi-modal remote sensing image dataset of the same detection target or scene; constructing a ground object classification model based on an uncertainty-weighted multi-modal fusion network, including: a feature embedding module, which is used to generate a feature embedding space for the input multi-modal remote sensing image data to obtain embedded features; an expression enhancement module, which is used to enhance the feature expression of the embedded features; a channel dimension feature extraction module, which is used to extract channel dimension features in the channel dimension for the enhanced features of hyperspectral; a spatial dimension feature extraction module, which is used to extract spatial dimension features in the spatial dimension for the enhanced features of the supplementary modality; an inter-modal feature fusion module, which is used to fuse the channel dimension features and the spatial dimension features; a decision-level fusion module, which is used to estimate the uncertainties of the channel dimension features, the spatial dimension features, and the fusion feature evidence to obtain the ground object classification results; and performing model training based on the constructed dataset to obtain a trained ground object classification model.
[0135] The ground object classification system based on the uncertainty-weighted multi-modal fusion network provided by the embodiments of the present invention addresses the problems of lack of uncertainty estimation of modal information and lack of interpretability. By using several modules and through a ground object classification model, under the explicit guidance of a fusion guidance loss function, it fuses channel dimension features and spatial dimension features based on an inter-modal feature fusion module to obtain fusion features that simultaneously retain spectral features and spatial structure features, realizing general and robust multi-modal fusion, and performing ground object classification on multi-modal remote sensing image data to be classified.
[0136] Based on the same inventive concept as the foregoing embodiments, the embodiments of the present invention also provide an electronic device, including a memory and a processor. The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement a ground object classification method based on an uncertainty-weighted multi-modal fusion network as proposed in the above embodiments.
[0137] The embodiments of the present invention also provide a computer-readable storage medium with a computer program stored thereon. When the program is executed by a processor, it overcomes the problems of lack of uncertainty estimation of modal information and lack of interpretability, filters out the correlation noise caused by inter-class similarity and intra-class difference, and avoids the damage to the multi-modal interaction process caused by the noise in the single-modal feature embedding space, realizing more reliable classification decisions. The storage medium can be any non-volatile storage device such as a hard disk, a solid-state drive, a flash drive, an optical disc, etc., for storing computer program code and necessary data files. The stored computer program includes: a data acquisition module and a ground object classification module.
[0138] The embodiments of the present invention also provide a computer program product containing instructions that, when running on a computer, wholly or partially generates a ground object classification method based on an uncertainty-weighted multi-modal fusion network as proposed in the above embodiments. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0139] Finally, it should be noted that the above specific embodiments are only relatively representative examples of the present invention. Obviously, the present invention is not limited to the above specific embodiments and there can be many variations. Any simple modification, equivalent change, and modification made to the above specific embodiments based on the technical essence of the present invention should be considered as belonging to the protection scope of the present invention.
Claims
1. A method for ground object classification based on an uncertainty-weighted multimodal fusion network, characterized in that include: Acquire multimodal remote sensing image data; Inputting the acquired multimodal remote sensing image data into the trained land object classification model to obtain land object classification results of the multimodal remote sensing image; wherein the training of the land object classification model includes: Construct a multi-modal remote sensing image dataset of the same detection target or scene; A land object classification model is constructed based on an uncertainty weighted multimodal fusion network, including: a feature embedding module, which is used to generate a feature embedding space for the input multimodal remote sensing image data to obtain embedded features; an expression enhancement module, which is used to enhance the feature expression of the embedded features; a channel dimension feature extraction module, which is used to enhance features for hyperspectral and extract channel dimension features in the channel dimension; a spatial dimension feature extraction module, which is used to enhance features for supplementary modalities and extract spatial dimension features in the spatial dimension; an inter-modal feature fusion module, which is used to fuse channel dimension features and spatial dimension features; a decision-level fusion module, which is used to estimate the uncertainty of channel dimension features, spatial dimension features and fused feature evidence to obtain land object classification results; Model training is performed based on the constructed data set to obtain a trained land feature classification model.
2. The method for classifying ground objects based on an uncertainty weighted multi-modal fusion network according to claim 1, wherein Enhance the feature expression of embedded features, including: Feature expansion process: By modeling high-order correlations between features, the nonlinear expression ability of features is enhanced; Internal feature fusion process: By explicitly utilizing the correlation information between features, the discriminative ability of intra-modal features is improved.
3. The method for classifying ground objects based on an uncertainty-weighted multimodal fusion network according to claim 1, wherein The construction of the inter-modality feature fusion module includes: Calculate the cosine similarity matrix for the input channel dimension features and spatial dimension features; The cosine similarity matrix is used to multiply the channel dimension features and the space dimension features row by row and column by column respectively to obtain the channel blending features and the space blending features, and then the features are spliced; The concatenated fusion features are input into the convolutional layer, and the fusion features are output to complete the construction of the inter-modal feature fusion module.
4. A method for classifying ground objects based on an uncertainty-weighted multi-modal fusion network according to claim 1, characterized in that, The construction of the decision-level fusion module includes: Based on the Dirichlet allocation framework, the uncertainty of channel dimension features, spatial dimension features and fusion feature evidence is estimated; The estimated uncertain results are used as weights to realize evidence fusion, obtain the object classification results, and complete the construction of the uncertainty weighted decision-level fusion module.
5. A method for classifying ground objects based on an uncertainty weighted multi-modal fusion network according to claim 1, characterized in that The object classification model also includes a loss function, which is divided into two parts: the first part is a fusion-guided loss function, and the second part is based on a mean square loss function.
6. A ground object classification system based on an uncertainty weighted multi-modal fusion network, characterized in that, include: A data acquisition module is used to acquire multimodal remote sensing image data; The object classification module is used to input the acquired multimodal remote sensing image data into the trained object classification model to obtain the object classification result of the multimodal remote sensing image; wherein the training of the object classification model includes: Construct a multi-modal remote sensing image dataset of the same detection target or scene; Constructing a ground object classification model based on an uncertainty-weighted multimodal fusion network, including: a feature embedding module for generating a feature embedding space for the input multimodal remote sensing image data to obtain embedded features; an expression enhancement module for enhancing the feature expression of the embedded features; a channel dimension feature extraction module for extracting channel dimension features in the channel dimension for the enhanced features of hyperspectral; a spatial dimension feature extraction module for extracting spatial dimension features in the spatial dimension for the enhanced features of the supplementary modality; an inter-modal feature fusion module for fusing the channel dimension features and the spatial dimension features; a decision-level fusion module for estimating the uncertainties of the channel dimension features, the spatial dimension features, and the fusion feature evidence to obtain a ground object classification result; Training the model based on the constructed dataset to obtain a trained ground object classification model.
7. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the ground object classification method based on the uncertainty-weighted multimodal fusion network according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the ground object classification method based on the uncertainty-weighted multimodal fusion network according to any one of claims 1 to 5.
9. A computer program product comprising instructions, characterized in that, When it runs on a computer, it causes the computer to execute the steps of the ground object classification method based on the uncertainty-weighted multimodal fusion network according to any one of claims 1 to 5.