A multi-modal feature adaptive fusion method based on confidence estimation
By using a multimodal feature adaptive fusion method based on confidence estimation, the modality weights are dynamically adjusted, which solves the quality difference and conflict problems in multimodal fusion, improves classification accuracy and robustness, and adapts to complex remote sensing scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHWEST JIAOTONG UNIV
- Filing Date
- 2025-12-03
- Publication Date
- 2026-07-24
AI Technical Summary
Existing multimodal fusion methods suffer from significant quality differences between modalities, lack of dynamic adaptability, and modal conflict redundancy, leading to decreased classification accuracy and insufficient robustness.
A multimodal feature adaptive fusion method based on confidence estimation is adopted. By constructing a modal feature encoder, an auxiliary classifier and a confidence prediction network, the modal weights are dynamically adjusted for fusion, the influence of low-quality modalities is suppressed and sample-level adaptive fusion is achieved.
It improves the classification accuracy and robustness of multimodal fusion, adapts to different numbers and types of modalities, and enhances the scalability and stability of the model.
Smart Images

Figure CN121564488B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal fusion technology, and in particular to an adaptive fusion method for multimodal features based on confidence estimation. Background Technology
[0002] With the development of remote sensing sensors and geographic information technology, data analysis has gradually shifted from relying on single-modal data analysis to comprehensive interpretation integrating multi-source heterogeneous data. Typical multimodal remote sensing data includes high-resolution optical imagery, synthetic aperture radar (SAR) images, street view images, and points of interest (POIs). Optical imagery provides rich spatial texture and structural features; SAR images offer all-weather, all-time observation capabilities; and POI textual information contains semantic information about regional functions and attributes. The introduction of multimodal data has provided new possibilities for improving the accuracy and stability of remote sensing analysis.
[0003] However, existing multimodal fusion methods still have significant shortcomings. First, there are significant differences in quality between modalities. In different geographical regions or under different observation conditions, some modalities may be noisy or lack information. For example, optical images perform poorly under shadow or fog conditions, while POI text may have annotation errors or semantic redundancy. Second, fusion strategies lack dynamic adaptability. Traditional methods usually use feature splicing, fixed weighting, or simple fusion based on attention mechanisms. These methods fail to explicitly characterize the reliability of each modality at the sample level, resulting in a significant decrease in overall model performance when some modalities fail or their quality fluctuates. At the same time, modal conflicts and redundancy problems are common. The feature spaces of different modalities differ greatly, and direct fusion can easily introduce irrelevant or even contradictory information, affecting the discriminative power of classification.
[0004] Therefore, there is an urgent need for a multimodal fusion framework that can adaptively sense the confidence of each modality and dynamically adjust the fusion weights, so as to enhance the robustness and scalability of the model while ensuring classification accuracy. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a multimodal feature adaptive fusion method based on confidence estimation.
[0006] The objective of this invention is achieved through the following technical solution: a multimodal feature adaptive fusion method based on confidence estimation, comprising the following steps:
[0007] S1: Acquire data input from multiple remote sensing modalities;
[0008] S2: Construct feature encoders for each modality and extract their respective modal feature representations;
[0009] S3: Construct an auxiliary classifier for each modality, predict the class based on the modality features, and output the corresponding class probability distribution;
[0010] S4: Calculate the pseudo confidence level of each modality on the current sample based on the probability distribution output by the classifier;
[0011] S5: Train a modality confidence prediction network using pseudo-confidence as the regression target;
[0012] S6: Concatenate the features of each modality and perform weighted fusion based on their respective confidence weights to obtain the fused feature representation;
[0013] S7: Input the fused feature representation into the downstream recognition or classification module to complete the semantic recognition task of land features or regions.
[0014] Preferably, in step S2, the modal feature encoders are as follows:
[0015] Optical images are used to extract visual semantic features using convolutional neural networks or visual Transformers;
[0016] A pre-trained language model based on Transformer is used to obtain the semantic vectors of the POI text.
[0017] Preferably, in step S4, the formula for calculating the pseudo-confidence level is:
[0018] ;
[0019] in, For the first The pseudo-confidence of each modality For the first Feature representation of each modality The true labels for the samples.
[0020] Preferably, in step S5, the modality confidence prediction network consists of two to three fully connected layers, with BatchNorm and ReLU activation functions set between layers, and sigmoid or scaled linear output used in the output layer to obtain scalar confidence.
[0021] Preferably, in step S5, the optimization objective of the modal confidence prediction network is the mean squared error loss function, and its loss form is:
[0022] ;
[0023] in, To predict confidence levels.
[0024] Preferably, in step S6, feature fusion uses channel-dimensional concatenation, and the formula for calculating the fused feature representation is:
[0025] .
[0026] The present invention has the following advantages:
[0027] 1. This invention guides the confidence prediction network with pseudo-confidence and adaptively adjusts the modality weights according to the samples, thereby suppressing the negative impact of low-quality modalities on fusion, and thus being compatible with different numbers and types of modalities, improving scalability and adapting to complex remote sensing scenarios.
[0028] 2. The present invention provides a fusion mechanism that can adapt to different modal qualities and types, and achieves dynamic and scalable multimodal feature fusion by explicitly modeling modal confidence. Attached Figure Description
[0029] Figure 1 This is a schematic diagram of the multimodal feature adaptive fusion method process;
[0030] Figure 2 This is a schematic diagram of a multimodal feature adaptive fusion method.
[0031] Figure 3 This is a schematic diagram of the confusion matrix under different methods. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0033] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0034] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other.
[0035] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0036] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0037] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0038] To facilitate understanding of the principles and implementation process of this invention by those skilled in the art, the invention will be described below with reference to a specific application scenario. Taking the classification of urban functional areas based on confidence-based fusion of high-resolution images and point of interest (POI) name information as an example, the complete implementation process of this invention will be explained.
[0039] In this embodiment, as Figures 1-2 As shown, a multimodal feature adaptive fusion method based on confidence estimation includes the following steps:
[0040] S1: Acquire data input from multiple remote sensing modalities, including high-resolution optical remote sensing image features and point of interest (POI) name sequence text features; specifically, select a typical urban center area (example: a 10×10 km area in a city center), and acquire high-resolution optical remote sensing images (resolution 0.3~1.0 m, multispectral bands optional) and POIs. Data (source: open map platform or commercial map API, fields must include at least name, latitude and longitude, category label optional), and attention must be paid to time consistency during collection (same time period or recent time period), and data version and collection time should be recorded; then, using street boundaries (or regular grids) as geographic units, for each unit, image blocks are cropped from the image according to street vectors, and they are resampled / orthorectified to 224×224 pixels to ensure geometric consistency. All POI names in the unit are sorted by lexicographical order or frequency of occurrence and then concatenated into a sequence string, separated by semicolons ";". If the street image exceeds the range of a single image, tiles are stitched together first and then cropped. If the street area is too small / too large, the cropping scale is adjusted using minimum / maximum boundary constraints; finally, POI names are deduplicated, outlier removed, and semantically normalized to reduce the impact of noise on semantic encoding.
[0041] S2: Construct feature encoders for each modality and extract their respective modal feature representations; further, the feature encoders for each modality are as follows:
[0042] Optical images are used to extract visual semantic features using convolutional neural networks or visual Transformers;
[0043] A Transformer-based pre-trained language model is used to obtain semantic vectors for POI text. Specifically, a ResNet50 network is used as the image encoder. A 224×224 image patch is input into the model to extract 512-dimensional visual features, which are then mapped to a 256-dimensional representation through a fully connected layer. The POI text encoder uses a BERT Chinese pre-trained language model. The concatenated POI name sequence is input into the encoder, and the [CLS] token is taken as the sequence-level representation to obtain 768-dimensional semantic features. These features are then mapped to a 256-dimensional representation through a fully connected layer. This ensures that the high-resolution image and POI text are represented as high-level semantic vectors of the same dimension, facilitating subsequent fusion. BatchNorm and L2 normalization are applied to the outputs of both modalities to ensure that the feature scales are comparable before fusion.
[0044] S3: Construct auxiliary classifiers for each modality, predict the class based on modality features, and output the corresponding class probability distribution; specifically, construct an independent classifier for each modality ( and The image modality classifier consists of two fully connected layers; the text modality classifier uses the same structure, outputting class probabilities through softmax, and training each modality classifier through supervision (cross-entropy).
[0045] S4: Calculate the pseudo-confidence of each modality on the current sample based on the probability distribution output by the classifier; further, the pseudo-confidence of a modality is defined as the probability of the true class in the prediction, therefore the formula for calculating the pseudo-confidence is:
[0046] ;
[0047] in, For the first The pseudo-confidence of each modality For the first Feature representation of each modality The true label for the sample is defined as follows: Specifically, for each block sample, the probability distribution of the functional area category is output for both the image and text modalities. The predicted probability of the category consistent with the true label is defined as the pseudo-confidence of that modality. For example, if the probability of an image predicting "residential area" for a block is 0.76 and the probability of a POI is 0.88, then the pseudo-confidences of the image and the POI are 0.76 and 0.88, respectively.
[0048] S5: Train a modal confidence prediction network using pseudo-confidence as the regression target; specifically, the modal confidence prediction network ( and The algorithm consists of two to three fully connected layers, with BatchNorm and ReLU activation functions applied between layers. A sigmoid function or a scaled linear output is used in the output layer to obtain a scalar confidence score. During training, the pseudo-confidence score is used as the regression target, and the mean squared error loss function is used for optimization. The loss function is as follows:
[0049] ;
[0050] in, To predict confidence levels, the network is jointly optimized with the modality classification loss and the fused task loss during actual training. This process enables the network to learn the mapping relationship between modality features and their reliability, thus allowing for automatic estimation of confidence levels during the inference phase without the need for real labels.
[0051] S6: Concatenate the features of each modality and perform weighted fusion based on their respective confidence weights to obtain the fused feature representation; furthermore, feature fusion adopts channel-dimensional concatenation, and the calculation formula for the fused feature representation is:
[0052] .
[0053] The fused features are input into a multilayer perceptron (MLP) for nonlinear mapping to improve the discriminative power of the fused representation.
[0054] S7: Input the fused feature representation into the downstream recognition or classification module to complete the semantic recognition task of land features or regions. Specifically, input the fused features into the final classifier (two-layer fully connected network) to complete the urban functional area classification task. During training, the classification loss and confidence loss are jointly optimized to ensure the synergistic improvement of classification accuracy and confidence prediction.
[0055] The validation results on a sample of 2499 blocks in the central urban area of a certain city are shown in Table 1.
[0056] Table 1
[0057] method Pub Com Edu Gre Ind Res I-Res Und OA (%) Kappa Channel stacking 64.00 85.71 66.67 88.46 62.50 81.98 100.00 85.71 80.40 0.7421 Element addition 100.00 88.89 52.94 92.31 83.33 79.13 84.62 43.75 79.60 0.7297 Element-wise multiplication 62.50 92.31 64.29 88.46 70.00 83.04 59.09 63.64 79.20 0.7268 Decision average 85.71 86.96 53.85 83.33 60.00 83.19 91.67 100.00 82.00 0.7618 This method 85.00 88.00 53.85 89.29 87.50 85.19 100.00 81.82 85.20 0.8053
[0058] The baseline method (channel stacking) achieves an overall accuracy (OA) of 80.4%; the static weighted method has a maximum OA of 82.00%, while the method of this invention achieves an OA of 85.2%, with the Kappa coefficient improved to 0.8053. Furthermore, the confusion matrices based solely on POI data, solely on image data, and based on this method are visualized, as shown below. Figure 3 As shown, the results demonstrate that the present invention achieves sample-level dynamic fusion, effectively overcoming the uncertainty caused by differences in modality quality. Specifically, the present invention guides the confidence prediction network through pseudo-confidence and adaptively adjusts the modality weights according to the samples, thereby suppressing the negative impact of low-quality modalities on fusion. This allows for compatibility with different numbers and types of modalities, improving scalability and adapting to complex remote sensing scenarios. In other words, the present invention is not only applicable to urban functional area classification but can also be extended to multimodal fusion tasks such as geological disaster identification, crop classification, and land cover monitoring, providing a general technical framework for intelligent interpretation in complex remote sensing environments.
[0059] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A multimodal feature adaptive fusion method based on confidence estimation, characterized in that: Includes the following steps: S1: Acquire data input from multiple remote sensing modalities; S2: Construct feature encoders for each modality and extract their respective modal feature representations; S3: Construct an auxiliary classifier for each modality, predict the class based on the modality features, and output the corresponding class probability distribution; S4: Calculate the pseudo confidence level of each modality on the current sample based on the probability distribution output by the classifier; S5: Train a modality confidence prediction network using pseudo-confidence as the regression target; S6: Concatenate the features of each modality and perform weighted fusion based on their respective confidence weights to obtain the fused feature representation; S7: Input the fused feature representation into the downstream recognition or classification module to complete the semantic recognition task of land features or regions; In step S2, the modal feature encoders are as follows: Optical images are used to extract visual semantic features using convolutional neural networks or visual Transformers; A pre-trained language model based on Transformer is used to obtain text semantic vectors for POI text; In step S4, the formula for calculating the pseudo-confidence level is: ; in, For the first The pseudo-confidence of each modality For the first Feature representation of each modality The true labels for the samples.
2. The multimodal feature adaptive fusion method based on confidence estimation according to claim 1, characterized in that: In step S5, the modality confidence prediction network consists of two to three fully connected layers, with BatchNorm and ReLU activation functions set between the layers, and sigmoid or scaled linear output used in the output layer to obtain scalar confidence.
3. The multimodal feature adaptive fusion method based on confidence estimation according to claim 2, characterized in that: In step S5, the optimization objective of the modal confidence prediction network is the mean square error loss function, and its loss form is: ; in, To predict confidence levels.
4. The multimodal feature adaptive fusion method based on confidence estimation according to claim 3, characterized in that: In step S6, feature fusion uses channel-dimensional concatenation, and the formula for calculating the fused feature representation is as follows: 。