Evidence fusion multi-modal classification model uncertainty measurement method
By using the Dirichlet distribution trust value method with two-layer fusion in multimodal data processing, the problem of poor effectiveness of multimodal data uncertainty measurement in the prior art is solved, and more efficient and accurate uncertainty measurement is achieved.
Patent Information
- Application Number
- CN202510170668.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-10
AI Technical Summary
The existing uncertainty measurement methods are difficult to fully integrate the characteristics of different modal data when processing multimodal data, resulting in poor uncertainty measurement.
By dividing the modal subset of multimodal data, the fully connected layer activation value is extracted as the original feature, the Softmax and ReLU evidence are generated using the Basic Belief Assignment (BBA) method and parameterized into the trust value of the Dirichlet distribution. Then, based on dynamic trust theory, the bilayer fusion within and between modals is performed, and the Dirichlet distribution entropy is calculated to quantify the uncertainty of the model.
The evidence fusion efficiency of multimodal data is improved, the information of multimodal data is fully utilized, the robustness and reliability of the model in complex environments is improved, and the accuracy of uncertainty measurement is significantly improved.
Smart Images

Figure CN120125937A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for measuring the uncertainty of a multi-modal classification model for evidence fusion, belonging to the field of artificial intelligence security. Background Art
[0002] Model uncertainty includes two sources: data quality and model structure. The entire data space is difficult to represent and generalize due to insufficient representativeness of the data set, and the information in complex data cannot be fully learned and expressed due to the limitations of the model structure. Ultimately, it is manifested as the randomness and low confidence probability of the model's prediction when dealing with unknown samples due to insufficient cognitive ability. Multi-modal data is a complex data that integrates information of multiple data types such as images, texts, and audios. Each modality expresses the information of a unique aspect of the same object or phenomenon, comprehensively displays the characteristics of different aspects, provides an overall view and complementary perspectives of the target, and helps to more comprehensively and accurately identify and predict the established target object. With the increasing application of multi-modal data, the inherent heterogeneity and complexity of such data pose higher requirements for the model's uncertainty measurement ability. In high-risk fields such as medical diagnosis and autonomous driving, any error may lead to serious consequences. Effective uncertainty measurement can help identify potential risks in model predictions and optimize the model decision-making process to provide necessary safety guarantees.
[0003] Current uncertainty measurement methods mainly include probability model methods, confidence methods, etc. These methods perform well in single-modal data processing and can effectively quantify the uncertainty of model predictions. However, multi-modal data usually has problems such as dimensionality differences and inconsistent feature spaces between modalities, making it difficult to align and fuse the information of different modal data, resulting in poor uncertainty measurement effects when these methods are used to process multi-modal data.
[0004] 1. Probability Model Method
[0005] Methods such as Bayesian neural networks (BNNs) and Gaussian processes (GPs) estimate uncertainty by defining a probability distribution for the model output. BNNs introduce a prior distribution on the network weights and use the posterior distribution to infer the uncertainty of the output. These methods have shown excellent performance on single-modal data, but on multi-modal data, they are often limited by the failure to fully fuse the features of different data sources, resulting in the inability to accurately reflect the interaction between different data, thus affecting the accuracy of uncertainty measurement.
[0006] 2. Evidence Deep Learning Method
[0007] Existing evidence deep learning frameworks usually adopt a single Basic Belief Assignment (BBA) method. Commonly used single BBA methods are usually the Rectified Linear Unit (ReLU) function and the Softmax function, which are used to connect the activation values output by the fully connected layer of the model to evidence fragments, thereby quantifying belief quality and uncertainty. The ReLU function preserves the original strength of the evidence, but often ignores negative evidence, resulting in the loss of evidence information. The Softmax output is overly confident. In multi-modal tasks, the quality of each modality varies with samples, and there are problems with both choosing an output that preserves the original quality and an output with an overconfident tendency.
[0008] In summary, the existing uncertainty measurement methods still have the following limitations: 1) The activation function is single, and the one-sided evidence transformed is difficult to comprehensively evaluate uncertainty; 2) There are dimensional differences in the data of each modality in multi-modal tasks, and it is difficult to align and fuse them during comprehensive quantification. Summary of the Invention
[0009] The purpose of the present invention is to address the problems that the single evidence of existing methods affects the measurement accuracy and is not applicable to multi-modal classification models. By establishing and fusing multiple types of evidence, it is not only applicable to the measurement of the uncertainty of multi-modal classification models but also can improve the measurement accuracy.
[0010] The design principle of the present invention is as follows: First, divide the multi-modal data into data subsets according to modalities, and input each modality's data subset into the corresponding target model respectively, and extract the activation values of the fully connected layer of the model as the original features of this modality; Second, calculate the Softmax evidence and ReLU evidence of the original features respectively based on the BBA method, and use the Dirichlet distribution to parameterize the Softmax and ReLU evidence into two types of trust values that conform to the Dirichlet distribution respectively; Then, based on the dynamic trust theory, perform the first fusion within the same modality, multiply the two types of trust values element by element to obtain the trust value of this modality, and then perform the second fusion among modalities, and calculate the weighted output of each modality to obtain the fusion trust value of the multi-modal classification model; Finally, calculate the Dirichlet distribution entropy of the fusion trust value to measure the uncertainty of the model.
[0011] The technical solution of the present invention is realized through the following steps:
[0012] Step 1, divide each modality's data subset and extract features respectively.
[0013] Step 1.1, first divide the multi-modal data into data subsets according to modalities, and preprocess each modality's data subset after division to ensure data quality and consistency.
[0014] Step 1.2: Input each subset of modal data into the target model respectively, and extract the activation values of the fully connected layer of the model as the original features of this modality.
[0015] Step 2: Use two BBA methods to transform the original features to obtain two types of evidence, and then convert the two types of evidence into belief values that conform to the Dirichlet distribution.
[0016] Step 2.1: Use the BBA method to convert the output original features into preliminary evidence.
[0017] Step 2.2: Based on the Dirichlet theoretical framework, parameterize the Softmax and ReLU evidence sources into two types of belief values that conform to the Dirichlet distribution respectively.
[0018] Step 3: Double-layer fusion of belief values within and between modalities.
[0019] Step 3.1: Within each modality, based on the dynamic trust theory, multiply the two types of belief values element by element and combine them to obtain the preliminary belief value in the single modality. This process is the first layer of fusion in the double-layer fusion.
[0020] Step 3.2: Again based on the dynamic trust theory, calculate the belief values of each modality with weights between modalities to obtain the fused belief value. This process is the second layer of fusion in the double-layer fusion.
[0021] Step 4: Calculate the Dirichlet distribution parameters and fusion probabilities for each class, and quantify the classification uncertainty of the model based on the entropy of the Dirichlet distribution.
[0022] Step 4.1: Use the Dirichlet framework to calculate the Dirichlet distribution parameters and probabilities for each class.
[0023] Step 4.2: Calculate the entropy of the Dirichlet distribution to measure the uncertainty of the model output.
[0024] Beneficial effects
[0025] Compared with the traditional evidence deep learning method, the present invention has significant advantages in the uncertainty measurement of the multi-modal classification model. The present invention enriches the evidence information source by using the evidence generated by a variety of activation functions for Dirichlet distribution transformation, and can evaluate the uncertainty information more comprehensively.
[0026] Compared with the traditional uncertainty measurement based on the probability model method, when dealing with multi-modal data, the present invention uses the dynamic trust theory to perform double-layer evidence fusion within and between modalities, improves the evidence fusion efficiency of multi-modal data and fully utilizes the data information, comprehensively considers the information expression of all modalities, and effectively improves the robustness and reliability of the model in complex environments. Description of the drawings
[0027] Figure 1 This is the schematic diagram of the uncertainty measurement method for the multi-modal classification model of evidence fusion in the present invention. Detailed implementation manners
[0028] To better illustrate the purpose and advantages of the present invention, the implementation manners of the method of the present invention will be further described in detail below with reference to examples.
[0029] The experimental data comes from publicly available multi-modal datasets: the COCO dataset and the Flickr30k dataset. The COCO dataset is a large-scale dataset for image recognition, segmentation, and image captioning, and is widely used in multiple core tasks in the field of computer vision research, including object detection, segmentation, human keypoint detection, and image captioning. The COCO dataset contains more than 200,000 images, 328,000 person descriptions, and 1.5 million object instances of 91 object categories. These images are mainly objects in daily scenes, and each image is accompanied by at least 5 different descriptive texts. The Flickr30k dataset contains approximately 31,000 images, each with 5 different English descriptions, and these images and descriptions are mainly collected from the Flickr website.
[0030] This experiment was carried out on a computer with the following specific configuration: Inter i7-6700, CPU 3.4GHz, 8G of memory, GPU is RTX6000, 24G of video memory, and the operating system is windows 11, 64-bit.
[0031] The specific process of this experiment is as follows:
[0032] Step 1, divide each modal data subset and extract features respectively.
[0033] Step 1.1, first divide the multi-modal data into data subsets according to the modality, and preprocess each divided modal data subset:
[0034] Directly extract the image data as the image data subset; extract the text from the text descriptions associated with the images and store it in JSON format.
[0035] Image data: Standardize the images, resize them to a unified resolution of 224x224 pixels, and normalize them so that the pixel values are in the range of [0,1]. The calculation method is shown in Equation (1).
[0036]
[0037] Where X represents the original image data, and mean(X) and std(X) represent the mean and standard deviation of the image data respectively.
[0038] Text data: Perform preprocessing steps such as word segmentation, stop word removal, and stemming on the text, and perform TF-IDF conversion to convert the text into a numerical representation. The calculation method is shown in Equation (2).
[0039]
[0040] Among them, TF i,j represents the frequency of word i in document j, N is the total number of documents, and DF i represents the number of documents containing word i.
[0041] Step 1.2: Input each modal data subset into the target model respectively, and extract the activation values of the fully connected layer of the model as the original features of this modality. The specific extraction method is as follows:
[0042] Image feature extraction: Use the pre-trained convolutional neural network ResNet-50 to extract image features. The calculation method is shown in Equation (3).
[0043] F image = CNN(X norm ) (3)
[0044] Among them, F image represents the image features, and X norm represents the preprocessed image data.
[0045] Text feature extraction: Use the pre-trained language model BERT to extract text features. The calculation method is shown in Equation (4).
[0046] F text = BERT(tokens) (4)
[0047] Among them, F text represents the text features, and tokens are the preprocessed text data.
[0048] Step 2: Use two BBA methods to transform the original features to obtain two types of evidence, and then transform the two types of evidence into belief values that conform to the Dirichlet distribution.
[0049] Step 2.1: Use the Basic Belief Assignment (BBA) method to transform the output original features into preliminary evidence.
[0050] BBA1 reassigns the output of the deep neural network after activation by the Softmax function to obtain evidence. The calculation method is shown in Equation (5).
[0051]
[0052] BBA2 uses ReLU as the link function to obtain evidence, and the calculation method is shown in Equation (6).
[0053]
[0054] Step 2.2, based on the Dirichlet theory framework, parameterize the Softmax and ReLU evidence sources into two types of trust values that conform to the Dirichlet distribution.
[0055] The concentration parameter of the Dirichlet distribution is a confidence distribution. In the case of K mutually exclusive singletons, assign a trust value b to each k = 1,..., K k , representing the trust value for the occurrence of each class. In addition, assign an uncertainty u to the trust value to represent the uncertainty in the prediction process. The sum of these K + 1 non-negative trust values is equal to 1, representing the complete assignment of trust values and uncertainties, as shown in Equation (7).
[0056]
[0057] The calculation methods of the trust value b and the uncertainty u are shown in Equation (8).
[0058]
[0059] where represents the Dirichlet strength corresponding to the Dirichlet distribution.
[0060] Step 3, double-layer fusion of trust values within and between modalities.
[0061] Step 3.1, the two kinds of evidence of the nth sample obtained in the previous step are respectively represented as and in the vth modality. and The corresponding trust values are and and the uncertainties are
[0062]
[0063] where ⊕ represents the Dempster-Shafer fusion rule, ⊙ represents the Hadamard product, and K represents the normalization coefficient.
[0064] Obtain the first-layer fusion trust values for each modality, denoted as and where n is the sample index and v is the modality index.
[0065] Step 3.2, based on the dynamic trust theory again, calculate the trust values and uncertainties of each modality with weights among modalities to obtain the fused trust values and uncertainties, and the calculation method is shown in Equation (9).
[0066]
[0067] The obtained fused trust value vector (b n ) ⊕ and uncertainty (u n ) ⊕ represent the fused trust values from different modalities in the nth sample.
[0068] Step 4, calculate the Dirichlet distribution parameters and fusion probabilities for each class, and calculate the uncertainty based on the Dirichlet distribution entropy.
[0069] Step 4.1, based on the fused trust value vector (b n ) ⊕ and uncertainty (u n ) ⊕ obtained in Step 3.2, calculate the Dirichlet distribution parameters and probabilities for each class, and the calculation method is shown in Equation (10).
[0070]
[0071] The resulting fusion probability p n and uncertainty (u n ) ⊕ can take into account information from multiple modalities and provide a more comprehensive and robust overall trust and uncertainty estimate.
[0072] Step 4.2, use the entropy H(α) of the Dirichlet distribution to measure the uncertainty of the model output, and the calculation method is shown in Equation (11).
[0073]
[0074] where is the beta function of the Dirichlet distribution, ψ(x) is the logarithmic derivative of the gamma function, called the Digamma function, and K is the number of classes.
[0075] If the calculated entropy is low, it indicates that the model is more confident in the prediction result, indicating less uncertainty; if the entropy is high, it indicates that the model is more hesitant about the prediction result, indicating greater uncertainty.
[0076] The specific description above further elaborates on the purpose, technical solution, and beneficial effects of the invention. It should be understood that the above description is only a specific embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. Uncertainty measurement method for multimodal classification model based on evidence fusion, characterized by The method comprises the following steps: Step 1: Divide the multimodal data into data subsets according to the modality, input each modality data subset into the corresponding target model, and extract the activation value of the fully connected layer of the model as the original feature of the modality; Step 2: Based on the BBA method, the Softmax evidence and ReLU evidence of the original features are calculated respectively. Using the Dirichlet distribution, the Softmax and ReLU evidence are parameterized into two types of trust values that conform to the Dirichlet distribution. Step 3: Based on the dynamic trust theory, the first fusion is performed within the same modality, and the two types of trust values are multiplied element by element to obtain the trust value of the modality. Then, the second fusion is performed between the modalities, and the output of each modality is weighted to obtain the fusion trust value of the multimodal classification model. Step 4: Calculate the uncertainty of the Dirichlet distribution entropy measurement model of the fused trust value.
2. The uncertainty measurement method for multimodal classification model based on evidence fusion according to claim 1 is characterized by: Step 2 uses two BBA methods to obtain multiple evidences from the neural network output. After the deep neural network output is redistributed by the Softmax function activation to obtain the first piece of evidence, the second BBA method is used to use ReLU as the link function to obtain the second piece of evidence, so as to make full use of the original features to enrich the evidence source and facilitate the fusion of the following steps. The calculation formula for evidence conversion is:
3. The uncertainty measurement method for multimodal classification model based on evidence fusion according to claim 1 is characterized by: In step 3, the converted trust values are double-layered and integrated within and between modalities: within each modality, the two types of trust values are multiplied and merged element by element based on the dynamic trust theory to form a more comprehensive and consistent decision-making basis. This process is the first layer of double-layer integration, and the calculation formula is: in, represents the Dempster-Shafer fusion rule, ⊙ represents the Hadamard product, K represents the normalization coefficient, and the first-layer fusion trust value of each modality is obtained, which is recorded as Where n is the sample index and v is the modal index. Based on the dynamic trust theory, the trust value of each modality is weighted and calculated between modalities to obtain the fusion trust value. The calculation formula is
Citation Information
Cited By
Multi-modal post-fusion method, device and equipment based on consistent auxiliary channel, medium and product
CN121389028A