Trustworthy Multi-Perspective Classification Method Based on Evidence Deep Learning

By introducing a degradation layer in the multi-view classification task, global information is propagated to each perspective, and the problem of difficulty in mining complementary information between perspectives in the prior art is solved, and higher prediction accuracy and more reasonable uncertainty output are achieved.

CN114492620BActive Publication Date: 2025-06-27XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210080384.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-06-27
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively explore complementary information between perspectives in multi-view classification tasks, resulting in the inability to output reasonable uncertainty and increase the possibility of misdiagnosis.

Method used

A trusted multi-perspective classification method based on deep learning of evidence is proposed. The global information is propagated to each perspective through the degradation layer, so that each perspective can learn evidence based on global information, and the model parameters are optimized using gradient descent algorithm.

Benefits of technology

It improves the accuracy of predictions, and successfully explores complementary information between deep perspectives, outputs uncertainty that is more in line with human cognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114492620B_ABST
    Figure CN114492620B_ABST
Patent Text Reader

Abstract

The present invention discloses a trustworthy multi-view classification method based on evidence deep learning, including the following method steps: S1. Sample definition, setting that there are N samples in the data set, and each sample has V views; S2. Single-view evidence, estimating the classification uncertainty of single-view data; S3. Multi-view evidence fusion, using a degradation layer to propagate global information to each view, so that each view can learn evidence based on global information; S4. Optimization objective, using the gradient descent algorithm to optimize all parameters in the model. The method of the present invention not only improves the prediction accuracy, but also uses the degradation layer designed in the present invention to mine complementary information between deep and easily overlooked views, so as to output a more human-cognitive uncertainty during prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and in particular to a reliable multi-view classification method based on evidence deep learning. Background Art

[0002] Multi-view classification means that each sample in the dataset contains features from multiple different views. For example, when a doctor makes a cancer diagnosis, he often needs to comprehensively consider features from multiple different views such as magnetic resonance imaging and clinical test results. Using deep learning to implement multi-view classification tasks can make up for the defect that traditional machine learning is difficult to mine deep information from multi-view data. Most past research has focused on how to improve prediction accuracy while ignoring the reliability of decisions. However, in many high-risk applications, not only does the model need to obtain the decision result, but also the confidence level of the decision. For example, in medical diagnosis, the confidence level of the decision is crucial. A decision with an unknown confidence level may be unreliable, thus misleading the doctor to make a wrong diagnosis and delaying the best treatment opportunity for the patient.

[0003] In recent years, some methods have emerged that can output the uncertainty of the prediction while making the prediction, but still cannot effectively mine the complementary information between multiple views, and thus cannot output reasonable uncertainty. Here is an example to illustrate. In a medical diagnosis event, based on magnetic resonance imaging, which is view 1, it can predict three categories: healthy, left lesion, and right lesion, but the prediction accuracy is often not high due to blurred images; based on clinical test results, which is view 2, it can only predict two categories: healthy and lesion, but cannot distinguish whether the lesion is on the left or the right, but the accuracy is relatively high. According to common sense, if view 2 predicts a lesion, then it will give view 1 a great deal of complementary information. However, in fact, in the past research, when dealing with view 2, since it cannot distinguish whether the lesion is on the left or the right, that is, the complete category information, it can only output a high uncertainty, and thus cannot mine the deep complementary information between these two views. Ignoring the information of any view in medical diagnosis increases the possibility of misdiagnosis.

[0004] In the multi-view classification problem, how to mine the complementary information between views and obtain an uncertainty that conforms to human cognition while ensuring that the accuracy is not reduced still poses a great challenge. Therefore, how to provide a reliable multi-view classification method based on evidence deep learning is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] An object of the present invention is to propose a reliable multi-view classification method based on evidence deep learning. The object of the present invention is to propose a deep learning method for multi-view classification problems that can effectively mine the complementary information between views and obtain reasonable uncertainty while ensuring that the accuracy is not reduced, so as to solve the problems existing in the above background art.

[0006] A trustworthy multi - perspective classification method based on evidence deep learning according to an embodiment of the present invention includes the following method steps:

[0007] S1. Sample definition: Assume there are N samples in the data set, and each sample has V perspectives;

[0008] S2. Single - perspective evidence: Estimate the classification uncertainty of single - perspective data;

[0009] S3. Multi - perspective evidence fusion: Use the degradation layer to propagate global information to each perspective, so that each perspective can learn evidence based on global information;

[0010] S4. Optimization objective: Use the gradient descent algorithm to optimize all parameters in the model.

[0011] Preferably, in step S1:

[0012] The representation vector of the n - th sample at the v - th perspective:

[0013]

[0014] where D v is its dimension;

[0015] Considering that the data set contains C categories, each category is represented as a one - hot vector as the true label of the sample. The true label of the n - th sample:

[0016] y n ∈{0,1} C ;

[0017] The goal is to, given a sample, predict its label and give the uncertainty u n ∈[0,1].

[0018] Preferably, in step S2:

[0019] For the classification problem of C categories, the Dirichlet distribution parameters of the n - th sample at the v - th perspective are:

[0020]

[0021] Then its probability density function is expressed as:

[0022]

[0023] where B(x) represents the probability density function of the multinomial distribution;

[0024] T C is a C - dimensional simplex:

[0025]

[0026] Use a learner based on a deep neural network to determine the parameters of the Dirichlet distribution

[0027]

[0028] where f v (x) represents a learner based on a deep neural network, represents the evidence vector of the nth sample from the vth perspective;

[0029] Use the Dirichlet distribution to describe the evidence distribution, and let Then the probability of predicting the nth sample as class c and the prediction uncertainty are respectively:

[0030]

[0031]

[0032] where, represents the intensity of the Dirichlet distribution.

[0033] Preferably, the evidence value The larger it is, the greater the possibility that the corresponding class is the true class; when the sum of the evidence values of the C classes is larger, the uncertainty will be smaller.

[0034] Preferably, in step S3:

[0035] There is a degradation relationship between the evidence of a specific perspective and the fused evidence:

[0036]

[0037] where the operation represents the element-wise multiplication of two vectors, and the meaning of ≈ is that it does not require the data on both sides to be strictly equal but expects them to be as close as possible, d v ∈R C represents the intensity degradation vector, which is used to scale the intensity between the evidence of a specific perspective and the fused evidence, U v ∈R C×C represents the class degradation relationship matrix;

[0038] The objective function of the degradation layer can be formulated as:

[0039]

[0040] where, is the normalized weight coefficient of the v-th perspective. The larger the weight, the higher the credibility of the corresponding perspective. When inputting a sample into the model, the perspective with high credibility will play a positive role in optimizing the model parameters; while the perspective with low credibility may affect the update direction of the parameters, thereby resulting in poor generalization ability of the model. Therefore, in the present invention, this weight coefficient is added to weaken the influence of the perspective with low credibility on the update of the model parameters.

[0041] Preferably, the meaning of the degradation loss is to express the difference between the fused evidence and the evidence of the feature perspective.

[0042] Preferably, the degradation layer includes two degradation stages: intensity degradation and category degradation. The intensity degradation means that the information intensity of the evidence of a specific perspective is inconsistent with the fused evidence, and the category degradation means that the clustering clarity of a specific perspective is inconsistent with the clarity of the fused evidence.

[0043] Preferably, in step S4:

[0044] According to the relationship between the Dirichlet distribution and the evidence theory in S2:

[0045]

[0046] It is expected that the true label y of the sample n has more evidence for the corresponding category than other categories. It is only necessary to minimize the log-likelihood function:

[0047]

[0048] However, the evidence value is not the larger the better. In the present invention, it is not desired that the evidence values of multiple categories compete with each other. Therefore, the KL divergence is introduced in the present invention as a regularization term:

[0049]

[0050] Among them, represents the Dirichlet distribution parameter after removing the evidence of the true category; on the one hand, the KL divergence can limit the growth of the total evidence, and on the other hand, it can promote the evidence distribution to tend to the direction with smaller information entropy. In other words, the evidence values between different categories are more different.

[0051] The overall optimization objective of the model is given:

[0052]

[0053] During optimization, set t represents the number of iterations. It is expected that the model can learn sufficient evidence at the beginning of training, and then gradually increase the penalty of the regularization term. Set δ = 0.1 to represent the coefficient of the degradation layer.

[0054] Preferably, the evidence value of the function excitation true category tends to the total evidence value.

[0055] The beneficial effects of the present invention are as follows:

[0056] The method of the present invention not only improves the prediction accuracy, but also uses the designed degradation layer in the present invention to mine the complementary information between deep and easily overlooked perspectives, so as to output a more human - cognitive uncertainty during prediction. Description of the Drawings

[0057] The drawings are used to provide further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:

[0058] Figure 1 is the overall flowchart of the deep network of the reliable multi - perspective classification method based on evidence deep learning proposed by the present invention;

[0059] Figure 2 is the visualization example diagram of the three - classification samples of Perspective 1 of the reliable multi - perspective classification method based on evidence deep learning proposed by the present invention;

[0060] Figure 3 is the visualization example diagram of the three - classification samples of Perspective 2 of the reliable multi - perspective classification method based on evidence deep learning proposed by the present invention;

[0061] Figure 4 is the bar chart of the evidence value data of the TMC model;

[0062] Figure 5 is the bar chart of the evidence value data of the reliable multi - perspective classification method based on evidence deep learning proposed by the present invention. Detailed Embodiments

[0063] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic way, so they only show the components related to the present invention. Refer to Figures 1-5 :

[0064] Embodiment 1:

[0065] This embodiment proposes a deep - learning method for solving the multi - perspective classification problem with uncertainty prediction by considering the consistency and complementarity of multi - perspective data, and verifies that this method significantly improves the classification accuracy and confidence. In the implementation of this method, a degradation layer is used to fuse multi - perspective information and model the semantic association between the fused evidence and the specific - perspective evidence. This fusion paradigm can be applied to other multi - perspective deep - learning models to generate reliable decisions.

[0066] A trustworthy multi-view classification method based on deep learning of evidence according to an embodiment of the present invention includes the following method steps:

[0067] S1. Sample definition: Assume that there are N samples in the data set, and each sample has V views;

[0068] S2. Single-view evidence: Estimate the classification uncertainty of single-view data;

[0069] S3. Multi-view evidence fusion: Use the degradation layer to propagate the global information to each view, so that each view can learn the evidence based on the global information;

[0070] S4. Optimization objective: Use the gradient descent algorithm to optimize all parameters in the model.

[0071] In step S1:

[0072] The representation vector of the nth sample from the vth view:

[0073]

[0074] D v is its dimension;

[0075] Considering that the data set contains C categories, each category is represented as a one-hot vector as the true label of the sample. The true label of the nth sample:

[0076] y n ∈{0,1} C ;

[0077] The goal is to predict the label of a given sample and give the uncertainty u of the prediction n ∈[0,1].

[0078] In step S2:

[0079] For the classification problem of C categories, the Dirichlet distribution parameters of the nth sample from the vth view are:

[0080]

[0081] Then its probability density function is expressed as:

[0082]

[0083] Among them, B(x) represents the probability density function of the multinomial distribution;

[0084] T C is a C-dimensional simplex:

[0085]

[0086] Use a learner based on a deep neural network to determine the parameters of the Dirichlet distribution

[0087]

[0088] where f v (x) represents a learner based on a deep neural network represents the evidence vector of the v-th perspective of the n-th sample

[0089] Use the Dirichlet distribution to describe the evidence distribution, and let Then the probability of predicting the n-th sample as class c and the prediction uncertainty are respectively

[0090]

[0091]

[0092] where represents the intensity of the Dirichlet distribution

[0093] Preferably, the evidence value The larger it is, the greater the possibility that the corresponding class is the true class; when the sum of the evidence values of the C classes is larger, the uncertainty will be smaller

[0094] The present invention aims to propose a degradation layer to capture complementary information between multiple perspectives. In multi-perspective learning, a specific perspective often only contains local information. When the number of clusterable clusters of local information is less than the number of final classification classes, the local information of a specific perspective is often less than that of the global information, that is, the information contained is more one-sided. The model proposed by the present invention includes two degradation stages: intensity degradation and class degradation. Intensity degradation indicates that the information intensity of the evidence of a specific perspective is inconsistent with that of the fused evidence, and class degradation indicates that the clustering clarity of a specific perspective is inconsistent with that of the fused evidence

[0095] Refer to Figures 2-3 , in the embodiment, for a three-classification task including two perspectives, if there are not enough features in perspective 1 to distinguish between class 2 and class 3, but it is possible to clearly distinguish between class 1 and classes 2 and 3, then perspective 1 should have useful information, that is, low uncertainty, to assist perspective 2 in more clearly classifying

[0096] In step S3

[0097] There is a degradation relationship between the evidence of a specific perspective and the fused evidence

[0098]

[0099] Among them, the operation represents the element-wise multiplication of two vectors. The meaning of ≈ is that it does not require the data on both sides to be strictly equal but expects them to be as close as possible. d v ∈R C represents the intensity degradation vector, which is used to scale the intensity between the evidence of a specific perspective and the fused evidence. U v ∈R C×C represents the category degradation relationship matrix;

[0100] Taking the sample distribution of perspective 1 as an example, it is expected that the model can learn the matrix U 1 :

[0101]

[0102] Because in perspective 1, category 1 is clearly separable, so while categories 2 and 3 are completely indistinguishable. Therefore, the fused evidence of category 2 or 3 must be evenly spread to the evidence of categories 2 and 3 in perspective 1. So

[0103] The objective function of the degradation layer can be formulated as:

[0104]

[0105] Among them, is the normalized weight coefficient of the v-th perspective. The larger the weight, the higher the credibility of the corresponding perspective.

[0106] When inputting a sample into the model, the perspective with high credibility will play a positive role in optimizing the model parameters; while the perspective with low credibility may affect the update direction of the parameters, which may lead to poor generalization ability of the model. Therefore, in the present invention, this weight coefficient is added to weaken the influence of the perspective with low credibility on the update of the model parameters.

[0107] The meaning of the said degradation loss is to express the difference between the fused evidence and the evidence of the feature perspective.

[0108] In step S4:

[0109] According to the relationship between the Dirichlet distribution and the evidence theory in S2:

[0110]

[0111] It is expected that the evidence of the category corresponding to the true label y of the sample n is more than that of other categories. It is only necessary to minimize the log-likelihood function:

[0112]

[0113] However, the evidence value is not the larger the better. In the present invention, it is not desirable to have a situation where the evidence values of multiple categories compete with each other. Therefore, the KL divergence is introduced as a regularization term:

[0114]

[0115] where represents the Dirichlet distribution parameter after removing the evidence of the true category; on the one hand, the KL divergence can limit the growth of the total evidence, and on the other hand, it can promote the evidence distribution to tend to a direction with a smaller information entropy. In other words, the evidence values between different categories are more different;

[0116] The overall optimization objective of the model is given:

[0117]

[0118] During optimization, set t represents the number of iterations. It is expected that the model can learn sufficient evidence at the beginning of training, and then gradually increase the penalty of the regularization term. Set δ = 0.1 to represent the coefficient of the degradation layer. The function encourages the evidence value of the true category to tend to the total evidence value.

[0119] After implementing the above S1 - S4, compared with the multi - view classification model proposed in the prior art, the method proposed in the present invention can not only ensure that the accuracy is not lower than that of the previous methods, but also calculate a more realistic and logical uncertainty by mining the complementary information between multiple views.

[0120] Referring to Figures 4-5 , to verify this, in this embodiment, test samples are designed, including 2 views and 3 categories, and their distribution is as Figure 2 shown. The best - proposed method TMC (the paper "Trusted Multi - view Classification") and the method proposed in the present invention are used for experiments respectively, and the evidence values learned by the model are recorded for comparison.

[0121] Through Figures 4-5 it can be seen that on view 1, when categories 2 and 3 are not easily distinguishable but are both clearly separable from category 1, TMC will have a situation of insufficient evidence mining, that is, it cannot fully mine the information carried by view 1, and thus outputs a relatively high uncertainty of 0.49, which is unreasonable. The method proposed in the present invention has mined sufficient evidence for both categories 2 and 3 on view 1. This is due to the information mining of view 1 by the degradation layer, and then fusing the view information to predict an accurate and reliable decision uncertainty of 0.22.

[0122] The present invention has been experimented on two publicly available datasets. The HandWritten dataset consists of handwritten digit pictures, and the extracted features include 6 perspectives, divided into 10 categories. The Scene15 dataset contains 4,485 scene pictures, divided into 15 categories. The present invention respectively uses the GIST, HOG, and LBP methods to extract features from three perspectives for multi-perspective classification experiments. Taking the accuracy rate as an index, for each method, 10 experiments are conducted on each dataset and the average value is taken. The experimental results are shown in Table 1:

[0123] Table 1

[0124] Method HandWritten Scene15 TMC 98.04% 63.85% The present invention 98.75% 66.93%

[0125] The method of the present invention is verified through examples, which not only improves the prediction accuracy, but also uses the designed degradation layer in the present invention to mine the complementary information between deep and easily overlooked perspectives, so as to output a more human-cognitive uncertainty during prediction.

[0126] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.

Claims

1. A trustworthy multi-view classification method based on evidence deep learning, characterized in that, It includes the following method steps: S1. Sample definition: Set that there are N samples in the data set, and each sample has V perspectives. The data set is an image data set; S2. Single-perspective evidence: Estimate the classification uncertainty of single-perspective data; S3. Multi-perspective evidence fusion: Use the degradation layer to propagate the global information to each perspective, so that each perspective can learn the evidence based on the global information; In step S3: There is a degradation relationship between the specific perspective evidence and the fusion evidence: Among them, the operation represents the element-wise multiplication of two vectors. The meaning of ≈ is that it does not require the data on both sides to be strictly equal but expects them to be as close as possible. d v ∈R C represents the intensity degradation vector, which is used to scale the intensity between the evidence of a specific perspective and the fused evidence. U v ∈R C×C represents the class degradation relationship matrix; Denote the evidence vector of the $n$-th sample from the $v$-th perspective; The objective function of the degradation layer can be formulated as: Among them, is the normalized weight coefficient of the v-th perspective. The larger the weight coefficient, the higher the credibility of the corresponding perspective; S4. Optimization objective: Use the gradient descent algorithm to optimize all parameters in the model.

2. The trustworthy multi-view classification method based on evidence deep learning according to claim 1, wherein In step S1: The representation vector of the nth sample at the vth perspective: where D v is its dimension; Considering that the data set contains C categories, each category is represented as a one-hot vector as the true label of the sample. The true label of the nth sample: y n ∈{0,1} C ; The goal is, given a sample, to predict its label and give the uncertainty u of the prediction n ∈ [0, 1].

3. The trustworthy multi-view classification method based on evidence deep learning according to claim 1, characterized in that In step S2: For the classification problem of C categories, the Dirichlet distribution parameter of the nth sample at the vth perspective is: Then its probability density function is expressed as: Among them, B(x) represents the probability density function of the multinomial distribution; T C is a C-dimensional simplex: Use a learner based on a deep neural network to determine the parameters of the Dirichlet distribution where, f v (x) represents a learner based on a deep neural network, represents the evidence vector of the nth sample from the vth perspective; Describe the evidence distribution using the Dirichlet distribution, and let Then the probability of predicting the nth sample as class c and the prediction uncertainty are respectively: Among them, represents the intensity of the Dirichlet distribution.

4. The trustworthy multi-view classification method based on evidence deep learning according to claim 3, wherein The evidence value The larger it is, the greater the possibility that the corresponding category is the true category; when the sum of the evidence values of the C categories is larger, the uncertainty will be smaller.

5. The trustworthy multi-perspective classification method based on evidence deep learning according to claim 1, wherein The degradation layer includes two degradation stages: intensity degradation and category degradation. The intensity degradation means that the information intensity of the specific perspective evidence and the fusion evidence is inconsistent, and the category degradation means that the clustering clarity of the specific perspective and the clarity of the fusion evidence are inconsistent.

6. The trustworthy multi-perspective classification method based on evidence deep learning according to claim 1, wherein In step S4: According to the relationship between the Dirichlet distribution in S2 and the evidence theory: True label y of the desired sample n If there is more evidence for the corresponding class than for other classes, only the log-likelihood function needs to be minimized: Introduce the KL divergence as a regularization term: Among them, represents the Dirichlet distribution parameter after removing the true class evidence. On the one hand, the KL divergence can limit the growth of the total evidence, and on the other hand, it can promote the evidence distribution to tend to the direction with a smaller information entropy; Give the overall optimization objective of the model: During optimization, set Let t represent the number of iterations. It is expected that the model can learn sufficient evidence at the beginning of training and then gradually increase the penalty of the regularization term. Setting δ = 0.1 represents the coefficient of the degradation layer.

7. The trustworthy multi-perspective classification method based on evidence deep learning according to claim 6, wherein The evidence value of the function excitation for the true category tends to the total evidence value.