Trustworthy conflict multi-view classification method based on double-layer evidence exploration
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING BIG DATA ADVANCED TECH RES INST
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]本申请实施例的目的是提供一种基于双层证据探索的可信冲突多视图分类方法,能够解决在实际的冲突多视图数据场景下,目前的可信多视图分类方法所存在的分类性能及不确定性估计可靠性不佳的问题
[0010]在本申请实施例中,引入了双层证据探索策略,在视图内层面联合挖掘特征证据与邻域结构证据并进行视图内融合,以生成特定于视图的高质量综合证据表示,由此从源头提升了证据可靠性,并降低了单一证据生成机制带来的有偏输出风险;在视图间层面依据多视图数据“一致性—互补性”并存的内在属性,将不同视图的综合证据表示动态解耦为共识证据与互补证据,以便后续据此协同生成分类结果与不确定性估计结果,由此实现了对跨视图共享信息与视图特异信息的有效协作建模;如此,本申请能够显著提升实际的冲突多视图数据场景下的分类性能与不确定性估计可靠性。
Smart Images

Figure CN122530673A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, specifically relating to a credible conflict multi-view classification method based on two-layer evidence exploration. Background Technology
[0002] Multi-view learning technology aims to improve the performance and robustness of tasks such as classification, retrieval, and clustering by jointly modeling multi-source heterogeneous information (e.g., images and text, magnetic resonance imaging (MRI) and computed tomography (CT) in medical diagnosis, multi-sensor signals, and different feature extraction methods) of the same object (i.e., samples). In safety-critical scenarios such as autonomous driving, intelligent monitoring, medical imaging diagnosis, and industrial inspection, models not only need to provide predicted categories but also must be able to measure the reliability of the prediction results and provide uncertainty estimates to support risk control and human-machine collaborative decision-making. Therefore, reliable multi-view classification has gradually become a hot research and application area.
[0003] In the process of developing this application, the inventors discovered that current reliable multi-view classification methods mainly rely on ideal assumptions (that each view's information is clean, does not interfere with each other, and is strictly aligned). They have the drawbacks of over-reliance on a single evidence generation mechanism, which may lead to biased evidence estimation, and neglecting in-depth mining of evidence within a view, focusing only on the exploration and fusion of evidence between views. As a result, in actual conflicting multi-view data scenarios, current reliable multi-view classification methods have problems with poor classification performance and reliability of uncertainty estimation. Summary of the Invention
[0004] The purpose of this application is to provide a credible conflict multi-view classification method based on two-layer evidence exploration, which can solve the problems of poor classification performance and uncertainty estimation reliability of current credible multi-view classification methods in actual conflict multi-view data scenarios.
[0005] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide a credible conflict multi-view classification method based on two-layer evidence exploration, the method comprising: In response to receiving multi-view data associated with a multi-view classification task, evidence is extracted from the feature space and structure space for each view in the multi-view data to obtain feature evidence and neighborhood structure evidence for each view. The feature evidence and neighborhood structure evidence of each view are fused within the view to obtain a comprehensive evidence representation for each view; The comprehensive evidence representation of different views associated with the multi-view data is dynamically decoupled into consensus evidence and complementary evidence. The consensus evidence is used to characterize the degree of consistent support of different views in the same category judgment, and the complementary evidence is used to characterize the information increment carried by the differences between views. Based on the consensus evidence and the complementary evidence, the classification results and uncertainty estimation results of the multi-view data are determined.
[0006] Secondly, embodiments of this application provide a credible conflict multi-view classification device based on two-layer evidence exploration, the device comprising: The in-view evidence processing module is used to respond to the multi-view data associated with the multi-view classification task, extract evidence from the feature space and structure space for each view in the multi-view data respectively, and obtain the feature evidence and neighborhood structure evidence of each view; and perform in-view fusion on the feature evidence and neighborhood structure evidence of each view to obtain the comprehensive evidence representation of each view. The inter-view evidence processing module is used to dynamically decouple the comprehensive evidence representation of different views associated with the multi-view data into consensus evidence and complementary evidence. The consensus evidence is used to characterize the degree of consistent support of different views in the same category judgment, and the complementary evidence is used to characterize the information increment carried by the differences between views. The credible classification module is used to determine the classification result and uncertainty estimation result of the multi-view data based on the consensus evidence and the complementary evidence.
[0007] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of the credible conflict multi-view classification method based on two-layer evidence exploration as described in the first aspect.
[0008] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the credible conflict multi-view classification method based on two-layer evidence exploration as described in the first aspect.
[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the credible conflict multi-view classification method based on two-layer evidence exploration as described in the first aspect.
[0010] In this embodiment, a two-layer evidence exploration strategy is introduced. At the intra-view level, feature evidence and neighborhood structure evidence are jointly mined and fused within the view to generate a high-quality comprehensive evidence representation specific to the view. This improves the reliability of evidence from the source and reduces the risk of biased output caused by a single evidence generation mechanism. At the inter-view level, based on the inherent attribute of "consistency-complementarity" coexisting in multi-view data, the comprehensive evidence representations of different views are dynamically decoupled into consensus evidence and complementary evidence. This allows for the subsequent collaborative generation of classification results and uncertainty estimation results, thereby achieving effective collaborative modeling of cross-view shared information and view-specific information. Thus, this application can significantly improve the classification performance and uncertainty estimation reliability in real-world conflict-ridden multi-view data scenarios. Attached Figure Description
[0011] Figure 1 A flowchart illustrating the implementation of a credible conflict multi-view classification method based on two-layer evidence exploration, provided in this application embodiment; Figure 2 A schematic diagram illustrating the implementation process of a credible conflict multi-view classification method based on two-layer evidence exploration, provided in an embodiment of this application; Figure 3 A schematic diagram of a credible conflict multi-view classification device based on two-layer evidence exploration provided in an embodiment of this application; Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0014] First, to facilitate understanding of the technical solutions provided in this application, the main technical concepts involved in the embodiments of this application will be briefly explained below.
[0015] The Dirichlet distribution is an important multidimensional continuous probability distribution in statistics and machine learning. It is mainly used to model the distribution of probabilities of multiple mutually exclusive classes. It can generate a multidimensional probability vector, and each element ∈ [0,1] and the sum is 1.
[0016] Radial Basis Function (RBF) kernel: also known as Gaussian kernel or squared exponential kernel, it is widely used in various kernel learning algorithms and is suitable for handling nonlinear classification and regression problems.
[0017] Multi-view data is widespread in the real world, such as images and text in social media tweets and MRI and CT images in medical diagnosis. Compared to single-view data, multi-view data can describe the same object from different data sources, modalities, or sensor perspectives, offering advantages in information complementarity and redundancy verification. Correspondingly, compared to single-view learning, multi-view learning can simultaneously utilize the complementarity and consistency between different views, thus often exhibiting stronger generalization ability in complex environments. Therefore, multi-view learning methods have been widely researched and applied in tasks such as cluster analysis, information retrieval, and classification.
[0018] However, traditional multi-view classification methods typically focus on improving classification accuracy, paying insufficient attention to the reliability and uncertainty assessment of the prediction results. In safety-critical applications such as autonomous driving, medical diagnostics, financial analysis, and the Internet of Things, systems not only need to output accurate classification results but also provide the reliability of those results (i.e., uncertainty estimation results) to support further manual review or risk control. Therefore, reliable multi-view learning is gradually becoming a research hotspot.
[0019] In recent years, uncertainty modeling methods based on evidence learning and subjective logic (or evidence theory) can provide beliefs and uncertainties while outputting prediction results, thereby supporting credible decision-making. They have also been extended to multi-view learning to achieve multi-view evidence fusion, achieving good results.
[0020] Specifically, current credible multi-view classification methods typically employ a basic framework of "intra-view evidence generation + inter-view evidence fusion": at the intra-view level, evidence neural networks or evidence collectors are constructed for each view to map the original features into evidence quantities or confidence levels for each category, thereby obtaining the belief distribution and uncertainty for each view; at the inter-view level, Dempster-Shafer evidence theory or weighted fusion rules are used to synthesize multi-view evidence, outputting the final classification result and global uncertainty.
[0021] However, in the process of implementing this application, the inventors discovered that the current reliable multi-view classification method still has significant shortcomings when facing noisy and conflicting multi-view data, resulting in a significant performance degradation in actual conflicting data scenarios.
[0022] Specifically, current reliable multi-view classification methods are mostly based on ideal assumptions (that each view's information is clean, non-interfering, and strictly aligned), and achieve result output and uncertainty estimation through evidence learning and evidence fusion mechanisms. However, in practical applications, multi-view data is often affected by acquisition errors, equipment failures, transmission errors, and human factors, resulting in noise, conflicts, and misalignment issues. This leads to contradictory information between views, making current methods prone to bias in evidence generation and conflict resolution, thus affecting classification reliability. Specifically, current methods have at least the following shortcomings: First, current methods often rely excessively on a single evidence generation mechanism, meaning they extract evidence from the original feature space using only a single view-specific neural network. This can lead to biased evidence estimation and easily produce biased evidence output. When a view is contaminated by noise, occluded, or its discriminative information is more reflected in the structural relationships between samples, a single feature mapping is insufficient to fully mine the deep evidence information within the view. This can result in problems such as "incorrect but confident" or "weak and difficult-to-discriminate evidence," which can accumulate and amplify in subsequent fusion processes and directly affect the credibility of the final decision.
[0023] Secondly, current methods typically neglect in-depth mining of evidence within a view, focusing only on the exploration and fusion of evidence between views. That is, they concentrate their main improvements on inter-view fusion and conflict suppression strategies, ignoring that the quality of evidence within a view is the root cause of conflict robustness. When the quality of evidence within a view itself is insufficient, the fusion end can only perform "reduction-style" weight reduction, making it difficult to improve the ability to express evidence from the root, resulting in difficulty in stable operation under conditions of high conflict ratio or strong noise.
[0024] Based on the above analysis, current credible multi-view classification technology faces shortcomings in conflict data scenarios, such as biased evidence generation and insufficient deep evidence mining within views. This makes it difficult for the technology to achieve high accuracy, strong robustness, and credible uncertainty output in real-world environments where noise, view misalignment, and semantic conflicts are prevalent.
[0025] To address the aforementioned issues, this application provides a credible conflict multi-view classification method based on two-layer evidence exploration. This method improves evidence quality by jointly mining feature evidence and neighborhood structure evidence at the intra-view level and performing intra-view fusion. It also dynamically decouples consensus evidence and complementary evidence at the inter-view level to balance consistency and complementarity, making classification more robust in conflict scenarios. This significantly improves classification performance and uncertainty estimation reliability in conflict multi-view data scenarios.
[0026] The following description, in conjunction with the accompanying drawings, details a credible conflict multi-view classification method based on two-layer evidence exploration provided by this application, through specific embodiments and application scenarios.
[0027] See Figure 1 The diagram shown is an implementation flowchart of a credible conflict multi-view classification method based on two-layer evidence exploration provided in this application embodiment. The method may include the following steps: Step S101: In response to receiving the multi-view data associated with the multi-view classification task, extract evidence from the feature space and structure space for each view in the multi-view data to obtain the feature evidence and neighborhood structure evidence for each view.
[0028] In this embodiment of the application, considering that relying solely on direct features may overlook potential higher-order relationships between samples, this application introduces a joint mining strategy of feature discrimination evidence (i.e., feature evidence) and higher-order structural evidence (i.e., neighborhood structural evidence) to achieve in-depth mining of evidence within the view.
[0029] Specifically, for each view, two parallel evidence extraction branches can be constructed: a feature evidence branch and a structural evidence branch.
[0030] The feature evidence branch can employ evidence neural networks or evidence collectors to extract evidence from the feature space for each view in the multi-view data, thereby obtaining basic discriminative information from the original features of each view as feature evidence.
[0031] Structural evidence branches can employ graph convolutional networks, etc., to capture potential higher-order relationships between different samples in each view, thereby generating neighborhood structural evidence that reflects these higher-order relationships.
[0032] Step S102: Perform intra-view fusion of the feature evidence and neighborhood structure evidence for each view to obtain a comprehensive evidence representation for each view.
[0033] In practice, attention mechanisms and other methods can be used to fuse the feature evidence and neighborhood structure evidence of each view, thereby obtaining a comprehensive evidence representation for each view.
[0034] Therefore, by generating and fusing dual-source evidence of "feature evidence + structural evidence" within the above-mentioned view, the evidence estimation bias caused by a single evidence mechanism can be effectively reduced, thereby improving the credibility of the subsequent final decision.
[0035] Step S103: Dynamically decouple the comprehensive evidence representation of different views associated with the multi-view data into consensus evidence and complementary evidence, wherein the consensus evidence is used to characterize the degree of consistent support of different views on the same category judgment, and the complementary evidence is used to characterize the information increment carried by the differences between views.
[0036] In practical implementation, at the level between views, based on the inherent characteristic of "consistency-complementarity" coexisting in multi-view data, the comprehensive evidence representation of multiple views can be decomposed into consensus evidence and complementary evidence to balance consistency and complementarity, making the classification more robust in conflict scenarios.
[0037] For example, in the comprehensive evidence representation of V views In the process of dynamically decoupling consensus evidence and complementary evidence, for the component corresponding to the k-th category in the consensus evidence... , We can take the minimum value of the component corresponding to the k-th category in the comprehensive evidence representation of each view, that is: Thus, consensus evidence was obtained. Where K represents the number of categories in the multi-view classification task. Complementary evidence This can be obtained by averaging the differences between the synthesized evidence from each view and the consensus evidence, i.e.: .
[0038] It should be noted that the consensus evidence focuses on characterizing the degree of consistent support of different views in the same category judgment, and can be regarded as a reliable lower bound for cross-view evidence, used to suppress inflated evidence caused by noise and conflict; complementary evidence focuses on characterizing the information increment carried by the differences between views, emphasizing the unique advantages of each view in different categories, and used to supplement the discrimination clues that are difficult to cover by a single view.
[0039] Therefore, this application implements a "dynamic evidence decoupling and exploration" mechanism oriented towards the intrinsic attributes of multiple views, dynamically decomposing multi-view evidence at the inter-view level into values reflecting cross-view characteristics. Figure 1 Consensus evidence and complementary evidence reflecting view-specific contributions are used to subsequently collaboratively generate final classification results and global uncertainty estimation results, thereby achieving effective collaborative modeling of cross-view shared information and view-specific information.
[0040] Step S104: Based on the consensus evidence and the complementary evidence, determine the classification result and uncertainty estimation result of the multi-view data.
[0041] In practice, weighted sums and other methods can be used to incorporate consensus evidence. complementary evidence Processed as fused evidence ,For example Then, based on methods such as Dirichlet distribution calculation, the classification results and uncertainty estimation results of the multi-view data are determined according to the fused evidence to achieve reliable classification.
[0042] As can be seen from the above technical solution, this application introduces a two-layer evidence exploration strategy. At the intra-view level, feature evidence and neighborhood structure evidence are jointly mined and fused within the view to generate a high-quality comprehensive evidence representation specific to the view. This improves the reliability of evidence from the source and reduces the risk of biased output caused by a single evidence generation mechanism. At the inter-view level, based on the inherent attribute of "consistency-complementarity" coexisting in multi-view data, the comprehensive evidence representations of different views are dynamically decoupled into consensus evidence and complementary evidence. This allows for the subsequent collaborative generation of classification results and uncertainty estimation results, thereby achieving effective collaborative modeling of cross-view shared information and view-specific information. Thus, this application can significantly improve the classification performance and uncertainty estimation reliability in actual conflicting multi-view data scenarios.
[0043] In some embodiments, the intra-view fusion of feature evidence and neighborhood structure evidence for each view to obtain a comprehensive evidence representation for each view includes: Based on the pre-defined mapping relationship between belief quality and uncertainty quality and evidence, the feature evidence and neighborhood structure evidence of each view are converted into subjective logical opinions. The subjective logical opinions include the belief quality and uncertainty quality mapped by the corresponding evidence. The belief quality is used to characterize the deterministic support of the corresponding evidence for different categories of the multi-view classification task, and the uncertainty quality is used to characterize the uncertainty of the corresponding evidence. Using uncertainty quality as a reliability metric, the subjective logical opinions derived from the feature evidence and neighborhood structure evidence of each view are fused to obtain a fused opinion for each view. Based on the pre-defined mapping relationship between the quality of belief and the quality of uncertainty and the evidence, the fused opinion of each view is converted into a comprehensive evidence representation of each view.
[0044] In this embodiment, to achieve credible output, the feature evidence and neighborhood structure evidence of each view are first converted into subjective logical opinions (i.e., a joint "belief-uncertainty" representation of the classification results) based on the preset mapping relationship between belief quality and uncertainty quality and evidence, respectively. For example, the subjective logical opinions corresponding to the feature evidence and neighborhood structure evidence of view v can be expressed as follows: and ,in, and This represents the quality of beliefs mapped by the feature evidence and neighborhood structure evidence of view v, respectively. and This represents the quality of uncertainty mapped by the feature evidence and the neighborhood structure evidence of view v, respectively.
[0045] Then, a conflict opinion aggregation strategy is adopted to fuse the subjective logical opinions transformed from the feature evidence and neighborhood structure evidence of each view. That is, uncertainty quality is used as a reliability metric to give greater weight to branch evidence with less uncertainty (i.e. more reliable) during fusion, thereby improving the stability of fusion.
[0046] As one possible implementation, the fused opinion of view v includes the fused belief quality and uncertainty quality, and the component corresponding to the k-th category in the fused belief quality... and the quality of uncertainty after fusion Determined by the following formula:
[0047] in, and This represents the quality of beliefs mapped by the feature evidence and neighborhood structure evidence of view v, respectively. and The component corresponding to the k-th category, and This represents the quality of uncertainty mapped by the feature evidence and the neighborhood structure evidence of view v, respectively.
[0048] Finally, based on the preset mapping relationship between the quality of belief and the quality of uncertainty and the evidence, the fused opinion of each view is converted back into evidence, thus obtaining the comprehensive evidence representation of each view.
[0049] Optionally, the pre-defined mapping relationships between the quality of belief and the quality of uncertainty and the evidence are expressed as follows:
[0050] in, This represents either the feature evidence or the neighborhood structure evidence of view v. and express The quality of beliefs and the quality of uncertainty that are mapped Indicates the use of The total concentration parameters of the constructed Dirichlet distribution, express The total Dirichlet strength, where K represents the number of categories in the multi-view classification task.
[0051] In practical implementation, for any view Non-negative evidence (For example, feature evidence or neighborhood structure evidence), construct Dirichlet parameters for each sample and each class. As a component, it forms the total concentration parameter of the Dirichlet distribution. ,in, for The component corresponding to the k-th category is used to characterize the support strength of samples in view v for the k-th category. Accordingly, let... The total strength of evidence (i.e., Dirichlet total strength) Then the quality of belief The component corresponding to the kth category (Also known as category belief quality, used to characterize the deterministic support assigned to the k-th category), and uncertainty quality. (Used to characterize uncertainty, which varies with the total strength of evidence) (Increases and decreases) can be represented as follows:
[0052] And it satisfies the following normalization constraints:
[0053] Therefore, based on the above and From the expression, we can obtain reliable expressions for the pre-defined mapping relationships between the aforementioned belief quality and uncertainty quality and the evidence, respectively.
[0054] It is understandable that, based on the aforementioned pre-defined mapping relationship, it is possible to... The fused opinions (including the fused belief quality b and uncertainty quality u) are transformed into a comprehensive evidence representation e, where S represents the total Dirichlet strength of e and K represents the number of categories.
[0055] In some embodiments, the step of extracting evidence from the feature space and structure space for each view in the multi-view data to obtain feature evidence and neighborhood structure evidence for each view includes: The feature matrix of each view is input into an evidence-based deep neural network pre-trained independently for each view to obtain feature evidence for each view. The evidence-based deep neural network is constructed based on a multilayer perceptron (MLP) and a Softplus function. The feature matrix of each view is used to characterize the feature information of different samples under each view. The adjacency matrix and feature matrix of each view are input into a pre-trained graph convolutional network to obtain neighborhood structure evidence for each view. The adjacency matrix of each view is used to characterize the similarity and neighborhood structure between samples in the feature space of each view.
[0056] In practical implementation, the feature evidence branch uses an evidence-based deep neural network to extract feature evidence. This evidence-based deep neural network can be constructed based on a single-layer MLP and the Softplus function. The Softplus function is used to ensure non-negative output evidence to meet the requirements of Dirichlet parameter construction. Specifically, the aforementioned process can be represented as follows:
[0057] in, The feature evidence representing view v (also known as the evidence value matrix for each class, where each element in the matrix represents a feature evidence for view v) (This represents the strength of evidence that the i-th sample in the v-th view belongs to the k-th category). This represents the feature matrix of view v, where n represents the number of samples. This represents the feature dimension of view v, and K represents the number of categories. This represents the evidence-based deep neural network corresponding to view v. express Network parameters.
[0058] Understandably, the feature matrix for each view can be obtained from multi-view data (or the public dataset used during training). The element in the i-th row of the feature matrix represents the feature of the i-th sample, and the feature matrix for each view represents the feature information of different samples. For example, the Mfeat handwritten digit dataset contains ten Arabic numerals, with 200 images in each class, for a total of 2000 samples. This dataset has six view representations: 76-dimensional Fourier coefficient features of character shape, 216-dimensional contour correlation features, 64-dimensional Karhunen-Love coefficient features, pixel mean of 240 images from a 2x3 window, 47 Zernike moments, and 6 morphological features.
[0059] The structural evidence branch employs a graph convolutional network and can introduce non-negative constraints to extract neighborhood structural evidence.
[0060] Understandably, the adjacency matrix of each view can be obtained using methods such as multi-kernel learning. By using the adjacency matrix of each view as input to the graph convolutional network, the graph structure and information propagation path of the graph convolutional network can be defined. This allows the graph convolutional network to aggregate relevant neighbor node information, thereby extracting domain structure evidence that reflects the relationships between samples under each view, thus improving robustness.
[0061] For example, the graph convolution propagation process of a graph convolutional network can be represented as follows: Propagation at layer l (excluding the Softplus function):
[0062] Input (Level 0): , This represents the feature matrix (i.e., sample features) under view v.
[0063] Output of layer l: This represents the node representation after the l-th graph convolution propagation.
[0064] Finally, nonnegativity constraints are applied to the output to obtain evidence of the neighborhood structure. ,Right now:
[0065] Among them, after adding self-loops to the adjacency matrix A of the input graph convolutional network, we have I represents the identity matrix, and the corresponding degree matrix is... ; Indicates the first Layer trainable parameters, This represents the node representation of the last layer (i.e., the Lth layer). , and This represents the node representations of layer 0, layer 1, and layer (l+1), where n represents the number of samples. This represents the feature dimension of view v. Let K represent the feature dimension of the l-th layer of the graph convolutional network, and K represent the number of categories.
[0066] Thus, evidence of neighborhood structure is generated by the output of the last layer of the graph convolutional network, and non-negativity is guaranteed by the Softplus function.
[0067] Optionally, the adjacency matrix of each view is determined by the following formula:
[0068] in, This represents the adjacency matrix of view v. and This represents the kernel fusion weight of view v. and This represents the RBF kernel and polynomial kernel of view v.
[0069] In this embodiment, the adjacency matrix of any view v is obtained by weighted fusion of RBF kernel and polynomial kernel. To ensure that the obtained It can accurately characterize the similarity and neighborhood structure between samples in the corresponding view feature space; among them, The element in the i-th row and j-th column This indicates the correlation strength between sample i and sample j in view v.
[0070] As one possible implementation, the RBF core is defined as:
[0071] in, Indicates the kernel width parameter. This represents the original features of sample i,j in each view.
[0072] As one possible implementation, the polynomial kernel is defined as:
[0073] in, Indicates the bias term. Indicates the degree of the polynomial. This represents the original features of sample i,j in each view.
[0074] In some embodiments, the loss functions used by the evidential deep neural network and the graph convolutional network during the training phase are represented as follows:
[0075]
[0076]
[0077]
[0078]
[0079] in, Indicates the total loss; This represents the classification loss for the v-th view; The evidence cross-entropy loss for the v-th view is used to constrain the correct category to generate greater evidence. This represents the KL regularization term for the v-th view, used to prevent the wrong category from generating too much evidence; This indicates a loss of consensus evidence; This indicates a loss of complementary evidence; This represents the conflict consistency loss, used to quantify and suppress cross-view conflicts; Indicates the loss weighting coefficient; This represents the annealing factor, where V represents the number of views in the multi-view data, and the set of all view pairs. , Indicates the evidence of the i-th view. Evidence from the j-th view (For example, the degree of conflict between synthesized evidence presented in different views) express and The predicted distribution distance between them express and The weighting terms between them and express and The strength of the evidence.
[0080] In this embodiment, a classification loss is introduced into the loss function, and an evidence cross-entropy loss is introduced into the classification loss to constrain the generation of greater evidence for the correct category. For example, based on samples from a certain view... The Dirichlet parameter (i.e., the total concentration parameter of the Dirichlet distribution). Total strength of Dirichlet Its evidence cross-entropy loss It can be represented as:
[0081] in, Indicates the truth label, Represents the Digamma function. This represents the Dirichlet concentration parameter of the nth sample in the jth class, i.e. The j-th component, where K represents the number of categories.
[0082] To avoid misclassification generating excessive evidence, a KL regularization term is further introduced into the classification loss. For example, based on samples from a certain view... Dirichlet parameters Its KL regularization term It can be represented as:
[0083] Among them, the adjusted Dirichlet parameters It is used to ensure that the model does not over-penalize evidence in the correct category, while appropriately penalizing evidence in the wrong category; Indicates For the Dirichlet distribution of the concentration parameter, This represents a uniform Dirichlet distribution.
[0084] For consensus evidence, a consensus evidence loss is further introduced into the loss function, for example, consensus evidence loss. It can be represented as:
[0085] in, Indicates sample The Dirichlet parameters of the corresponding consensus evidence. This represents the annealing coefficient.
[0086] For complementary evidence, a complementary evidence loss is further introduced into the loss function, for example, complementary evidence loss. It can be represented as:
[0087] in, and Indicates sample The Dirichlet parameters and total Dirichlet strength of the corresponding complementary evidence.
[0088] Given that current conflict measurement methods mostly characterize the degree of conflict based on the differences in the probability distribution of view outputs (such as KL divergence, JS divergence, or cosine distance), they often fail to fully consider the impact of evidence quality factors (such as certainty and strength) on conflict assessment, resulting in insufficient precision in conflict quantification: when a view has low evidence strength and high uncertainty, its distribution difference does not necessarily represent the real risk of conflict; conversely, when a view has high evidence strength but points to the wrong category, relying solely on distribution distance may not accurately assess its harm to fusion.
[0089] The root cause of the above problems is that multi-view conflicts are not just "different outputs", but also the result of the intertwining and superposition of "inconsistencies between high-quality evidence" and "noise disturbances from low-quality evidence". Current solutions lack a mechanism to incorporate evidence quality-related information into conflict measurement and consistency constraints, causing conflict quantification and fusion scheduling to deviate from actual reliability requirements.
[0090] To address the aforementioned issues, this application proposes a conflict measurement and consistency constraint mechanism based on evidence quality. Specifically, to quantify and suppress cross-view conflicts, this application further introduces the aforementioned conflict consistency loss based on evidence quality into the loss function. This utilizes a conflict measurement weighted by evidence quality to suppress cross-view conflicts, effectively handling conflict information between multiple views, thereby achieving more accurate conflict modeling and more robust and reliable fusion.
[0091] After calculating the total loss, the Adam optimizer can be used to update the model parameters of the evidence-based deep neural network and the graph convolutional network to achieve model training.
[0092] As one possible implementation method, and Predicted distribution distance between It can be represented as:
[0093] in, and This represents the predicted class probability vectors for the same sample from the i-th and j-th views, used to measure the difference in the predicted distributions of the two views.
[0094] As one possible implementation, the evidentiary strength of the v-th view evidence It is expressed as follows:
[0095] in, Let K represent the total Dirichlet intensity of the v-th view, and K represent the number of categories.
[0096] Optionally, the method further includes: During the training phase, the features of each view associated with each multi-view training data are concatenated along the feature dimension to obtain the augmented view associated with each multi-view training data. After taking the augmented views associated with each multi-view training data as the newly added views for each multi-view training data, the evidence-based deep neural network and the graph convolutional network are trained using the training set composed of the multi-view training data.
[0097] In practice, the features of each view are first concatenated along the feature dimension to construct the enhanced view. This process can be represented as follows:
[0098] in, Indicates an enhanced view. This represents the view features of each of the V views in the multi-view training data, where n represents the number of samples. This represents the feature dimension of the v-th view.
[0099] Then, the enhanced view will be used as the first Each view input, compared to the original All views participate in the model training process to enhance the correlation between views.
[0100] Understandably, after introducing augmented views, the loss functions used in the training phase of the aforementioned evidence-based deep neural networks and graph convolutional networks can be adjusted accordingly:
[0101] In this embodiment of the application, it is considered that traditional methods generally lack effective data augmentation or global view construction mechanisms to enhance cross-view functionality. Figure 1 Consistency and complementarity modeling makes the model more prone to overfitting view-specific noise in conflict environments, making it difficult to stably capture consensus information across views and make full use of complementary information, thus limiting the effectiveness and stability of the algorithm in real complex scenarios.
[0102] To address the aforementioned issues, this application constructs a global fusion view by stitching together features from multiple views as a data augmentation strategy. This involves stitching together features from each view along the feature dimension to obtain a global fusion view as an augmented view. This augmented view is then used as an additional view in training and inference to enhance cross-view correlation and improve fusion robustness. This can further improve classification performance and uncertainty estimation reliability in conflicting multi-view data scenarios.
[0103] In some embodiments, before determining the classification result and uncertainty estimation result of the multi-view data based on the consensus evidence and the complementary evidence, the method further includes: Determine the belief quality of the consensus evidence, which characterizes the deterministic support of assigning the corresponding evidence to different categories of the multi-view classification task; The belief quality of the consensus evidence is adjusted based on exponentiation to amplify the relative advantage of high-confidence categories and suppress low-confidence categories.
[0104] In this embodiment of the application, the separation degree between classes is defined as follows:
[0105] in, and The values represent the belief quality of the i-th and j-th categories, K represents the number of categories, and the separation degree between categories. SD ( b The sum of the absolute differences in confidence quality between all categories is the sum of the absolute differences in confidence quality. The higher the value, the stronger the discriminative power of the evidence.
[0106] Based on the above definition of separation between classes, in order to improve the separation of consensus viewpoints, this application enhances the confidence level (i.e., belief quality) of consensus evidence by exponentiation. For example, the aforementioned process can be represented as follows:
[0107] in, This represents the quality of category belief under the k-th category corresponding to the consensus evidence. Indicates that the separation has been enhanced. , This indicates the preset index.
[0108] Therefore, by exponentializing the quality of consensus beliefs, we can amplify the relative advantage of high-confidence categories and suppress low-confidence categories, thereby increasing the differences between different categories and making consensus evidence more discriminative.
[0109] As one possible implementation, the quality of consensus evidence belief enhanced by separation can be standardized to maintain a constant total evidence. For example, the aforementioned process can be represented as follows:
[0110] Where K represents the number of categories.
[0111] In some embodiments, the uncertainty estimation result is expressed as follows:
[0112] in, The uncertainty estimation results represent fused evidence comprised of consensus evidence and complementary evidence. Uncertainty quality; express Total Dirichlet strength; Indicates the use of The k-th component of the total concentration parameter in the constructed Dirichlet distribution; K represents the number of categories in the multi-view classification task.
[0113] For example, see Figure 2 The diagram illustrates the implementation process of the credible conflict multi-view classification method based on two-layer evidence exploration. This implementation process includes the following steps: S1: Collect multi-view data and preprocess it to obtain the training dataset and the test dataset.
[0114] In this step, for all datasets, normalization is first performed, and then 80% of the data is randomly selected as the training set and 20% as the test set. In order to generate a conflict test set, Gaussian noise with a specific standard deviation can be added to 10% of the data in the test set, and the view information of 40% of the data in the test set can be randomly modified to cause view information misalignment and conflict.
[0115] S2: Construct a two-layer evidence exploration framework, including in-view evidence exploration and fusion and inter-view evidence decoupling and fusion.
[0116] In this step, a hierarchical evidence mining strategy is adopted. First, a dual-branch intra-view evidence exploration module is constructed. At the intra-view level, a deep neural network for evidence and a graph convolutional network are combined to extract evidence from the feature space and structure space for each view, obtaining feature evidence and neighborhood structure evidence. Intra-view fusion is then performed by calculating the corresponding belief values and uncertainties to generate a view-specific comprehensive evidence representation. Next, an inter-view evidence decoupling and fusion module is constructed. This dynamically decomposes the multi-view comprehensive evidence representation into consensus evidence and complementary evidence, and enhances the separability of the consensus evidence, ultimately obtaining fused evidence. This achieves the systematic exploration and fusion of evidence information at both the intra-view and inter-view levels.
[0117] S3: Input the training set into the above network structure for training, and optimize the model parameters through the loss function.
[0118] In this step, a conflict metric and consistency constraint module based on evidence quality is constructed. During model training, the conflict metric weighted by evidence quality is used to suppress cross-view conflicts, so as to make conflict quantification more accurate and promote training stability.
[0119] S4: Input test sample multi-view data, output classification results and corresponding uncertainty estimates.
[0120] In this step, the trained neural network model predicts the category of each sample in the test dataset containing conflict views and outputs the uncertainty of the corresponding classification results.
[0121] Specifically, the test set is input into the trained model to obtain fused evidence. This leads to the Dirichlet parameters. Based on this, the classification results and uncertainty estimation results are output to improve the reliability of the results and achieve reliable classification output.
[0122] To verify the performance of this application in the conflict multi-view classification scenario, as shown in Table 1 below, systematic experimental evaluations were conducted on seven publicly available multi-view benchmark datasets (Scene, PIE, UCI-3view, Mfeat, ALOI, NUS3w, Animal).
[0123] In the experiment, following a unified data partitioning strategy, 80% of the samples were randomly selected as the training set and 20% as the test set. Simultaneously, to simulate noise pollution and view misalignment conflicts commonly found in real-world applications, this application further constructed a conflict test scenario: Gaussian noise (e.g., proportionally adding noise with a standard deviation of σ) was injected into a certain proportion of samples in the test set, and the information of a certain view was randomly modified in some samples in the test set to create view misalignment, thus forming a test set setting containing noise and conflict. Regarding evaluation metrics, this application uses classification accuracy as the main performance indicator. Experimental results show that under normal data conditions, this application achieves optimal or near-optimal classification performance on most datasets. Under conflict data conditions, compared with normal testing, all comparative methods showed varying degrees of performance degradation, but the degradation of this application was significantly smaller across all datasets, demonstrating stronger conflict robustness and stability. Furthermore, it can effectively distinguish the uncertainty levels of samples under different conflict intensities, verifying the ability of this application to output reliable predictions and risk measurements in conflict environments.
[0124] Table 1 Experimental Evaluation Results
[0125] In summary, this application proposes a credible conflict multi-view classification method based on two-layer evidence exploration for scenarios involving noise and view misalignment. At the intra-view level, it jointly utilizes a deep neural network for evidence and a graph convolutional network to mine feature-based discriminative evidence and neighborhood structure evidence, respectively. An intra-view fusion strategy is then employed to generate a high-quality, comprehensive evidence representation specific to each view, thereby improving evidence reliability from the source and reducing the risk of biased output from a single evidence generation mechanism. Simultaneously, at the inter-view level, based on the inherent property of "consistency-complementarity" coexisting in multi-view data, multi-view evidence is dynamically decoupled into consensus evidence and complementary evidence for collaborative generation of the final classification result and global uncertainty estimation. This achieves effective collaborative modeling of cross-view shared information and view-specific information. Experimental results fully demonstrate that, under multiple publicly available multi-view datasets and normal / conflict dual test settings, the method provided in this application outperforms existing state-of-the-art methods in key indicators such as classification accuracy. Furthermore, under conflict test conditions, the performance degradation is smaller, and the output uncertainty and conflict measurement are more discriminative and consistent. This proves its comprehensive advantages in reliable classification and credible risk assessment in conflict environments, providing a reliable technical path and broad application prospects for building a credible multi-view intelligent decision-making system for real, complex, multi-source data.
[0126] It should be noted that this application addresses the real-world pain point of "conflicting multi-view data caused by noise and view misalignment" by creatively constructing a unified framework for "intra-view and inter-view" two-layer evidence exploration and credible fusion. This framework enables the model to not only output classification results but also provide stable uncertainty estimates in conflict environments and significantly reduce the negative impact of conflicting views on decision-making.
[0127] Compared to the limitations of credible multi-view methods in related technologies, which mainly focus on evidence fusion between views or rely on a single evidence generation mechanism, this application achieves end-to-end collaborative optimization from evidence generation and conflict measurement to fusion decision-making through hierarchical mining, multi-granularity evidence representation and evidence quality-driven conflict modeling, thereby possessing stronger robustness and credibility in conflict data scenarios.
[0128] Specifically, this application proposes a "dual-layer evidence exploration" intra-view evidence mining mechanism. For each view, it simultaneously introduces discriminative feature evidence learning based on deep neural networks and neighborhood structure evidence learning based on graph convolutional networks, and synthesizes these two complementary types of evidence into a view-level comprehensive evidence representation through an intra-view fusion strategy. Secondly, this application introduces a "dynamic evidence decoupling and exploration" mechanism oriented towards the intrinsic attributes of multiple views, dynamically decomposing multi-view evidence at the inter-view level into representations reflecting cross-view... Figure 1 This application presents both consistent consensus evidence and complementary evidence reflecting the unique contributions of each view. Finally, it designs a "conflict measurement and consistency constraint based on evidence quality" mechanism, breaking away from the traditional approach of measuring conflict solely based on probability distribution differences, thus achieving more accurate conflict quantification. Simultaneously, this application further constructs a globally fused view by stitching together multi-view features as a data augmentation strategy to strengthen the explicit expression of cross-view correlation information, improve the model's ability to capture consensus information, and enhance stable fusion in conflict scenarios. Therefore, this application avoids the problems of classification bias and unreliable uncertainty estimation in noisy, misaligned, or cross-view conflict scenarios, making it difficult to apply in practical key scenarios.
[0129] It should be noted that the credible conflict multi-view classification method based on two-layer evidence exploration provided in this application embodiment can be executed by a credible conflict multi-view classification device based on two-layer evidence exploration, or by a control module in that device for executing the loading of the credible conflict multi-view classification method based on two-layer evidence exploration. This application embodiment uses the execution of the loading of the credible conflict multi-view classification method based on two-layer evidence exploration by a credible conflict multi-view classification device as an example to illustrate the credible conflict multi-view classification method based on two-layer evidence exploration provided in this application embodiment.
[0130] This application provides a credible conflict multi-view classification device based on two-layer evidence exploration, such as... Figure 3 As shown, the device includes: The in-view evidence processing module is used to respond to the multi-view data associated with the multi-view classification task, extract evidence from the feature space and structure space for each view in the multi-view data respectively, and obtain the feature evidence and neighborhood structure evidence of each view; and perform in-view fusion on the feature evidence and neighborhood structure evidence of each view to obtain the comprehensive evidence representation of each view. The inter-view evidence processing module is used to dynamically decouple the comprehensive evidence representation of different views associated with the multi-view data into consensus evidence and complementary evidence. The consensus evidence is used to characterize the degree of consistent support of different views in the same category judgment, and the complementary evidence is used to characterize the information increment carried by the differences between views. The credible classification module is used to determine the classification result and uncertainty estimation result of the multi-view data based on the consensus evidence and the complementary evidence.
[0131] Optionally, the in-view evidence processing module is further configured to perform the following steps: Based on the pre-defined mapping relationship between belief quality and uncertainty quality and evidence, the feature evidence and neighborhood structure evidence of each view are converted into subjective logical opinions. The subjective logical opinions include the belief quality and uncertainty quality mapped by the corresponding evidence. The belief quality is used to characterize the deterministic support of the corresponding evidence for different categories of the multi-view classification task, and the uncertainty quality is used to characterize the uncertainty of the corresponding evidence. Using uncertainty quality as a reliability metric, the subjective logical opinions derived from the feature evidence and neighborhood structure evidence of each view are fused to obtain a fused opinion for each view. Based on the pre-defined mapping relationship between the quality of belief and the quality of uncertainty and the evidence, the fused opinion of each view is converted into a comprehensive evidence representation of each view.
[0132] Optionally, the pre-defined mapping relationships between the quality of belief and the quality of uncertainty, and the evidence, respectively, are expressed as follows:
[0133] in, This represents either the feature evidence or the neighborhood structure evidence of view v. and express The quality of beliefs and the quality of uncertainty that are mapped Indicates the use of The total concentration parameters of the constructed Dirichlet distribution, express The total Dirichlet strength, where K represents the number of categories in the multi-view classification task.
[0134] Optionally, the in-view evidence processing module is further configured to perform the following steps: The feature matrix of each view is input into an evidence-based deep neural network pre-trained independently for each view to obtain feature evidence for each view. The evidence-based deep neural network is constructed based on a multilayer perceptron and a Softplus function. The feature matrix of each view is used to characterize the feature information of different samples under each view. The adjacency matrix and feature matrix of each view are input into a pre-trained graph convolutional network to obtain neighborhood structure evidence for each view. The adjacency matrix of each view is used to characterize the similarity and neighborhood structure between samples in the feature space of each view.
[0135] Optionally, the adjacency matrix of each view is determined by the following formula:
[0136] in, This represents the adjacency matrix of view v. and This represents the kernel fusion weight of view v. and This represents the RBF kernel and polynomial kernel of view v.
[0137] Optionally, the loss function used by the evidential deep neural network and the graph convolutional network during the training phase is expressed as follows:
[0138]
[0139]
[0140]
[0141]
[0142] in, Indicates the total loss; This represents the classification loss for the v-th view; This represents the evidence cross-entropy loss, used to constrain the correct category to generate more evidence; This indicates a KL regularization term, used to prevent incorrect categories from generating excessive evidence; This indicates a loss of consensus evidence; This indicates a loss of complementary evidence; This represents the conflict consistency loss, used to quantify and suppress cross-view conflicts; Indicates the loss weighting coefficient; This represents the annealing factor, where V represents the number of views in the multi-view data, and the set of all view pairs. , Indicates the evidence of the i-th view. Evidence from the j-th view The degree of conflict between them express and The predicted distribution distance between them express and The weighting terms between them and express and The strength of the evidence.
[0143] Optionally, the device further includes a model training module for performing the following steps: During the training phase, the features of each view associated with each multi-view training data are concatenated along the feature dimension to obtain the augmented view associated with each multi-view training data. After taking the augmented views associated with each multi-view training data as the newly added views for each multi-view training data, the evidence-based deep neural network and the graph convolutional network are trained using the training set composed of the multi-view training data.
[0144] Optionally, the trusted classification module is further configured to perform the following steps: Determine the belief quality of the consensus evidence, which characterizes the deterministic support of assigning the corresponding evidence to different categories of the multi-view classification task; The belief quality of the consensus evidence is adjusted based on exponentiation to amplify the relative advantage of high-confidence categories and suppress low-confidence categories.
[0145] Optionally, the uncertainty estimation result is expressed as follows:
[0146] in, The uncertainty estimation results represent fused evidence comprised of consensus evidence and complementary evidence. Uncertainty quality; express Total Dirichlet strength; Indicates the use of The k-th component of the total concentration parameter in the constructed Dirichlet distribution; K represents the number of categories in the multi-view classification task.
[0147] As can be seen from the above technical solution, this application introduces a two-layer evidence exploration strategy. At the intra-view level, feature evidence and neighborhood structure evidence are jointly mined and fused within the view to generate a high-quality comprehensive evidence representation specific to the view. This improves the reliability of evidence from the source and reduces the risk of biased output caused by a single evidence generation mechanism. At the inter-view level, based on the inherent attribute of "consistency-complementarity" coexisting in multi-view data, the comprehensive evidence representations of different views are dynamically decoupled into consensus evidence and complementary evidence. This allows for the subsequent collaborative generation of classification results and uncertainty estimation results, thereby achieving effective collaborative modeling of cross-view shared information and view-specific information. Thus, this application can significantly improve the classification performance and uncertainty estimation reliability in actual conflicting multi-view data scenarios.
[0148] The trusted conflict multi-view classification device based on two-layer evidence exploration in this application embodiment can be a device, or it can be a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network attached storage (NAS), PCs, televisions (TVs), ATMs, or self-service machines, etc., and this application embodiment does not specifically limit the scope.
[0149] The credible conflict multi-view classification device based on two-layer evidence exploration in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0150] The credible conflict multi-view classification device based on two-layer evidence exploration provided in this application embodiment can achieve... Figure 1 The various processes implemented in the embodiment of the credible conflict multi-view classification method based on two-layer evidence exploration shown are not described in detail here to avoid repetition.
[0151] Optionally, this application embodiment also provides an electronic device, including a processor 110, a memory 109, and a program or instructions stored in the memory 109 and executable on the processor 110. When the program or instructions are executed by the processor 110, they implement the various processes of the above-described embodiments of the credible conflict multi-view classification method based on two-layer evidence exploration and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0152] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0153] Figure 4 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0154] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.
[0155] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0156] The processor 110 is used to perform the following steps: In response to receiving multi-view data associated with a multi-view classification task, evidence is extracted from the feature space and structure space for each view in the multi-view data to obtain feature evidence and neighborhood structure evidence for each view. The feature evidence and neighborhood structure evidence of each view are fused within the view to obtain a comprehensive evidence representation for each view; The comprehensive evidence representation of different views associated with the multi-view data is dynamically decoupled into consensus evidence and complementary evidence. The consensus evidence is used to characterize the degree of consistent support of different views in the same category judgment, and the complementary evidence is used to characterize the information increment carried by the differences between views. Based on the consensus evidence and the complementary evidence, the classification results and uncertainty estimation results of the multi-view data are determined.
[0157] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described embodiments of the credible conflict multi-view classification method based on two-layer evidence exploration, and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0158] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0159] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described embodiments of the credible conflict multi-view classification method based on two-layer evidence exploration, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0160] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0161] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0163] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A credible conflict multi-view classification method based on two-layer evidence exploration, characterized in that, The method includes: In response to receiving multi-view data associated with a multi-view classification task, evidence is extracted from the feature space and structure space for each view in the multi-view data to obtain feature evidence and neighborhood structure evidence for each view. The feature evidence and neighborhood structure evidence of each view are fused within the view to obtain a comprehensive evidence representation for each view; The comprehensive evidence representation of different views associated with the multi-view data is dynamically decoupled into consensus evidence and complementary evidence. The consensus evidence is used to characterize the degree of consistent support of different views in the same category judgment, and the complementary evidence is used to characterize the information increment carried by the differences between views. Based on the consensus evidence and the complementary evidence, the classification results and uncertainty estimation results of the multi-view data are determined.
2. The method according to claim 1, characterized in that, The in-view fusion of feature evidence and neighborhood structure evidence for each view to obtain a comprehensive evidence representation for each view includes: Based on the pre-defined mapping relationship between belief quality and uncertainty quality and evidence, the feature evidence and neighborhood structure evidence of each view are converted into subjective logical opinions. The subjective logical opinions include the belief quality and uncertainty quality mapped by the corresponding evidence. The belief quality is used to characterize the deterministic support of the corresponding evidence for different categories of the multi-view classification task, and the uncertainty quality is used to characterize the uncertainty of the corresponding evidence. Using uncertainty quality as a reliability metric, the subjective logical opinions derived from the feature evidence and neighborhood structure evidence of each view are fused to obtain a fused opinion for each view. Based on the pre-defined mapping relationship between the quality of belief and the quality of uncertainty and the evidence, the fused opinion of each view is converted into a comprehensive evidence representation of each view.
3. The method according to claim 2, characterized in that, The pre-defined mapping relationships between the quality of belief and the quality of uncertainty, and evidence, are expressed as follows: in, This represents either the feature evidence or the neighborhood structure evidence of view v. and express The quality of beliefs and the quality of uncertainty that are mapped Indicates the use of The total concentration parameters of the constructed Dirichlet distribution, express The total Dirichlet strength, where K represents the number of categories in the multi-view classification task.
4. The method according to claim 1, characterized in that, The step of extracting evidence from the feature space and structure space for each view in the multi-view data to obtain feature evidence and neighborhood structure evidence for each view includes: The feature matrix of each view is input into an evidence-based deep neural network pre-trained independently for each view to obtain feature evidence for each view. The evidence-based deep neural network is constructed based on a multilayer perceptron and a Softplus function. The feature matrix of each view is used to characterize the feature information of different samples under each view. The adjacency matrix and feature matrix of each view are input into a pre-trained graph convolutional network to obtain neighborhood structure evidence for each view. The adjacency matrix of each view is used to characterize the similarity and neighborhood structure between samples in the feature space of each view.
5. The method according to claim 4, characterized in that, The adjacency matrix of each view is determined by the following formula: in, This represents the adjacency matrix of view v. and This represents the kernel fusion weight of view v. and This represents the RBF kernel and polynomial kernel of view v.
6. The method according to claim 4, characterized in that, The loss functions used by the evidential deep neural network and the graph convolutional network during the training phase are expressed as follows: in, Indicates the total loss; This represents the classification loss for the v-th view; The evidence cross-entropy loss for the v-th view is used to constrain the correct category to generate greater evidence. This represents the KL regularization term for the v-th view, used to prevent the wrong category from generating too much evidence; This indicates a loss of consensus evidence; This indicates a loss of complementary evidence; This represents the conflict consistency loss, used to quantify and suppress cross-view conflicts; Indicates the loss weighting coefficient; This represents the annealing factor, where V represents the number of views in the multi-view data, and the set of all view pairs. , Indicates the evidence of the i-th view. Evidence from the j-th view The degree of conflict between them express and The predicted distribution distance between them express and The weighting terms between them and express and The strength of the evidence.
7. The method according to claim 4, characterized in that, The method further includes: During the training phase, the features of each view associated with each multi-view training data are concatenated along the feature dimension to obtain the augmented view associated with each multi-view training data. After taking the augmented views associated with each multi-view training data as the newly added views for each multi-view training data, the evidence-based deep neural network and the graph convolutional network are trained using the training set composed of the multi-view training data.
8. The method according to any one of claims 1-7, characterized in that, Before determining the classification result and uncertainty estimation result of the multi-view data based on the consensus evidence and the complementary evidence, the method further includes: Determine the belief quality of the consensus evidence, which characterizes the deterministic support of assigning the corresponding evidence to different categories of the multi-view classification task; The belief quality of the consensus evidence is adjusted based on exponentiation to amplify the relative advantage of high-confidence categories and suppress low-confidence categories.
9. The method according to any one of claims 1-7, characterized in that, The uncertainty estimation results are expressed as follows: in, The uncertainty estimation results represent fused evidence comprised of consensus evidence and complementary evidence. Uncertainty quality; express Total Dirichlet strength; Indicates the use of The k-th component of the total concentration parameter in the constructed Dirichlet distribution; K represents the number of categories in the multi-view classification task.
10. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the credible conflict multi-view classification method based on two-layer evidence exploration as described in any one of claims 1-9.