Multi-label medical diagnosis man-machine cooperation method based on confusion matrix and hybrid expert system

By combining a multi-label medical diagnostic method based on confusion matrix and hybrid expert system, and using a maximum entropy model and multi-label confusion matrix to generate transition probability matrix and calibration weights, the prediction results of AI and human experts are integrated, which solves the problem of insufficient label correlation modeling in multi-label diagnosis, improves diagnostic accuracy and reduces misdiagnosis and missed diagnosis rates.

CN121075604APending Publication Date: 2025-12-05NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511168089.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing human-machine decision fusion methods suffer from problems such as ineffective modeling of label correlation, high computational cost, and weak generalization ability in multi-label medical diagnosis. In particular, they cannot accurately assess the confidence of human experts in multi-label scenarios, resulting in limited improvement in diagnostic accuracy and performance.

Method used

The transition probability matrix is ​​generated by the maximum entropy model as a weighted multi-label confusion matrix. The initial weights are generated by using the prediction results of AI and human experts. The calibration weights are then generated. Finally, the prediction results of AI and human experts are fused through the multi-label confusion matrix to generate the diagnostic results.

Benefits of technology

It achieves more precise weight allocation, improves the accuracy of multi-label medical diagnosis, reduces misdiagnosis and missed diagnosis, and assesses the confidence of human experts through multi-label confusion matrix, quantifies the systematic bias of expert prediction, optimizes the independent decision threshold, and reduces missed diagnosis and misdiagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075604A_ABST
    Figure CN121075604A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of man-machine decision fusion. The invention provides a multi-label medical diagnosis man-machine cooperation method based on a confusion matrix and a hybrid expert system. According to the embodiment of the invention, a man-machine decision fusion framework based on expert mixing is provided, the human experts are combined with the global modeling capability of the AI model, the initial man-machine weight is generated by the gating network, and the initial weight is calibrated through the multi-label confusion matrix. The correlation between the labels is fully considered to construct a confusion matrix to evaluate the confidence of the human experts, a maximum entropy model which does not need to be assumed on the premise is used to better conform to a real scene, the dependency relationship between the labels is captured, systematic deviation of expert prediction is quantified, and accurate evaluation of the confidence of the human experts is achieved. For heterogeneity of multi-label classification, an independent decision threshold is optimized for each label, Hamming loss is minimized through grid search, probability output is converted into a binary diagnosis result, and the phenomena of missed diagnosis and misdiagnosis are effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the present disclosure relates to the technical field of human-computer decision fusion, in particular to a multi-label medical diagnosis human-computer collaboration method based on a confusion matrix and a hybrid expert system. BACKGROUND

[0002] With the rapid development and increasing popularity of artificial intelligence (AI) technology, it provides help to improve the diagnosis and treatment level of primary medical institutions. However, for a diagnosis report or chest film, multiple diseases can often be diagnosed, which belongs to the problem of multi-label classification (Mulit-label Classification). In multi-label medical diagnosis, accurately modeling label correlation is the key to improving classification accuracy. Existing AI models explore label correlation through different strategies, including modeling the spatial context of label co-occurrence patterns, and utilizing latent contextual information to embed the semantic relationship between labels. However, in the medical diagnosis task, spatial or latent contextual clues are often subtle or missing, limiting the effectiveness of these methods. In addition, such methods usually have high computational cost and weak generalization ability in different clinical scenarios.

[0003] Therefore, we advocate the introduction of human experts, whose experience judgments are irreplaceable when dealing with complex information. This human-computer collaboration has been successfully applied in multiple medical fields, including complex decisions in online disease diagnosis, skin cancer recognition, and sepsis diagnosis. The human-computer collaboration mode provides broad potential for MLC problems, especially for dealing with complex correlations between labels. Existing methods mainly focus on single-label classification tasks, including combining human predictions with model probabilities through confusion matrices and calibration, delegating tasks to human experts, and assigning tasks to human experts or AI models. However, these methods cannot be directly extended to the MLC task of medical diagnosis because they usually lack explicit modeling of label correlation. This limitation leads to limited performance improvement in complex MLC scenarios.

[0004] However, the existing human-computer decision fusion method still has the following deficiencies: (1) It mainly targets single-label scenarios and needs to assume independence between labels, which does not conform to the actual situation and ignores the complex relationships between labels in multi-label problems in diagnosis.

[0005] (2) The confusion matrix used to evaluate the ability of human experts is also a confusion matrix designed for single-label problems extended to multi-label problems. When dealing with large label sets, it will lead to exponential growth in computational complexity and face the problem of data sparsity, and cannot capture different dependency strengths between different pairs of labels, limiting its effectiveness in real multi-label scenarios.

[0006] (3) The method for exploring the correlation between labels is difficult to capture the correlation between labels without spatial / semantic context relationship, has high computational cost, and has weak generalization ability in different clinical scenarios.

[0007] Therefore, it is necessary to improve one or more problems existing in the above related technical solutions.

[0008] It should be noted that this section aims to provide background or context for the technical solutions of the disclosure stated in the claims. The description herein is not admitted to be prior art merely because it is included in this section. SUMMARY

[0009] The purpose of the embodiments of the present disclosure is to provide a multi-label medical diagnosis human-computer collaboration method based on confusion matrix and hybrid expert system, thereby at least partially overcoming one or more problems caused by the limitations and defects of the related art.

[0010] According to the embodiments of the present disclosure, a multi-label medical diagnosis human-computer collaboration method based on confusion matrix and hybrid expert system is provided, which comprises: Modeling label correlation by a maximum entropy model to generate a transition probability matrix, and using the transition probability matrix as a basis for weighting to construct a multi-label confusion matrix; Extracting visual features of medical images, and generating initial weights according to the visual features, AI prediction results and expert prediction results; Generating confusion influence factors according to the multi-label confusion matrix; Fusing the initial weights and the confusion influence factors, and dynamically calibrating the initial weights of human experts and AI models to generate calibrated weights; Fusing the AI prediction results, the expert prediction results and the calibrated weights to generate probability outputs; Converting the probability outputs into binary predictions using label-specific thresholds to obtain diagnosis results.

[0011] Further, in the step of modeling label correlation by a maximum entropy model to generate a transition probability matrix, and using the transition probability matrix as a basis for weighting to construct a multi-label confusion matrix, it comprises: Defining the label correlation as the probability of label i being misclassified as label j, estimating the conditional probability by a maximum entropy model, and generating a transition probability matrix; Using the transition probability matrix as a basis for weighting to construct a multi-label confusion matrix, and updating the multi-label confusion matrix according to a preset rule; wherein the mode of the multi-label confusion matrix includes perfect match, missed diagnosis, misdiagnosis and mixed error.

[0012] Further, the preset rule comprises: Perfect match: when the real label is consistent with the predicted label, the multi-label confusion matrix is updated as follows:

[0013] wherein m is the weight of perfect match, and the calculation formula is:

[0014] wherein n is the number of categories, is the number of real labels, is the number of predicted labels; when the real label and the predicted label are both zero, the multi-label confusion matrix is updated as:

[0015] wherein, is no real label, is no predicted label; Missed diagnosis: when the expert misses the real label, the multi-label confusion matrix update rule is:

[0016] wherein, is the probability that the i-th category is misdiagnosed as the j-th category in the transition probability matrix, and when the predicted label is all zero, the multi-label confusion matrix is updated as:

[0017] Misdiagnosis: when the expert predicts a non-existent disease, the multi-label confusion matrix is updated as:

[0018] wherein, is the probability that the k-th category is misdiagnosed as the i-th category in the transition probability matrix; Mixed error: when there is both missed diagnosis and misdiagnosis, only missed diagnosis is processed according to the missed diagnosis rule.

[0019] Further, the method further comprises: constructing a MoE-based human-machine decision fusion framework, the MoE-based human-machine decision fusion framework comprising an AI model, an expert prediction module, a Transformer encoder, a gating network and a fusion module; wherein the gating network comprises a fusion layer and a feedforward neural network layer, and the feedforward neural network layer comprises four independent feedforward neural networks.

[0020] Further, in the step of extracting visual features of the medical image and generating initial weights according to the visual features, the AI prediction result and the expert prediction result, the step comprises: extracting visual features of the medical image using a pre-trained ResNet-50 network; The AI model and the expert prediction module are used to obtain an AI prediction result and an expert prediction result respectively; The visual features, the AI prediction result and the expert prediction result are input into a concatenation layer for fusion, and then input into four independent feedforward neural networks to generate a first weight, a second weight, a third weight and a fourth weight; The first weight, the second weight, the third weight and the fourth weight are fused to obtain an initial weight.

[0021] Further, in the step of generating confusion influence factors according to the multi-label confusion matrix, the step comprises: The self-correlation matrix is calculated according to the multi-label confusion matrix, and the self-correlation matrix is input into a Transformer encoder to further capture the complex correlation between labels to obtain an output matrix; The attention weight of the multi-label confusion matrix is calculated based on the learning query vector and the key vector to determine the mode of the multi-label confusion matrix; Based on the mode of the multi-label confusion matrix, the output matrix is projected to an expert-specific decision space through linear transformation to generate a set of confusion influence factors to quantify the trust degree of each label on the decision of the human expert.

[0022] Further, in the step of fusing the initial weight with the confusion influence factors and dynamically calibrating the initial weight of the human expert and the AI model to generate a calibrated weight, the step comprises: The confusion influence factors and the initial weight are fused through Hadamard product to generate a calibrated weight.

[0023] Further, the expression of the initial weight is:

[0024] wherein, the first weight is, the second weight is, the third weight is, the first bias term is, the second bias term is; The expression of the confusion influence factor is:

[0025] wherein, E is the number of experts, the corresponding bias term is, the linear transformation is, and M is the output matrix; The expression of the calibrated weight is:

[0026] wherein, This represents element-wise multiplication with broadcast mechanism. The function is a Sigmoid function, and the confusion factor is constrained to an interval. .

[0027] Furthermore, the expression for the probability output is:

[0028] in, This represents the weight of the human expert corresponding to each label. This represents the AI ​​model weights corresponding to each label. Based on expert predictions, For AI to predict results; The expression for binary prediction is:

[0029]

[0030] in, For probability output, For a specific threshold of the label, For the threshold set, To verify the actual annotation of the set, This is for outputting label information.

[0031] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: In the embodiments of this disclosure, the multi-label medical diagnosis human-machine collaboration method based on confusion matrix and hybrid expert system, as described above, firstly proposes a novel human-machine decision fusion framework based on expert hybridity (MoE). This framework combines experienced human experts with the global modeling capabilities of AI models. Initial human-machine weights are generated by a gating network, and the initial weights are calibrated using a multi-label confusion matrix (MLCM) to achieve more accurate weight allocation, improve the accuracy of multi-label medical diagnosis, and reduce misdiagnosis and missed diagnosis. Secondly, a confusion matrix is ​​constructed to fully consider the correlation between labels to evaluate the confidence of human experts. The maximum entropy model, which does not require preconditions, is more in line with real-world scenarios, capturing the dependencies between labels and quantifying the systematic bias of expert predictions, thus achieving an accurate assessment of the confidence of human experts. Thirdly, to address the heterogeneity of multi-label classification, an independent decision threshold is optimized for each label. Hamming loss is minimized through grid search, and the probability output is converted into a binary diagnostic result, effectively reducing missed diagnosis and misdiagnosis. Attached Figure Description

[0032] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0033] Figure 1 This diagram illustrates the steps of a multi-label medical diagnostic human-computer collaboration method based on a confusion matrix and a hybrid expert system, as shown in an exemplary embodiment of this disclosure. Figure 2 This diagram illustrates a multi-label problem diagnosis process in an exemplary embodiment of this disclosure. Figure 3 This diagram illustrates a specific flowchart of a multi-label medical diagnostic human-computer collaboration method based on a confusion matrix and a hybrid expert system, as shown in an exemplary embodiment of this disclosure. Figure 4 A schematic diagram illustrating the construction process of a multi-label confusion matrix (MLCM) in an exemplary embodiment of this disclosure is shown. Figure 5 A detailed architecture diagram of the MoE human-machine decision fusion module in an exemplary embodiment of this disclosure is shown; Figure 6 This invention illustrates experimental results for various aspects of performance evaluation in exemplary embodiments of this disclosure; Figure 7 The evolution of the multi-label confusion matrix (MLCM) in an exemplary embodiment of this disclosure is illustrated. Detailed Implementation

[0034] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0035] Furthermore, the accompanying drawings are merely illustrative diagrams of embodiments of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.

[0036] This example implementation provides a multi-label medical diagnostic human-computer collaboration method based on a confusion matrix and a hybrid expert system. (Reference) Figure 1 As shown, this multi-label medical diagnostic human-computer collaboration method based on confusion matrix and hybrid expert system may include: Step S101: generating a transition probability matrix by modeling label correlation through a maximum entropy model, and taking the transition probability matrix as a basis for weighting to construct a multi-label confusion matrix; Step S102: extracting visual features of the medical image, and generating initial weights according to the visual features, AI prediction results, and expert prediction results; Step S103: generating a confusion influence factor according to the multi-label confusion matrix; Step S104: fusing the initial weights and the confusion influence factor, and dynamically calibrating the initial weights of the human experts and the AI model to generate calibrated weights; Step S105: fusing the AI prediction results, the expert prediction results, and the calibrated weights to generate a probability output; Step S106: converting the probability output into a binary prediction using label-specific thresholds to obtain a diagnosis result.

[0037] Through the above multi-label medical diagnosis human-computer collaboration method based on a confusion matrix and a mixed expert system, in a first aspect, a novel human-computer decision fusion framework based on an expert mixing (MoE) is proposed, which combines experienced human experts and the global modeling capability of an AI model, generates initial human-computer weights through a gating network, and calibrates the initial weights through a multi-label confusion matrix (MLCM) to achieve more accurate weight distribution, improve the accuracy of multi-label medical diagnosis, and reduce misdiagnosis and missed diagnosis. In a second aspect, the label correlation is fully considered to construct a confusion matrix to evaluate the confidence of human experts, and a maximum entropy model without prior assumptions is used to better meet real-world scenarios, capture the dependency between labels, quantify the systematic bias of expert predictions, and accurately evaluate the confidence of human experts. In a third aspect, in view of the heterogeneity of multi-label classification, independent decision thresholds are optimized for each label, the Hamming loss is minimized through grid search, the probability output is converted into a binary diagnosis result, and the missed diagnosis and misdiagnosis phenomena are effectively reduced.

[0038] In the following, reference will be made to Figures 1 to 7 The above multi-label medical diagnosis human-computer collaboration method based on a confusion matrix and a mixed expert system in the present example embodiment will be described in more detail.

[0039] In step S101, a transition probability matrix is generated by modeling label correlation through a maximum entropy model, and the transition probability matrix is taken as a basis for weighting to construct a multi-label confusion matrix.

[0040] Specifically, the present application focuses on the human-machine collaboration problem of multi-label disease diagnosis, the relationship between labels is complex, and the goal is to design a human-machine decision fusion scheme that not only considers the complex relationship between labels and the confidence of human experts, but also considers the heterogeneity between labels. The present application involves human experts and AI models, human experts evaluate confidence through the confusion matrix MLCM, and the ultimate goal is to minimize the Hamming loss, the missed diagnosis rate and the misdiagnosis rate of the decision result at the same time, and improve the decision performance.

[0041] In the multi-label medical diagnosis task, each instance may contain multiple diseases with strong label correlation (containing n categories). For a given instance, a human expert provides a binary prediction , while an AI model generates a probability prediction based on instance features . To evaluate the confidence of human experts, the present application constructs a multi-label confusion matrix (MLCM) that considers label correlation . Each element in the matrix represents the confidence of the human expert in predicting the instance with the true label i as label j, and the diagonal element reflects the confidence of correct prediction. In addition, a row and a column are added to the matrix to handle the case of no true label (NTL) or no predicted label (NPL).

[0042] The expert mixing (MoE) method is used to fuse the predictions of human and AI models to generate the fused probability output . Under this framework, MoE dynamically allocates different experts' weights through input features, enabling the system to utilize the complementary advantages of each expert to optimize decision-making. The fused output is converted to a binary prediction through label-specific thresholds . The main goal of this method is to minimize the Hamming loss, the missed diagnosis rate (FNR) and the misdiagnosis rate (FPR) at the same time:

[0043] where and are the weights of human experts and AI models in the fusion, respectively.

[0044] Human expert confidence evaluation The method comprises two key modules: 1) modeling label correlation by maximum entropy, and 2) constructing a probability-weighted multi-label confusion matrix (MLCM). The label correlation is modeled by estimating the probability of label prediction error, which is calculated using a maximum entropy model without prior assumptions, and the result is stored in a conversion probability matrix that captures all possible prediction errors. In the confusion matrix construction process, common prediction scenarios are divided into four categories, and the label correlation is explicitly considered by combining the previously calculated conversion probability matrix to construct a more realistic confusion matrix, thereby enabling accurate evaluation of the confidence of human experts and laying the foundation for subsequent human-machine decision fusion.

[0045] Maximum entropy model for label correlation The label correlation is defined as the probability of misclassifying class i as class j, and is quantified by the conversion probability matrix:

[0046] When , the diagonal element represents the probability of correctly predicting class i samples. Traditional methods use Bayesian methods to calculate conditional probabilities, which require the assumption of label independence (i.e. ), which has , is expressed as:

[0047] However, in the medical disease diagnosis task, there is complex correlation between labels, and assuming label independence will introduce evaluation bias. To solve this problem, a maximum entropy model is used, which selects the highest entropy distribution from all probability distributions that satisfy the known constraints, avoiding making arbitrary independence assumptions. This principle ensures that the model is as unbiased as possible given the information. For any label pair , an independent binary logistic regression model is trained to estimate , with the following steps: Feature representation construction: To avoid redundant calculations, focus on specific areas, extract all true labels and human predictions samples (corresponding to misclassification areas). For each such sample, a simple input feature vector is constructed to capture the context association and potential confusion patterns between classes. The generated feature-label pairs are used to train the maximum entropy model to learn the confusion or dependency relationships between different classes.

[0048] Prediction target definition: For each label pair , the prediction target is defined as (i.e. human predicted outcome of class j). This objective serves as the supervisory signal for training the maximum entropy model, enabling it to capture the likelihood of predicting class j when class i is misclassified.

[0049] Conditional probability estimation: model the conditional probability using a logistic regression function The estimation formula of the maximum entropy model is:

[0050] where is the sigmoid function, denotes the parameters trained for class j.

[0051] Maximum entropy model training: optimize the model using L-BFGS to learn the parameters by minimizing the L2 regularized logistic loss and bias b:

[0052] where N is the number of training samples and C is the regularization hyperparameter that controls the strength of the L2 penalty.

[0053] Maximum entropy model inference: after training, evaluate the model using a fixed input (corresponding to the condition) and store the output at the position of the conditional probability matrix. This value reflects the likelihood of misclassifying class i as class j when class i is missed.

[0054] By the above steps, we obtain a transformed probability matrix that captures the label correlation.

[0055] Constructing a probability-weighted multi-label confusion matrix (MLCM) To avoid the unrealistic assumption that the correlation between each label is equal in the original method, we introduce the previously calculated transformed probability matrix as a weighting basis into the construction process of the confusion matrix, constructing a more realistic multi-label confusion matrix (MLCM). According to existing methods of constructing multi-label confusion matrices, additional rows and columns are reserved to handle the boundary conditions of no true label (NTL) and no predicted label (NPL). Through empirical observation of human expert predictions and true labels, predictions are usually classified into four categories: (1) perfect match, (2) missed detection, (3) false detection, and (4) mixed error. When constructing the expert confusion matrix, we classify these four types of prediction patterns to avoid redundant calculations and achieve accurate assessment of expert confidence.

[0056] Class 1: Perfect match (true label and predicted label are consistent) When the true label and the predicted label are consistent (i.e. When the number of true labels is zero, the confusion matrix is updated as follows:

[0057] where m is the weight of perfect match, and the calculation formula is:

[0058] NT and NP represent the number of true labels and the number of predicted labels of the instance, respectively. In particular, when both the true label and the predicted label are zero (i.e. ), the confusion matrix is updated as:

[0059] for handling the case of no true label or no predicted label.

[0060] Category 2: Missed diagnosis (there is an un-predicted true label) When the expert misses the true label (i.e. ), the confusion matrix update rule is:

[0061] where is the probability of the i-th misjudgment being the j-th in the transition probability matrix. When the predicted label is all zero (i.e. ), the confusion matrix is updated as:

[0062] Category 3: Misdiagnosis (there is an extra false positive prediction) When the expert predicts a non-existent disease (i.e. ), the confusion matrix is updated as:

[0063] Category 4: Mixed error (both missed diagnosis and misdiagnosis exist) When both missed diagnosis and misdiagnosis exist (i.e. and ), to avoid redundant calculation of mixed errors, only the missed diagnosis case is handled according to the rules of category 2.

[0064] In step S102, the visual features of the medical image are extracted, and the initial weights are generated according to the visual features, the AI prediction result and the expert prediction result.

[0065] Specifically, the visual features of the medical image are extracted using a pre-trained ResNet-50 network; The AI prediction result and the expert prediction result are obtained using an AI model and an expert prediction module, respectively; The visual features, AI prediction results and expert prediction results are input into a concatenation layer for fusion, and then input into four independent feedforward neural networks to generate first weights, second weights, third weights and fourth weights. The first weights, second weights, third weights and fourth weights are fused to obtain initial weights.

[0066] In steps S103 to S106, a confusion influence factor is generated according to a multi-label confusion matrix; the initial weights are fused with the confusion influence factor, and the initial weights of the human experts and the AI model are dynamically calibrated to generate calibrated weights; the AI prediction results, the expert prediction results and the calibrated weights are fused to generate a probability output; and the probability output is converted into a binary prediction by using a label-specific threshold to obtain a diagnosis result.

[0067] Specifically, the MoE-based human-machine decision fusion framework A decision fusion framework based on a mixed expert model (MoE) is proposed, which adopts a modular learning architecture and can dynamically select and combine multiple expert models based on input correlation. In the scenario of the present application, "expert" refers to a human expert and an AI model. This method is suitable for multi-label classification tasks, in which different labels can benefit from different combinations of expert inputs. The framework integrates human expertise and AI model intelligence to improve diagnostic performance. The method of the present application includes two modules: 1) recalibration of human-machine weights through a confusion matrix; and 2) fusion of human-machine decisions through a gating network. Before the fusion of decisions, first, a visual feature is extracted using a pre-trained backbone deep convolutional neural network (ResNet-50) on ImageNet-1k, and then the features are processed by a Transformer encoder to capture potential information embedded in a multi-label confusion matrix (MLCM), so that the model can learn human-specific decision tendencies and uncertainties. These learned information is then adaptively weighted by the human decision through a gating network. By explicitly modeling the inherent label correlation in multi-label classification (MLC), the framework effectively integrates the complementary advantages of human expert and AI model decisions while mitigating their respective weaknesses, thereby improving overall diagnostic accuracy.

[0068] The MoE architecture proposed in the present application aggregates input features at the connection layer and inputs them to an independent feedforward neural network (FNN) corresponding to each label to generate initial weights. The MLCM iteratively extracts information through N layers of Transformer blocks, and then the refined representation is projected through a linear transformation layer to recalibrate the initial weights, especially to adjust the contribution weight of the human expert. Finally, the fusion weights of the human expert and the AI model are generated.

[0069] 1. Calibrate the initial weights of the human expert To effectively capture the complex information in the confusion matrix C, we first compute the autocorrelation matrix to enrich the relationship information embedded in the confusion matrix and achieve a more comprehensive representation of confusion-related features. This augmented matrix is then used as input to a two-layer multi-head self-attention Transformer encoder to further capture complex relationships between labels. The entire encoding process aims to capture subtle label correlations and expert-specific biases, providing support for the subsequent decision fusion stage. The formal expression is as follows:

[0070] where M is the output matrix, each row is a d-dimensional embedding representing the augmented features of the label, and is used for subsequent human-machine fusion. In the Transformer encoder, attention mechanisms are applied to each label to quantify misclassification relationships. Specifically, for each label i, the model computes its attention score for other labels j based on learned query and key vectors to assess the degree of attention that label i pays to j, thereby capturing confusion patterns:

[0071] where and are the query and key vectors for labels i and j, respectively. The attention weight softly measures the confusion likelihood between labels i and j, capturing expert-specific misclassification patterns and potential label correlations. The learned attention distribution is then used to compute context-aware label embeddings, resulting in the final output M.

[0072] In the human-machine fusion process, to more effectively determine the weight distribution of human experts, we recalibrate expert-specific confidence based on the information extracted by the Transformer encoder. Specifically, the output features are projected to an expert-specific decision space through a linear transformation to generate a set of confusion influence factors that quantify the degree of trust that each label has in the decision-making of human experts:

[0073] where , E is the number of experts, is the corresponding bias term.

[0074] The generated confusion influence factors are input into a gating network to dynamically recalibrate the initial weights of human experts, enhancing the influence of experts on high-confidence labels while reducing the influence on error-prone labels. This fine-grained, label-specific weight adjustment improves the reliability of human-machine fusion and ultimately enhances the overall prediction performance.

[0075] 2. Gated network for human-machine decision fusion In multi-label classification (MLC) problems, both human experts and AI models are heterogeneous, and the difference in their feature focus leads to fluctuations in the prediction accuracy of different cases. To alleviate the performance fluctuations caused by this heterogeneity, a mixture-of-experts (MoE) model for human-machine decision fusion is introduced. The HQS method improves MLC performance by combining task-specific experts and shared experts. The core of the MoE architecture is the gating network, which plays a key role in adaptively selecting experts and assigning weights. In the method of the present application, the gating network dynamically assigns weights based on input features and human-machine decision, realizing personalized and context-aware fusion. Specifically, an independent four-layer neural gating network is designed for each disease, and a feature compression and confusion matrix recalibration mechanism is integrated to generate expert weights.

[0076] Initial weight generation: Since the prediction ability of human experts and AI models differs on different labels, a dedicated gating network is used for each label category. A four-layer fully connected network maps d-dimensional image features to the expert weight space, and uses Activation functions enhance gradient flow. The initial weight is calculated as follows:

[0077] where is the weight of the i-th fully connected layer.

[0078] During training, layer normalization (LN) is applied after the second linear transformation, and Dropout is applied after the first nonlinear transformation, to improve model generalization and training stability.

[0079] Confusion matrix calibration of initial weights: the confusion influence factor is fused with the dynamically generated initial weight through Hadamard product to dynamically adjust the calibrated human expert weight:

[0080] where represents element-wise multiplication with a broadcast mechanism, is a Sigmoid function that constrains the confusion factor to the interval .

[0081] Human-machine fusion mechanism: fusion prediction is obtained by combining the expert output and with the learned weights to get the final probability fusion result :

[0082] where is the first column, is the second column. Specifically, denotes the human expert weight corresponding to each label, denotes the AI model weight corresponding to each label. Specific threshold for each class of label is the probability output

[0083] is converted to binary prediction, a label-specific threshold strategy is adopted. Due to label imbalance and difficulty difference, applying a uniform threshold (such as 0.5) to all labels may lead to poor performance, so an independent threshold is optimized for each label i. This method can customize more fine-grained decision boundaries for each class. Formally, the threshold set is represented as where n is the total number of labels. These thresholds are optimized by grid search on a reserved validation set to minimize the Hamming loss commonly used in multi-label tasks:

[0084] The final binary prediction for label i is as follows:

[0085] This method ensures that the binary decision boundary for each label is calibrated by experience, improving the overall performance of the multi-label task.

[0086] In the process of human-machine decision fusion, it is crucial to jointly consider image features and all human-machine decisions to assign weights. In addition, due to the heterogeneity of multi-label classification, it is necessary to design label-specific gating networks and thresholds.

[0087] In one specific embodiment, the present application proposes a new human-machine decision fusion method based on confusion matrix and expert mixing (MoE) model. For a given instance, human experts and AI models provide their respective predictions. To evaluate the reliability of human expert predictions, a method is introduced to evaluate expert confidence by constructing a multi-label confusion matrix (MLCM), which serves as a proxy for expert confidence. Given the prediction results of human experts and AI models, a decision fusion method based on MoE gating network is proposed. This method assigns different weights to the predictions of human experts and AI models, and generates the final fusion output by calculating the weighted combination. In addition, label-specific thresholds are used to convert probability output to binary (0 / 1) classification decisions.

[0088] ​​a) Assessing human expert confidence. To assess the confidence of human experts, the present method proposes a method to construct a multi-label confusion matrix (MLCM) that can capture label correlations. Specifically, label correlation is defined as the probability of label i being misclassified as label j, and a maximum entropy model without prior assumption is used to model this relationship to derive a transition probability matrix. Subsequently, a probability-weighted multi-label confusion matrix (MLCM) construction method is introduced that integrates this transition probability matrix to quantify expert confidence.

[0089] b) MoE-based human-AI decision fusion. To fuse the predictions of human experts and AI models, a dedicated gating network is designed for each label that generates initial weights based on prediction results and image features. To further optimize these weights, a Transformer encoder equipped with multi-head self-attention mechanism is adopted to extract information from the MLCM. The extracted representation is linearly transformed and input into the gating network to recalibrate the weights of human experts. Subsequently, a weighted combination is calculated based on the recalibrated weights to generate the probability output.

[0090] In one specific embodiment, as shown in FIG. 1, a schematic diagram of the diagnosis process for multi-label problems is shown. For a patient’s CT image, both a human doctor expert and an AI model are given to make disease diagnosis. The human doctor expert gives a 0 / 1 label to represent whether the disease exists or not, and the AI model outputs the corresponding probability for each disease. The decision results of the human doctor expert and the AI model are fused to obtain the final CT image diagnosis result. A CT image can have multiple diseases at the same time. Figure 2

[0091] In one specific embodiment, as shown in FIG. 2, a schematic diagram of the diagnosis process for multi-label problems is shown. For a patient’s CT image, both a human doctor expert and an AI model are given to make disease diagnosis. The human doctor expert gives a 0 / 1 label to represent whether the disease exists or not, and the AI model outputs the corresponding probability for each disease. The decision results of the human doctor expert and the AI model are fused to obtain the final CT image diagnosis result. A CT image can have multiple diseases at the same time. Figure 3 ​As shown, it is a whole framework of multi-label medical diagnosis human-computer collaboration method, including evaluation of human expert confidence and MoE-based decision fusion. Among them, the evaluation of human expert confidence module combines the correlation between labels, models the label correlation by estimating the conditional probability of label prediction error, calculates using the maximum entropy model without prior assumptions according to the historical decision results of human experts, and stores the results in the conversion probability matrix. In the process of constructing the confusion matrix, common prediction situations are divided into 4 categories, combined with the previously calculated conversion probability matrix, the label correlation is explicitly considered, and a more realistic confusion matrix is constructed, so that the confidence of human experts can be accurately evaluated, laying a foundation for subsequent human-computer decision fusion. The MoE-based decision fusion module considers the heterogeneity of labels, designs an independent gating network for each label category, inputs image features and human expert and AI model decision results, and generates different initial weights through the gating network. Process the multi-label confusion matrix (MLCM) through the Transformer encoder to obtain the latent information and adjust the initial weight in the gating network part to generate the final weight. According to the weight, the human-computer classification results are fused, and according to the threshold of each label obtained in the training process, the probability value obtained by fusion is converted into 0 / 1 label.

[0092] In a specific embodiment, as shown in Figure 4 As shown, it is a process diagram for constructing a multi-label confusion matrix (MLCM). According to the misdiagnosis and missed diagnosis, the confusion matrix is constructed in 4 categories. It is divided into 4 categories: perfect match (true label and predicted label are consistent), missed diagnosis (there are true labels that are not predicted), misdiagnosis (there are additional false positive predictions), and mixed errors (both missed diagnosis and misdiagnosis).

[0093] In a specific embodiment, as shown in Figure 5 As shown, it is a detailed architecture diagram of the MoE human-computer decision fusion module. In the splicing layer, the image features, human expert decision results and AI model decision results are aggregated and then input into the independent feedforward neural network (FNN) corresponding to each label to generate the initial weight. The multi-label classification module (MLCM) iteratively extracts information through N layers of Transformer blocks, and then projects the optimized representation through a linear transformation layer to recalibrate the initial weight, especially to adjust the contribution of human experts. Finally, the final fusion weight of human experts and AI model is generated.

[0094] In a specific embodiment, the algorithm for constructing the feature representation of the maximum entropy model is as shown in Table 1: Table 1

[0095] First, focus on the area of classification error, that is, the true label and human expert prediction is For each such sample, a simple input feature vector is constructed in order to capture the context-dependent correlation and potential confusion patterns among classes. The resulting feature-label pairs will be used to train a maximum entropy model, which is responsible for learning the confusion or dependency relationships among different classes.

[0096] In one specific embodiment, the algorithm of the multi-label medical diagnosis human-machine collaboration method is shown in Table 2: Table 2

[0097] The human-machine collaboration method is divided into three stages: evaluation of human expert confidence, fusion training based on MoE model, and inference stage. The input is the real label situation, the prediction results of human experts and AI model and the image, and the final output is the prediction label for the image. In the evaluation of human expert confidence stage, the mutual influence between disease classes is modeled by calculating the conversion probability matrix, which is added to the construction process of the confusion matrix, and finally a multi-label confusion matrix considering the correlation between labels is obtained. In the fusion training stage with the MoE model, the image features are extracted, the initial human-machine allocation weight is obtained by using the gating network, the potential relationship in the confusion matrix is obtained by using the Transformer encoder, and the initial weight is calibrated, and then the final probability fusion result is obtained. In the inference stage, the test image and the known human expert multi-label confusion matrix are input to obtain the final binary classification output result.

[0098] In one specific embodiment, as shown in Table 3, the performance advantage in four indicators of Hamming loss, AUC value, MAP value and F1 score. On each dataset, the application framework (MoE+MLCM) is compared with human experts, AI models (ResNet18) and existing human-machine collaboration methods (CHM, HAIT, JSF) for evaluation, and ablation experiments are conducted to prove the effectiveness of each module. On the ChestX-ray dataset, compared with human experts, the framework reduces the Hamming loss by 13.7%; compared with AI models (ResNet18), it reduces by 26.4%; at the same time, its performance exceeds existing fusion methods (CHM, HEIT, JSF), with an improvement of 5.6% to 46.1%. Similar trends are observed on the S12L-ECG dataset: compared with human experts, the Hamming loss is reduced by 77.6%; compared with AI models, it is reduced by 83.9%; compared with the most advanced fusion method, the improvement is 44.4% to 77.6%. Notably, from the comparison of F1 score, the method of the present application reduces the missed diagnosis and misdiagnosis rate by at least 7.9% on the ChestX-ray dataset and at least 2.8% on the S12L-ECG dataset, which highlights its ability to reduce diagnostic errors through effective human-machine collaboration. In addition, the variance analysis of the experimental results shows that the method of the present application has significantly enhanced robustness compared with existing solutions. These findings verify the effectiveness of the proposed framework in reducing diagnostic inaccuracies and its strong generalization ability on datasets of different scales and clinical scenarios. In the ablation experiment section, the influence of multi-label confusion matrix (MLCM), gating network and label-specific threshold on model performance is evaluated respectively. On the ChestX-ray dataset, the performance decreases significantly after removing MLCM: the Hamming loss increases by 14.1% to 0.0707, and the F1 score decreases by 7.4% to 0.6641; on the S12L-ECG dataset, removing MLCM causes the Hamming loss to increase by 547% to 0.0097. This confirms the core role of MLCM in modeling expert bias and label correlation. When replacing the specially designed gating network with a traditional stacking method, the F1 score of the ChestX-ray dataset decreases by 23.9% to 0.5460, and the F1 score of the S12L-ECG dataset directly decreases to 0.0000 (completely fails), which fully demonstrates the necessity of the gating network proposed in the present application.The most significant performance degradation occurs when replacing label-specific thresholds with a fixed value of 0.5: the Hamming loss of the ChestX-ray dataset increases by 1248% to 0.8231, and the F1 score of the S12L-ECG dataset decreases by 67.0% to 0.2667, which also leads to an 82.3% increase in the rate of clinically critical false negatives. The above experiments verify the unique contribution of each component to the robust performance of the framework: MLCM is used to correct the bias of human experts, the gating network dynamically generates and adjusts human-machine weights, and label-specific thresholds guarantee the reliability of clinical predictions.

[0099] Table 3

[0100] In a specific embodiment, as shown in Figure 6 , the experimental results of the performance evaluation of the present application are shown. Figure 6 In Fig. (a), the multi-index performance radar chart of the ChestX-ray dataset is intuitively displayed, all indexes are uniformly standardized, and the Hamming loss is inverted, so the area of the polygon in the figure is positively correlated with the performance. The polygon formed by the method of the present application has the largest area and the most rounded shape, indicating that it has achieved balanced and excellent performance in multiple indexes, which reflects the comprehensive advantage of the method of the present application. Figure 6 In Figs. (b) and (c), the reliability of the multi-label classification module (MLCM) in reflecting the confidence of human experts is evaluated, and the diagonal elements thereof are compared with the recall rate of each label. The calculation formula of the recall rate is equivalent to the normalized result of the diagonal elements of the MLCM. Therefore, by comparing the two, a standardized method can be provided for evaluating the accuracy of the MLCM. For all labels, the difference between the two indicators is less than 0.025, and they show a strong linear correlation (Pearson correlation coefficient r = 0.984). The above results show that the MLCM proposed in the present application can accurately reflect the confidence of human experts. Figure 6 In Fig. (d), the scalability of the present application is verified. Within a certain range, as the number of AI models increases, the overall performance improves. Specifically, compared with combination 1, after adding VGG19, the F1 score is increased by about 2%, and the Hamming loss is reduced by 4%; further adding AlexNet, compared with the baseline, the F1 score is additionally increased by 2.8%, and the Hamming loss is overall reduced by 6%. This result verifies that the proposed method has good scalability.

[0101] In a specific embodiment, as shown in Figure 7 , the evolution process of the multi-label classification module (MLCM) in the present application is shown. Figure 7Fig. 3(a) shows the MLCM obtained based on the ChestX-ray dataset. Figure 7 Fig. 3(b) is a self-correlation matrix $C^TC$, which integrates the relationship embedding into the confusion matrix, thus more comprehensively representing the features related to confusion. Figure 7 Fig. 3(c) presents the semantic correlation matrix generated by the Transformer encoder, which captures richer and asymmetric label relationships to better guide the weight distribution in the gating network. Figure 7 Fig. 3(d) is the attention weight normalized by SoftMax, which optimizes the specialist weight. This operation "softens" the output of the specialist, especially for the easily confused classes, to improve the robustness of the model. From Figure 7 The transformation process from Fig. 3(a) to Fig. 3(d) embodies the explainability of this method.

[0102] Through the above multi-label medical diagnosis human-computer collaboration method based on confusion matrix and mixed expert system, in the first aspect, a novel human-computer decision fusion framework based on expert mixing (MoE) is proposed, which combines the experience of human experts with the global modeling ability of AI models, generates initial human-computer weights by a gating network, and calibrates the initial weights through a multi-label confusion matrix (MLCM) to achieve more accurate weight distribution, improve the accuracy of multi-label medical diagnosis, and reduce misdiagnosis and missed diagnosis. In the second aspect, the label correlation is fully considered to construct the confusion matrix to evaluate the confidence of human experts, and the maximum entropy model without prior assumptions is used to better meet the real scene, capture the dependency between labels, quantify the systematic bias of expert prediction, and accurately evaluate the confidence of human experts. In the third aspect, in view of the heterogeneity of multi-label classification, the independent decision threshold is optimized for each label, the Hamming loss is minimized through grid search, the probability output is converted into binary diagnostic results, and the missed diagnosis and misdiagnosis phenomenon is effectively reduced.

[0103] In addition, the terms "first", "second" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, the meaning of "multiple" is two or more, unless otherwise explicitly specified.

[0104] In the description of the disclosure, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the disclosure. In the description of the disclosure, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in the specification.

[0105] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses or adaptations of the disclosure that follow the general principles thereof and include the general principles thereof and include other known or customary features not specifically mentioned herein. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the disclosure are indicated by the appended claims.

Claims

1. A multi-label medical diagnosis human-machine collaboration method based on confusion matrix and mixed expert system, characterized in that, The method comprises: A transition probability matrix is generated by modeling label correlation through a maximum entropy model, and the transition probability matrix is used as a basis for weighting to construct a multi-label confusion matrix; Visual features of the medical image are extracted, and initial weights are generated according to the visual features, the AI prediction result and the expert prediction result; A confusion influence factor is generated according to the multi-label confusion matrix; The initial weights are fused with the confusion influence factor, and the initial weights of the human experts and the AI model are dynamically calibrated to generate calibrated weights; The AI prediction result, the expert prediction result and the calibrated weights are fused to generate a probability output; The probability output is converted into a binary prediction by using a label-specific threshold to obtain a diagnosis result.

2. The multi-label medical diagnosis man-machine collaboration method based on confusion matrix and mixed expert system according to claim 1, characterized in that, In the step of generating a transition probability matrix by modeling label correlation through a maximum entropy model and using the transition probability matrix as a basis for weighting to construct a multi-label confusion matrix, the step comprises: The label correlation is defined as the probability of label i being misclassified as label j, and the transition probability matrix is generated by estimating the conditional probability through a maximum entropy model; The transition probability matrix is used as a basis for weighting to construct a multi-label confusion matrix, and the multi-label confusion matrix is updated according to a preset rule; wherein the modes of the multi-label confusion matrix include perfect match, missed diagnosis, misdiagnosis and mixed error.

3. The multi-label medical diagnosis man-machine collaboration method based on confusion matrix and mixed expert system according to claim 2, characterized in that, The preset rule comprises: Perfect match: when the real label is consistent with the predicted label, the multi-label confusion matrix is updated as follows: wherein m is the weight of the perfect match, and the calculation formula is: wherein n is the number of classes, is the number of true labels, is the number of predicted labels; when both the true label and the predicted label are zero, the multi-label confusion matrix is updated as: wherein, is no real label, is no predicted label; Missed diagnosis: when the expert misses the real label, the multi-label confusion matrix updating rule is: wherein, is the probability of the i-th type of misjudgment in the transition probability matrix as the j-th type, and when the predicted label is all zero, the multi-label confusion matrix is updated as: Misdiagnosis: when the expert predicts a non-existent disease, the multi-label confusion matrix is updated as: wherein, is the probability of a kth type of misclassification being an ith type in the transition probability matrix; Mixed error: when there are missed diagnosis and misdiagnosis at the same time, only the missed diagnosis is processed according to the rule of missed diagnosis.

4. The multi-label medical diagnosis man-machine collaboration method based on confusion matrix and mixed expert system according to claim 3, characterized in that, The method further comprises: A human-machine decision fusion framework based on MoE is constructed, and the human-machine decision fusion framework based on MoE comprises an AI model, an expert prediction module, a Transformer encoder, a gating network and a fusion module; wherein the gating network comprises a fusion layer and a feedforward neural network layer, and the feedforward neural network layer comprises four independent feedforward neural networks.

5. The multi-label medical diagnosis man-machine collaboration method based on confusion matrix and mixed expert system according to claim 4, characterized in that, In the step of extracting visual features of the medical image and generating initial weights according to the visual features, the AI prediction result and the expert prediction result, the step comprises: Visual features of the medical image are extracted by using a pre-trained ResNet-50 network; The AI prediction result and the expert prediction result are obtained by using the AI model and the expert prediction module, respectively; After the visual features, the AI prediction result and the expert prediction result are input into a concatenation layer for fusion, the first weight, the second weight, the third weight and the fourth weight are generated by inputting them into four independent feedforward neural networks, respectively; The first weight, the second weight, the third weight and the fourth weight are fused to obtain the initial weight.

6. The multi-label medical diagnosis man-machine cooperation method based on the confusion matrix and the mixed expert system according to claim 5, characterized in that, In the step of generating a confusion influence factor according to the multi-label confusion matrix, the step comprises: A self-correlation matrix is calculated according to the multi-label confusion matrix, and the self-correlation matrix is input into a Transformer encoder to further capture the complex correlation between labels to obtain an output matrix; The learning-based query vector and the key vector calculate attention weights of the multi-label confusion matrix to determine a pattern of the multi-label confusion matrix; Based on the pattern of the multi-label confusion matrix, the output matrix is projected to an expert-specific decision space through a linear transformation to generate a set of confusion impact factors to quantify the degree of trust of each label on the human expert decision.

7. The multi-label medical diagnosis human-machine collaboration method based on confusion matrix and hybrid expert system according to claim 6, characterized in that, In the step of fusing the initial weight with the confusion impact factor and dynamically calibrating the initial weight of the human expert and the AI model to generate the calibrated weight, the step includes: The confusion impact factor and the initial weight are fused through the Hadamard product to generate the calibrated weight.

8. The multi-label medical diagnosis man-machine cooperation method based on the confusion matrix and the mixed expert system according to claim 7, characterized in that, The expression of the initial weight is: wherein, is a first weight, is a second weight, is a third weight, is a first bias term, is a second bias term; The expression of the confusion impact factor is: wherein, E is the number of experts, is a corresponding bias term, is a linear transformation, M is an output matrix; The expression of the calibrated weight is: wherein denotes an element-wise multiplication with a broadcast mechanism, is a Sigmoid function and constrains the confusion factor to the interval .

9. The multi-label medical diagnosis human-machine collaboration method based on a confusion matrix and a hybrid expert system according to claim 8, characterized in that, The expression of the probability output is: wherein, represents the human expert weight corresponding to each label, represents the AI model weight corresponding to each label, is the expert prediction result, is the AI prediction result; The expression of the binary prediction is: wherein, is a probability output, is a label-specific threshold, is a set of thresholds, is a validation set true label situation, is an output label situation.