Entity-level credibility evaluation method, device and system and storage medium
By using dynamic inconsistency threshold calibration and information entropy quantification, combined with visual heatmaps, the static limitations of medical entity-level uncertainty assessment are addressed, enhancing the interpretability and robustness of medical diagnostic models and supporting clinical decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JILIN UNIVERSITY
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-17
AI Technical Summary
Existing medical diagnostic models lack dynamic calibration in measuring uncertainty at the medical entity level, which prevents the models from accurately identifying high-risk descriptions within a diagnosis and limits the interpretability and credibility of clinical decisions.
By generating a dynamic inconsistency threshold through adversarial perturbation and Bayesian variance weighting, and combining information entropy to quantify the credibility index of medical entity descriptions, and by integrating visualization heatmaps with clinical diagnosis, entity-level credibility assessment is provided.
It significantly improves the robustness of the model under data distribution shifts and noise interference, provides interpretability support, and assists in clinical decision-making to validate high-risk areas.
Smart Images

Figure CN121884047A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, specifically relating to an entity-level credibility assessment method, apparatus, system, and storage medium. Background Technology
[0002] In recent years, with the rapid evolution of deep learning technology and its deep integration with medical artificial intelligence applications, AI-based medical diagnostic technology has become a research hotspot for assisting clinical decision-making. However, existing medical diagnostic models suffer from insufficient interpretability, making it difficult for clinicians to fully understand and trust their decision logic and results, thus limiting the practical application value of AI diagnostic results. Against this backdrop, medical entity-level credibility assessment has become a core element in improving the interpretability of AI. Existing methods face three major technical challenges in achieving this goal: First, the semantic complexity of medical symptoms requires model outputs to possess both accuracy and fine-grained interpretability; second, existing methods typically rely on probabilistic estimations to quantify prediction confidence, but such static estimations cannot dynamically reflect the model's uncertainty regarding the described medical entity, especially when the input data has distributional biases or noise interference, making confidence assessment prone to bias; third, existing methods rely on empirically set thresholds, lacking targeted calibration for medical entity-level prediction bias, which may lead to the omission of high-uncertainty areas in clinical decision-making, increasing the risk of misdiagnosis.
[0003] Conformal prediction methods generate reliable prediction sets through statistical calibration, significantly improving the robustness of models under data perturbations and ensuring the reliability of prediction coverage within a pre-set confidence level. However, existing conformal prediction methods mostly focus on the static assessment of the overall confidence of AI-generated content, making it difficult to provide strong support for clinical decision-making. A key challenge lies in the lack of targeted calibration for fine-grained medical entity uncertainty measurement, resulting in the model's inability to accurately locate high-risk descriptions within a diagnosis, thus limiting its practicality for clinical decision validation. Therefore, there is an urgent need for a method that integrates dynamic calibration, adversarial robustness enhancement, and visual interaction to accurately measure medical entity-level uncertainty and provide interpretable support for clinical validation. Summary of the Invention
[0004] To address the problems existing in the prior art, this invention provides an entity-level credibility assessment method, apparatus, system, and storage medium. It solves the problem that existing AI-generated medical diagnostic methods rely on static thresholds for measuring medical entity-level uncertainty and lack dynamic calibration mechanisms. By generating a dynamic inconsistency threshold through adversarial perturbation and Bayesian variance weighting, and combining information entropy to quantify the credibility index of medical entity descriptions, and integrating with clinical diagnosis using visual heatmaps, the robustness of the model under data distribution shifts and noise interference is significantly improved, providing interpretable support for clinical decision-making.
[0005] To achieve the above objectives, the present invention provides the following solution: An entity-level credibility assessment method includes: Step 101: Divide the medical image dataset into a training set, a calibration set, and a test set; Step 102: Fine-tune the medical diagnostic model based on the training set, embed a Bayesian adapter at the end of its decoding layer, and output the probability estimate of the medical entity and its Bayesian variance. Step 103: Based on the Bayesian adapter-based medical diagnostic model, calculate the inconsistency score of each medical entity in the calibration set and generate a dynamic inconsistency threshold. Step 104: Based on the dynamic inconsistency threshold, generate multiple candidate predicted diagnoses for the test images, extract the medical entity-level prediction set, and filter the candidate results; Step 105: Calculate the medical entity-level information entropy based on the candidate results and generate a visual heatmap; Step 106: Extract the description with the highest predicted probability from the candidate results to form a complete medical diagnosis. Integrate the heat map with the medical diagnosis to assist in clinical decision verification.
[0006] Preferably, in step 102, a pre-trained medical diagnostic model HuatuoGPT-Vision is used as the initial architecture. This model consists of a visual encoder from LLaVA-1.5 and a text decoder from LLaMA-3-8B. A Bayesian adapter is embedded in the last layer of the model decoder, and the final output includes the probability estimate of medical entities and their Bayesian variance.
[0007] Preferably, in step 103, a slight adversarial perturbation is applied to the medical image, and the inconsistency score of each medical entity in the calibration set is calculated based on the magnitude of the change in the predicted probability, in conjunction with Bayesian variance.
[0008] The present invention also provides an entity-level credibility assessment device, comprising: The first processing module is used to divide the medical image dataset into a training set, a calibration set, and a test set. The second processing module is used to fine-tune the medical diagnostic model based on the training set, embed a Bayesian adapter at the end of its decoding layer, and output the probability estimate of the medical entity and its Bayesian variance. The third processing module is used to calculate the inconsistency score of each medical entity in the calibration set based on the Bayesian adapter-based medical diagnostic model and generate a dynamic inconsistency threshold. The fourth processing module is used to generate multiple candidate predictive diagnoses for the test images based on the dynamic inconsistency threshold, extract the medical entity-level prediction set, and filter the candidate results. The fifth processing module is used to calculate the medical entity-level information entropy based on the candidate results and generate a visual heatmap. The sixth processing module is used to extract the description with the highest prediction probability from the candidate results to form a complete medical diagnosis, integrating the heat map with the medical diagnosis to assist in clinical decision verification.
[0009] Preferably, the second processing module uses the pre-trained medical diagnostic model HuatuoGPT-Vision as the initial architecture. This model consists of a visual encoder from LLaVA-1.5 and a text decoder from LLaMA-3-8B. A Bayesian adapter is embedded in the last layer of the model decoder, and the final output includes the probability estimate of medical entities and their Bayesian variance.
[0010] Preferably, the third processing module applies a slight adversarial perturbation to the medical images and, based on the magnitude of the change in the predicted probability, calculates the inconsistency score for each medical entity in the calibration set using Bayesian variance.
[0011] The present invention also provides an entity-level credibility assessment system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program performs an entity-level credibility assessment method when executed by the processor.
[0012] The present invention also provides a storage medium storing a computer program, which executes an entity-level credibility assessment method when running.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention utilizes dynamic inconsistency threshold calibration technology to generate statistically significant confidence boundaries based on Bayesian variance and perturbation response analysis, overcoming the static limitations of existing methods that rely on empirical assumptions for fixed thresholds. It combines information entropy hierarchical quantification of the predictive uncertainty of medical entities and uses graded color markings in a warning heatmap to intuitively locate high-risk description areas. This invention integrates dynamic threshold generation, uncertainty quantification, and a visual interaction mechanism to provide verifiable entity-level credibility assessment for AI-driven medical diagnosis, assisting doctors in quickly focusing on key abnormal areas and optimizing the clinical validation process. Attached Figure Description
[0014] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1This is a flowchart of the entity-level credibility assessment method according to an embodiment of the present invention; Figure 2 This is a schematic diagram showing the integrated display of diagnostic results and heat map in an embodiment of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] Example 1 like Figure 1 As shown, the present invention provides an entity-level credibility assessment method, comprising: Step 101: Divide the medical image dataset into a training set, a calibration set, and a test set; Step 102: Fine-tune the medical diagnostic model based on the training set, embed a Bayesian adapter at the end of its decoding layer, and output the probability estimate of the medical entity and its Bayesian variance. Step 103: Calculate the inconsistency score for each medical entity based on the calibration set, and generate a dynamic inconsistency threshold; Step 104: Generate multiple candidate predicted diagnoses for the test images, extract the medical entity-level prediction set, and filter the candidate results; Step 105: Calculate the medical entity-level information entropy based on the candidate results and generate a visual heatmap; Step 106: Extract the description with the highest predicted probability from the candidate results to form a complete medical diagnosis. Integrate the heat map with the medical diagnosis to assist in clinical decision verification.
[0019] In one embodiment of the present invention, step 101 involves dividing the existing medical image dataset into a training set, a calibration set, and a test set, and then performing data preprocessing on them. The calibration set has no data overlap with the model training set, and the image source institutions and equipment models are independent; the test set contains image-diagnosis pairs that were not involved in training and calibration, used to verify the model's generalization performance; data preprocessing includes standardization, random rotation, and random cropping to enhance the model's robustness to input data.
[0020] Step 102: The structure of the designed medical diagnostic model is as follows: The pre-trained medical diagnostic model HuatuoGPT-Vision is used as the initial architecture. This model consists of a visual encoder from LLaVA-1.5 and a text decoder from LLaMA-3-8B. A Bayesian adapter is embedded in the last layer of the model decoder, and the final output is a probability estimate containing medical entities (such as "pulmonary nodules" and "cardiac hypertrophy"). and its Bayesian variance During fine-tuning, model parameters are frozen, and variational inference is used to optimize the Bayesian adapter. The loss function is... The KL divergence term, which combines the cross-entropy and weight distribution, is: (1) in For the learnable weight parameters of the Bayesian adapter, For variational posterior distribution, For weights In variational posterior distribution The expectations below For predicting labels, For input data, For standard Gaussian priors, This is the balance coefficient.
[0021] Step 103: Apply a slight adversarial perturbation to the input medical image, and calculate the variance of each medical entity in the calibration set based on the magnitude of the change in the predicted probability. Inconsistency score: (2) in The model represents medical entities Corresponding real tags The predicted probability, For the added perturbation, For the permissible disturbance range, The variance and inconsistency score of the Bayesian adapter output in step two. This value is used to quantify the deviation between the model's prediction and the actual value for a single medical entity; a higher value indicates greater uncertainty for that medical entity. It is based on a pre-set reliability level. Through the quantile function Determine the dynamic threshold : (3) in This indicates the total number of medical entities in the calibration center, and introduces... Indicates based on preset confidence level The minimum sample size determined by the conformal prediction coverage guarantee theory must satisfy the following conditions. To ensure the statistical significance of the dynamic threshold generation; The calculation process represents selecting the calibration set fractions sorted from largest to smallest. One value is used as the threshold.
[0022] Step 104: Screening Traditional Chinese Medicine Entities in Test Images candidate prediction set : (4) in Indicates for medical entities The One acceptable description; for the candidate set Elements according to predicted probability Before selecting descending order One result; for candidate medical entities The description includes a guiding task prefix "Generate Diagnosis:", which is then encoded into a token sequence using the T5Tokenizer tokenizer as model input. A pre-trained Clinical-T5 model is used to extract deep semantic features from the entity description, and Clinical-T5 is used to extract elements within the candidate set. Text features Calculate the semantic similarity between all features. : (5) in ; Calculate dynamic semantic similarity threshold based on semantic similarity set : (6) in The average value of the elements in the semantic similarity set. Standard deviation, To preset a sensitivity coefficient, semantic similarity is used. The candidates are merged, and the result with the highest probability is retained; for a certain medical entity, if The medical entity is marked as high uncertainty, and a manual review process is triggered.
[0023] Step 105: For each of the filtered medical entity candidate sets Calculate the information entropy of its probability distribution. : (7) in, Higher values indicate greater uncertainty in medical entities. MedCLIP-SAMv2 is used to segment regions corresponding to medical images and medical entity text. Warning heatmaps with different colors and transparency are overlaid on these segmented regions to visually demonstrate the degree of uncertainty of the medical entities. A three-level warning heatmap is defined based on the entropy value range, and the entropy value is mapped to color. Piecewise functions: (8) Among them, green Red indicates a low-risk area. Indicating high-risk areas, the yellow gradient channel value changes from... Calculations were performed to achieve an exponential decay transition from yellow to red; transparency It is negatively correlated with entropy, and the calculation formula is: (9) in, Based on the basic transparency parameter, The attenuation coefficient is... The number of candidates retained for step 104 is calculated to ensure that high-entropy medical entities are more significant. Step 106: From each of the filtered candidate medical entities Extract the description with the highest prediction probability. This ultimately forms a complete medical diagnosis; a credibility index is displayed next to the medical entity description in the diagnosis. The calculation formula is: (10) in, To calibrate the maximum variance of centralized medical entities.
[0024] Complete diagnosis and thermodynamics Figure 1 Same as the display, the rendering is as follows Figure 2 As shown, this provides more comprehensive information support for clinical decision-making.
[0025] In summary, this invention effectively overcomes the static limitations of existing methods that rely on empirical assumptions for fixed thresholds by employing dynamic inconsistency threshold calibration technology, accurately capturing the confidence boundaries of medical entity descriptions. It introduces a Bayesian adapter to model the uncertainty of weight distribution, distinguishing between cognitive uncertainty and data noise. By generating dynamic thresholds through adversarial perturbation and Bayesian variance weighting, it enhances the model's robustness to equipment noise and image artifacts, ensuring the statistical significance of threshold calculations. Furthermore, by combining information entropy and heatmap visualization, it intuitively quantifies medical entity-level uncertainty and uses warning colors to mark high-risk areas, providing interpretability support for clinical decision-making.
[0026] Example 2 The present invention also provides an entity-level credibility assessment device, comprising: The first processing module is used to divide the medical image dataset into a training set, a calibration set, and a test set. The second processing module is used to fine-tune the medical diagnostic model based on the training set, embed a Bayesian adapter at the end of its decoding layer, and output the probability estimate of the medical entity and its Bayesian variance. The third processing module is used to calculate the inconsistency score of each medical entity in the calibration set based on the Bayesian adapter-based medical diagnostic model and generate a dynamic inconsistency threshold. The fourth processing module is used to generate multiple candidate predictive diagnoses for the test images based on the dynamic inconsistency threshold, extract the medical entity-level prediction set, and filter the candidate results. The fifth processing module is used to calculate the medical entity-level information entropy based on the candidate results and generate a visual heatmap. The sixth processing module is used to extract the description with the highest prediction probability from the candidate results to form a complete medical diagnosis, integrating the heat map with the medical diagnosis to assist in clinical decision verification.
[0027] As one embodiment of the present invention, the second processing module uses the pre-trained medical diagnostic model HuatuoGPT-Vision as the initial architecture. This model consists of a visual encoder from LLaVA-1.5 and a text decoder from LLaMA-3-8B. A Bayesian adapter is embedded in the last layer of the model decoder, and the final output includes the probability estimate of medical entities and their Bayesian variance.
[0028] As one embodiment of the present invention, the third processing module applies a slight adversarial perturbation to the medical image, and calculates the inconsistency score of each medical entity in the calibration set based on the change in the predicted probability and in conjunction with Bayesian variance.
[0029] Example 3 The present invention also provides an entity-level credibility assessment system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program performs an entity-level credibility assessment method when executed by the processor.
[0030] Example 4 The present invention also provides a storage medium storing a computer program, which executes an entity-level credibility assessment method when running.
[0031] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. An entity level trustworthiness evaluation method, characterized in that, include: Step 101: Divide the medical image dataset into a training set, a calibration set, and a test set; Step 102: Fine-tune the medical diagnostic model based on the training set, embed a Bayesian adapter at the end of its decoding layer, and output the probability estimate of the medical entity and its Bayesian variance. Step 103: Based on the Bayesian adapter-based medical diagnostic model, calculate the inconsistency score of each medical entity in the calibration set and generate a dynamic inconsistency threshold. Step 104: Based on the dynamic inconsistency threshold, generate multiple candidate predicted diagnoses for the test images, extract the medical entity-level prediction set, and filter the candidate results; Step 105: Calculate the medical entity-level information entropy based on the candidate results and generate a visual heatmap; Step 106: Extract the description with the highest predicted probability from the candidate results to form a complete medical diagnosis. Integrate the heat map with the medical diagnosis to assist in clinical decision verification.
2. The entity-level credibility assessment method as described in claim 1, characterized in that, In step 102, the pre-trained medical diagnostic model HuatuoGPT-Vision is used as the initial architecture. This model consists of a visual encoder from LLaVA-1.5 and a text decoder from LLaMA-3-8B. A Bayesian adapter is embedded in the last layer of the model decoder, and the final output includes the probability estimate of medical entities and their Bayesian variance.
3. The entity-level credibility assessment method as described in claim 2, characterized in that, In step 103, a slight adversarial perturbation is applied to the medical images, and the inconsistency score of each medical entity in the calibration set is calculated based on the magnitude of the change in the predicted probability, in conjunction with Bayesian variance.
4. An entity-level credibility assessment device, characterized in that, include: The first processing module is used to divide the medical image dataset into a training set, a calibration set, and a test set. The second processing module is used to fine-tune the medical diagnostic model based on the training set, embed a Bayesian adapter at the end of its decoding layer, and output the probability estimate of the medical entity and its Bayesian variance. The third processing module is used to calculate the inconsistency score of each medical entity in the calibration set based on the Bayesian adapter-based medical diagnostic model and generate a dynamic inconsistency threshold. The fourth processing module is used to generate multiple candidate predictive diagnoses for the test images based on the dynamic inconsistency threshold, extract the medical entity-level prediction set, and filter the candidate results. The fifth processing module is used to calculate the medical entity-level information entropy based on the candidate results and generate a visual heatmap. The sixth processing module is used to extract the description with the highest prediction probability from the candidate results to form a complete medical diagnosis, integrating the heat map with the medical diagnosis to assist in clinical decision verification.
5. The entity-level credibility assessment device as described in claim 4, characterized in that, The second processing module uses the pre-trained medical diagnostic model HuatuoGPT-Vision as the initial architecture. This model consists of a visual encoder from LLaVA-1.5 and a text decoder from LLaMA-3-8B. A Bayesian adapter is embedded in the last layer of the model decoder, and the final output includes the probability estimate of medical entities and their Bayesian variance.
6. The entity-level credibility assessment device as described in claim 5, characterized in that, The third processing module applies a slight adversarial perturbation to the medical images and, based on the magnitude of the change in the predicted probability, calculates the inconsistency score for each medical entity in the calibration set using Bayesian variance.
7. An entity-level credibility assessment system, characterized in that, include: A memory and a processor, wherein the memory stores a computer program executed by the processor, the computer program performing the entity-level trust assessment method as described in any one of claims 1-3 when executed by the processor.
8. A storage medium, characterized in that, The storage medium stores a computer program that, when executed, performs the entity-level credibility assessment method as described in any one of claims 1-3.