Medical question and answer method based on big language model illusion detection
By combining evidential deep learning and structural entropy methods, the large language model is optimized, the problem of hallucination phenomena in medical Q&A is solved, the model's self-awareness and output reliability are improved, and the accuracy and safety of medical information are ensured.
Patent Information
- Application Number
- CN202510164981.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-13
AI Technical Summary
The large language model has an ‘aliambic phenomenon’ in medical Q&A scenarios, generating false or inaccurate outputs, causing patients to receive misleading medical information, seriously threatening health and life safety.
A method based on evidence deep learning and structural entropy is used to optimize the large language model, and the final answer of the output is judged and the threshold is set to detect the hallucination output through the weighted fusion of the loss function and the structural entropy loss function of evidence deep learning.
It effectively enhances the model's self-awareness in the medical field, improves the reliability of medical Q&A output, promptly detects and prevents hallucinatory output, and reduces the risk of misleading patient information.
Smart Images

Figure CN120144701A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical Q&A, and particularly relates to a medical Q&A method based on hallucination detection of large language models. Background Art
[0002] With the rapid development of natural language processing (NLP) technology, large language models (LLMs) are gradually being applied to medical Q&A systems, bringing a higher level of intelligence to human-computer interaction. These models have achieved remarkable results in tasks such as text generation and Q&A, being able to quickly process questions and provide corresponding answers.
[0003] However, the "hallucination phenomenon" existing in LLMs is particularly intractable in the high-risk field of medicine. The so-called hallucination phenomenon refers to the model generating false or inaccurate outputs. In the medical Q&A scenario, this may lead to patients receiving misleading medical information, such as incorrect diagnostic suggestions, medication guides, etc., which may further cause serious consequences, seriously threatening the health and life safety of patients, not only greatly affecting the user experience, but also causing irreparable serious consequences.
[0004] The generation of the hallucination phenomenon mainly stems from the fact that the model fails to accurately grasp its own knowledge boundary during the training process. When facing problems with insufficient knowledge in the medical field, it lacks a clear understanding and accurate assessment of its own capabilities. Currently, the commonly used methods for detecting hallucinations mostly rely on semantic entropy to judge the credibility of the output. However, since it relies on the model itself for semantic detection, it may instead introduce new hallucinations. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a medical Q&A method based on hallucination detection of large language models, which optimizes the large language model by combining the theory of evidential deep learning and structural entropy to effectively solve the problems brought by the hallucination phenomenon in the medical Q&A scenario.
[0006] To achieve the above object, the technical solution adopted by the present invention is: a medical Q&A method based on hallucination detection of large language models, including the steps of:
[0007] S10 Data collection: Obtain questions and use the large language model to generate multiple answers for the input questions;
[0008] S20 Perform data preprocessing on the multiple answers given by the large language model collected;
[0009] S30 Use evidential deep learning to process the questions to obtain the loss function of evidential deep learning;
[0010] S40 Use structural entropy and enhanced semantic clustering to process the generated answers to obtain the structural entropy loss function;
[0011] S50 Knowledge boundary hallucination detection, integrating the loss function of evidential deep learning and the structural entropy loss function to obtain the total weighted loss function, and judging the final answer output by setting a threshold.
[0012] Furthermore, the data collection process includes the steps of:
[0013] S101 Input question definition: First, a series of input questions covering different medical topics and fields need to be defined, including: The question forms include open-ended questions, closed-ended questions, and hypothetical questions; the number of questions;
[0014] S102 Answer generation: Use a large language model to generate answers, including: Diversity generation: For each input question, the model will generate multiple possible answers;
[0015] During the generation process, a multi-round dialogue method is adopted to gradually guide the model to provide answers.
[0016] Furthermore, the data preprocessing process includes the steps of:
[0017] Delete duplicate answers;
[0018] Check whether the logic and grammar of the generated answers are correct.
[0019] Furthermore, use evidential deep learning to process the questions, including:
[0020] S301 Direction of the evidence vector: Convert the input medical question text into vector form through word embedding technology; Softmax transformation: Use the Softmax function to transform the above input vector;
[0021] S302 Magnitude of the evidence vector: Use the input question vector processed by word embedding; Perform linear layer and transformation calculations: First train a linear layer to perform a linear transformation on the input vector; After passing the input vector into this linear layer, apply the Sigmoid function to the output of the linear layer;
[0022] S303 Conversion of evidence values and application of Dirichlet distribution: The Dirichlet distribution is used to simulate the probability distribution of multinomial distribution parameters; After combining the model output with the Dirichlet distribution, the model output is transformed into a distribution containing rich information, which contains various uncertainty information of the model when processing input questions; Based on the direction d i and magnitude S of the evidence vector, convert the logits output by the model into evidence values and define the parameters of the Dirichlet distribution;
[0023] S304 Introduce a regularization term and determine the loss function: Introduce a regularization term to further optimize the judgment of the knowledge boundary;
[0024] The loss function is:
[0025]
[0026] where is the regularization term; α i is the Dirichlet distribution parameter, β and γ are balance coefficients, n is the number of samples, i and j are the sample index and class index related to the Dirichlet distribution parameter, and K is the number of classes.
[0027] Furthermore, using structural entropy and enhanced semantic clustering to process the generated answers, including the steps:
[0028] S401 Construct a graph structure: Represent the generated answers as a graph structure G = (V, E), where: The vertex set V is each generated answer s (i) corresponding to a vertex in the graph; The edge set E is to add edges by calculating the similarity between answers, and the weight of the edge is calculated based on the similarity score between answers, and the weight of the edge reflects the relationship strength between two answers;
[0029] S402 Definition and calculation of structural entropy: Establish an encoding tree T of graph G, which represents graph G in a tree structure; Each node of the encoding tree corresponds to a non-empty subset of vertices in V, and each node represents a group of related answers; The root node λ of the encoding tree represents the entire vertex set of graph G, representing all possible answers; Each leaf node corresponds to a vertex in graph G, that is, a specific generated answer; Calculate the structural entropy of each node, and obtain the total structural entropy of graph G by calculating the sum of the entropies of all nodes in tree T; By minimizing the structural entropy, identify the semantic equivalent groups of the generated answers;
[0030] S403 Enhance semantic clustering: Improve the clustering ability of the generated answers through deeper semantic analysis, construct a similarity graph between the generated answers, and add edges in the similarity graph according to a preset threshold, thereby forming a network structure reflecting the semantic relationship of the answers;
[0031] S404 Calculation of structural entropy and loss function: For each node α in the similarity graph, calculate its structural entropy and obtain the structural entropy; The loss function obtained from the total structural entropy is:
[0032]
[0033] where λ 1 is a hyperparameter, and H T (G) is the total structural entropy.
[0034] Furthermore, for a certain node α, the calculation formula for its structural entropy H T (G; α) is:
[0035]
[0036] In this formula: g α is the sum of the weights of the edges connected to node α; m is the sum of the weights of all the edges in the graph; V α is the number of vertices in node α, representing the number of answers related to this node; is the number of vertices in the parent node of node α, representing the number of answers directly connected to the parent node of node α;
[0037] The total structural entropy H T (G) of graph G is the sum of the entropies of all the nodes in tree T, expressed as:
[0038] H T (G) = ∑ α∈T H T (G; α).
[0039] Furthermore, to enhance the clustering ability of the generated answers through deeper semantic analysis, a similarity graph is constructed among the generated answers, and edges are added in the similarity graph according to a preset threshold, thereby forming a network structure reflecting the semantic relationships of the answers, including the steps of:
[0040] Semantic reconstruction: Jointly encode the generated answers with the input questions, and supplement the possibly missing information with the help of a semantic reconstruction model;
[0041] Generate answer embedding representations: Use a pre-trained sentence embedding model to convert each generated answer s (i) into an embedding vector e (i) ;
[0042] Calculate semantic similarity: For each pair of generated answers s (i) and s (j) , calculate the semantic similarity score by combining the entailment score and the embedding vector similarity;
[0043] Construct a similarity graph: According to the calculation results of the similarity scores, construct a similarity graph:
[0044] If the similarity of two answers is higher than the preset threshold θ, add a weighted edge for them in graph G, and the weight is their similarity value.
[0045] Furthermore, the loss function of evidential deep learning and the structural entropy loss function are weighted and fused to obtain the total loss function
[0046] Set the loss threshold θ t , if is greater than θ t, indicating that the answer may be hallucinatory or beyond the knowledge boundary, and the model needs to re-evaluate the answer or prompt the user about the uncertainty; if is less than or equal to θ t , then the answer can be provided to the user as a relatively reliable output.
[0047] Furthermore, the total loss function
[0048] where L edl_loss is the loss function of evidence-based deep learning, is the structural entropy loss function; λ 2 and λ 3 are weight hyperparameters used to adjust the contribution ratio of the two loss functions in the total loss.
[0049] Furthermore, hallucination detection: Set a loss threshold θ t , if it is greater than θ t , it indicates that there are problems with the model output in terms of disease diagnosis, judgment of the knowledge boundary of symptom description, and semantic consistency with the patient's medical record information, and there may be hallucinatory content such as fictional diagnoses and incorrect medication suggestions, or it exceeds the medical knowledge boundary mastered by the model; at this time, re-evaluate the generated answer; if it is less than or equal to θ t , then it is considered that the answer generated by the model is within the acceptable range and is output to the user.
[0050] Beneficial effects of adopting this technical solution:
[0051] 1. Enhance the self-awareness of the model in the medical field: When dealing with medical Q&A, large language models often have difficulty clearly defining their own knowledge scope. By introducing EDL, this invention helps the model establish self-awareness and enables it to accurately identify the boundaries in medical knowledge. For example, when faced with questions related to rare diseases or complex medical conditions and the model lacks sufficient knowledge reserves and has uncertainties, the model can actively refuse to answer the question, avoiding giving incorrect medical advice due to blind answering and reducing the risk of misleading patient information, thereby enhancing the trust of patients and medical staff in the model.
[0052] 2. Improving the reliability of medical Q&A outputs: By combining structural entropy and semantic clustering, the present invention can effectively evaluate the credibility of the answers generated by the model for medical questions. As an indicator to measure system complexity and information content, structural entropy helps the model identify semantic relationships and similarities between answers. In the medical Q&A scenario, when the model faces multiple possible outputs, by minimizing structural entropy, the system can screen out the most representative and accurate answers and discard those with high uncertainty. This mechanism ensures the reliability of the model outputs in this high-risk field, enabling it to provide more accurate medical information for patients and medical staff.
[0053] 3. Effectively detecting hallucination outputs in medical Q&A: The optimization method proposed by the present invention not only focuses on the self-awareness of the model in the medical field and the definition of knowledge boundaries, but also particularly emphasizes the detection of hallucination outputs in medical Q&A. Through the combination of structural entropy and semantic clustering, the model can accurately identify potential hallucination outputs within its knowledge range and reject them in a timely manner. For example, when the treatment plan generated by the model contradicts known guidelines, or the diagnosis result deviates from common medical cognition, the detection mechanism can effectively prevent the spread of incorrect medical information based on the model's clear understanding of its own medical knowledge and sensitivity to uncertainty, which is particularly crucial in medical application scenarios with extremely high accuracy requirements.
[0054] In summary, the purpose of the present invention is to construct LLM optimization by combining evidence-based deep learning and structural entropy to enhance the model's self-awareness in the medical field, improve the reliability of medical Q&A outputs, effectively detect hallucination outputs in Q&A, promote the research of deep learning algorithms in the medical NLP field, enhance the usage experience of patients and medical staff, and effectively address the challenges in high-risk fields such as medicine. This innovative method not only provides a new solution for existing medical NLP technologies, but also lays a solid foundation for the development of future intelligent medical applications. By achieving the above goals, the present invention expects to promote the development of medical NLP technologies while enhancing the intelligent level of medical human-computer interaction and providing more reliable support for medical industry applications. Brief Description of the Drawings
[0055] Figure 1 It is a schematic flowchart of a medical Q&A method based on large language model hallucination detection according to the present invention;
[0056] Figure 2 It is a schematic system framework diagram of a medical Q&A method based on large language model hallucination detection in an embodiment of the present invention. Detailed Embodiments
[0057] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings.
[0058] In this embodiment, referring to Figure 1 and Figure 2 as shown, the present invention proposes a medical Q&A method based on large language model hallucination detection, including the steps of:
[0059] S10 Data collection: Obtain questions and use a large language model to generate multiple answers for the input questions;
[0060] S20 Perform data preprocessing on the multiple answers given by the large language model collected;
[0061] S30 Use evidential deep learning to process the questions to obtain the loss function of evidential deep learning;
[0062] S40 Use structural entropy and enhanced semantic clustering to process the generated answers to obtain the structural entropy loss function;
[0063] S50 Knowledge boundary hallucination detection, integrate the loss function of evidential deep learning and the structural entropy loss function to obtain the total weighted loss function, and judge the final answer output by setting a threshold.
[0064] As an optimized solution of the above embodiment, in the data collection process, it includes the steps of:
[0065] S101 Input question definition: First, a series of input questions covering different medical topics and fields need to be defined; these questions should be extensive and representative to ensure that the generated answers can reflect diverse viewpoints and information. Including: Question form: Ensure that the expression of the questions is diverse, including open-ended questions, closed-ended questions, and hypothetical questions, which can prompt the model to generate answers with different styles and information depths; Number of questions. Determine a reasonable number of questions to ensure that the generated answer set is rich enough, and at the same time, it will not cause an excessive burden on processing and analysis.
[0066] S102 Answer generation: Use a large language model to generate answers, including: Diversity generation: For each input question, the model will generate multiple possible answers; these answers reflect different narrative ways and information contents, thus enriching the answer set.
[0067] During the generation process, adopt a multi-round dialogue method to gradually guide the model to provide answers. This helps to explore the same question from different angles and increase the depth of the answers.
[0068] As an optimized solution of the above embodiment, in the data preprocessing process, it includes the steps of:
[0069] Delete duplicate answers to ensure sufficient differences between different answers, avoid duplicate content, and improve the diversity of data;
[0070] Check whether the logic and grammar of the generated answers are correct. When problems are found, make corrections or regenerate in a timely manner.
[0071] As an optimized solution to the above embodiments, use evidential deep learning to process the problem, including:
[0072] S301 Direction of the evidence vector: Convert the input medical problem text into vector form through word embedding technology; for example, for the problem "How to manage the daily diet of hypertensive patients", the model will map each word in the text, such as "hypertension", "patient", "how", "carry out", "daily", "diet", "management", into corresponding vectors respectively, and these vectors together form the input vector sequence. This vector sequence contains the semantic information of the medical problem and provides a basis for subsequent calculations. Softmax transformation: Use the Softmax function to transform the above input vector; the Softmax function is often used in multi-classification problems, and it can transform the input vector into a probability distribution, making the sum of all output values equal to 1. The calculation formula is
[0073]
[0074] where z is the input vector and K represents the number of categories.
[0075] Through this calculation, the "direction" of the decoupled evidence vector is obtained. It quantifies the relative confidence of the model's prediction for each category and helps the model initially judge the tendency of the input problem in different categories.
[0076] S302 Magnitude of the evidence vector: Use the input problem vector processed by word embedding; since in practical applications, when the number of categories is large, the target probability carrier is often sparse, in order to more accurately evaluate the magnitude of the model output, this step needs further processing. Perform linear layer and transformation calculations: First train a linear layer to perform a linear transformation on the input vector; after passing the input vector into this linear layer, apply the Sigmoid function to the output of the linear layer. The Sigmoid function can compress the output value into the interval (0,1), and then through specific mathematical transformations, a lossless non-negative magnitude is obtained. The calculation formula is
[0077]
[0078] The S obtained through this series of operations can evaluate the magnitude of the model output in a more refined way, enabling the model to better understand its own tendency to predict the input problem and providing an important basis for subsequent judgment of uncertainty.
[0079] Conversion of Evidence Values and Application of Dirichlet Distribution: The Dirichlet distribution is used to model the probability distribution of multinomial distribution parameters, and it can express the uncertainty of the model in a natural way. After combining the model output with the Dirichlet distribution, the output of the model is transformed into a distribution containing rich information, which contains various uncertainty information of the model when dealing with input problems. By analyzing this information, the uncertainty of the model when facing problems can be better captured; in the direction d i and size S of the evidence vector, the logits output by the model are converted into evidence values, and the parameters of the Dirichlet distribution are defined. The core purpose of this step is to transform the model output into a distribution form that can better reflect uncertainty, so as to more accurately detect the knowledge boundary.
[0080] Calculation of Dirichlet Distribution: The Dirichlet distribution is calculated by the following formula:
[0081] α i = e i + 1 = S·d i + 1, i = 1, 2, …, K,
[0082]
[0083] The predicted probability of the k-th unit set is the average value of the corresponding Dirichlet distribution
[0084] S304 Introduction of Regularization Terms and Determination of Loss Function: Introducing regularization terms further optimizes the judgment of the knowledge boundary;
[0085] Taking the Dirichlet distribution parameter α i as the input, calculate the Dirichlet parameter after removing non-misleading evidence from the prediction parameters:
[0086]
[0087] And calculate according to the formula:
[0088]
[0089] It can punish the divergence when the model is in the "I don't know" state, prompting the model to be more cautious when facing uncertainty problems, reducing misjudgments, and thus improving the accuracy of knowledge boundary detection.
[0090] The loss function is:
[0091]
[0092] Among them, is a regular term that penalizes the divergence of the model when it is in the “I don’t know” state; α i is the Dirichlet distribution parameter, β and γ are the balance coefficients, n is the number of samples, i, j are the sample index and category index related to the Dirichlet distribution parameter, and K represents the number of categories.
[0093] The second half of the loss function measures the actual distribution of the parameters of the Dirichlet distribution and the uniform distribution (all parameters are equal, that is, ) differences.
[0094] If the model's judgment on the problem is close to a uniform distribution, it means that the model is uncertain, and the loss value of the second half increases. The addition of the two parts prompts the model to avoid divergent judgments when facing problems, and to give reasonable feedback when uncertain, effectively improving the accuracy of knowledge boundary detection.
[0095] Applying evidential deep learning to large language models can significantly enhance their self-awareness when faced with unknown problems. Large language models can keenly identify uncertainties in inputs with layers of evidence.
[0096] As an optimization solution for the above embodiment, after the judgment of the knowledge boundary is completed, in order to ensure the accuracy of the model output, the generated answer will be further considered. Here, the structural entropy and enhanced semantic clustering are used. Traditional hallucination detection methods often rely on the answers generated by the model for judgment. However, when facing complex natural language generation tasks, this method is easily affected by the limitations of the model itself, resulting in inaccurate recognition of hallucination output.
[0097] In order to solve this problem, the structural entropy and enhanced semantic clustering methods introduced in the present invention, as a new measurement tool, can more comprehensively evaluate the similarity and uncertainty between generated answers. This method not only focuses on the content of the generated answers, but also considers the relationship between the answers, thereby providing a more accurate basis for judgment. In addition, the enhanced semantic clustering technology optimizes the effect of answer clustering by performing detailed semantic analysis and reconstruction of the generated answers. This technology can effectively identify answers that are semantically similar but have large differences in form, thereby reducing misjudgments caused by different expressions. This method significantly improves the ability to identify semantically equivalent groups of generated answers, thereby improving the overall reliability of the model output.
[0098] Using structural entropy and enhanced semantic clustering, the generated answers are processed, including the following steps:
[0099] S401 constructs a graph structure: the generated answers are represented as a graph structure G = (V, E), where: the vertex set V is each generated answer s (i)Corresponds to a vertex in the graph; the edge set E adds edges by calculating the similarity between answers (such as the combined entailment score and the embedding vector similarity), and the weight of the edge is calculated based on the similarity score between answers, and the weight of the edge reflects the strength of the relationship between two answers. The edge weights between answers with high similarity scores are higher, and vice versa. This step helps to construct a relationship network between answers and lays the foundation for subsequent structural entropy calculation.
[0100] Definition and calculation of S402 structural entropy: Establish the encoding tree T of graph G, which represents graph G in a tree structure; each node of the encoding tree corresponds to a non-empty subset of vertices in V, and each node represents a group of related answers; the root node λ of the encoding tree represents the entire vertex set of graph G, representing all possible answers; each leaf node corresponds to a vertex in graph G, that is, a specific generated answer; this structure enables each layer of the tree to reflect the hierarchical relationship and similarity between answers. Calculate the structural entropy of each node, and obtain the total structural entropy of graph G by calculating the sum of the entropies of all nodes in tree T; by minimizing the structural entropy, identify the semantic equivalent groups of generated answers;
[0101] S403 Strengthen semantic clustering: Improve the clustering ability of generated answers through deeper semantic analysis, construct a similarity graph between the generated answers, and add edges in the similarity graph according to a preset threshold, thereby forming a network structure reflecting the semantic relationship of answers.
[0102] S404 Calculation of structural entropy and loss function: For each node α in the similarity graph, calculate its structural entropy and obtain the structural entropy; the loss function obtained from the total structural entropy is:
[0103]
[0104] where λ 1 is a hyperparameter, and H T (G) is the total structural entropy.
[0105] where, for a certain node α, its structural entropy H T (G; α) is calculated by the formula:
[0106]
[0107] In this formula: g α is the sum of the weights of the edges connected to node α; m is the sum of the weights of all edges in the graph; V α is the number of vertices in node α, representing the number of answers related to this node; is the number of vertices in the parent node of node α, representing the number of answers directly connected to the parent node of node α;
[0108] The total structural entropy H of graph GT (G) is the sum of the entropies of all nodes in tree T, expressed as:
[0109] H T (G)=∑ α∈T H T (G;α).
[0110] In free-form generation, traditional semantic clustering methods often fail to effectively handle different answers to the same question. Although these answers may have different expressions, they may essentially convey the same semantic information. This leads to poor clustering results and increases the risk of hallucination output. To address this issue, we propose an enhanced semantic clustering method. By performing deeper semantic analysis to improve the clustering ability of the generated answers, a similarity graph is constructed among the generated answers, and edges are added to the similarity graph according to a preset threshold, thereby forming a network structure that reflects the semantic relationships of the answers.
[0111] This enhanced semantic clustering method effectively overcomes the deficiencies of traditional methods, ensuring that in free-form generation, even answers with different expressions but similar semantics can be accurately identified and classified. This not only improves the clustering accuracy but also lays a foundation for subsequent structural entropy calculation and hallucination detection.
[0112] The core of this step is to ensure that answers with different expressions can be classified into the same semantic group through more precise semantic analysis of the generated answers.
[0113] By performing deeper semantic analysis to improve the clustering ability of the generated answers, a similarity graph is constructed among the generated answers, and edges are added to the similarity graph according to a preset threshold, thereby forming a network structure that reflects the semantic relationships of the answers, including the steps of:
[0114] Semantic reconstruction: Jointly encode the generated answers with the input questions, and supplement the possibly missing information with the help of a semantic reconstruction model;
[0115] It is possible to select models such as Sentence-Transformers or T5 models to encode relevant medical contexts such as questions and medical records, so as to guide the rewriting and improvement of answers. These models can accurately capture the semantic information of the answers and help generate more accurate answer descriptions.
[0116] Example: Suppose the question is "What is the patient's current blood glucose level?", and the context is "The patient had a comprehensive physical examination today, and the blood glucose test result showed 6.5 mmol / L." After semantic reconstruction, the original answer "6.5 mmol / L" can be changed to "The patient's current blood glucose level is 6.5 mmol / L."
[0117] Generate answer embedding representations: Use a pre-trained sentence embedding model (such as MiniLM) to convert each generated answer s (i) into an embedding vector e (i) . The embedding vector can effectively capture the semantic information of the answer, enabling the evaluation of the similarity between answers through these vectors in subsequent similarity calculations.
[0118] Calculate semantic similarity: For each pair of generated answers s (i) and s (j) , calculate the semantic similarity score by combining the entailment score and the embedding vector similarity;
[0119] The formula for the semantic similarity score is:
[0120] F(s (i) |s (j) ) = (1 - T e ) cos s im(e (i) , e (j) ) + T e E(s (i) , s (j) )
[0121] where cos s im(e (i) , e (j) ) represents the cosine similarity calculated based on the embedding vectors; E(s (i) , s (j) is the joint entailment score, reflecting the semantic relationship between the answers; T e is the balance parameter, used to control the weights of the two similarity metrics.
[0122] Construct a similarity graph: Based on the calculation results of the similarity scores, construct a similarity graph:
[0123] If the similarity between two answers is higher than a preset threshold θ, add a weighted edge for them in the graph G, and the weight is their similarity value. This process helps to form a network structure reflecting the semantic relationships between answers.
[0124] As an optimized solution to the above embodiment, the loss function of evidential deep learning and the structural entropy loss function are weighted and fused to obtain the total loss function
[0125] Set the loss threshold θ t , if is greater than θ t , it indicates that the answer may have hallucinations or exceed the knowledge boundary, and the model needs to re-evaluate the answer or prompt the user about the uncertainty; if is less than or equal to θ t, the answer can be provided to the user as relatively reliable output.
[0126] Among them, the evidential deep learning theory can enable the model to accurately perceive its confidence in different inputs when processing input data by introducing an additional evidence layer, thereby effectively judging the knowledge boundary. However, when facing complex natural language generation tasks, simply relying on evidential deep learning to judge the knowledge boundary cannot completely avoid the model from generating hallucinated content.
[0127] On the other hand, the structural entropy and enhanced semantic clustering method detect hallucinated outputs from the perspectives of the similarity and semantic relationship of the generated answers, but there are deficiencies in judging the model's understanding and confidence in the input questions. Therefore, fusing the loss functions of the two can make up for each other's strengths and weaknesses, provide more comprehensive and accurate constraints for the model, and improve the reliability and accuracy of the model.
[0128] To achieve the collaborative optimization of knowledge boundary judgment and hallucination detection, the above two loss functions are weighted and fused.
[0129] Total loss function
[0130] Among them, L edl_loss is the loss function of evidential deep learning, is the structural entropy loss function; λ 2 and λ 3 are weight hyperparameters used to adjust the contribution ratio of the two loss functions in the total loss. By reasonably adjusting these two hyperparameters, the model can achieve a better balance in both knowledge boundary judgment and hallucination detection.
[0131] Among them, for hallucination detection: set a loss threshold θ t , if it is greater than θ t , it indicates that there are problems with the model output in terms of disease diagnosis, symptom description knowledge boundary judgment, and semantic consistency with the patient's medical record information, and there may be hallucinated content such as fictional diagnoses and incorrect medication suggestions, or it exceeds the medical knowledge boundary mastered by the model; at this time, the generated answer is re-evaluated, such as double-checking the medical record data again, calling more medical knowledge base information, or directly prompting the user that the answer may be uncertain; if it is less than or equal to θ t , then it is considered that the answer generated by the model is within the acceptable range and is output to the user.
[0132] Through such a judgment process, combined with the fused loss function, the model can more effectively define the boundary in medical knowledge, screen out possible hallucinated content, and greatly improve the quality and reliability of the output results in the medical scenario.
[0133] The integrated loss function takes into account both the model's confidence in the input question and the semantic relationships between the generated answer and medical standard terms, medical record information, etc. When faced with unknown questions such as rare diseases and complex medical conditions, the model can accurately determine its boundaries in medical knowledge reserves based on the evidence-based deep learning loss function, avoiding giving incorrect diagnostic conclusions or treatment suggestions. At the same time, the loss functions based on structural entropy and enhanced semantic clustering can perform strict hallucination detection on the generated answers, such as diagnostic results and medication instructions, to ensure that the output content conforms to medical facts and is reliable. This synergy enables the model to provide more valuable and accurate medical information services to doctors, patients, etc. when dealing with various medical natural language question-and-answer tasks, effectively improving the performance and practicality of the model in the medical field.
[0134] The large language model optimization method based on structural entropy and evidence-based deep learning proposed in this invention has achieved remarkable results in solving the hallucination problem of medical large language models and enhancing the awareness of medical knowledge boundaries, bringing new breakthroughs and inspirations to the development of the medical natural language processing field.
[0135] In enhancing the model's self-awareness, the application of the evidence-based deep learning theory plays a key role. By introducing an additional evidence layer, the model can accurately perceive its confidence in various medical question inputs. When faced with unknown medical problems such as rare disease diagnosis and complex medical condition analysis, the model can keenly identify uncertainties and avoid blindly giving incorrect diagnostic conclusions, treatment suggestions, and other information. This improvement in self-awareness not only significantly reduces the risk of misleading patient information but also enhances the trust of users such as doctors and patients when interacting with the model. For example, when faced with some diseases with atypical symptoms, the model will not easily give inaccurate diagnoses but will instead indicate possible uncertainties and guide doctors to conduct further examinations.
[0136] Another important achievement of this study is the improvement in the reliability of the model output. The combination of structural entropy and enhanced semantic clustering methods provides a powerful means for evaluating the credibility of the answers generated by medical models. Structural entropy helps the model identify the semantic relationships and similarities between medical answers. When analyzing issues such as treatment plans for comorbid diseases and drug interactions, the process of minimizing structural entropy can screen out the most representative medical suggestions and effectively reject outputs with higher uncertainties. The enhanced semantic clustering technology, through in-depth semantic analysis and reconstruction, accurately identifies answers that are semantically similar but have different forms for patient symptom descriptions, medical record interpretations, etc., reducing misjudgments and greatly enhancing the overall reliability of the model output in medical scenarios.
[0137] The loss function that fuses evidential deep learning and structural entropy comprehensively considers the confidence of the model in the input question and the semantic relationship between the generated answer and the medical knowledge system and the patient's medical record information. When facing complex medical natural language generation tasks, the model can accurately define the boundaries of medical knowledge and effectively filter out hallucinations such as fictional diagnoses and incorrect medication suggestions. This mechanism is particularly crucial in the application of the high-risk medical field, effectively preventing the spread of incorrect medical information, ensuring the accuracy and reliability of the information obtained by patients, and directly related to the treatment effect and life health of patients.
[0138] Overall, this method improves the performance of large language models in tasks such as medical Q&A and medical generative conversations, and promotes the improvement of the intelligent level of medical human-computer interaction. Although certain achievements have been made currently, there is still room for optimization in the future. For example, further explore the optimization strategy of hyperparameters, improve the adaptability of the model in complex medical scenarios, such as multi-disciplinary joint diagnosis and treatment, cross-regional medical data differences, etc., and expand this method to more medical natural language processing tasks, such as medical literature review generation, medical image report interpretation assistance, etc. It is believed that with the continuous in-depth research, this method will provide more solid and reliable technical support for various application scenarios in the medical industry, help the medical natural language processing technology reach a new height, and ultimately benefit the vast number of patients.
[0139] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A medical question-answering method based on large language model hallucination detection, characterized in that: Includes steps: S10 Data Collection: Obtain questions and use a large language model to generate multiple answers for the input questions; S20 performs data preprocessing on the multiple answers collected by the large language model; S30 uses evidential deep learning to process the problem and obtains the loss function of evidential deep learning; S40 uses structural entropy and enhanced semantic clustering to process the generated answers and obtain the structural entropy loss function; S50 knowledge boundary hallucination detection integrates the loss function of evidential deep learning and the structural entropy loss function to obtain the total weighted loss function, and determines the final answer of the output by setting a threshold.
2. A medical question-answering method based on large language model hallucination detection according to claim 1, characterized in that: The data collection process includes the following steps: S101 Input question definition: First, a series of input questions covering different medical topics and fields need to be defined, including: question formats including open-ended questions, closed-ended questions, and hypothetical questions; number of questions; S102 Answer Generation: Generate answers using a large language model, including: Diversity Generation: For each input question, the model will generate multiple possible answers; During the generation process, multiple rounds of dialogue are used to gradually guide the model to provide answers.
3. A medical question-answering method based on large language model hallucination detection according to claim 1, characterized in that: The data preprocessing process includes the following steps: Remove duplicate answers; Check the generated answers for logical and grammatical correctness.
4. A medical question-answering method based on large language model hallucination detection according to claim 1, characterized in that: Using evidence-based deep learning to tackle problems, including: S301 Direction of evidence vector: convert the input medical question text into vector form through word embedding technology; Softmax conversion: use Softmax function to convert the above input vector; S302 Size of evidence vector: Use the input question vector after word embedding; perform linear layer and transformation calculations: first train a linear layer to perform a linear transformation on the input vector; after passing the input vector into the linear layer, apply the Sigmoid function to the output of the linear layer; S303 Conversion of evidence value and application of Dirichlet distribution: Dirichlet distribution is used to simulate the probability distribution of multiple distribution parameters; after combining the model output with Dirichlet distribution, the model output is converted into a distribution containing rich information, which contains various uncertainty information when the model processes the input problem; in the direction of the evidence vector d i Based on the size S, the logits output by the model are converted into evidence values, and the parameters of the Dirichlet distribution are defined; S304 introduces a regularization term and determines a loss function: introducing a regularization term to further optimize the judgment of the knowledge boundary; The loss function is: Among them, l edl_reg is the regularization term; α i is the Dirichlet distribution parameter, β and γ are the balance coefficients, n is the number of samples, i, j are the sample index and category index related to the Dirichlet distribution parameter, and K represents the number of categories.
5. The medical question-answering method based on large language model hallucination detection according to claim 1, characterized in that: Using structural entropy and enhanced semantic clustering, the generated answers are processed, including the following steps: S401 constructs a graph structure: the generated answers are represented as a graph structure G = (V, E), where: the vertex set V is each generated answer s (i) Corresponding to a vertex in the graph; the edge set E is added by calculating the similarity between answers. The weight of the edge is calculated based on the similarity score between the answers. The weight of the edge reflects the strength of the relationship between the two answers. S402 Definition and calculation of structural entropy: Establish a coding tree T of graph G, which represents graph G in a tree structure; each node of the coding tree corresponds to a non-empty vertex subset in V, and each node represents a set of related answers; the root node λ of the coding tree represents the entire vertex set of graph G, representing all possible answers; each leaf node corresponds to a vertex in graph G, that is, a specific generated answer; calculate the structural entropy of each node, and obtain the total structural entropy of graph G by calculating the sum of the entropies of all nodes in the tree T; identify the semantically equivalent groups of generated answers by minimizing the structural entropy; S403 Enhanced semantic clustering: Improve the clustering ability of generated answers through deeper semantic analysis, build a similarity graph between generated answers, add edges to the similarity graph according to a preset threshold, and thus form a network structure that reflects the semantic relationship of the answers; S404 Calculation of structural entropy and loss function: For each node α in the similarity graph, calculate its structural entropy and obtain the structural entropy; the total structural entropy loss function is: l Structure =λ1·H T (G); Among them, λ1 is a hyperparameter, H T (G) is the total structural entropy.
6. A medical question-answering method based on large language model hallucination detection according to claim 7, characterized in that: For a node α, its structural entropy H T The calculation formula of (G; α) is: In this formula: g α is the sum of the weights of the edges connected to node α; m is the sum of the weights of all edges in the graph; V α is the number of vertices in node α, indicating the number of answers related to the node; V α - is the number of vertices in the parent node of node α, indicating the number of answers directly connected to the parent node of node α; The total structural entropy H of the graph G T (G) is the sum of the entropy of all nodes in the tree T, expressed as: H T (G)=∑ α∈T H T (G;α)。 7. A medical question-answering method based on large language model hallucination detection according to claim 6, characterized in that: The clustering ability of generated answers is improved through deeper semantic analysis, a similarity graph is constructed between the generated answers, and edges are added to the similarity graph according to a preset threshold, thereby forming a network structure that reflects the semantic relationship of the answers, including the following steps: Semantic reconstruction: The generated answer is semantically co-encoded with the input question, and the semantic reconstruction model is used to supplement the missing information. Generate answer embedding representation: Use the pre-trained sentence embedding model to embed each generated answer s (i) Convert to embedding vector e (i) ; Compute semantic similarity: For each pair of generated answers s (i) and (j) , the semantic similarity score is calculated by combining the implication score and the embedding vector similarity; Construct a similarity graph: Based on the calculation results of the similarity scores, construct a similarity graph: If the similarity between two answers is higher than the preset threshold θ, a weighted edge is added to them in graph G, and the weight is their similarity value.
8. The medical question-answering method based on large language model hallucination detection according to claim 1, characterized in that: The loss function of evidential deep learning is weighted and fused with the structural entropy loss function to obtain the total loss function l total ; Set the loss threshold θ t , if l total Greater than θ t , indicating that the answer may be an illusion or exceed the knowledge boundary, and the model needs to re-evaluate the answer or prompt the user for uncertainty; if l total Less than or equal to θ t , then the answer can be provided to the user as a relatively reliable output.
9. A medical question-answering method based on large language model hallucination detection according to claim 8, characterized in that: The total loss function l total =λ2·L edl_loss +λ3·l Structure ; Among them, L edl_loss is the loss function of evidential deep learning, l Structure is the structural entropy loss function; λ2 and λ3 are weight hyperparameters used to adjust the contribution ratio of the two loss functions in the total loss.
10. The medical question-answering method based on large language model hallucination detection according to claim 1, characterized in that: Hallucination detection: setting a loss threshold θ t , if it is greater than θ t , indicating that there are problems with the model output in terms of disease diagnosis, symptom description knowledge boundary judgment and semantic consistency with patient medical record information. There are fictitious diagnoses, wrong medication suggestions, hallucination-like content, or beyond the boundaries of medical knowledge mastered by the model. At this time, the generated answers are re-evaluated; If less than or equal to θ t , the answer generated by the model is considered to be within an acceptable range and the output is provided to the user.
Citation Information
Cited By
Closed source LLM agent illusion monitoring method and system based on incremental generalization boundary
CN120744075A
Large model illusion detection method, system and device based on PCA contribution rate and medium
CN120781171A
Knowledge-enhanced medical illusion static detection and correction method and system
CN120911443A
Big model question and answer method and system for hallucination suppression
CN120929579A
Question and answer processing method, system and equipment based on three-dimensional entropy evaluation and medium
CN121009182A