Medical suggestion obtaining method based on big language model illusion elimination and related assembly
By integrating preset medical knowledge bases and historical case data from specific medical examination institutions in a large language model, the problem of logical errors and inaccuracy of the model when generating medical suggestions is solved, significantly improving the scientificity and credibility of the suggestions.
Patent Information
- Application Number
- CN202510090306.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
AI Technical Summary
Large language models often experience logical errors and ‘hedonies’ problems that do not conform to actual medical knowledge when generating medical advice, resulting in insufficient accuracy and reliability in clinical applications.
By searching the current medical condition data of the target object in the preset medical knowledge base, obtaining related knowledge points, and integrating them into the text generation algorithm to obtain supplementary knowledge points. Then, the current condition data and supplementary knowledge points are input into a target medical large language model based on the historical case data and historical reasoning path training of a specific medical examination institution, and thinking chain derivation is performed to generate accurate medical advice.
It effectively reduces the 'aliasm' phenomenon in generating medical suggestions by large language models, improves the accuracy and reliability of suggestions, and makes the generated medical suggestions more in line with medical logic and actual situations.
Smart Images

Figure CN119943353A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical technology, and in particular to a method, device, equipment and medium for obtaining medical advice based on large language model hallucination elimination. Background Art
[0002] In recent years, the application of artificial intelligence (AI) technology in the medical field has continued to deepen, especially the rapid development of large language models (LLMs), which has attracted widespread attention to application scenarios such as intelligent diagnosis and decision support. These models can generate diagnostic recommendations and treatment plans similar to those of human doctors by learning a large amount of medical data.
[0003] Large language models usually rely on public or general datasets for training. The models usually generate text based on statistical probability without a deep understanding of medical logic and causality. Therefore, with the popularity of these models in actual clinical applications, an important problem has gradually emerged, that is, when generating medical advice, the models often encounter the "hallucination" problem that the generated advice does not match the actual medical knowledge or is logically incorrect. This has become one of the main obstacles to its further development.
[0004] In summary, how to obtain accurate medical advice that complies with medical logic is a problem to be solved in this field. Summary of the invention
[0005] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for obtaining medical advice based on large language model hallucination elimination, so as to obtain accurate medical advice that conforms to medical logic. The specific scheme is as follows:
[0006] In a first aspect, the present application discloses a method for obtaining medical advice based on large language model hallucination elimination, comprising:
[0007] Searching the preset medical knowledge base according to the current medical condition information of the target object to obtain related knowledge points;
[0008] Integrating the related knowledge points into the text generation algorithm to obtain supplementary knowledge points;
[0009] Inputting the current medical condition data and the supplementary knowledge points into a target medical big language model; wherein the target medical big language model is acquired based on historical case data and historical reasoning paths of a specific medical examination institution, and the specific medical examination institution is an institution that examines the target object;
[0010] The target medical big language model uses the current medical condition data and the supplementary knowledge points to deduce a thought chain, so as to output a target medical suggestion after the hallucination information is eliminated.
[0011] Optionally, obtain the target medical language model, including:
[0012] The initial large language model is adjusted using historical case data of a specific medical examination institution to obtain an initial medical large language model;
[0013] The initial medical big language model is trained based on the historical reasoning path and the historical case data to obtain a target medical big language model; wherein the historical reasoning path includes symptom recognition, etiology analysis, diagnosis hypothesis formation, and treatment plan selection.
[0014] Optionally, the adjusting the initial large language model using historical case data of a specific medical examination institution to obtain an initial medical large language model includes:
[0015] Collecting initial historical case data of a specific medical examination institution, and preprocessing the initial historical case data to obtain first target historical case data;
[0016] Extracting abnormal symptom data from the first target historical case data to obtain second target historical case data; wherein the abnormal symptom data is symptom data that does not belong to the corresponding preset range;
[0017] The initial large language model is trained using the first target historical case data to obtain a trained medical large language model, and the trained medical large language model is fine-tuned using the second target historical case data to obtain an initial medical large language model.
[0018] Optionally, the collection of initial historical case data from a specific medical examination institution includes:
[0019] Collect original inpatient case data from specific medical examination institutions;
[0020] Extracting historical medical condition information and historical medical advice from the original inpatient case data;
[0021] Initial historical case data is constructed with the historical medical condition data as model input and the historical medical advice as model output.
[0022] Optionally, the fine-tuning the trained medical language model using the second target historical case data to obtain an initial medical language model includes:
[0023] Taking medical logic rules and preset expert knowledge as constraints, and using the second target historical case data and a multi-task learning framework, the trained medical big language model is fine-tuned to obtain an initial medical big language model.
[0024] Optionally, the target medical large language model uses the current medical condition data and the supplementary knowledge points to perform thought chain deduction to output target medical advice after the hallucination information is eliminated, including:
[0025] The target medical big language model utilizes the current medical condition data and the supplementary knowledge points to perform thought chain deduction to generate tasks corresponding to each sub-reasoning path, and combines the current medical condition data and the supplementary knowledge points to generate sub-conclusions corresponding to each sub-reasoning path task, so as to output target medical recommendations after the hallucination information is eliminated based on the sub-conclusions.
[0026] Optionally, the architecture of the preset medical knowledge base is a hybrid architecture of retrieval and generation, and the vector index of the preset medical knowledge base is constructed based on a semantic vector model;
[0027] Accordingly, the current medical condition data of the target object is searched in a preset medical knowledge base to obtain related knowledge points, including:
[0028] The current medical condition data of the target object is compared with the vector index of each knowledge point in the preset medical knowledge base to obtain the related knowledge points.
[0029] In a second aspect, the present application discloses a medical advice acquisition device based on large language model hallucination elimination, comprising:
[0030] A knowledge base retrieval module is used to search the preset medical knowledge base according to the current medical condition information of the target object to obtain related knowledge points;
[0031] A text generation module, used to integrate the related knowledge points into a text generation algorithm to obtain supplementary knowledge points;
[0032] An information input module, used to input the current medical condition data and the supplementary knowledge points into a target medical big language model; wherein the target medical big language model is acquired based on historical case data and historical reasoning paths of a specific medical examination institution, and the specific medical examination institution is an institution that examines the target object;
[0033] The medical advice output module is used to use the current medical condition data and the supplementary knowledge points to perform thought chain deduction through the target medical large language model to output the target medical advice after the hallucination information is eliminated.
[0034] In a third aspect, the present application discloses an electronic device, comprising:
[0035] Memory, used to store computer programs;
[0036] A processor is used to execute the computer program to implement the steps of the aforementioned disclosed method for obtaining medical advice based on large language model hallucination elimination.
[0037] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the aforementioned method for obtaining medical advice based on large language model hallucination elimination are implemented.
[0038] The beneficial effects of the present application are as follows: the present application searches a preset medical knowledge base according to the current medical condition information of the target object to obtain related knowledge points; integrates the related knowledge points into a text generation algorithm to obtain supplementary knowledge points; inputs the current medical condition information and the supplementary knowledge points into a target medical big language model; wherein the target medical big language model is acquired based on historical case data and historical reasoning paths of a specific medical examination institution, and the specific medical examination institution is an institution that examines the target object; the current medical condition information and the supplementary knowledge points are used by the target medical big language model to deduce a thinking chain to output a target medical recommendation after the illusion information is eliminated. It can be seen that the present application combines the associated knowledge points in the preset medical knowledge base to realize real-time retrieval of relevant medical information and ensure the scientificity and authenticity of the generated content; the target medical big language model is acquired based on the historical case data and historical reasoning paths of a specific medical examination institution, that is, by using the historical medical record data of a specific medical examination institution to train a customized large language model, the target medical big language model is more adaptable to the specific medical environment; further, the target medical big language model is also acquired based on the historical reasoning path, that is, in the process of acquiring the target medical big language model, the model learns the doctor's step-by-step reasoning process according to the historical reasoning path, so in actual application, the target medical big language model can use the current medical condition information and supplementary knowledge points to deduce the thinking chain, simulate the doctor's step-by-step reasoning process, enhance the logic of the generated content, and effectively reduce the "hallucination" phenomenon in the medical big language model. The accuracy and reliability of the generated medical advice are significantly improved, which provides more reliable technical support for the application of the medical big language model in clinical decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0040] Figure 1 A flow chart of a method for obtaining medical advice based on large language model hallucination elimination disclosed in this application;
[0041] Figure 2 A schematic diagram of obtaining medical advice based on large language model hallucination elimination disclosed in this application;
[0042] Figure 3 A schematic diagram of a specific model training based on historical case data disclosed in this application;
[0043] Figure 4 A specific medical thinking chain learning diagram disclosed in this application;
[0044] Figure 5 A schematic diagram of constructing a specific preset medical knowledge base disclosed in this application;
[0045] Figure 6 A schematic diagram of the structure of a medical advice acquisition device based on large language model hallucination elimination disclosed in this application;
[0046] Figure 7 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0048] In recent years, the application of artificial intelligence technology in the medical field has continued to deepen, especially the rapid development of large language models, which has attracted widespread attention to application scenarios such as intelligent diagnosis and decision-making assistance. These models can generate diagnostic recommendations and treatment plans similar to those of human doctors by learning a large amount of medical data.
[0049] Large language models usually rely on public or general datasets for training. The models usually generate text based on statistical probability without a deep understanding of medical logic and causality. Therefore, with the popularity of these models in actual clinical applications, an important problem has gradually emerged, that is, when generating medical advice, the models often encounter the "hallucination" problem that the generated advice does not match the actual medical knowledge or is logically incorrect. This has become one of the main obstacles to its further development.
[0050] To this end, this application provides a medical advice acquisition solution to obtain accurate medical advice that complies with medical logic.
[0051] See also Figure 1 As shown, the embodiment of the present application discloses a method for obtaining medical advice based on large language model hallucination elimination, including:
[0052] Step S11: searching the preset medical knowledge base according to the current medical condition information of the target object to obtain related knowledge points.
[0053] Obtain the current medical condition information of the target object. The medical condition information can specifically include the target object's condition description information or related medical examination reports. The preset medical knowledge base collects a wide range of medical information, including but not limited to disease diagnosis criteria, treatment methods, drug information, etc., to provide high-quality data support for the subsequent generation process. Specifically, 200 medical-related data can be extracted from external resources such as medical literature and clinical diagnosis and treatment guidelines, covering common organs, symptoms, diseases and corresponding diagnosis and treatment plans. This part of the data is used for factual verification of medical retrieval enhancement generation to further improve the accuracy and verifiability of the model in medical content generation. Design a medical knowledge retrieval system that can be seamlessly connected with the generative artificial intelligence model. The retrieval system will search for the most relevant medical knowledge information in the preset medical knowledge base based on the input context information (i.e., the current medical condition information), and pass this medical knowledge information to the generative artificial intelligence model, which outputs related knowledge points.
[0054] In this embodiment, the architecture of the preset medical knowledge base is a hybrid architecture of retrieval and generation, and the vector index of the preset medical knowledge base is constructed based on a semantic vector model. The LightRAG model is a hybrid architecture that combines the advantages of both retrieval and generation methods. By combining the retriever and the generator, LightRAG can retrieve the most relevant fragments from a large number of pre-trained corpora, and use the generator to optimize the combination of these fragments to produce natural, fluent and informative answers. Therefore, when the architecture of the preset medical knowledge base is a hybrid architecture of retrieval and generation, it can reduce the demand for computing resources and improve the processing speed, so that the preset medical knowledge base can be deployed in a resource-constrained environment. BGE (BAAI General Embeddings, i.e., semantic vector model) can better understand the query intent by introducing a retrieval mechanism based on semantic vectors, build a vector index for each knowledge base for subsequent retrieval enhancement generation, and find the most relevant fragments from a large number of documents as the basis for generating content, so the vector index of the preset medical knowledge base is constructed based on the semantic vector model, which helps to ensure that the generated content is closer to the facts and reduce the risk of fictitious or inaccurate information.
[0055] In this embodiment, the searching in the preset medical knowledge base according to the current medical condition data of the target object to obtain the associated knowledge points includes: comparing the current medical condition data of the target object with the vector index of each knowledge point in the preset medical knowledge base to obtain the associated knowledge points. It can be understood that because the vector index of the preset medical knowledge base is constructed based on the semantic vector model, when searching in the preset medical knowledge base according to the current medical condition data of the target object, it can be specifically comparing the current medical condition data of the target object with the vector index of each knowledge point in the preset medical knowledge base to obtain the associated knowledge points.
[0056] Step S12: Incorporating the associated knowledge points into the text generation algorithm to obtain supplementary knowledge points.
[0057] The retrieved related knowledge points are fragmented. Based on this knowledge, the text generation algorithm is used to integrate the knowledge points so that these related knowledge points are logically connected together. In other words, by dynamically integrating the retrieved related knowledge points into the text generation algorithm to obtain supplementary knowledge points, the supplementary knowledge points not only meet the requirements of the medical profession, but also are closer to the actual situation. In this way, the possible errors or misleading information in the model-generated content is effectively reduced.
[0058] Step S13: Input the current medical condition information and the supplementary knowledge points into the target medical big language model; wherein the target medical big language model is acquired based on the historical case data and historical reasoning paths of a specific medical examination institution, and the specific medical examination institution is the institution that examines the target object.
[0059] In this embodiment, obtaining a target medical big language model includes: adjusting an initial big language model using historical case data of a specific medical examination institution to obtain an initial medical big language model; training the initial medical big language model based on a historical reasoning path and the historical case data to obtain a target medical big language model; wherein the historical reasoning path includes symptom recognition, etiology analysis, diagnosis hypothesis formation, and treatment plan selection.
[0060] Different hospitals may have significant differences in abnormal symptoms, diagnosis names, etc., and the large language model may lack logic and accuracy when generating medical content. Therefore, in this embodiment, the target medical big language model incorporates the historical case data and historical reasoning paths of a specific medical examination institution, where the specific medical examination institution is an institution that examines the target object, and the historical reasoning path includes symptom recognition, etiology analysis, diagnosis hypothesis formation, and treatment plan selection. In this way, by using the historical medical record data of a specific medical examination institution to train a customized large language model, the target medical big language model is more adaptable to the specific medical environment; further, the target medical big language model is also acquired based on the historical reasoning path, that is, in the process of acquiring the target medical big language model, the model learns the doctor's step-by-step reasoning process according to the historical reasoning path, so as to learn reasoning paths such as symptom recognition, etiology analysis, diagnosis hypothesis formation, and treatment plan selection, simulate the doctor's step-by-step reasoning process, and improve logic and accuracy.
[0061] In this embodiment, the use of historical case data from a specific medical examination institution to adjust the initial large language model to obtain an initial medical large language model includes: collecting initial historical case data from a specific medical examination institution, and preprocessing the initial historical case data to obtain first target historical case data; extracting abnormal symptom data from the first target historical case data to obtain second target historical case data; wherein the abnormal symptom data is symptom data that does not fall within a corresponding preset range; training the initial large language model using the first target historical case data to obtain a trained medical large language model, and fine-tuning the trained medical large language model using the second target historical case data to obtain an initial medical large language model.
[0062] A large amount of proprietary data such as electronic medical records accumulated within a specific medical examination institution (i.e., initial historical case data) is used to build a customized initial medical big language model. These data are not only huge in quantity, but also contain rich medical information, covering multiple departments and multiple disease types within the hospital, providing valuable learning resources for the model. The initial historical case data is preprocessed to obtain the first target historical case data, and abnormal symptom data is extracted from the first target historical case data to obtain the second target historical case data, that is, symptom data that does not belong to the corresponding preset range is extracted, such as sinus bradycardia, fatty liver, and test indicators deviating from the normal range. The second target historical case data is used to fine-tune the trained medical big language model to obtain the initial medical big language model, as shown in Table 1:
[0063] Table 1
[0064]
[0065] In this embodiment, the collection of initial historical case data of a specific medical examination institution includes: collecting original inpatient case data of a specific medical examination institution; extracting historical medical condition information and historical medical advice from the original inpatient case data; and constructing initial historical case data with the historical medical condition information as model input and the historical medical advice as model output.
[0066] 10,000 original inpatient case data from a specific medical examination institution were collected, covering information from multiple departments and various types of patients. Historical medical condition data and historical medical advice were extracted from each original inpatient case data. Specifically, historical medical condition data can be various examination reports of patients during hospitalization, such as test reports, pathology reports, admission records, etc. Historical medical advice is specifically the "discharge doctor's instructions" in the "discharge summary". In this way, the initial historical case data with historical medical condition data as model input and historical medical advice as model output was constructed. Multiple experts annotated the historical medical condition data to obtain accurate real labels, that is, the model input and model output in the question-answer pair, so as to verify the accuracy of different models. The specific details are shown in Table 2:
[0067] Table 2
[0068]
[0069] The initial historical case data is preprocessed. The specific process may be: desensitizing the initial historical case data to obtain desensitized data to ensure that the privacy and security of patients are not affected; removing low-quality medical data in the desensitized data to obtain the first target historical case data. In order to improve the quality of the data, the data is finely cleaned, including removing noise, duplicate samples, and irrelevant medical information, to ensure the quality of the data, wherein low-quality medical data includes noise data, redundant data, and invalid medical data.
[0070] In this embodiment, the use of the second target historical case data to fine-tune the trained medical big language model to obtain an initial medical big language model includes: taking medical logic rules and preset expert knowledge as constraints, and using the second target historical case data and a multi-task learning framework to fine-tune the trained medical big language model to obtain an initial medical big language model.
[0071] Taking medical logic rules and preset expert knowledge as constraints, and using the second target historical case data and multi-task learning framework to fine-tune the trained medical language model, the initial medical language model obtained is an abnormal symptom recognition and diagnosis model covering all departments and all types. In addition to using the second target historical case data, medical logic rules and preset expert knowledge are also introduced as constraints to standardize the model output and ensure the scientificity and reliability of the symptom descriptions and diagnostic results generated. Taking bacterial infection as an example, when any two indicators of the physical examination items exceed the normal range, bacterial infection can be determined, as shown in Table 3:
[0072] Table 3
[0073]
[0074] Existing large language models have problems with insufficient logic and accuracy when generating medical content. In this embodiment, the Medical Chain-of-Thought (MCoT) is introduced. The Medical Chain-of-Thought imitates the thinking mode of human doctors when facing medical problems, that is, problem analysis and solution are carried out through a series of orderly reasoning paths. By introducing MCoT, the model can be guided to follow a step-by-step decomposition and reasoning strategy in the process of generating medical information, thereby effectively improving the rationality and credibility of the model output. First, the basic framework of MCoT is clarified, including but not limited to key steps such as symptom recognition, etiology analysis, diagnostic hypothesis formation, and treatment plan selection. Each step needs to set clear goals and evaluation criteria to ensure that the model can accurately understand and perform the corresponding reasoning tasks. Secondly, taking the follow-up recommendation generation task in medical advice as an example, the "basic medical thinking chain" is the structuring of the original report, while the "medical diagnosis thinking chain" is based on the abnormalities of each organ (part) after structuring to generate corresponding follow-up recommendations and merge the follow-up recommendations. Next, the initial medical big language model is trained to adapt to MCoT. The model is specially trained using a data set containing rich medical cases and reasoning paths. The historical reasoning path includes but is not limited to symptom recognition, etiology analysis, diagnostic hypothesis formation, and treatment plan selection. During the training process, the focus is on allowing the initial medical big language model to learn how to decompose complex medical problems into several simple small problems, and solve these problems in turn, and finally draw reasonable conclusions. In this way, the target medical big language model is obtained; among them, the data set example of the historical reasoning path is shown in Table 4:
[0075] Table 4
[0076]
[0077] After the training of the target medical language model is completed, a series of designed test cases are used to evaluate the performance of the target medical language model, especially its improvement in logical reasoning ability and answer accuracy. Based on the evaluation results, the model parameters are further tuned to enhance its reliability and accuracy in processing medical information. The evaluation is mainly carried out from the following three dimensions:
[0078] 1) Model accuracy: Compare the performance of the target medical language model with the baseline model in generating treatment recommendations, and evaluate their effectiveness in reducing hallucinations. Specifically, compare the treatment recommendations generated by the model with the medical recommendations of actual doctors, and calculate indicators such as Precision, Recall, and F1-Score (the harmonic mean of precision and recall);
[0079] 2) Logic of the model: Analyze the reasoning process of the model to generate medical content and evaluate the contribution of the "medical thinking chain" to improving logical reasoning ability. The experiment will focus on observing whether the model can reasonably derive correct medical advice according to medical logic, and use manual verification to verify whether its reasoning process is accurate;
[0080] 3) Factualness of the model: By detecting the factuality of the content generated by the model, the effectiveness of combining medical knowledge retrieval enhancement technology in reducing hallucinations is evaluated. Specifically, the experiment conducts a factual evaluation of the suggestions generated by the model. Multiple medical experts can score the generated results on a scale of 1 to 10 to verify the effectiveness of the traceable generation scheme based on retrieval enhancement technology.
[0081] Step S14: using the target medical large language model to deduce a thought chain using the current medical condition data and the supplementary knowledge points, so as to output a target medical suggestion after eliminating the hallucination information.
[0082] In this embodiment, the target medical big language model uses the current medical condition data and the supplementary knowledge points to perform a chain of thought deduction to output a target medical recommendation after the hallucination information is eliminated, including: the target medical big language model uses the current medical condition data and the supplementary knowledge points to perform a chain of thought deduction to generate tasks corresponding to each sub-reasoning path, and combines the current medical condition data and the supplementary knowledge points to generate each sub-conclusion corresponding to each sub-reasoning path task, so as to output the target medical recommendation after the hallucination information is eliminated based on the each sub-conclusion.
[0083] The current medical condition data and supplementary knowledge points are input into the target medical big language model. Because the medical thinking chain is introduced in the acquisition process of the target medical big language model, the target medical big language model learns how to decompose complex medical problems into several simple small problems and solve these problems in sequence. In other words, the target medical big language model can use the current medical condition data and supplementary knowledge points to deduce the thinking chain to generate tasks corresponding to each sub-reasoning path, and combine the current medical condition data and supplementary knowledge points to generate each sub-conclusion corresponding to each sub-reasoning path task, so as to output the target medical advice after the illusion information is eliminated based on each sub-conclusion. The generation method based on the medical reasoning chain enhances the rigor and logic of the generated content by simulating medical logical reasoning, eliminates illusion information, helps reduce the risk of misdiagnosis and missed diagnosis, and improves the accuracy and safety of clinical decision-making.
[0084] The beneficial effects of the present application are as follows: the present application searches a preset medical knowledge base according to the current medical condition information of the target object to obtain related knowledge points; integrates the related knowledge points into a text generation algorithm to obtain supplementary knowledge points; inputs the current medical condition information and the supplementary knowledge points into a target medical big language model; wherein the target medical big language model is acquired based on historical case data and historical reasoning paths of a specific medical examination institution, and the specific medical examination institution is an institution that examines the target object; the current medical condition information and the supplementary knowledge points are used by the target medical big language model to deduce a thinking chain to output a target medical recommendation after the illusion information is eliminated. It can be seen that the present application combines the associated knowledge points in the preset medical knowledge base to realize real-time retrieval of relevant medical information and ensure the scientificity and authenticity of the generated content; the target medical big language model is acquired based on the historical case data and historical reasoning paths of a specific medical examination institution, that is, by using the historical medical record data of a specific medical examination institution to train a customized large language model, the target medical big language model is more adaptable to the specific medical environment; further, the target medical big language model is also acquired based on the historical reasoning path, that is, in the process of acquiring the target medical big language model, the model learns the doctor's step-by-step reasoning process according to the historical reasoning path, so in actual application, the target medical big language model can use the current medical condition information and supplementary knowledge points to deduce the thinking chain, simulate the doctor's step-by-step reasoning process, enhance the logic of the generated content, and effectively reduce the "hallucination" phenomenon in the medical big language model. The accuracy and reliability of the generated medical advice are significantly improved, which provides more reliable technical support for the application of the medical big language model in clinical decision-making.
[0085] Below Figure 2 Taking a specific schematic diagram of obtaining medical advice based on large language model hallucination elimination as an example, the present application is explained accordingly.
[0086] The training process of the large language model requires the use of historical case data and historical reasoning paths of specific medical examination institutions, including, for example, Figure 3 A specific model training diagram based on historical case data is shown, and the process is as follows:
[0087] 1) Private data collation: Collect the original inpatient case data of specific medical examination institutions, and perform pre-processing operations such as data desensitization, data cleaning, quality inspection, and format unification on the data, and finally obtain the first target historical case data;
[0088] 2) Model post-pre-training: Use the open-source big language model as the initial big language model, and use the first target historical case data to train the initial big language model to obtain a trained medical big language model. Furthermore, the performance of the trained medical big language model can be evaluated to adjust parameters and optimize hyperparameters of the trained medical big language model.
[0089] 3) Model fine-tuning: extract abnormal symptom data from the first target historical case data to obtain second target historical case data, use the first target historical case data to train the initial big language model to obtain a trained medical big language model, use medical logic rules and preset expert knowledge as constraints, and use the second target historical case data and a multi-task learning framework to fine-tune the trained medical big language model to obtain an initial medical big language model. It can be understood that the performance of the initial medical big language model can also be evaluated to adjust parameters and optimize hyperparameters of the initial medical big language model.
[0090] Furthermore, after obtaining the initial medical language model, it is necessary to train the initial medical language model based on the historical reasoning path and historical case data to obtain the target medical language model, for example Figure 4 The figure shows a specific medical thinking chain learning diagram. The medical thinking chain can be specifically divided into the medical basic thinking chain and the medical diagnosis thinking chain. The generation of medical suggestions is divided into multiple sub-reasoning path tasks, thereby generating sub-conclusions corresponding to each task, thereby learning to simulate medical logical reasoning and enhancing the rigor and logic of the generated content.
[0091] For example Figure 5 A specific schematic diagram of building a preset medical knowledge base is shown. Before actually obtaining medical advice, a preset medical knowledge base needs to be built. The content extraction system can obtain medical information from medical classics. Medical classics specifically include medical books, electronic document libraries, clinical guidelines, and expert consensus. Among them, the content extraction system can perform layout analysis, structured analysis, optical character recognition, data cleaning and other operations on medical classics to obtain medical information. The architecture of the preset medical knowledge base is a hybrid architecture of retrieval and generation. The vector index of the preset medical knowledge base is constructed based on a semantic vector model, and medical information, medical knowledge graphs, etc. are included in the preset medical knowledge base, so that the preset medical knowledge base includes disease diagnosis standards, treatment methods, drug information and other contents, providing high-quality data support for the subsequent generation process.
[0092] like Figure 2As shown, in the actual process of obtaining medical advice, the preset medical knowledge base is searched according to the current medical condition information of the target object to obtain related knowledge points; the related knowledge points are integrated into the text generation algorithm to obtain supplementary knowledge points; the current medical condition information and the supplementary knowledge points are input into the target medical big language model; the current medical condition information and the supplementary knowledge points are used through the target medical big language model to deduce the thinking chain to output the target medical advice after the hallucination information is eliminated.
[0093] See also Figure 6 As shown, the embodiment of the present application discloses a medical advice acquisition device based on large language model hallucination elimination, comprising:
[0094] The knowledge base search module 11 is used to search the preset medical knowledge base according to the current medical condition information of the target object to obtain related knowledge points;
[0095] A text generation module 12, used to integrate the related knowledge points into a text generation algorithm to obtain supplementary knowledge points;
[0096] An information input module 13 is used to input the current medical condition information and the supplementary knowledge points into a target medical language model; wherein the target medical language model is acquired based on historical case data and historical reasoning paths of a specific medical examination institution, and the specific medical examination institution is an institution that examines the target object;
[0097] The medical advice output module 14 is used to use the current medical condition data and the supplementary knowledge points to perform thought chain deduction through the target medical large language model to output the target medical advice after the hallucination information is eliminated.
[0098] The beneficial effects of the present application are as follows: the present application searches a preset medical knowledge base according to the current medical condition information of the target object to obtain related knowledge points; integrates the related knowledge points into a text generation algorithm to obtain supplementary knowledge points; inputs the current medical condition information and the supplementary knowledge points into a target medical big language model; wherein the target medical big language model is acquired based on historical case data and historical reasoning paths of a specific medical examination institution, and the specific medical examination institution is an institution that examines the target object; the current medical condition information and the supplementary knowledge points are used by the target medical big language model to deduce a thinking chain to output a target medical recommendation after the illusion information is eliminated. It can be seen that the present application combines the associated knowledge points in the preset medical knowledge base to realize real-time retrieval of relevant medical information and ensure the scientificity and authenticity of the generated content; the target medical big language model is acquired based on the historical case data and historical reasoning paths of a specific medical examination institution, that is, by using the historical medical record data of a specific medical examination institution to train a customized large language model, the target medical big language model is more adaptable to the specific medical environment; further, the target medical big language model is also acquired based on the historical reasoning path, that is, in the process of acquiring the target medical big language model, the model learns the doctor's step-by-step reasoning process according to the historical reasoning path, so in actual application, the target medical big language model can use the current medical condition information and supplementary knowledge points to deduce the thinking chain, simulate the doctor's step-by-step reasoning process, enhance the logic of the generated content, and effectively reduce the "hallucination" phenomenon in the medical big language model. The accuracy and reliability of the generated medical advice are significantly improved, which provides more reliable technical support for the application of the medical big language model in clinical decision-making.
[0099] Furthermore, an embodiment of the present application also provides an electronic device. Figure 7 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.
[0100] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the medical advice acquisition method based on large language model hallucination elimination performed by the electronic device disclosed in any of the aforementioned embodiments.
[0101] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0102] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0103] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system 221, a computer program 222 and data 223, etc. The storage method can be temporary storage or permanent storage.
[0104] Among them, the operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device to realize the operation and processing of the massive data 223 in the memory 22 by the processor 21, which can be Windows, Unix, Linux, etc. In addition to including a computer program that can be used to complete the medical advice acquisition method based on large language model illusion elimination performed by the electronic device disclosed in any of the aforementioned embodiments, the computer program 222 can also further include a computer program that can be used to complete other specific tasks. In addition to data transmitted from an external device received by the electronic device, the data 223 can also include data collected by its own input and output interface 25.
[0105] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the medical advice acquisition method based on large language model hallucination elimination disclosed above is implemented. For the specific steps of the method, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, and no further description will be given here.
[0106] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0107] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. The steps of the method or algorithm described in conjunction with the embodiments disclosed herein can be implemented directly with hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable EPROM (Erasable Programmable Read Only Memory), electrically erasable programmable EEPROM (Electrically Erasable Programmable read only memory), register, hard disk, removable disk, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium known in the technical field.
[0108] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0109] The above is a detailed introduction to the method, device, equipment and medium for obtaining medical advice based on large language model hallucination elimination provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A method for obtaining medical advice based on large language model hallucination elimination, characterized in that: include: Searching the preset medical knowledge base according to the current medical condition information of the target object to obtain related knowledge points; Integrating the related knowledge points into the text generation algorithm to obtain supplementary knowledge points; Inputting the current medical condition data and the supplementary knowledge points into a target medical big language model; wherein the target medical big language model is acquired based on historical case data and historical reasoning paths of a specific medical examination institution, and the specific medical examination institution is an institution that examines the target object; The target medical big language model uses the current medical condition data and the supplementary knowledge points to deduce a thought chain, so as to output a target medical suggestion after the hallucination information is eliminated.
2. The method for obtaining medical advice based on large language model hallucination elimination according to claim 1, characterized in that: Obtain the target medical language model, including: The initial large language model is adjusted using historical case data of a specific medical examination institution to obtain an initial medical large language model; The initial medical big language model is trained based on the historical reasoning path and the historical case data to obtain a target medical big language model; wherein the historical reasoning path includes symptom recognition, etiology analysis, diagnosis hypothesis formation, and treatment plan selection.
3. The method for obtaining medical advice based on large language model hallucination elimination according to claim 2, characterized in that: The adjusting the initial large language model by using the historical case data of a specific medical examination institution to obtain the initial medical large language model includes: Collecting initial historical case data of a specific medical examination institution, and preprocessing the initial historical case data to obtain first target historical case data; Extracting abnormal symptom data from the first target historical case data to obtain second target historical case data; wherein the abnormal symptom data is symptom data that does not belong to the corresponding preset range; The initial large language model is trained using the first target historical case data to obtain a trained medical large language model, and the trained medical large language model is fine-tuned using the second target historical case data to obtain an initial medical large language model.
4. The method for obtaining medical advice based on large language model hallucination elimination according to claim 3, characterized in that: The initial historical case data collected from a specific medical examination institution includes: Collect original inpatient case data from specific medical examination institutions; Extracting historical medical condition information and historical medical advice from the original inpatient case data; Initial historical case data is constructed with the historical medical condition data as model input and the historical medical advice as model output.
5. The method for obtaining medical advice based on large language model hallucination elimination according to claim 3, characterized in that: The fine-tuning of the trained medical language model using the second target historical case data to obtain an initial medical language model includes: Taking medical logic rules and preset expert knowledge as constraints, and using the second target historical case data and a multi-task learning framework, the trained medical big language model is fine-tuned to obtain an initial medical big language model.
6. The method for obtaining medical advice based on large language model hallucination elimination according to claim 2, characterized in that: The target medical big language model uses the current medical condition data and the supplementary knowledge points to perform thought chain deduction to output target medical suggestions after the illusion information is eliminated, including: The target medical big language model utilizes the current medical condition data and the supplementary knowledge points to perform thought chain deduction to generate tasks corresponding to each sub-reasoning path, and combines the current medical condition data and the supplementary knowledge points to generate sub-conclusions corresponding to each sub-reasoning path task, so as to output target medical recommendations after the hallucination information is eliminated based on the sub-conclusions.
7. The method for obtaining medical advice based on large language model hallucination elimination according to any one of claims 1 to 5, characterized in that: The architecture of the preset medical knowledge base is a hybrid architecture of search and generation, and the vector index of the preset medical knowledge base is constructed based on a semantic vector model; Accordingly, the current medical condition data of the target object is searched in a preset medical knowledge base to obtain related knowledge points, including: The current medical condition data of the target object is compared with the vector index of each knowledge point in the preset medical knowledge base to obtain the related knowledge points.
8. A medical advice acquisition device based on large language model hallucination elimination, characterized in that: include: A knowledge base retrieval module is used to search the preset medical knowledge base according to the current medical condition information of the target object to obtain related knowledge points; A text generation module, used to integrate the related knowledge points into a text generation algorithm to obtain supplementary knowledge points; An information input module, used to input the current medical condition data and the supplementary knowledge points into a target medical big language model; wherein the target medical big language model is acquired based on historical case data and historical reasoning paths of a specific medical examination institution, and the specific medical examination institution is an institution that examines the target object; The medical advice output module is used to use the current medical condition data and the supplementary knowledge points to perform thought chain deduction through the target medical large language model to output the target medical advice after the hallucination information is eliminated.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is used to execute the computer program to implement the steps of the method for obtaining medical advice based on large language model hallucination elimination as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein, when the computer program is executed by a processor, the steps of the method for obtaining medical advice based on large language model hallucination elimination are implemented as described in any one of claims 1 to 7.