Liver disease diagnosis and treatment auxiliary system and method based on large language model

By constructing a liver disease-specific dataset and iteratively training the model, the data bias problem of general-purpose large language models in the field of liver disease was solved, improving the accuracy and professionalism of auxiliary analysis for liver disease diagnosis and treatment, and ensuring the accuracy and security of the output.

CN121768630APending Publication Date: 2026-03-31BEIJING LIUYUAN SPACE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing general-purpose large language models in the field of liver disease suffer from insufficient information accuracy and specialty adaptability, and lack of professional data processing, resulting in low accuracy of auxiliary analysis.

Method used

Collect multi-source liver disease data, preprocess it to generate a task-driven liver disease-specific dataset, iteratively train a basic large language model to generate a liver disease-specific large language model, derive preliminary diagnostic and treatment assistance results through thought chain deduction, and perform preset output verification to obtain the target diagnostic and treatment assistance results.

Benefits of technology

It improves the accuracy and professionalism of auxiliary analysis for liver disease diagnosis and treatment, ensures that the output content meets the needs of liver disease diagnosis and treatment, enhances the interpretability of decisions, and guarantees the accuracy and security of the output through a dual verification mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768630A_ABST
    Figure CN121768630A_ABST
Patent Text Reader

Abstract

The invention discloses a liver disease diagnosis and treatment auxiliary system and method based on a large language model, and relates to the technical field of medical data processing, the method comprises the following steps: collecting multi-source liver disease data, and preprocessing the multi-source liver disease data to generate a task-driven special liver disease data set; performing iterative training on a preset basic large language model based on the task-driven special hepatopathy data set to obtain a special hepatopathy large language model; inputting the diagnosis and treatment data of the patient into a special liver disease large language model for thinking chain derivation to generate a preliminary diagnosis and treatment auxiliary result; and performing preset output verification on the preliminary diagnosis and treatment auxiliary result to obtain a target diagnosis and treatment auxiliary result. According to the method, data bias is solved by constructing the special hepatopathy data set, a high-quality training basis is provided for the model, and the output content is ensured to meet the hepatopathy diagnosis and treatment requirements by improving the special adaptation capability of the model. And meanwhile, the decision interpretability is enhanced through thinking chain derivation, and the output accuracy and safety are guaranteed through preset output verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical data processing technology, and in particular to a liver disease diagnosis and treatment auxiliary system and method based on a large language model. Background Technology

[0002] While existing Large Language Models (LLMs) can achieve basic medical record generation and medical knowledge Q&A in the medical field, they have significant shortcomings in liver disease scenarios: First, the training data lacks representativeness for liver disease specialties, making it prone to inaccuracy issues such as fabricated information and outdated treatment plans, such as confusing antiviral drugs for hepatitis B and hepatitis C; Second, data processing is not optimized for liver disease subtypes and disease stages, and its ability to analyze complex disease courses such as the progression from compensated to decompensated cirrhosis is weak.

[0003] Therefore, existing general-purpose large language models suffer from insufficient accuracy and specialization in the field of liver diseases, as well as a lack of professionalism and security in data processing, resulting in low accuracy in auxiliary analysis. Summary of the Invention

[0004] The main objective of this application is to provide a liver disease diagnosis and treatment auxiliary system and method based on a large language model, aiming to solve the technical problem of low accuracy in auxiliary analysis of liver diseases using existing general-purpose large language models.

[0005] To achieve the above objectives, this application proposes a liver disease diagnosis and treatment assistance method based on a large language model, the method comprising: Collect multi-source liver disease data and preprocess the multi-source liver disease data to generate a task-driven liver disease-specific dataset; Based on the task-driven liver disease dataset, the preset basic large language model is iteratively trained to obtain a liver disease-specific large language model. The patient's diagnosis and treatment data are input into the liver disease-specific big language model to perform thought chain deduction and generate preliminary diagnosis and treatment auxiliary results. The preliminary diagnostic and treatment assistance results are verified by a preset output to obtain the target diagnostic and treatment assistance results.

[0006] In one embodiment, the step of collecting multi-source liver disease data and preprocessing the multi-source liver disease data to generate a task-driven liver disease dataset includes: Connect to hospital information systems and laboratory information systems to collect multi-source liver disease data, consisting of electronic medical records, laboratory indicator data, clinical guidelines, research literature, and standardized medical terminology; The multi-source liver disease data was cleaned using a medical data quality rule base to obtain valid liver disease data; The valid liver disease data are then standardized to obtain standardized liver disease data. The standardized liver disease data is correlated with corresponding diagnostic and treatment task parameters to generate a task-driven liver disease dataset consisting of several specialized subset datasets.

[0007] In one embodiment, the step of iteratively training a preset basic large language model based on the task-driven liver disease-specific dataset to obtain a liver disease-specific large language model includes: Based on the task-driven liver disease dataset, the preset basic large language model is supervised and fine-tuned to obtain an optimized large language model. The data output format of the optimized large language model is optimized based on the prompting engineering consisting of step-by-step prompt words to obtain a large language model for liver disease.

[0008] In one embodiment, the step of inputting patient diagnosis and treatment data into the liver disease-specific large language model for thought chain deduction to generate preliminary diagnosis and treatment assistance results includes: A liver disease guideline knowledge graph with a ternary structure is constructed based on pre-set clinical guidelines for liver diseases. The liver disease guideline knowledge graph includes guideline ternaries with elements of applicable conditions, core evidence, and recommendation level. Using the aforementioned liver disease-specific big language model, and based on the patient's diagnosis and treatment data and the liver disease guideline knowledge graph, a thought chain deduction is performed to generate preliminary diagnostic and treatment assistance results.

[0009] In one embodiment, the step of generating preliminary diagnostic and treatment assistance results by performing thought chain deduction based on the patient's medical data and the liver disease guideline knowledge graph includes: The patient's medical data is broken down into a set of medical elements; The set of diagnostic and treatment elements is subjected to a thought chain deduction to obtain a decision subset associated according to a preset thought chain order; the decision subset includes a disease feature subset and an examination subset; Based on the retrieval enhancement mechanism, the set of diagnostic and treatment elements is dynamically matched with the guideline triples in the liver disease guideline knowledge graph, and an interpretable decision chain is generated based on the matched triples, the disease feature subset, and the examination subset. Preliminary diagnostic and treatment assistance results are generated based on the subset of disease characteristics, the subset of examinations, and the interpretable decision chain.

[0010] In one embodiment, the preliminary diagnostic and treatment assistance results include structured electronic medical records, laboratory early warning information, and guideline push content; The step of generating preliminary diagnostic and treatment support results based on the interpretable decision chain, which includes at least structured electronic medical records, laboratory early warning information, and guideline push content, includes: The structured electronic medical record is generated by calling the medical record structure template, filling the disease feature subset and the examination subset into the corresponding fields of the preset electronic medical record; Perform a two-dimensional threshold anomaly analysis on the test indicator data to generate the test early warning information; The matching guidelines corresponding to the explainable decision chain are sorted according to the recommendation level of the explainable decision chain, and the sorted matching guidelines and the corresponding explainable decision chains are associated and output as the guide push content.

[0011] In one embodiment, the step of performing a preset output verification on the preliminary diagnostic and treatment assistance result to obtain the target diagnostic and treatment assistance result includes: The local liver disease knowledge base and the preliminary diagnostic and treatment assistance results are aggregated and compared to obtain the diagnostic and treatment element matching deviation rate. If the detection shows that the matching deviation rate of the diagnosis and treatment elements exceeds a preset threshold, an expert review process is triggered, and manual review opinions based on the feedback from the expert review process are obtained. The preliminary diagnostic and treatment assistance results are optimized based on the manual review opinions to obtain the target diagnostic and treatment assistance results.

[0012] Furthermore, to achieve the above objectives, this application also proposes a liver disease diagnosis and treatment assistance system based on a large language model, the system comprising: The data processing module is used to collect multi-source liver disease data and preprocess the multi-source liver disease data to generate a task-driven liver disease-specific dataset. The model training module is used to iteratively train the preset basic large language model based on the task-driven liver disease dataset to obtain a liver disease-specific large language model. The reasoning generation module is used to input patient diagnosis and treatment data into the liver disease-specific big language model to perform thought chain reasoning and generate preliminary diagnosis and treatment assistance results; The security verification module is used to perform preset output verification on the preliminary diagnostic and treatment assistance results to obtain the target diagnostic and treatment assistance results.

[0013] Furthermore, to achieve the above objectives, this application also proposes a liver disease diagnosis and treatment auxiliary device based on a large language model. The device includes: a memory, a processor, and a liver disease diagnosis and treatment auxiliary program based on a large language model stored in the memory and capable of running on the processor. The liver disease diagnosis and treatment auxiliary program based on a large language model is configured to implement the steps of the liver disease diagnosis and treatment auxiliary method based on a large language model as described above.

[0014] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium, on which a program implementing a liver disease diagnosis and treatment assistance method based on a large language model is stored. The program implementing the liver disease diagnosis and treatment assistance method based on a large language model is executed by a processor to implement the steps of the liver disease diagnosis and treatment assistance method based on a large language model as described above.

[0015] This application provides a liver disease diagnosis and treatment assistance system and method based on a large language model. The method includes: collecting multi-source liver disease data and preprocessing the data to generate a task-driven liver disease-specific dataset; iteratively training a pre-set basic large language model based on the task-driven liver disease-specific dataset to obtain a liver disease-specific large language model; inputting patient diagnosis and treatment data into the liver disease-specific large language model for thought chain deduction to generate preliminary diagnosis and treatment assistance results; and performing pre-set output verification on the preliminary diagnosis and treatment assistance results to obtain the target diagnosis and treatment assistance results. This application first addresses data bias by constructing a liver disease-specific dataset, providing a high-quality training foundation for the model. Then, iterative training with professional and rich data enhances the model's specialty adaptability, ensuring that the output content aligns with the needs of liver disease diagnosis and treatment. Simultaneously, this application enhances decision interpretability through thought chain deduction and ensures output accuracy and security through pre-set output verification. Therefore, this application can effectively improve the accuracy of liver disease diagnosis and treatment assistance analysis, providing reliable diagnostic and treatment support for clinical diagnosis of liver diseases. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a first flowchart illustrating the first embodiment of the liver disease diagnosis and treatment assistance method based on a large language model according to this application. Figure 2 This is a second flowchart illustrating the first embodiment of the liver disease diagnosis and treatment assistance method based on a large language model of this application. Figure 3 This is a schematic diagram of the first process of the second embodiment of the auxiliary method for liver disease diagnosis and treatment based on a large language model in this application; Figure 4 This is a second flowchart illustrating a second embodiment of the liver disease diagnosis and treatment assistance method based on a large language model according to this application; Figure 5 This is a schematic diagram of the third process of the second embodiment of the auxiliary method for liver disease diagnosis and treatment based on a large language model in this application; Figure 6 This is a schematic diagram of the module structure of the liver disease diagnosis and treatment auxiliary system based on a large language model, as described in an embodiment of this application. Figure 7 This is a schematic diagram of the hardware operating environment of the liver disease diagnosis and treatment assistance method based on a large language model in the embodiments of this application.

[0019] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0022] The main solution of this application is as follows: collect multi-source liver disease data, preprocess the multi-source liver disease data to generate a task-driven liver disease-specific dataset; iteratively train a pre-set basic large language model based on the task-driven liver disease-specific dataset to obtain a liver disease-specific large language model; input patient diagnosis and treatment data into the liver disease-specific large language model to perform thought chain deduction and generate preliminary diagnosis and treatment assistance results; perform pre-set output verification on the preliminary diagnosis and treatment assistance results to obtain the target diagnosis and treatment assistance results.

[0023] While existing general-purpose large language models can achieve functions such as basic medical record generation and medical knowledge Q&A in the medical field, they have significant shortcomings in the context of hepatology: First, the training data lacks representativeness for hepatology, making it prone to inaccuracies such as fabricated information and outdated treatment plans, such as confusing antiviral drugs for hepatitis B and hepatitis C; Second, data processing is not optimized for hepatology subtypes and disease stages, resulting in weak analytical capabilities for complex disease progressions such as the progression from compensated to decompensated cirrhosis; Third, there is a lack of mechanisms for dynamically integrating the latest hepatology guidelines, with recommended treatments often lagging behind clinical norms; Fourth, the decision-making process is "black box" in nature, with insufficient interpretability leading to ambiguity in the definition of clinical responsibility; Fifth, the mechanisms for protecting medical data privacy and ensuring compliance are inadequate, posing a risk of sensitive information leakage.

[0024] To address the aforementioned issues, this application constructs a four-layer architecture: "Data Layer - Model Layer - Application Layer - Security Layer." Through the construction of a disease-specific data system, specialized model optimization, end-to-end clinical function adaptation, and security compliance mechanisms, it achieves professional, efficient, and safe assistance in liver disease diagnosis and treatment. First, this application addresses data bias by constructing a disease-specific dataset, providing a high-quality training foundation for the model. It enhances the model's specialty adaptability through multi-strategy iterative training, ensuring that the output aligns with the needs of liver disease diagnosis and treatment. It strengthens decision interpretability through thought chain derivation. Finally, it ensures the accuracy and security of the output through a dual-verification mechanism, providing reliable diagnostic and treatment support for clinical liver disease diagnosis. This effectively improves the accuracy and professionalism of liver disease diagnosis and treatment assistance analysis based on a large language model.

[0025] It should be noted that the executing entity in this embodiment can be a liver disease diagnosis and treatment auxiliary system based on a large language model, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a liver disease diagnosis and treatment auxiliary device based on a large language model capable of achieving the above functions. This embodiment does not specifically limit it in this way. The following uses a liver disease diagnosis and treatment auxiliary device based on a large language model (hereinafter referred to as the auxiliary device) as the executing entity as an example to describe this embodiment and the following embodiments.

[0026] Based on this, embodiments of this application provide a liver disease diagnosis and treatment assistance method based on a large language model, referring to... Figure 1 , Figure 1 This is a first flowchart illustrating the first embodiment of the liver disease diagnosis and treatment assistance method based on a large language model according to this application.

[0027] In this embodiment, the liver disease diagnosis and treatment assistance method based on a large language model includes steps S10 to S40: Step S10: Collect multi-source liver disease data and preprocess the multi-source liver disease data to generate a task-driven liver disease dataset. Understandably, the aforementioned multi-source liver disease data can be a collection of data related to the diagnosis and treatment of liver diseases from multiple channels, including but not limited to electronic medical records in hospital information systems (HIS) (covering chief complaints, present medical history, physical examinations, etc. for subtypes of viral hepatitis, cirrhosis, and liver cancer), laboratory indicator data (liver function, viral markers, tumor markers, etc.) in laboratory information systems (LIS), authoritative domestic and international clinical guidelines (such as the "Guidelines for the Prevention and Treatment of Chronic Hepatitis B"), research literature (such as papers in the field of liver diseases in databases such as PubMed and CNKI), and standardized medical terminology (such as ICD-11 (International Statistical Classification of Diseases and Related Health Problems-11th Revision) or SNOMED CT (Systematized Nomenclature of Medicine-Clinical Terms) codes, etc.). Therefore, before conducting diagnostic and treatment-aided analysis, the auxiliary equipment can perform a series of standardized processing operations on multi-source liver disease data, including data cleaning (removing invalid records and repairing outliers), desensitization (encrypting sensitive patient information), structure conversion (converting unstructured text into a format that conforms to medical data exchange standards), and data association (associating historical and real-time data according to diagnostic and treatment tasks), etc., to improve the accuracy of model analysis.

[0028] The aforementioned task-driven liver disease datasets can be specialized datasets built for specific liver disease diagnosis and treatment tasks (such as generating admission records, writing discharge summaries, recommending treatment plans, etc.). They can contain several specialized sub-datasets, and each sub-dataset is associated with all-dimensional data required to complete the corresponding task. For example, the admission record dataset needs to be associated with the patient's previous admission / discharge records and new information from the current visit.

[0029] Therefore, in one feasible implementation, refer to Figure 2 , Figure 2 This is a second flowchart illustrating the first embodiment of the liver disease diagnosis and treatment assistance method based on a large language model according to this application. In this embodiment, step S10 may include steps A1 to A4: Step A1: Connect to the hospital information system and laboratory information system to collect multi-source liver disease data consisting of electronic medical records, laboratory indicator data, clinical guidelines, research literature, and standardized medical terminology; Step A2: Clean the multi-source liver disease data using a medical data quality rule base to obtain valid liver disease data; Step A3: Standardize the valid liver disease data to obtain standardized liver disease data; Step A4: Perform data association on the standardized liver disease data for corresponding diagnosis and treatment task parameters to generate a task-driven liver disease dataset consisting of several specialized subset datasets.

[0030] It is easy to understand that the aforementioned hospital information system can be the core system used by a hospital to store and manage basic patient information, medical records, inpatient information, etc. In this embodiment, it is connected through a standardized interface to collect electronic medical record data. The laboratory information system (LIS) can be a system used to manage clinical laboratory data, which can obtain test results such as liver function, viral markers, and tumor markers. Therefore, assistive devices can connect to the HIS system via the HL7FHIR (Health Level Seven International-Fast Healthcare Interoperable Resources) standardized interface to collect patients' electronic medical records (including chief complaints, present illness, past medical history, etc.); connect to the LIS system to obtain liver function (ALT (Alanine Aminotransferase), AST (Aspartate Aminotransferase), etc.), virology (HBV DNA (Hepatitis B Virus Deoxyribonucleic Acid), etc.), tumor markers (AFP (Alpha-Fetoprotein), etc.) and other test indicator data; collect clinical guidelines from the WHO (World Health Organization) website and the database of the Chinese Society of Hepatology; download research literature in the field of liver disease from the PubMed platform and CNKI; obtain standardized terminology such as ICD-11 and SNOMED CT from the World Health Organization database, and summarize them into the above multi-source liver disease data.

[0031] The aforementioned medical data quality rule base can be a pre-defined set of rules for filtering high-quality data. It can include rules for field completeness (such as "test data must include the test date and result value") and rules for data reasonableness (such as "ALT normal range 0-40 U / L (Unit / Liter)"). Therefore, during the data cleaning process, the auxiliary equipment can use the medical data quality rule base to remove invalid data and repair abnormal data from multi-source liver disease data, achieving data denoising and standardization to obtain valid liver disease data that meets quality rules and has no key missing or obvious abnormalities. For example, by calling the medical data quality rule base, records missing patient IDs (identification) and test dates can be automatically removed; data with ALT, AST, and other indicators exceeding the normal range by more than 10 times can be marked as abnormal and trigger review and repair by laboratory physicians; and duplicate medical record data can be deleted to ensure data uniqueness.

[0032] Understandably, the aforementioned standardization process can refer to the process of structuring, de-identifying, and / or encrypting valid liver disease data. For example, natural language processing (NLP) technology can be used to convert free-text medical records into JSON format, extracting structured fields such as "chief complaint" and "present illness"; standardizing the mapping of terms like "palmar erythema" and "spider angioma" using SNOMED CT encoding; and unifying the units of laboratory indicators to international standard units (e.g., ALT units to U / L). The aforementioned diagnostic and treatment task parameters can be characteristic parameters that distinguish different diagnostic and treatment tasks. For example, the admission record task requires association with the "historical medical history" parameter, while the discharge record task requires association with the "entire process of this diagnosis and treatment" parameter. Therefore, auxiliary equipment can link and integrate standardized data with relevant historical data and auxiliary data according to the diagnostic and treatment task parameters, i.e., perform the aforementioned data association to obtain datasets for a single diagnostic and treatment task, such as admission record datasets, discharge record datasets, and treatment plan recommendation datasets, forming the aforementioned specialized sub-datasets. For example, in the process of data association and subset generation, for the admission record task, the patient's previous admission / discharge record can be associated with the new information of this visit to generate an admission record subset; for the discharge record task, the information of the entire process of this hospitalization, such as the initial assessment, treatment intervention, and discharge summary, can be integrated to generate a discharge record subset; similarly, other specialized subsets such as laboratory indicator analysis and guideline push can be constructed to jointly form a task-driven liver disease dataset.

[0033] Therefore, existing technologies address the problems of multi-source heterogeneity, inconsistent quality, privacy sensitivity, and lack of task specificity in liver disease data, leading to biased model training data, high risk of privacy leakage, and low training efficiency. This embodiment integrates multi-source data to cover all dimensions of liver disease diagnosis and treatment; improves data quality through cleaning and standardization, laying the foundation for model training; and ensures data privacy and security while adapting the dataset to specific diagnostic and treatment tasks through de-identification and task-specific association, thereby improving the relevance and efficiency of model training.

[0034] Step S20: Iteratively train the preset basic large language model based on the task-driven liver disease dataset to obtain a liver disease specific large language model; Step S30: Input the patient's diagnosis and treatment data into the liver disease-specific big language model to perform thought chain deduction and generate preliminary diagnosis and treatment assistance results; Step S40: Perform preset output verification on the preliminary diagnostic and treatment assistance results to obtain the target diagnostic and treatment assistance results.

[0035] It is important to understand that the aforementioned pre-set basic large language model can be a pre-trained large language model with long context processing capabilities and adapted to medical texts. In this embodiment, the Qwen2-7B model is preferred as the pre-set basic large language model, as it can adapt to the needs of ultra-long medical records and multimodal data processing. The aforementioned iterative training process can refer to a multi-round, multi-strategy model optimization process, including supervised fine-tuning (SFT), direct preference optimization (DPO), and prompt word engineering optimization, thereby gradually adapting the basic model to hepatology-specific tasks. Therefore, the aforementioned hepatology-specific large language model can be a large language model that, after training and strategy optimization with hepatology-specific data, possesses capabilities such as understanding hepatology-specific terminology, diagnostic and treatment logic reasoning, and generating structured content.

[0036] It is easy to understand that the aforementioned patient diagnosis and treatment data can be real-time data generated during the diagnosis and treatment process, including chief complaint, present medical history, physical examination results, laboratory test reports, and imaging examination conclusions. The reasoning process can be described as a model breaking down liver disease diagnosis and treatment decisions into atomized steps, such as "tumor diameter analysis → vascular invasion assessment → liver function grading → treatment plan matching," making the decision-making logic traceable.

[0037] It should be understood that the aforementioned preliminary diagnostic and treatment assistance results may be unverified diagnostic and treatment assistance content generated by the model through reasoning, which may include structured electronic medical records, laboratory indicator analysis reports, disease warning information, clinical guideline content, and preliminary treatment suggestions, etc.

[0038] The preset output verification can use a dual mechanism of "knowledge base comparison + expert review" to verify the compliance and accuracy of the preliminary results, including calculation of the matching degree of diagnostic and treatment elements and manual review of high-risk content. This results in final diagnostic and treatment assistance content that has been verified and optimized to meet clinical standards and safety requirements, which can be directly pushed to the doctor's terminal for clinical reference.

[0039] Therefore, in a feasible implementation, step S40 may include steps B1 to B3: Step B1: Perform a comparative analysis of diagnostic and treatment elements on the local liver disease knowledge base and the preliminary diagnostic and treatment assistance results to obtain the diagnostic and treatment element matching deviation rate; Step B2: If the detection of the matching deviation rate of the diagnosis and treatment elements exceeds a preset threshold, the expert review process is triggered, and the manual review opinion based on the feedback of the expert review process is obtained. Step B3: Optimize the preliminary diagnostic and treatment assistance results based on the manual review opinions to obtain the target diagnostic and treatment assistance results.

[0040] It is easy to understand that the aforementioned local liver disease knowledge base can be a pre-built database that stores standardized diagnosis and treatment plans for liver diseases, including treatment drugs, dosages, courses of treatment, contraindications, and guidelines for each subtype of liver disease, and is regularly synchronized with the latest guidelines.

[0041] The process of aggregating and comparing diagnostic and treatment elements can be achieved by classifying and aggregating the diagnostic and treatment elements (such as drug selection, dosage, and examination items) in the preliminary results with standard elements in the knowledge base, and calculating the matching degree. Algorithms such as cosine similarity and clinical terminology mapping can be used. Furthermore, the aforementioned diagnostic and treatment element matching deviation rate can be the proportion of the number of elements in the preliminary results that do not match the standard elements in the knowledge base to the total number of elements. For example, if a core drug mismatch or a dosage exceeding the guideline range is detected, it can be included in the deviation rate.

[0042] It is important to understand that the aforementioned preset thresholds can be critical deviation rate values ​​set based on clinical risk, including overall deviation rate thresholds and key element deviation rate thresholds. In this embodiment, the overall deviation rate threshold can be set to 30%, and the key element deviation rate threshold can be set to 10%. The specific values ​​can be dynamically adjusted according to the liver disease subtype; this embodiment does not impose any restrictions on this. The aforementioned expert review process can be a process where a review team composed of chief physicians specializing in liver diseases manually reviews high-deviation results, including result evaluation, deviation cause analysis, and feedback on corrective actions. Correspondingly, the aforementioned manual review opinions can be corrective suggestions given after expert review, such as "replacing drug A with drug B" or "supplementing liver function Child-Pugh classification assessment," etc.

[0043] Therefore, in this embodiment, the auxiliary device can first perform a comparative analysis of diagnostic and treatment elements to break down the preliminary diagnostic and treatment assistance results into a set of diagnostic and treatment elements such as treatment drugs, dosages, examination items, and follow-up frequencies; retrieve the standard set of diagnostic and treatment elements for the corresponding disease (such as chronic hepatitis B) from the local liver disease knowledge base; calculate the element matching degree pair by pair through a semantic matching algorithm, count the number of mismatched elements, and calculate the overall deviation rate and the deviation rate of key elements.

[0044] If the overall deviation rate is ≤30% and the deviation rate of key elements is ≤10%, the process proceeds directly to the next step of result output. If any threshold is exceeded, the auxiliary device can automatically push the preliminary results and deviation analysis report to the expert terminal, triggering the expert review process. This allows experts to combine clinical experience with guideline standards to evaluate the results, mark the deviation points, and provide correction suggestions.

[0045] Then, the auxiliary equipment can automatically correct the deviations in the preliminary results based on the received human review opinions, such as replacing drugs that do not conform to the guidelines or supplementing missing test items; after correction, it is compared with the knowledge base again to ensure that the deviation rate is below the threshold before proceeding to the next step of result output.

[0046] Finally, in the results output stage, the assistive device integrates and corrects the structured medical records, early warning information, and guideline content to generate the target diagnostic and treatment assistance results.

[0047] Therefore, existing general-purpose language models may output results that do not conform to clinical norms, and the lack of effective verification mechanisms to ensure their safety and reliability makes them difficult to apply directly in clinical practice. This embodiment can automate the verification of preliminary diagnostic and treatment assistance results through knowledge base comparison, improving verification efficiency. Expert review allows for manual verification of high-risk content, reducing the clinical risk of erroneous output. Simultaneously, a dual-threshold deviation rate verification mechanism ensures that the target results conform to clinical norms, providing doctors with safe and reliable diagnostic and treatment assistance and reducing the potential for medical disputes.

[0048] This embodiment provides a liver disease diagnosis and treatment assistance method based on a large language model. The method includes: collecting multi-source liver disease data and preprocessing the data to generate a task-driven liver disease-specific dataset; iteratively training a pre-set basic large language model based on the task-driven liver disease-specific dataset to obtain a liver disease-specific large language model; inputting patient diagnosis and treatment data into the liver disease-specific large language model to generate preliminary diagnosis and treatment assistance results through thought chain deduction; and performing pre-set output verification on the preliminary diagnosis and treatment assistance results to obtain the target diagnosis and treatment assistance results. This application addresses data bias by constructing a liver disease-specific dataset, providing a high-quality training foundation for the model, and further ensuring that the output content aligns with liver disease diagnosis and treatment needs by improving the model's specialty adaptability. Simultaneously, thought chain deduction enhances decision interpretability, and pre-set output verification ensures output accuracy and security.

[0049] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the first embodiment described above can be referred to the above description, and will not be repeated hereafter.

[0050] Based on the first embodiment, please refer to Figure 3 , Figure 3 This is a first flowchart illustrating the second embodiment of the liver disease diagnosis and treatment assistance method based on a large language model according to this application. In this embodiment, step S20 includes steps C1 to C2: Step C1: Based on the task-driven liver disease dataset, supervised fine-tuning of the preset basic large language model is performed to obtain an optimized large language model; Step C2: Optimize the data output format of the optimized large language model based on the prompting engineering consisting of step-by-step prompt words to obtain a liver disease-specific large language model.

[0051] It is easy to understand that the aforementioned supervised fine-tuning (SFT) can be a process of targeted training of a pre-trained model based on labeled disease-specific data. By inputting liver disease diagnosis and treatment samples into the model, comparing the differences between the output and the labeled results, and adjusting the parameters, the model is adapted to the specialized task. Optimizing a large language model can be a model that, after SFT fine-tuning, possesses preliminary liver disease specialized knowledge and reasoning ability, but its output format may not conform to clinical document standards.

[0052] At this point, the aforementioned step-by-step prompts can be prompt statements broken down according to the structure of the medical document or the decision-making logic, such as "When generating the admission record, first clearly state the chief complaint (no more than 20 characters), then describe the present illness in chronological order, including the causes of onset, the evolution of symptoms, and previous treatments." Prompt engineering can be a technique that guides the model to generate output content that meets requirements through the design and optimization of prompts. Its core lies in breaking down complex tasks into step-by-step instructions that the model can understand. Therefore, this embodiment can optimize the model's data output format, that is, constrain the structure and content of the model's output through prompts, ensuring that the generated medical records, reports, etc., comply with clinical standards and format requirements.

[0053] For example, during the process of obtaining an optimized model through supervised fine-tuning, the assistive device can first extract labeled samples from a task-driven dataset, such as "input patient basic information + test results → output preliminary diagnosis" or "input current medical history + physical signs → output epidemiological history". The samples are divided into training and validation sets in an 8:2 ratio, and the Qwen2-7B model is trained using SFT. The learning rate, batch size and other parameters are iteratively adjusted until the model achieves a specialist terminology accuracy of over 95% and a diagnosis and treatment logic matching rate of over 90% on the validation set, thereby obtaining an optimized large language model.

[0054] Then, by optimizing the output format through prompting engineering, step-by-step prompts are designed for different clinical document types. For example, when generating discharge records, the prompts could be: "1. Chief complaint: summarize the patient's core symptoms and duration of hospitalization; 2. Treatment process: describe the examinations, treatments, and changes in the condition in chronological order; 3. Discharge diagnosis: clarify the primary and secondary diagnoses; 4. Discharge recommendations: include medication, follow-up, and dietary precautions." The prompts are embedded into the model reasoning process to guide the optimization model to output content in a standardized format, thus obtaining a large language model for liver disease.

[0055] Therefore, this embodiment addresses the problem that general-purpose basic models lack liver disease specialist knowledge, and their output content does not conform to clinical document format standards, thus failing to directly meet the needs of diagnostic and treatment assistance. It utilizes SFT fine-tuning to enable the model to quickly grasp liver disease specialist knowledge and diagnostic and treatment logic, improving output accuracy. Furthermore, by providing hints and engineering constraints on the output format, it ensures that the generated medical records and reports are directly adapted to clinical application scenarios, reducing subsequent modification costs for doctors and improving diagnostic and treatment efficiency.

[0056] In one feasible implementation, refer to Figure 4 , Figure 4 This is a second flowchart illustrating the second embodiment of the liver disease diagnosis and treatment assistance method based on a large language model according to this application. In this embodiment, step S30 includes steps D1 to D2: Step D1: Construct a liver disease guideline knowledge graph with a ternary structure based on preset clinical guidelines for liver diseases. The liver disease guideline knowledge graph includes guideline ternaries with elements of applicable conditions, core evidence, and recommendation level. Step D2: Using the liver disease-specific big language model, based on the patient's diagnosis and treatment data and the liver disease guideline knowledge graph, a thought chain deduction is performed to generate preliminary diagnostic and treatment assistance results.

[0057] It is easy to understand that the aforementioned pre-defined clinical guidelines for liver diseases can be authoritative domestic and international guidelines for the diagnosis and treatment of liver diseases, such as the "Guidelines for the Prevention and Treatment of Chronic Hepatitis B (2024 Edition)" and the "Guidelines for the Diagnosis and Treatment of Primary Liver Cancer (2024 Edition)". The ternary structure can be a knowledge unit structure composed of three elements: "applicable conditions - core evidence - recommendation level", which is the basic data unit for constructing a knowledge graph of liver disease guidelines.

[0058] Therefore, in this embodiment, the aforementioned liver disease guideline knowledge graph can be a visualized knowledge network constructed around triples, achieving real-time association between patient conditions and guideline content through ClinicalBERT semantic retrieval. In this structured data, the applicable conditions can be the patient's clinical characteristics corresponding to the guideline-recommended treatment, such as the applicable conditions for antiviral treatment of chronic hepatitis B, which may include "HBV DNA > 2000 IU / mL, persistently elevated ALT, and / or no decompensated cirrhosis," etc.; core evidence can be the medical basis supporting the guideline recommendations, including randomized controlled trials (RCT) results, pathophysiological mechanisms, real-world research data, etc., such as "evidence that TAF (Tenofovir Alafenamide) can inhibit HBV DNA replication comes from a multicenter RCT study"; the recommendation level can be the strength of recommendation and level of evidence explicitly stated in the guideline, such as Grade A (strong recommendation supported by high-quality RCTs), Grade B (weak recommendation supported by moderate-quality cohort studies), etc.

[0059] Therefore, the above-mentioned guideline triad can be a combination of "applicable conditions - core evidence - recommendation level" that includes a single recommended regimen, such as "applicable conditions: compensated hepatitis B cirrhosis + HBV DNA > 2000 IU / mL; core evidence: HBV DNA inhibition rate of more than 90% after 5 years of TAF treatment; recommendation level: strong recommendation of Class IA".

[0060] For example, the assistive device can first parse the preset authoritative guidelines during the construction of the liver disease guideline knowledge graph, extract the applicable conditions, core evidence and recommendation level of each recommended option, and construct triples; use a graph database (such as Neo4j) to store the triples, establish the association between the patient's disease characteristics and "applicable conditions-evidence type-recommendation level", and form the liver disease guideline knowledge graph; at the same time, optimize the semantic retrieval function through the ClinicalBERT model to support accurate matching based on the patient's condition.

[0061] Then, the patient's medical data (such as "Hepatitis B surface antigen positive for 10 years, discomfort in the liver area for 1 week, HBV DNA 5×") will be used. The patient's blood alcohol content (IU / mL, International Unit / millilite, ALT 80U / L) is input into a liver disease-specific language model. The model then initiates CoT (Chain of Thought) reasoning, first analyzing the patient's disease type (hepatitis B cirrhosis) and disease stage (compensated phase), and identifying abnormal key indicators (elevated HBV DNA, elevated ALT) to generate patient symptom characteristics. Using RAG (Retrieval-Augmented Generation) technology, the liver disease guideline knowledge graph is retrieved to obtain a triplet for "antiviral treatment in the compensated phase of hepatitis B cirrhosis" that matches the patient's symptom characteristics. Finally, combining the reasoning chain with the matched guideline triplet evidence, preliminary diagnostic and treatment assistance results are generated, including structured medical records, disease analysis, and treatment recommendations.

[0062] Therefore, existing models suffer from a lack of authoritative guidelines for reasoning, resulting in untraceable decision-making logic, low reliability of output results, and ambiguous clinical responsibility delineation. This embodiment addresses these issues by using knowledge graph construction to structure and visualize guideline content, providing authoritative evidence for reasoning. Furthermore, by employing a "thinking chain deduction + guideline association" approach, the decision-making process becomes deconstructible and traceable, overcoming the clinical application obstacles caused by the "black box nature" of existing general-purpose large language models, improving the reliability of results, and clarifying the basis for decision-making and the boundaries of responsibility.

[0063] In one feasible implementation, refer to Figure 5 , Figure 5 This is a third flowchart illustrating the second embodiment of the liver disease diagnosis and treatment assistance method based on a large language model according to this application. In this embodiment, step D2 includes steps D21 to D24: Step D21: Decompose the patient's medical data into a set of medical elements; Step D22: Perform thought chain deduction on the set of diagnostic and treatment elements to obtain a decision subset associated according to a preset thought chain order; the decision subset includes a disease feature subset and an examination subset; It is important to understand that the aforementioned set of diagnostic and treatment elements can be a collection of key information extracted from patient diagnostic and treatment data, covering elements such as disease type, stage, complications, laboratory indicators, past medical history, and medication history, such as "{Disease type: Chronic hepatitis B, HBV DNA: 5×10 4 "IU / mL, ALT: 80U / L, Previous medication: None" or "HBV DNA elevated + ALT elevated".

[0064] As the above analysis shows, the derivation of the thought chain can be a reasoning process that breaks down the diagnostic and treatment decisions corresponding to the set of diagnostic and treatment elements into atomic steps. Each step can correspond to the generation of a subset of decisions, ensuring logical coherence. Accordingly, the above-mentioned preset thought chain order can be a reasoning order pre-set according to the liver disease diagnosis and treatment guidelines, such as "condition assessment → examination recommendations → treatment plan → follow-up plan".

[0065] Therefore, the aforementioned decision subset can be a set of interim results generated during the derivation of the thought process. The disease characteristic subset can be key clinical information describing the patient's current liver disease-related condition. Its content can be sourced from electronic medical records, including chief complaints, present illness, past medical history, physical examination, and preliminary diagnosis, to provide a basis for subsequent examinations, treatments, and follow-ups. For example, the disease characteristic subset may include "Hepatitis B surface antigen positive for 10 years" or "Recent onset of discomfort in the liver area." The examination subset can be derived from the disease characteristic subset, identifying the necessary examinations or laboratory indicators. Its source can be examination procedures, laboratory indicators, and imaging examinations recommended in clinical guidelines to clarify the diagnosis or assess the severity of the condition. For example, the examination subset may include "urgent testing of AFP or PIVKA-II (Protein Induced by Vitamin K Absence or Antagonist-II)," "Enhanced abdominal CT (Computed Tomography)," or "Testing liver function + coagulation function." In summary, the aforementioned disease feature subsets and examination subsets are sequentially related decision subsets formed by logically decomposing the diagnostic and treatment elements according to the clinical diagnosis and treatment pathway for liver diseases. They can respectively correspond to the complete clinical process of "diagnostic input - examination deduction", such as "tumor diameter analysis → vascular invasion examination → liver function classification → treatment plan matching".

[0066] Step D23: Based on the retrieval enhancement mechanism, dynamically match the set of diagnostic and treatment elements with the guideline triples in the liver disease guideline knowledge graph, and generate an interpretable decision chain based on the matching triples, the disease feature subset, and the examination subset; Step D24: Generate preliminary diagnostic and treatment assistance results based on the subset of disease characteristics, the subset of examinations, and the interpretable decision chain.

[0067] Understandably, the aforementioned retrieval enhancement mechanism can be a technique that optimizes the model output by real-time retrieval of external knowledge bases (such as guideline knowledge graphs), i.e., the aforementioned RAG technology, to ensure that the reasoning process is supported by authoritative evidence. During the dynamic matching process, the auxiliary device can perform real-time association based on the semantic similarity between the set of diagnostic and treatment elements analyzed from the patient's real-time medical data and the guideline triples in the liver disease guideline knowledge graph. For example, "HBV DNA elevation + ALT elevation" can be matched with the "antiviral treatment triple." Therefore, the content of the aforementioned matching triple can consist of the recommendation level, core evidence, and applicable conditions obtained from the liver disease guideline knowledge graph that match the set of diagnostic and treatment elements, in order to generate treatment recommendations or interventions. At this point, the determined interpretable decision chain can be a structured decision logic that includes "reasoning steps + evidence sources" based on the matching triplet to generate each disease feature subset and examination subset. The structure can be "patient disease feature subset / examination subset → guideline applicability matching degree → core evidence cited literature ID → recommendation level corresponding to evidence strength". For example, the interpretable decision chain can be: "ALT elevation + HBV DNA positive" → disease assessment: 85% matching degree of compensated hepatitis B cirrhosis → data from the "Guidelines for the Prevention and Treatment of Chronic Hepatitis B (2024 Edition)" → recommendation level A.

[0068] In summary, the auxiliary equipment can first break down the set of diagnostic and treatment elements, perform structured analysis on the input patient diagnostic and treatment data, and use entity recognition algorithms to extract elements such as disease type, test index values, and past medical history to form a set of diagnostic and treatment elements and standardize its expression.

[0069] Then, based on the set of diagnostic and treatment elements, a thought chain deduction is performed to generate a decision subset model, and reasoning is carried out in the order of "disease characteristics → examination plan". For example, first, a disease characteristic subset (such as "active phase of chronic hepatitis B, compensated cirrhosis") is generated by combining the set of elements with specialist knowledge; based on the disease characteristics, examination items such as liver function re-examination and abdominal ultrasound are recommended to form an examination subset.

[0070] Subsequently, the guideline dynamically matches decision subsets to generate an interpretable decision chain. In this process, a retrieval enhancement mechanism can be used to semantically match the set of diagnostic and treatment elements with triples in the knowledge graph. For example, "HBV DNA elevation + ALT elevation" matches "hepatitis B antiviral treatment triple" to obtain the corresponding indications, core evidence (RCT study data), and recommendation level of the matched "hepatitis B antiviral treatment triple". The decision subset generated in the reasoning step is then integrated with the matched evidence to form an interpretable decision chain.

[0071] Finally, based on the interpretable decision chain, the content of each decision subset is integrated to generate preliminary diagnostic and treatment assistance results. Therefore, this embodiment breaks down the "set of diagnostic and treatment elements" into several sequential workstations like an "assembly line." Each workstation first generates an "intermediate inference representation" (a machine-readable intermediate conclusion), and immediately uses search enhancement (RAG) to compare it with the set of diagnostic and treatment elements and the real-time updated liver disease guideline knowledge graph. The most suitable "applicable conditions-core evidence-recommendation level" matching triple is then attached to form an "interpretable decision chain," so that doctors can see the corresponding original guideline text at any subsequent step, achieving traceability of responsibility.

[0072] This embodiment addresses the problems of fragmented reasoning logic, lack of structured decision chains, and lack of clear evidence supporting the output results in existing large language models used for auxiliary analysis in liver disease diagnosis and treatment, leading to low clinical applicability and credibility. This embodiment addresses these issues by breaking down diagnostic and treatment elements and using structured reasoning to make the decision-making logic clear and coherent. Furthermore, it enhances the credibility of results by providing authoritative evidence for each decision-making step through enhanced retrieval and guideline matching. Simultaneously, it solves the "black box" problem of model reasoning by constructing an interpretable decision chain, meeting the clinical need for transparency in decision-making.

[0073] In one feasible implementation, the preliminary diagnostic and treatment auxiliary results include structured electronic medical records, laboratory early warning information, and guideline push content; in this embodiment, step D24 includes steps D241~D243: Step D241: Call the medical record structure template, fill the disease feature subset and the examination subset into the corresponding fields of the preset electronic medical record, and generate the structured electronic medical record; Step D242: Perform a two-dimensional threshold anomaly analysis on the test index data to generate the test early warning information; Step D243: Sort the matching guidelines corresponding to the interpretable decision chain according to the recommendation level of the interpretable decision chain, and output the sorted matching guidelines and the corresponding interpretable decision chain as the guide push content.

[0074] It is easy to understand that the aforementioned structured electronic medical record can be an electronic medical record that conforms to clinical norms and data standards, with clearly defined fields and a unified format, including fixed modules such as chief complaint, present illness, physical examination, auxiliary examinations, and preliminary diagnosis. Correspondingly, the aforementioned structured medical record template can be a preset electronic medical record format template, or a liver disease specialty medical record template with customized fields according to the needs of liver disease specialty, such as a template for liver cancer patients highlighting liver area signs, AFP levels, etc. Therefore, the liver disease specialty big language model can call liver disease specialty medical record templates (such as a template specifically for hepatitis B cirrhosis), and at the same time, the liver disease specialty big language model can extract content such as chief complaint (e.g., "15-year history of hepatitis B, discomfort in the liver area for 1 week") and present illness (e.g., "diagnosed with hepatitis B in 2010, diagnosed with cirrhosis in 2019") from the disease feature subset, extract laboratory / imaging results from the examination subset, automatically fill them into the corresponding fields of the template; and organize the content in the order of "chief complaint → present illness → physical examination → auxiliary examinations → preliminary diagnosis" to generate a structured electronic medical record.

[0075] It is important to understand that the aforementioned early warning information can be risk alerts generated for abnormal test indicators, including abnormal values, associated disease risks, and recommended measures, such as "ALT 80U / L increase, indicating hepatocellular damage, recommending screening for active hepatitis." To obtain effective and accurate early warning information, in this embodiment, the liver disease-specific big data model can analyze abnormalities in test indicators from two dimensions: "static threshold + dynamic trend," i.e., perform the aforementioned dual-dimensional threshold anomaly analysis. The static threshold can be set based on guidelines, and the dynamic trend can be predicted based on historical data.

[0076] Therefore, the liver disease-specific big data model can perform two-dimensional analysis of laboratory test data: In the static dimension, if AFP > 400 ng / mL is detected for one month, a red alert for liver cancer risk is triggered; in the dynamic dimension, laboratory test data (including ALT, AST, AFP, etc.) can be directly input into a pre-trained time-series / risk prediction model, such as an LSTM (Long Short-Term Memory) model, and the LSTM model will output "abnormal trend results"—such as "bilirubin may increase by >50% in the next 7 days" or "probability of continued deterioration of coagulation function is 0.78"; then, the abnormal indicators, risk levels, and recommended measures are integrated into laboratory test warning information. For example, the above results can be converted into warning text / color / level to form "laboratory test warning information," which can be displayed as a pop-up or written into the medical record to remind doctors to intervene in a timely manner.

[0077] Understandably, the matching guidelines corresponding to the aforementioned explainable decision chains are the guideline documents corresponding to the core evidence in the matching triples of the explainable decision chains. Since the matching process of the liver disease guideline knowledge graph may simultaneously provide multiple matching triples, and each matching triple contains a corresponding recommendation level (such as "TAF antiviral grade A," "low molecular weight heparin anticoagulation grade B," or "liver biopsy grade A"), correspondingly, each matching triple's corresponding explainable decision chain also has a corresponding recommendation level. Therefore, the aforementioned guideline push content can be a guideline document recommendation scheme matched according to the patient's condition, and can be presented in order of recommendation level, facilitating doctors to prioritize high-level evidence. At this time, the auxiliary device can first obtain the matching triples in the liver disease guideline knowledge graph associated with each explainable decision chain, and obtain the guideline documents and recommendation levels (such as strong recommendation level A) corresponding to the core evidence; then, the liver disease-specific large language model can present the relevant guideline content in order of recommendation level from high to low, including the recommended scheme, core evidence, and applicable conditions. Simultaneously, it can also summarize and integrate the latest research literature, annotate applicable scenarios, and push it as supplementary guideline content.

[0078] Therefore, compared to existing models whose output is fragmented, lacks structured organization, and whose analysis of test indicators only remains at the numerical description level, failing to provide risk warnings and precise guidance support, and thus unable to directly assist clinical decision-making, this embodiment reduces doctors' writing and organization time through structured electronic medical records, improving the standardization of medical records; it proposes a two-dimensional test warning system to achieve early identification of disease risks, reducing the risk of missed diagnoses / misdiagnoses; and it provides doctors with evidence-based support through tiered guideline pushes, assisting in the formulation of scientific treatment plans and improving the quality and efficiency of diagnosis and treatment.

[0079] This embodiment discloses a supervised fine-tuning of a pre-set basic large language model based on a task-driven liver disease-specific dataset to obtain an optimized large language model; and optimization of the data output format of the optimized large language model based on a prompting engineering consisting of step-by-step prompt words to obtain a liver disease-specific large language model. A liver disease guideline knowledge graph with a ternary structure is constructed based on pre-defined clinical guidelines for liver diseases. This knowledge graph includes guideline triplets with applicable conditions, core evidence, and recommendation levels as elements. Patient diagnosis and treatment data are decomposed into sets of diagnostic and treatment elements using a liver disease-specific large language model. A thought chain deduction is performed on these sets to obtain decision subsets associated according to a pre-defined thought chain order. These decision subsets include disease feature subsets and examination subsets. A retrieval enhancement mechanism is used to dynamically match the sets of diagnostic and treatment elements with the guideline triplets in the liver disease guideline knowledge graph, generating interpretable decision chains based on the matched triplets, disease feature subsets, and examination subsets. A structured medical record template is invoked, and the disease feature subsets and examination subsets are filled into the corresponding fields of a pre-defined electronic medical record to generate a structured electronic medical record. A two-dimensional threshold anomaly analysis is performed on laboratory indicator data to generate laboratory warning information. The matching guidelines corresponding to the interpretable decision chains are sorted according to their recommendation levels, and the sorted matching guidelines and their corresponding interpretable decision chains are output as guideline push content. Compared to existing models, which output fragmented content lacking structure and whose analysis of test indicators only provides numerical descriptions and fails to offer risk warnings or precise guidance, thus hindering direct clinical decision-making, this embodiment addresses these issues. It reduces doctors' writing and organization time through structured electronic medical records, improving record standardization. The proposed dual-dimensional test warning system enables early identification of disease risks, reducing the risk of missed or misdiagnosed diagnoses. Simultaneously, tiered guideline delivery provides doctors with evidence-based support, assisting in the development of scientific treatment plans and improving the quality and efficiency of diagnosis and treatment.

[0080] This application also provides a liver disease diagnosis and treatment assistance system based on a large language model; please refer to [reference needed]. Figure 6 , Figure 6 This is a schematic diagram of the module structure of the liver disease diagnosis and treatment assistance system based on a large language model, according to an embodiment of this application. In this embodiment, the liver disease diagnosis and treatment assistance system based on a large language model includes: The data processing module 601 is used to collect multi-source liver disease data and preprocess the multi-source liver disease data to generate a task-driven liver disease-specific dataset. Model training module 602 is used to iteratively train a preset basic large language model based on the task-driven liver disease dataset to obtain a liver disease specific large language model. The reasoning generation module 603 is used to input patient diagnosis and treatment data into the liver disease-specific big language model to perform thought chain deduction and generate preliminary diagnosis and treatment auxiliary results. The security verification module 604 is used to perform preset output verification on the preliminary diagnostic and treatment assistance results to obtain the target diagnostic and treatment assistance results.

[0081] In a feasible implementation, in this embodiment, the data processing module 601 is used to interface with the hospital information system and the laboratory information system to collect multi-source liver disease data composed of electronic medical records, laboratory indicator data, clinical guidelines, research literature, and standardized medical terminology; to clean the multi-source liver disease data through a medical data quality rule base to obtain effective liver disease data; to standardize the effective liver disease data to obtain standardized liver disease data; and to perform data association on the standardized liver disease data with corresponding diagnostic and treatment task parameters to generate a task-driven liver disease dataset composed of several specialized subset datasets.

[0082] In a feasible implementation, in this embodiment, the model training module 602 is used to supervise and fine-tune a preset basic large language model based on the task-driven liver disease dataset to obtain an optimized large language model; and to optimize the data output format of the optimized large language model according to a prompting project composed of step-by-step prompt words to obtain a liver disease-specific large language model.

[0083] In a feasible implementation, in this embodiment, the reasoning generation module 603 is used to construct a liver disease guideline knowledge graph with a triple structure based on a preset liver disease clinical guideline. The liver disease guideline knowledge graph includes guideline triples with applicable conditions, core evidence, and recommendation levels as elements. Through the liver disease-specific big language model, based on the patient's diagnosis and treatment data and the liver disease guideline knowledge graph, a thought chain inference is performed to generate preliminary diagnosis and treatment assistance results.

[0084] In a feasible implementation, in this embodiment, the reasoning generation module 603 is used to decompose the patient's diagnosis and treatment data into a set of diagnosis and treatment elements; perform thought chain deduction on the set of diagnosis and treatment elements to obtain a decision subset associated according to a preset thought chain order; the decision subset includes a disease feature subset and an examination subset; dynamically match the set of diagnosis and treatment elements with the guideline triples in the liver disease guideline knowledge graph based on a retrieval enhancement mechanism, and generate an interpretable decision chain based on the matching triples, the disease feature subset, and the examination subset; and generate preliminary diagnosis and treatment assistance results based on the disease feature subset, the examination subset, and the interpretable decision chain.

[0085] In a feasible implementation, in this embodiment, the reasoning generation module 603 is used to call the medical record structure template, fill the disease feature subset and the examination subset into the corresponding fields of the preset electronic medical record, and generate the structured electronic medical record; perform two-dimensional threshold anomaly analysis on the test indicator data to generate the test early warning information; sort the matching guidelines corresponding to the interpretable decision chain according to the recommendation level corresponding to the interpretable decision chain, and output the sorted matching guidelines and the corresponding interpretable decision chain as the guide push content.

[0086] In a feasible implementation, in this embodiment, the security verification module 604 is used to perform a comparative analysis of diagnostic and treatment elements on the local liver disease knowledge base and the preliminary diagnostic and treatment assistance results to obtain a diagnostic and treatment element matching deviation rate; if the diagnostic and treatment element matching deviation rate exceeds a preset threshold, an expert review process is triggered, and manual review opinions based on the expert review process are obtained; the preliminary diagnostic and treatment assistance results are optimized according to the manual review opinions to obtain the target diagnostic and treatment assistance results.

[0087] The liver disease diagnosis and treatment assistance system based on a large language model provided in this application, employing the liver disease diagnosis and treatment assistance method based on a large language model in the above embodiments, can solve the technical problem of low accuracy in liver disease auxiliary analysis using existing general-purpose large language models. Compared with the prior art, the beneficial effects of the liver disease diagnosis and treatment assistance system based on a large language model provided in this application are the same as those of the liver disease diagnosis and treatment assistance method based on a large language model provided in the above embodiments, and other technical features of the liver disease diagnosis and treatment assistance system based on a large language model are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0088] This application provides a liver disease diagnosis and treatment auxiliary device based on a large language model. The liver disease diagnosis and treatment auxiliary device based on a large language model includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the liver disease diagnosis and treatment auxiliary method based on a large language model in the above embodiment 1.

[0089] The following is for reference. Figure 7This document illustrates a structural schematic diagram of a liver disease diagnosis and treatment auxiliary device based on a large language model, suitable for implementing embodiments of this application. The liver disease diagnosis and treatment auxiliary device based on a large language model in this application embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The liver disease diagnosis and treatment auxiliary device based on a large language model shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0090] like Figure 7 As shown, the liver disease diagnosis and treatment auxiliary device based on a large language model may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to programs stored in read-only memory (ROM) 1002 or programs loaded from storage device 1003 into random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the liver disease diagnosis and treatment auxiliary device based on the large language model. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the large language model-based liver disease diagnostic and treatment auxiliary device to exchange data wirelessly or via wired communication with other devices. Although the figure shows a large language model-based liver disease diagnostic and treatment auxiliary device with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0091] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment disclosed in this application includes a liver disease diagnosis and treatment assistance program product based on a large language model, which includes a liver disease diagnosis and treatment assistance program based on a large language model carried on a computer-readable medium, the large language model-based liver disease diagnosis and treatment assistance program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the liver disease diagnosis and treatment assistance program based on a large language model can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the liver disease diagnosis and treatment assistance program based on a large language model is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0092] The liver disease diagnosis and treatment auxiliary device based on a large language model provided in this application, employing the liver disease diagnosis and treatment auxiliary method based on a large language model in the above embodiments, can solve the technical problem of low accuracy in auxiliary liver disease analysis using existing general-purpose large language models. Compared with the prior art, the beneficial effects of the liver disease diagnosis and treatment auxiliary device based on a large language model provided in this application are the same as those of the liver disease diagnosis and treatment auxiliary method based on a large language model provided in the above embodiments, and other technical features in this liver disease diagnosis and treatment auxiliary device based on a large language model are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0093] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0094] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0095] This application provides a storage medium having computer-readable program instructions (i.e., a liver disease diagnosis and treatment assistance program based on a large language model) stored thereon, which are used to execute the liver disease diagnosis and treatment assistance method based on a large language model in the above embodiments.

[0096] The storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of the storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0097] The aforementioned storage medium may be included in a liver disease diagnosis and treatment auxiliary device based on a large language model; or it may exist independently and not be assembled into a liver disease diagnosis and treatment auxiliary device based on a large language model.

[0098] The aforementioned storage medium carries one or more programs. When the aforementioned one or more programs are executed by the liver disease diagnosis and treatment auxiliary device based on a large language model, the liver disease diagnosis and treatment auxiliary device based on a large language model becomes: liver disease diagnosis and treatment auxiliary device based on a large language model.

[0099] The liver disease diagnosis and treatment auxiliary program code based on a large language model, used to perform the operations of this application, can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0100] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and liver disease diagnosis and treatment auxiliary programs based on large language models according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0101] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0102] The readable storage medium provided in this application is a storage medium storing computer-readable program instructions (i.e., a liver disease diagnosis and treatment assistance program based on a large language model) for executing the aforementioned liver disease diagnosis and treatment assistance method based on a large language model. This solves the technical problem of low accuracy in liver disease auxiliary analysis using existing general-purpose large language models. Compared with the prior art, the beneficial effects of the storage medium provided in this application are the same as those of the liver disease diagnosis and treatment assistance method based on a large language model provided in the above embodiments, and will not be repeated here.

[0103] The above are only some embodiments of this application and do not limit the scope of the solution of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included within the protection scope of this application.

Claims

1. A liver disease diagnosis and treatment auxiliary method based on a large language model, characterized in that, The method includes: Collect multi-source liver disease data and preprocess the multi-source liver disease data to generate a task-driven liver disease-specific dataset; Based on the task-driven liver disease dataset, the preset basic large language model is iteratively trained to obtain a liver disease-specific large language model. The patient's diagnosis and treatment data are input into the liver disease-specific big language model to perform thought chain deduction and generate preliminary diagnosis and treatment auxiliary results. The preliminary diagnostic and treatment assistance results are verified by a preset output to obtain the target diagnostic and treatment assistance results.

2. The liver disease diagnosis and treatment assistance method based on a large language model as described in claim 1, characterized in that, The steps of collecting multi-source liver disease data, preprocessing the multi-source liver disease data, and generating a task-driven liver disease-specific dataset include: Connect to hospital information systems and laboratory information systems to collect multi-source liver disease data, consisting of electronic medical records, laboratory indicator data, clinical guidelines, research literature, and standardized medical terminology; The multi-source liver disease data was cleaned using a medical data quality rule base to obtain valid liver disease data; The valid liver disease data are then standardized to obtain standardized liver disease data. The standardized liver disease data is correlated with corresponding diagnostic and treatment task parameters to generate a task-driven liver disease dataset consisting of several specialized subset datasets.

3. The liver disease diagnosis and treatment assistance method based on a large language model as described in claim 1, characterized in that, The step of iteratively training a pre-defined basic large language model based on the task-driven liver disease dataset to obtain a liver disease-specific large language model includes: Based on the task-driven liver disease dataset, the preset basic large language model is supervised and fine-tuned to obtain an optimized large language model. The data output format of the optimized large language model is optimized based on the prompting engineering consisting of step-by-step prompt words to obtain a large language model for liver disease.

4. The liver disease diagnosis and treatment assistance method based on a large language model as described in claim 2, characterized in that, The steps of inputting patient diagnosis and treatment data into the liver disease-specific large language model for thought chain deduction to generate preliminary diagnosis and treatment auxiliary results include: A liver disease guideline knowledge graph with a ternary structure is constructed based on pre-set clinical guidelines for liver diseases. The liver disease guideline knowledge graph includes guideline ternaries with elements of applicable conditions, core evidence, and recommendation level. Using the aforementioned liver disease-specific big language model, and based on the patient's diagnosis and treatment data and the liver disease guideline knowledge graph, a thought chain deduction is performed to generate preliminary diagnostic and treatment assistance results.

5. The liver disease diagnosis and treatment assistance method based on a large language model as described in claim 4, characterized in that, The step of generating preliminary diagnostic and treatment auxiliary results by reasoning based on the patient's medical data and the liver disease guideline knowledge graph includes: The patient's medical data is broken down into a set of medical elements; The set of diagnostic and treatment elements is subjected to a thought chain deduction to obtain a decision subset associated according to a preset thought chain order; the decision subset includes a disease feature subset and an examination subset; Based on the retrieval enhancement mechanism, the set of diagnostic and treatment elements is dynamically matched with the guideline triples in the liver disease guideline knowledge graph, and an interpretable decision chain is generated based on the matched triples, the disease feature subset, and the examination subset. Preliminary diagnostic and treatment assistance results are generated based on the subset of disease characteristics, the subset of examinations, and the interpretable decision chain.

6. The liver disease diagnosis and treatment assistance method based on a large language model as described in claim 5, characterized in that, The preliminary diagnostic and treatment support results include structured electronic medical records, laboratory early warning information, and guideline push content; The step of generating preliminary diagnostic and treatment assistance results based on the subset of disease characteristics, the subset of examinations, and the interpretable decision chain includes: The structured electronic medical record is generated by calling the medical record structure template, filling the disease feature subset and the examination subset into the corresponding fields of the preset electronic medical record; Perform a two-dimensional threshold anomaly analysis on the test indicator data to generate the test early warning information; The matching guidelines corresponding to the explainable decision chain are sorted according to the recommendation level of the explainable decision chain, and the sorted matching guidelines and the corresponding explainable decision chains are associated and output as the guide push content.

7. The liver disease diagnosis and treatment assistance method based on a large language model as described in claim 1, characterized in that, The step of performing preset output verification on the preliminary diagnostic and treatment assistance results to obtain the target diagnostic and treatment assistance results includes: The local liver disease knowledge base and the preliminary diagnostic and treatment assistance results are aggregated and compared to obtain the diagnostic and treatment element matching deviation rate. If the detection shows that the matching deviation rate of the diagnosis and treatment elements exceeds a preset threshold, an expert review process is triggered, and manual review opinions based on the feedback from the expert review process are obtained. The preliminary diagnostic and treatment assistance results are optimized based on the manual review opinions to obtain the target diagnostic and treatment assistance results.

8. A liver disease diagnosis and treatment auxiliary system based on a large language model, characterized in that, The liver disease diagnosis and treatment assistance system based on a large language model includes: The data processing module is used to collect multi-source liver disease data and preprocess the multi-source liver disease data to generate a task-driven liver disease-specific dataset. The model training module is used to iteratively train the preset basic large language model based on the task-driven liver disease dataset to obtain a liver disease-specific large language model. The reasoning generation module is used to input patient diagnosis and treatment data into the liver disease-specific big language model to perform thought chain deduction and generate preliminary diagnosis and treatment assistance results; The security verification module is used to perform preset output verification on the preliminary diagnostic and treatment assistance results to obtain the target diagnostic and treatment assistance results.

9. A liver disease diagnosis and treatment auxiliary device based on a large language model, characterized in that, The device includes: a memory, a processor, and a liver disease diagnosis and treatment assistance program based on a large language model stored in the memory and executable on the processor, the liver disease diagnosis and treatment assistance program based on a large language model configured to implement the steps of the liver disease diagnosis and treatment assistance method based on a large language model as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a liver disease diagnosis and treatment assistance program based on a large language model. When the liver disease diagnosis and treatment assistance program based on a large language model is executed by the processor, it implements the steps of the liver disease diagnosis and treatment assistance method based on a large language model as described in any one of claims 1 to 7.