Diagnostic and treatment pathway mining method and system based on dialogue data of internet hospital

WO2026200275A1PCT designated stage Publication Date: 2026-10-01RENJI HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/076485
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-02-02
Publication Date
2026-10-01

Smart Images

  • Figure CN2026076485_01102026_PF_FP_ABST
    Figure CN2026076485_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the interdisciplinary field of natural language processing and biomedicine, and discloses a diagnostic and treatment pathway mining method and system based on dialogue data of an Internet hospital. The method comprises: acquiring medical dialogue data, and preprocessing the data; designing a prompt to guide a large language model to identify and annotate medical entities in a doctor-patient dialogue, and assigning a preset label to each medical entity, wherein the medical entities are divided into four categories in the present invention: symptom, examination, disease and treatment; analyzing contextual relationships between the entities by means of the large language model, and extracting sub-chains having logical associations; performing strict validation and logical rule-based filtering on the sub-chains to ensure the medical rationality and clinical significance; and fusing the sub-chains to generate a standardized diagnostic and treatment pathway. The present disclosure effectively mines diagnostic and treatment processes in medical dialogues, which not only conforms to the standards in the medical field, but also exhibits strong operability and broad applicability.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for mining treatment pathways based on dialogue data from internet hospitals Technical Field

[0001] This invention relates to the interdisciplinary fields of natural language processing and biomedicine, and in particular to a method and system for mining diagnostic and treatment pathways based on dialogue data from internet hospitals. Background Technology

[0002] In modern internet hospitals, the dialogue data between doctors and patients contains a wealth of diagnostic and treatment knowledge, covering the entire process from patient symptom descriptions, examination information, doctor diagnoses to treatment plans. Utilizing this data for diagnostic and treatment process analysis can help identify key steps in diagnosis and treatment, providing scientific decision support.

[0003] Information extraction from medical dialogues is an emerging research direction in the field of medical information extraction. It aims to extract key information (such as clinical terminology and its attributes, treatment plans, etc.) from doctor-patient conversations to automatically generate electronic medical records (EMRs), thereby reducing the burden on doctors writing narrative reports. Early research introduced information extraction methods based on heuristic rules (such as character matching, regular expressions, etc.), but the results were unsatisfactory. To capture more complex interdependencies between entity pairs, recent research aims to enhance information extraction models through external modules. For example, LogicRE, a framework based on state-of-the-art (SOTA) rules, first learns logical rules based on the output logarithm of a trained neural model, and then refines the predicted relationships of the neural model using the learned rules. MILR first learns logical rules from annotated data, and then trains a neural model that is penalized by auxiliary loss for reflecting violations of the learned rules. However, both of these frameworks suffer from error propagation problems due to their pipeline nature.

[0004] Large Language Models (LLMs), with their powerful reasoning and linguistic capabilities combined with extensive knowledge reserves, have demonstrated enormous potential in the field of medical information extraction. These models can understand diverse terminology and extract information efficiently according to predetermined rules. However, existing research on LLM-based information extraction mainly focuses on entity relation extraction, and the mining of treatment processes from medical dialogue data remains to be explored. How to design reasonable extraction methods to extract treatment processes from internet hospital dialogue data to meet the needs of scientific decision support remains a challenging problem.

[0005] In conclusion, there is an urgent need for a method that can mine the diagnosis and treatment process based on medical dialogue data to solve the above problems.

[0006] In view of this, this invention proposes a method for mining diagnosis and treatment processes based on dialogue data from internet hospitals. First, to address the error propagation problem inherent in traditional pipeline-based extraction methods, this invention proposes an LLM-based review mechanism. Based on prior rules, filtering is performed before the fusion stage to block error propagation during extraction. Second, to address the current lack of diagnosis and treatment process mining methods based on medical dialogues, this invention proposes a multi-stage extraction and fusion method for diagnosis and treatment processes based on a large model, capable of identifying the physical location and logical relationships of entities in the dialogue, extracting sub-chains, and then generating high-quality process links through filtering and fusion strategies. A method for mining diagnosis and treatment paths based on dialogue data from internet hospitals includes the following steps:

[0007] The invention acquires medical dialogue data and preprocesses it. Instructions are designed to guide a large language model to identify and label medical entities in doctor-patient dialogues, assigning a pre-defined label to each entity. Medical entities are categorized into four types: features, examinations, diseases, and treatments. The large language model analyzes the contextual relationships between entities, extracting logically related sub-chains. These sub-chains undergo rigorous review and logical rule filtering to ensure their medical rationality and clinical significance. Finally, the sub-chains are merged to generate standardized diagnostic and treatment pathways.

[0008] The preprocessing includes: removing noisy data: using natural language processing (NLP) technology to remove irrelevant content, including system prompts, meaningless pauses, repeated sentences, emoticons, and irrelevant topics; removing sensitive information: desensitizing information involving patient privacy; removing redundant information: merging repeated consultation questions or content that asks about the same symptoms multiple times; standardizing expressions: unifying the format of medical terminology; and optimizing the storage format: transforming the cleaned data into a structured format that is easy to analyze.

[0009] The fusion process includes: traversing the disease set, fusing the examination-disease and disease-treatment sub-chains with the same disease, and then, based on the examination set, fusing the feature-examination and examination-disease sub-chains with the same examination to form a main chain from feature to examination, and then from examination to disease to treatment; for the remaining isolated sub-chains, first calculating the similarity between the nodes in the sub-chain and the nodes with the corresponding labels in the main chain; if the similarity exceeds a preset threshold, then the large model determines whether the sub-chain can be fused with the main chain; if the similarity does not reach the threshold, the sub-chain is directly discarded; finally, a complete diagnosis and treatment process is formed.

[0010] The design instructions guide the large language model to identify and label medical entities in doctor-patient dialogues, assigning a pre-defined label to each medical entity. Medical entities in this invention are subdivided into four categories: features, examinations, diseases, and treatments. Features are symptoms, signs, or clinical manifestations, including at least headache and fever; examinations are medical examinations or tests, including at least blood tests and CT scans; diseases are diagnoses made by doctors, including at least pneumonia and diabetes; and treatments are treatment plans, including at least medication, surgery, physical therapy, and lifestyle modifications. The LLM (Limited Language Model) is used to automatically extract medical entities from the text; a predefined label is assigned to each identified entity; and medical terms with synonyms or different expressions are standardized.

[0011] By analyzing the contextual relationships between entities using a large language model, logically related sub-chains are extracted. This includes: using the model to analyze the dialogue context and identify potential logical relationships between entities; standardizing sub-chains with similar or diverse expressions, retaining frequently occurring sub-chains, and filtering out low-frequency specific sub-chains; setting frequency thresholds to retain broadly representative sub-chains; and removing isolated sub-chains that lack statistical significance.

[0012] Subchains undergo rigorous review and logical rule filtering to ensure their medical rationality and clinical significance. This includes defining a constraint set C = {symptom → examination, examination → disease, disease → treatment} to ensure the subchains conform to medical diagnostic logic; and validating the rationality of subchains using a medical knowledge base and clinical guidelines. The following abnormal subchains are removed: those lacking an examination step (directly from "symptom → disease") and those with an incorrect order ("treatment → disease").

[0013] A second aspect of the present invention provides a system for mining treatment pathways based on dialogue data from internet hospitals, comprising:

[0014] The system includes a preprocessing unit for acquiring and preprocessing medical dialogue data; an entity labeling unit for designing instructions to guide a large language model to identify and label medical entities in the doctor-patient dialogue, assigning a pre-defined label to each medical entity; medical entities are subdivided into four categories in this invention: features, examinations, diseases, and treatments; a sub-chain extraction unit for analyzing the contextual relationships between entities using a large language model to extract logically related sub-chains; a review and filtering unit for rigorously reviewing and filtering the sub-chains using logical rules to ensure their medical rationality and clinical significance; and a sub-chain fusion unit for fusing the sub-chains to generate standardized diagnostic and treatment pathways.

[0015] A third aspect of the present invention provides an electronic device, comprising: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a circuit; the at least one processor invokes the instructions in the memory to cause the electronic device to execute the above-described method for mining treatment paths based on Internet hospital dialogue data.

[0016] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the above-described method for mining treatment pathways based on Internet hospital dialogue data.

[0017] The present invention has the following beneficial effects:

[0018] This invention addresses the field of medical information extraction by proposing a novel method for mining treatment processes based on dialogue data from internet hospitals. This method solves the error propagation problem inherent in traditional pipeline-based extraction methods. Furthermore, based on a unique review mechanism and a multi-stage large-model extraction and fusion strategy, this method effectively mines treatment processes within medical dialogues, conforming to medical standards and possessing strong operability and universality. Attached Figure Description

[0019] Figure 1 is a flowchart of a method for mining diagnosis and treatment paths based on dialogue data from Internet hospitals provided in an embodiment of the present invention. Detailed Implementation

[0020] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” or “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0021] feature:

[0022] In medicine, a "characteristic" refers to the distinctive features of a disease, condition, or health state, typically manifested as symptoms, signs, or physiological changes. Characteristics can be patient-reported symptoms (e.g., pain, fatigue) or signs discovered by a physician through examination (e.g., a lump, fever). Characteristic features are crucial for diagnosing and differentiating between diseases. Characteristics can also include genetic, molecular markers, or imaging findings, which often aid in further clinical judgment.

[0023] examine:

[0024] Examination refers to a systematic assessment of a patient using various clinical methods to diagnose, monitor, and evaluate disease progression or treatment effectiveness. Examinations can be physical (e.g., physical examination) or laboratory or imaging (e.g., blood tests, X-rays, MRI). Examinations help physicians understand a patient's health status, identify potential pathological changes, and provide a basis for treatment decisions. Examinations typically include the following types:

[0025] Physical examination: The assessment of the patient's physical condition through methods such as palpation, auscultation, and inspection.

[0026] Laboratory tests: These involve analyzing samples such as blood, urine, and sputum.

[0027] Imaging examinations: using imaging techniques such as X-rays, ultrasound, CT, and MRI to examine the condition of internal organs.

[0028] disease

[0029] Also known as diagnosis, it refers to an abnormality or disorder in the body's physiology, psychology, or function. It usually manifests as a set of abnormal symptoms, signs, and / or biochemical indicators, and may be caused by infection, genetics, immune dysregulation, environmental factors, etc. Disease may affect one or more systems of the body, leading to impairment or damage to health functions, affecting the patient's quality of life and survival.

[0030] deal with

[0031] Also known as treatment, it refers to a series of medical interventions and methods aimed at alleviating, curing, or managing a disease to improve the patient's health. Treatment includes various approaches such as medication, surgery, physical therapy, psychotherapy, nutritional intervention, and lifestyle modifications. The goal of treatment is to control symptoms, restore function, or prevent further deterioration of the disease.

[0032] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to Figure 1. The first embodiment of the treatment path mining method based on Internet hospital dialogue data in the present invention includes:

[0033] Acquire medical dialogue data and preprocess the data;

[0034] The design instructions guide the large language model to identify and label medical entities in doctor-patient dialogues, assigning a pre-defined label to each medical entity; medical entities are subdivided into four categories in this invention: features, examinations, diseases, and treatments.

[0035] By analyzing the contextual relationships between entities using a large language model, logically related sub-chains are extracted.

[0036] Subchains undergo rigorous review and logical rule filtering to ensure their medical rationality and clinical significance;

[0037] By integrating the various subchains, standardized diagnostic and treatment pathways are generated.

[0038] In the data preprocessing stage, the doctor-patient dialogue data from the internet hospital undergoes comprehensive cleaning and noise reduction to ensure the accuracy and effectiveness of subsequent analysis. The main task of this stage is to remove redundant information and noise that is meaningless to the analysis, improve data quality, and ensure that the processed data reflects the authentic doctor-patient interactions. Specific operations include: removing noisy data: using natural language processing techniques to identify and eliminate irrelevant and unstructured content, such as system prompts, meaningless pauses, repeated sentences, emoticons, and irrelevant topics; removing sensitive information: to protect patient privacy and comply with laws and regulations, sensitive information involving personal privacy (such as names, ID numbers, and contact information) in the doctor-patient dialogue is anonymized during preprocessing; removing redundant information: identifying and removing redundant information in the doctor-patient dialogue. For example, for repeated consultation questions or situations where the same symptoms are asked multiple times, these duplicates can be merged or filtered out to reduce information redundancy in the dataset and ensure the efficiency and representativeness of the analysis. Through these cleaning and noise reduction steps, the final dataset is ensured to have high-quality language structure and medical relevance, providing reliable foundational data for subsequent entity labeling, path mining, and deep analysis.

[0039] In the entity labeling stage, Large Language Modeling (LLM) is used to identify and label medical entities in doctor-patient dialogues. This process aims to accurately extract key medical information from natural language text for further analysis and processing. Leveraging the powerful reasoning capabilities of LLM, various medical entities in the text can be identified, and each entity is assigned a pre-defined label. Medical entities in this invention are subdivided into four categories: features, examinations, diseases, and treatments. These categories cover various types of information commonly found in medical dialogues, specifically as follows: Features: These refer to symptoms, signs, or other clinical manifestations described by the patient in the dialogue, such as headache, cough, fever, etc. Features are often the initial clues for doctors to make a diagnosis. Examinations: These include various medical examinations recommended by doctors to confirm or rule out a disease, such as blood tests, imaging examinations (e.g., X-rays, CT scans), physiological tests, etc. Diseases: These refer to the doctor's judgment of a disease or pathological state based on the patient's symptoms, signs, and examination results, such as pneumonia, diabetes, coronary heart disease, etc. Treatments: These refer to treatment plans recommended or implemented by doctors based on the diagnostic results, including drug therapy, surgery, physical therapy, lifestyle modifications, etc.

[0040] In the sub-chain extraction stage, a large language model is first used to deeply analyze medical entities in doctor-patient dialogues, uncovering the potential logical relationships between entities within the context. The model identifies connections and dependencies between entities, extracting sub-chains with semantic relevance. Further, terminology normalization is performed on the extracted sub-chains, unifying similar or different medical terms into standardized expressions to reduce the diversity and complexity of language expression and improve the accuracy and consistency of subsequent analysis. During processing, for frequently occurring sub-chains, the system retains one to avoid redundancy and ensure data conciseness and efficiency; while sub-chains with fewer repetitions are considered to represent individual-specific medical phenomena and may not have broad statistical significance, thus being filtered out. To ensure the rationality of this process, the system sets filtering thresholds, selecting sub-chains based on their frequency of occurrence and the universality of the medical facts they represent, ensuring that the final retained sub-chains have high representativeness and validity, and avoiding interference from individual anomalies in the analysis results.

[0041] During the review and filtering phase, the extracted sub-chains are rigorously screened and filtered according to pre-defined prior rules. The core objective of this process is to ensure that the final extracted diagnostic and treatment pathways conform to both medical common sense and reflect real clinical diagnostic and treatment procedures. To achieve this goal, this invention employs a multi-dimensional rule-based filtering strategy, primarily including the internal constraints of the process and the rationality verification of the diagnostic logic. Specifically, the set of constraints within the process is precisely defined as: C = {symptoms → examination, examination → disease, disease → treatment}. This set of constraints reflects the typical diagnostic and treatment sequence in the medical field: first, symptoms guide the doctor to conduct examinations; then, a disease diagnosis is made based on the examination results; and finally, a treatment plan is formulated based on the diagnosis. The application of filtering rules requires that the extracted sub-chains must follow this basic logic. In addition, the review and filtering phase also conducts in-depth verification of the rationality of the diagnostic logic. For example, with the support of a medical knowledge base and clinical guidelines, the system can identify and eliminate sub-chains that are inconsistent with medical facts. For example, some subchains may show "symptoms → disease" without going through the necessary examination steps, or show "treatment → disease" which reverses the diagnosis and treatment process. These subchains will be considered to be inconsistent with reasonable medical diagnosis and treatment logic and will be filtered out.

[0042] In the sub-chain fusion stage, this invention employs a two-stage fusion strategy. Through progressive reasoning and refined processing, it ensures that the final generated diagnostic and treatment path conforms to medical logic and has a high degree of standardization. First, in the first-stage fusion, reasoning and connections are performed using the disease as the core node, expanding as far as possible to both ends to form the main chain. The system merges existing diagnostic and treatment chains to simplify the path, making the chain for each disease node more standardized and clear. In the second-stage fusion, the similarity between isolated sub-chain entities and main chain entities is first calculated. If the similarity exceeds a preset threshold, a larger model is used to further determine whether the sub-chain can be fused with the main chain; otherwise, the sub-chain is discarded. For the retained isolated sub-chains, the similarity between the sub-chain nodes and the corresponding tag nodes of the main chain is further calculated. When the similarity exceeds the threshold, the larger model determines whether fusion can be completed. This process effectively eliminates diagnostic and treatment paths with excessively strong individual specificity, transforming them into standard paths that conform to general medical practice.

[0043] As a preferred approach, the specific steps include:

[0044] S1. In the data preprocessing stage, the doctor-patient dialogue data in the Internet hospital is thoroughly cleaned and denoised, including removing noisy data: identifying and removing irrelevant and unstructured content, such as system prompts, meaningless pauses, emoticons, and irrelevant topics; removing sensitive information: desensitizing sensitive information involving personal privacy in the doctor-patient dialogue (such as names, ID numbers, contact information, etc.); and removing redundant information: identifying and removing redundant information in the doctor-patient dialogue, such as in situations where repeated consultation questions or multiple inquiries about the same symptoms are made.

[0045] S11. Remove noise data such as emoticons, special symbols, and garbled characters from the dialogue; construct a noise vocabulary. Remove noisy vocabulary data from the dialogue using regular expressions.

[0046] ,

[0047] in This indicates the filtered dialogue. This indicates that regular expressions are used to filter noise. Indicates the original dialogue. This indicates that the content should be removed. ;

[0048] S12. Process sensitive information: Utilize an open-source Entity Recognition (NER) model for sensitive information detection to detect and replace sensitive information in the dialogue.

[0049]

[0050] in This indicates a dialogue after processing sensitive information, where identified sensitive entities are replaced to ensure anonymity;

[0051] S13. Remove redundant information, and based on contextual similarity, vectorize adjacent dialogue sentences and calculate cosine similarity.

[0052] ,

[0053] in This indicates two adjacent conversations. express The cosine similarity is calculated if... If it is redundant, it is determined to be redundant. In this invention, it is set to... The value is 0.65.

[0054] S2. In the entity labeling stage, a prompt-guided Large Language Model (LLM) is used to identify and label medical entities in the doctor-patient dialogue. Based on LLM, various medical entities in the text are identified, and a pre-set label is assigned to each entity. In this invention, medical entities are subdivided into four categories: features, examinations, diseases, and treatments.

[0055] The prompt contains the following information:

[0056] "You are a professional doctor, and your job is to add appropriate labels to medical entities in medical consultation conversations."

[0057] Please adhere to the following points:

[0058] The tag must be one of the following: [feature, examination, disease, treatment], and other tags are prohibited.

[0059] Features include symptoms, signs or other clinical manifestations; examinations include various medical examinations or tests; disease refers to a disease in Western medicine; treatment includes drug therapy, surgery, physical therapy, lifestyle modifications, etc.

[0060] Add tags to medical entities in the conversation, such as 'I've been feeling a bit of a headache lately';

[0061] Each entity can have at most one tag;

[0062] Here is an example:

[0063]

[0064] # Please add appropriate labels to the medical entities in the medical consultation conversation based on the information provided.

[0065] S3. In the sub-chain extraction stage, the potential logical relationships between entities in the context are mined using a large language model to identify the connections and dependencies between entities and extract sub-chains with semantic relevance. Further, medical entities in the sub-chains are extracted and terminology normalization is performed to unify similar or different medical terms into standardized expressions. Then, for sub-chains with a repetition count greater than a threshold, only one of them is retained; for sub-chains with a repetition count less than a threshold, they are considered to represent individual-specific medical phenomena and do not have broad statistical significance, so they are filtered out.

[0066] S31. By mining the potential logical relationships between entities in the context through a large language model, identify the connections and dependencies between entities, and extract the semantically related sub-chains between entities.

[0067] The prompt contains the following information:

[0068] "You are a professional doctor, and your job is to identify medical entities in medical consultation conversations and associate those entities;

[0069] Please adhere to the following points:

[0070] Medical entities will be marked with a <label> annotation. You only need to identify the logical relationship between the entities after the label in the conversation.

[0071] Identify the flow relationships between entities and represent them as a chain of [<label A>Entity A-<label B>Entity B]. For example, in the case of 'Patient: Doctor, I've been experiencing severe headaches lately, what should I do?\nDoctor: Have you had your blood pressure checked?', the extraction should be [<feature>headache-<check>blood pressure check].

[0072] The physical location of an entity usually indicates the procedural relationship in the diagnosis and treatment process.

[0073] You need to consider whether your chain satisfies the logical relationship in the conversation and whether it conforms to medical common sense; this is very important.

[0074] Here is an example:

[0075]

[0076] # Please extract the entity sub-chain from the medical consultation dialogue based on the provided information.

[0077] S32. Divide the entity set by label, and merge the terms of the entities in each entity set based on the instruction prompts of the large model; for example, in the disease set, "knee arthritis" and "hip arthritis" will be merged into "osteoarthritis";

[0078] The prompt contains the following information:

[0079] "You are a professional doctor, and your job is to identify medical entities in the list and merge them into a smaller list;

[0080] Please adhere to the following points:

[0081] The medical entities in the list are all nouns representing {label};

[0082] These entities may have the same meaning but different nouns. You can merge them, for example, "knee arthritis" and "hip arthritis" can be merged into "osteoarthritis".

[0083] You can appropriately broaden the scope of the noun's definition to improve your merging efficiency;

[0084] The medical list is as follows:

[0085]

[0086] S33. Based on the prompt suggestion model, the entity terms in the sub-chains are unified into a standardized expression.

[0087] The prompt contains the following information:

[0088] "You are a professional doctor, and your job is to map medical entities in the subchain into standardized representations; note that you can only select the standardized representation that is closest to the entity in the current subchain from the given list;

[0089] # List of standardized representations of {label1}:

[0090] {tag1_list}

[0091] # List of standardized representations of {tag2}:

[0092] {tag2_list}

[0093] # Input:

[0094] {Subchain to be converted}

[0095] The following is an example for your reference.

[0096] {Sample}”

[0097] S34. For subchains with a repetition count greater than the threshold, only one of them is retained; for subchains with a repetition count less than the threshold, they are considered to represent individual-specific medical phenomena and do not have broad statistical significance, so they are filtered out.

[0098] S4. This invention employs a multi-dimensional rule filtering strategy, mainly including the constraint relationships within the process and the rationality verification of diagnostic logic. Define the set of constraints within the process:

[0099] ,

[0100] in, Symptoms refer to the patient's complaints or manifestations. The term "examination" refers to various examinations, including clinical examinations, imaging examinations, and laboratory tests. This refers to a disease, diagnosed through examination and analysis. This refers to treatment, an intervention for a diagnosed disease, requiring the extracted subchains to adhere to one of the constraints. Furthermore, supported by medical knowledge bases and clinical guidelines, the rationality of the diagnostic logic is thoroughly examined. Subchains that are inconsistent with medical facts are identified and eliminated.

[0101] S41. Define the set of constraints within the process, and verify whether the sub-chains conform to the constraints based on regular expressions.

[0102] ,

[0103] in, This represents the filtered subchain. This indicates that subchains that do not meet the constraints are filtered according to the regular expression. This represents the subchain to be filtered. Represents a set of constraints.

[0104] S42. Using the instruction prompts of the large model, and based on compliance with process constraints, further validate the sub-chains semantically and in terms of content by incorporating the medical knowledge base k:

[0105] ,

[0106] in This is a response from a large model. This indicates that knowledge is retrieved from a medical knowledge base.

[0107] The prompt contains the following information:

[0108] "You are a professional doctor, and your job is to determine whether the subchain conforms to medical facts; you will receive some relevant documents, please refer to the documents to make a professional judgment;

[0109] # document

[0110] {Retrieved Medical Knowledge}

[0111] # Input:

[0112] {Subchain to be judged}

[0113] Please determine whether the given subchain conforms to medical facts. Note that the output can only be [conforms / does not conform]. Any extra or incorrect output may cause the system to crash.

[0114] S5. In the sub-chain fusion stage, a two-stage fusion strategy is adopted. In the first stage of fusion, the disease is used as the core node for reasoning and connection, and the chain is expanded to both ends as much as possible to form the main chain. In the second stage of fusion, the similarity between isolated sub-chain entities and entities in the main chain is calculated first. If the similarity is greater than the threshold, the large model is used to further determine whether the sub-chain can be fused; otherwise, the sub-chain is discarded.

[0115] S51. Traverse the disease set, merge the check-disease and disease-treatment sub-chains with the same disease, and then based on the check set, merge the feature-check and check-disease sub-chains with the same check to form the main chain from feature to check, and then from check to disease to treatment.

[0116] S52. For the remaining isolated subchains, first calculate the similarity between the nodes in the subchain and the nodes with the corresponding labels in the main chain:

[0117] ,

[0118] in Represents the nodes in the subchain. This indicates nodes in the main chain that have the same label; if the similarity exceeds a threshold... If the similarity does not reach the threshold, the large model determines whether the sub-chain can be merged with the main chain; if the similarity does not reach the threshold, the sub-chain is discarded. In this invention, here... Set to 0.65;

[0119] The main model prompt details are as follows:

[0120] "You are a professional doctor, and your job is to determine whether two chains can merge;"

[0121] # Input:

[0122] {Subchain to be judged 1}

[0123] {Subchain 2 at the main chain entry point}

[0124] ## Determine if two chains can be merged. Note that the output can only be [Yes / No]. Any extra or incorrect output may cause the system to crash.

[0125] The above describes the method for mining treatment paths based on dialogue data from internet hospitals in embodiments of the present invention. The following describes the device for mining treatment paths based on dialogue data from internet hospitals in embodiments of the present invention:

[0126] This invention also provides an electronic device, which can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) (e.g., one or more processors) and memory, and one or more storage media (e.g., one or more mass storage devices) for storing applications or data. The memory and storage media can be temporary or persistent storage. The program stored in the storage media may include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the processor may be configured to communicate with the storage media and execute the series of instruction operations stored in the storage media on the electronic device.

[0127] The electronic device may also include one or more power supplies, one or more wired or wireless network interfaces, one or more input / output interfaces, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that the electronic device structure in this embodiment does not constitute a limitation on the electronic device itself, and may include more or fewer components, or combinations of certain components, or different component arrangements.

[0128] This invention provides an electronic device structure that can vary significantly depending on configuration or performance. It may include one or more central processing units (CPUs) (e.g., one or more processors) and memory, and one or more storage media (e.g., one or more mass storage devices) for storing applications or data. The memory and storage media can be temporary or persistent storage. The program stored in the storage media may include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the processor may be configured to communicate with the storage media and execute the series of instruction operations stored in the storage media on the electronic device.

[0129] The electronic device may also include one or more power supplies, one or more wired or wireless network interfaces, one or more input / output interfaces, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that the structure of the electronic device does not constitute a limitation on the electronic device itself, and may include more or fewer components than described above, or combine certain components, or have different component arrangements.

[0130] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the aforementioned method.

[0131] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system, device, or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0132] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0133] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for mining treatment pathways based on dialogue data from internet hospitals, characterized in that, Includes the following steps: Acquire medical dialogue data and preprocess the data; The design instructions guide the large language model to identify and label medical entities in doctor-patient dialogues, assigning a pre-defined label to each medical entity; medical entities are subdivided into four categories: features, examinations, diseases, and treatments; entity sets are divided according to labels; By analyzing the contextual relationships between entities using a large language model, logically related sub-chains are extracted. Subchains undergo rigorous review and logical rule filtering to ensure their medical rationality and clinical significance; Reasoning and connection are performed with diseases as the core nodes, and the main chain is extended to both ends to form the main chain. Specifically, the disease set is traversed, and the check-disease and disease-treatment sub-chains with the same disease are merged. Then, based on the check set, the feature-check and check-disease sub-chains with the same check are merged to form the main chain from feature to check, and then from check to disease to treatment. For the remaining isolated subchains, the similarity between the nodes in the subchain and the nodes with the corresponding labels in the main chain is first calculated. If the similarity exceeds a preset threshold, the large model will then determine whether the subchain can be merged with the main chain. If the similarity does not reach the threshold, the subchain is discarded directly. Finally, a complete diagnosis and treatment process is formed.

2. The method for mining treatment paths based on dialogue data from internet hospitals according to claim 1, characterized in that, The preprocessing includes: Remove noisy data: Use natural language processing (NLP) techniques to remove irrelevant content, including system prompts, meaningless pauses, repeated sentences, emoticons, and irrelevant topics; Sensitive information removal: Desensitize information involving patient privacy; Remove redundant information: merge duplicate consultation questions or questions asking the same symptoms multiple times; Standardized expression: Unifying the format of medical terminology; Storage format optimization: Transform the cleaned data into a structured format.

3. The method for mining treatment paths based on dialogue data from internet hospitals according to claim 1, characterized in that, The design instructions guide the large language model to identify and label medical entities in doctor-patient dialogues, assigning a pre-defined label to each medical entity; including: Medical entities are classified into four categories: characteristics, examinations, diseases, and treatments. The characteristics are: symptoms, signs, or clinical manifestations, including at least headache and fever; The examination includes: medical examinations or tests, including at least blood tests and CT scans; The illness is defined as a doctor's diagnosis, which includes at least pneumonia and diabetes. Treatment includes: treatment plans, which at least include medication, surgery, physical therapy, and lifestyle modifications; Use LLM to automatically extract medical entities from text; Assign a predefined label to each identified entity; Standardize medical terms that are synonyms or have different expressions.

4. The method for mining treatment paths based on dialogue data from internet hospitals according to claim 1, characterized in that, The step of analyzing the contextual relationships between entities using a large language model and extracting logically related sub-chains includes: Use models to analyze the dialogue context and identify potential logical relationships between entities; Subchains with similar or diverse expressions are standardized, retaining frequently occurring subchains and filtering out low-frequency specific subchains; Set a frequency threshold to retain subchains that are widely representative; Remove isolated subchains that lack statistical significance.

5. The method for mining treatment paths based on dialogue data from internet hospitals according to claim 1, characterized in that, The rigorous review and logical rule filtering of the subchains to ensure their medical rationality and clinical significance include: Ensure that the subchain conforms to medical diagnosis and treatment logic; The rationale for the subchain was validated using medical knowledge bases and clinical guidelines; Remove the following abnormal subchain: "symptoms → disease".

6. A system for mining treatment pathways based on dialogue data from internet hospitals, characterized in that, The system includes: The preprocessing unit acquires medical dialogue data and preprocesses the data. The entity labeling unit is designed to guide the large language model to identify and label medical entities in doctor-patient dialogues, assigning a pre-defined label to each medical entity; medical entities are subdivided into four categories: features, examinations, diseases, and treatments; entity sets are divided according to labels; The subchain extraction unit analyzes the contextual relationships between entities using a large language model to extract subchains with logical connections. The review and filtering unit rigorously reviews and filters subchains using logical rules to ensure their medical rationality and clinical significance. The sub-chain fusion unit uses the disease as the core node for reasoning and connection, and expands to both ends to form the main chain. Specifically, it traverses the disease set, merges the check-disease and disease-treatment sub-chains with the same disease, and then, based on the check set, merges the feature-check and check-disease sub-chains with the same check, forming the main chain from feature to check, and then from check to disease to treatment. For the remaining isolated subchains, the similarity between the nodes in the subchain and the nodes with the corresponding labels in the main chain is first calculated. If the similarity exceeds a preset threshold, the large model will then determine whether the subchain can be merged with the main chain. If the similarity does not reach the threshold, the subchain is discarded directly. Finally, a complete diagnosis and treatment process is formed.

7. An electronic device comprising a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the electronic device to execute the steps of the diagnosis and treatment path mining method based on Internet hospital dialogue data as described in any one of claims 1-5.

8. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement each step of the diagnosis and treatment path mining method based on Internet hospital dialogue data as described in any one of claims 1-5.