A method and system for mining diagnosis and treatment paths based on Internet hospital dialogue data
By applying a multi-stage extraction and fusion method based on large language models in medical dialogue data, combined with the review mechanism, the error transmission problem in the mining of diagnosis and treatment process is solved, and high-quality diagnosis and treatment process generation is achieved, which is in line with medical standards.
Patent Information
- Application Number
- CN202510379840.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The prior art has problems with mistransmission in the mining of diagnostic and treatment process in medical dialogue data, and lacks effective methods to identify and extract diagnostic and treatment processes.
A multi-stage extraction and fusion method based on large language model (LLM) is proposed, and combined with the review mechanism, it generates standardized diagnosis and treatment paths by identifying medical entities, analyzing contextual relationships, filtering and fusion sub-chains.
It effectively solves the problem of error transmission in the traditional pipeline framework, realizes high-quality diagnosis and treatment process mining, complies with medical standards, and is highly operable and universal.
Smart Images

Figure CN119905276B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - field of natural language processing and biomedicine, and particularly to a method and system for mining diagnosis and treatment paths based on Internet hospital dialogue data. Background Art
[0002] In modern Internet hospitals, the dialogue data between doctors and patients contains a large amount of diagnosis and treatment knowledge, covering the entire process of information from patient symptom description, examination information, doctor diagnosis to treatment plan. Mining the diagnosis and treatment process using these data can help identify key steps in diagnosis and treatment and provide scientific decision - making support.
[0003] Medical dialogue information extraction is an emerging research direction in the field of medical information extraction, aiming to extract key information (such as clinical terms and their attributes, treatment plans, etc.) from doctor - patient conversations to automatically generate electronic medical records (EMRs), thereby reducing the burden on doctors in writing narrative reports. Early research introduced information extraction methods based on heuristic rules (such as character matching, regular expressions, etc.), but the effect was not ideal. To capture more complex interdependent relationships between entity pairs, recent research aims to enhance information extraction models through external modules. For example, the LogicRE framework based on SOTA rules first learns logical rules according to the output logits of a trained neural model, and then refines the prediction relationships of the neural model through the learned rules; MILR first learns logical rules from annotated data, and then trains a neural model that is penalized by an auxiliary loss for reflecting violations of the learned rules. However, the above two frameworks have error propagation problems due to their pipeline characteristics.
[0004] Large language models (LLMs), with their powerful reasoning and language capabilities and extensive knowledge reserves, have shown great potential in the field of medical information extraction. These models can understand diverse term expressions and perform efficient information extraction according to predefined rules. However, existing research on LLM - based information extraction mainly focuses on entity relationship extraction, and the mining of diagnosis and treatment processes in medical dialogue data still needs to be explored. How to design a reasonable extraction method to extract the diagnosis and treatment process from Internet hospital dialogue data to meet the needs of scientific decision - making support is still a challenging problem.
[0005] In summary, there is an urgent need for a method that can realize the mining of diagnosis and treatment processes based on medical dialogue data to solve the above problems. Summary of the Invention
[0006] In view of this, the present invention proposes a diagnosis and treatment process mining method based on Internet hospital conversation data. First, in order to solve the problem of error propagation in traditional pipeline framework extraction methods, the present invention proposes an LLM-based review mechanism, which performs filtering before the fusion stage according to a priori rules to block the error propagation in the extraction stage. Second, in order to solve the current lack of diagnosis process mining methods based on medical conversations, the present invention proposes a diagnosis and treatment process mining method based on multi-stage extraction and fusion of large models for conversation data, which can identify the physical location and logical relationship of entities in the conversation, extract sub-chains, and then generate high-quality process links through filtering and fusion strategies. A diagnosis and treatment path mining method based on Internet hospital conversation data, comprising the following steps:
[0007] Acquire medical conversation data and pre-process the data; design instructions to guide the large language model to identify and annotate medical entities in doctor-patient conversations, and assign a pre-set label to each medical entity; medical entities are subdivided into four categories in the present invention, namely: characteristics, examinations, diseases and treatments; analyze the contextual relationship between entities through the large language model, and extract sub-chains with logical associations; strictly review and filter the sub-chains according to logical rules to ensure their medical rationality and clinical significance; integrate the sub-chains to generate a standardized diagnosis and treatment path.
[0008] The preprocessing includes: removing noise data: using natural language processing (NLP) technology to eliminate irrelevant content, including system prompts, meaningless pauses, repeated sentences, emoticons, and irrelevant topics; removing sensitive information: desensitizing information involving patient privacy; removing redundant information: merging repeated consultation questions or multiple inquiries about the same symptoms; standardizing expressions: unifying the format of medical terminology; optimizing storage formats: converting cleaned data into a structured format that is easy to analyze.
[0009] The fusion includes: traversing the disease set, fusing the examination-disease and disease-treatment sub-chains with the same disease, and then based on the examination set, fusing the feature-examination and examination-disease sub-chains with the same examination, forming a main chain from feature to examination, and then from examination to disease to treatment; for the remaining isolated sub-chains, first calculate the similarity between the nodes in the sub-chain and the nodes with corresponding labels in the main chain. If the similarity exceeds the preset threshold, the large model is used to determine whether the sub-chain can be fused with the main chain; if the similarity does not reach the threshold, the sub-chain is directly discarded; and finally a complete diagnosis and treatment process is formed.
[0010] The described design instructions guide the large language model to identify and label medical entities in the doctor-patient dialogue, and assign a preset label to each medical entity; medical entities are subdivided into four categories in the present invention, namely: features, examinations, diseases, and treatments; including: classifying medical entities into four categories: features, examinations, diseases, and treatments; features are: symptoms, signs, or clinical manifestations, at least including headache and fever; examinations are: medical examinations or tests, at least including blood tests and CT scans; diseases are: diagnoses made by doctors, at least including pneumonia and diabetes; treatments are: treatment plans, at least including medications, surgeries, physical therapies, lifestyle adjustments; using the LLM to automatically extract medical entities from the text; assigning predefined labels to each identified entity; unifying medical terms with synonyms or different expressions.
[0011] By analyzing the context relationships between entities through the large language model, sub-chains with logical relevance are extracted, including: using the model to analyze the dialogue context and identify potential logical relationships between entities; standardizing sub-chains with similar or diverse expressions, retaining frequently occurring sub-chains, and filtering out low-frequency specific sub-chains. Set a frequency threshold to retain sub-chains with broad representativeness; delete isolated sub-chains lacking statistical significance.
[0012] The sub-chains are strictly reviewed and filtered by logical rules to ensure their medical rationality and clinical significance, including: defining a constraint set C = {symptom → examination, examination → disease, disease → treatment} to ensure that the sub-chains conform to the medical diagnosis and treatment logic; using medical knowledge bases and clinical guidelines to verify the rationality of the sub-chains. Eliminate the following abnormal sub-chains: lack of examination link: directly from "symptom → disease"; wrong order: "treatment → disease".
[0013] The second aspect of the present invention provides a diagnosis and treatment path mining system based on Internet hospital dialogue data, including:
[0014] A preprocessing unit that obtains medical dialogue data and preprocesses the data; an entity marking unit that designs instructions to guide the large language model to identify and label medical entities in the doctor-patient dialogue, and assigns a preset label to each medical entity; medical entities are subdivided into four categories in the present invention, namely: features, examinations, diseases, and treatments; a sub-chain extraction unit that extracts sub-chains with logical relevance by analyzing the context relationships between entities through the large language model; a review and filtering unit that strictly reviews and filters the sub-chains to ensure their medical rationality and clinical significance; a sub-chain fusion unit that fuses each sub-chain to generate a standardized diagnosis and treatment path.
[0015] The third aspect of the present invention provides an electronic device, including: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor invokes the instructions in the memory to enable the electronic device to execute the above-mentioned method for mining the diagnosis and treatment path based on the Internet hospital dialogue data.
[0016] The fourth aspect of the present invention provides a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it enables the computer to execute the above-mentioned method for mining the diagnosis and treatment path based on the Internet hospital dialogue data.
[0017] The present invention has the following beneficial effects:
[0018] The present invention proposes a novel method for mining the diagnosis and treatment process based on Internet hospital dialogue data in the field of medical information extraction, which solves the problem of error propagation in the traditional pipeline framework extraction method. In addition, based on a unique review mechanism and a multi-stage large model extraction fusion strategy, this method effectively mines the diagnosis and treatment process in medical conversations, which not only conforms to the norms in the medical field but also has strong operability and universality. Description of the Drawings
[0019] Figure 1 It is a flowchart of the method for mining the diagnosis and treatment path based on Internet hospital dialogue data provided by an embodiment of the present invention. Detailed Embodiments
[0020] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that illustrated or described here. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0021] Features:
[0022] In medicine, a "feature" refers to a characteristic manifestation of a disease, disorder, or health condition, usually presented as symptoms, signs, or physiological changes. Features can be symptoms reported by the patient themselves (e.g., pain, fatigue) or signs detected by a doctor through examination (e.g., a mass, fever). The features of a disease are the key basis for helping with diagnosis and differentiating between different diseases. Features can also include genes, molecular markers, or imaging findings, etc., which usually contribute to further clinical judgment.
[0023] Examination:
[0024] Examination refers to the systematic assessment of a patient through various clinical means for the purpose of diagnosing, monitoring, and evaluating the progression of a disease or the effectiveness of treatment. An examination can be physical (such as a physical examination), or laboratory or imaging-based (such as a blood test, X-ray, MRI, etc.). Examinations help doctors understand the patient's health status, identify potential pathological changes, and provide a basis for treatment decisions. Examinations generally include the following types:
[0025] Physical examination: Assessment of the patient's physical condition through means such as palpation, auscultation, and inspection.
[0026] Laboratory examination: Analysis of samples such as blood, urine, sputum, etc. through laboratory tests.
[0027] Imaging examination: Using imaging techniques such as X-ray, ultrasound, CT, MRI, etc. to view the condition of internal organs.
[0028] Disease
[0029] It can also be called a diagnosis, which refers to an abnormality or disorder in the physiology, psychology, or function of the body, usually manifested as a set of abnormal symptoms, signs, and / or biochemical indicators, and may be caused by reasons such as infection, genetics, immune disorders, environmental factors, etc. A disease may affect one or more systems of the body, resulting in impairment or damage to healthy functions, and affecting the patient's quality of life and survival.
[0030] Treatment
[0031] It can also be called therapy, which refers to a series of medical interventions and means aimed at alleviating, curing, or managing a disease to improve the patient's health status. Treatment includes various methods such as drug therapy, surgery, physical therapy, psychotherapy, nutritional intervention, and lifestyle changes. The purpose of treatment is to control symptoms, restore function, or prevent the further deterioration of the disease.
[0032] For ease of understanding, the specific process of the embodiment of the present invention is described below. Please refer to Figure 1 , the first embodiment of the method for mining the diagnosis and treatment path based on the dialogue data of the Internet hospital in the embodiment of the present invention includes:
[0033] Obtain medical conversation data and pre-process the data;
[0034] Design instructions to guide the large language model to identify and label medical entities in doctor-patient dialogues, and assign a pre-set label to each medical entity; medical entities are subdivided into four categories in the present invention, namely: characteristics, examinations, diseases, and treatments;
[0035] Analyze the contextual relationship between entities through a large language model and extract sub-chains with logical associations;
[0036] Strictly review and filter the subchains according to logical rules to ensure their medical rationality and clinical significance;
[0037] Integrate each sub-chain to generate a standardized diagnosis and treatment pathway.
[0038] In the data preprocessing stage, the doctor-patient dialogue data in the Internet hospital is firstly cleaned and denoised to ensure the accuracy and effectiveness of the subsequent analysis process. The main task of this stage is to remove redundant information and noise that are meaningless to the analysis, improve data quality, and ensure that the processed data can reflect the real doctor-patient interaction content. Specific operations include: Removing noise data: Identifying and eliminating irrelevant and unstructured content such as system prompts, meaningless pauses, repeated sentences, emoticons, and irrelevant topics through natural language processing technology; Removing sensitive information: In order to protect patient privacy and comply with laws and regulations, sensitive information involving personal privacy in doctor-patient dialogues (such as name, ID number, contact information, etc.) must be desensitized during the preprocessing process. Removing redundant information: Identifying and removing redundant information in doctor-patient dialogues. For example, for repeated consultation questions or situations where the same symptoms are asked multiple times, these duplicate items can be merged or filtered out to reduce information redundancy in the data set and ensure the efficiency and representativeness of the analysis. Through these cleaning and denoising steps, it is ensured that the final data set has high-quality language structure and medical relevance, providing reliable basic data for subsequent entity labeling, path mining, and in-depth analysis.
[0039] In the entity tagging stage, a large language model (LLM) is used to identify and annotate medical entities in the doctor-patient dialogue. This process aims to accurately extract key medical information from natural language text for further analysis and processing. Through the powerful reasoning ability of LLM, multiple medical entities in the text can be identified and a pre-set label can be assigned to each entity. Medical entities are subdivided into four categories in the present invention, namely: features, examinations, diseases and treatments. These categories cover various types of information commonly seen in medical dialogues, as follows: Features: refers to the symptoms, signs or other clinical manifestations described by the patient in the dialogue, such as headache, cough, fever, etc. Features are usually the initial clues for doctors to make a diagnosis. Examinations: include various medical examinations recommended by doctors to confirm or exclude a certain disease, such as blood tests, imaging examinations (such as X-rays, CT scans), physiological tests, etc. Disease: refers to the judgment of the disease or pathological state made by the doctor based on the patient's symptoms, signs and examination results, for example, pneumonia, diabetes, coronary heart disease, etc. Treatment: refers to the treatment plan recommended or implemented by the doctor based on the diagnosis results, including drug therapy, surgery, physical therapy, lifestyle adjustments, etc.
[0040] In the subchain extraction stage, the medical entities in the doctor-patient dialogue are first analyzed in depth through the large language model to explore the potential logical relationship between entities in the context. The model identifies the connection and dependency of entities and extracts subchains with semantic relevance between entities. Furthermore, the term normalization process is performed on the medical entities in the extracted subchains to unify similar or different medical terms into standardized expressions to reduce the diversity and complexity of language expression and improve the accuracy and consistency of subsequent analysis. During the processing, for those subchains that appear frequently, the system will retain one of them to avoid redundancy and ensure the simplicity and efficiency of data; for subchains with fewer repetitions, they are considered to represent individual-specific medical phenomena and may not have broad statistical significance, so they are filtered out. In order to ensure the rationality of this process, the system sets a filtering threshold to screen according to the frequency of subchain occurrence and the universality of the medical facts it represents, ensuring that the subchains that are finally retained have high representativeness and effectiveness, and avoiding the interference of individual anomalies on the analysis results.
[0041] In the review and filtering stage, the extracted sub-chains are strictly screened and filtered according to preset prior rules. The core goal of this process is to ensure that the finally mined diagnosis and treatment path not only conforms to medical common sense but also reflects the real clinical diagnosis and treatment process. To achieve this goal, the present invention adopts a multi-dimensional rule filtering strategy, mainly including the internal constraint relationship of the process and the rationality test of the diagnostic logic. Specifically, the internal constraint set of the process is precisely defined as: C = {symptom → examination, examination → disease, disease → treatment}. This constraint set reflects the typical diagnosis and treatment sequence in the medical field: first, symptoms guide doctors to conduct examinations, then disease diagnoses are made based on the examination results, and finally treatment plans are formulated according to the diagnosis results. The application of the filtering rules requires that the extracted sub-chains must follow this basic logic. In addition, the review and filtering stage also deeply examines the rationality of the diagnostic logic. For example, with the support of medical knowledge bases and clinical guidelines, the system can identify and eliminate those sub-chains that do not conform to medical facts. For example, some sub-chains may show "symptom → disease" without going through the necessary examination steps, or show "treatment → disease" which reverses the diagnosis and treatment process, and these sub-chains will be considered not to conform to the reasonable medical diagnosis and treatment logic and thus be filtered out.
[0042] In the sub-chain fusion stage, the present invention adopts a two-stage fusion strategy. Through step-by-step reasoning and refined processing, it ensures that the finally generated diagnosis and treatment path not only conforms to medical logic but also has a high degree of standardization. First, in the first-stage fusion, reasoning and connection are carried out with the disease as the core node, and the main chain is extended as much as possible to both ends. The system merges the existing diagnosis and treatment chains to simplify the path and make the chains of each disease node more standardized and clear. In the second-stage fusion, first, the similarity between the isolated sub-chain entity and the main-chain entity is calculated. If the similarity exceeds the preset threshold, the large model is further used to determine whether the sub-chain can be fused with the main chain; otherwise, the sub-chain is discarded. For the retained isolated sub-chains, the similarity between the sub-chain nodes and the corresponding label nodes of the main chain is further calculated. When the similarity exceeds the threshold, it is handed over to the large model to determine whether the fusion can be completed. This process effectively eliminates the diagnosis and treatment paths with too strong individual specificity and transforms them into standard paths that conform to general medical practice.
[0043] As a preferred method, the specific steps include:
[0044] S1. In the data preprocessing stage, comprehensively clean and denoise the doctor-patient dialogue data in the Internet hospital, including removing noise data: identifying and removing those unimportant and unstructured contents, such as system prompts, meaningless pauses, emojis, and irrelevant topics, etc.; removing sensitive information: performing desensitization processing on sensitive information (such as names, ID numbers, contact information, etc.) related to personal privacy in doctor-patient dialogues; removing redundant information: identifying and removing redundant information in doctor-patient dialogues, such as in the situation of repeated consultation questions or multiple inquiries about the same symptoms.
[0045] S11. Remove noise data such as emojis, special symbols, and garbled characters in the dialogue; construct a noise word list , and remove the data in the noise word list from the dialogue through regular expressions,
[0046] ,
[0047] where represents the filtered dialogue, represents regular expression filtering of noise, represents the original dialogue, represents the removed content, ;
[0048] S12. Process sensitive information, use an open-source named entity recognition model (NER) for sensitive information detection to detect and replace sensitive information in the dialogue,
[0049] ,
[0050] where represents the dialogue after processing sensitive information, replacing the identified sensitive entities to ensure anonymity;
[0051] S13. Remove redundant information. Based on context similarity, vectorize and represent two adjacent sentences in the dialogue and calculate the cosine similarity,
[0052] ,
[0053] where represents two adjacent dialogues, represents the cosine similarity of, calculate if , then it is determined to be redundant. In the present invention, the value is set to 0.65.
[0054] S2. In the entity tagging stage, use a large language model (LLM) guided by a prompt to identify and annotate medical entities in doctor-patient conversations; identify multiple medical entities in the text based on the LLM, and assign a preset label to each entity; medical entities are subdivided into four categories in the present invention, namely: features, examinations, diseases, and treatments;
[0055] The specific content of the prompt is as follows: "You are a professional doctor, and your job is to add appropriate labels before medical entities in medical consultation conversations;
[0056] Please follow the following key points:
[0057] The label needs to be one of [feature, examination, disease, treatment], and no other label content is allowed;
[0058] Features include symptoms, signs, or other clinical manifestations, examinations include various medical examinations or tests, diseases refer to Western medicine diseases, and treatments include drug treatments, surgeries, physical therapies, lifestyle adjustments, etc.
[0059] Add it before the medical entity in the conversation in the form of <label>, such as 'I have been feeling a bit <feature> dizzy recently';
[0060] Each entity can have at most one label added;
[0061] The following is an example:
[0062] {dialogue}
[0063] # Please add appropriate labels before the medical entities in the medical consultation conversation according to the information provided.";
[0064] S3. In the sub-chain extraction stage, use a large language model to mine the potential logical relationships between entities in the context, identify the connection and dependency relationships of entities, and extract sub-chains with semantic relevance between entities; further, extract the medical entities in the sub-chains, perform term normalization processing, and unify similar or different medical terms into a standardized expression; then, for sub-chains with a repetition count greater than the threshold, only keep one of them; for sub-chains with a repetition count less than the threshold, it is considered that they represent individual-specific medical phenomena and have no broad statistical significance, and they are filtered;
[0065] S31. Use a large language model to mine the potential logical relationships between entities in the context, identify the connection and dependency relationships of entities, and extract sub-chains with semantic relevance between entities;
[0066] The specific content of the prompt is as follows:
[0067] "You are a professional doctor, and your job is to identify medical entities in medical consultation conversations and associate the entities;
[0068] Please follow the following key points:
[0069] There will be <tag> annotations before medical entities. You only need to identify the logical relationships of the entities after the tags in the conversation;
[0070] Identify the process relationships between entities and represent them as chains of [<tagA> entity A - <tagB> entity B]. For example, in 'Patient: Doctor, I've been having a really bad <feature> headache lately. What should I do? \nDoctor: Have you <examination> taken your blood pressure?', you should extract [<feature> headache - <examination> take blood pressure];
[0071] The physical location where an entity appears usually indicates the process relationship in the diagnosis and treatment process.
[0072] You need to consider whether your chain satisfies the logical relationship in the conversation and whether it conforms to medical common sense. This is very important;
[0073] The following is an example:
[0074] {dialogue}
[0075] # Please extract the entity sub-chain in the medical consultation conversation according to the information provided.";
[0076] S32. Divide the entity set by tags and merge the entities in each entity set based on the instruction prompt for the large model; for example, in the disease set, "knee arthritis" and "hip arthritis" will be merged into "osteoarthritis";
[0077] The specific content of the prompt is as follows:
[0078] "You are a professional doctor, and your job is to identify the medical entities in the list and merge them into a smaller list;
[0079] Please follow the following key points:
[0080] The medical entities in the list are all nouns representing {tag};
[0081] There may be cases where the entities have the same meaning but different noun representations. You can merge them. For example, "knee arthritis" and "hip arthritis" will be merged into "osteoarthritis";
[0082] You can appropriately expand the scope of the noun explanation to improve your merging efficiency;
[0083] The medical list is as follows:
[0084] {list}"
[0085] S33. Unify the entity terms in the sub-chain into a standardized expression based on the large model with prompts;
[0086] The specific content of the prompt is as follows:
[0087] "You are a professional doctor, and your job is to map the medical entities in the sub-chain into standardized expressions; note that you can only select the standardized representation closest to the entity in the current sub-chain from the given list;
[0088] # Standardized representation list of {tag1}: {tag1_list}
[0089] # Standardized representation list of {tag2}: {tag2_list}
[0090] # Input:
[0091] {Sub-chain to be converted}
[0092] ## The following is an example for your reference
[0093] {Example}"
[0094] S34. For sub-chains with a repetition count greater than the threshold, only keep one of them; for sub-chains with a repetition count less than the threshold, they are considered to represent individual-specific medical phenomena and have no broad statistical significance, so they are filtered out.
[0095] S4. The present invention adopts a multi-dimensional rule filtering strategy, mainly including the constraint relationship within the process and the rationality test of the diagnostic logic. Define the constraint set within the process:
[0096] ,
[0097] Among them, represents Symptom, that is, the patient's main complaint or manifestation. represents Examination, including clinical examinations, imaging examinations, laboratory examinations, etc., represents Disease, the disease diagnosed through examinations and analyses, represents Treatment, the intervention measures taken for the diagnosed disease. It is required that the extracted sub-chain must follow one of the constraints. In addition, with the support of medical knowledge bases and clinical guidelines, the rationality of the diagnostic logic is deeply tested. Identify and eliminate those sub-chains that do not conform to medical facts.
[0098] S41. Define the constraint set within the process and verify whether the sub-chain follows the constraints based on regular expressions,
[0099] ,
[0100] Among them, represents the filtered sub-chain, represents filtering out sub-chains that do not meet the constraints according to regular expressions, represents the sub-chain to be filtered, represents the set of constraints.
[0101] S42. Use the instruction prompt large model to further perform semantic and content-level verification on the sub-chain based on meeting the process constraints and in combination with the medical knowledge base k:
[0102] ,
[0103] Among them is the reply of the large model, represents retrieving knowledge from the medical knowledge base.
[0104] The specific content of the prompt is as follows:
[0105] "You are a professional doctor, and your job is to judge whether the sub-chain conforms to medical facts; you will be given some relevant documents, and please make a professional judgment with reference to the documents;
[0106] # Document {Retrieved medical knowledge}
[0107] # Input:
[0108] {Sub-chain to be judged}
[0109] ## Please judge whether this sub-chain conforms to medical facts. Note that the output can only be [Conforms / Does not conform], and any extra or incorrect output may cause the system to crash."
[0110] S5. In the sub-chain fusion stage, adopt a two-stage fusion strategy. In the first-stage fusion process, use the disease as the core node for reasoning connection, expand the chain as much as possible to both ends to form the main chain. In the second-stage fusion process, first calculate the similarity between the isolated sub-chain entity and the entity in the main chain. If the similarity is greater than the threshold, use the large model to further judge whether the sub-chain can be fused, otherwise discard the sub-chain;
[0111] S51. Traverse the disease set, fuse the examination-disease and disease-treatment sub-chains with the same disease, and then based on the examination set, fuse the feature-examination and examination-disease sub-chains with the same examination to form the main chain from features to examinations, then from examinations to diseases to treatments.
[0112] S52. For the remaining isolated sub-chains, first calculate the similarity between the nodes in the sub-chain and the nodes with corresponding labels in the main chain:
[0113] ,
[0114] where represents a node in the sub-chain, represents a node in the main chain with the same label. If the similarity exceeds the threshold , it is handed over to the large model to determine whether the sub-chain can be fused with the main chain; if the similarity does not reach the threshold, the sub-chain is directly discarded; in the present invention, here is set to 0.65.
[0115] The specific content of the large model prompt is as follows:
[0116] "You are a professional doctor, and your job is to determine whether two chains can be fused;
[0117] # Input:
[0118] {Sub-chain 1 to be judged}
[0119] {Sub-chain 2 at the access point in the main chain}
[0120] ## Determine whether the two chains can be fused. Note that the output can only be [Yes / No], and any extra or incorrect output may cause the system to crash."
[0121] The above describes the diagnosis and treatment path mining method based on Internet hospital dialogue data in the embodiments of the present invention. Next, the diagnosis and treatment path mining device based on Internet hospital dialogue data in the embodiments of the present invention will be described:
[0122] The embodiments of the present invention also provide an electronic device, which may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPUs) (for example, one or more processors) and a memory, and one or more storage media for storing applications or data (for example, one or more mass storage devices). Among them, the memory and the storage media can be short-term storage or persistent storage. The program stored in the storage media may include one or more modules, and each module may include a series of instruction operations on the electronic device. Further, the processor may be configured to communicate with the storage media and execute a series of instruction operations in the storage media on the electronic device.
[0123] The electronic device may further include one or more power supplies, one or more wired or wireless network interfaces, one or more input / output interfaces, and / or one or more operating systems, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that the structure of the electronic device in this embodiment does not constitute a limitation on the electronic device, and it may include more or fewer components, or combine certain components, or have different component arrangements.
[0124] An electronic device structure provided by an embodiment of the present invention may vary greatly due to different configurations or performances. It may include one or more processors (central processing units, CPUs) (for example, one or more processors) and a memory, and one or more storage media for storing applications or data (for example, one or more mass storage devices). Among them, the memory and the storage media may be transient storage or persistent storage. The program stored in the storage media may include one or more modules, and each module may include a series of instruction operations on the electronic device. Further, the processor may be configured to communicate with the storage media and execute a series of instruction operations in the storage media on the electronic device.
[0125] The electronic device may further include one or more power supplies, one or more wired or wireless network interfaces, one or more input / output interfaces, and / or one or more operating systems, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that the electronic device structure does not constitute a limitation on the electronic device, and it may include more or fewer components than the foregoing, or combine certain components, or have different component arrangements.
[0126] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the foregoing method.
[0127] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, or units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0128] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0129] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A diagnosis and treatment path mining method based on Internet hospital conversation data, characterized in that: The following steps are involved: Obtain medical conversation data and pre-process the data; Design instructions to guide the large language model to identify and label medical entities in doctor-patient conversations, assigning a pre-set label to each medical entity; subdividing medical entities into four categories: characteristics, examinations, diseases, and treatments; dividing entity sets by labels; Analyze the contextual relationship between entities through a large language model and extract sub-chains with logical associations; Strictly review and filter the subchains according to logical rules to ensure their medical rationality and clinical significance; Reasoning and connection are performed with diseases as the core nodes, and the main chain is expanded to both ends. Specifically, the disease set is traversed, and the examination-disease and disease-treatment sub-chains with the same disease are integrated. Then, based on the examination set, the feature-examination and examination-disease sub-chains with the same examination are integrated to form a main chain from feature to examination, and then from examination to disease to treatment. For the remaining isolated sub-chains, the similarity between the nodes in the sub-chain and the nodes with corresponding labels in the main chain is first calculated. If the similarity exceeds the preset threshold, the large model will determine whether the sub-chain can be merged with the main chain; if the similarity does not reach the threshold, the sub-chain will be directly discarded; finally, a complete diagnosis and treatment process is formed.
2. According to claim 1, a diagnosis and treatment path mining method based on Internet hospital conversation data is characterized in that: The pre-processing comprises: Remove noisy data: Use natural language processing (NLP) technology to remove irrelevant content, including system prompts, meaningless pauses, repeated sentences, emoticons, and irrelevant topics; Remove sensitive information: Desensitize information involving patient privacy; Remove redundant information: merge repeated consultation questions or content that asks about the same symptoms multiple times; Standardized expression: unified medical terminology format; Storage format optimization: Convert the cleaned data into a structured format.
3. According to claim 1, a diagnosis and treatment path mining method based on Internet hospital conversation data is characterized in that: The design instructions guide the large language model to identify and label medical entities in the doctor-patient dialogue and assign a pre-set label to each medical entity; including: Medical entities are divided into four categories: characteristics, examinations, diseases, and treatments; Characterized by: symptoms, signs or clinical manifestations, including at least headache and fever; Examinations are: medical examinations or tests, including at least blood tests and CT scans; Diseases are: diagnoses made by doctors, including at least pneumonia and diabetes; Treatment is: a treatment plan that includes at least medication, surgery, physical therapy, and lifestyle adjustments; Automatically extract medical entities from text using LLM; Assign a predefined label to each identified entity; Unify medical terms that are synonymous or have different expressions.
4. According to claim 1, a diagnosis and treatment path mining method based on Internet hospital conversation data is characterized in that: The process of analyzing the contextual relationship between entities through a large language model and extracting subchains with logical associations includes: Use the model to analyze the conversation context and identify the potential logical relationships between entities; Standardize similar or diversely expressed subchains, retain frequently occurring subchains, and filter low-frequency specific subchains; Set frequency thresholds to retain subchains that are broadly representative; Isolated subchains lacking statistical significance were removed.
5. According to claim 1, a diagnosis and treatment path mining method based on Internet hospital conversation data is characterized in that: The subchains are strictly reviewed and filtered by logical rules to ensure their medical rationality and clinical significance, including: Ensure that the subchain complies with medical diagnosis and treatment logic; Use medical knowledge base and clinical guidelines to verify the rationality of subchains; Eliminate the following abnormal sub-chain: "symptom→disease".
6. A diagnosis and treatment path mining system based on Internet hospital conversation data, characterized in that: The system comprises: A preprocessing unit, which obtains medical conversation data and preprocesses the data; The entity tagging unit designs instructions to guide the large language model to identify and label medical entities in doctor-patient conversations, assigning a pre-set label to each medical entity; medical entities are subdivided into four categories: characteristics, examinations, diseases, and treatments; entity sets are divided by labels; The subchain extraction unit analyzes the contextual relationship between entities through a large language model and extracts subchains with logical associations; Review and Filter Unit: strictly review and filter the subchains according to logical rules to ensure their medical rationality and clinical significance; The subchain fusion unit uses the disease as the core node for reasoning and connection, and expands to both ends to form a main chain. Specifically, it traverses the disease set, fuses the examination-disease and disease-treatment subchains with the same disease, and then, based on the examination set, fuses the feature-examination and examination-disease subchains with the same examination, forming a main chain from feature to examination, and then from examination to disease to treatment; For the remaining isolated sub-chains, the similarity between the nodes in the sub-chain and the nodes with corresponding labels in the main chain is first calculated. If the similarity exceeds the preset threshold, the large model will determine whether the sub-chain can be merged with the main chain; if the similarity does not reach the threshold, the sub-chain will be directly discarded; finally, a complete diagnosis and treatment process is formed.
7. An electronic device, comprising a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the electronic device executes each step of the diagnosis and treatment path mining method based on Internet hospital conversation data as described in any one of claims 1-5.
8. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the various steps of the diagnosis and treatment path mining method based on Internet hospital conversation data as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Disease knowledge map construction method and platform system, device, and storage medium
CN109271530A
DRG high-compilation high-suit behavior identification method and device, electronic equipment and medium
CN117786123A