Patient information generation method and device based on questions and answers

By introducing a question-and-answer method in the health pre-diagnosis system, combining knowledge graphs and generative large language models, the shortcomings of the existing system in symptom recognition and disease reasoning are solved, and more efficient and accurate medical services and better user experience are achieved.

CN120072253APending Publication Date: 2025-05-30GUANGDONG KAMFU TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411913627.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing health prediagnosis methods have shortcomings in the accuracy of symptom recognition, the depth of disease reasoning and user experience, resulting in low system efficiency and patient satisfaction.

Method used

A question-and-answer method is adopted, combining knowledge graph entity alignment, generative large language model and dynamic problem generation to improve the accuracy and user experience of pre-diagnosis methods.

Benefits of technology

By accurately converting and aligning the patient's natural language descriptions, highly relevant symptoms collections are generated, improving the accuracy and efficiency of disease reasoning, and providing a more intelligent, precise and personalized user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072253A_ABST
    Figure CN120072253A_ABST
Patent Text Reader

Abstract

The invention discloses a patient information generation method and device based on questions and answers. The method comprises the steps that patient condition information is converted into medical terms; the medical terms belong to a language library in a medical knowledge graph; aligning the medical terms according to a conversion condition to obtain an aligned medical knowledge graph entity; generating a first candidate symptom set according to the aligned medical knowledge graph entity; determining a first candidate disease set and a second candidate symptom set corresponding to the first candidate disease set according to the first patient question and answer information; the second candidate symptom set is greater than or equal to the first candidate symptom set; generating a health question set according to the second patient question and answer information; and according to answer information of the patient to the health question set, generating patient health information. By adopting the technical means of combining knowledge graph entity alignment, language large model and dynamic problem generation, the accuracy of the pre-inquiry method is improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical data processing, and particularly to a method and device for generating patient information based on question and answer. Background Art

[0002] Existing health pre-consultation methods face multiple technical challenges and user experience problems in practical applications, which limit the effectiveness of the system and patient satisfaction. The following are the existing problems described comprehensively:

[0003] First, the accuracy of symptom recognition is insufficient: The symptom information described by patients is usually in natural language form, containing a large number of colloquial, vague or inaccurate expressions. These non-standardized descriptions make it difficult for the system to accurately extract key symptom information. There is a significant difference between the professionalism of medical terms and the daily language of patients, resulting in the system's difficulty in understanding and correctly converting the patient's description into standard medical terms. Traditional pre-consultation systems often lack the understanding of conversation history and context information and cannot accurately capture the subtle changes in the patient's continuous description. Second, the depth and accuracy of disease inference are insufficient: Existing pre-consultation systems usually rely on simple rule matching to infer diseases and lack in-depth understanding and application of medical knowledge. This shallow inference is prone to misjudgment or missed diagnosis.

[0004] In addition, the interaction method of traditional pre-consultation is relatively fixed and cannot dynamically adjust questions according to the specific situation of patients, resulting in poor user experience. For example, unnecessary repeated questions may be asked or important information may be missed. Summary of the Invention

[0005] Embodiments of the present invention provide a method and device for generating patient information based on question and answer, which adopt a combination of knowledge graph entity alignment, large language model and dynamic question generation to improve the accuracy of the pre-consultation method and enhance the user experience.

[0006] To achieve the above object, the first aspect of the embodiments of the present application provides a method for generating patient information based on question and answer, including:

[0007] Converting the patient's condition information into medical terms; the medical terms belong to the corpus in the medical knowledge graph;

[0008] Aligning the medical terms according to conversion conditions to obtain aligned medical knowledge graph entities;

[0009] Generating a first candidate symptom set according to the aligned medical knowledge graph entities;

[0010] Determining a first candidate disease set and a second candidate symptom set corresponding to the first candidate disease set according to the first patient question-and-answer information; the second candidate symptom set is greater than or equal to the first candidate symptom set;

[0011] Generate a set of health questions based on the second patient's Q&A information;

[0012] Generate patient health information based on the patient's answers to the set of health questions.

[0013] In a possible implementation manner of the first aspect, before converting the patient's condition information into medical terms, it further includes:

[0014] Obtain the patient's self-reported information;

[0015] Judge the rationality of the patient's self-reported information according to a preset rule. If the patient's self-reported information meets the preset rule, use the patient's self-reported information as the patient's condition information.

[0016] In a possible implementation manner of the first aspect, converting the patient's condition information into medical terms specifically includes:

[0017] Apply a generative large language model to convert the patient's condition information into medical terms.

[0018] In a possible implementation manner of the first aspect, aligning the medical terms according to the conversion conditions to obtain aligned medical knowledge graph entities specifically includes:

[0019] Calculate the similarity between the medical terms and all medical knowledge graph entities;

[0020] Select the medical knowledge graph entity with the largest similarity as the target medical knowledge graph entity;

[0021] If the maximum similarity value is greater than a preset similarity threshold, use the target medical knowledge graph entity as the aligned medical knowledge graph entity; if the maximum similarity value is less than or equal to the preset similarity threshold, perform a synonym judgment according to the generative large language model, and use the target medical knowledge graph entity as the aligned medical knowledge graph entity after passing the judgment.

[0022] In a possible implementation manner of the first aspect, generating the first candidate symptom set according to the aligned medical knowledge graph entity specifically includes:

[0023] According to the aligned medical knowledge graph entity, determine several other medical knowledge graph entities in the medical knowledge graph whose shortest path lengths do not exceed a preset distance threshold;

[0024] Select several of the other medical knowledge graph entities as the first candidate symptom set according to the occurrence frequencies of the other medical knowledge graph entities.

[0025] In a possible implementation of the first aspect, selecting a number of the other medical knowledge graph entities as the first candidate symptom set according to the occurrence frequencies of the other medical knowledge graph entities specifically includes:

[0026] Arrange all the other medical knowledge graph entities in descending order according to the occurrence frequencies of the other medical knowledge graph entities; if there are other medical knowledge graph entities with equal occurrence frequencies, count the number of occurrences of the other medical knowledge graph entities in various diseases in the medical knowledge graph, and the other medical knowledge graph entities with larger numbers of occurrences are ranked earlier in the descending order;

[0027] Select a preset number of other medical knowledge graph entities starting from the first order according to the order of the arrangement result.

[0028] In a possible implementation of the first aspect, determining the first candidate disease set and the second candidate symptom set corresponding to the first candidate disease set according to the first patient's Q&A information specifically includes:

[0029] Obtain the first patient's Q&A information after the patient selects the first candidate symptom set;

[0030] Match the first patient's Q&A information through the medical knowledge graph to determine the set of self-reported symptoms corresponding to the first patient's Q&A information and the sets of symptoms corresponding to several suspected diseases;

[0031] Calculate the IOU value of each suspected disease according to the set of self-reported symptoms corresponding to the patient and the sets of symptoms corresponding to several suspected diseases; the IOU value is the ratio of the number of symptom intersections between two sets to the number of symptom unions;

[0032] Determine the first candidate disease set and the second candidate symptom set corresponding to the first candidate disease set according to the IOU value.

[0033] In a possible implementation of the first aspect, generating a set of health problems according to the second patient's Q&A information specifically includes:

[0034] Obtain the second patient's Q&A information after the patient selects the second candidate symptom set;

[0035] Process the second patient's Q&A information using a language model to generate disease inference questions;

[0036] Obtain the third patient's Q&A information after the patient selects the disease inference questions;

[0037] Judging the score threshold of the third patient's Q&A information through the medical knowledge graph, generating corresponding disease questions for all diseases with scores exceeding the score threshold, and forming a disease question set;

[0038] Regarding the disease question set, the preset allergy history question set, and the preset medical history question set as the health question set.

[0039] In a possible implementation manner of the first aspect, the generating of the patient health information according to the patient's answer information to the health question set specifically includes:

[0040] Obtaining the patient's answer information to the health question set;

[0041] Using natural language generation technology to process the answer information and generating a patient health information report including a diagnosis suggestion and a health suggestion.

[0042] A second aspect of the embodiments of the present application provides a patient information generation device based on Q&A, including:

[0043] A conversion module, configured to convert patient condition information into medical terms; the medical terms belong to the corpus in the medical knowledge graph;

[0044] An alignment module, configured to align the medical terms according to conversion conditions to obtain aligned medical knowledge graph entities;

[0045] A first set generation module, configured to generate a first candidate symptom set according to the aligned medical knowledge graph entities;

[0046] A second set generation module, configured to determine a first candidate disease set and a second candidate symptom set corresponding to the first candidate disease set according to the first patient Q&A information; the second candidate symptom set is greater than or equal to the first candidate symptom set;

[0047] A question generation module, configured to generate a health question set according to the second patient Q&A information;

[0048] An information generation module, configured to generate patient health information according to the patient's answer information to the health question set.

[0049] Compared with the prior art, a method and device for generating patient information based on question and answer provided by an embodiment of the present invention use a generative large language model specially trained in the medical field to convert the patient's colloquial or non-standardized descriptions into standard medical terms; during the process of aligning medical terms, it is ensured that the medical terms can accurately match the corresponding entities in the knowledge graph, and a set of symptoms highly relevant to the patient's initial description (the first candidate symptom set) is generated based on the aligned medical knowledge graph entities to help the patient supplement important information that may be missing. Through multiple rounds of interaction, the scope of suspected diseases is gradually narrowed, and the accuracy of disease inference is improved.

[0050] By making multi-level use of the medical knowledge graph, the limitations of the existing pre-consultation system in aspects such as symptom recognition, disease inference, and user experience are solved. It not only improves the efficiency and accuracy of medical services but also provides a more intelligent, precise, and personalized experience for patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a schematic flowchart of a method for generating patient information based on question and answer provided by an embodiment of the present invention;

[0052] Figure 2 is a schematic diagram of the operation of a medical consultation system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] To solve the above problems, please refer to Figure 1 , an embodiment of the present invention provides a method for generating patient information based on question and answer, including:

[0055] S10. Convert the patient's condition information into medical terms; the medical terms belong to the corpus in the medical knowledge graph.

[0056] S11. Align the medical terms according to the conversion conditions to obtain aligned medical knowledge graph entities.

[0057] S12. Generate a first candidate symptom set according to the aligned medical knowledge graph entities.

[0058] S13. Determine a first set of candidate diseases and a second set of candidate symptoms corresponding to the first set of candidate diseases according to the first patient's Q&A information; the second set of candidate symptoms is greater than or equal to the first set of candidate symptoms.

[0059] S14. Generate a set of health questions according to the second patient's Q&A information.

[0060] S15. Generate patient health information according to the patient's response information to the set of health questions.

[0061] S10 converts the patient's natural language description into professional medical terms to improve the accuracy and consistency of symptom recognition; in S11, the medical terms are aligned according to the conversion conditions to ensure that the medical terms can accurately match the corresponding entities in the knowledge graph, providing a reliable basis for subsequent symptom association and disease reasoning; S12 generates a set of symptoms (the first set of candidate diseases) highly relevant to the patient's initial description to help the patient supplement important information that may be missed and reduce unnecessary questions; after the patient selects the symptoms in the first set of candidate symptoms, in S13, the patient's answers are recorded, and according to the answers, the first set of candidate diseases and its corresponding second set of candidate symptoms are determined. The second set of candidate symptoms contains more potential symptoms. Through one round of interaction, the scope of suspected diseases is narrowed, the accuracy of disease reasoning is improved, and more relevant symptoms are provided for the patient to confirm. After the patient selects the symptoms in the second set of candidate symptoms, the answers are recorded as the second patient's Q&A information. In S14, through dynamically generated questions, the possibility of the disease is further confirmed, and necessary medical history and allergy information are collected to ensure the comprehensiveness and accuracy of the information. Finally, in S15, natural language generation technology is used to process the patient's answers to generate a personalized health information report including diagnosis suggestions and health suggestions.

[0062] The above method solves the limitations of existing pre-consultation systems in aspects such as symptom recognition, disease reasoning, user experience, and report generation. It not only improves the efficiency and accuracy of medical services but also provides a more intelligent, precise, and personalized experience for patients.

[0063] Exemplarily, before converting the patient's condition information into several medical terms, it further includes:

[0064] Obtain the patient's self-reported information.

[0065] Judge the rationality of the patient's self-reported information according to preset rules. If the patient's self-reported information meets the preset rules, use the patient's self-reported information as the patient's condition information.

[0066] Patients can input their self-reported information in various ways, including but not limited to text boxes, voice input, etc. The system should support multimodal input to meet the needs of different patients. Patient self-reported information usually includes health-related information such as symptom descriptions, medical histories, past medical histories, allergy histories, and living habits. For example, a patient may describe, "I've been having headaches lately, accompanied by nausea, which has lasted for three days. I had migraines before, but they're fine now, and I have no drug allergy history." This can collect all relevant information provided by the patient and provide comprehensive data support for subsequent processing.

[0067] A series of rules are preset to judge the reasonableness and completeness of the patient's self-reported information. These rules can be based on medical common sense, clinical guidelines, and the internal logic of the system. This is to ensure the accuracy and completeness of the input information and avoid deviations in subsequent processing due to incomplete or unreasonable information. Through a reasonable judgment mechanism, the quality of the data can be improved, thereby enhancing the overall diagnostic accuracy.

[0068] By obtaining the patient's self-reported information and making a reasonableness judgment, the accuracy and completeness of the input information can be ensured, providing a reliable basis for subsequent symptom recognition and disease reasoning. This process not only improves the quality of the data but also helps patients better understand how to describe their conditions, thereby enhancing the efficiency and accuracy of the entire question-and-answer-based patient information generation method.

[0069] Exemplarily, the conversion of the patient's condition information into several medical terms specifically includes:

[0070] Applying a generative large language model to convert the patient's condition information into medical terms.

[0071] Exemplarily, the alignment of the medical terms according to the conversion conditions to obtain aligned medical knowledge graph entities specifically includes:

[0072] Calculating the similarity between the medical terms and all medical knowledge graph entities;

[0073] Selecting the medical knowledge graph entity with the maximum similarity as the target medical knowledge graph entity;

[0074] If the maximum similarity value is greater than the preset similarity threshold, the target medical knowledge graph entity is used as the aligned medical knowledge graph entity; if the maximum similarity value is less than or equal to the preset similarity threshold, a synonym judgment is made according to the generative large language model, and after passing the judgment, the target medical knowledge graph entity is used as the aligned medical knowledge graph entity.

[0075] Self-reported information from patients after rationality judgment, such as "I have been having headaches and nausea recently, which has lasted for three days. I had migraines before, but I am fine now. I have no history of drug allergies."

[0076] Generative large language models (such as BERT, T5, etc.) that have been specially trained in the medical field can be used to convert the patient's natural language description into standard medical terms. These terms belong to the lexicon in the medical knowledge graph.

[0077] In practical applications, the patient's self-reported information can first be preprocessed, including word segmentation, stop word removal, word form restoration and other operations, to ensure that the text format input into the model is standardized.

[0078] Because the generative large language model can not only recognize individual words, but also understand the overall semantics of the sentence, it can more accurately convert the patient's description into medical terms. For example, "headache" may be converted to "headache", "nausea" may be converted to "nausea", and "migraine" may be converted to "migraine". For words with multiple meanings, the model will select the most appropriate medical term based on the context. For example, "apple" may refer to fruit or electronic devices in different contexts, and the model will select the correct term based on the patient's description. If the patient's description contains more vague expressions, the model can infer more specific medical terms based on the context and existing medical knowledge. For example, "dizziness" may be expanded to "dizziness" or "vertigo", depending on the description of other symptoms.

[0079] The above sub-process converts the patient's natural language description into professional medical terms, improves the accuracy and consistency of symptom identification, and ensures that subsequent processing can be carried out based on standardized language.

[0080] The converted medical terms are similar to all entities in the medical knowledge graph. Similarity calculation can be achieved through a variety of methods, such as cosine similarity, Jaccard coefficient, etc. If cosine similarity is used, then the medical knowledge graph entity with the greatest similarity is selected as the target entity based on the similarity score. If the maximum similarity is greater than the preset similarity threshold, the target entity is directly used as the aligned medical knowledge graph entity; if the maximum similarity is less than or equal to the preset similarity threshold, further synonym judgment is performed.

[0081] If the maximum similarity is less than or equal to a preset similarity threshold, the system will further use a generative large language model to determine synonymy. For example, if the similarity between "headache" and "migraine" is 0.6, which is lower than the preset threshold of 0.7, the system will use the model to determine whether these two terms are synonyms. If the judgment result is synonymous, "migraine" will still be used as the aligned medical knowledge graph entity.

[0082] The above sub - process ensures that medical terms can be accurately matched to the corresponding entities in the knowledge graph, avoiding misjudgment or missed diagnosis caused by inconsistent terms. Through similarity calculation and synonym judgment, the system can establish correct associations between different expressions and improve the accuracy of alignment.

[0083] In summary, by applying a generative large language model to convert patient condition information into medical terms and aligning these terms according to the conversion conditions, the system can ensure the standardization and accuracy of the input information. This process not only improves the accuracy of symptom recognition but also provides a reliable basis for subsequent disease reasoning. Through similarity calculation and synonym judgment, correct associations can be established between different expressions, avoiding misjudgment or missed diagnosis caused by inconsistent terms. Ultimately, it helps to improve the efficiency and accuracy of the patient information generation method based on question - answering, thus providing more intelligent and precise medical services for patients.

[0084] Exemplarily, generating the first candidate symptom set according to the aligned medical knowledge graph entity specifically includes:

[0085] According to the aligned medical knowledge graph entity, determine several other medical knowledge graph entities in the medical knowledge graph whose shortest path lengths do not exceed a preset distance threshold;

[0086] Select several of the other medical knowledge graph entities as the first candidate symptom set according to the occurrence frequencies of the other medical knowledge graph entities.

[0087] In the patient information generation method based on question - answering, generating the first candidate symptom set according to the aligned medical knowledge graph entity is one of the key steps. This process aims to provide a set of most likely relevant symptoms for the patient through the association relationships in the medical knowledge graph to help the system more accurately identify potential diseases.

[0088] In a medical knowledge graph, centered on the aligned entities, find other entities whose shortest path length to them does not exceed a preset distance threshold. Here, the "shortest path" refers to the minimum number of edges or nodes passed between one entity and another. The interrogation system can preset a distance threshold (such as 2 or 3), indicating the maximum distance to other entities in the knowledge graph that can be reached starting from the aligned entity. This threshold can be adjusted according to the application scenario and requirements. For example, a smaller threshold can ensure that the found entities are highly relevant to the aligned entity, while a larger threshold can extend to a wider range of associations.

[0089] By querying the historical data in the medical knowledge graph, the co-occurrence times of each entity in different diseases can be counted. For example, "nausea" and "migraine" often appear together in the medical records of migraine patients, so their co-occurrence frequency is relatively high. In addition, the system can also refer to resources such as medical literature and clinical guidelines to obtain more information about the entity occurrence frequency. By selecting a number of the most relevant entities according to the occurrence frequency, the interrogation system of this embodiment can generate a first candidate symptom set that is both representative and not overly verbose, helping the patient to confirm whether there are other related symptoms, so as to provide more comprehensive data support for subsequent disease reasoning.

[0090] Exemplarily, the step of selecting a number of the other medical knowledge graph entities as the first candidate symptom set according to the occurrence frequency of the other medical knowledge graph entities specifically includes:

[0091] Sort all the other medical knowledge graph entities in descending order according to the occurrence frequency of the other medical knowledge graph entities; if there are other medical knowledge graph entities with equal occurrence frequencies, count the occurrence times of the other medical knowledge graph entities in various diseases in the medical knowledge graph, and the other medical knowledge graph entities with larger occurrence times are ranked higher in the descending order;

[0092] Select a preset number of the other medical knowledge graph entities starting from the first order according to the order of the sorting result.

[0093] In some cases, there may be multiple other entities with the same occurrence frequency. To further distinguish these entities, it is necessary to count their occurrence times in various diseases in the medical knowledge graph and perform a secondary sorting according to the occurrence times.

[0094] For entities with the same occurrence frequency, the system counts their co-occurrence times in different types of diseases. For example, assume the occurrence frequencies of "dizziness" and "lightheadedness" are both 60%, but "dizziness" frequently appears in multiple diseases such as migraine and vertigo, while "lightheadedness" mainly appears in a few diseases such as low blood pressure. Then, the occurrence times of "dizziness" are larger.

[0095] These entities are sorted again according to the occurrence times to ensure that entities with larger occurrence times are ranked in the front. For example, assume the occurrence frequencies of "dizziness" and "lightheadedness" are both 60%, but "dizziness" appears in 10 diseases, while "lightheadedness" only appears in 5 diseases. Then, "dizziness" will be ranked before "lightheadedness".

[0096] By sorting in descending order according to the occurrence frequencies of other medical knowledge graph entities and performing a secondary sort based on disease relevance when the occurrence frequencies are equal, it can be ensured that the generated first candidate symptom set not only has a higher frequency but also has a stronger disease relevance. Finally, according to the order of the sorting results, a preset number of other entities are selected starting from the first order to generate a first candidate symptom set that is both representative and not overly lengthy. This process not only improves the accuracy of symptom recognition but also provides a reliable basis for subsequent disease reasoning.

[0097] Exemplarily, the determining of the first candidate disease set and the second candidate symptom set corresponding to the first candidate disease set according to the first patient's Q&A information specifically includes:

[0098] Obtain the first patient's Q&A information after the patient selects the first candidate symptom set;

[0099] Match the first patient's Q&A information through the medical knowledge graph to determine the set of symptoms described by the patient corresponding to the first patient's Q&A information and the sets of symptoms corresponding to several suspected diseases;

[0100] According to the set of symptoms described by the patient and the sets of symptoms corresponding to several suspected diseases, calculate the IOU value of each suspected disease; the IOU value is the ratio of the number of symptom intersections between two sets to the number of symptom unions;

[0101] Determine the first candidate disease set and the second candidate symptom set corresponding to the first candidate disease set according to the IOU value.

[0102] The medical interview system applying this embodiment presents the first set of candidate symptoms to the patient and asks the patient to confirm whether these symptoms actually exist. For example, the system may ask, "Do you have the following symptoms? Nausea, vomiting, dizziness, photophobia, phonophobia." The patient can choose "Yes" or "No" according to the actual situation. The system records the patient's selection result to form the first patient Q&A information. The key to this step is to ensure that the information provided by the patient is as accurate as possible to avoid subsequent reasoning deviations caused by misunderstandings or false reports.

[0103] The system can use the entity relationships in the medical knowledge graph to match the symptoms in the first patient Q&A information with the medical terms in the knowledge graph to determine the set of symptoms self-reported by the patient. At the same time, the system can also search for suspected diseases related to these symptoms and extract the set of typical symptoms for each suspected disease. The system searches for suspected diseases related to these symptoms in the medical knowledge graph according to the set of symptoms self-reported by the patient. For example, assuming the patient selects "nausea", "dizziness", and "photophobia", the system may find the following suspected diseases:

[0104] Migraine, Vertigo, Hypotension, Anemia, and Depression.

[0105] For each suspected disease, the system extracts its set of typical symptoms. For example, the set of typical symptoms of migraine may include "nausea", "vomiting", "dizziness", "photophobia", "phonophobia", etc.

[0106] By matching the first patient Q&A information with the medical knowledge graph, the system can determine the set of symptoms self-reported by the patient and find several suspected diseases related to these symptoms and their corresponding sets of symptoms.

[0107] It should be noted that Intersection over Union (IOU) is a commonly used metric to measure the similarity between two sets. Here, the IOU value is used to measure the similarity between the set of symptoms self-reported by the patient and the set of symptoms corresponding to each suspected disease. Specifically, the IOU value is equal to the ratio of the number of symptom intersections between the two sets to the number of symptom unions.

[0108] Suppose the set of symptoms self-reported by the patient is {"nausea", "dizziness", "photophobia"}, and the set of typical symptoms of migraine is {"nausea", "vomiting", "dizziness", "photophobia", "phonophobia"}. Then the intersection is {"nausea", "dizziness", "photophobia"}, and the union is {"nausea", "vomiting", "dizziness", "photophobia", "phonophobia"}. According to the above formula, calculate the IOU value for each suspected disease. For example, for migraine, the IOU value is 0.6. Similarly, the system can calculate the IOU values for other suspected diseases. For example, suppose the set of typical symptoms of vertigo is {"vertigo", "dizziness", "tinnitus"}, then the intersection is {"dizziness"}, and the union is {"nausea", "dizziness", "photophobia", "vertigo", "tinnitus"}, so the IOU value is 0.2. By calculating the IOU value, the system can quantify the similarity between the set of symptoms self-reported by the patient and each suspected disease, thereby providing a numerical basis for subsequent disease reasoning. The higher the IOU value, the more consistent the symptoms self-reported by the patient are with the typical symptoms of the suspected disease, and vice versa.

[0109] Exemplarily, generating the set of health questions according to the second patient Q&A information specifically includes:

[0110] Obtain the second patient Q&A information after the patient makes selections from the second candidate symptom set;

[0111] Use a language model to process the second patient Q&A information to generate disease reasoning questions;

[0112] Obtain the third patient Q&A information after the patient makes selections from the disease reasoning questions;

[0113] Through the medical knowledge graph, perform a score threshold judgment on the third patient Q&A information, and generate corresponding disease questions for all diseases whose scores exceed the score threshold to form a set of disease questions;

[0114] Use the set of disease questions, the preset allergy history question set, and the preset medical history question set as the set of health questions.

[0115] Display the second candidate symptom set to the patient and ask the patient to confirm whether these symptoms actually exist. For example, the system may ask: "Do you have the following symptoms? Nausea, vomiting, dizziness, photophobia, phonophobia." The patient can choose "yes" or "no" according to the actual situation.

[0116] Then record the patient's selection results to form the second patient Q&A information. The key to this step is to further refine the patient's symptom description to ensure the accuracy of subsequent reasoning.

[0117] Subsequently, a pre-trained language model (such as BERT, GPT, etc.) can be used to process the second patient's Q&A information to generate a series of disease inference questions. These questions are designed to guide the patient to provide more detailed information about their symptoms so that the system can perform disease inference more accurately.

[0118] Show the generated disease inference questions to the patient and ask the patient to answer them one by one. The patient's answers will form the third patient's Q&A information, providing more basis for subsequent disease inference. Record the patient's answers, especially those closely related to disease diagnosis. For example, if the patient indicates that dizziness worsens when standing, the system will take this information as an important diagnostic clue.

[0119] Finally, the entity relationships in the medical knowledge graph and the patient's answers can be used to calculate the scores for each suspected disease. The score calculation can be based on multiple factors, including symptom matching degree, symptom severity, symptom duration, etc. For example, if the patient indicates that dizziness worsens when standing, the system may assign a higher score to hypotension because it highly matches the typical symptoms of hypotension.

[0120] In addition, a score threshold can be set, and only diseases with scores exceeding this threshold will be included in the final set of disease questions. This can ensure that the generated questions have high relevance and likelihood, avoiding unnecessary interference. For each disease with a score exceeding the threshold, generate a set of specific questions related to that disease. These questions are designed to further confirm the patient's condition or guide the patient to perform more self-examinations. For example, for hypotension, the following questions may be generated: "Do you have symptoms such as fatigue or palpitations? These symptoms are usually related to hypotension."

[0121] By relying on the second patient's Q&A information, a set of health questions can be effectively generated, ensuring the accuracy and comprehensiveness of disease inference. Specifically, first obtain the patient's selection results for the second set of candidate symptoms, then use the language model to generate a series of disease inference questions to further refine the patient's symptom description. Next, the system calculates the scores for each suspected disease based on the patient's answers, and generates corresponding disease questions for diseases with scores exceeding the threshold, forming a set of disease questions. Finally, the system integrates the set of disease questions, the preset set of allergy history questions, and the preset set of medical history questions into a complete set of health questions to ensure that doctors can comprehensively understand the patient's health status. This method not only improves the accuracy of diagnosis but also provides more intelligent and precise medical services for patients.

[0122] Exemplarily, generating patient health information according to the patient's answer information to the set of health questions specifically includes:

[0123] Obtain the patient's answer information to the set of health questions;

[0124] Process the said response information using natural language generation technology to generate a patient health information report including diagnostic suggestions and health suggestions.

[0125] Present a set of health questions to the patient and ask the patient to answer them one by one. The patient's answers will form the final response information, providing a basis for subsequent health information generation. Record the patient's answers, especially those closely related to disease diagnosis. For example, if the patient indicates that dizziness worsens when standing, the system will take this information as an important diagnostic clue.

[0126] Process the patient's response information using a pre-trained language model (such as BERT, GPT, etc.) to generate a patient health information report including diagnostic suggestions and health suggestions. These reports not only provide preliminary diagnostic opinions but also offer personalized health management suggestions for the patient.

[0127] Based on the patient's response information to the set of health questions, a patient health information report including diagnostic suggestions and health suggestions can be effectively generated. Specifically, first obtain the patient's response information to the set of health questions, and then use natural language generation (NLG) technology to process this information to generate a personalized health information report. This report not only provides preliminary diagnostic opinions but also offers personalized health management suggestions for the patient to help them better manage their health. In addition, the system can also recommend further examinations or consultations with professional doctors as needed to ensure the accuracy of the diagnosis.

[0128] Compared with the prior art, a question-and-answer-based patient information generation method provided by an embodiment of the present invention uses a generative large language model specially trained in the medical field to convert the patient's colloquial or non-standard descriptions into standard medical terms; during the medical term alignment process, ensure that the medical terms can accurately match the corresponding entities in the knowledge graph, and generate a set of symptoms (the first candidate symptom set) highly relevant to the patient's initial description based on the aligned medical knowledge graph entities to help the patient supplement important information that may be missed. Through multiple rounds of interaction, gradually narrow the scope of suspected diseases and improve the accuracy of disease reasoning.

[0129] By making multi-level use of the medical knowledge graph, the limitations of the existing pre-consultation system in aspects such as symptom recognition, disease reasoning, and user experience are solved. It not only improves the efficiency and accuracy of medical services but also provides a more intelligent, precise, and personalized experience for patients.

[0130] An embodiment of the present application provides a question-and-answer-based patient information generation device, including: a conversion module, an alignment module, a first set generation module, a second set generation module, a question generation module, and an information generation module.

[0131] A conversion module for converting patient condition information into medical terms; the medical terms belong to the corpus in the medical knowledge graph.

[0132] An alignment module for aligning the medical terms according to conversion conditions to obtain aligned medical knowledge graph entities.

[0133] A first set generation module for generating a first candidate symptom set according to the aligned medical knowledge graph entities.

[0134] A second set generation module for determining a first candidate disease set and a second candidate symptom set corresponding to the first candidate disease set according to first patient Q&A information; the second candidate symptom set is greater than or equal to the first candidate symptom set.

[0135] A question generation module for generating a set of health questions according to second patient Q&A information.

[0136] An information generation module for generating patient health information according to the patient's answer information to the set of health questions.

[0137] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.

[0138] Compared with the prior art, a patient information generation device based on Q&A provided by an embodiment of the present invention uses a generative large language model specially trained in the medical field to convert the patient's colloquial or non-standard description into standard medical terms; during the alignment process of medical terms, it is ensured that the medical terms can accurately match the corresponding entities in the knowledge graph, and a set of symptoms (the first candidate symptom set) highly relevant to the patient's initial description is generated according to the aligned medical knowledge graph entities, helping the patient supplement important information that may be missed. Through multiple rounds of interaction, the scope of suspected diseases is gradually narrowed, and the accuracy of disease inference is improved.

[0139] By making multi-level use of the medical knowledge graph, the limitations of the existing pre-consultation system in aspects such as symptom recognition, disease inference, and user experience are solved. It not only improves the efficiency and accuracy of medical services but also provides a more intelligent, precise, and personalized experience for patients.

[0140] Please refer to Figure 2 , an embodiment of the present application provides a consultation system applying the above method embodiment, including a user side, a front end, and an algorithm side. Among them, the algorithm side includes the above-mentioned patient information generation device based on Q&A, and a possible operation process of the entire system can be referred to as follows:

[0141] Step 1: Patient Information Input and Preliminary Verification. The patient inputs their basic information and a self-description of their condition through the system interface (user side). The system receives and records the data entered by the patient. The system uses predefined logical rules to determine whether the basic information provided by the patient is reasonable. If it is not reasonable, the system prompts the patient to fill it in again; if it is reasonable, proceed to the next step.

[0142] Step 2: Symptom Extraction and Entity Alignment. Using natural language processing techniques, extract potential symptom descriptions from the patient's self-description. Align the extracted symptom descriptions with the standard medical terms in the medical knowledge graph to obtain symptom descriptions in professional terms.

[0143] Step 3: Symptom Confirmation and High-Frequency Symptom Recommendation. Based on the medical knowledge graph, infer high-frequency symptoms related to the patient's symptoms. The system displays the recommended high-frequency symptoms for the patient to select and confirm. The patient selects the corresponding symptoms from the recommended high-frequency symptoms according to their own situation.

[0144] Step 4: Preliminary Disease Inference and Symptom Inquiry. The system conducts preliminary disease inference through the knowledge graph based on the symptoms selected by the patient. The system asks the patient whether they have the symptoms of the inferred disease.

[0145] Step 5: Dynamic Question Generation and Disease Inference. Based on the symptoms selected by the patient for the second time, use a language model (LLM) to generate symptom-related questions. The patient answers the generated questions, and the system collects more information. The system conducts disease inference again according to the knowledge graph. If there is a disease score exceeding the preset threshold, proceed to the next step; otherwise, skip the generation of disease-related questions. For diseases with scores exceeding the threshold, use the LLM to generate specific questions for that disease. To facilitate the patient's answering, the generated questions are all multiple-choice questions. The patient answers these questions to further confirm the disease possibility.

[0146] Step 6: Inquiry about Allergies and Medical History. The system uses a fixed question template to ask the patient about their allergies and medical history. The system integrates information such as the patient's self-description, symptom confirmation, question answers, disease inference, and allergies. Using natural language generation technology, generate an AI report containing diagnostic suggestions and health advice based on the conversation content. Display the generated AI report to the patient through the system interface and provide options for printing or downloading the electronic version.

[0147] Through the above implementation process, the interrogation system applying the question-based patient information generation method can provide personalized pre-interrogation services for patients and provide auxiliary diagnostic information for medical professionals, thereby improving the efficiency and accuracy of medical services.

[0148] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned patient information generation method based on question and answer.

[0149] The computer device may be a computing device such as a smart phone, a tablet computer, a desktop computer, and a cloud server. The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the figure is only an example of the computer device, and does not constitute a limitation on the computer device. It may include more or fewer components than shown, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0150] The so-called processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0151] The memory may be an internal storage unit of the computer device in some embodiments, such as the hard disk or memory of the computer device. The memory may also be an external storage device of the computer device in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory may include both the internal storage unit and the external storage device of the computer device. The memory is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory may also be used to temporarily store data that has been output or will be output.

[0152] An embodiment of the present application provides a computer program product, which when running on a computer device enables the computer device to implement the steps in the above-mentioned method embodiments when executed.

[0153] In several embodiments provided in this application, it can be understood that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, the program segment, or the part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved.

[0154] If the described functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0155] The above is the preferred implementation manner of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A method for generating patient information based on question and answer, characterized in that: include: Convert patient condition information into medical terms; The medical terms belong to the lexicon in the medical knowledge graph; Aligning the medical terms according to the conversion conditions to obtain aligned medical knowledge graph entities; Generating a first candidate symptom set according to the aligned medical knowledge graph entity; Determine, according to the first patient question and answer information, a first candidate disease set and a second candidate symptom set corresponding to the first candidate disease set; the second candidate symptom set is greater than or equal to the first candidate symptom set; Generate a set of health questions based on the second patient's question and answer information; Patient health information is generated based on the patient's answers to the set of health questions.

2. A method for generating patient information based on question and answer as claimed in claim 1, characterized in that: Before converting the patient's condition information into medical terms, the method further includes: Obtain patient self-reported information; The rationality of the patient's self-reported information is judged according to preset rules. If the patient's self-reported information meets the preset rules, the patient's self-reported information is used as the patient's condition information.

3. A method for generating patient information based on question and answer as claimed in claim 1, characterized in that: The converting of the patient's condition information into medical terms specifically includes: A generative large language model is applied to convert the patient's condition information into medical terms.

4. The method for generating patient information based on question and answer as claimed in claim 1, characterized in that: The medical terms are aligned according to the conversion conditions to obtain aligned medical knowledge graph entities, specifically including: Calculating the similarity between the medical term and all medical knowledge graph entities; Select the medical knowledge graph entity with the greatest similarity as the target medical knowledge graph entity; If the maximum similarity value is greater than a preset similarity threshold, the target medical knowledge graph entity is used as an aligned medical knowledge graph entity; if the maximum similarity value is less than or equal to the preset similarity threshold, a synonym judgment is performed based on a generative large language model, and after the judgment, the target medical knowledge graph entity is used as an aligned medical knowledge graph entity.

5. The method for generating patient information based on question and answer as claimed in claim 1, characterized in that: Generating a first candidate symptom set according to the aligned medical knowledge graph entity specifically includes: According to the aligned medical knowledge graph entity, determining in the medical knowledge graph a number of other medical knowledge graph entities whose shortest path lengths do not exceed a preset distance threshold; According to the occurrence frequency of the other medical knowledge graph entities, a number of the other medical knowledge graph entities are selected as the first candidate symptom set.

6. A method for generating patient information based on question and answer as claimed in claim 5, characterized in that: The selecting, according to the occurrence frequency of the other medical knowledge graph entities, a plurality of the other medical knowledge graph entities as the first candidate symptom set specifically includes: According to the frequency of occurrence of other medical knowledge graph entities, all other medical knowledge graph entities are arranged in descending order; if there are other medical knowledge graph entities with equal frequency of occurrence, the number of occurrences of other medical knowledge graph entities for various diseases in the medical knowledge graph is counted, and other medical knowledge graph entities with larger number of occurrences are arranged earlier in descending order; According to the order of the arrangement results, a preset number of other medical knowledge graph entities are selected starting from the first order.

7. The method for generating patient information based on question and answer as claimed in claim 1, characterized in that: The determining, according to the first patient question and answer information, a first candidate disease set and a second candidate symptom set corresponding to the first candidate disease set specifically includes: Acquire first patient question and answer information after the patient selects the first candidate symptom set; Matching the first patient question and answer information through the medical knowledge graph to determine a patient self-reported symptom set corresponding to the first patient question and answer information and a symptom set corresponding to a plurality of suspected diseases; According to the patient's self-reported symptom set and the symptom sets corresponding to several suspected diseases, the IOU value of each suspected disease is calculated; the IOU value is the ratio of the number of symptom intersections between the two sets to the number of symptom unions; Determine a first candidate disease set and a second candidate symptom set corresponding to the first candidate disease set according to the IOU value.

8. The method for generating patient information based on question and answer as claimed in claim 1, characterized in that: The generating of a set of health questions according to the second patient's question and answer information specifically includes: Acquire second patient question and answer information after the patient selects the second candidate symptom set; Processing the second patient question and answer information using a language model to generate a disease reasoning question; Obtaining third patient question and answer information after the patient selects the disease reasoning question; Performing a score threshold judgment on the third patient question and answer information through the medical knowledge graph, generating corresponding disease questions for all diseases with scores exceeding the score threshold, and forming a disease question set; The disease question set, the preset allergy history question set and the preset medical history question set are taken as the health question set.

9. The method for generating patient information based on question and answer as claimed in claim 1, characterized in that: The generating of the patient health information according to the patient's answer information to the health question set specifically includes: Obtaining the patient's answer information to the set of health questions; The answer information is processed using natural language generation technology to generate a patient health information report including diagnosis suggestions and health suggestions.

10. A patient information generation device based on question and answer, characterized in that: include: A conversion module, used to convert patient condition information into medical terms; The medical terms belong to the lexicon in the medical knowledge graph; An alignment module, used for aligning the medical terms according to the conversion conditions to obtain aligned medical knowledge graph entities; A first set generation module, used to generate a first candidate symptom set according to the aligned medical knowledge graph entity; A second set generating module is used to determine, according to the first patient question and answer information, a first candidate disease set and a second candidate symptom set corresponding to the first candidate disease set; the second candidate symptom set is greater than or equal to the first candidate symptom set; A question generation module, used for generating a set of health questions according to the second patient's question and answer information; The information generation module is used to generate patient health information based on the patient's answer information to the set of health questions.

Citation Information

Cited By

  • Knowledge graph-driven lightweight medical large model system

    CN120376117A

  • Intelligent question and answer method and device suitable for lung health and electronic equipment

    CN120849551A