Medical guide consensus retrieval question and answer method based on knowledge graph and LLM
By constructing a medical guide consensus search question-and-answer method based on knowledge graphs and large language models, the problem of complex consensus information in medical guide is solved, and the effect of quickly obtaining relevant information is achieved, the clarity and accuracy of answers is improved, and reading time is saved.
Patent Information
- Application Number
- CN202411553554.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-02
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The information in the consensus on medical guidelines is complicated and it is difficult to quickly obtain relevant information, which leads to the waste of time for information search and understanding when diagnosed and treated.
The medical guide consensus search question-and-answer method based on knowledge graph and large language model is adopted. By constructing a medical guide consensus knowledge graph and adaptive Chinese input recommendation algorithm module, the intelligent question-and-answer system is constructed in combination with the GPT big model, and the user's questions are quickly retrieved and answer generation.
It improves the clarity and focus of answers, saves reading time, provides answers that are more in line with user needs, and ensures the accuracy and timeliness of answers through the context enhancement of the knowledge graph.
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] The present invention relates to a method for retrieving and answering questions about medical guideline consensus, and particularly to a method for retrieving and answering questions about medical guideline consensus based on a knowledge graph and an LLM (Large Language Model), belonging to the technical field of natural language processing. Background Art
[0002] Medical guideline consensus provides a scientific basis for clinicians, aiming to help doctors understand the best methods for diagnosing, treating, and preventing diseases. By following the guidelines, doctors can make well-founded diagnoses and treatments when faced with various diseases, thereby reducing the probability of misdiagnosis and mistreatment and ensuring that patients receive homogeneous, fair, and reasonable medical care. However, the literature content is complex and cumbersome, inconvenient to read, and difficult to find key information. Therefore, some computer technologies have been continuously applied to the medical field.
[0003] The technology of knowledge graph is now relatively mature. Facing the strong correlations among complex symptoms, diseases, and drugs in the medical field, the knowledge graph can be well applied. Constructing a knowledge graph of medical guideline consensus to structure and associate scattered medical knowledge (such as diseases, symptoms, treatment methods, etc.) helps doctors quickly obtain relevant information and improve the efficiency of diagnosis and treatment. By analyzing and mining medical data, it is also possible to discover new research directions and potential treatment methods, promoting the progress of medical research.
[0004] With the continuous development of computer technology, the combination of medicine and deep learning has become a development trend. GPT (Generative Pre-trained Transformer) series models can quickly integrate and summarize a large amount of information on medical guideline consensus, helping medical professionals and patients obtain the required knowledge more quickly and reducing the time for searching and understanding information. GPT shows a high degree of intelligence and flexibility in the interaction process with users, and further improves the integrity of the user experience.
[0005] The present invention is a method for retrieving and answering questions about medical guideline consensus based on a knowledge graph and an LLM, which combines the knowledge graph, large language model with medicine to help users obtain authoritative and comprehensive answers. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for retrieving and answering questions about medical guideline consensus based on a knowledge graph and an LLM.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is: A method for retrieving and answering questions about medical guideline consensus based on a knowledge graph and an LLM, characterized by comprising the following steps.
[0008] Step 1: Construct a medical guideline consensus knowledge graph, which consists of the following steps: Step 1-1: Summarize medical guideline consensus data; Step 1-2: Perform data cleaning. Use regular expressions, keyword filtering, sentiment analysis, text cleaning tools, etc. to remove inappropriate sentences, names, etc., ensure the quality of the data, and establish a medical guideline consensus database; Step 1-3: Conduct medical domain entity recognition, such as important entities, such as diseases, symptoms, drugs, treatment methods, etc.; Step 1-4: Conduct relationship extraction to construct the relationships between entities, such as the treatment relationship between drugs and diseases, the association between symptoms and diseases, etc.; Step 1-5: Regularly update the knowledge graph to maintain the timeliness and accuracy of information; Step 1-6: Graph optimization. Extract triples (entity 1, relationship, entity 2) from the knowledge graph as training data, and use the Node2Vec embedding model for training to optimize the embedding space so that under specific relationships, the embedding vectors of related entities are closer.
[0009] Step 2: Construct an adaptive Chinese input recommendation algorithm module, which consists of the following steps: Step 2-1: Construct a model training database; Step 2-2: Use the Transformer pre-training model to generate vectors with semantic information from the pinyin and Chinese characters of the sentences in the training set and store them; Step 2-3: Retrieve through the user's input in the stored vectors with semantic information; Step 2-4: Sort the retrieved content and prompt relevant questions for the user input in order.
[0010] Step 3: Construct an intelligent question-answering system, which consists of the following steps: Step 3-1: Intent recognition of user questions. Intent recognition provides model support for subsequent entity linking to the knowledge graph; this model consists of an input layer, a convolutional layer 1, a pooling layer 1, a dropout layer 1, a BiLSTM, a dropout layer 2, a Self-Attention, and an output layer; Step 3-2: Use the medical guideline consensus database in Step 1-2, and use the Embedding model to convert all the text data in the database into high-dimensional vector forms to obtain a medical guideline consensus vector library; Step 3-3: Convert the user's question into a high-dimensional vector form, calculate the cosine similarity between the question vector and the knowledge vectors in the knowledge base, sort the obtained knowledge vectors in descending order according to the cosine similarity, return the texts corresponding to the top k knowledge vectors, and screen out the k answers closest to the user's question; Step 3-4: The relevant answers retrieved from the knowledge graph are used as the context relationship output by the model result; Step 3-5: Input the context relationship generated by the knowledge graph, the user's question, and the answer into the prompt template to construct the prompt input information, input it into the GPT large model, and through the understanding, analysis, and polishing of the large language model, generate the final answer; Step 3-6: If no relevant answers are retrieved, prohibit the model from generating relevant answers to avoid misleading users.
[0011] Furthermore, Step 3-1 also includes the intention recognition model training step: Traditional machine learning methods cannot understand the deep semantic information of the user's words. Therefore, use the intention recognition model to classify and train the question dataset, preprocess the text segmentation into word vectors, use the sentence vector as the input sequence, and predict its intention label. Each input sequence corresponds to an intention type.
[0012] Furthermore, Step 3-5 also includes converting the results retrieved from the knowledge graph into a context relationship, inputting them together with the question and the answer into the large model for comprehensive analysis, which can add specific domain knowledge and fill in the knowledge not learned by the GPT pre-trained model.
[0013] Furthermore, Step 3-6 also includes adjusting the prompt template. The optimized prompt template is: prompt_template = "
Instruction
Known Information
Question
[0014] Adopting the above technical solutions, the beneficial effects of the present invention are: 1. The present invention combines the large model language for retrieval, making the answer more clear and the key points more prominent, saving reading time; 2. The present invention uses the knowledge graph for context enhancement, not only realizing the tracing of facts but also aggregating and connecting facts, making the obtained answers more in line with the user's needs; 3. The present invention has an adaptive Chinese input recommendation algorithm module, which recommends questions according to the user's input, making the user's questions more standardized and the answers more in line with the user's needs. Brief Description of the Drawings
[0015] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments; Figure 1 is the structural diagram of the medical guideline consensus method of the present invention; Figure 2 is the structural diagram of the intention recognition model. Embodiment
[0016] In order to more clearly demonstrate the purpose, technical solution and advantages of the present invention, the technical solution of the present invention will be described in detail and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the present invention, rather than all embodiments. Based on these embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative labor shall fall within the protection scope of the present invention.
[0017] A medical guideline consensus retrieval and question-answering method based on knowledge graph and LLM, by using natural language processing technology and combining knowledge graph technology, comprehensively analyzes medical guideline consensus to obtain multiple answers corresponding to user questions.
[0018] As Figure 1 shown, the overall method structure is as follows.
[0019] Step 1: Construct a medical guideline consensus knowledge graph, which consists of the following steps: Step 1-1: Summarize medical guideline consensus data; Step 1-2: Perform data cleaning, using regular expressions, keyword filtering, text cleaning tools, etc., to remove inappropriate sentences, names, etc., ensure the quality of the data, and establish a medical guideline consensus database; Regular expression: Use regular expressions to match specific patterns of names and inappropriate remarks; Keyword filtering: A keyword list of inappropriate remarks can be maintained and the text can be traversed to delete these words; Text cleaning tool: Use NLTK, spaCy, or other text processing libraries to perform comprehensive text cleaning on text data; Step 1-3: Perform medical domain entity recognition, such as important entities, such as diseases, symptoms, drugs, treatment methods, etc.; Step 1-4: Perform relationship extraction to construct relationships between entities, such as the treatment relationship between drugs and diseases, the association between symptoms and diseases, etc.; Step 1-5: Regularly update the knowledge graph to maintain the timeliness and accuracy of information; Step 1-6: Atlas Optimization. Extract triples (entity 1, relationship, entity 2) from the knowledge graph as training data, and use the Node2Vec embedding model for training to optimize the embedding space so that the embedding vectors of related entities are closer under a specific relationship.
[0020] Step 2: Build an adaptive Chinese input recommendation algorithm module, which consists of the following steps: Step 2-1: Build a model training database; Step 2-2: Use the Transformer pre-trained model to generate vectors with semantic information from the pinyin and Chinese characters of the sentences in the training set and store them; Step 2-3: Retrieve from the stored vectors with semantic information through the user's input; Step 2-4: Sort the retrieved content and prompt the user input with relevant questions in order.
[0021] Step 3: Build an intelligent question-answering system, which consists of the following steps: Step 3-1: Input the user's question into the intent recognition model to provide model support for subsequent entity linking to the knowledge graph; As Figure 2 shown, it is the structure diagram of the intent recognition model. This model consists of an input layer, a convolutional layer 1, a pooling layer 1, a dropout layer 1, a BiLSTM, a dropout layer 2, a Self-Attention, and an output layer; Use this model to train the intent recognition model for the user's question. Traditional machine learning methods cannot understand the deep semantic information of the user's words. Therefore, use the BiLSTM network to classify and train the question dataset. Preprocess the text segmentation into word vectors, use the sentence vectors as the input sequence, and predict its intent label. Each input sequence corresponds to an intent type; Convolutional layer 1: Used for preliminary feature extraction to extract low-level features of the text; Pooling layer 1: Perform downsampling operations to reduce the size of the feature map, retain the main features, and improve the anti-overfitting ability of the model; Dropout layer 1: Prevent too many input and output neurons from causing overfitting problems during the training phase; BiLSTM: Bidirectional long short-term memory network, with the feature of model parameter sharing. By processing in two directions of the input sequence, that is, forward and backward, the model can capture the context information before and after the current position at the same time, which is the key module for intent recognition of users by combining context information; Dropout layer 2: Because the long short-term memory network is relatively complex, the dropout layer 2 is used to prevent overfitting; Self-Attention: The core function of the Self-Attention model is that it can capture the dependencies between elements in the sequence, whether local or global, thus helping the model better understand the input sequence and improve the model's performance ability. It is a key module for accurately classifying the user's intention; Step 3-2: Utilize the medical guideline consensus database in Step 1-2 and use the Embedding model to convert all the text data in the database into high-dimensional vector form to obtain a medical guideline consensus vector library; Step 3-3: Convert the user's question into high-dimensional vector form, calculate the cosine similarity between the question vector and the knowledge vectors in the knowledge base, sort the obtained knowledge vectors in descending order according to the cosine similarity, return the texts corresponding to the top k knowledge vectors, and screen out the k answers closest to the user's question; Step 3-4: Retrieve the relevant answers in the knowledge graph as the context relationship output by the model; Step 3-5: Input the context relationship generated by the knowledge graph, the user's question, and the relevant answers into the prompt template to construct the prompt input information, and input it into the GPT large model. Input the context relationship, together with the question and answers, into the large model for comprehensive analysis, which can add domain-specific knowledge and fill in the knowledge not learned by the GPT pre-trained model. Through the understanding, analysis, and polishing of the large language model, generate the final answer; Step 3-6: If no relevant answers are retrieved, prohibit the model from generating relevant answers to avoid misleading users; The prompt template has been adjusted, and the optimized prompt template is: prompt_template = "
Instruction
Known Information
Question
Claims
1. A medical guideline consensus retrieval question-answering method based on knowledge graph and LLM, characterized by: The following steps are involved: Step 1: Construct a medical guidelines consensus knowledge graph, which consists of the following steps: Step 1-1: Summarize medical guideline consensus data; Step 1-2: Perform data cleaning, using regular expressions, keyword filtering, sentiment analysis, text cleaning tools, etc. to remove inappropriate sentences and names, ensure data quality, and establish a medical guideline consensus database; Step 1-3: Perform entity recognition in the medical field, such as important entities such as diseases, symptoms, drugs, treatment methods, etc. Step 1-4: Perform relationship extraction to build relationships between entities, such as the therapeutic relationship between drugs and diseases, the association between symptoms and diseases, etc. Step 1-5: Update the knowledge graph regularly to keep the information timely and accurate; Step 1-6: Graph optimization: extract triples (entity 1, relationship, entity 2) from the knowledge graph as training data, use the Node2Vec embedded model for training, and optimize the embedding space so that the embedding vectors of related entities are closer under a specific relationship. Step 2: Build an adaptive Chinese input recommendation algorithm, which consists of the following steps: Step 2-1: Build a model training database; Step 2-2: Use the Transformer pre-training model to generate vectors with semantic information from the pinyin and Chinese characters of the sentences in the training set and store them; Step 2-3: Retrieve the stored vector with semantic information through user input; Step 2-4: Sort the retrieved content and prompt the user with relevant questions in order; Step 3: Build an intelligent question-answering system, which consists of the following steps: Step 3-1: Intent recognition of user questions. Intent recognition provides model support for subsequent entity linking to the knowledge graph. The model consists of input layer, convolution layer 1, pooling layer 1, random dropout layer 1, BiLSTM, random dropout layer 2, Self-Attention, and output layer.
2. Step 3-2: Using the medical guideline consensus database in step 1-2, use the Embedding model to convert all text data in the database into high-dimensional vector form to obtain the medical guideline consensus vector library; Step 3-3: Convert the user's question into a high-dimensional vector form, calculate the cosine similarity between the question vector and the knowledge vector in the knowledge base, and sort the obtained knowledge vectors from high to low according to the cosine similarity, return the text corresponding to the first k knowledge vectors, and select the k answers closest to the user's question; Step 3-4: Relevant answers retrieved from the knowledge graph are used as contextual relations for the model output; Step 3-5: Input the contextual relationship, user questions and answers generated by the knowledge graph into the prompt template, construct the prompt input information, input it into the GPT large model, and generate the final answer through the understanding, analysis and polishing of the large language model; Step 3-6: If no relevant answer is retrieved, prohibit the model from generating relevant answers to avoid misleading users.
3. The medical guideline consensus retrieval question-answering method based on knowledge graph and LLM according to claim 1, characterized in that: Step 3-1 also includes the step of training the intent recognition model: traditional machine learning methods cannot understand the deep semantic information of the user's speech, so the intent recognition model is used to classify and train the question dataset, preprocess the text into word vectors, and use the sentence vectors as input sequences to predict its intent labels. Each input sequence corresponds to an intent type.
4. The medical guideline consensus retrieval question-answering method based on knowledge graph and LLM according to claim 1, characterized in that: The steps 3-5 also include converting the results retrieved from the knowledge graph into contextual relationships, and inputting them into the large model together with the questions and answers for comprehensive analysis, which can add specific domain knowledge and fill in the knowledge that has not been learned by the GPT pre-training model.
5. The medical guideline consensus retrieval question-answering method based on knowledge graph and LLM according to claim 1, characterized in that: The steps 3-6 also include adjusting the prompt template. The prompt optimization template is: prompt_template="【Instruction】Find and integrate the answer based on the known information. It is not allowed to add fabricated elements to the answer. Give a clear answer and corresponding explanation. If the answer cannot be found, please answer "Unable to obtain relevant information about this question." \n\n【Known information】{context} \n\n【Question】{question}".
Citation Information
Patent Citations
Large language model medical question-answering system based on medical record knowledge graph
CN117056493A
Retrieval enhancement generation system based on medical knowledge graph
CN118394952A
Auxiliary medical management method and system based on natural questions and answers and knowledge graph
CN118568228A
Traditional Chinese medicine question-answering system construction method based on large language model and knowledge graph
CN118838996A
Systems for controllable summarization of content
US12008332B1