Intelligent question answering system and method based on knowledge graph and large language model

Through an intelligent question-and-answer system based on knowledge graphs and large language models, the problem of insufficient knowledge resources for Kawasaki disease is solved, accurate and rich answers to Kawasaki disease-related problems are achieved, and the efficiency of diagnosis and treatment is improved.

CN120011506APending Publication Date: 2025-05-16CHONGQING MEDICAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510085205.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, Kawasaki disease knowledge resources are insufficient, which makes it difficult for doctors and patients to obtain timely and detailed disease knowledge, affecting the timeliness and effectiveness of diagnosis and treatment.

Method used

An intelligent question-and-answer system based on knowledge graph and large language model is adopted. Relevant knowledge is extracted from the text through the knowledge graph construction module. The knowledge retrieval filtering module recognizes user intentions and calculates the similarity between knowledge and questions. The question-and-answer module uses relevant knowledge as a prompt and performs questions and answers based on user questions.

Benefits of technology

It achieves accurate and rich answers to specific areas such as Kawasaki disease, improves the interpretability and accuracy of answers to large language models, provides a convenient channel for obtaining knowledge, and improves the efficiency of diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011506A_ABST
    Figure CN120011506A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical instruments, and particularly discloses an intelligent question answering system and method based on a knowledge graph and a large language model.The system comprises a knowledge graph construction module, a knowledge retrieval filtering module and a question answering module; the knowledge graph construction module extracts related knowledge from a text, and the knowledge retrieval filtering module calculates cosine similarity between user question features and each piece of knowledge by identifying user intentions, and finds knowledge highly related to user questions in a knowledge graph; the question-answering module takes related knowledge as prompts and answers questions and answers according to user questions. By adopting the technical scheme, based on the knowledge graph and the large language model, questions in specific fields such as medical science and the like can be answered more accurately and richly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical devices and relates to an intelligent question-answering system and method based on a knowledge graph and a large language model. Background Art

[0002] For some rare diseases (such as Kawasaki disease), ordinary people lack understanding, which may lead to patients not seeking medical treatment in time, resulting in serious consequences. If the doctor happens to lack relevant clinical experience and medical knowledge, it is easy to delay diagnosis, increase the incidence of complications, and affect the patient's prognosis (the incidence of complications of coronary artery disease in children with Kawasaki disease who are not treated in time is as high as 15% to 25%, and in developed countries, coronary artery disease has become the main cause of acquired heart disease in children. After children with Kawasaki disease see a doctor, timely intravenous immunoglobulin (IVIG) treatment can effectively reduce the incidence of coronary artery disease to about 4%. Therefore, promoting timely medical treatment for children with Kawasaki disease is an important part of diagnosis and treatment).

[0003] In addition, due to the large amount and complexity of medical professional knowledge, it is cumbersome for medical students to acquire complete and detailed knowledge of diseases. At present, disease knowledge is mainly presented in the form of professional books, literature, and videos, and the knowledge is relatively scattered and lacks organization. It is difficult for medical students to obtain overall and detailed knowledge related to the disease from one channel. In systematic learning or specific analysis of a related issue, it is difficult for medical students to acquire relevant knowledge, especially for a specific case. The learning willingness is not strong and the learning efficiency is low, which will greatly affect the diagnosis and treatment effect. Summary of the invention

[0004] The purpose of the present invention is to provide an intelligent question-answering system and method based on knowledge graph and large language model to solve the problem of insufficient construction of knowledge resources on Kawasaki disease.

[0005] In order to achieve the above-mentioned purpose, the basic scheme of the present invention is: an intelligent question-answering system based on knowledge graph and large language model, including a knowledge graph construction module, a knowledge retrieval and filtering module and a question-answering module;

[0006] The knowledge graph construction module extracts relevant knowledge from the text;

[0007] The knowledge retrieval and filtering module identifies the user's intention, calculates the cosine similarity between the user's question feature and each piece of knowledge, and finds knowledge with high relevance to the user's question in the knowledge graph;

[0008] The question-answering module uses relevant knowledge as prompts and answers questions based on user questions.

[0009] The working principle and beneficial effects of this basic solution are as follows: This technical solution constructs a knowledge graph of related diseases, and by inputting the queried relevant medical knowledge as prompts and user questions into the large language model, accurate and rich answers to user questions can be obtained.

[0010] Furthermore, the knowledge graph construction module includes a knowledge extraction module, a knowledge proofreading module, and a knowledge fusion module;

[0011] The knowledge extraction module collects data related to the field to be answered, and the knowledge extraction module includes an entity recognition module, a relationship extraction module and an attribute extraction module;

[0012] The knowledge proofreading module includes an entity proofreading module and a relationship cleaning module, which respectively perform entity proofreading and relationship cleaning;

[0013] The knowledge fusion module includes an ontology matching module and an alignment module. The knowledge fusion module integrates the entities, relations and attributes processed by different data sources.

[0014] Improve the quality of the knowledge graph by organizing data, proofreading entities, and cleaning relationships.

[0015] Furthermore, the knowledge retrieval and filtering module identifies the user's intention, calculates the cosine similarity between the user's question feature and each piece of knowledge, and finds knowledge with high relevance to the user's question in the knowledge graph;

[0016] The specific calculation process is:

[0017] Represent the user's question features and each piece of knowledge as a vector;

[0018] Calculate the cosine similarity between two knowledge vectors by the angle between them;

[0019] Use the question and the queried knowledge triples to calculate the similarity, and add the relationship type information that the question may involve. Since the knowledge triples contain relationship fields, adding relationship type information improves the similarity between the question and the related relationship triples, and improves the relevance of the query results.

[0020] Since knowledge triples already contain relationship fields, the addition of relationship type information improves the similarity between questions and related relationship triples, and the relevance of query results is improved.

[0021] Furthermore, the question-answering module uses the question-related knowledge triples obtained from the knowledge retrieval and filtering module as input to the large language model, and by setting prompts, enables the large model to answer user questions based on relevant knowledge, thereby reducing the occurrence of "hallucination phenomenon" of the large language model and improving the interpretability of the large language model's answers.

[0022] Simple structure and easy to use.

[0023] The present invention also provides an intelligent question-answering method for the system of the present invention, comprising the following steps:

[0024] Collect data related to the problem, build a data model, that is, an entity and relationship framework, and build a knowledge graph based on the entity and relationship framework;

[0025] Get input question;

[0026] Based on the input question, identify the user's intention, calculate the cosine similarity between the user's question features and each piece of knowledge, find the knowledge that is highly relevant to the user's question in the knowledge graph, filter the knowledge, and obtain knowledge triples;

[0027] Knowledge triples and user questions are input into a large language model to generate answers.

[0028] This method realizes intelligent question and answer based on knowledge graph, providing better Kawasaki disease knowledge services to patients and their families, medical students, and ordinary people.

[0029] Furthermore, the method of constructing a knowledge graph based on the entity and relationship framework is:

[0030] Extract medical entities and relationships from the data model. Entities include symptoms or signs, groups, medical devices, instruments and materials, abnormal test results, tests, drugs, treatments, prognosis, anatomical parts and substances, semantics, institutions, etiology classes, and diseases.

[0031] Proofread the extracted entities and clean the extracted relationships;

[0032] Perform entity alignment;

[0033] Import the cleaned entities, relationships, and attributes into the graph database;

[0034] Use the predefined entity and relationship framework to annotate data, train the information extraction model, evaluate the actual information extraction effect, optimize the entity and relationship framework, repeat the above process multiple times, and finally determine the entity and relationship framework;

[0035] According to the defined entity and relationship types, randomly select sentences containing the specified type of entities and relationships from the collected relevant data for annotation. The annotated data basically covers all entity and relationship types, and the UIE unified information extraction model is trained using this annotated data.

[0036] Select the uie-base and uie-m-base models as training models, set the learning rate to 1e-5, and change the number of training cycles to train the uie-base and uie-m-base models respectively;

[0037] The number of training cycles is set to 300 to train the uie-base model; the number of training cycles is set to 150 to train the uie-base model. These two models are selected as the entity extraction and relationship extraction models respectively.

[0038] Medical knowledge is extracted from medical texts through information extraction models, and knowledge graphs enable computers to truly understand knowledge and achieve cognitive intelligence.

[0039] Furthermore, the database is queried to see if there is a corresponding entity. If so, 10 different expressions in the same language are selected from the concepts corresponding to the queried entity to calculate the entity similarity. If they are considered similar, it means that the entity and the concept represent the same concept. A bidirectional alias relationship is established in the knowledge graph for the entities representing the same concept, and entity alignment is completed.

[0040] In the database, different expressions of the same thing and its corresponding concept use triples: thing expression, representative, concept;

[0041] The weighted average of the Jaccard coefficient and the Levenshetein ratio is used to calculate the similarity between entities. The Jaccard coefficient weight is set to 0.3, the Levenshetein ratio weight is set to 0.7, and the threshold is set to 0.8 to achieve entity alignment.

[0042] A weighted average of the Jaccard coefficient and the Levenshetein ratio is used to achieve better entity alignment.

[0043] Furthermore, when retrieving and filtering knowledge, the relationship type in the knowledge graph is used as the classification basis to construct a query statement template, and the user's question is identified for user intent, medical named entity recognition, and key entity extraction. If it is recognized that the user's question may contain a certain relationship type, a query template is constructed based on this relationship type, and the relationship type is fixed. If the question contains an entity of the entity type limited by this relationship type, it will be parsed one by one and its entity name and entity type will be passed into the query template being constructed. No restrictions are set for the other entity, and then all the constructed query templates are merged, and finally this query statement is used to query the knowledge graph;

[0044] Use the big language model to build a big model agent, input the relationship type and its description into the big language model as prompts, and identify the relationships involved in the user's question;

[0045] Use all constructed query templates to perform joint query and return the queried knowledge triples and entity attribute triples;

[0046] In the case where no knowledge is found, a large model agent, namely a large language model application, is constructed. The key entities identified from the user questions are used to query using query templates that do not specify the relationship type. Multiple triples containing any key entities are checked and returned to the next process for knowledge filtering.

[0047] Knowledge query is performed by merging query templates built according to the relationship types involved in user questions, utilizing the natural language understanding capabilities of large language models. It has low requirements on the form of questions, can identify the intent of complex questions, and has strong applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a structural block diagram of the intelligent question-answering method of the present invention;

[0049] Figure 2 It is a structural schematic diagram of a knowledge graph construction module of an intelligent question-answering system based on a knowledge graph and a large language model of the present invention;

[0050] Figure 3 It is a structural schematic diagram of a BIOS database of the intelligent question-answering method of the present invention;

[0051] Figure 4 It is a flow chart of intelligent question and answering of the intelligent question and answering method of the present invention. DETAILED DESCRIPTION

[0052] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0053] In the description of the present invention, it is necessary to understand that the terms "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0054] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal connection between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.

[0055] The present invention discloses an intelligent question-answering system based on a knowledge graph and a large language model, which enables the large model to be more accurate and rich in answering questions in specific fields such as Kawasaki disease. The intelligent question-answering system based on a knowledge graph and a large language model includes a knowledge graph construction module, a knowledge retrieval and filtering module, and a question-answering module.

[0056] The knowledge graph construction module extracts relevant knowledge from the text, and the knowledge retrieval and filtering module identifies the user's intention, calculates the cosine similarity between the user's question features and each piece of knowledge, and finds knowledge in the knowledge graph that is highly relevant to the user's question.

[0057] The question-and-answer module uses relevant knowledge as prompts and answers questions based on user questions.

[0058] When in use, the user inputs a question into the intelligent question-answering system, and the knowledge retrieval and filtering module will identify the information therein, including keywords, medical entities, and possible relationships involved. This information will be used to query and filter the medical knowledge related to the question. The filtered medical knowledge will be input into the large language model together with the user's question, and the large language model will give the answer.

[0059] By inputting the retrieved relevant medical knowledge as prompts and user questions into the large language model, accurate and rich answers to user questions can be obtained. The construction of the Kawasaki disease intelligent question-answering system has improved the accuracy and richness of the large language model's answers to Kawasaki disease-related questions. It can serve as a smart medical assistant for patients and their families, a Kawasaki disease knowledge learning tool for medical students, and a popular science assistant for people who care about health.

[0060] In a preferred embodiment of the present invention, Figure 2 As shown in the figure, the knowledge graph construction module includes a knowledge extraction module, a knowledge proofreading module, and a knowledge fusion module. The knowledge extraction module collects data related to the field to be answered (such as guidelines, literature, etc.). The knowledge extraction module includes an entity recognition module, a relationship extraction module, and an attribute extraction module to extract medical entities and relationships from the collected data.

[0061] The knowledge proofreading module includes an entity proofreading module and a relationship cleaning module, which perform entity proofreading and relationship cleaning respectively. The knowledge fusion module includes an ontology matching module and an alignment module. The knowledge fusion module integrates the entities, relationships and attributes processed by different data sources to obtain a knowledge graph.

[0062] Knowledge graph is a structured data storage method that usually uses triples of the form (subject, predicate, object) to store knowledge. Different from the traditional keyword search method, knowledge graph enables computers to truly understand knowledge and achieve cognitive intelligence.

[0063] In a preferred embodiment of the present invention, the knowledge retrieval and filtering module identifies the user's intention, calculates the cosine similarity between the user's question feature and each piece of knowledge, and finds knowledge with high relevance to the user's question in the knowledge graph;

[0064] The specific calculation process is:

[0065] Represent the user's question features and each piece of knowledge as a vector;

[0066] Calculate the cosine similarity between two knowledge vectors by the angle between them;

[0067] Use the question and the queried knowledge triples to calculate the similarity, and add the relationship type information that the question may involve. Since the knowledge triples contain relationship fields, adding relationship type information improves the similarity between the question and the related relationship triples, and improves the relevance of the query results.

[0068] A knowledge retrieval filtering module is set up in the knowledge query. Compared with only using the question and the queried knowledge triples for similarity calculation, the relationship type information that the question may involve is also added. Since the knowledge triples already contain relationship fields, the addition of relationship type information improves the similarity between the question and the related relationship triples, and the relevance of the query results is improved.

[0069] In a preferred embodiment of the present invention, the question-answering module uses the question-related knowledge triples obtained from the knowledge retrieval and filtering module as the input of the large language model, and by setting prompts, the large model answers the user's question based on the relevant knowledge, thereby reducing the occurrence of the "hallucination phenomenon" of the large language model and improving the interpretability of the large language model's answers.

[0070] The present invention also provides an intelligent question-answering method for the system of the present invention, such as Figure 1 and Figure 4 As shown, the following steps are included:

[0071] Collect data related to the problem, build a data model, that is, an entity and relationship framework, and build a knowledge graph based on the entity and relationship framework;

[0072] Get input question;

[0073] Based on the input question, identify the user's intention, calculate the cosine similarity between the user's question features and each piece of knowledge, find the knowledge that is highly relevant to the user's question in the knowledge graph, filter the knowledge, and obtain knowledge triples;

[0074] Knowledge triples and user questions are input into a large language model to generate answers.

[0075] In a preferred embodiment of the present invention, the method for constructing a knowledge graph based on an entity and relationship framework is:

[0076] Extract medical entities and relationships from the data model. Entities include symptoms or signs, groups, medical devices, instruments and materials, abnormal test results, tests, drugs, treatments, prognosis, anatomical parts and substances, semantics, institutions, etiology classes, and diseases.

[0077] Proofread the extracted entities and clean the extracted relationships;

[0078] Perform entity alignment;

[0079] Import the cleaned entities, relationships, and attributes into a graph database (such as Neo4j graph database, BIOS database, etc.);

[0080] After repeatedly using the predefined entity and relationship frameworks to annotate data, train information extraction models, and compare information extraction results, we finally determined the entity description framework (as shown in Table 1) and the entity relationship description framework (as shown in Table 2);

[0081] According to the defined entity and relationship types, randomly select sentences containing the specified type of entities and relationships from the collected relevant data for annotation. The annotated data basically covers all entity and relationship types. Use the doccano tool to annotate a part of the data, and use this annotated data to train the UIE unified information extraction model.

[0082] Select the uie-base and uie-m-base models as training models, set the learning rate to 1e-5, and change the number of training cycles to train the uie-base and uie-m-base models respectively;

[0083] The model was evaluated using training data, and its performance on a real guide data was also evaluated using three indices: micro-average score, V-measures, and ARI (Adjusted Rand Index). Since the micro-average precision, micro-average recall, and micro-average f1 are equal each time the model is evaluated, the micro-average score only needs to be represented by one value. The training effect and actual performance of the UIE model are shown in Table 3. The number of training cycles is set to 300, and the uie-base model is trained, and the entity extraction effect is the best; the number of training cycles is set to 150, and the uie-base model is trained, and the relationship extraction effect is the best. These two models are selected as the entity extraction and relationship extraction models respectively.

[0084] Table 1 Entity description framework

[0085]

[0086]

[0087] Table 2 Entity relationship description framework

[0088]

[0089]

[0090] Table 3 UIE model training results

[0091]

[0092] In a preferred embodiment of the present invention, a BIOS database (such as Figure 3 As shown in the figure, taking "ischemic changes" as an example), entity alignment is performed. Whether there is a corresponding entity in the database, if so, 10 different expressions in the same language are selected from the concepts corresponding to the queried entity to calculate the entity similarity. If it is determined to be similar, it means that the entity and the concept represent the same concept. A bidirectional alias relationship is established for the entities representing the same concept in the knowledge graph, and the entity alignment is completed;

[0093] In the database, different representations of the same thing and its corresponding concepts use triples: thing representation, representative, concept; a concept number can correspond to multiple different thing representations, and different concepts may have the same thing representation, which is why entity alignment is required.

[0094] The Jaccard coefficient is defined as the ratio of the size of the intersection of two texts to the size of the union of the two texts. The Levenshtein ratio is defined as the ratio of the value of the larger length of the two texts minus the edit distance to the larger length of the two texts. For both methods, the closer the value is to 1, the more similar the texts are.

[0095] Using the Jaccard coefficient alone for calculation, the entity alignment accuracy is high, but because the Jaccard coefficient is overly dependent on the surface similarity of the text, it performs poorly when the same concept is expressed differently. Using the Levenshetein ratio alone has a significantly better entity alignment effect in this case, but the scope of things expressed in the concepts of entity alignment may be large or small, and cannot accurately correspond to entities within the same scope of things.

[0096] Therefore, the weighted average of the Jaccard coefficient and the Levenshetein ratio is used to calculate the similarity between entities. The Jaccard coefficient weight is set to 0.3, the Levenshetein ratio weight is set to 0.7, and the threshold is set to 0.8 to achieve the entity alignment effect.

[0097] In a preferred embodiment of the present invention, during knowledge retrieval and filtering, a query statement template is constructed based on the relationship type in the knowledge graph as a classification basis, and the query template is merged according to the relationship type involved in the user question to perform knowledge query, and user intent recognition, medical named entity recognition and key entity extraction are performed on the user question. If it is recognized that the user question may contain a certain relationship type, a query template is constructed based on this relationship type, and the relationship type is fixed. If the question contains an entity of an entity type limited by this relationship type, it will be parsed one by one and its entity name and entity type will be passed into the query template being constructed. No restrictions are set for the other entity, and all the constructed query templates are merged, and finally the knowledge graph is queried using this query statement. This method utilizes the natural language understanding ability of the large language model, can accurately identify user intent, does not require the definition of a question template, and has strong applicability.

[0098] Use the big language model to build a big model agent, input the relationship type and its description into the big language model as prompts, and identify the relationships involved in the user's question;

[0099] Use all constructed query templates to perform joint queries (after the user asks a question and the query template is constructed) and return the queried knowledge triples and entity attribute triples;

[0100] In the case where no knowledge is found, a large model agent, namely a large language model application, is constructed. The key entities identified from the user questions are used to query using query templates that do not specify the relationship type. Multiple triples containing any key entities are queried and returned to the next process for knowledge filtering.

[0101] The question feature-triplet semantic matching method is adopted. The pre-trained model MacBert is used to generate the embedding of question features and triplets, and then the cosine similarity is calculated. The question feature is a concatenation of the key entity list identified from the user's question and the relationship type identified from it. Considering the cost of tokens consumption in the large model, only the top 300 knowledge triplets are selected and input into the large language model.

[0102] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0103] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

Claims

1. An intelligent question-answering system based on knowledge graph and large language model, characterized in that: It includes knowledge graph construction module, knowledge retrieval and filtering module and question-answering module; The knowledge graph construction module extracts relevant knowledge from the text; The knowledge retrieval and filtering module identifies the user's intention, calculates the cosine similarity between the user's question feature and each piece of knowledge, and finds knowledge with high relevance to the user's question in the knowledge graph; The question-answering module uses relevant knowledge as prompts and answers questions based on user questions.

2. The intelligent question-answering system based on knowledge graph and large language model according to claim 1, characterized in that: The knowledge graph construction module includes a knowledge extraction module, a knowledge proofreading module, and a knowledge fusion module; The knowledge extraction module collects data related to the field to be answered, and the knowledge extraction module includes an entity recognition module, a relationship extraction module and an attribute extraction module; The knowledge proofreading module includes an entity proofreading module and a relationship cleaning module, which respectively perform entity proofreading and relationship cleaning; The knowledge fusion module includes an ontology matching module and an alignment module. The knowledge fusion module integrates the entities, relations and attributes processed by different data sources.

3. The intelligent question-answering system based on knowledge graph and large language model as claimed in claim 1, characterized in that: The knowledge retrieval and filtering module identifies the user's intention, calculates the cosine similarity between the user's question feature and each piece of knowledge, and finds knowledge with high relevance to the user's question in the knowledge graph; The specific calculation process is: Represent the user's question features and each piece of knowledge as a vector; Calculate the cosine similarity between two knowledge vectors by the angle between them; Use the question and the queried knowledge triples to calculate the similarity, and add the relationship type information that the question may involve. Since the knowledge triples contain relationship fields, adding relationship type information improves the similarity between the question and the related relationship triples, and improves the relevance of the query results.

4. The intelligent question-answering system based on knowledge graph and large language model according to claim 1, characterized in that: The question-answering module uses the question-related knowledge triples obtained from the knowledge retrieval and filtering module as inputs to the large language model, and enables the large model to answer user questions based on relevant knowledge by setting prompts.

5. An intelligent question-answering method for the system according to any one of claims 1 to 4, characterized in that: The steps include: Collect data related to the problem, build a data model, that is, an entity and relationship framework, and build a knowledge graph based on the entity and relationship framework; Get input question; Based on the input question, identify the user's intention, calculate the cosine similarity between the user's question features and each piece of knowledge, find the knowledge that is highly relevant to the user's question in the knowledge graph, filter the knowledge, and obtain knowledge triples; Knowledge triples and user questions are input into a large language model to generate answers.

6. The intelligent question-answering method according to claim 5, characterized in that: The method of constructing a knowledge graph based on the entity and relationship framework is: Extract medical entities and relationships from the data model. Entities include symptoms or signs, groups, medical devices, instruments and materials, abnormal test results, tests, drugs, treatments, prognosis, anatomical parts and substances, semantics, institutions, etiology classes, and diseases. Proofread the extracted entities and clean the extracted relationships; Perform entity alignment; Import the cleaned entities, relationships, and attributes into the graph database; After repeatedly using the predefined entity and relationship frameworks to annotate data, train information extraction models, and compare information extraction results, the entity description framework and entity relationship description framework are finally determined; According to the defined entity and relationship types, randomly select sentences containing the specified type of entities and relationships from the collected relevant data for annotation. The annotated data basically covers all entity and relationship types, and the UIE unified information extraction model is trained using this annotated data. Select the uie-base and uie-m-base models as training models, set the learning rate to 1e-5, and change the number of training cycles to train the uie-base and uie-m-base models respectively; The number of training cycles is set to 300 to train the uie-base model; the number of training cycles is set to 150 to train the uie-base model. These two models are selected as the entity extraction and relationship extraction models respectively.

7. The intelligent question-answering method according to claim 6, characterized in that: Query the database to see if there is a corresponding entity. If there is, select 10 different expressions in the same language from the concepts corresponding to the queried entity to calculate the entity similarity. If they are considered similar, it means that the entity and the concept represent the same concept. In the knowledge graph, a bidirectional alias relationship is established for the entities representing the same concept, and entity alignment is completed. In the database, different expressions of the same thing and its corresponding concept use triples: thing expression, representative, concept; The weighted average of the Jaccard coefficient and the Levenshetein ratio is used to calculate the similarity between entities. The Jaccard coefficient weight is set to 0.3, the Levenshetein ratio weight is set to 0.7, and the threshold is set to 0.8 to achieve entity alignment.

8. The intelligent question-answering method according to claim 5, characterized in that: When retrieving and filtering knowledge, the query statement template is constructed based on the relationship type in the knowledge graph as the classification basis, and the user's question is identified for user intent, medical named entity recognition, and key entity extraction. If it is recognized that the user's question may contain a certain relationship type, a query template is constructed based on this relationship type, and the relationship type is fixed. If the question contains an entity of the entity type limited by this relationship type, it will be parsed one by one and its entity name and entity type will be passed into the query template being constructed. No restrictions are set for the other entity, and then all the constructed query templates are merged, and finally this query statement is used to query the knowledge graph; Use the big language model to build a big model agent, input the relationship type and its description into the big language model as prompts, and identify the relationships involved in the user's question; Use all constructed query templates to perform joint query and return the queried knowledge triples and entity attribute triples; In the case where no knowledge is found, a large model agent, namely a large language model application, is constructed. The key entities identified from the user questions are used to query using query templates that do not specify the relationship type. Multiple triples containing any key entities are checked and returned to the next process for knowledge filtering.

Citation Information

Cited By

  • Traditional Chinese medicine rehabilitation diagnosis system based on multi-modal knowledge graph and large language model

    CN120221058A