Large language model question and answer method based on knowledge graph enhancement
By combining knowledge graphs and graph neural networks, evidence subgraphs are constructed and feature extraction and relational reasoning are performed, the problem of insufficient interpretability and accuracy of large language models in the medical field is solved, and higher quality and credible medical Q&A results are generated.
Patent Information
- Application Number
- CN202510579730.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-07
AI Technical Summary
Large language models lack interpretability and deep understanding of professional fields in the reasoning process in the medical field, resulting in inaccurate or misleading information in the output results.
Combining knowledge graphs and graph neural networks, by extracting key entities and constructing evidence subgraphs, using graph neural networks to perform feature extraction and relationship reasoning, enhancing the reasoning ability of large language models, and adjusting the relationship between nodes through global smooth feature initialization and dynamic weight modeling to generate natural language explanations.
It improves the accuracy and interpretability of large language models in the medical field, reduces the risk of inaccurate information, and enhances users' trust in the decision-making process.
Smart Images

Figure CN120492581A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graph technology, and specifically to a large language model question answering method based on knowledge graph enhancement. Background Art
[0002] With the growth and complexity of medical data and the increasing demand for precision medicine, traditional diagnostic methods face challenges. To address this, large language models (LLMs) have been introduced into medical-assisted diagnosis systems. By processing natural language through deep learning, LLMs can rapidly analyze medical literature and medical records, helping doctors identify disease patterns, predict conditions, and provide recommendations, significantly improving clinical decision-making efficiency. However, large language models can sometimes produce seemingly plausible but inaccurate "hallucinations," which is unacceptable in the medical field, which demands high accuracy and reliability. To address this issue, knowledge graphs (KGs) have been introduced. They construct a structured network of concepts such as diseases, symptoms, and medications, and their relationships, enhancing case understanding and supporting logical reasoning. However, relying solely on knowledge graphs cannot fully meet the complex needs of the medical field. While KGs provide a relational structure between entities, they still have limitations in processing unstructured data and capturing complex contextual information, and their reasoning capabilities are also somewhat limited.
[0003] In order to further improve the reasoning accuracy and interpretability of the system, combining graph neural networks (GNNs) has become a new development direction. GNNs can effectively process graph-structured data in KGs and achieve deeper feature extraction and relationship reasoning through a message passing mechanism between nodes. In addition, GNNs can combine unstructured text information with structured KGs, allowing the system to not only obtain rich semantics from the text, but also use the clear relationships provided by KGs for accurate reasoning. This not only improves the model's ability to understand complex medical scenarios, but also increases the interpretability of the decision-making process. This method significantly reduces the risk of inaccurate or misleading information in the output results of large language models, which is critical to ensuring patient safety. Summary of the Invention
[0004] In order to overcome the shortcomings of the above technologies, the present invention provides a large language model question answering method based on knowledge graph enhancement, which can achieve higher-level reasoning, improve the accuracy and interpretability of answers, and make up for the shortcomings of LLMs.
[0005] The technical solution adopted by the present invention to overcome the technical problems is:
[0006] A large language model question answering method based on knowledge graph enhancement, including:
[0007] a) Obtain n medical problems and form a medical problem dataset Q, where Q = {q1, q2, ..., qi ,...,q n}, where q i is the i-th medical problem;
[0008] b) The i-th medical problem q i Input into the Qwen2.5-14B-Instruct model and output the i-th medical problem q i A preliminary answer i , all n preliminary answers form a preliminary answer set A, A={a1,a2,...,a i ,...,a n};
[0009] c) From the i-th medical problem q i and the i-th medical problem q i A preliminary answer i Extract z key entities from the dataset to form a key entity set E, where E = {e1, e2, ..., e k ,...,e z}, where e k is the kth key entity;
[0010] d) Using the Med-BERT model, we extract t triple entities from the DiseaseKG knowledge graph and obtain the triple entity set G, G = {u1,u2,...,u j ,...,u t}, where u j is the jth triple entity;
[0011] e) Using the kth key entity e k and the jth triplet entity u j Construct the highest similarity entity set V q ;
[0012] f) According to the highest similarity entity set V q Construct evidence subgraph K q ;
[0013] g) According to the evidence subgraph K q Construct the final node feature matrix J′;
[0014] h) Obtain the contextual enhanced feature T of the i-th neighbor node based on the final node feature matrix J′ i ctx ;
[0015] i) According to the evidence subgraph K q Construct a new subgraph Z, input the new subgraph Z into the Qwen2.5-14B-Instruct model, and output the natural language explanation Tm ;
[0016] j) Use Apollo Server to build a GraphQL interface and call the evidence subgraph K through HTTP requests in the Apollo Server resolver q The API enables the Qwen2.5-14B-Instruct model to be linked to the evidence subgraph K q , the context enhanced feature T i ctx and natural language interpretation T m Input to link to evidence subgraph K q In the Qwen2.5-14B-Instruct model, the output is the final generated answer A Q .
[0017] Furthermore, in step a), n medical questions are obtained from the CliMedBench medical question dataset.
[0018] Furthermore, in step c), the i-th medical problem q i and the i-th medical problem q i A preliminary answer i Input into the Med-BERT model and output z key entities.
[0019] Furthermore, step e) comprises the following steps:
[0020] e-1) The kth key entity e k Input into the 300-dimensional GloVe model and output the 300-dimensional key entity vector m k , k∈{1,...,z}, the j-th triplet entity u j Input into the 300-dimensional GloVe model and output the 300-dimensional triple entity vector g j , j∈{1,...,t};
[0021] e-2) The vector m of the key entity k Input into the encoder of the 12-layer Tradsformer, and output the text entity embedding h′ with context information e , the vector g of the triple entity j Input into a layer of graph convolutional network GCN, and output the knowledge graph entity embedding h′ with structural information g ;
[0022] e-3) Embed text entities with contextual information into h′ e Input into the multi-layer perceptron MLP, and output the projected embedding h″ e, embed the knowledge graph entity with structural information into h′ g Input into the multi-layer perceptron MLP, and output the projected embedding h″ g ;
[0023] e-4) Calculate the vector m of the key entity k The vector g of the triple entity j The cosine similarity of the local similarity S is obtained. local (m k ,g j );
[0024] e-5) Embed the projected image into h″ e and the projected embedding h″ g Input into the attention mechanism and calculate the global similarity S global (h″ e ,h″ g );
[0025] e-6) by formula Calculate the similarity S(m k ,g j ), where α and β are weight values;
[0026] e-7) Create a similarity matrix S with z rows and t columns. The value of the kth row and jth column of the similarity matrix S is the similarity S(m k ,g j ), k∈{1,...,z}, j∈{1,...,t}, select all vectors m containing key entities in the similarity matrix S k If the selected similarity is greater than or equal to the similarity threshold τ, it will be taken out, all the similarities taken out are sorted in descending order, and the vectors of the triple entities in the first k similarities after descending order are taken out, -1≤τ≤1, k=10, the m triple entities selected by the vectors of all z key entities constitute the highest similarity entity set V q , V q ={g1,g2,...,g i ,...,g m}, g i is the vector of the i-th triplet entity filtered out.
[0027] Further, step f) comprises the following steps:
[0028] f-1) Establish a queue set D, D = φ, and establish a visited node set V visited , V visited =φ, φ is an empty set;
[0029] f-2) The highest similarity entity set Vq The vector g of the i-th triplet entity in i In tuple (g i ,0) in the form of queue set D, i∈{1,...,m}, 0 is the vector g of the i-th triple entity i The initial depth of the i-th triplet entity is i Add as a node to the visited node set V visited middle;
[0030] f-3) Get the vector g of the i-th triple entity in the DiseaseKG knowledge graph i The x triple entities that have a relationship are the vector g of the i-th triple entity i Neighbor nodes, vector g of the i-th triplet entity i The jth neighbor node of j , set the jth neighbor node to g j In tuple (g j ,0+1) in the form of queue set D, 0+1 is the jth neighbor node g j The current depth of the ith triplet entity is i All x neighbor nodes are added to the visited node set V visited where j∈{1,...,x}, j≠i;
[0031] f-4) Set the visited node set V visited Cleaning is performed, and when cleaning, the visited node set V visited Only one duplicate node is retained and the rest of the redundant duplicate nodes are deleted;
[0032] f-5) Establish subgraph node structure K JD , K JD =φ, establish subgraph relationship set K GX ,
[0033] K GX =φ;
[0034] f-6) Take a tuple (g) from the queue set D i ,d), i∈{1,...,z}, d is the vector g of the i-th triplet entity i The current depth of d is determined to determine whether d is greater than or equal to the maximum depth d max If so, then the tuple (g i ,d) Delete from the queue set D, if otherwise check the vector g of the i-th triplet entity i Is the visited node set V visited In the example, if the vector g of the i-th triplet entity iNot in the visited node set V visited In the example, the vector g of the i-th triple entity is i Add to the visited node set V visited and will be composed of the vector g of the i-th triple entity i The edges formed by the relationship with each neighbor node are added to the subgraph relationship set K GX In the detection, the vector g of the i-th triplet entity i The jth neighbor node g j Whether it exists in the visited node set V visited If yes, then the jth neighbor node g j Add to subgraph node structure K JD And the tuple (g j ,d+1) is added to the queue set D;
[0035] f-7) Repeat step f-6) until the queue set D is empty;
[0036] f-8) Subgraph node structure K JD Cleaning is performed, and when cleaning, the subgraph node structure K JD Only one repeated neighbor node is retained, and the remaining redundant repeated neighbor nodes are deleted to obtain the cleaned subgraph node structure K JD ′, the subgraph relationship set K GX Cleaning is performed, and when cleaning, the subgraph relationship set K GX Only one duplicate edge is retained and the rest of the redundant duplicate edges are deleted to obtain the cleaned subgraph relationship set K GX ';
[0037] f-9) Create evidence subgraph K q =(K JD ′,K GX ′).
[0038] Preferably, d max =8.
[0039] Further, step g) comprises the following steps:
[0040] g-1) Evidence subgraph K q Input into the BERT model, and output the node feature matrix J with N rows and N columns, where N is the subgraph node structure K after cleaning JD The number of neighbor nodes in ′;
[0041] g-2) Construct the adjacency matrix A, where the value of the i-th row and j-th column of the adjacency matrix A is a ij , i∈{1,...,N}, j∈{1,...,N}, if the node structure K of the subgraph after cleaning JD′, there is an edge consisting of the relationship between the i-th neighbor node and the j-th neighbor node, then a ij =1, if the node structure K of the cleaned subgraph JD ′ does not contain an edge formed by the relationship between the i-th neighbor node and the j-th neighbor node, then a ij =0,a ii =0,a jj =0;
[0042] g-3) Add the adjacency matrix A to the N-order identity matrix to obtain the extended adjacency matrix A′, which can be obtained by the formula Calculate the smooth weight matrix A ^ , where D is the degree matrix corresponding to the extended adjacency matrix A′;
[0043] g-4) Through the formula Calculate the first smoothing feature Where r is the degree of influence of the neighbor node feature, r∈(0,1);
[0044] g-5) Through the formula Calculate the second smoothing feature
[0045] g-6) Repeat step g-5) t-2 times to obtain the smoothed features of the tth time
[0046] g-7) Solve the smoothing characteristics of the tth time The Frobenius norm of the t-1th smoothing characteristic The difference in the Frobenius norm of Then execute step g-8). If the difference is greater than Then execute step g-6) until the difference is less than or equal to
[0047] g-8) by formula Calculate the high-order characteristic matrix J (1) , where W (1) is the weight matrix, (·) is the nonlinear activation function, and the formula J (2) =(A^J (1) W (2) ) calculate the high-order characteristic matrix J (2) , where W (2) is the weight matrix;
[0048] g-9) The smooth feature of the tth time and the high-order characteristic matrix J (2) Perform residual connection to obtain the final node feature matrix J′ with N rows and N columns.
[0049] Preferably, in step g-7) Further, step h) comprises the following steps:
[0050] h-1) Input the element in the i-th row and j-th column of the final node feature matrix J′ into the multi-head attention mechanism with 8 attention heads, and output the attention score in is the attention score of the element in the i-th row and j-th column of the final node feature matrix J′ output by the i-th attention head of the multi-head attention mechanism, m∈{1,2,...,8}, i∈{1,...,N}, j∈{1,...,N}, Input into the softmax function and output the attention weight
[0051] h-2) The attention weight Multiply it with the value vector V of the i-th attention head to get the feature
[0052] h-3) Through the formula Calculate the features after time-varying weight adjustment Where γ (m) is the time-varying weight vector of the ith attention head, m∈{1,2,...,8}, γ (m) ∈(0,1), ⊙ is the element-by-element multiplication operation;
[0053] h-4) by formula Calculate the final node features Where, α i is the i-th weight value, i∈{1,2,...,8};
[0054] h-5) Node characteristics is the node feature matrix J final The element in the i-th row and j-th column of the node feature matrix J final Perform residual connection with the final node feature matrix J′ to obtain the feature matrix J out ; h-6) The feature matrix J out The global eigenvector P is obtained by taking the average value of the N×N elements in global , the evidence subgraph K q Input into the BERT model, and output the feature vectors of all neighbor nodes. The feature vector of the i-th neighbor node is T i , the feature vector of the jth neighbor node is T j ;
[0055] h-7) by formula Calculate the attention weight d after weighted average ij , β i is the i-th weight value, i∈{1,2,...,8}, through the formula
[0056] s ij =d ij ·cosine_similarity(T i ·P global )·cosine_similarity(T j ·P global )
[0057] Calculate the importance score s of the jth neighbor node to the ith neighbor node ij , where cosine_similarity(T i ·P global ) is the characteristic vector of the i-th neighbor node T i and the global eigenvector P global The cosine similarity, cosine_similarity(T j ·P global ) is the feature vector of the jth neighbor node T j and the global eigenvector P global The cosine similarity of , i∈{1,...,N}, j∈{1,...,N};
[0058] h-8) Sort the importance scores of all N-1 neighbor nodes to the i-th neighbor node in descending order, and select the neighbor nodes with the first k importance scores that are opposite to the i-th neighbor node as the neighborhood set N of the i-th neighbor node. k (i), k = 10, by formula Calculate the context enhancement feature T of the i-th neighbor node i ctx .
[0059] Furthermore, in step i), according to the evidence subgraph K q Construct a new subgraph Z i The method is as follows: i-1) construct the visited node set F, F = φ, and construct the edge set R, R = φ;
[0060] i-2) From the cleaned subgraph relationship set K GX ′, find y vectors g with the i-th triple entity i The edge with the starting point is the vector g of the ith triplet entity. i Neighbor node g j, i∈{1,...,m}, q∈{1,...,y}, j∈{1,...,x}, the neighbor node g j Add to the visited node set F, and add the qth edge to the edge set R;
[0061] i-3) Create a new subgraph Z = (F, R).
[0062] The beneficial effects of the present invention are as follows: first, key entities are extracted from medical texts, and evidence subgraphs are constructed using knowledge graphs to enhance the reasoning ability of large language models. Graph neural networks are used to perform feature extraction and relationship reasoning on subgraphs to enhance the model's understanding of medical scenarios. Through global smooth feature initialization and dynamic weight modeling, the relationship between nodes is precisely adjusted, and the reasoning space is narrowed to the most relevant part. In addition, the subgraph is formatted as an entity chain, converted into a natural language description, and integrated into a unified reasoning graph to provide a comprehensive perspective for the model. Ultimately, the model generates the final answer based on the reasoning graph and key reasoning elements, while constructing an explanatory context to enhance the interpretability of the model. The present invention effectively solves the problem of insufficient accuracy and interpretability of large language models in medical applications, and provides an innovative solution for the development of medical question-answering systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0064] The following is combined with Figure 1 The present invention is further described.
[0065] A large language model question answering method based on knowledge graph enhancement, including:
[0066] a) Obtain n medical problems and form a medical problem dataset Q, where Q = {q1, q2, ..., q i ,...,q n}, where q i is the i-th medical problem.
[0067] b) The i-th medical problem q i Input into the Qwen2.5-14B-Instruct model and output the i-th medical problem q i A preliminary answer i , all n preliminary answers form a preliminary answer set A, A={a1,a2,...,a i ,...,a n}.
[0068] c) From the i-th medical problem q i and the i-th medical problem q iA preliminary answer i Extract z key entities from the dataset to form a key entity set E, where E = {e1, e2, ..., e k ,...,e z}, where e k is the kth key entity.
[0069] d) Using the Med-BERT model, we extract t triple entities from the DiseaseKG knowledge graph and obtain the triple entity set G, G = {u1,u2,...,u j ,...,u t}, where u j is the jth triplet entity.
[0070] e) Using the kth key entity e k and the jth triplet entity u j Construct the highest similarity entity set V q ; f) According to the highest similarity entity set V q Construct evidence subgraph K q .
[0071] g) According to the evidence subgraph K q The final node feature matrix J′ is constructed.
[0072] h) Obtain the contextual enhanced feature T of the i-th neighbor node based on the final node feature matrix J′ i ctx .
[0073] i) According to the evidence subgraph K q Construct a new subgraph Z, input the new subgraph Z into the Qwen2.5-14B-Instruct model, and output the natural language explanation T m .
[0074] j) Use Apollo Server to build a GraphQL interface and call the evidence subgraph K through HTTP requests in the Apollo Server resolver q The API enables the Qwen2.5-14B-Instruct model to be linked to the evidence subgraph K q , the context enhanced feature T i ctx and natural language interpretation T m Input to link to evidence subgraph K q In the Qwen2.5-14B-Instruct model, the output is the final generated answer A Q .
[0075] This approach addresses the lack of explainability and deep domain understanding inherent in large language models (LLMs) during reasoning. First, key entities—factual assertions requiring further verification—are extracted from the input question and the LLM's initial response. The LLM's inherent capabilities are leveraged to extract key information from the text. Next, the extracted key entities are encoded with entities in an external knowledge graph, and a cosine similarity matrix is calculated to obtain the set of entities with the highest similarity scores, which are then used to construct the evidence subgraph.
[0076] Then, an evidence subgraph is constructed from the knowledge graph based on the extracted entities, generating a compact and highly relevant evidence graph.
[0077] When processing evidence subgraphs, we use a global smoothing feature initialization method and a dynamic weighting approach to model local relationships. We extract context-enhanced features from the evidence subgraphs as key reasoning elements, capturing global dependency information and local relationship changes, precisely narrowing the reasoning space to the most relevant knowledge components. We format all subgraphs as entity chains, and leverage a large language model to translate knowledge graph information into natural language interpretations, providing a comprehensive perspective on all evidence subgraphs.
[0078] Based on natural language interpretation of facts, the selected information and key reasoning elements are selected, and a large language model is used to verify each factual assertion and make revision suggestions. The preliminary answer is adjusted according to the verification results, which improves the reasoning ability and accuracy of the large language model in the medical field and enhances its interpretability, enabling users to better understand and trust the model's decision-making process.
[0079] In one embodiment of the present invention, in step a), n medical questions are obtained from the CliMedBench medical question dataset.
[0080] In one embodiment of the present invention, in step c), the i-th medical question q i and the i-th medical problem q i A preliminary answer i Input into the Med-BERT model and output z key entities.
[0081] In one embodiment of the present invention, step e) comprises the following steps:
[0082] e-1) The kth key entity e k Input into the 300-dimensional GloVe model and output the 300-dimensional key entity vector m k , k∈{1,...,z}, the j-th triplet entity u j Input into the 300-dimensional GloVe model and output the 300-dimensional triple entity vector g j , j∈{1,...,t}.
[0083] e-2) The vector m of the key entity k Input into the encoder of the 12-layer Tradsformer, and output the text entity embedding h′ with context information e , the vector g of the triple entity j Input into a layer of graph convolutional network GCN, and output the knowledge graph entity embedding h′ with structural information g .
[0084] e-3) Embed text entities with contextual information into h′ e Input into the multi-layer perceptron MLP, and output the projected embedding h″ e , embed the knowledge graph entity with structural information into h′ g Input into the multi-layer perceptron MLP, and output the projected embedding h″ g .
[0085] e-4) Calculate the vector m of the key entity k The vector g of the triple entity j The cosine similarity of the capture substructure and fine-grained features is used to obtain the local similarity S local (m k ,g j );
[0086] e-5) Embed the projected image into h″ e and the projected embedding h″ g Input into the attention mechanism to capture the matching relationship of the overall semantic features and calculate the global similarity S global (h″ e ,h″ g ).
[0087] e-6) by formula Calculate the similarity S(m k ,g j ), where α and β are weight values.
[0088] e-7) Create a similarity matrix S with z rows and t columns. The value of the kth row and jth column of the similarity matrix S is the similarity S(m k ,g j ), k∈{1,...,z}, j∈{1,...,t}, select all vectors m containing key entities in the similarity matrix S kIf the selected similarity is greater than or equal to the similarity threshold τ, it will be taken out, all the similarities taken out are sorted in descending order, and the vectors of the triple entities in the first k similarities after descending order are taken out, -1≤τ≤1, k=10, the m triple entities selected by the vectors of all z key entities constitute the highest similarity entity set V q , V q ={g1,g2,...,g i ,...,g m}, g i is the vector of the i-th triplet entity filtered out.
[0089] In one embodiment of the present invention, step f) comprises the following steps:
[0090] f-1) Establish a queue set D, D = φ, and establish a visited node set V visited , V visited =φ, φ is an empty set.
[0091] f-2) The highest similarity entity set V q The vector g of the i-th triplet entity in i In tuple (g i ,0) in the form of queue set D, i∈{1,...,m}, 0 is the vector g of the i-th triple entity i The initial depth of the i-th triplet entity is i Add as a node to the visited node set V visited middle.
[0092] f-3) Get the vector g of the i-th triple entity in the DiseaseKG knowledge graph i The x triple entities that have a relationship are the vector g of the i-th triple entity i Neighbor nodes, vector g of the i-th triplet entity i The jth neighbor node of j , set the jth neighbor node to g j In tuple (g j ,0+1) in the form of queue set D, 0+1 is the jth neighbor node g j The current depth of the ith triplet entity is i All x neighbor nodes are added to the visited node set V visited where j∈{1,...,x}, j≠i.
[0093] f-4) Set the visited node set V visited Cleaning is performed, and when cleaning, the visited node set V visitedOnly one duplicate node is retained and the rest of the duplicate nodes are deleted.
[0094] f-5) Establish subgraph node structure K JD , K JD =φ, establish subgraph relationship set K GX ,
[0095] K GX =φ.
[0096] f-6) Take a tuple (g) from the queue set D i ,d), i∈{1,...,z}, d is the vector g of the i-th triplet entity i The current depth of d is determined to determine whether d is greater than or equal to the maximum depth d max If so, then the tuple (g i ,d) Delete from the queue set D, if otherwise check the vector g of the i-th triplet entity i Is the visited node set V visited In the example, if the vector g of the i-th triplet entity i Not in the visited node set V visited In the example, the vector g of the i-th triple entity is i Add to the visited node set V visited and will be composed of the vector g of the i-th triple entity i The edges formed by the relationship with each neighbor node are added to the subgraph relationship set K GX In the detection, the vector g of the i-th triplet entity i The jth neighbor node g j Whether it exists in the visited node set V visited If yes, then the jth neighbor node g j Add to subgraph node structure K JD And the tuple (g j ,d+1) is added to the queue set D.
[0097] f-7) Repeat step f-6) until the queue set D is empty.
[0098] f-8) Subgraph node structure K JD Cleaning is performed, and when cleaning, the subgraph node structure K JD Only one repeated neighbor node is retained, and the remaining redundant repeated neighbor nodes are deleted to obtain the cleaned subgraph node structure K JD ′, the subgraph relationship set K GX Cleaning is performed, and when cleaning, the subgraph relationship set K GX Only one duplicate edge is retained and the rest of the redundant duplicate edges are deleted to obtain the cleaned subgraph relationship set KGX ′.
[0099] f-9) Create evidence subgraph K q =(K JD ′,K GX ′).
[0100] In this embodiment, preferably, d max =8.
[0101] In one embodiment of the present invention, step g) comprises the following steps:
[0102] g-1) When processing the evidence subgraph K q Through comparative experiments, the method of the present invention is significantly superior to the traditional LLM method and the method using only KG in terms of generation quality, reasoning accuracy and hallucination in the medical field. In particular, it can generate higher quality and more accurate answers in the treatment of complex problems in the medical field, and provide stronger interpretability. Specifically, the evidence subgraph K q Input into the BERT model, and output the node feature matrix J with N rows and N columns, where N is the subgraph node structure K after cleaning JD ′ is the number of neighbor nodes. g-2) In order to propagate node feature information in the graph, an adjacency matrix A is constructed. The value of the i-th row and j-th column of the adjacency matrix A is a ij , i∈{1,...,N}, j∈{1,...,N}, if the node structure K of the subgraph after cleaning JD ′, there is an edge consisting of the relationship between the i-th neighbor node and the j-th neighbor node, then a ij =1, if the node structure K of the cleaned subgraph JD ′ does not contain an edge formed by the relationship between the i-th neighbor node and the j-th neighbor node, then a ij =0,a ii =0,a jj =0.
[0103] g-3) Add the adjacency matrix A to the N-order identity matrix to obtain the extended adjacency matrix A′, which can be obtained by the formula Calculate the smooth weight matrix A ^ , where D is the degree matrix corresponding to the extended adjacency matrix A′, which is used to balance the amplitude of feature propagation.
[0104] g-4) Through the formula Calculate the first smoothing feature Where r is the degree of influence of the neighbor node features, r∈(0,1).
[0105] g-5) Through the formula Calculate the second smoothing feature
[0106] g-6) Repeat step g-5) t-2 times to obtain the smoothed features of the tth time
[0107] g-7) Solve the smoothing characteristics of the tth time The Frobenius norm of the t-1th smoothing characteristic The difference in the Frobenius norm of Then execute step g-8). If the difference is greater than Then execute step g-6) until the difference is less than or equal to g-8) by formula Calculate the high-order characteristic matrix J (1) , where W (1) is the weight matrix, (·) is the nonlinear activation function, and the formula J (2) =(A^J (1) W (2) ) calculate the high-order characteristic matrix J (2) , where W (2) is the weight matrix.
[0108] g-9) The smooth feature of the tth time and the high-order characteristic matrix J (2) Perform residual connection to retain the original feature information and obtain the final node feature matrix J′ with N rows and N columns.
[0109] In this embodiment, preferably, in step g-7)
[0110] In one embodiment of the present invention, the relationships between nodes may vary across different stages and tasks of the graph. Therefore, a method for modeling local relationships with dynamic weights is proposed to more carefully model the relationships between nodes in the evidence subgraph. Therefore, we introduce time-varying relationship modeling, where the strength of the relationship between a node and its neighbors changes dynamically as training progresses. This allows the network to gradually adjust the weights of local relationships based on the requirements of the task and the structure of the graph. Specifically, step h) includes the following steps:
[0111] h-1) Input the element in the i-th row and j-th column of the final node feature matrix J′ into the multi-head attention mechanism with 8 attention heads, and output the attention score in is the attention score of the element in the i-th row and j-th column of the final node feature matrix J′ output by the i-th attention head of the multi-head attention mechanism, m∈{1,2,...,8}, i∈{1,...,N}, j∈{1,...,N}, Input into the softmax function and output the attention weight
[0112] h-2) The attention weight Multiply it with the value vector V of the i-th attention head to get the feature
[0113] h-3) Through the formula Calculate the features after time-varying weight adjustment Where γ (m) is the time-varying weight vector of the i-th attention head, which is used to adjust the relationship strength between nodes, m∈{1,2,...,8}, γ (m) ∈(0,1), ⊙ is the element-by-element multiplication operation.
[0114] h-4) by formula Calculate the final node features Where, α i is the i-th weight value, i∈{1,2,...,8}. h-5) Node features is the node feature matrix J final The element in the i-th row and j-th column of the node feature matrix J final Perform residual connection with the final node feature matrix J′ to obtain the feature matrix J out h-6) The feature matrix J out The global eigenvector P is obtained by taking the average value of the N×N elements in global , the evidence subgraph K q Input into the BERT model, and output the feature vectors of all neighbor nodes. The feature vector of the i-th neighbor node is T i , the feature vector of the jth neighbor node is T j .
[0115] h-7) by formula Calculate the attention weight d after weighted average ij , β i is the i-th weight value, i∈{1,2,...,8}, through the formula
[0116] s ij =d ij ·cosine_similarity(T i ·P global )·cosine_similarity(T j ·P global ) Calculate the importance score s of the jth neighbor node to the ith neighbor node ij, where cosine_similarity(T i ·P global ) is the characteristic vector of the i-th neighbor node T i and the global eigenvector P global The cosine similarity, cosine_similarity(T j ·P global ) is the feature vector of the jth neighbor node T j and the global eigenvector P global The cosine similarity of , i∈{1,...,N}, j∈{1,...,N}.
[0117] h-8) Finally, when selecting key neighborhoods and stacking multiple layers, we introduce a recursive neighborhood expansion mechanism combined with sparse selection to gradually expand the neighborhood of the evidence subgraph at each layer while ensuring that the information aggregation of each layer is more efficient. Specifically, all N-1 neighbor nodes are sorted in descending order of importance scores to the i-th neighbor node, and the neighbor nodes with the top k importance scores that are opposite to the i-th neighbor node are selected as the neighborhood set N of the i-th neighbor node. k (i), k = 10, by formula Calculate the context enhancement feature T of the i-th neighbor node i ctx .
[0118] In one embodiment of the present invention, in step i), according to the evidence subgraph K q Construct a new subgraph Z i The method is:
[0119] i-1) Construct the visited node set F, F = φ, and construct the edge set R, R = φ.
[0120] i-2) From the cleaned subgraph relationship set K GX ′, find y vectors g with the i-th triple entity i The edge with the starting point is the vector g of the ith triplet entity. i Neighbor node g j , i∈{1,…,m},q∈{1,…,y},j∈{1,…,x}, the neighbor node g j Add it to the visited node set F and add the qth edge to the edge set R.
[0121] i-3) Create a new subgraph Z = (F, R).
[0122] In order to verify the effectiveness of the knowledge graph-enhanced large language model medical question answering method of the present invention, we evaluated the performance of the present invention in complex medical question answering tasks using two medical question answering datasets, GenMedGPT-5k (as shown in Table 1) and CMCQA (as shown in Table 2). We used two indicators, BERTScore and GPT-4Rating, for quantitative evaluation. BERTScore measures the semantic similarity between the generated answer and the reference answer. GPT4 is used to rank the quality of the answer based on the actual situation, and compares the answers based on four criteria: response diversity and completeness, overall factual correctness, disease diagnosis correctness, and drug recommendation correctness. As well as a calculation indicator of the hallucination phenomenon, by calculating the tfidf similarity score between different output sentences. The lower the score, the more hallucinations there are in the answer.
[0123] Table 1 BERT scores and GPT-4 rankings of all methods on the GenMedGPT-5k dataset
[0124] method BERT score GPT-4 ranking hallucinations Method of the present invention 0.7896 0.8189 0.8154 GPT-3.5 0.7612 0.7945 0.8001 BM25 Retriever 0.7583 0.7816 0.7425 KG Retriever 0.7698 0.8030 0.7866 Tree-of-thought 0.7202 0.7949 0.7658 GPT-4 0.7789 0.8102 0.8065
[0125] Table 2. GPT-4 ranking of BERT scores of all methods on the CMCQA dataset
[0126] method BERT score GPT-4 ranking hallucinations Method of the present invention 0.9516 0.9324 0.8962 GPT-3.5 0.9226 0.9310 0.8156 BM25 Retriever 0.9102 0.9289 0.7861 KG Retriever 0.9356 0.9246 0.8068 Tree-of-thought 0.9286 0.9166 0.8234 GPT-4 0.9396 0.9331 0.8315
[0127] Through comparative experiments, the method of the present invention is significantly superior to the traditional LLM method and the method using only KG in terms of generation quality, reasoning accuracy and hallucination in the medical field. Especially in the processing of complex problems in the medical field, it can generate higher quality and more accurate answers and provide stronger interpretability.
[0128] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A large language model question answering method based on knowledge graph enhancement, characterized in that: include: a) Obtain n medical problems and form a medical problem dataset Q, where Q = {q1, q2, ..., q i ,...,q n }, where q i is the i-th medical problem; b) The i-th medical problem q i Input into the Qwen2.5-14B-Instruct model and output the i-th medical problem q i A preliminary answer i , all n preliminary answers form a preliminary answer set A, A={a1,a2,...,a i ,…,a n }; c) From the i-th medical problem q i and the i-th medical problem q i A preliminary answer i Extract z key entities from the key entity set E, E = {e1, e2, ..., e k ,...,e z }, where e k is the kth key entity; d) Using the Med-BERT model, we extract t triple entities from the DiseaseKG knowledge graph and obtain the triple entity set G, G = {u1,u2,...,u j ,...,u t }, where u j is the jth triple entity; e) Using the kth key entity e k and the jth triplet entity u j Construct the highest similarity entity set V q ; f) According to the highest similarity entity set V q Construct evidence subgraph K q ; g) According to the evidence subgraph K q Construct the final node feature matrix J′; h) Obtain the contextual enhanced feature T of the i-th neighbor node based on the final node feature matrix J′ i ctx ; i) According to the evidence subgraph K q Construct a new subgraph Z, input the new subgraph Z into the Qwen2.5-14B-Instruct model, and output the natural language explanation T m ; j) Use Apollo Server to build a GraphQL interface and call the evidence subgraph K through HTTP requests in the Apollo Server resolver q The API enables the Qwen2.5-14B-Instruct model to be linked to the evidence subgraph K q , the context enhanced feature T i ctx and natural language interpretation T m Input to link to evidence subgraph K q In the Qwen2.5-14B-Instruct model, the output is the final generated answer A Q .
2. The large language model question answering method based on knowledge graph enhancement according to claim 1 is characterized by: In step a), n medical problems are obtained from the CliMedBench medical problem dataset.
3. The large language model question answering method based on knowledge graph enhancement according to claim 1 is characterized by: In step c), the i-th medical problem q i and the i-th medical problem q i A preliminary answer i Input into the Med-BERT model and output z key entities.
4. The large language model question answering method based on knowledge graph enhancement according to claim 1 is characterized in that Step e) comprises the following steps: e-1) The kth key entity e k Input into the 300-dimensional GloVe model and output the 300-dimensional key entity vector m k , k∈{1,...,z}, the j-th triplet entity u j Input into the 300-dimensional GloVe model and output the 300-dimensional triple entity vector g j , j∈{1,...,t}; e-2) The vector m of the key entity k Input into the encoder of the 12-layer Tradsformer, and output the text entity embedding h′ with context information e , the vector g of the triple entity j Input into a layer of graph convolutional network GCN, and output the knowledge graph entity embedding h′ with structural information g ; e-3) Embed text entities with contextual information into h′ e Input into the multi-layer perceptron MLP, and output the projected embedding h″ e , embed the knowledge graph entity with structural information into h′ g Input into the multi-layer perceptron MLP, and output the projected embedding h″ g ; e-4) Calculate the vector m of the key entity k Vector g of triple entity j The cosine similarity of the local similarity S is obtained. local (m k ,g j ); e-5) Embed the projected image into h″ e and the projected embedding h″ g Input into the attention mechanism and calculate the global similarity S global (h″ e ,h″ g ); e-6) by formula Calculate the similarity S(m k ,g j ), where α and β are weight values; e-7) Create a similarity matrix S with z rows and t columns. The value of the kth row and jth column of the similarity matrix S is the similarity S(m k ,g j ), k∈{1,...,z}, j∈{1,...,t}, select all vectors m containing key entities in the similarity matrix S k If the selected similarity is greater than or equal to the similarity threshold τ, it will be taken out, all the similarities taken out are sorted in descending order, and the vectors of the triple entities in the first k similarities after descending order are taken out, -1≤τ≤1, k=10, the m triple entities selected by the vectors of all z key entities constitute the highest similarity entity set V q , V q ={g1,g2,...,g i ,...,g m }, g i is the vector of the i-th triplet entity filtered out.
5. The large language model question answering method based on knowledge graph enhancement according to claim 4 is characterized in that Step f) comprises the following steps: f-1) Establish a queue set D, D = φ, and establish a visited node set V visited , V visited =φ, φ is an empty set; f-2) The highest similarity entity set V q The vector g of the i-th triplet entity in i In tuple (g i ,0) in the form of queue set D, i∈{1,...,m}, 0 is the vector g of the i-th triple entity i The initial depth of the i-th triplet entity is i Add as a node to the visited node set V visited middle; f-3) Get the vector g of the i-th triple entity in the DiseaseKG knowledge graph i The x triple entities that have a relationship are the vector g of the i-th triple entity i Neighbor nodes, vector g of the i-th triplet entity i The jth neighbor node of j , set the jth neighbor node to g j In tuple (g j ,0+1) in the form of queue set D, 0+1 is the jth neighbor node g j The current depth of the ith triplet entity is i All x neighbor nodes are added to the visited node set V visited where j∈{1,...,x}, j≠i; f-4) Set the visited node set V visited Cleaning is performed, and when cleaning, the visited node set V visited Only one duplicate node is retained and the rest of the redundant duplicate nodes are deleted; f-5) Establish subgraph node structure K JD , K JD =φ, establish subgraph relationship set K GX , K GX =φ; f-6) Take a tuple (g) from the queue set D i ,d), i∈{1,...,z}, d is the vector g of the i-th triplet entity i The current depth of d is determined to determine whether d is greater than or equal to the maximum depth d max If so, then the tuple (g i ,d) Delete from the queue set D, if otherwise check the vector g of the i-th triplet entity i Is the visited node set V visited In the example, if the vector g of the i-th triplet entity i Not in the visited node set V visited In the example, the vector g of the i-th triple entity is i Add to the visited node set V visited and will be composed of the vector g of the i-th triple entity i The edges formed by the relationship with each neighbor node are added to the subgraph relationship set K GX In the detection, the vector g of the i-th triplet entity i The jth neighbor node g j Whether it exists in the visited node set V visited If yes, then the jth neighbor node g j Add to subgraph node structure K JD And the tuple (g j ,d+1) is added to the queue set D; f-7) Repeat step f-6) until the queue set D is empty; f-8) Subgraph node structure K JD Cleaning is performed, and when cleaning, the subgraph node structure K JD Only one repeated neighbor node is retained, and the remaining redundant repeated neighbor nodes are deleted to obtain the cleaned subgraph node structure K JD ′, the subgraph relationship set K GX Cleaning is performed, and when cleaning, the subgraph relationship set K GX Only one duplicate edge is retained and the rest of the redundant duplicate edges are deleted to obtain the cleaned subgraph relationship set K GX '; f-9) Create evidence subgraph K q =(K JD ′,K GX ′).
6. The large language model question answering method based on knowledge graph enhancement according to claim 5 is characterized by: d max =8。 7. The large language model question answering method based on knowledge graph enhancement according to claim 1 is characterized in that Step g) comprises the following steps: g-1) Evidence subgraph K q Input into the BERT model, and output the node feature matrix J with N rows and N columns, where N is the subgraph node structure K after cleaning JD The number of neighbor nodes in ′; g-2) Construct the adjacency matrix A, where the value of the i-th row and j-th column of the adjacency matrix A is a ij , i∈{1,...,N}, j∈{1,...,N}, if the node structure K of the subgraph after cleaning JD ′, there is an edge consisting of the relationship between the i-th neighbor node and the j-th neighbor node, then a ij =1, if the node structure K of the cleaned subgraph JD ′ does not contain an edge formed by the relationship between the i-th neighbor node and the j-th neighbor node, then a ij =0,a ii =0,a jj =0; g-3) Add the adjacency matrix A to the N-order identity matrix to obtain the extended adjacency matrix A′, which can be obtained by the formula The smoothed weight matrix A^ is calculated, where D is the degree matrix corresponding to the extended adjacency matrix A′; g-4) Through the formula Calculate the first smoothing feature Where r is the degree of influence of the neighbor node feature, r∈(0,1); g-5) Through the formula Calculate the second smoothing feature g-6) Repeat step g-5) t-2 times to obtain the smoothed features of the tth time g-7) Solve the smoothing characteristics of the tth time The Frobenius norm of the t-1th smoothing characteristic The difference in the Frobenius norm of Then execute step g-8), if the difference is greater than Then execute step g-6) until the difference is less than or equal to g-8) by formula Calculate the high-order characteristic matrix J (1) , where W (1) is the weight matrix, (·) is the nonlinear activation function, and the formula J (2) =(A^J (1) W (2) ) calculate the high-order characteristic matrix J (2) , where W (2) is the weight matrix; g-9) The smooth feature of the tth time and the high-order characteristic matrix J (2) Perform residual connection to obtain the final node feature matrix J′ with N rows and N columns.
8. The large language model question answering method based on knowledge graph enhancement according to claim 1 is characterized by: In step g-7) 9. The large language model question answering method based on knowledge graph enhancement according to claim 1 is characterized in that Step h) comprises the following steps: h-1) Input the element in the i-th row and j-th column of the final node feature matrix J′ into the multi-head attention mechanism with 8 attention heads, and output the attention score in is the attention score of the element in the i-th row and j-th column of the final node feature matrix J′ output by the i-th attention head of the multi-head attention mechanism, m∈{1,2,...,8}, i∈{1,...,N}, j∈{1,...,N}, Input into the softmax function and output the attention weight h-2) The attention weight Multiply it with the value vector V of the i-th attention head to get the feature h-3) Through the formula Calculate the features after time-varying weight adjustment Where γ (m) is the time-varying weight vector of the ith attention head, m∈{1,2,...,8}, γ (m) ∈(0,1), ⊙ is the element-by-element multiplication operation; h-4) by formula Calculate the final node features Where, α i is the i-th weight value, i∈{1,2,...,8}; h-5) Node characteristics is the node feature matrix J final The element in the i-th row and j-th column of the node feature matrix J final Perform residual connection with the final node feature matrix J′ to obtain the feature matrix J out ; h-6) The feature matrix J out The global eigenvector P is obtained by taking the average value of the N×N elements in global , the evidence subgraph K q Input into the BERT model, and output the feature vectors of all neighbor nodes. The feature vector of the i-th neighbor node is T i , the feature vector of the jth neighbor node is T j ; h-7) by formula Calculate the attention weight d after weighted average ij , β i is the i-th weight value, i∈{1,2,...,8}, through the formula s ij =d ij ·cosine_similarity(T i ·P global )·cosine_similarity(T j ·P global ) Calculate the importance score s of the jth neighbor node to the ith neighbor node ij , where cosine_similarity(T i ·P global ) is the characteristic vector of the i-th neighbor node T i and the global eigenvector P global The cosine similarity, cosine_similarity(T j ·P global ) is the feature vector of the jth neighbor node T j and the global eigenvector P global The cosine similarity of , i∈{1,...,N}, j∈{1,...,N}; h-8) Sort the importance scores of all N-1 neighbor nodes to the i-th neighbor node in descending order, and select the neighbor nodes with the first k importance scores that are opposite to the i-th neighbor node as the neighborhood set N of the i-th neighbor node. k (i), k = 10, by formula Calculate the context enhancement features of the i-th neighbor node 10. The large language model question answering method based on knowledge graph enhancement according to claim 5 is characterized in that In step i), according to the evidence subgraph K q Construct a new subgraph Z i The method is: i-1) Construct the visited node set F, F = φ, and the edge set R, R = φ; i-2) From the cleaned subgraph relationship set K GX ′, find y vectors g with the i-th triple entity i The edge with the starting point is the vector g of the ith triplet entity. i Neighbor node g j , i∈{1,…,m}, q∈{1,…,y}, j∈{1,…,x}, the neighbor node g j Add to the visited node set F, and add the qth edge to the edge set R; i-3) Create a new subgraph Z = (F, R).
Citation Information
Patent Citations
Open domain natural language reasoning question-answering system and method driven by large language model
CN116932708A
Knowledge graph-based few-sample multi-hop reasoning optimization method
CN118095445A
Knowledge graph completion method and system fusing entity description information and graph attention
CN119312891A
Medical knowledge graph question and answer method based on GAT and probabilistic reasoning
CN119578543A
Method and system for generating questions and answers based on knowledge graph
KR102697127B1
Cited By
Medical field large language model training method and system based on knowledge graph enhancement
CN121998101A