Method for enhancing large language model by knowledge graph during testing

The non-invasive knowledge graph enhancement of large language models is achieved through GraphRAG technology and multi-head attention mechanism, which solves the problems of high cost, poor dynamics and semantic alignment deviation of the knowledge graph enhancement method, and improves the dynamic knowledge update and complex reasoning capabilities of the model.

CN120492637AActive Publication Date: 2025-08-15SOUTHEAST UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510652032.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-15
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The prior art has problems in large language models with high cost, poor dynamics, disruption of original attention patterns, semantic alignment deviations and low retrieval efficiency.

Method used

Using a non-invasive knowledge graph enhancement method, triplets are retrieved in real time through GraphRAG technology, combining multi-head attention mechanism and feedforward neural network, two-way information interaction between text and knowledge graphs is realized, and knowledge representation is dynamically updated.

Benefits of technology

Without changing the model parameters, the dynamic knowledge update ability and complex inference performance of large language models are improved, and the fluency and accuracy of language generation are maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492637A_ABST
    Figure CN120492637A_ABST
Patent Text Reader

Abstract

The invention provides a method for enhancing a large language model through a knowledge graph during testing. According to the method, a novel attention mechanism based on the knowledge graph is provided, the effect of enhancing a large language model by the knowledge graph is improved in a bidirectional information flow aggregation mode, and the method comprises the following steps that triples are retrieved from the knowledge graph according to a text input sequence to serve as knowledge graph triple input sequences; the text input sequence and the knowledge graph triple input sequence are converted into embedded vectors, and a query matrix, a key matrix and a value matrix are calculated; mixed attention is calculated; richer feature representation is obtained through a feedforward neural network; and generating model prediction probability distribution according to the feature representation so as to predict the next word.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing and knowledge graph enhancement, and specifically relates to a method for dynamically integrating knowledge graphs (KGs) in the reasoning stage of large language models (LLMs). By designing a bidirectional hybrid attention mechanism, non-invasive knowledge enhancement is achieved, thereby improving the model's factual reasoning ability while retaining its original language generation performance. Background Art

[0002] The Knowledge Graph provides interpretable symbolic knowledge support for language models through a structured semantic network of entities, relationships, and entities. While large language models (LLMs) have demonstrated excellent performance in text generation and comprehension tasks, they still face core issues such as factual errors, delayed temporal knowledge updates, and broken chains of complex reasoning. Research on knowledge graph enhancement for large language models has been extensively conducted. By combining the explicit knowledge system of the knowledge graph with the implicit semantic representation of the LLM, the proportion of fictitious content in model generation can be reduced, improving the credibility of answers in complex question-answering tasks. Therefore, knowledge graph enhancement for large language models has become an important research direction for improving the capabilities of large language models.

[0003] Currently, academia has proposed several methods for augmenting large language models with knowledge graphs. The first and most direct approach is to incorporate a large amount of structured knowledge from the knowledge graph into the pre-training data. This naturally achieves the goal of knowledge graph augmentation, but the investment cost is prohibitive, and the knowledge is converted into static feature representations, which cannot solve the problem of temporal knowledge updating. Similarly, the second approach involves intrusively injecting knowledge into the knowledge graph through additional parameter fine-tuning, which is also the main KG augmentation method currently relied upon by academia. This method reduces training costs, but such intrusive architectural modifications disrupt the original attention pattern of the LLM, impairing the model's inherent knowledge and capabilities, and resulting in reduced language generation fluency. In addition, graph-based retrieval-augmented generation (GRAGE) has also been widely used in this area. Due to its lower cost and greater flexibility, it has become a new paradigm for knowledge injection methods. However, issues such as graph retrieval efficiency, semantic misalignment between structured data and natural language text, and the retrieval quality and computational overhead of traditional RAGs remain.

[0004] Therefore, the existing technical bottlenecks can be roughly summarized as follows:

[0005] Pre-training injection: Converting KG triples into text for pre-training (such as ERNIE) results in static knowledge solidification, inability to dynamically update, and extremely high training costs.

[0006] Fine-tuning injection: Fine-tuning the model by adding an adapter module (such as K-Adapter) and modifying the model parameters to destroy the original attention distribution and impair language fluency.

[0007] Retrieval-augmented generation (RAG): concatenating KG triples into prompt input (such as KG-GPT) faces problems such as semantic alignment deviation, high retrieval latency, and insufficient support for multi-hop reasoning.

[0008] Therefore, the core problems solved by the present invention are as follows:

[0009] Dynamicity: Retrieve KG in real time during testing to avoid knowledge obsolescence.

[0010] Non-invasive: LLM parameters are not modified, preserving its native capabilities.

[0011] Deep Fusion: Realizes two-way interaction between KG and text at the attention layer, surpassing the shallow splicing of traditional RAG. Summary of the Invention

[0012] The purpose of the present invention is to address the deficiencies in the above-mentioned background technology and to provide a method for enhancing a large language model with a knowledge graph during testing, which can effectively utilize knowledge graph knowledge to enhance the capabilities of a large language model.

[0013] The technical solution adopted by the present invention is: a method for enhancing a large language model with a knowledge graph during testing, comprising the following steps:

[0014] Extract triples from the knowledge graph according to the text input sequence as the knowledge graph triple input sequence;

[0015] Convert text input sequences and knowledge graph triple input sequences into embedding vectors and calculate query matrix, key matrix and value matrix;

[0016] Computing mixed attention;

[0017] Get richer feature representation through feedforward neural networks;

[0018] Based on the above feature representation, the model is generated to predict the probability distribution and thus predict the next word. In particular, GraphRAG technology is used to retrieve triples from the knowledge graph structured data according to the text input sequence as the knowledge graph triple input sequence.

[0019] The text input sequence and the knowledge graph triple input sequence are converted into embedding vectors. For a given network hidden layer, assuming that its text input embedding matrix is X and the knowledge graph triple embedding matrix is z, according to the same parameter matrix W Q , W K , W V, then the query matrix, key matrix, and value matrix of the input text and knowledge graph are calculated as follows:

[0020]

[0021] Among them, Q X , K X and V X Denote the query matrix, key matrix and value matrix generated by the input embedding matrix, Q Z , K Z and V Z It represents the query matrix, key matrix and value matrix generated by the triple embedding matrix.

[0022] Among them, a multi-head attention mechanism is adopted, and the mixed attention is calculated by the method of bidirectional information aggregation.

[0023] Split the transformed query matrix, key matrix, and value matrix into multiple heads respectively. Assume there are h heads and the dimension of each head is d k , then:

[0024]

[0025] Among them, d dim is the embedding dimension of the model, so the query matrix, key matrix and value matrix after segmentation are as follows, where i represents the i-th attention head:

[0026] Q i =split(Q,i);K i =split(K,i);V i =split(V,i)

[0027] Among them, Q i , K i and V i denote the query matrix, key matrix, and value matrix of the i-th attention head after segmentation, while Q, K, and V denote the original query matrix, key matrix, and value matrix, respectively.

[0028] Among them, the hybrid attention feature is combined with the feedforward neural network to obtain a richer feature representation. The calculation process is as follows:

[0029] FFN(X)=ReLU(X T W1)W2

[0030] Among them, FFN is the abbreviation of Feed-Forward Network, which means feedforward neural network, X is the input feature matrix, ReLU is an activation function, defined as Max(0,X), which means taking the maximum value between the input and 0, X TRepresents the transpose of the input matrix, and W1 and W2 are two linear transformation matrices in the feedforward neural network.

[0031] Among them, for any attention head of the multi-head attention mechanism, hybrid attention aggregates three information flows: text input information flow based on self-attention, output information flow enhanced by knowledge graph, and knowledge graph information flow;

[0032] The three information flow output matrices are as follows:

[0033] The output matrix of the self-attention-based text input information flow is composed of the text input query matrix and text input key matrix The self-attention weight matrix and text input value matrix Multiplying together gives:

[0034]

[0035] The output information flow of knowledge graph enhancement is output by the query matrix based on the knowledge graph triple input and text input key matrix The attention weight matrix and text input value matrix Multiplying together gives:

[0036]

[0037] The knowledge graph information flow output matrix is composed of the text input query matrix and the knowledge graph triple input key matrix The attention weight matrix and the knowledge graph triple input value matrix Multiplying together gives:

[0038]

[0039] Among them, GraphRAG is used as a tool to retrieve specific relevant knowledge graph triples from the knowledge graph.

[0040] Among them, the text input information flow and the output information flow enhanced by the knowledge graph in the hybrid attention introduce nonlinear complex changes through the feedforward neural network to obtain a new feature representation. The knowledge graph information flow skips the feedforward neural network layer and is directly aggregated with the changed input information flow. The aggregated feature representation is then fused. The improved fusion process is as follows:

[0041]

[0042] Among them, x n It represents the representation vector of the nth word after integrating the knowledge graph triples. Exp refers to the exponential function. represents the inner product of the nth input word query feature and the ith input word key feature, represents the inner product of the nth input word query feature and the jth triple word key feature, and They represent the value feature vectors of the i-th input word and the j-th triple word respectively.

[0043] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, a method for enhancing a large language model with a knowledge graph during testing is implemented.

[0044] A computer-readable storage medium stores computer instructions, which, when executed by a processor, implement a method for enhancing a large language model with a knowledge graph during testing.

[0045] Furthermore, the present invention makes full use of GraphRAG technology to retrieve triples from knowledge graph structured data according to the text input sequence as the knowledge graph triple input sequence. This non-invasive knowledge injection method can solve the problem of temporal knowledge updating while retaining the original expression ability of the model.

[0046] Furthermore, the present invention does not change the parameter matrix W Q , W K , W V , using them to directly calculate the query matrix, key matrix and value matrix, eliminating the cost of training or fine-tuning, and compared to fine-tuning, it does not destroy the inherent knowledge and capabilities of the model.

[0047] Furthermore, the present invention adopts a multi-head attention mechanism to mix attention. Compared with the traditional single-head attention, the complexity of the multi-head attention mechanism does not increase, but it greatly improves the model's ability to model complex relationships.

[0048] Furthermore, the present invention achieves bidirectional mutual checking between the input stream and the knowledge graph stream through a bidirectional information aggregation method and hybrid attention, ensuring the coherence of adaptive knowledge retrieval and context.

[0049] Furthermore, the present invention rationally combines hybrid attention features with feedforward neural network layers to obtain new feature representations, making the model output more reasonable.

[0050] Furthermore, the hybrid attention proposed in this paper aggregates three information flows: a self-attention-based text input information flow, a knowledge graph-enhanced output information flow, and a knowledge graph information flow. In addition to the traditional self-attention-based text input information flow, the knowledge graph-enhanced output information flow in hybrid attention leverages the structured semantics of the knowledge graph to locate specific fields in the text input and increase the attention weights for these fields. The knowledge graph information flow, on the other hand, is a query of the knowledge graph information by the text input. Combining these three, a hybrid attention mechanism for bidirectional information aggregation is achieved.

[0051] Furthermore, the hybrid attention model proposed in this paper introduces nonlinear and complex changes to the text input information flow and the knowledge graph-enhanced output information flow through a feedforward neural network, resulting in a new feature representation. The knowledge graph information flow skips the feedforward neural network layer and is directly aggregated with the remaining two information flows after the changes. The aggregated feature representation is then normalized, making the model's output more reasonable for complex problems.

[0052] Compared with the prior art, the advantages of the present invention are as follows:

[0053] 1. The knowledge graph enhancement method proposed in this invention is a test-time enhancement method that does not change any parameters of the model and therefore does not face the catastrophic forgetting problem brought about by traditional fine-tuning methods.

[0054] 2. The knowledge graph enhancement method proposed in the present invention has very strong adaptability and can very effectively adapt to the dynamic changes of knowledge graphs.

[0055] 3. Because it does not involve parameter updates and has strong adaptability, the method proposed in the present invention can be very effectively and quickly integrated into the current large model and has high practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 Schematic diagram of the attention module guided by the knowledge graph of the present invention;

[0057] Figure 2 Schematic diagram of the overall structure of the knowledge graph-guided attention module. DETAILED DESCRIPTION

[0058] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments to facilitate a clear understanding of the present invention, but they do not constitute a limitation to the present invention.

[0059] Example 1

[0060] like Figure 2 As shown, the present invention provides a method for enhancing a large language model with a knowledge graph during testing, comprising the following steps:

[0061] Extract triples from the knowledge graph according to the text input sequence as the knowledge graph triple input sequence;

[0062] Convert text input sequences and knowledge graph triple input sequences into embedding vectors and calculate query matrix, key matrix and value matrix;

[0063] Computing mixed attention;

[0064] Get richer feature representation through feedforward neural networks;

[0065] Based on the above feature representation, the generative model predicts the probability distribution and thus predicts the next word.

[0066] Specifically, GraphRAG technology is used to retrieve triples from knowledge graph structured data according to the text input sequence as the knowledge graph triple input sequence.

[0067] Figure 1 The core module of the present invention, the knowledge graph guided attention module, is described. As shown in the figure, this module contains three information flows. The first information flow is the information flow in the original self-attention module, which is the information flow of the model aggregating the input value matrix V through self-attention. X The second information flow is the knowledge graph fusion information flow, which is the query feature matrix Q X To reaggregate the triple value matrix V Z The third information flow is to use the knowledge graph query matrix Q Z To re-aggregate the input value matrix V X , thereby completing the re-aggregation of input features, removing noise and retaining important information.

[0068] Figure 2 This describes the overall structure of our proposed knowledge graph-guided attention module. Its fusion process consists of two steps. The first step uses knowledge graph attention (KGA) to aggregate and normalize the knowledge graph information. The normalized features are then fed into a feedforward neural network (FFN) for further feature enrichment. Finally, the results are normalized and output.

[0069] Specifically, the text input sequence and the knowledge graph triple input sequence are converted into embedding vectors. Assuming that the text input embedding vector is X and the knowledge graph triple embedding vector is Z, according to the same parameter matrix W Q , W K , W V , then the query matrix, key matrix and value matrix are calculated as follows:

[0070]

[0071] Specifically, the present invention adopts a multi-head attention mechanism and calculates hybrid attention through a bidirectional information aggregation method. The transformed query matrix, key matrix and value matrix are divided into multiple heads respectively. Assume that we have h heads, and the dimension of each head is d k , then:

[0072]

[0073] Among them, d dim is the embedding dimension of the model. Therefore, the query matrix, key matrix, and value matrix after segmentation are as follows, where i represents the i-th attention head:

[0074] Q i =split(Q,i);K i =split(K,i);V i =split(V,i)

[0075] Specifically, for any attention head of the multi-head attention mechanism, hybrid attention aggregates three information flows: text input information flow based on self-attention, output information flow enhanced by knowledge graph, and knowledge graph information flow.

[0076] Specifically, the three information flow output matrices are calculated as follows:

[0077] The output matrix of the self-attention-based text input information flow is composed of the text input query matrix and text input key matrix The self-attention weight matrix and text input value matrix Multiplying together gives:

[0078]

[0079] The output information flow of knowledge graph enhancement is output by the query matrix based on the knowledge graph triple input and text input key matrix The attention weight matrix and text input value matrix Multiplying together gives:

[0080]

[0081] The knowledge graph information flow output matrix is composed of the text input query matrix and the knowledge graph triple input key matrix The attention weight matrix and the knowledge graph triple input value matrix Multiplying together gives:

[0082]

[0083] Specifically, a new feature representation is derived by combining hybrid attention features with a feedforward neural network. The text input information stream and the knowledge graph augmented output information stream undergo nonlinear and complex transformations through a feedforward neural network, resulting in a new feature representation. The knowledge graph information stream skips one layer of the feedforward neural network and is directly aggregated with the remaining two transformed information streams. The aggregated feature representation is then normalized.

[0084] The contents not described in detail in this specification belong to the prior art known to those skilled in the art.

[0085] It should be noted that the above embodiments are not intended to limit the scope of protection of the present invention, and equivalent changes or substitutions made on the basis of the above technical solutions fall within the scope of protection of the claims of the present invention.

Claims

1. A method for enhancing a large language model with a knowledge graph during testing, characterized by: The following steps are involved: Retrieve triples from the knowledge graph as knowledge graph input information based on the text input sequence; Convert text input sequences and knowledge graph triple input sequences into embedding vectors and calculate query matrix, key matrix and value matrix; Computing mixed attention; Get richer feature representation through feedforward neural networks; Based on the above feature representation, the generative model predicts the probability distribution and thus predicts the next word.

2. The method for enhancing a large language model with a knowledge graph during testing according to claim 1, characterized in that: GraphRAG technology is used to retrieve triples from knowledge graph structured data according to the text input sequence as the knowledge graph triple input sequence.

3. The method for enhancing a large language model with a knowledge graph during testing according to claim 1, characterized in that: The text input sequence and the knowledge graph triple input sequence are converted into embedding vectors. For a given network hidden layer, assuming that its text input embedding matrix is X and the knowledge graph triple embedding matrix is Z, according to the same parameter matrix W Q , W K , W V , then the query matrix, key matrix, and value matrix of the input text and knowledge graph are calculated as follows: Among them, Q X , K X and V X Denote the query matrix, key matrix and value matrix generated by the input embedding matrix, Q Z , K Z and V Z It represents the query matrix, key matrix and value matrix generated by the triple embedding matrix.

4. The method for enhancing a large language model with a knowledge graph during testing according to claim 1, characterized in that: Adopting a multi-head attention mechanism and calculating mixed attention through a bidirectional information aggregation method, Split the transformed query matrix, key matrix, and value matrix into multiple heads respectively. Assume there are h heads and the dimension of each head is d k , then: Among them, d dim is the embedding dimension of the model, so the query matrix, key matrix and value matrix after segmentation are as follows, where i represents the i-th attention head: Q i =split(Q,i);K i =split(K,i);V i =split(V,i) Among them, Q i , K i and V i denote the query matrix, key matrix, and value matrix of the i-th attention head after segmentation, while Q, X, and V denote the original query matrix, key matrix, and value matrix, respectively.

5. The method for enhancing a large language model with a knowledge graph during testing according to claim 1, characterized in that: Combined with the mixed attention features, a feedforward neural network is used to obtain a richer feature representation. The calculation process is as follows: FFN(X)=ReLU(X T W1) W2 Among them, FFN is the abbreviation of Feed-Forward Network, which means feedforward neural network, X is the input feature matrix, ReLU is an activation function, defined as Max(0,X), which means taking the maximum value between the input and 0, X T Represents the transpose of the input matrix, and W1 and W2 are two linear transformation matrices in the feedforward neural network.

6. The method for enhancing a large language model with a knowledge graph during testing according to claim 4, characterized in that: For any attention head of the multi-head attention mechanism, hybrid attention aggregates three information flows: the text input information flow based on self-attention, the output information flow enhanced by the knowledge graph, and the knowledge graph information flow; The three information flow output matrices are as follows: The output matrix of the self-attention-based text input information flow is composed of the text input query matrix and text input key matrix The self-attention weight matrix and text input value matrix Multiplying together gives: The output information flow of knowledge graph enhancement is output by the query matrix based on the knowledge graph triple input and text input key matrix The attention weight matrix and text input value matrix Multiplying together gives: The knowledge graph information flow output matrix is composed of the text input query matrix and the knowledge graph triple input key matrix The attention weight matrix and the knowledge graph triple input value matrix Multiplying together gives:

7. The method for enhancing a large language model with a knowledge graph during testing according to claim 2, characterized in that: GraphRAG is used as a tool to retrieve specific relevant knowledge graph triples from the knowledge graph.

8. The method for enhancing a large language model with a knowledge graph during testing according to claim 5, characterized in that: In hybrid attention, the text input information flow and the output information flow enhanced by the knowledge graph are introduced into nonlinear and complex changes through the feedforward neural network to obtain a new feature representation. The knowledge graph information flow skips the feedforward neural network layer and is directly aggregated with the changed input information flow. The aggregated feature representation is then fused. The improved fusion process is as follows: Among them, x n It represents the representation vector of the nth word after integrating the knowledge graph triples. Exp refers to the exponential function. represents the inner product of the nth input word query feature and the ith input word key feature, represents the inner product of the nth input word query feature and the jth triple word key feature, and They represent the value feature vectors of the i-th input word and the j-th triple word respectively.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements a method for enhancing a large language model with a knowledge graph during testing as described in any one of claims 1 to 8.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, a method for enhancing a large language model with a knowledge graph during testing as described in any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Fine adjustment method, system and equipment based on large language model and medium

    CN117290480A

  • Natural language understanding algorithm based on semantic matching and knowledge graph

    CN117350378A

  • Intelligent task sequence planning method based on language vision large model and knowledge graph

    CN117874258A

  • Knowledge graph representation learning method of fusion graph structure based on pre-training language model

    CN118210927A

  • Large language model question and answer generation method based on knowledge graph enhancement

    CN118227769A