Knowledge graph question and answer method based on double-cascade graph attention framework

By introducing a dual-cascade graph attention framework and a cyclic self-attention mechanism in knowledge graph questions and answers, the problem of insufficient understanding of problems and relationship semantics in multi-hop reasoning is solved, and more efficient inference performance and finer-grained semantic capture are achieved.

CN120031127AActive Publication Date: 2025-05-23CHONGQING UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510037284.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-23
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The prior art lacks understanding of problems and relationship semantics in multi-hop inference, resulting in low inference performance and lacks comprehensive attention to problem semantic information.

Method used

A knowledge graph question-and-answer method based on the dual-cascade graph attention framework is proposed. The problem is preprocessed and encoded through the initialization layer, the semantic contribution of each word is calculated, and the initial problem vector is fused into the initial problem vector, combining the cyclic self-attention mechanism and vector graph attention network to enhance the representation ability of problems and relationships.

Benefits of technology

It effectively solves the problem of insufficient understanding of problems and relationship semantics in multi-hop reasoning, improves inference performance, can capture the semantic information of the problem more granularly, and reduces interference from non-critical paths in the inference process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031127A_ABST
    Figure CN120031127A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph question and answer method based on a double-cascade graph attention framework, which comprises the following steps of: S1, preprocessing and coding a question through an initialization layer to obtain a word-level question vector and a sentence vector of the question; s2, calculating the semantic contribution degree of each word, and fusing the semantic contribution degree into an initial problem vector; and S3, calculating a relation score, and calculating a score of the tail entity by using the head entity and the relation score. The problem of insufficient understanding of problems and relation semantics in multi-hop reasoning can be effectively solved, and reasoning performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of language processing technology, and in particular to a knowledge graph question answering method based on a dual-cascade graph attention framework. Background Art

[0002] Knowledge graph question answering (KBQA) is a key upstream task for applications such as intelligent search, personalized recommendation, and human-computer dialogue systems. Its goal is to accurately locate answers from the knowledge base by deeply analyzing the problem and combining the structure and content of the knowledge graph. In the KBQA scenario of multi-hop questions, since the grammatical structure of the question is usually more complex and involves more entities and relations, question answering often requires multi-hop reasoning, which requires not only a strong semantic understanding ability of the model, but also the ability to handle relational reasoning across multiple entities. Its research has attracted widespread attention.

[0003] The challenges faced by the KBQA task include information loss during encoding and compression, lack of fine-grained problem analysis, and weak interpretability of the reasoning process. Research in recent years has mainly included methods based on semantic parsing, embedding, and path-based methods. For example, Lan proposed a method based on generating query graphs, but it faces the problems of huge query space and insufficient query graph effectiveness. The EmbedKGQA method proposed by Saxena et al. is a typical embedding-based method. Its performance is limited by the embedding quality of the knowledge base and is highly dependent on the embedding of entities and relations. The path-based method proposed by Shi et al. can better improve the interpretability of multi-hop reasoning, but this method lacks comprehensive attention to the semantic information of the problem. Overall, the path-based method can capture complex relationship paths and efficiently handle multi-hop reasoning problems, becoming a mainstream method that has attracted much attention.

[0004] Although existing research has made positive progress, it still lacks attention to the interaction of word and sentence scale information in complex problems, as well as information transmission at different reasoning time steps. Therefore, it is not possible to fully understand the semantics of the problem, especially not conducive to reasoning of multi-hop problems. In addition, it is very important to capture the complex relationship between the question and the answer during the reasoning process. Summary of the invention

[0005] The present invention aims to solve the technical problems existing in the prior art, and particularly innovatively proposes a knowledge graph question answering method based on a dual cascade graph attention framework, which can effectively solve the problem of insufficient semantic understanding of questions and relationships in multi-hop reasoning and improve the reasoning performance.

[0006] In order to achieve the above object, the present invention provides a knowledge graph question answering method based on a dual cascade graph attention framework, comprising the following steps:

[0007] S1: Preprocess and encode the question through the initialization layer to obtain the word-level question vector and sentence vector of the question;

[0008] S2: Calculate the semantic contribution of each word and merge it into the initial question vector;

[0009] S3: Calculate the relationship score and use the head entity and relationship score to calculate the score of the tail entity.

[0010] In the above scheme: step S1 also includes the following steps:

[0011] S1-1: Pre-train the input question through BERT to obtain the word-level question vector;

[0012] S1-2: The word-level question vector Q obtained after BERT pre-training 0 , and the word vector Q at the current step size t (t∈[0,T]) is fused to obtain the self-attention matrix M s ; The formula is as follows:

[0013] M s =MLP(Q t )×Q 0

[0014] Among them, M s is the self-attention matrix, MLP(·) represents the linear fully connected layer, t is the number of steps in the current reasoning stage, and T is the total number of reasoning steps;

[0015] S1-3: The self-attention matrix M s Through Softmax conversion, we can get the key information distribution M of the problem under the current number of hops. s' , the formula is as follows:

[0016] M s' =Softmax(M s )

[0017] Among them, M s' represents the key information distribution of the problem at the current reasoning stage, and Softmax(·) is the activation function;

[0018] S1-4: M s' Embedded in Q t In the above example, we get the inference representation information R t , the formula is as follows:

[0019] R t =M s' ×Q t .

[0020] In the above scheme: step S2 also includes the following contents:

[0021] S2-1: Perform word-level reasoning on the input question;

[0022] S2-2: Calculate the weight of the cyclic self-attention;

[0023] S2-3: Fusion sentence vector reasoning;

[0024] S2-4: The semantic information q represented by fine-grained information t Processed by VGAT.

[0025] In the above scheme: step S2-1 also includes the following steps:

[0026] When the first reasoning is performed, execute S2-1-1, otherwise execute S2-1-2;

[0027] S2-1-1: The word-level question vector Q obtained by BERT pre-training is processed through the word inference layer 0 and the reasoning representation information R of the first reasoning stage 0 Processing and transformation, the formula is as follows:

[0028] q 0 =Q 0 +R 0

[0029] Among them, q 0 is the semantic information for initialization;

[0030] S2-1-2: store the reasoning representation information of the current reasoning stage;

[0031] The formula is as follows:

[0032] U 0 =R 0

[0033] Among them, U 0 is the initial storage unit; return to step S2-1;

[0034] S2-1-3: store the reasoning representation information of the current reasoning stage;

[0035] The formula is as follows:

[0036] U t =Z t ×U t-1 +R t

[0037] Among them, U t is the storage unit of the current reasoning stage, U t-1 is the storage unit of the last inference stage, Z trepresents the key information extracted in the previous reasoning stage, and Z t Calculate using the following formula:

[0038] Z t =Sigmoid(MLP(U t-1 ))

[0039] Among them, MLP(·) represents the linear fully connected layer, Sigmoid(·) is the activation function;

[0040] S2-1-4: Calculate the semantic information q of the current reasoning stage t ;

[0041] q t =q 0 +Z t ×U t-1

[0042] Among them, q 0 is the initialized semantic information, U t-1 It is the storage unit of the previous reasoning stage;

[0043] When t>0, the information storage unit needs to be integrated into the t-1 step through Z t Filtered information to obtain a more fine-grained dynamic problem representation q t , and incorporate this representation into the initialization semantic information q 0 middle.

[0044] In the above scheme: step S2-2 also includes the following contents:

[0045] S2-2-1: Calculate the initial attention weight;

[0046] S2-2-1-1: Distribution of key information on the problem M s' Perform one-dimensional summation and normalization to obtain the attention score;

[0047] S2-2-1-2: Multiply the attention score by the attention mask att_mask to ensure that the model only pays attention to the attention scores of valid words;

[0048] S2-2-1-3: Perform two-dimensional summation on the attention scores again and normalize them to obtain the final initialized attention weights;

[0049] That is, it is calculated by the following formula:

[0050]

[0051] in, To initialize the attention weight, ∑ col To sum by column, ∑row is the row-wise summation, and att_mask is the attention mask.

[0052] S2-2-2: Calculate the final attention weight

[0053] In each cycle, the final weight obtained from the previous calculation is used as the initial value for the next calculation and is calculated using the following formula:

[0054]

[0055] Among them, the number of cycles n is a hyperparameter;

[0056] S2-2-3: Calculate the problem vector of the prediction relationship;

[0057] The formula is as follows:

[0058]

[0059] in, is the problem vector for predicting the relation.

[0060] In the above scheme: step S2-3 also includes the following contents:

[0061] S2-3-1: Calculate the initial relationship prediction vector;

[0062] The sentence vector q obtained by BERT pre-training s Vector with question Add them together to get the initial relationship prediction vector r 0 , the formula is as follows:

[0063]

[0064] q s is the sentence vector of the question,

[0065] S2-3-2: storing the problem vector of the prediction relationship of the current reasoning stage;

[0066] The formula is as follows:

[0067]

[0068] in, is the initial storage unit, represents the relation representation selected in the previous reasoning phase, and Calculate using the following formula:

[0069]

[0070] Among them, MLP(·) represents the linear fully connected layer, Sigmoid(·) is the activation function;

[0071] S2-3-3: Calculate the final relationship prediction vector r t ;

[0072] The formula is as follows:

[0073]

[0074] Among them, r t is the final relationship prediction vector;.

[0075] In the above scheme: step S2-4 also includes the following steps:

[0076] S2-4-1: Each node in the vector graph is recorded as:

[0077] N={token 1 ,token 2 ,token 3 ,...,token i}, where i∈[1,N], token i represents the feature representation of the i-th node;

[0078] S2-4-2: Construct the self-correlation feature adjacency matrix e-adj;

[0079] S2-4-3: Convert using the following formula:

[0080] W k =e-adj

[0081] W k is the weight matrix of the corresponding input linear transformation;

[0082] S2-4-4: Calculate the attention of all nodes’ neighboring nodes;

[0083] Calculate using the following formula:

[0084]

[0085] Among them, α ij is the attention between vertices i and j, a T is the transposed matrix of the 2F'-dimensional vector; LeakyReLU is the activation function, exp is the exponential function, W is the trainable parameter, Ni is the set of adjacent nodes of the i-th node; || represents connection;

[0086] S2-4-5: Calculate the final representation of the i-th node through adjacent nodes;

[0087] Calculate using the following formula:

[0088]

[0089] Where K represents the number of attention mechanisms; is the attention between vertices i and j, which is generated by the kth attention mechanism (α k ) calculated by the normalized attention coefficient; W k is the weight matrix of the corresponding input linear transformation; σ represents the coefficient;

[0090] S2-4-6: Aggregation through VGAT layer;

[0091] Calculate using the following formula:

[0092]

[0093] Among them, sum_token means sum;

[0094] S2-4-7: Problem vector of the relationship between the features aggregated by the VGAT layer and the prediction Fusion

[0095] The formula is as follows:

[0096]

[0097] r v Represents the relationship vector after fusion by the VGAT layer;

[0098] S2-4-8: The final relationship prediction vector r is obtained by jointly representing the dual cascade framework and the vector graph attention layer. final ;

[0099] The formula is as follows:

[0100] The formula is as follows:

[0101] r final =r v +r t

[0102] Among them, r v represents the relationship vector after fusion by the VGAT layer, r final Represents the final relationship prediction vector.

[0103] In the above scheme: step S3 also includes the following steps:

[0104] S3-1: Perform multi-label classification on the relationship of the knowledge graph to obtain the answer entity vector;

[0105] Calculate using the following formula:

[0106] e t = f(e t-1 , Sigmoid(MLP(r final )))

[0107] where e t is the answer entity vector, and f(·) represents calculating the scores of all entities at the current inference stage;

[0108] S3-2: Predict the final answer;

[0109] It is calculated through the following formula:

[0110]

[0111] where Y is the predicted final answer, and e t is the answer entity vector; r final is the relation prediction vector;

[0112] S3-3: Train the final answer;

[0113] It is trained through the following formula:

[0114] L = α(Y - y) 2 + βD kl (Y, y)

[0115] where y is the true value, Y is the predicted final answer, L is the loss function, and D kl (·) is the kl divergence, and α and β are adjustable parameters.

[0116] In summary, the beneficial effects of the present invention are as follows: A double-cascade inference framework is proposed, which fuses word-level information to enhance the problem representation, embeds sentence information to enhance the relation representation, and uses recurrent self-attention to connect the two levels, which can effectively solve the problem of insufficient semantic understanding of problems and relations in multi-hop reasoning and improve the performance of reasoning. Between each inference time step, the recurrent self-attention mechanism can also be used to calculate the attention weights multiple times and continuously update the problem representation, so as to capture the semantic information of the problem in a more fine-grained manner. In addition, a vector graph attention network for relation enhancement is designed, which effectively alleviates the interference problem of non-critical paths in the reasoning process. BRIEF DESCRIPTION OF THE DRAWINGS

[0117] Figure 1 is the structural schematic diagram of the present invention.

[0118] Figure 2 is the structural schematic diagram of the VGAT layer.

[0119] Figure 3This is a statistical chart of the rise and fall of the n-th reasoning relationship score of the n-th reasoning problem after adding the VGAT layer.

[0120] Figure 4 It is a flowchart of the vector graph attention network.

[0121] Figure 5 It is a bar chart of the amount of relation score improvement on the MetaQA dataset.

[0122] Figure 6 It is a histogram of the amount of relation score reduction on the MetaQA dataset. DETAILED DESCRIPTION

[0123] The present invention will be further described below by way of embodiments and in conjunction with the accompanying drawings:

[0124] like Figure 1 to Figure 6 As shown, a knowledge graph question answering method based on a dual-cascade graph attention framework includes the following steps:

[0125] S1: Preprocess and encode the question through the initialization layer to obtain the word-level question vector and sentence vector of the question;

[0126] S1-1: Pre-train the input question through BERT to obtain the word-level question vector;

[0127] S1-2: The word-level question vector Q obtained after BERT pre-training 0 , and the word vector Q at the current step size t (t∈[0,T]) is fused to obtain the self-attention matrix M s ; The formula is as follows:

[0128] M s =MLP(Q t )×Q 0

[0129] Among them, M s is the self-attention matrix, MLP(·) represents the linear fully connected layer, t is the number of steps in the current reasoning stage, and T is the total number of reasoning steps;

[0130] S1-3: The self-attention matrix M s Through Softmax conversion, we can get the key information distribution M of the problem under the current number of hops. s' , the formula is as follows:

[0131] M s' =Softmax(M s )

[0132] Among them, M s'represents the key information distribution of the problem at the current reasoning stage, and Softmax(·) is the activation function;

[0133] S1-4: M s' Embedded in Q t In the above example, we get the inference representation information R t , the formula is as follows:

[0134] R t =M s' ×Q t

[0135] S2: Calculate the semantic contribution of each word and merge it into the initial question vector;

[0136] S2-1: Perform word-level reasoning on the input question; when performing the first reasoning, execute S2-1-1, otherwise execute S2-1-2;

[0137] S2-1-1: The word-level question vector Q obtained by BERT pre-training is processed through the word inference layer 0 and the reasoning representation information R of the first reasoning stage 0 Processing and transformation, the formula is as follows:

[0138] q 0 =Q 0 +R 0

[0139] Among them, q 0 is the semantic information for initialization;

[0140] S2-1-2: store the reasoning representation information of the current reasoning stage;

[0141] The formula is as follows:

[0142] U 0 =R 0

[0143] Among them, U 0 is the initial storage unit; return to step S2-1;

[0144] S2-1-3: store the reasoning representation information of the current reasoning stage;

[0145] The formula is as follows:

[0146] U t =Z t ×U t-1 +R t

[0147] Among them, U t is the storage unit of the current reasoning stage, U t-1is the storage unit of the last inference stage, Z t represents the key information extracted in the previous reasoning stage, and Z t Calculate using the following formula:

[0148] Z t =Sigmoid(MLP(U t-1 ))

[0149] Among them, MLP(·) represents the linear fully connected layer, Sigmoid(·) is the activation function;

[0150] S2-1-4: Calculate the semantic information q of the current reasoning stage t ;

[0151] q t =q 0 +Z t ×U t-1

[0152] Among them, q 0 is the initialized semantic information, U t-1 It is the storage unit of the previous reasoning stage;

[0153] When t>0, the information storage unit needs to be integrated into the t-1 step through Z t Filtered information to obtain a more fine-grained dynamic problem representation q t , and incorporate this representation into the initialization semantic information q 0 middle.

[0154] S2-2: Calculate the weight of the cyclic self-attention;

[0155] S2-2-1: Calculate the initial attention weight;

[0156] S2-2-1-1: Distribution of key information on the problem M s' Perform one-dimensional summation and normalization to obtain the attention score;

[0157] S2-2-1-2: Multiply the attention score by the attention mask att_mask to ensure that the model only pays attention to the attention scores of valid words;

[0158] S2-2-1-3: Perform two-dimensional summation on the attention scores again and normalize them to obtain the final initialized attention weights;

[0159] That is, it is calculated by the following formula:

[0160]

[0161] in, To initialize the attention weight, ∑ col To sum by column, Σ row is the sum by row, att_mask is the attention mask;

[0162] In this way, the semantic information in the question weight matrix can be accurately captured and applied to reasoning.

[0163] S2-2-2: Calculate the final attention weight

[0164] In each cycle, the final weight obtained from the previous calculation is used as the initial value for the next calculation and is calculated using the following formula:

[0165]

[0166] The number of loops n is a hyperparameter, which depends on different datasets. In this embodiment, n is 3 on the WebQSP and MetaQA datasets;

[0167] S2-2-3: Calculate the problem vector of the prediction relationship; weight Incorporating into the representation of dynamic problems t In the above example, we obtain the problem vector of the prediction relationship, which improves the model’s expressiveness and reasoning accuracy.

[0168] The formula is as follows:

[0169]

[0170] in, is the problem vector for predicting the relation;

[0171] S2-3: Fusion sentence vector reasoning;

[0172] S2-3-1: Calculate the initial relationship prediction vector;

[0173] The sentence vector q obtained by BERT pre-training s Vector with question Add them together to get the initial relationship prediction vector r 0 , the formula is as follows:

[0174]

[0175] q s is the sentence vector of the question,

[0176] S2-3-2: storing the problem vector of the prediction relationship of the current reasoning stage;

[0177] The formula is as follows:

[0178]

[0179] in, is the initial storage unit, represents the relation representation selected in the previous reasoning phase, and Calculate using the following formula:

[0180]

[0181] Among them, MLP(·) represents the linear fully connected layer, Sigmoid(·) is the activation function;

[0182] S2-3-3: Calculate the final relationship prediction vector r t ;

[0183] The formula is as follows:

[0184]

[0185] Among them, r t Predict the final relationship vector;

[0186] S2-4: The semantic information q represented by fine-grained information t Processed by VGAT;

[0187] The problem vector q represented by fine-grained information t Each token in is regarded as a virtual node, and the hidden layer output of each token is regarded as the node feature. We use q t The token represents the construction of the self-correlation feature adjacency matrix e-adj, such as Figure 2 As shown in Figure 1. Although graph neural networks usually aggregate the information of neighboring nodes to learn global features, the association matrix ensures that each node can at least obtain its own information. While considering the context information of the node, it is also necessary to retain the feature information of the node itself. t and e-adj are embedded into the GAT layer to obtain the problem representation after VGAT.

[0188] GAT uses the attention mechanism to calculate the weights between each node in the graph and its adjacent nodes. The training process of GAT depends on the relationship between node pairs and is not affected by the specific network structure.

[0189] S2-4-1: Each node in the vector graph is recorded as:

[0190] N={token 1 ,token 2 ,token 3 ,...,token i}, where i∈[1,N], token i represents the feature representation of the i-th node;

[0191] S2-4-2: Construct the self-correlation feature adjacency matrix e-adj;

[0192] S2-4-3: Convert using the following formula:

[0193] W k =e-adj

[0194] W k is the weight matrix of the corresponding input linear transformation; the autocorrelation feature adjacency matrix e-adj is regarded as the weight matrix W k ;

[0195] S2-4-4: Calculate the attention of all nodes’ neighboring nodes;

[0196] Calculate using the following formula:

[0197]

[0198] Among them, α ij is the attention between vertices i and j, a T is the transposed matrix of the 2F'-dimensional vector; LeakyReLU is the activation function, exp is the exponential function, W is the trainable parameter, Ni is the set of adjacent nodes of the i-th node; || represents connection;

[0199] S2-4-5: Calculate the final representation of the i-th node through adjacent nodes;

[0200] Calculate using the following formula:

[0201]

[0202] Where K represents the number of attention mechanisms; is the attention between vertices i and j, which is generated by the kth attention mechanism (α k ) calculated by the normalized attention coefficient; W k is the weight matrix of the corresponding input linear transformation; σ represents the coefficient;

[0203] S2-4-6: Aggregation through VGAT layer;

[0204] Calculate using the following formula:

[0205]

[0206] Among them, sum_token means sum.

[0207] S2-4-7: Problem vector of the relationship between the features aggregated by the VGAT layer and the prediction Fusion

[0208] The formula is as follows:

[0209]

[0210] r v Represents the relationship vector after fusion by the VGAT layer;

[0211] S2-4-8: The final relationship prediction vector r is obtained by jointly representing the dual cascade framework and the vector graph attention layer. final ;

[0212] The formula is as follows:

[0213] r final =r v +r t

[0214] Among them, r v represents the relationship vector after fusion by the VGAT layer, r final represents the final relationship prediction vector;

[0215] S3: Calculate the relationship score and use the head entity and relationship score to calculate the score of the tail entity;

[0216] In order to enhance the expressive power of the current path and improve the accuracy of relationship prediction, this technical solution uses a method of directly fusing the question vector to process the relationship vector. This fusion operation aims to enhance the semantics of the relationship and further improve the expressive power of the current path. At the same time, it enables the model to better understand the relationship between the question and the relationship, thereby improving the accuracy of relationship prediction.

[0217] S3-1: Perform multi-label classification on the relationship of the knowledge graph to obtain the answer entity vector;

[0218] Calculate using the following formula:

[0219] e t =f(e t-1 ,Sigmoid(MLP(r final )))

[0220] Among them, e t is the answer entity vector, f(·) represents the score of all entities in the current reasoning stage;

[0221] After T rounds of reasoning iterations, the relationship prediction vector r is calculated t To predict the final answer entity. Each iteration calculates a relationship prediction vector r final, which contains the predicted relations after reasoning. The relation prediction vector rt is used to predict the answer entity. If the correct answer can be obtained in round t-1 reasoning, the semantics of the question in round t reasoning is considered to be wrong. This means that in each iteration, the model will focus on the part where the correct answer was not found in the previous iteration, and try to find a more accurate answer through the next iteration.

[0222] S3-2: predict the final answer;

[0223] Calculate using the following formula:

[0224]

[0225] Among them, Y is the final answer of the prediction, e t is the answer entity vector; r final Predict vectors for relations;

[0226] S3-3: training the final answer;

[0227] The training is performed using the following formula:

[0228] L=α(Yy) 2 +βD kl (Y,y)

[0229] Among them, y is the true value, Y is the final predicted answer, L is the loss function, and D kl (·) is the KL divergence, which is a statistical measure of the difference between two probability distributions P and Q. α and β are adjustable parameters.

[0230] This technical solution encodes the question with a pre-trained model, and then starts from the topic entity of the question, and jumps along the relationship triple to find the answer entity by calculating the relationship score and entity score of the next hop. This technical solution is different from the existing technology in that it constructs a dual-cascade reasoning framework with cyclic self-attention. The first layer of the framework is word level, which integrates the word-level question vector into the initial reasoning information, and updates the reasoning information and filters the noise at each hop to more accurately calculate the relationship score. The second layer of the framework is sentence level, which embeds the sentence-level question vector into the relationship vector to deepen the model's understanding of the overall semantics of the question. Between each reasoning time step, we propose a cyclic self-attention mechanism to capture fine-grained semantics by learning the word-level question vector multiple times and enhance the subsequent relationship vector representation. At the same time, we also propose a graph attention network based on virtual nodes of the question vector to model the question vector to fully capture the potential connections between distant entities in complex problems.

[0231] In order to better illustrate this technical solution, the following is verified by using WebQuestionsSP as a training set, where WebQuestionsSP (WebQSP) contains 2998 training sets, 100 validation sets, and 1639 test sets. Its questions come from the FreeBase knowledge base, and each question is 1 round of reasoning or 2 rounds of reasoning. The knowledge graph corresponding to this dataset is large in scale, with millions of entities and triples. Following pruned the knowledge graph so that it only contains the mentioned relations and the 2-hop triples of the mentioned entities, and also added reverse relations. The processed knowledge graph contains 1.8 million entities, 1144 relations, and 11.4 million triples. MetaQA is a large multi-hop knowledge graph question-answering dataset. Its knowledge graph comes from the film field. The scale of questions exceeds 400,000. These questions are generated by more than a dozen templates. The questions have a maximum of three hops, containing 43,000 entities, 9 relations, and 135,000 triples.

[0232] The DCGA framework is implemented through the Pytorch deep learning framework. For the WebQSP dataset, BERT is used as the pre-trained model for question encoding. The version is bert-base-uncased. The model is downloaded from the HuggingFace official website. It has a 12-layer Transformer encoder and contains 110M parameters. The hyperparameter T indicates that the model is reasoning within the range of T inferences. The WebQSP dataset sets T = 2. For the MetaQA dataset, BIGRU is used as the encoder of the question and T = 3 is set. The RAdam optimizer is used for training, the learning rate is 0.001, the batch_size of MetaQA is 128, and the batch_size of WebQSP is 16. The evaluation indicators are Hits@1 and F1, and the model is trained on an RTX 3090.

[0233] Figure 4 This is a schematic diagram of the role of the Vector Graph Attention Network (VGAT) in our reasoning process. Since the official website of the Freebase knowledge base has been terminated, we use entity codes to represent entities.

[0234] As shown in Table 1, on MetaQA, TransferNet and GFC achieve 100% Hit@1 in 2 and 3 inferences, while our DCGA achieves 100% in 2 inferences. In 1 inference, DCGA reaches 97.6%, exceeding TransferNet's 97.5% and tying with GFC. In 3 inferences, DCGA reaches 99.9%. We believe that this is because DCGA increases the complexity of the model, while the problems in MetaQA are relatively simple, with fewer constrained entities and similarity relationships. Therefore, it is easy to filter out some key information when filtering noise, resulting in performance fluctuations.

[0235] On the WebQSP dataset, DCGA's Hit@1 is 78.9%, which is better than GFC and RE-KBQA, 17.7% higher than the large model ChatGpt, 6.3% higher than StructGpt, and 0.4% higher than ReasoningLM, reaching the SOTA on this dataset. DCGA's F1 value on the WebQSP dataset is 75.9%, higher than RnG-KBQA's 75.6%. We analyze the reasons through case studies, such as Figure 4 As shown. It can be seen that after the enhancement of VGAT, the relationship score in one-hop reasoning is improved compared with the case without VGAT, which shows that VGAT can effectively capture the relationship between nodes in single-hop reasoning, thereby improving the accuracy of the model. In two-hop reasoning, the advantage of VGAT is more obvious, which reduces the score of r3 in the non-critical path "m.014j017→m.03btywc→m.014y6", while the r3 of the critical path "m.014j017→m.02_96cv→m.014y6" is maintained. This is because the model measures the relationship and entity matching degree involved in the reasoning path with the question through the semantic clues of the question. The score of the reasoning path that is irrelevant or weakly related to the question will be lower than the score of the reasoning path related to the question. This shows that VGAT can better focus on the critical path in reasoning, thereby reducing the impact of noise and improving the accuracy of reasoning.

[0236] Table 1: Hit@1 values ​​on MetaQA, WebQSP, and F1 values ​​on WebQSP compared with 12 baseline models.

[0237]

[0238] We replicate the results of Jiang et al.’s LLMs and other baselines from Xie et al. DCGA significantly outperforms the baselines. Table 2: Comparison of Hit@1 values ​​for 1-pass and 2-pass inference for TransferNet, GFC, and DCGA models on WebQSP.

[0239]

[0240] As shown in Table 2, compared with GFC, DCGA improves by 1.6% in single-reasoning problems and 3.1% in double-reasoning problems. That is, the improvement of single-reasoning problems is slightly lower than that of double-reasoning problems, indicating that the DCGA method is more suitable for multiple-reasoning problems. Analyzing the reason, we believe that single-reasoning problems only consider the relationship between connected nodes, and do not need to capture more distant associations and complex information propagation between nodes, while multiple-reasoning problems require multiple information propagation. Since the dual-cascade reasoning framework of DCGA in this paper can effectively enhance key reasoning information in the reasoning process of entity jumps, it can improve the performance of multi-hop question answering.

[0241] In many real scenarios, the knowledge graph is usually incomplete, which requires the model to have stronger reasoning capabilities. We test the WebQSP dataset on an incomplete knowledge graph (50% KG) and compare the performance of DCGA with other baseline methods, as shown in Tables 3 and 4.

[0242] Table 3: Hit@1 values ​​of different models on the MetaQA dataset with 50% KG.

[0243]

[0244] Table 4: Hit@1 values ​​of different models on the WebQSP dataset with 50% KG.

[0245]

[0246] Since some entities, relations or attribute information are missing in the incomplete knowledge graph, it may be difficult to find matching entities or relations during the reasoning process, causing the embedding-based method to learn inaccurate representations. As can be seen from Table 4, the performance of DCGA is 58.8%, which is better than TransferNet's 52.4%, indicating that the DCGA in this paper still has stable and good reasoning ability in incomplete knowledge graphs.

[0247] To further verify the impact of this technical solution, we conducted ablation experiments on the WebQSP dataset, and the results are shown in the table. After removing the dual cascade framework (w / o dual cascade framework), Hit@1 dropped by 1.2% and F1 dropped by 6.7%, indicating that the embedding of sentence vectors in this paper can enrich the question vectors for relation path prediction. After deleting the vector graph attention network (w / o vector graph attention network), Hit@1 dropped by 1.4% and F1 dropped by 6.8%. This large drop reflects the importance of introducing question semantics to the overall performance of the model, and also shows the effectiveness of the vector graph attention network proposed in this paper. After deleting the cyclic self-attention (w / o deleting the cyclic self-attention), Hit@1 dropped by 1.5% and F1 dropped by 6.0%, indicating that the model's performance in accuracy and reasoning is significantly impaired, especially in fine-grained reasoning ability.

[0248] Table 5: Ablation study on WebQSP.

[0249]

[0250] To further verify the impact of our VGAT on the reasoning process, we conducted a statistical analysis of the relationship scores of each reasoning question on the MetaQA dataset. The relationship score is crucial to the reasoning result. In the MetaQA test set, the ratio of the number of first-order reasoning, second-order reasoning, and third-order reasoning is 1:1.5:1.4. We statistically analyze the relationship scores of questions under different numbers of reasoning, as shown in the following formula:

[0251]

[0252] a1 represents a single reasoning problem, a2 represents a double reasoning problem, a3 represents a triple reasoning problem, Qn represents the number of n reasoning problems, d represents the reasoning method with different numbers of reasoning, D = [1,3], j represents different types of relations, rv represents the relation score when using VGAT, and rb represents the relation score when not using VGAT.

[0253] Therefore, from Figure 5 We can see that in terms of the amount of relationship score improvement, first inference > second inference > third inference, this is because the accuracy of information will gradually decay in each inference. Figure 6 It is known that among the number of relationship scores that have decreased, whether it is a single reasoning or multiple reasoning problem, the relationship whose score has decreased the most in a single reasoning, combined with Figure 6It can be seen that in the single-inference problem, the number of relationship scores for each reasoning is higher than that of the model without VGAT; in the second and third reasoning problems, the number of improved scores for the second and third reasoning is the largest. Since the nth reasoning in the n-inference problem is directly related to the selection of the final entity, the vector graph attention network (VGAT) can handle complex cross-entity long-range dependencies and significantly improve the relationship scores on the reasoning path through the context information passed layer by layer in multiple reasoning, thereby enhancing the accuracy and robustness of the final reasoning.

[0254] like Figure 3 As shown in the figure, taking the three-reasoning problem as an example, the number of relationships with decreased scores in the first reasoning is the largest, and the relationships in the first and second reasoning are weakened and filtered. The first two reasonings are usually used to narrow the scope of candidate entities and guide the model to gradually approach the correct answer, while the third reasoning is the confirmation of the final entity. The score of the last reasoning has a key impact on the final reasoning quality. VGAT also suppresses relationships that are irrelevant to the problem. Therefore, in the three-reasoning problem, the number of relationships with increased scores in the third reasoning is the largest.

Claims

1. A knowledge graph question answering method based on a dual-cascade graph attention framework, characterized by: The following steps are involved: S1: Preprocess and encode the question through the initialization layer to obtain the word-level question vector and sentence vector of the question; S2: Calculate the semantic contribution of each word and merge it into the initial question vector; S3: Calculate the relationship score and use the head entity and relationship score to calculate the score of the tail entity.

2. A knowledge graph question answering method based on a dual cascade graph attention framework according to claim 1, characterized in that: Step S1 also includes the following steps: S1-1: Pre-train the input question through BERT to obtain the word-level question vector; S1-2: Compare the word-level question vector Q0 obtained after BERT pre-training with the word vector Q at the current step size t (t∈[0,T]) is fused to obtain the self-attention matrix M s ; The formula is as follows: M s =MLP(Q t )×Q0 Among them, M s is the self-attention matrix, MLP(·) represents the linear fully connected layer, t is the number of steps in the current reasoning stage, and T is the total number of reasoning steps; S1-3: The self-attention matrix M s Through Softmax conversion, we can get the key information distribution M of the problem under the current number of hops. s' , the formula is as follows: M s' =Softmax(M s ) Among them, M s' represents the key information distribution of the problem at the current reasoning stage, and Softmax(·) is the activation function; S1-4: M s' Embedded in Q t In the above example, we get the inference representation information R t , the formula is as follows: R t =M s' ×Q t 。 3. A knowledge graph question answering method based on a dual cascade graph attention framework according to claim 2, characterized in that: Step S2 also includes the following contents: S2-1: Perform word-level reasoning on the input question; S2-2: Calculate the weight of the cyclic self-attention; S2-3: Fusion sentence vector reasoning; S2-4: The semantic information q represented by fine-grained information t Processed by VGAT.

4. A knowledge graph question answering method based on a dual cascade graph attention framework according to claim 1, characterized in that: Step S2-1 also includes the following steps: When the first reasoning is performed, execute S2-1-1, otherwise execute S2-1-2; S2-1-1: The word-level question vector Q0 obtained after BERT pre-training and the reasoning representation information R0 of the first reasoning stage are processed and transformed through the word reasoning layer. The formula is as follows: q0=Q0+R0 Among them, q0 is the initialized semantic information; S2-1-2: store the reasoning representation information of the current reasoning stage; The formula is as follows: U0=R0 Wherein, U0 is the initial storage unit; return to step S2-1; S2-1-3: store the reasoning representation information of the current reasoning stage; The formula is as follows: U t =Z t ×U t-1 +R t Among them, U t is the storage unit of the current reasoning stage, U t-1 is the storage unit of the last inference stage, Z t represents the key information extracted in the previous reasoning stage, and Z t Calculate using the following formula: Z t =Sigmoid(MLP(U t-1 )) Among them, MLP(·) represents the linear fully connected layer, Sigmoid(·) is the activation function; S2-1-4: Calculate the semantic information q of the current reasoning stage t ; q t =q0+Z t ×U t-1 Among them, q0 is the initialized semantic information, U t-1 It is the storage unit of the previous reasoning stage; When t>0, the information storage unit needs to be integrated into the t-1 step through Z t Filtered information to obtain a more fine-grained dynamic problem representation q t , and incorporate this representation into the initialization semantic information q0.

5. A knowledge graph question answering method based on a dual cascade graph attention framework according to claim 1, characterized in that: Step S2-2 also includes the following contents: S2-2-1: Calculate the initial attention weight; S2-2-1-1: Distribution of key information on the problem M s' Perform one-dimensional summation and normalization to obtain the attention score; S2-2-1-2: Multiply the attention score by the attention mask att_mask to ensure that the model only pays attention to the attention scores of valid words; S2-2-1-3: Perform two-dimensional summation on the attention scores again and normalize them to obtain the final initialized attention weights; That is, it is calculated by the following formula: in, To initialize the attention weight, ∑ col To sum by column, ∑ row is the row-wise summation, and att_mask is the attention mask. S2-2-2: Calculate the final attention weight In each cycle, the final weight obtained from the previous calculation is used as the initial value for the next calculation and is calculated using the following formula: Among them, the number of cycles n is a hyperparameter; S2-2-3: Calculate the problem vector of the prediction relationship; The formula is as follows: in, is the problem vector for predicting the relation.

6. A knowledge graph question answering method based on a dual cascade graph attention framework according to claim 1, characterized in that: Step S2-3 also includes the following: S2-3-1: Calculate the initial relationship prediction vector; The sentence vector q obtained by BERT pre-training s Vector with question Add them together to get the initial relationship prediction vector r0, the formula is as follows: q s is the sentence vector of the question, S2-3-2: storing the problem vector of the prediction relationship of the current reasoning stage; The formula is as follows: in, is the initial storage unit, represents the relation representation selected in the previous reasoning phase, and Calculate using the following formula: Among them, MLP(·) represents the linear fully connected layer, Sigmoid(·) is the activation function; S2-3-3: Calculate the final relationship prediction vector r t ; The formula is as follows: Among them, r t is the final relationship prediction vector;.

7. A knowledge graph question answering method based on a dual cascade graph attention framework according to claim 1, characterized in that: Step S2-4 also includes the following steps: S2-4-1: Each node in the vector graph is recorded as: N={token1,token2,token3,...,token i }, where i∈[1,N], token i represents the feature representation of the i-th node; S2-4-2: Construct the self-correlation feature adjacency matrix e-adj; S2-4-3: Convert using the following formula: W k =e-adj W k is the weight matrix of the corresponding input linear transformation; S2-4-4: Calculate the attention of all nodes’ neighboring nodes; Calculate using the following formula: Among them, α ij is the attention between vertices i and j, a T is the transposed matrix of the 2F'-dimensional vector; LeakyReLU is the activation function, exp is the exponential function, W is the trainable parameter, Ni is the set of adjacent nodes of the i-th node; || represents connection; S2-4-5: Calculate the final representation of the i-th node through adjacent nodes; Calculate using the following formula: Where K represents the number of attention mechanisms; is the attention between vertices i and j, which is generated by the kth attention mechanism (α k ) calculated by the normalized attention coefficient; W k is the weight matrix of the corresponding input linear transformation; σ represents the coefficient; S2-4-6: Aggregation through VGAT layer; Calculate using the following formula: Among them, sum_token means sum; S2-4-7: Problem vector of the relationship between the features aggregated by the VGAT layer and the prediction relationship Fusion The formula is as follows: r v Represents the relationship vector after fusion by the VGAT layer; S2-4-8: The final relationship prediction vector r is obtained by jointly representing the dual cascade framework and the vector graph attention layer. final ; The formula is as follows: r final =r v +r i Among them, r v represents the relationship vector after fusion by the VGAT layer, r final Represents the final relationship prediction vector.

8. A knowledge graph question answering method based on a dual cascade graph attention framework according to claim 1, characterized in that: Step S3 also includes the following steps: S3-1: Perform multi-label classification on the relationship of the knowledge graph to obtain the answer entity vector; Calculate using the following formula: e t =f(e t-1 ,Sigmoid(MLP(r final ))) Among them, e t is the answer entity vector, f(·) represents the score of all entities in the current reasoning stage; S3-2: predict the final answer; Calculate using the following formula: Among them, Y is the final answer of the prediction, e t is the answer entity vector; r final Predict vector for relationship; S3-3: training the final answer; The training is performed using the following formula: L=α(Yy) 2 +βD kl (And,and) Among them, y is the true value, Y is the final predicted answer, L is the loss function, and D kl (·) is the KL divergence, and α and β are adjustable parameters.

Citation Information

Patent Citations

  • Mongolian multi-hop question and answer method based on three-channel cognitive map and map attention network

    CN113779220A

  • Commodity information automatic question and answer method based on e-commerce knowledge graph

    CN116881409A

  • Medical knowledge graph-oriented entity alignment method and related device

    CN119025682A

  • Method and apparatus for automatically generating inference questions and answers

    WO2021184311A1