Question Generation Method and System Based on Knowledge Enhancement and Dual-Graph Interaction
Through a deep learning network model based on knowledge enhancement and dual-graph interaction, the problem of insufficient processing of text structure information in the existing technology is solved, and more accurate and syntax-correct problems are generated, which improves the effect of problem generation.
Patent Information
- Application Number
- CN202211721517.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-12-30
AI Technical Summary
When generating questions related to answers, existing problem generation techniques lack effective processing of text structure information and fail to make full use of external knowledge, resulting in insufficient accuracy of the generated question sentences.
Using a deep learning network model based on knowledge enhancement and dual-graph interaction, we use the training data set and knowledge graph, combined with graph convolutional network and long and short-term memory network to perform representation vector learning of context sentences and answers to generate grammatical correct questions related to answers.
It improves the accuracy of problem generation, and the generated problems are more in line with user needs, with higher syntax accuracy and information richness.
Smart Images

Figure CN116186216B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a question generation method and system based on knowledge enhancement and dual-graph interaction. Background Technique
[0002] Question Generation (QG) aims to create human-like reverse multi-angle and informative questions with knowledge points / facts as answers through a series of inputs such as text, knowledge base, pictures, structured SQL, etc. In the field of open-domain dialogue, it is used as a cold start to initiate new topics to reduce uninformative chitchat, or to obtain feedback by asking questions to better understand users and make the human-computer dialogue more interesting; QG can increase the number of high-quality question-and-answer pairs in the database to improve the recall rate in a question-and-answer system based on text similarity calculation; in a search system, asking questions about the search results can quickly find the required information and improve the retrieval efficiency; in the field of education, intelligently generating high-quality questions according to the teaching content to help students understand details, and generating corresponding questions according to the materials to help teachers prepare lessons intelligently, etc. In summary, the question generation technology has wide applications, and related research work still faces several challenges. Therefore, the research on question generation technology has academic research significance and practical application value.
[0003] With the application of the encoder-decoder neural network model to natural language generation tasks, the research work on question generation tasks has also started to use the encoder-decoder neural network to achieve "sequence-to-sequence" question generation, integrating the two steps of "what to ask" and "how to ask" based on the grammar rule template method into an end-to-end training task. The question generation method based on the "sequence-to-sequence" encoder-decoder model has gradually become the mainstream of question generation technology. Du et al. first applied the "sequence-to-sequence" encoder-decoder neural network model to the question generation task, realizing the generation of questions according to a given sentence or paragraph, but unable to give an answer, only generating questions based on the given sentence or paragraph, lacking answer information, so the applicable range is not wide. Zhou et al. added the answer position to the input of the model to realize the generation of questions according to the given sentence and answer. Duan et al. proposed two semantic enhancement strategies, capturing semantic knowledge at a finer granularity than word vectors by integrating sememe external knowledge in the embedding layer, and at the same time integrating bidirectional attention flow after the encoding layer to enhance the semantic representation between the text and the answer. The above methods process by adding additional information to the context semantic information, but ignore the hidden rich text structure information.
[0004] With the rapid development of graph neural networks, graph neural networks have also been applied to the field of question generation. Huang et al. used a graph convolutional (GCN) network and a bidirectional long short-term memory network (LSTM) network to capture both the structural information and sequence information of the context simultaneously. Chai et al. constructed two graphs, where the nodes of the passage graph are sentences and the nodes of the answer graph are each answer, focusing on the information from the passage and the answer respectively. During the graph interaction process, each graph updates its node representation under the control of the gates calculated by the other graph. Pan et al. adopted an SRL-based or dependency parsing-based method to construct a semantic-level graph for the input text, then encoded the semantic graph by introducing an attention-based GGNN
[36] , and then fused the document-level and graph-level representations for joint training of content selection and question decoding.
[0005] In addition, some researchers have tried to introduce external knowledge into the question generation task and achieved good performance. Jia et al. augmented knowledge through a simple back-translation method, using sentence-level paraphrase generation (PG) as an auxiliary task for multi-task learning to supplement the paraphrase knowledge of the task-sharing encoder. On the other hand, a new diversity training loss function was adopted to introduce more question patterns into question generation. Jia et al. extracted knowledge-based structure triples from two databases, ConceptNet and WordNet, and selected triples containing the conversion information from text to question. Two auxiliary tasks, relation classification and tail concept generation, were designed to incorporate the commonsense triples into the main question generation model for training. Majumder et al. first discovered the missing content by distinguishing the patterns of global knowledge and local context, input this missing pattern into a fine-tuned BART
[29] model, and then used PPLM to decode and generate questions. Summary of the Invention
[0006] The purpose of the present invention is to provide a question generation method and system based on knowledge enhancement and dual-graph interaction, which is beneficial to improving the accuracy of question generation.
[0007] To achieve the above purpose, the technical solution adopted by the present invention is: a question generation method based on knowledge enhancement and dual-graph interaction, comprising the following steps:
[0008] Step A: Collect context sentences and corresponding question-answer pairs to construct a training dataset DR for question generation;
[0009] Step B: Use the training dataset DR and the knowledge graph to train a deep learning network model T based on the knowledge enhancement and dual-graph interaction network;
[0010] Step C: Input the context sentence and the answer into the trained deep learning network model T, and output the generated question.
[0011] Further, the specific steps of step B include the following steps:
[0012] Step B1: Perform initial encoding on all training samples in the training dataset DR to obtain the initial representation vector E of the context sentence c and the context sentence adjacency matrix A c ; Concatenate the predicted question words behind the answer to obtain the initial representation vector E of the answer a ; At the same time, based on the syntactic dependency tree of the context sentence and the syntactic dependency tree of the answer, perform node expansion from the knowledge graph to obtain the knowledge-enhanced context sentence graph and the knowledge-enhanced answer graph, and their adjacency matrices are respectively represented as A ckg 、A akg , and use the knowledge graph to map each node information into continuous low-dimensional vectors, and finally form the knowledge-expanded representation vectors E of the context sentence and the answer ck 、E ak ;
[0013] Step B2: Input the initial representation vector E of the context sentence obtained in step B1 c and the initial representation vector E of the answer a into two different bidirectional long short-term memory networks respectively to obtain the sentence representation vector H s and the answer representation vector H a ;
[0014] Step B3: Input the graph adjacency matrix A obtained in step B1 ckg and the knowledge-expanded representation vector E of the context sentence ck into a K-layer graph convolutional network, and input the graph adjacency matrix A akg and the knowledge-expanded representation vector E of the answer ak into another K-layer graph convolutional network. Through the graph convolutional network, perform node information update to learn external knowledge to obtain the knowledge-enhanced context sentence representation G ckg 、the knowledge-enhanced answer representation G akg , input the sentence representation vector H obtained in step B2 s and the knowledge-enhanced context sentence representation G ckg into the multi-head attention to obtain the enhanced context sentence representation vector H ckg , input the answer representation vector H a and the knowledge-enhanced answer representation G akg into another multi-head attention to obtain the enhanced answer context representation vector H akg ;
[0015] Step B4: Combine the context sentence representation vector H obtained in step B3 ckg with the Glove embedding E of the context sentence cs , the BERT embedding E cb to obtain the sentence representation H cgb . Combine the answer representation vector H obtained in step B3 akg with the Glove embedding E of the answer as , the BERT embedding E ab to obtain the answer representation H agb . Input the sentence representation H cgb and the answer representation H agb into the attention mechanism to obtain the context-aware answer representation H ac . Finally, combine the enhanced context sentence representation vector H ckg with H ac to obtain the representation vector H cka . Then input H cka into a bidirectional long short-term memory network to obtain the context sentence representation vector H related to the answer d ;
[0016] Step B5: Input the adjacency matrix A in step B1 c and the representation vector H obtained in step B4 d into a K-layer graph convolutional network to learn and obtain the structural information representation G of the context sentence cca ; Construct a semantic-related weight graph through the representation vector H d and the attention mechanism, and its adjacency matrix is denoted as A w . Input the matrix A w and the representation vector H d into a K-layer graph convolutional network to learn and obtain the semantic information representation G of the context sentence wca ; Construct a virtual graph for inter-graph information interaction, that is, duplicate a sentence and splice it behind the same sentence to form a graph. Connect the same words between the two sentences with an edge, and its value is 1. Do not connect different words directly, and its value is 0. The adjacency matrix of the virtual graph is denoted as A v . Concatenate the structural information feature G cca and the semantic information feature G wca and denote it as G cw . Input the matrix A v and G cw into a K-layer graph convolutional network to learn and obtain the final context sentence representation H e ;
[0017] Step B6: Calculate the context vector c using the final context sentence representation H e t Using a long short - term memory network as the decoder to generate questions word by word, embedding the previous word \(w\) t-1 and the context vector \(c\) t and inputting them into the long short - term memory network to calculate the new hidden state \(s\) t According to \(s\) t and the context vector \(c\) t calculate the probability \(p\) of generating a word gen At the same time, according to the hidden state \(s\) t and the context vector \(c\) t calculate a copy switch \(g\) to decide whether to copy a word from the input sentence or generate a new word. According to the target loss function \(loss\), use the backpropagation method to calculate the gradients of the parameters in the deep learning network model, and use the stochastic gradient descent method to update the parameters;
[0018] Step B7: When the number of times the BLEU - 4 score of the deep learning network model does not update reaches the given threshold or reaches the maximum number of iterations, terminate the training process of the deep learning network model.
[0019] Furthermore, the specific steps of step B1 are as follows:
[0020] Step B11: Traverse the training dataset \(DR\). After tokenizing the context sentences, answers, and questions in it and removing stop words, each training sample in \(DR\) is represented as \(ds=(s,a,q)\); where \(s\) is the text content of the context sentence, \(a\) is the content of the answer; \(q\) is the target question to be generated by the model.
[0021] The context sentence \(s\) is represented as:
[0022]
[0023] where is the \(i\) - th word in the context sentence, \(i = 1,2,\cdots,m\), and \(m\) is the number of words in the context sentence \(s\); the answer \(a\) is represented as:
[0024]
[0025] where is the \(i\) - th word in the answer, \(i = 1,2,\cdots,n\), and \(n\) is the number of words in the answer \(a\);
[0026] The target question \(q\) is represented as:
[0027]
[0028] where is the \(i\) - th word in the target question, \(i = 1,2,\cdots,l\), and \(l\) is the number of words in the target question \(q\);
[0029] Step B12: Encode the context sentence obtained in Step B11 to obtain the initial representation vector E of the context sentence cs ;
[0030] where E cs is expressed as:
[0031]
[0032] where is the word vector corresponding to the i-th word obtained by looking up in the pre-trained word vector matrix , where d represents the dimension of the word vector and |V| is the number of words in the vocabulary V;
[0033] Step B13: Perform part-of-speech tagging and named entity recognition on the context sentence to obtain the continuous representation vector E cp after part-of-speech tagging conversion and the continuous representation vector E cn after named entity recognition conversion;
[0034] where E cp is expressed as:
[0035]
[0036] where is the part-of-speech vector corresponding to the i-th word obtained by looking up in the part-of-speech vector matrix , where d1 represents the dimension of the part-of-speech vector and |V1| is the number of part-of-speech tags in the tag set V1;
[0037] E cn is expressed as:
[0038]
[0039] where is the entity vector corresponding to the i-th word obtained by looking up in the entity vector matrix , where d2 represents the dimension of the entity vector and |V2| is the number of entity tags in the tag set V2;
[0040] Step B14: Insert the [CLS] token at the front of the context sentence s and the [SEP] token at the end to construct the sentence input sequence s1 for the BERT model; input s1 into the BERT model to obtain the output sequence of the i-th layer of the model the output sequence of the last layer of the model E b ;
[0041] Among them, the sentence input sequence is represented as:
[0042]
[0043] Among them, the output sequence of the i-th layer of the model is represented as:
[0044]
[0045] Among them, the output sequence E of the last layer of the model cb is represented as:
[0046]
[0047] Where is the BERT word embedding vector corresponding to the i-th word obtained by looking up in the BERT word embedding vector matrix where d3 represents the dimension of the entity vector and |V3| is the number of words in the vocabulary V3;
[0048] Step B15: Concatenate the E cs obtained in step B12, the E cp and E cn obtained in step B13, and the E cb obtained in step B14 to get the initial representation vector E of the final context sentence c ;
[0049] Among them, E c is represented as:
[0050]
[0051] where m is the number of words, is the initial representation vector of the i-th word, which can be represented as:
[0052]
[0053] where [;] represents the concatenation operation;
[0054] Step B16: Perform syntactic dependency parsing on the context text to generate an undirected syntactic dependency structure graph and encode it into a corresponding m-order syntactic structure dependency adjacency matrix A c ;
[0055] Among them, A c is represented as:
[0056]
[0057]
[0058] Step B17: Generate interrogative words. First, perform named entity recognition on the answer to obtain the corresponding interrogative words; after obtaining the interrogative words, splice them behind the answer to obtain a new answer representation. For the answer representation perform encoding to obtain the initial representation vector E of the context sentence. as ; Insert [CLS] in front of the answer representation and [SEP] at the end to construct the answer input sequence of the BERT model. Input into the BERT model to obtain the output sequence of the i-th layer of the model. The output sequence E of the last layer of the model. ab ; Finally, splice E as and E ab to obtain the final initial answer representation vector E. a ;
[0059] Among them, is expressed as:
[0060]
[0061] Among them, is the original answer word, and
[0062] E as is expressed as:
[0063]
[0064] Among them, is the word vector corresponding to the i-th word , obtained by looking up in the pre-trained word vector matrix , where d represents the dimension of the word vector and |V| is the number of words in the vocabulary V;
[0065] Among them, the answer input sequence can be expressed as:
[0066]
[0067] Among them, the output sequence of the i-th layer of the model is expressed as:
[0068]
[0069] Among them, the output sequence E ab of the last layer of the model is expressed as:
[0070]
[0071] in, For the i-th word The corresponding BERT word embedding vector is obtained by using the BERT word embedding vector matrix where d3 represents the dimension of the entity vector and |V3| is the number of words in the vocabulary V3.
[0072] E a It can be expressed as:
[0073]
[0074] Among them, n+1 is the number of words in the answer, is the initial representation vector of the i-th word, It is expressed as:
[0075]
[0076] Where [;] represents the splicing operation;
[0077] Step B18: Based on the context syntactic dependency tree and the answer syntactic dependency tree, knowledge nodes are expanded from the knowledge graph. For sentences, only one of nouns, adjectives, numerals, and adverbs is expanded. For answers, all words are expanded due to their short length. For each node in the dependency tree, five relationships, ISA, Synonym, SymbolOf, SimilarTo, and DefinedAs, are selected from the knowledge graph for expansion. The expanded knowledge nodes are sorted according to the weights given by the knowledge graph, and the two words with the largest weight values are selected for expansion. The expanded nodes and the original sentence nodes form a knowledge-enhanced context sentence graph, whose adjacency matrix is represented as A. ckg , the expanded nodes and the original answer nodes form a knowledge-enhanced answer graph, whose adjacency matrix is represented by A akg ; Use the knowledge embedding of the knowledge graph itself to map each node into a continuous low-dimensional vector, and the final sentence knowledge embedding vector E ck and answer knowledge embedding vector E ak ;
[0078] Among them, E ck It is expressed as:
[0079]
[0080] in, z is the total number of context sentence nodes after expansion, d is the dimension of the knowledge word vector, is the knowledge embedding vector of the original context sentence, Knowledge embedding representation for the expanded sentence node;
[0081] Among them, A ckg is represented as:
[0082]
[0083] Among them, 1 indicates that there is an edge connection between nodes, and 0 indicates no;
[0084] Among them, E ak is represented as:
[0085]
[0086] Among them r is the total number of answer nodes after expansion, d is the dimension of the knowledge word vector, is the knowledge embedding vector of the original context sentence, is the knowledge embedding representation of the expanded answer node;
[0087] Among them, A akg is represented as:
[0088]
[0089] Among them, 1 indicates that there is an edge connection between nodes, 0 indicates no, and r is the total number of sentence nodes after expansion.
[0090] Furthermore, the specific steps of step B2 include the following steps:
[0091] Step B21: Input the sequence c of the initial representation vector E of the context sentence s into the forward layer and the backward layer of the bidirectional long short-term memory network respectively, and obtain the state vector sequence of the forward hidden layer and the state vector sequence of the backward hidden layer respectively, where [;] represents the vector connection operation;
[0092] Step B22: Input the sequence a of the initial representation vector E of the answer a into the forward layer and the backward layer of the bidirectional long short-term memory network respectively, and obtain the state vector sequence of the forward hidden layer and the state vector sequence f is the activation function; concatenation is used to obtain the context sentence representation vector [;] represents the vector concatenation operation.
[0093] Furthermore, step B3 specifically includes the following steps:
[0094] Step B31: Use the sentence knowledge embedding vector E ck as the initial representation vector of the graph node, and perform graph convolution operations on the adjacency matrix A ckg to update the graph node information; the update process of the hidden state of node i in the K-th layer of the graph convolution network is as follows:
[0095]
[0096] where k ∈ [0, k] represents the layer number of the graph convolution network, is the hidden state output by node i in the K-th layer network, σ is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector;
[0097] Step B32: Concatenate the hidden states of the K-th layer of the graph convolution network, take the first m nodes, and obtain the knowledge-enhanced context sentence representation G ckg , G ckg is represented as follows:
[0098]
[0099] where m is the number of words in the original sentence, and d is the dimension of the knowledge word vector;
[0100] Step B33: Use the answer knowledge embedding vector E ak as the initial representation vector of the graph node, and perform graph convolution operations on the adjacency matrix A akg to update the graph node information; the update process of the hidden state of node i in the K-th layer of the graph convolution network is as follows:
[0101]
[0102] where k ∈ [0, K - 1] represents the layer number of the graph convolution network, is the hidden state output by node i in the K-th layer network, σ is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector;
[0103] Step B34: Concatenate the hidden states of the K-th layer graph convolutional network, and take the first n + 1 nodes to obtain the knowledge-enhanced answer representation G akg , G akg is expressed as follows:
[0104]
[0105] where n + 1 is the number of words in the original answer, and d is the dimension of the knowledge word vector;
[0106] Step B35: Input the sentence representation vector H s obtained in Step B21 and the knowledge-enhanced context sentence representation G ckg into the multi-head attention to obtain the enhanced context sentence representation vector H ckg , H ckg is expressed as follows:
[0107] H ckg = MultiHead(G ckg , H s , H s )
[0108] MultiHead(G ckg , H s , H s ) = Concat(head1, head2, …, Head h )W3
[0109]
[0110] where MultiHead represents the multi-head attention mechanism, Concat() represents the concatenation operation, h is the number of parallel heads, is the trainable parameter matrix, head i represents the output function of the i-th attention function, i ∈ [1, h], and softmax() is the activation function;
[0111] Step B36: Input the answer representation vector H a obtained in Step B22 and the knowledge-enhanced answer representation G akg into another multi-head attention to obtain the enhanced answer context representation vector H akg , H akg is expressed as:
[0112] H akg = MultiHead(G akg , H a , H a )
[0113] MultiHead(G akg ,H a ,H a ) = Concat(head1, head2, …, head h )W7
[0114]
[0115] Among them, MultiHead represents the multi - head attention mechanism, Concat() represents the concatenation operation, h is the number of parallel heads, is the trainable parameter matrix, head i represents the output function of the i - th attention function, i ∈ [1, h], and softmax() is the activation function.
[0116] Furthermore, the specific steps of step B4 are as follows:
[0117] Step B41: Concatenate the context sentence representation vector H ckg obtained in step B35 with the Glove embedding E cs of the context sentence obtained in step B12 and the BERT embedding E cb obtained in step B14 to get the sentence representation H cgb , H cgb is expressed as follows:
[0118]
[0119] Among them d is the dimension of the word vector, [;] represents the vector concatenation operation, is the i - th hidden state representation of the enhanced context sentence representation vector H ckg ;
[0120] Step B42: Concatenate the answer representation vector H akg obtained in step B36 with the Glove embedding E as and the BERT embedding E ab obtained in step B17 to get the answer representation H agb , H agb is expressed as follows:
[0121]
[0122] Among them d is the dimension of the word vector, [;] represents the vector concatenation operation, The enhanced answer context representation vector H akg is the i-th hidden state representation of
[0123] Step B43: Input the sentence representation H cgb , the answer representation H agb and the answer representation vector H akg into the attention mechanism to obtain the context-aware answer representation H ac , and H ac is represented as follows:
[0124] H ac = softmax(Relu(W 11 H cgb )Relu(W 12 H agb ) T )H akg
[0125] where, softmax() and Relu() are activation functions, and W 11 , W 12 are learnable parameter matrices;
[0126] Step B44: Concatenate the enhanced context sentence representation vector H ckg with H ac to obtain the representation vector H cka ;
[0127] where, H cka is represented as follows:
[0128]
[0129] where, are respectively the i-th word hidden state representations of H ckg and H ac ;
[0130] Step B45: Input the sequence of the representation vector H cka into the forward layer and the backward layer of the bidirectional long short-term memory network respectively, and obtain the state vector sequence of the forward hidden layer and the state vector sequence of the backward hidden layer respectively. where f is the activation function; connect them to obtain the context sentence representation vector [;] represents the vector concatenation operation.
[0131] Furthermore, step B5 specifically includes the following steps:
[0132] Step B51: Use the context sentence representation vector H obtained in step B45 d as the initial representation vector of the graph node, and perform graph convolution operations on the adjacency matrix A obtained in step B16 using a K-layer graph convolutional network c to update the graph node information; the update process of the hidden state of node i in the K-layer graph convolutional network is as follows:
[0133]
[0134] where k ∈ [0, K-1] represents the layer number of the graph convolutional network, is the hidden state output by node i in the K-th layer network, σ is the non-linear activation function Relu(), is a trainable parameter matrix, is the bias vector;
[0135] Step B52: Connect the hidden states of the K-th layer graph convolutional network to obtain the structural information representation G of the context sentence cca , G cca is expressed as follows:
[0136]
[0137] where m is the number of words in the original sentence, and d is the dimension of the knowledge word vector;
[0138] Step B53: Use the attention mechanism to calculate the semantic similarity score for the context sentence representation vector H d to obtain the edge weight matrix A w , and its calculation process is as follows:
[0139]
[0140] where m is the length of the context sentence, d is the dimension of the H d vector, is a trainable parameter matrix;
[0141] Step B54: Use the context sentence representation vector H d as the initial representation vector of the graph node, and perform graph convolution operations on the adjacency matrix A using a K-layer graph convolutional network w to update the graph node information; the update process of the hidden state of node i in the K-layer graph convolutional network is as follows:
[0142]
[0143] where \(k\in[0, K - 1]\) represents the number of layers of the graph convolutional network, is the hidden state output by node \(i\) in the \(K\)-th layer network, \(\sigma\) is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector;
[0144] Step B55: Concatenate the hidden states of the \(K\)-th layer graph convolutional network to obtain the semantic information representation \(G\) of the context sentence wca , \(G\) wca is represented as follows:
[0145]
[0146] where \(m\) is the number of words in the original sentence, and \(d\) is the dimension of the knowledge word vector;
[0147] Step B56: Construct a virtual graph for inter-graph information interaction, and represent a context sentence as connected to the same context sentence and represented as and then construct a graph, that is, construct an \((m + m)\times(m + m)\) adjacency matrix \(A\) v , which is represented as follows:
[0148]
[0149]
[0150] where \((m + m)\) represents the sum of the lengths of the two context sentences;
[0151] Step B57: Concatenate the structural information representation \(G\) cca of the context sentence with the semantic information representation \(G\) wca of the context sentence to obtain the representation vector \(G\) cw ;
[0152] where \(G\) cw is represented as follows:
[0153]
[0154] where are respectively the \(i\)-th word hidden state representations of \(G\) cca and \(G\) wca ;
[0155] Step B58: Use the concatenated representation vector \(G\) cw as the initial representation vector of the graph nodes, and use the \(K\)-layer graph convolutional network on the adjacency matrix \(A\) vPerform graph convolution operations to update the information of graph nodes; the update process of the hidden state of node i in the K-th layer graph convolution network is as follows:
[0156]
[0157] where k ∈ [0, K - 1] represents the layer number of the graph convolution network, is the hidden state output by node i in the K-th layer network, σ is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector;
[0158] Step B59: Concatenate the hidden states of the K-th layer graph convolution network, and the final context sentence representation H e is represented as follows:
[0159]
[0160] where, m is the number of words in the original sentence, and d is the dimension of the knowledge word vector.
[0161] Furthermore, the specific steps of the said Step B6 include the following steps:
[0162] Step B61: Use a long short-term memory network as a decoder to generate questions word by word; at the decoding time step t, the long short-term memory network decoder reads the previous word embedding w t-1 and the context vector c t together to calculate the new hidden state representation s t :
[0163] s t = LSTM(s t-1 , [w t-1 ; c t )
[0164]
[0165] c t = ∑ i a t,i H e
[0166] where s t-1 is the hidden state of the previous long short-term memory network, and W 18 is the learnable parameter matrix;
[0167] Step B62: Calculate the probability p of generating a word from the long short-term memory network hidden representation s t and the context vector c t gen ;
[0168] p gen = softmax(W 19 [s t ; c t ) + b6)
[0169] where softmax() is the activation function, [;] represents the concatenation operation, W 19 is the learnable parameter matrix, and b6 is the learnable bias vector;
[0170] Step B63: To handle rare and unknown words, calculate g according to the pointer network to decide whether to copy a word from the input sentence or generate a new word, where g is expressed as follows:
[0171] g = σ(W 20 s t + W 21 c t ) + b7)
[0172] where σ is the non - linear activation function, W 20 , W 21 are the learnable parameter matrices, and b7 is the learnable bias vector;
[0173] Step B64: Finally, generate the question probability p final , and its calculation formula is as follows:
[0174] p final (y t |y<t; θ) = gp copy (y t , θ1) + (1 - g)p gen (y t , θ2)
[0175] where θ, θ1, θ2 are the learnable network parameters;
[0176] Step B65: Use cross - entropy as the loss function to calculate the loss value and penalize the repetition phenomenon through covloss. Update the learning rate through the gradient optimization algorithm Adam, and use backpropagation to iteratively update the model parameters to train the model by minimizing the loss function; the calculation formula for minimizing the loss function L is as follows:
[0177] L = -∑logp final (y t |y<t; θ) + βcovloss t
[0178] covloss t = ∑ i min(at,i , c t,i )
[0179] where p final (y t |y < t; θ) is the calculated probability of generating a question word, β is a hyperparameter, a t,i is the attention vector representation of the i-th generated word at time step t calculated in step B61, c t,i is the context vector representation of the i-th generated word at time step t.
[0180] The present invention also provides a question generation system based on knowledge enhancement and dual-graph interaction using the above method, including:
[0181] A data collection module for collecting context sentences and corresponding question-answer pairs to construct a training data set for question generation;
[0182] A text preprocessing module for preprocessing the training samples in the training data set;
[0183] A text encoding module for initially encoding all the training samples in the training data set to obtain the initial representation vectors of the context sentences, the context sentence adjacency matrix, the initial representation vectors of the answers, and at the same time expanding the knowledge nodes of the context sentences and answer texts from the knowledge graph to construct knowledge-enhanced context sentences and a knowledge-enhanced answer adjacency matrix, and querying to obtain the knowledge-expanded representation vectors of the context sentences and answers;
[0184] A network training module for inputting the initial representation vectors of the context sentences, the context sentence adjacency matrix, the initial representation vectors of the answers, the context sentence adjacency matrix, and the initial representation vectors of the answers into a deep learning network to obtain the final representation vectors of the user comments, inputting the knowledge-enhanced context sentences, the knowledge-enhanced answer adjacency matrix, and the knowledge-expanded representation vectors of the context sentences and answers into the deep learning network to obtain the final representation vectors of the context sentences, using the final representation vectors to decode and generate questions, and calculating the loss by comparing with the reference questions in the training data set, and training the entire deep learning network model with the goal of minimizing the loss to obtain a deep learning network model based on knowledge enhancement and dual-graph interaction network;
[0185] A question generation module for using the trained deep learning network model to process the input context sentences and answers to generate questions related to the answers, with correct grammar and answerable by the sentences.
[0186] Compared with the prior art, the present invention has the following beneficial effects: The present invention constructs a deep learning network model based on knowledge enhancement and dual-graph interaction network, and trains this deep learning network model through a training dataset and a knowledge graph. Finally, the trained deep learning network model is used to generate questions related to the answer, with correct grammar and answerable by the sentence for the input context sentence and answer, improving the accuracy of question generation, and having strong practicability and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0187] Figure 1 is a flowchart of the method implementation of the embodiment of the present invention;
[0188] Figure 2 is an architecture diagram of the deep learning network model in the embodiment of the present invention;
[0189] Figure 3 is a schematic diagram of the system structure of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0190] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0191] It should be noted that the following detailed description is exemplary and is intended to provide further description of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0192] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0193] As Figure 1 shown, this embodiment provides a question generation method based on knowledge enhancement and dual-graph interaction, including the following steps:
[0194] Step A: Collect context sentences and corresponding question-and-answer pairs, and construct a training dataset DR for question generation.
[0195] Step B: Use the training dataset DR and the knowledge graph ConceptNet to train a deep learning network model T based on knowledge enhancement and dual-graph interaction network, and generate corresponding questions through this model. The architecture of the deep learning network model T is as Figure 2 shown.
[0196] Step C: Input the context sentence and the answer into the trained deep learning network model T, and output the generated question.
[0197] In this embodiment, step B specifically includes the following steps:
[0198] Step B1: Perform initial encoding on all training samples in the training dataset DR to obtain the initial representation vector E of the context sentence c and the context sentence adjacency matrix A c ; Concatenate the predicted interrogative word behind the answer to obtain the initial representation vector E of the answer a ; At the same time, based on the syntactic dependency tree of the context sentence and the syntactic dependency tree of the answer, perform node expansion from the knowledge graph ConceptNet to obtain the knowledge-enhanced context sentence graph and the knowledge-enhanced answer graph, and their adjacency matrices are respectively represented as A ckg 、A akg , and use the knowledge graph to map each node information into a continuous low-dimensional vector, and finally form the knowledge-expanded representation vectors E ck 、E ak .
[0199] In this embodiment, step B1 specifically includes the following steps:
[0200] Step B11: Traverse the training dataset DR. After tokenizing the context sentence, the answer, and the question in it and removing the stop words, each training sample in DR is represented as ds = (s, a, q); where s is the text content of the context sentence, a is the content of the answer; q is the target question to be generated by the model.
[0201] The context sentence s is represented as:
[0202]
[0203] where is the i-th word in the context sentence, i = 1, 2,..., m, and m is the number of words in the context sentence s.
[0204] The answer a is represented as:
[0205]
[0206] where is the i-th word in the answer, i = 1, 2,..., n, and n is the number of words in the answer a.
[0207] The target question q is represented as:
[0208]
[0209] Among them, is the i-th word in the target question, where i = 1, 2, …, l, and l is the number of words in the target question q.
[0210] Step B12: Encode the context sentence obtained in step B11 to get the initial representation vector E of the context sentence cs .
[0211] Among them, E cs is expressed as:
[0212]
[0213] Among them, is the word vector corresponding to the i-th word , obtained by looking up in the pre-trained word vector matrix , where d represents the dimension of the word vector, and |V| is the number of words in the vocabulary V.
[0214] Step B13: Use a tool to perform part-of-speech tagging and named entity recognition on the context sentence to obtain the continuous representation vector E cp of the part-of-speech tagging conversion and the continuous representation vector E cn of the named entity recognition conversion.
[0215] Among them, E cp is expressed as:
[0216]
[0217] Among them, is the part-of-speech vector corresponding to the i-th word , obtained by looking up in the part-of-speech vector matrix , where d1 represents the dimension of the part-of-speech vector, and |V1| is the number of part-of-speech tags in the tag set V1.
[0218] E cn is expressed as:
[0219]
[0220] Among them, is the entity vector corresponding to the i-th word , obtained by looking up in the entity vector matrix , where d2 represents the dimension of the entity vector, and |V2| is the number of entity tags in the tag set V2.
[0221] Step B14: Insert the [CLS] token at the front of the context sentence s and the [SEP] token at the end to construct the sentence input sequence s1 for the BERT model; input s1 into the BERT model to obtain the output sequence of the i-th layer of the model The output sequence E of the last layer of the model b .
[0222] Among them, the sentence input sequence is expressed as:
[0223]
[0224] Among them, the output sequence of the i-th layer of the model is expressed as:
[0225]
[0226] Among them, the output sequence E of the last layer of the model cb is expressed as:
[0227]
[0228] Among them is the BERT word embedding vector corresponding to the i-th word obtained by looking up in the BERT word embedding vector matrix , where d3 represents the dimension of the entity vector and |V3| is the number of words in the vocabulary V3.
[0229] Step B15: Concatenate the E cs obtained in step B12, the E cp obtained in step B13, and the E cn obtained in step B13, and the E cb obtained in step B14 to obtain the initial representation vector E of the final context sentence c .
[0230] Among them, E c is expressed as:
[0231]
[0232] Among them, m is the number of words, is the initial representation vector of the i-th word, which can be expressed as:
[0233]
[0234] Among them, [;] represents the concatenation operation.
[0235] Step B16: Perform syntactic dependency parsing on the context text to generate an undirected syntactic dependency structure graph, and encode it into a corresponding m-order syntactic structure dependency adjacency matrix A c 。
[0236] where A c is represented as:
[0237]
[0238]
[0239] Step B17: Generate question words. First, perform named entity recognition on the answer. If the answer is a person's name, location, time, or number, the corresponding question words are who, where, when, and how. If the answer is a noun or noun phrase, the question word is what. If the answer is an adjective or ordinal word, the question word is which. If the answer contains "reason", "because", etc., the question word is why. In other cases, it is set to other. After obtaining the question word, splice it to the back of the answer to obtain a new answer representation Encode the answer representation to obtain the initial representation vector E of the context sentence as ; Insert the [CLS] at the front of the answer representation and insert the [SEP] marker at the end to construct the answer input sequence for the BERT model Input into the BERT model to obtain the output sequence of the i-th layer of the model The output sequence E of the last layer of the model ab ; Finally, splice E as and E ab to obtain the final initial representation vector E of the answer a 。
[0240] where is represented as:
[0241]
[0242] where is the original answer word is the predicted question word
[0243] E as is represented as:
[0244]
[0245] where is the i-th word The corresponding word vectors are obtained by looking up in the pre-trained word vector matrix where d represents the dimension of the word vector and |V| is the number of words in the vocabulary V.
[0246] Among them, the answer input sequence can be expressed as:
[0247]
[0248] Among them, the output sequence of the i-th layer of the model is expressed as:
[0249]
[0250] Among them, the output sequence E ab of the last layer of the model is expressed as:
[0251]
[0252] Among them, is the BERT word embedding vector corresponding to the i-th word and is obtained by looking up in the BERT word embedding vector matrix where d3 represents the dimension of the entity vector and |V3| is the number of words in the vocabulary V3.
[0253] E a can be expressed as:
[0254]
[0255] Among them, n + 1 is the number of words in the answer, is the initial representation vector of the i-th word, which is expressed as:
[0256]
[0257] where [;] represents the concatenation operation.
[0258] Step B18: Expand knowledge nodes from the knowledge graph based on the context syntactic dependency tree and the answer syntactic dependency tree. For sentences, only expand one of nouns, adjectives, numerals, and adverbs. Since the answer is short, expand all words. For each node of the dependency tree, select five relationships in the knowledge graph, namely ISA, Synonym, SymbolOf, SimilarTo, and DefinedAs, for expansion. Sort the expanded knowledge nodes according to the weights given by the knowledge graph, and select the two words with the largest weight values for expansion. The expanded nodes and the original sentence nodes form a knowledge-enhanced context sentence graph, and its adjacency matrix is represented as A ckg, the extended nodes and the original answer nodes form a knowledge-enhanced answer graph, whose adjacency matrix is denoted as A akg ; Map each node to a continuous low-dimensional vector by using the knowledge embedding of the knowledge graph itself, and the final sentence knowledge embedding vector E ck and the answer knowledge embedding vector E ak .
[0259] Among them, E ck can be expressed as:
[0260]
[0261] Among them, z is the total number of context sentence nodes after expansion, d is the dimension of the knowledge word vector, is the knowledge embedding vector of the original context sentence, is the knowledge embedding representation of the expanded sentence node.
[0262] Among them, A ckg can be expressed as:
[0263]
[0264] Among them, 1 indicates that there is an edge connection between nodes, and 0 indicates no.
[0265] Among them, E ak can be expressed as:
[0266]
[0267] Among them r is the total number of answer nodes after expansion, d is the dimension of the knowledge word vector, is the knowledge embedding vector of the original context sentence, is the knowledge embedding representation of the expanded answer node.
[0268] Among them, A akg can be expressed as:
[0269]
[0270] Among them, 1 indicates that there is an edge connection between nodes, 0 indicates no, and r is the total number of sentence nodes after expansion.
[0271] Step B2: Input the initial representation vectors E c of the context sentences obtained in Step B1 and the initial representation vector E a of the answer into two different bidirectional long short-term memory networks respectively to obtain the sentence representation vector H sand the answer representation vector H a 。
[0272] In this embodiment, step B2 specifically includes the following steps:
[0273] Step B21: Input the initial representation vector E c of the context sentence s into the forward layer and the backward layer of the bidirectional long short-term memory network respectively, and obtain the state vector sequence of the forward hidden layer and the state vector sequence of the backward hidden layer where f is the activation function; connect to obtain the context sentence representation vector [;] represents the vector connection operation.
[0274] Step B22: Input the initial representation vector E a of the answer a into the forward layer and the backward layer of the bidirectional long short-term memory network respectively, and obtain the state vector sequence of the forward hidden layer and the state vector sequence of the backward hidden layer where f is the activation function; connect to obtain the context sentence representation vector [;] represents the vector connection operation.
[0275] Step B3: Input the graph adjacency matrix A ckg obtained in step B1 and the knowledge extension representation vector E ck of the context sentence into a K-layer graph convolutional network, and input the graph adjacency matrix A akg and the knowledge extension representation vector E ak of the answer into another K-layer graph convolutional network, and perform node information update through the graph convolutional network to learn external knowledge, and obtain the knowledge-enhanced context sentence representation G ckg , the knowledge-enhanced answer representation G akg . Input the sentence representation vector H s obtained in step B2 and the knowledge-enhanced context sentence representation G ckg into the multi-head attention, and obtain the enhanced context sentence representation vector H ckg . Input the answer representation vector H a and the knowledge-enhanced answer representation G akg into another multi-head attention, and obtain the enhanced answer context representation vector H akg 。
[0276] In this embodiment, step B3 specifically includes the following steps:
[0277] Step B31: Take the sentence knowledge embedding vector E ck as the initial representation vector of the graph node, and use the K-layer graph convolutional network to perform graph convolution operations on the adjacency matrix A ckg to update the graph node information; the update process of the hidden state of node i in the K-layer graph convolutional network is as follows:
[0278]
[0279] where k ∈ [0, k] represents the number of layers of the graph convolutional network, is the hidden state output by node i in the K-th layer network, σ is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector.
[0280] Step B32: Connect the hidden states of the K-th layer graph convolutional network, take the first m nodes, and obtain the knowledge-enhanced context sentence representation G ckg , G ckg is expressed as follows:
[0281]
[0282] where m is the number of words in the original sentence, and d is the dimension of the knowledge word vector.
[0283] Step B33: Take the answer knowledge embedding vector E ak as the initial representation vector of the graph node, and use the K-layer graph convolutional network to perform graph convolution operations on the adjacency matrix A akg to update the graph node information; the update process of the hidden state of node i in the K-layer graph convolutional network is as follows:
[0284]
[0285] where k ∈ [0, K - 1] represents the number of layers of the graph convolutional network, is the hidden state output by node i in the K-th layer network, σ is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector.
[0286] Step B34: Connect the hidden states of the K-th layer graph convolutional network, take the first n + 1 nodes, and obtain the knowledge-enhanced answer representation G akg , G akg is expressed as follows:
[0287]
[0288] wherein n + 1 is the number of words in the original answer, and d is the dimension of the knowledge word vector.
[0289] Step B35: Input the sentence representation vector H s and the knowledge-enhanced context sentence representation G ckg into the multi-head attention to obtain the enhanced context sentence representation vector H ckg , H ckg is expressed as follows:
[0290] H ckg = MultiHead(G ckg , H s , H s
[0291] MultiHead(G ckg , H s , H s ) = Concat(head1, head2,..., head h )W3
[0292]
[0293] wherein, MultiHead represents the multi-head attention mechanism, Concat() represents the concatenation operation, h is the number of parallel heads, is the trainable parameter matrix, and head i represents the output function of the i-th attention function, i ∈ [1, h], and softmax() is the activation function.
[0294] Step B36: Input the answer representation vector H a and the knowledge-enhanced answer representation G akg into another multi-head attention to obtain the enhanced answer context representation vector H akg , H akg is expressed as:
[0295] H akg = MultiHead(G akg , H a , H a
[0296] MultiHead(G akg , H a , H a ) = Conca(head1, head2,..., head h)W7
[0297]
[0298] Among them, MultiHead represents the multi-head attention mechanism, Concat() represents the concatenation operation, h is the number of parallel heads, is a trainable parameter matrix, head i represents the output function of the i-th attention function, i ∈ [1, h], and softmax() is the activation function.
[0299] Step B4: The context sentence representation vector H obtained in step B3 ckg is concatenated with the Glove embedding E of the context sentence cs and the BERT embedding E cb to obtain the sentence representation H cgb , the answer representation vector H obtained in step B3 akg is concatenated with the Glove embedding E of the answer as and the BERT embedding E ab to obtain the answer representation H agb , the sentence representation H cgb and the answer representation H agb are input into the attention mechanism to obtain the context-aware answer representation H ac , and finally the enhanced context sentence representation vector H ckg is concatenated with H ac to obtain the representation vector H cka , and then H cka is input into the bidirectional long short-term memory network to obtain the context sentence representation vector H related to the answer d .
[0300] In this embodiment, step B4 specifically includes the following steps:
[0301] Step B41: The context sentence representation vector H obtained in step B35 ckg is concatenated with the Glove embedding E of the context sentence obtained in step B12 cs and the BERT embedding E obtained in step B14 cb to obtain the sentence representation H cgb , H cgb is expressed as follows:
[0302]
[0303] Among them d is the dimension of the word vector, [;] represents the vector concatenation operation, The enhanced context sentence representation vector H ckg is the i-th hidden state representation of
[0304] Step B42: Concatenate the answer representation vector H akg obtained in step B36 with the Glove embedding E as and the BERT embedding E ab to obtain the answer representation H agb , and H agb is represented as follows:
[0305]
[0306] where d is the dimension of the word vector, [;] represents the vector concatenation operation, is the i-th hidden state representation of the enhanced answer context representation vector H akg of
[0307] Step B43: Input the sentence representation H cgb , the answer representation H agb and the answer representation vector H akg into the attention mechanism to obtain the context-aware answer representation H ac , and H ac is represented as follows:
[0308] H ac = softmax(Relu(W 11 H cgb )Relu(W 12 H agb ) T )H akg
[0309] where, softmax(), Relu() are activation functions, and W 11 , W 12 are learnable parameter matrices.
[0310] Step B44: Concatenate the enhanced context sentence representation vector H ckg with H ac to obtain the representation vector H cka .
[0311] where, H cka is represented as follows:
[0312]
[0313] where, are respectively H ckg and Hac The hidden state representation of the i-th word.
[0314] Step B45: Input the representation vector H cka sequence of into the forward layer and the backward layer of a bidirectional long short-term memory network respectively, and obtain the state vector sequence of the forward hidden layer state vector sequence of the backward hidden layer where f is an activation function; concatenate to obtain the context sentence representation vector [;] represents the vector concatenation operation.
[0315] Step B5: Input the adjacency matrix A c obtained in Step B1 and the representation vector H d obtained in Step B4 into a K-layer graph convolutional network to learn the structural information representation G cca of the context sentence; construct a semantic correlation weight graph through the representation vector H d and the attention mechanism, and its adjacency matrix is denoted as A w , input the matrix A w and the representation vector H d into a K-layer graph convolutional network to learn the semantic information representation G wca of the context sentence; construct a virtual graph for inter-graph information interaction, that is, duplicate a sentence and splice it behind the same sentence to form a graph, connect the same words in the two sentences with an edge, and its value is 1, and do not connect different words directly, and its value is 0. The adjacency matrix of the virtual graph is denoted as A v , splice the structural information feature G cca and the semantic information feature G wca to be denoted as G cw , input the matrix A v and G cw into a K-layer graph convolutional network to learn the final context sentence representation H e .
[0316] In this embodiment, Step B5 specifically includes the following steps:
[0317] Step B51: Use the context sentence representation vector H d obtained in Step B45 as the initial representation vector of the graph nodes, and perform graph convolution operations on the adjacency matrix A c obtained in Step B16 using a K-layer graph convolutional network to update the graph node information. The update process of the hidden state of node i in the K-th layer graph convolutional network is as follows:
[0318]
[0319] where \(k\in[0, K - 1]\) represents the number of layers of the graph convolutional network, is the hidden state output by node \(i\) in the \(K\)-th layer network, \(\sigma\) is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector.
[0320] Step B52: Connect the hidden states of the \(K\)-th layer graph convolutional network to obtain the structural information representation \(G\) of the context sentence cca , \(G\) cca is expressed as follows:
[0321]
[0322] where \(m\) is the number of words in the original sentence, and \(d\) is the dimension of the knowledge word vector.
[0323] Step B53: Use the attention mechanism to calculate the semantic similarity score for the context sentence representation vector \(H\) d to obtain the edge weight matrix \(A\) w , and its calculation process is as follows:
[0324]
[0325] where \(m\) is the length of the context sentence, \(d\) is the dimension of the vector \(H\) d , is the trainable parameter matrix.
[0326] Step B54: Use the context sentence representation vector \(H\) d as the initial representation vector of the graph nodes, and use the \(K\)-layer graph convolutional network to perform graph convolution operations on the adjacency matrix \(A\) w to update the graph node information; the update process of the hidden state of node \(i\) in the \(K\)-th layer graph convolutional network is as follows:
[0327]
[0328] where \(k\in[0, K - 1]\) represents the number of layers of the graph convolutional network, is the hidden state output by node \(i\) in the \(K\)-th layer network, \(\sigma\) is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector.
[0329] Step B55: Concatenate the hidden states of the K-th layer graph convolutional network to obtain the semantic information representation G of the context sentence wca , G wca is expressed as follows:
[0330]
[0331] where m is the number of words in the original sentence, and d is the dimension of the knowledge word vector.
[0332] Step B56: Construct a virtual graph for inter-graph information interaction, and represent a context sentence as connected to the same context sentence and represented as followed by graph construction, that is, construct a (m + m) × (m + m) adjacency matrix A v , which is expressed as follows:
[0333]
[0334]
[0335] where (m + m) represents the sum of the lengths of two context sentences.
[0336] Step B57: Concatenate the structural information representation G of the context sentence cca with the semantic information representation G of the context sentence wca to obtain the representation vector G cw .
[0337] where, G cw is expressed as follows:
[0338]
[0339] where are respectively the i-th word hidden state representations of G cca and G wca .
[0340] Step B58: Use the concatenated representation vector G cw as the initial representation vector of the graph node, and perform graph convolution operations on the adjacency matrix A v using the K-th layer graph convolutional network to update the graph node information; the update process of the hidden state of node i in the K-th layer graph convolutional network is as follows:
[0341]
[0342] where k ∈ [0, K - 1] represents the number of layers of the graph convolutional network, is the hidden state output by node i in the K-th layer network, σ is the non-linear activation function Relu(), is a trainable parameter matrix, is the bias vector.
[0343] Step B59: Concatenate the hidden states of the K-th layer graph convolutional network, and the final context sentence representation H e is expressed as follows:
[0344]
[0345] wherein, m is the number of words in the original sentence, and d is the dimension of the knowledge word vector.
[0346] Step B6: Use the final context sentence representation H e to calculate the context vector c t , use the long short-term memory network as the decoder to generate words one by one, and embed the previous word w t-1 and the context vector c t are input into the long short-term memory network together to calculate the new hidden state s t , according to s t and the context vector c t calculate the probability p of generating a word gen , and at the same time calculate a copy switch g according to the hidden state s t and the context vector c t to determine whether to copy a word from the input sentence or generate a new word. According to the target loss function loss, use the backpropagation method to calculate the gradients of the parameters in the deep learning network model, and use the stochastic gradient descent method to update the parameters.
[0347] In this embodiment, the step B6 specifically includes the following steps:
[0348] Step B61: Use the long short-term memory network as the decoder to generate words one by one; at the decoding time step t, the long short-term memory network decoder reads the previous word embedding w t-1 and the context vector c t to calculate the new hidden state representation s t :
[0349] s t = LSTM(s t-1 , [w t-1 ; c t )
[0350]
[0351]
[0352] Among them, s t-1 is the hidden state of the previous long short-term memory network, and W 18 is a learnable parameter matrix.
[0353] Step B62: Calculate the probability p t of generating a word from the hidden representation s t of the long short-term memory network and the context vector c gen .
[0354] p gen = softmax(W 19 [s t ; c t ) + b6
[0355] where softmax() is the activation function, [;] represents the concatenation operation, W 19 is a learnable parameter matrix, and b6 is a learnable bias vector.
[0356] Step B63: To handle rare and unknown words, determine whether to copy a word from the input sentence or generate a new word according to the calculation of g by the pointer network, where g is expressed as follows:
[0357] g = σ(W 20 s t + W 21 c t ) + b7
[0358] where σ is the non-linear activation function, W 20 , W 21 are learnable parameter matrices, and b7 is a learnable bias vector.
[0359] Step B64: Finally, generate the problem probability p final , and its calculation formula is as follows:
[0360] p final (y t |y<t; θ) = gp copy (y t , θ1) + (1 - g)p gen (y t , θ2)
[0361] where θ, θ1, and θ2 are learnable network parameters.
[0362] Step B65: Calculate the loss value using cross - entropy as the loss function and penalize the repetition phenomenon through covloss. Update the learning rate through the gradient optimization algorithm Adam, and use backpropagation to iteratively update the model parameters to train the model by minimizing the loss function. The calculation formula for minimizing the loss function L is as follows:
[0363] L = -∑logp final (y t |y<t; θ)+βcovloss t
[0364] covloss t =∑ i min(a t,i ,c t,i )
[0365] where p final (y t |y<t; θ) is the calculated probability of generating the problem word, β is a hyperparameter, a t,i is the attention vector representation of the i - th generated word at time step t calculated in Step B61, and c t,i is the context vector representation of the i - th generated word at time step t.
[0366] Step B7: When the number of times the BLEU - 4 score of the deep learning network model does not update reaches the given threshold or reaches the maximum number of iterations, terminate the training process of the deep learning network model.
[0367] As Figure 3 shown, this embodiment provides a question generation system based on knowledge enhancement and dual - graph interaction using the above - mentioned method, including: a data collection module, a text pre - processing module, a text encoding module, a network training module, and a question generation module.
[0368] The data collection module is used to collect context sentences and corresponding question - answer pairs to construct a training data set for question generation.
[0369] The text pre - processing module is used to pre - process the training samples in the training data set, including word segmentation, stop - word removal, question - word generation, syntactic dependency parsing, named - entity recognition, part - of - speech tagging, virtual graph construction, etc.
[0370] The text encoding module is used to perform initial encoding on all training samples in the training data set to obtain the initial representation vectors of context sentences, the adjacency matrix of context sentences, the initial representation vectors of answers, and at the same time expand the knowledge nodes of context sentences and answer texts from the knowledge graph, construct knowledge - enhanced context sentences and knowledge - enhanced answer adjacency matrices, and query to obtain the knowledge - extended representation vectors of context sentences and answers.
[0371] The network training module is used to input the initial representation vector of the context sentence, the context sentence adjacency matrix, the initial representation vector of the answer, the context sentence adjacency matrix, and the initial representation vector of the answer into the deep learning network to obtain the final representation vector of the user comment. The knowledge-enhanced context sentence, the knowledge-enhanced answer adjacency matrix, and the knowledge-expanded representation vectors of the context sentence and the answer are input into the deep learning network to obtain the final representation vector of the context sentence. The final representation vector is used to decode and generate questions, and the loss is calculated by comparing with the reference questions in the training dataset. The entire deep learning network model is trained with the goal of minimizing the loss to obtain a deep learning network model based on knowledge enhancement and dual-graph interaction network.
[0372] The question generation module is used to process the input context sentence and answer using the trained deep learning network model to generate questions related to the answer, with correct grammar and answerable by the sentence.
[0373] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0374] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0375] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0376] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the steps in a process Figure 1 a process or multiple processes and / or blocks Figure 1 or the steps of the functions specified in multiple blocks.
[0377] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation of the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A question generation method based on knowledge enhancement and bi-graph interaction, characterized in that It includes the following steps: Step A: Collect context sentences and corresponding question-answer pairs, and construct a training dataset DR for question generation; Step B: Use the training dataset DR and the knowledge graph to train a deep learning network model T based on knowledge enhancement and dual-graph interaction network; Step C: Input the context sentence and the answer into the trained deep learning network model T, and output the generated question; The specific steps of step B include the following steps: Step B1: Perform initial encoding on all training samples in the training dataset DR to obtain the initial representation vector E of the context sentence c and the context sentence adjacency matrix A c ; Concatenate the predicted interrogative word after the answer to obtain the initial representation vector E of the answer a ; At the same time, based on the syntactic dependency tree of the context sentence and the syntactic dependency tree of the answer, perform node expansion from the knowledge graph to obtain the knowledge-enhanced context sentence graph and the knowledge-enhanced answer graph, and their adjacency matrices are respectively represented as A ckg 、A akg , and use the knowledge graph to map each node information into a continuous low-dimensional vector, and finally form the knowledge-expanded representation vectors E ck 、E ak ; Step B2: Input the initial representation vector E of the context sentence obtained in Step B1 c and the initial representation vector E of the answer a into two different bidirectional long short-term memory networks respectively to obtain the sentence representation vector H s and the answer representation vector H a ; Step B3: Input the graph adjacency matrix A obtained in Step B1 ckg and the knowledge expansion representation vector E of the context sentence ck into a K-layer graph convolutional network. Input the graph adjacency matrix A akg and the knowledge expansion representation vector E of the answer ak into another K-layer graph convolutional network. Update node information through the graph convolutional network to learn external knowledge, and obtain the knowledge-enhanced context sentence representation G ckg and the knowledge-enhanced answer representation G akg . Input the sentence representation vector H obtained in Step B2 s and the knowledge-enhanced context sentence representation G ckg into the multi-head attention to obtain the enhanced context sentence representation vector H ckg . Input the answer representation vector H a and the knowledge-enhanced answer representation G akg into another multi-head attention to obtain the enhanced answer context representation vector H akg ; Step B4: Take the context sentence representation vector H obtained in step B3 ckg and concatenate it with the Glove embedding E of the context sentence cs and the BERT embedding E cb to obtain the sentence representation H cgb . Take the answer representation vector H obtained in step B3 akg and concatenate it with the Glove embedding E of the answer as and the BERT embedding E ab to obtain the answer representation H agb . Input the sentence representation H cgb and the answer representation H agb into the attention mechanism to obtain the context-aware answer representation H ac . Finally, concatenate the enhanced context sentence representation vector H ckg with H ac to obtain the representation vector H cka . Then input H cka into a bidirectional long short-term memory network to obtain the context sentence representation vector H related to the answer d ; Step B5: Take the adjacency matrix A in Step B1 c and the representation vector H obtained in Step B4 d and input them into a K-layer graph convolutional network to learn the structural information representation G of the context sentence cca ; Construct a semantic correlation weight graph through the representation vector H d and the attention mechanism, and its adjacency matrix is denoted as A w , take the matrix A w and the representation vector H d and input them into a K-layer graph convolutional network to learn the semantic information representation G of the context sentence wca ; Construct a virtual graph for inter-graph information interaction, that is, copy a sentence once and splice it behind the same sentence to form a graph. Connect the same words in the two sentences with an edge, and its value is 1. Do not connect different words directly, and its value is 0. The adjacency matrix of the virtual graph is denoted as A v , splice the structural information feature G cca and the semantic information feature G wca to be denoted as G cw , take the matrix A v and G cw and input them into a K-layer graph convolutional network to learn the final context sentence representation H e ; Step B6: Represent H with the final context sentence e Calculate the context vector c t , use a long short-term memory network as a decoder to generate the question word by word, and embed the previous word w t-1 and the context vector c t into the long short-term memory network together to calculate the new hidden state s t , according to s t and the context vector c t calculate the probability p of generating a word gen , at the same time, according to the hidden state s t and the context vector c t calculate a copy switch g to determine whether to copy a word from the input sentence or generate a new word. According to the target loss function loss, use the backpropagation method to calculate the gradients of the parameters in the deep learning network model, and use the stochastic gradient descent method to update the parameters; Step B7: When the number of times the BLEU-4 score of the deep learning network model does not update reaches the given threshold or reaches the maximum number of iterations, terminate the training process of the deep learning network model.
2. The problem generation method based on knowledge enhancement and dual-graph interaction according to claim 1, wherein The specific steps of step B1 include the following steps: Step B11: Traverse the training dataset DR. After tokenizing the context sentences, answers, and questions in it and removing stop words, each training sample in DR is represented as ds=(s, a, q); where s is the text content of the context sentence, a is the content of the answer; q is the target question to be generated by the model; The context sentence s is represented as: wherein, is the i-th word in the context sentence, where i = 1, 2, …, m, and m is the number of words in the context sentence s; The answer a is represented as: Among them, is the i-th word in the answer, where i = 1, 2, …, n, and n is the number of words in the answer a; The target question q is represented as: Among them, is the i-th word in the target question, where i = 1, 2, …, l, and l is the number of words in the target question q; Step B12: Encode the context sentence obtained in step B11 to obtain an initial representation vector E of the context sentence cs ; Among them, E cs is expressed as: Among them, is the word vector corresponding to the i-th word obtained by looking up in the pre-trained word vector matrix where d represents the predefined vector dimension and |V| is the number of words in the dictionary V; Step B13: For the context sentence perform part-of-speech tagging and named entity recognition to obtain the continuous representation vector E of the part-of-speech tagging conversion cp and the continuous representation vector E of the named entity recognition conversion cn ; Among them, E cp is expressed as: wherein, is the part-of-speech vector corresponding to the i-th word, obtained by looking up in the part-of-speech vector matrix , where d1 represents the dimension of the part-of-speech vector, and |V1| is the number of part-of-speech in the vocabulary V1; E cn Expressed as: Among them, is the entity vector corresponding to the i-th word, which is obtained by looking up in the entity vector matrix where d2 represents the dimension of the entity vector and |V2| is the number of part-of-speech tags in the vocabulary V2; Step B14: Insert the [CLS] token in front of the context sentence s and the [SEP] token at the end to construct the sentence input sequence s1 for the BERT model; input s1 into the BERT model to obtain the output sequence of the i-th layer of the model The output sequence E of the last layer of the model b ; Among them, the sentence input sequence is represented as: Among them, the output sequence of the i-th layer of the model is expressed as: Among them, the output sequence E of the last layer of the model cb is expressed as: where is the BERT word embedding vector corresponding to the i-th word, obtained by looking up in the BERT word embedding vector matrix where d3 represents the dimension of the entity vector and |V3| is the number of words in the vocabulary V3; Step B15: Combine E obtained in step B12 cs , E obtained in step B13 cp and E cn , E obtained in step B14 cb to obtain the initial representation vector E of the final context sentence c ; Among them, E c is expressed as: where m is the number of words, is the initial representation vector of the i-th word, can be expressed as: Among them, [;] represents the concatenation operation; Step B16: Perform syntactic dependency parsing on the context text to generate an undirected syntactic dependency structure graph, and encode it into a corresponding m-order syntactic structure dependency adjacency matrix A c ; Among them, A c is represented as: Step B17: Generate question words. First, perform named entity recognition on the answer to obtain the corresponding question words. After obtaining the question words, splice them after the answer to obtain a new answer representation For the answer representation Perform encoding to obtain the initial representation vector E of the context sentence as ; At the answer representation Insert [CLS] at the front and [SEP] at the end to construct the answer input sequence of the BERT model Input into the BERT model to obtain the output sequence of the i-th layer of the model The output sequence E of the last layer of the model ab ; Finally, splice E as and E ab to obtain the final initial representation vector E of the answer a ; Among them, It is expressed as: Among them, is the original answer word, is the predicted interrogative word; E as Expressed as: Among them, is the word vector corresponding to the i-th word obtained by looking up in the pre-trained word vector matrix where d represents the predefined vector dimension and |V| is the number of words in the vocabulary V; Among them, the answer input sequence can be expressed as: Among them, the output sequence of the i-th layer of the model is expressed as: Among them, the output sequence E of the last layer of the model ab is expressed as: Among them, is the BERT word embedding vector corresponding to the i-th word, which is obtained by looking up in the BERT word embedding vector matrix where d3 represents the dimension of the entity vector, and |V3| is the number of words in the vocabulary V3; E a can be expressed as: where n + 1 is the number of words in the answer, is the initial representation vector of the i-th word, which is expressed as: Among them, [;] represents the concatenation operation; Step B18: Expand knowledge nodes from the knowledge graph based on the context syntactic dependency tree and the answer syntactic dependency tree. For a sentence, expand only one of nouns, adjectives, numerals, and adverbs. For the answer, since it is short, expand all words. For each node in the dependency tree, select five relationships in the knowledge graph, namely ISA, Synonym, SymbolOf, SimilarTo, and DefinedAs, for expansion. Sort the expanded knowledge nodes according to the weights given by the knowledge graph, and select the two words with the largest weight values for expansion. The expanded nodes and the original sentence nodes form a knowledge-enhanced context sentence graph, whose adjacency matrix is denoted as A ckg , and the expanded nodes and the original answer nodes form a knowledge-enhanced answer graph, whose adjacency matrix is denoted as A akg ; Map each node into a continuous low-dimensional vector using the knowledge embedding of the knowledge graph itself, and finally obtain the sentence knowledge embedding vector E ck and the answer knowledge embedding vector E ak ; Among them, E ck is expressed as: Among them, z is the total number of context sentence nodes after expansion, and d is the predefined vector dimension. is the knowledge embedding vector of the original context sentence. is the knowledge embedding representation of the expanded sentence node. Among them, A ckg is expressed as: Among them, 1 indicates that there is an edge connection between nodes, and 0 indicates that there is no edge connection; Among them, E ak is expressed as: Among them r is the total number of answer nodes after expansion, and d is the predefined vector dimension is the knowledge embedding vector of the original context sentence is the knowledge embedding representation of the expanded answer node Among them, A akg is expressed as: Among them, 1 indicates that there is an edge connection between nodes, 0 indicates no, and r is the total number of sentence nodes after expansion.
3. The problem generation method based on knowledge enhancement and dual-graph interaction according to claim 2, wherein The specific steps of step B2 include the following steps: Step B21: Input the initial representation vector E of the context sentence s c of the sequence into the forward layer and the backward layer of the bidirectional long short-term memory network respectively, and obtain the state vector sequence of the forward hidden layer and the state vector sequence of the backward hidden layer respectively, where f is the activation function; connect to obtain the context sentence representation vector [;] represents the vector connection operation; Step B22: The initial representation vector E of the answer a a sequence is respectively input into the forward layer and the backward layer of the bidirectional long short-term memory network to obtain the state vector sequence of the forward hidden layer and the state vector sequence of the backward hidden layer where f is the activation function; connecting to obtain the context sentence representation vector [;] represents the vector connection operation.
4. The problem generation method based on knowledge enhancement and dual-graph interaction according to claim 3, wherein The specific steps of step B3 include the following steps: Step B31: Embed the sentence knowledge vector E ck As the initial representation vector of the graph node, use the K-layer graph convolutional network to perform graph convolution operations on the adjacency matrix A ckg to update the graph node information; the update process of the hidden state of node i in the K-layer graph convolutional network is as follows: where \(k\in[0,k]\) represents the number of layers of the graph convolutional network, is the hidden state output by node \(i\) in the \(K\)-th layer network, \(\sigma\) is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector; Step B32: Concatenate the hidden states of the K-th layer graph convolutional network, and take the first m nodes to obtain a knowledge-enhanced context sentence representation G ckg , G ckg is expressed as follows: where m is the number of words in the original sentence, and d is the predefined vector dimension; Step B33: Embed the answer knowledge into vector E ak As the initial representation vector of the graph node, use the K-layer graph convolutional network to perform graph convolution operations on the adjacency matrix A akg to update the graph node information; the update process of the hidden state of node i in the K-layer graph convolutional network is as follows: where \(k\in[0, K - 1]\) represents the number of layers of the graph convolutional network, is the hidden state output by node \(i\) in the \(K\)-th layer network, \(\sigma\) is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector; Step B34: Concatenate the hidden states of the K-th layer graph convolutional network, and take the first n + 1 nodes to obtain the knowledge-enhanced answer representation G akg , G agkg is expressed as follows: where n + 1 is the number of words in the original answer, and d is the predefined vector dimension; Step B35: Take the sentence representation vector H obtained in Step B21 s and the knowledge-enhanced context sentence representation G ckg and input them into multi-head attention to obtain the enhanced context sentence representation vector H ckg , H ckg is represented as follows: H ckg = MultiHead(G ckg , H s , H s ) MultiHead(G ckg ,H s ,H s ) = Concat(head1, head2, …, head h )W3 Among them, MultiHead represents the multi-head attention mechanism, Concat() represents the concatenation operation, h is the number of parallel heads, is a trainable parameter matrix, head i represents the output function of the i-th attention function, i ∈ [1, h], and softmax() is the activation function; Step B36: Input the answer representation vector H obtained in Step B22 a and the knowledge-enhanced answer representation G akg into another multi-head attention to obtain the enhanced answer context representation vector H akg , H akg which is expressed as: H akg = MultiHead(G akg , H a , H a ) MultiHead(G akg ,H a ,H a ) = Concat(head1, head2, …, head h )W7 Among them, MultiHead represents the multi-head attention mechanism, Concat() represents the concatenation operation, h is the number of parallel heads, is a trainable parameter matrix, head i represents the output function of the i-th attention function, i ∈ [1, h], and softmax() is the activation function.
5. The problem generation method based on knowledge enhancement and dual-graph interaction according to claim 4, characterized in that The specific steps of step B4 include the following steps: Step B41: Take the context sentence representation vector H obtained in step B35 ckg and the Glove embedding E of the context sentence obtained in step B12 cs , the BERT embedding e obtained in step B14 cb and splice them to obtain the sentence representation H cgb , where H cgb is represented as follows: wherein d is a predefined vector dimension, [;] represents a vector concatenation operation, is the enhanced context sentence representation vector H ckg the i-th hidden state representation; Step B42: The answer representation vector h obtained in step B36 akg is concatenated with the Glove embedding E obtained in step B17 as and the BERT embedding E ab to obtain the answer representation H agb , where H agb is expressed as follows: where d is the predefined vector dimension, i ∈ [1, n + 1], and [;] represents the vector concatenation operation, is the enhanced answer context representation vector H akg the i-th hidden state representation of; Step B43: Input the sentence representation H cgb , the answer representation H agb and the answer representation vector H akg into the attention mechanism to obtain the context-aware answer representation H ac , where H ac is represented as follows: H ac = softmax(Relu(W 11 H cgb )Relu(W 12 H agb ) T )H akg Among them, softmax() and Relu() are activation functions, and W 11 , W 12 are learnable parameter matrices; Step B44: Concatenate the enhanced context sentence representation vector H ckg with H ac to obtain the representation vector h cka ; Among them, H cka is expressed as follows: Among them, are respectively the hidden state representations of the i-th word of H ckg and H ac ; Step B45: Input the sequence of the characterization vector H cka into the forward layer and the backward layer of the bidirectional long short-term memory network respectively, and obtain the state vector sequence of the forward hidden layer and the state vector sequence of the backward hidden layer respectively, where f is an activation function; connect them to obtain the context sentence characterization vector [;] represents the vector connection operation. 6. The problem generation method based on knowledge enhancement and dual-graph interaction according to claim 5, wherein The specific steps of step B5 include the following steps: Step B51: Use the context sentence representation vector H obtained in step B45 d as the initial representation vector of the graph node, and use the K-layer graph convolutional network to perform graph convolution operations on the adjacency matrix A obtained in step B16 c to update the graph node information; the update process of the hidden state of node i in the K-layer graph convolutional network is as follows: where \(k\in[0, K - 1]\) represents the number of layers of the graph convolutional network, is the hidden state output by node \(i\) in the \(K\)-th layer network, \(\sigma\) is the non-linear activation function Relu(), is a trainable parameter matrix, is the bias vector; Step B52: Connect the hidden states of the K-th layer graph convolutional network to obtain a structural information representation G of the context sentence cca , G cca is expressed as follows: Among them m is the number of words in the original sentence, and d is the predefined vector dimension; Step B53: Use the attention mechanism to calculate the semantic similarity score for the context sentence representation vector H d to obtain the edge weight matrix A w , and the calculation process is as follows: where m is the length of the context sentence, and d is the predefined vector dimension, is the trainable parameter matrix; Step B54: Use the context sentence representation vector H d as the initial representation vector of the graph node, and use a K-layer graph convolutional network to perform graph convolution operations on the adjacency matrix A w to update the graph node information; the update process of the hidden state of node i in the K-layer graph convolutional network is as follows: where \(k\in[0, K - 1]\) represents the number of layers of the graph convolutional network, is the hidden state output by node \(i\) in the \(K\)-th layer network, \(\sigma\) is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector; Step B55: Connect the hidden states of the K-th layer graph convolutional network to obtain the semantic information representation G of the context sentence wca , G wca is expressed as follows: where m is the number of words in the original sentence, and d is the predefined vector dimension; Step B56: Construct a virtual graph for information interaction between graphs, and represent a context sentence as Append the same context sentence as and then construct a graph, that is, construct an (m + m) × (m + m) adjacency matrix A v , which is represented as follows: Among them, (m + m) represents the sum of the lengths of two context sentences; Step B57: Represent the structural information G of the context sentence cca and the semantic information G of the context sentence wca are concatenated to obtain the representation vector G cw ; Among them, G cw is expressed as follows: Among them, are respectively the i-th word hidden state representations of G cca and G wca ; Step B58: Concatenate to obtain the representation vector G cw As the initial representation vector of the graph node, use the K-layer graph convolutional network to perform graph convolution operations on the adjacency matrix A v to update the graph node information; the update process of the hidden state of node i in the K-layer graph convolutional network is as follows: where \(k\in[0, K - 1]\) represents the number of layers of the graph convolutional network, is the hidden state output by node \(i\) in the \(K\)-th layer network, \(\sigma\) is the non-linear activation function Relu(), is the trainable parameter matrix, is the bias vector; Step B59: Concatenate the hidden states of the K-th layer graph convolutional network to obtain the final context sentence representation H e It is expressed as follows: Among them, m is the number of words in the original sentence, and d is the predefined vector dimension.
7. The problem generation method based on knowledge enhancement and dual-graph interaction according to claim 6, wherein The specific steps of step B6 include the following steps: Step B61: Use a long short-term memory network as a decoder to generate questions word by word; at decoding time step t, the long short-term memory network decoder reads the previous word embedding w t-1 and the context vector c t to jointly calculate a new hidden state representation s t : s t = LSTM(s t-1 , [w t-1 ; c t ) c t = ∑ i a t,i H e where s t-1 is the hidden state of the previous long short-term memory network, and W 18 is a learnable parameter matrix; Step B62: Calculate the probability p of generating a word from the long short-term memory network hidden representation s t and the context vector c t ; gen ; p gen = softmax(W 19 [s t ; c t + b6) where, softmax( ) is the activation function, [;] represents the concatenation operation, W 19 is a learnable parameter matrix, and b6 is a learnable bias vector; Step B63: To handle rare and unknown words, calculate g according to the pointer network to decide whether to copy words from the input sentence or generate a new word, where g is expressed as follows: g = σ(W 20 s t + W 21 c t + b7) where σ is a non-linear activation function, W 20 , W 21 are learnable parameter matrices, and b7 is a learnable bias vector; Step B64: Finally generate the problem probability p final , and its calculation formula is as follows: p final (y t |y < t; θ) = gp copy (y t , θ1) + (1 - g)p gen (y t , θ2) Among them, θ, θ1, θ2 are learnable network parameters; Step B65: Use cross-entropy as the loss function to calculate the loss value and punish the duplication phenomenon through covloss. Update the learning rate through the gradient optimization algorithm Adam, and use backpropagation to iteratively update the model parameters to train the model by minimizing the loss function; the calculation formula for minimizing the loss function L is as follows: L = -∑logp final (y t |y < t; θ) + βcovloss t covloss t = ∑ i min(a t,i , c t,i ) where p final (y t |y < t; θ) is the calculated probability of generating a problem word, β is a hyperparameter, a t,i is the attention vector representation of the i-th generated word at time step t calculated in step B61, c t,i is the context vector representation of the i-th generated word at time step t.
8. A question generation system based on knowledge enhancement and dual-graph interaction using the method according to any one of claims 1-7, characterized in that, It includes: A data collection module for collecting context sentences and corresponding question-answer pairs and constructing a training dataset for question generation; A text preprocessing module for preprocessing the training samples in the training dataset; A text encoding module for initially encoding all training samples in the training dataset to obtain the initial representation vector of the context sentence, the context sentence adjacency matrix, the initial representation vector of the answer, and at the same time expand the knowledge nodes of the context sentence and the answer text from the knowledge graph, construct a knowledge-enhanced context sentence and a knowledge-enhanced answer adjacency matrix, and query to obtain the knowledge-expanded representation vector of the context sentence and the answer; A network training module, which is used to input the initial representation vector of the context sentence, the context sentence adjacency matrix, the initial representation vector of the answer, the context sentence adjacency matrix, and the initial representation vector of the answer into a deep learning network to obtain the final representation vector of the user comment. The knowledge-enhanced context sentence, the knowledge-enhanced answer adjacency matrix, and the knowledge-expanded representation vectors of the context sentence and the answer are input into the deep learning network to obtain the final representation vector of the context sentence. The final representation vector is used to decode and generate questions, and the loss is calculated by comparing with the reference questions in the training dataset. The entire deep learning network model is trained with the goal of minimizing the loss to obtain a deep learning network model based on knowledge enhancement and dual-graph interaction network; A question generation module, which is used to process the input context sentence and answer using the trained deep learning network model to generate questions that are related to the answer, grammatically correct, and can be answered by the sentence.
Citation Information
Patent Citations
Automatic reply generation method for open domain dialogue system
CN110334190A
Comment sentiment analysis method and system based on knowledge graph and interaction graph convolution
CN114492459A