A Method and System for Answer Selection Based on Knowledge-Enhanced Graph Convolutional Network

Through the method of a knowledge-enhanced graph convolution network, combined with knowledge graph and syntactic structure, the problem that existing models fail to make full use of grammatical structure and knowledge graph are solved, and higher answer selection accuracy is achieved.

CN116028604BActive Publication Date: 2025-07-29FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211464352.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2025-07-29
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

The existing answer selection model fails to fully consider the grammatical structure dependence information between the question and the answer, limits the understanding of text semantic information, and the knowledge graph lacks context semantic correlation in question-and-answer matching, affecting the accuracy of answer selection.

Method used

Using a method based on knowledge-enhanced graph convolution network, the graph convolution neural network combines knowledge graph ConceptNet to perform text-knowledge matching and multi-hop node expansion of questions and answers, uses the syntactic structure to rely on the adjacency matrix and attention mechanism to enhance semantic representation, combines the BiGRU network for feature aggregation, and finally generates the answer correlation score through the softmax function.

Benefits of technology

The accuracy of answer selection is improved, and by comprehensively considering the grammatical structure and knowledge graph information of the questions and answers, the model's ability to understand text semantics is enhanced and the accuracy of answer selection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116028604B_ABST
    Figure CN116028604B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for answer selection based on a knowledge-enhanced graph convolutional network, including the following steps: Step A: Collect the questions and answer records of users in a Q&A platform, and label the true labels of each question-answer pair to construct a training set DS; Step B: Use the training data set DS and the knowledge graph ConceptNet to train a deep learning network model M based on a knowledge-enhanced graph convolutional neural network, and analyze the given question through this model to determine the correctness of the corresponding candidate answers; Step C: Input the user's question into the trained deep learning network model M and output the matching answer; Applying this technical solution is beneficial to improving the accuracy of answer selection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to an answer selection method and system based on a knowledge-enhanced graph convolutional network. Background Art

[0002] Answer Selection (Answer Selection) is an important subtask in the field of question answering and plays a very important role in many applications of information retrieval (IR) and natural language processing (NLP). With the rapid development of the Internet, a large number of question-and-answer communities have emerged on the Internet, such as: Zhihu, Quora, Stack Overflow, etc. People are keen to ask questions and obtain answers in question-and-answer communities. With the long-term and extensive participation of users, a huge amount of question-answer data pairs have been generated on the Internet. Along with the explosion of information, it has become difficult to continue filtering and screening the information in the question-and-answer system by manual means; at the same time, due to the sharp increase in network information in the question-and-answer system, the questions raised by current users in the question-and-answer system are often overwhelmed by the constantly emerging new questions and cannot get a quick response. Therefore, there is an urgent need for an automated method that can effectively perform answer selection, judge the matching relationship between questions and numerous candidate answers, select the best answer from them and rank it as high as possible in the answer list.

[0003] With the continuous in - depth research on deep learning methods, many researchers have also applied deep learning models to the field of answer selection. Question - answering matching models based on deep learning usually rely on convolutional neural networks (CNNs), recurrent neural networks (RNNs), graph neural networks (GNNs), or pre - trained language models that incorporate attention mechanisms. CNNs are used to obtain local semantic information of question and answer texts. RNNs can build semantic dependencies of text sequences. The attention mechanism enables the model to pay more attention to the key semantic parts in the question - answer pair. GNNs can abstract the question - answer pair into a graph data structure according to the text relationships between different words, such as syntactic relationships, and model the dependencies between graph nodes. The emergence of pre - trained language models has greatly promoted the development of the natural language processing field. Pre - trained language models can learn latent semantic information from a large amount of unlabeled text. Some researchers have carried out research on applying pre - trained language models to the answer - selection task. Devlin et al. proposed a general model BERT for natural language processing trained based on the Transformer architecture and applied it to the answer - selection task. However, existing answer - selection models, whether based on neural networks or pre - trained language models, mainly focus on obtaining feature representations of the contextual semantic association information between words in the question and answer texts, and do not fully consider mining the dependency information between questions and answers from the perspective of syntactic structure, which limits the model's understanding of text semantic information.

[0004] In addition, some research work has introduced knowledge graphs into the answer - selection task and also made certain progress. The factual background in the knowledge graph contains a large amount of entity information, which can provide effective commonsense reasoning information during the question - answering matching process and improve the accuracy of answer selection. Li and Wu et al. proposed a word - net - enhanced hierarchical model, which uses synsets and hypernyms in WordNet to enhance the word - embedding representation in the question - answer sentences and designed two attention mechanisms based on the relationship scores of synsets and hypernyms to capture richer question - answer interaction information. However, in some existing answer - selection models, although knowledge graphs are introduced, there is a lack of contextual semantic association between knowledge entities and the entity information is not effectively guided to help the model learn the correct semantic representation in different contexts, which limits the improvement of the performance of the answer - selection model. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide an answer - selection method and system based on a knowledge - enhanced graph convolutional network, which is beneficial to improving the accuracy of selecting the correct answer.

[0006] To achieve the above - mentioned purpose, the present invention adopts the following technical solutions: An answer - selection method based on a knowledge - enhanced graph convolutional network, comprising the following steps:

[0007] Step A: Collect the questions and answer records of users in the Q&A platform, and label the true labels of each question-answer pair to construct the training set DS;

[0008] Step B: Use the training dataset DS and the knowledge graph ConceptNet to train the deep learning network model M based on the knowledge-enhanced graph convolutional neural network, and analyze the correctness of the corresponding candidate answers for the given questions through this model;

[0009] Step C: Input the user's question into the trained deep learning network model M and output the matching answer.

[0010] In a preferred embodiment, step B specifically includes the following steps:

[0011] Step B1: Perform initial encoding on all training samples in the training dataset DS to obtain the initial features E of the question and answer text content q , E a , the global semantic feature sequence E of the question-answer pair cls , the syntactic structure dependency adjacency matrix A of the question-answer pair. At the same time, perform text-knowledge matching and multi-hop knowledge node extension query on the question and answer text from the knowledge graph ConceptNetc, connect the text-matched knowledge nodes and extended nodes to obtain the knowledge extension sequence, and map the information of each knowledge node in the knowledge extension sequence into continuous low-dimensional vectors to finally form the knowledge extension sequence feature C of the question and answer q , C a ;

[0012] Step B2: Connect the initial features E of the question and answer text content q , E a to obtain the text feature E of the question-answer pair qa . Through mask calculation on E qa , obtain the question-answer edge weight matrix M a . Multiply M a by the syntactic structure dependency adjacency matrix A to obtain the syntactic structure dependency adjacency matrix with edge association weights

[0013] Step B3: Input the text feature E of the question-answer pair obtained in step B2 qa and the syntactic structure dependency adjacency matrix with edge association weights into a K-layer graph convolutional network, and guide the propagation of node information through the syntactic structure dependency relationship between graph nodes to learn the text feature of the question-answer pair Then, for the semantic representation E of the question-answer pair qa and the original structural information feature of the question-answer text Semantic enhancement is carried out in an attention-based manner to ensure the accuracy of node semantic information, and the semantic structure information features of the question-answer are obtained.

[0014] Step B4: The initial features E q , E a of the question and answer text content obtained in step B1 q , C a , and the knowledge extension sequence features C of the question and answer are input into two attention calculation mechanisms guided by text semantics to obtain the semantic guidance knowledge features of the question q and the answer a. Then the semantic guidance knowledge representation is input into two multi-head self-attention mechanisms to obtain the self-attention knowledge representation. The semantic guidance knowledge features

[0015] and the self-attention knowledge features q , H a are input into two feed-forward neural network layers to obtain the context features H q , H a of the knowledge; the context features H qa of the knowledge are filtered and fused using a gating mechanism to obtain the context features H

[0016] of the question-answer; qa Step B5: The context features H of the question-answer and the semantic structure information features of the question-answer are fused by means of attention calculation to obtain the semantic structure information features cls of the knowledge-enhanced question-answer pair. Then the local semantic feature matrix E

[0017] obtained in step B1 is input into a multi-size convolutional neural network to obtain a multi-granularity global semantic feature representation. Step B6: The semantic structure information features of the knowledge-enhanced question-answer pair are input into a BiGRU network, and an average pooling operation is performed on the sequence of the hidden state outputs of the BiGRU to obtain the aggregated features of the question-answer pair. The aggregated features of the question-answer pair and the multi-granularity global semantic feature representation final are concatenated to obtain the final Q&A feature E finalInput it into a linear classification layer and normalize it using the softmax function to generate the correlation score f(q, a) ∈ [0, 1] between the question and the answer; then, according to the objective loss function loss, calculate the gradients of the parameters in the deep learning network model through backpropagation, and update the parameters using the stochastic gradient descent method;

[0018] Step B7: When the change in the loss value generated by the deep learning network model in each iteration is less than the given threshold or reaches the maximum number of iterations, terminate the training process of the deep learning network model.

[0019] In a preferred embodiment, step B1 specifically includes the following steps:

[0020] Step B11: Traverse the training set DS. After tokenizing the questions and candidate answer texts in it and removing the stop words, each training sample in DS is represented as ds = (q, a, p); where q is the text content of the question, a is the content of the candidate answer corresponding to the question; p is the correct or incorrect label corresponding to the question-answer pair, p ∈ [0, 1], 0: the candidate answer is the wrong answer, 1: the candidate answer is the correct answer;

[0021] The question q is represented as:

[0022]

[0023] where, is the i-th word in the question q, i = 1, 2,..., m, and m is the number of words in the question q;

[0024] The answer a is represented as:

[0025]

[0026] where, is the i-th word in the answer a, i = 1, 2,..., n, and n is the number of words in the answer a;

[0027] Step B12: Concatenate the question and the answer obtained in step B11, insert the [CLS] tag in front of the question q, and insert the [SEP] tags before and after the answer a to construct the question-answering input sequence X s ;

[0028] The question-answering input sequence can be represented as:

[0029]

[0030] Among them, m and n respectively represent the number of words in question q and answer a;

[0031] Step B13: Input X s into the BERT model to obtain the output sequence of the i-th layer of the model The output sequence E of the last layer of the model s ; According to the positions of the [CLS] and [SEP] tags in the E s sequence, segment the initial representation vectors of the question and the answer, so as to obtain the initial representation vectors E q and E a ; Connect the [CLS] token in cls to obtain the global semantic feature E of the question and the answer;

[0032] Among them, the output sequence of the i-th layer of the model is expressed as:

[0033]

[0034] Among them, the output sequence E of the last layer of the model s is expressed as:

[0035]

[0036] The initial feature E of question q q is expressed as:

[0037]

[0038] Where is the word vector corresponding to the i-th word , m is the length of the question sequence, and d is the dimension of the word vector;

[0039] The initial feature E of answer a a is expressed as:

[0040]

[0041] Where is the word vector corresponding to the i-th word , n is the length of the answer sequence, and d is the dimension of the word vector;

[0042] The global semantic feature E of the question and the answer cls is expressed as:

[0043]

[0044] Where where is the [CLS] token output by the i-th layer model, l1 is the number of encoder layers of BERT, and d is the dimension of the [CLS] vector;

[0045] Step B14: Concatenate the question text and the answer text to obtain the question-answer text sequence Perform syntactic dependency parsing on the question-answer text sequence X qa to generate an undirected syntactic structure dependency graph and encode it into a corresponding (m + n)-order syntactic structure dependency adjacency matrix A;

[0046] where the representation of A is:

[0047]

[0048]

[0049] Step B15: Perform question text-knowledge matching and multi-hop node expansion for each word in the question q and the answer a in the knowledge graph ConceptNet; first, perform text-knowledge matching for each word in the question q in the knowledge graph to obtain its corresponding knowledge node Similarly, the corresponding knowledge nodes for each word in the answer a can be obtained Secondly, in the process of multi-hop expanding knowledge nodes, perform multi-hop node selection according to the relationship between the text-matched knowledge nodes and the nodes in the knowledge graph; sort the multi-hop selected knowledge nodes according to their initial weights in the knowledge graph, and select the max_n extended knowledge nodes with the largest weights; connect the extended nodes and the text-matched knowledge nodes to form a knowledge expansion sequence; use knowledge embedding to map each knowledge node in the knowledge expansion sequence into a continuous low-dimensional vector, and finally form the knowledge expansion sequence feature C of the question q and the answer a , C q , C a ;

[0050] where the knowledge expansion sequence feature C of the question q q is represented as:

[0051]

[0052] where, l2 = (m + max_n × m) is the length of the question knowledge expansion sequence, and d is the dimension of the knowledge word vector; is The extended knowledge nodes, where max_n is the number of extended nodes;

[0053] Answer a knowledge expansion sequence feature C a Expressed as:

[0054]

[0055] Among them, l3 = (n + max_n × n) is the length of the answer knowledge expansion sequence, and d is the dimension of the knowledge word vector; is The extended knowledge nodes, where max_n is the number of extended nodes;

[0056] In a preferred embodiment, step B2 specifically includes the following steps:

[0057] Step B21: The initial features of the question and answer text content Are connected to obtain the text feature of the question-answer Among them m + n is the length of the question-answer text sequence, and d is the dimension of the word vector;

[0058] Step B22: Perform masked edge weight calculation on the question-answer text feature E obtained in step B21 qa To obtain the edge weight value matrix M a , and its calculation process is as follows:

[0059]

[0060] Among them m + n is the length of the sequence X qa of qa vector, W1 and W2 are trainable parameter matrices;

[0061] Step B23: Perform a dot product operation on the edge weight value matrix M a and the syntactic structure dependency adjacency matrix A obtained in step B14 to obtain the syntactic structure dependency adjacency matrix with edge weights Its calculation process is as follows:

[0062]

[0063] Among them, ⊙ is the matrix-by-element dot product operation.

[0064] In a preferred embodiment, step B3 specifically includes the following steps:

[0065] Step B31: Take the text feature E of the question-answer qa as the initial representation vector of the graph node, and use the K-layer graph convolutional network to perform graph convolution operations on the adjacency matrix to update the graph node information; the update process of the hidden state of node i in the k-th layer of the graph convolutional network is as follows:

[0066]

[0067]

[0068] where k ∈ [1, K], representing the number of layers of the graph convolutional network, is the hidden state output by node i in the k-th layer, Relu() is a non-linear activation function, is a trainable parameter matrix, is a bias vector, d i represents the dimension of the initial representation vector of node i;

[0069] Step B32: Connect the hidden states of the K-th layer of the graph convolutional network to obtain the original structural information feature of the question-answer which is expressed as follows:

[0070]

[0071] where, m + n is the length of the question-answer text sequence, and d is the dimension of the initial representation vector of the node;

[0072] Step B33: Semantically enhance the text feature E of the question-answer qa and the original structural information feature of the question-answer in an attention calculation manner to obtain the semantic structural information feature of the question-answer The calculation formula is as follows:

[0073]

[0074]

[0075] where, m + n is the length of the question-answer text sequence, d is the dimension of the initial representation vector of the node, W4 and W5 are trainable parameter matrices.

[0076] In a preferred embodiment, the step B4 specifically includes the following steps:

[0077] Step B41: Take the initial features E q , E aand the knowledge expansion sequence feature C of questions and answers is obtained in step B15 q , C a , which is input into two attention calculation mechanisms guided by text semantics to obtain the semantic guidance features of question q and answer a

[0078] where The calculation formula is as follows:

[0079]

[0080]

[0081] where l2 is the length of the knowledge expansion sequence feature C q ; W6 and W7 are trainable parameter matrices; similarly, the semantic guidance knowledge representation of the answer can be obtained

[0082] Step B42: The semantic guidance knowledge representations of question q and answer a are respectively input into two different multi-head attention mechanisms to obtain the self-attention knowledge features of the question and the answer

[0083] where The calculation formula is as follows:

[0084]

[0085]

[0086] where MHA represents the multi-head attention mechanism, num is the number of parallel heads, and Q (query), k (key), and V (value) are all semantic guidance question knowledge features are trainable parameter matrices,, head i represents the output of the i-th attention function, i ∈ [1, num]; similarly, the self-attention knowledge feature of the answer is obtained

[0087] Step B43: The self-attention knowledge features of the question and the answer and the semantic guidance knowledge features are input into two linear feed-forward layer networks for fusion to obtain the context feature H of knowledge q , H a ;

[0088] where H qThe calculation formula is as follows:

[0089]

[0090] Among them, is a trainable parameter matrix, is a bias vector;

[0091] Step B45: Input the knowledge context features H q and H a of the question and answer into a gating mechanism for filtering and fusion, so as to suppress knowledge noise and obtain the knowledge context feature H qa ;

[0092] Among them, the calculation formula of H qa is as follows:

[0093] g = sigmoid(H q W 15 : H a W 16 )

[0094] H qa = (1 - g) ⊙ H q + g t ⊙ H a

[0095] Among them l2 is the length of C q and l3 is the length of C a ; is a trainable parameter, and ":" is a concatenation operation.

[0096] In a preferred embodiment, step B5 specifically includes the following steps:

[0097] Step B51: Perform knowledge enhancement on the knowledge context feature H qa of the question-answer pair and the semantic structure information feature of the question-answer pair in the way of attention calculation to obtain the semantic structure information feature of the knowledge-enhanced question-answer pair. The calculation formula is as follows:

[0098]

[0099]

[0100] Among them, m + n is the length of the text sequence X qa of the question-answer pair, is a trainable parameter;

[0101] Step B52: Input the global semantic feature Ec obtained in Step B1 ls into a multi-scale convolutional neural network to obtain a multi-granularity global semantic feature representation which is expressed as:

[0102]

[0103] where MCNN() represents the multi-scale CNN.

[0104] In a preferred embodiment, Step B6 specifically includes the following steps:

[0105] Step B61: Input the semantic structure information feature of the knowledge-enhanced question-answer pair into the forward layer and the backward layer of a bidirectional GRU network to respectively obtain the state vector sequence of the forward hidden layer and the state vector sequence of the backward hidden layer where

[0106] Step B62: Concatenate and and pass through a linear layer to obtain the output sequence E of the BiGRU of the question-answer pair gru ; Perform average pooling on E gru to obtain the aggregated feature of the question-answer The calculation formula is as follows:

[0107]

[0108]

[0109] where is a trainable parameter, meanpool() is the average pooling function;

[0110] Step B63: Connect the aggregated feature of the question-answer and the multi-granularity global semantic feature representation to obtain the final question-answer feature representation Ef final ; Ef final is expressed as follows:

[0111]

[0112] Step B64: Use the final question-answer feature Ef inalIt is input into a linear classification layer and normalized using the softmax function to generate the relevance score f(q, a) ∈ [0, 1] between the question and the answer. The calculation formula is as follows:

[0113] f(q, a) = softamx(E final W 19 + b4)

[0114] Among them, is a trainable parameter matrix, is a bias vector;

[0115] Step B65: Use cross-entropy as the loss function to calculate the loss value. Update the learning rate through the gradient optimization algorithm Adam, and use backpropagation to iteratively update the model parameters to train the model by minimizing the loss function. The calculation formula for minimizing the loss function L is as follows:

[0116]

[0117] Among them, f(q, a) i ∈ [0, 1] is the relevance score between the question and the answer calculated by the softmax classifier, and y i ∈ [0, 1] is the binary classification label.

[0118] The present invention also provides an answer selection system based on a knowledge-enhanced graph convolutional network. The system implements the above-mentioned answer selection method based on a knowledge-enhanced graph convolutional network, including:

[0119] A data collection module that collects users' questions and answer records on a Q&A platform, and annotates the true labels of each question-answer pair to construct a training set DS;

[0120] A text preprocessing module for preprocessing the training samples in the training set, including word segmentation and stop word removal;

[0121] A text encoding module that initially encodes all the training samples in the training data set DS to obtain the initial features of the question and answer text content, the global semantic feature sequence of the question-answer pair, the syntactic structure dependency adjacency matrix of the question-answer pair, and at the same time performs text-knowledge matching and multi-hop knowledge node extension queries on the question and answer text from the knowledge graph ConceptNetc to obtain the knowledge extension sequence features of the question and answer;

[0122] The network model training module is used to input the initial features of questions and answer texts, the global semantic feature sequence of question-answer pairs, the syntactic structure dependency adjacency matrix of question-answer pairs, and the knowledge extension sequence features of questions and answers into a deep learning network to obtain the final representation vector of question-answer pairs. The probability of answer correctness is predicted using this representation vector, and the loss is calculated by comparing with the true category annotations in the training set. The entire deep learning network is trained with the goal of minimizing the loss to obtain a deep learning network model based on a knowledge-enhanced graph convolutional network;

[0123] The answer selection module selects a correct answer for a given question, analyzes and processes the input question using the deep learning network model of the knowledge-enhanced graph convolutional network, and outputs the candidate answer with the highest question-answer pair correlation score, indicating the correct answer selected for the question.

[0124] Compared with the prior art, the present invention has the following beneficial effects: it is beneficial to improve the accuracy of selecting the correct answer. BRIEF DESCRIPTION OF THE DRAWINGS

[0125] Figure 1 is the flowchart of the method implementation of the preferred embodiment of the present invention;

[0126] Figure 2 is the model architecture diagram in the preferred embodiment of the present invention;

[0127] Figure 3 is the schematic diagram of the system structure of the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0128] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0129] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0130] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0131] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0132] Such as Figures 1-3As shown in the figure, this embodiment provides an answer selection method based on a knowledge-enhanced graph convolutional network, including the following steps:

[0133] Step A: Collect the questions and answer records of users on the Q&A platform, and label the true labels of each question-answer pair to construct a training set DS;

[0134] Step B: Use the training data set DS and the knowledge graph ConceptNet to train a deep learning network model M based on a knowledge-enhanced graph convolutional neural network, and use this model to analyze the correctness of the given question and the corresponding candidate answers;

[0135] Step C: Input the user's question into the trained deep learning network model M and output the matching answer. This method and system are beneficial to improving the accuracy of answer selection;

[0136] In this embodiment, the specific steps of step B include the following steps:

[0137] Step B1: Perform initial encoding on all training samples in the training data set DS to obtain the initial features E of the question and answer text content q , E a , the global semantic feature sequence E of the question-answer pair cls , the syntactic structure dependency adjacency matrix A of the question-answer pair. At the same time, perform text-knowledge matching and multi-hop knowledge node extension query on the question and answer text from the knowledge graph ConceptNetc, connect the text-matched knowledge nodes and extended nodes to obtain a knowledge extension sequence, and map each knowledge node information in the knowledge extension sequence to a continuous low-dimensional vector, and finally form the knowledge extension sequence features C of the question and answer q , C a ; The specific steps of step B1 include the following steps:

[0138] Step B11: Traverse the training set DS. After tokenizing the questions and candidate answer texts in it and removing stop words, each training sample in DS is represented as ds=(q, a, p); where q is the text content of the question, a is the content of the candidate answer corresponding to the question; p is the correct or incorrect label corresponding to the question-answer pair, p∈[0, 1], 0: the candidate answer is the wrong answer, 1: the candidate answer is the correct answer;

[0139] The question q is expressed as:

[0140]

[0141] Among them, is the i-th word in the question q, i = 1, 2,..., m, and m is the number of words in the question q;

[0142] Answer a is expressed as:

[0143]

[0144] where is the i-th word in answer a, i = 1, 2,..., n, and n is the number of words in question a:

[0145] Step B12: Concatenate the question and the answer , insert the [CLS] token in front of question q, and insert the [SEP] tokens before and after answer a to construct the question-answering input sequence X for the BERT encoding model s ;

[0146] The question-answering input sequence can be expressed as:

[0147]

[0148] where m and n respectively represent the number of words in question q and answer a;

[0149] Step B13: Input X s into the BERT model to obtain the output sequence of the i-th layer of the model The output sequence E of the last layer of the model s ; According to the positions of the [CLS] and [SEP] tags in E s sequence, slice the initial representation vectors of the question and the answer to respectively obtain the initial representation vectors E q and E a ; Connect the [CLS] token in cls ;

[0150] where the output sequence of the i-th layer of the model is expressed as:

[0151]

[0152] where the output sequence E s of the last layer of the model is expressed as:

[0153]

[0154] The initial feature E q of question q is expressed as:

[0155]

[0156] where is the i-th word corresponding word vector, m is the length of the question sequence, and d is the dimension of the word vector.

[0157] Initial feature E of question a a is expressed as:

[0158]

[0159] where is the i-th word corresponding word vector, n is the length of the answer sequence, and d is the dimension of the word vector.

[0160] Global semantic feature E of question and answer cls is expressed as:

[0161]

[0162] where is the [CLS] token output by the i-th layer of the model, l1 is the number of encoder layers of BERT, and d is the dimension of the [CLS] vector.

[0163] Step B14: Concatenate the question text and the answer text to obtain the question-answer text sequence Perform syntactic dependency parsing on the question-answer text sequence X qa to generate an undirected syntactic structure dependency graph and encode it into a corresponding (m + n)-order syntactic structure dependency adjacency matrix A;

[0164] where the representation of A is:

[0165]

[0166]

[0167] Step B15: Perform question text-knowledge matching and multi-hop node expansion for each word in the question q and the answer a in the knowledge graph ConceptNet. First, for each word in the question q perform text-knowledge matching in the knowledge graph to obtain its corresponding knowledge node Similarly, the corresponding knowledge node of each word in the answer a can be obtained Secondly, in the process of multi-hop expanding the knowledge nodes, according to the text matching knowledge nodes ​Perform multi-hop node selection for the relationships between nodes in the knowledge graph; sort the knowledge nodes selected by multi-hop according to their initial weights in the knowledge graph, and select the max_n extended knowledge nodes with the largest weights from them. Connect the extended nodes with the text-matching knowledge nodes to form a knowledge extension sequence. Map each knowledge node in the knowledge extension sequence into a continuous low-dimensional vector using knowledge embedding, and finally form the knowledge extension sequence feature C of the question q and the answer a

[0168] . The extended nodes and the text-matching knowledge nodes are connected to form a knowledge extension sequence. Use knowledge embedding to map each knowledge node in the knowledge extension sequence into a continuous low-dimensional vector, and finally form the knowledge extension sequence feature C of the question q and the answer a q , C a ;

[0169] Among them, the knowledge extension sequence feature C of the question q q is expressed as:

[0170]

[0171] Among them, l2 = (m + max_n × m) is the length of the question knowledge extension sequence, and d is the dimension of the knowledge word vector. is the extended knowledge node of, and max_n is the number of extended nodes.

[0172] The knowledge extension sequence feature C of the answer a a is expressed as:

[0173]

[0174] Among them, l3 = (n + max_n × n) is the length of the answer knowledge extension sequence, and d is the dimension of the knowledge word vector. is the extended knowledge node of, and max_n is the number of extended nodes.

[0175] Step B2: Connect the initial features E q , E a of the question and answer text content to obtain the text feature E qa of the question-answer. Through mask calculation on E qa , obtain the question-answer edge weight matrix M a . Multiply M a with the syntactic structure dependency adjacency matrix A to obtain the syntactic structure dependency adjacency matrix with edge-associated weights The specific steps of the said Step B2 include the following steps:

[0176] Step B21: Connect the initial features of the question and answer text content to obtain the text feature Among them m + n is the length of the question-answer text sequence, and d is the dimension of the word vector;

[0177] Step B22: Perform masked edge weight calculation on the text feature E of the question-answer obtained in B21 qa to obtain the edge weight value matrix M a , and its calculation process is as follows:

[0178]

[0179] Among them m + n is the length of the sequence X qa and d is the dimension of the E qa vector, W1 and W2 are trainable parameter matrices;

[0180] Step B23: Perform a dot product operation on the edge weight value matrix M a and the syntactic structure dependency adjacency matrix A obtained in step B14 to obtain the syntactic structure dependency adjacency matrix with edge weights and its calculation process is as follows:

[0181]

[0182] Among them, ⊙ is the matrix-by-element dot product operation;

[0183] Step B3: Input the text feature E of the question-answer obtained in step B2 qa and the syntactic structure dependency adjacency matrix with edge association weights into a K-layer graph convolutional network, and guide the propagation of node information through the syntactic structure dependency relationship between graph nodes to learn the original structural information features of the question-answer text Then, perform semantic enhancement on the text feature E of the question-answer qa and the original structural information features of the question-answer text in an attention manner to ensure the accuracy of node semantic information, and obtain the semantic structure information features of the question-answer The specific steps of step B3 are as follows:

[0184] Step B31: Use the text feature E of the question-answer qa as the initial representation vector of the graph nodes, and use the K-layer graph convolutional network to perform graph convolution operations on the adjacency matrix to update the graph node information. The update process of the hidden state of node i in the k-th layer graph convolutional network is as follows:

[0185]

[0186]

[0187] where \(k\in[1,K]\), representing the number of layers of the graph convolutional network, is the hidden state output by node \(i\) in the \(k\)-th layer network, and \(Relu()\) is a non-linear activation function, is a trainable parameter matrix, is a bias vector, \(d\) i represents the dimension of the initial representation vector of node \(i\).

[0188] Step B32: Concatenate the hidden states of the \(K\)-th layer graph convolutional network to obtain the original structural information features of the question-answer It is expressed as follows:

[0189]

[0190] where, \(m + n\) is the length of the question-answer text sequence, and \(d\) is the dimension of the initial representation vector of the node:

[0191] Step B33: Semantically enhance the text features \(E\) qa of the question-answer and the original structural information features of the question-answer in an attention calculation manner to obtain the semantic structural information features of the question-answer. The calculation formula is as follows:

[0192]

[0193]

[0194] where, \(m + n\) is the length of the question-answer text sequence, \(d\) is the dimension of the initial representation vector of the node, \(W4\), \(W5\) are trainable parameter matrices;

[0195] Step B4: Input the initial features \(E\) q , \(E\) a of the question and answer text content obtained in Step B1, and the knowledge expansion sequence features \(C\) q , \(C\) a into two attention calculation mechanisms guided by text semantics to obtain the semantic-guided knowledge features of the question \(q\) and the answer \(a\). Then input the semantic-guided knowledge features into two multi-head self-attention mechanisms to obtain the self-attention knowledge representation To ensure that the semantic features of the knowledge entity itself are not lost, the semantic-guided knowledge representation and self-attention knowledge features are input into two feed-forward neural network layers to obtain the context features H of knowledge q , H a ; the context features H of knowledge q , H a are filtered and fused using a gating mechanism to obtain the context features H of question-answer knowledge qa ; step B4 specifically includes the following steps:

[0196] Step B41: The initial features E q , E a of the question and answer text content obtained in step B13 q , C a and the knowledge expansion features C

[0197] where the calculation formula is as follows:

[0198] α q = softmax(tanh(E q W6 × (C q W7) T ))

[0199]

[0200] where l2 is the length of the knowledge expansion sequence feature C q , W6, W7 are trainable parameter matrices. Similarly, the semantic-guided knowledge representation of the answer can be obtained

[0201] Step B42: The semantic-guided knowledge representations of question q and answer a

[0202] are respectively input into two different multi-head attention mechanisms to obtain the self-attention knowledge features of question and answer

[0203]

[0204]

[0205] ​Among them, MHA represents the multi-head attention mechanism, num is the number of parallel heads, and Q (query), k (key), and V (value) are all semantic-guided question knowledge features are trainable parameter matrices,, head i represents the output of the i-th attention function, i ∈ [1, num]; similarly, the self-attention knowledge features of the answer can be obtained

[0206] Step B43: Input the self-attention knowledge features of the question and the answer and the semantic-guided knowledge features into two linear feed-forward layer networks for fusion to obtain the context features H of the knowledge q , H a ;

[0207] where the calculation formula of H q is as follows:

[0208]

[0209] Among them, is a trainable parameter matrix, is a bias vector;

[0210] Step B45: Input the context features H of the question and the answer q , H a into a gating mechanism for filtering and fusion, so as to suppress knowledge noise and obtain the context features H of the question-answer qa ;

[0211] where the calculation formula of H qa is as follows:

[0212] g = sigmoid(H q W 15 : H a W 16 )

[0213] H qa = (1 - g) ⊙ H q + g t ⊙ H a

[0214] Among them l2 is the length of C q l3 is the length of C a length. is a trainable parameter, and ":" is the concatenation operation.

[0215] Step B5: Combine the knowledge context feature H of the question-answer pair qa and the semantic structure information feature of the question-answer pair using attention calculation to obtain the semantic structure information feature of the knowledge-enhanced question-answer pair Then, input the local semantic feature matrix E obtained in Step B1 cls into a multi-scale convolutional neural network to obtain a multi-granularity global semantic feature representation Step B5 specifically includes the following steps:

[0216] Step B51: Enhance the knowledge of the knowledge context feature H of the question-answer pair qa and the semantic structure information feature of the question-answer pair using attention calculation to obtain the semantic structure information feature of the knowledge-enhanced question-answer pair The calculation formula is as follows:

[0217]

[0218]

[0219] where m + n is the length of the text sequence X of the question-answer pair qa and are trainable parameters

[0220] Step B52: Input the global semantic feature E obtained in Step B1 cls into a multi-scale convolutional neural network to obtain a multi-granularity global semantic feature representation It is expressed as:

[0221]

[0222] where MCNN() represents a multi-scale CNN

[0223] Step B6: Input the semantic structure information feature of the knowledge-enhanced question-answer pair into a BiGRU network and perform average pooling on the sequence of the hidden state outputs of the BiGRU to obtain the aggregated feature of the question-answer pair Concatenate the aggregated feature of the question-answer pair and the multi-granularity global semantic feature representation to obtain the final Q&A feature E final ; Subsequently, use the final Q&A feature E finalInput it into a linear classification layer and normalize it using the softmax function to generate the correlation score f(q, a) ∈ [0, 1] between the question and the answer; then, according to the objective loss function loss, calculate the gradients of the parameters in the deep learning network model through the backpropagation method, and update the parameters using the stochastic gradient descent method; the specific steps of step B6 are as follows:

[0224] Step B61: Input the semantic structure information features of the knowledge-enhanced question-answer pairs into the forward layer and the backward layer of a bidirectional GRU network, and respectively obtain the state vector sequence of the forward hidden layer and the state vector sequence of the backward hidden layer where

[0225] Step B62: Concatenate and and pass through a linear layer to obtain the output sequence E of the BiGRU of the question-answer pair gru ; perform average pooling on E gru to obtain the aggregated feature of the question-answer pair The calculation formula is as follows:

[0226]

[0227]

[0228] where is a trainable parameter, meanpool() is the average pooling function;

[0229] Step B63: Connect the aggregated feature of the question-answer pair and the multi-granularity global semantic feature representation to obtain the final question-answer feature representation E final ; E final is represented as follows:

[0230]

[0231] Step B64: Input the final question-answer feature Ef inal into a linear classification layer and normalize it using the softmax function to generate the correlation score f(q, a) ∈ [0, 1] between the question and the answer. The calculation formula is as follows:

[0232] f(q, a) = softamx(E final W 19 + b4)

[0233] Among them, is a trainable parameter matrix, is a bias vector:

[0234] Step B65: Use cross-entropy as the loss function to calculate the loss value, update the learning rate through the gradient optimization algorithm Adam, and iteratively update the model parameters using backpropagation to train the model by minimizing the loss function; the calculation formula for minimizing the loss function L is as follows:

[0235]

[0236] where f(q, a) i ∈ [0, 1] is the relevance score of the question-answer calculated by the softmax classifier, and y i ∈ [0, 1] is the binary classification label.

[0237] Step B7: When the change in the loss value generated by each iteration of the deep learning network model is less than the given threshold or the maximum number of iterations is reached, terminate the training process of the deep learning network model.

[0238] As Figure 3 shown, this embodiment provides a rumor answer selection system for implementing the above method, including:

[0239] A data collection module that collects users' questions and answer records on the Q&A platform and annotates the true labels of each question-answer pair to construct a training set DS.

[0240] A text preprocessing module for preprocessing the training samples in the training set, including word segmentation, stop word removal, etc.;

[0241] A text encoding module that performs initial encoding on all training samples in the training dataset DS to obtain the initial features of the question and answer text content, the global semantic feature sequence of the question-answer pair, the syntactic structure dependency adjacency matrix of the question-answer pair, and at the same time performs text-knowledge matching and multi-hop knowledge node extension query on the question and answer text from the knowledge graph ConceptNetc to obtain the knowledge extension sequence features of the question and answer;

[0242] The network model training module is used to input the initial features of the question and answer text, the global semantic feature sequence of the question and answer pair, the syntactic structure dependency adjacency matrix of the question-answer pair, and the knowledge expansion sequence features of the question and answer into a deep learning network to obtain the final representation vector of the question and answer pair. The probability of the correctness of the answer is predicted using this representation vector, and the loss is calculated by comparing it with the true class label in the training set. The entire deep learning network is trained with the goal of minimizing the loss to obtain a deep learning network model based on the knowledge-enhanced graph convolutional network;

[0243] The answer selection module selects a correct answer for a given question. It analyzes and processes the input question using the deep learning network model of the knowledge-enhanced graph convolutional network and outputs the candidate answer with the highest question-answer pair correlation score, representing the correct answer selected for the question.

[0244] As described above, it is only a preferred embodiment of the present invention and not a limitation of the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for answer selection based on a knowledge-enhanced graph convolutional network, characterized in that It includes the following steps: Step A: Collect the questions and answer records of users in the Q&A platform, and label the true labels of each question-answer pair to construct the training dataset DS; Step B: Use the training dataset DS and the knowledge graph ConceptNet to train the deep learning network model M based on the knowledge-enhanced graph convolutional neural network, and analyze the correctness of the corresponding candidate answers for the given questions through this model; Step C: Input the user's question into the trained deep learning network model M and output the matching answer; The specific steps of Step B include the following steps: Step B1: Perform initial encoding on all training samples in the training dataset DS to obtain the initial features E of the question and answer text content q , E a , the global semantic feature sequence E of the question-answer pair cls , the syntactic structure dependency adjacency matrix A of the question-answer pair. At the same time, perform text-knowledge matching and multi-hop knowledge node expansion query on the question and answer text from the knowledge graph ConceptNet, connect the text-matched knowledge nodes and the expanded nodes to obtain the knowledge expansion sequence, and map the information of each knowledge node in the knowledge expansion sequence to a continuous low-dimensional vector, and finally form the knowledge expansion sequence feature C of the question and answer q , C a ; Step B2: Connect the initial features E q and E a of the question and answer text content to obtain the text feature E qa of the question-answer. By performing a masking calculation on E qa , obtain the question-answer edge weight matrix M a . Multiply M a by the syntactic structure dependency adjacency matrix A to obtain the syntactic structure dependency adjacency matrix with edge association weights Step B3: Take the text features E of the question-answer obtained in Step B2 qa and the syntactic structure dependency adjacency matrix with edge association weights and input them into a K-layer graph convolutional network. Through the propagation of node information guided by the syntactic structure dependency relationship between graph nodes, learn the original structural information features of the question-answer text Then, for the text features E of the question-answer qa and the original structural information features of the question-answer text perform semantic enhancement in an attention-based manner to obtain the semantic structure information features of the question-answer Step B4: Input the initial features E q and E a of the question and answer text content obtained in Step B1, and the knowledge extension sequence features C q and C a of the question and answer into two attention calculation mechanisms guided by text semantics to obtain the semantic-guided knowledge features of the question q and the answer a Then input the semantic-guided knowledge features into two multi-head self-attention mechanisms to obtain the self-attention knowledge representations Input the semantic-guided knowledge representations and the self-attention knowledge features into two feed-forward neural network layers to obtain the context features H q and H a ; Use a gating mechanism to filter and fuse the context features H q and H a of the knowledge to obtain the context features H qa of the question-answer knowledge; Step B5: Fuse the knowledge context feature H of the question-answer qa and the semantic structure information feature of the question-answer by means of attention calculation to obtain the knowledge-enhanced semantic structure information feature of the question-answer Then, input the local semantic feature matrix E obtained in Step B1 cls into a multi-scale convolutional neural network to obtain multi-granularity global semantic features Step B6: The semantic structure information features of the knowledge-enhanced question-answer are input into a BiGRU network, and an average pooling operation is performed on the sequence of the hidden state outputs of the BiGRU to obtain the aggregated features of the question-answer The aggregated features of the question-answer and the multi-granularity global semantic features are concatenated to obtain the final Q&A feature E final ; subsequently, E final is input into a linear classification layer and normalized using the softmax function to generate the correlation score f(q,a) ∈ [0,1] between the question-answer; then, according to the objective loss function loss, the gradients of the parameters in the deep learning network model are calculated by the backpropagation method, and the parameters are updated using the stochastic gradient descent method; Step B7: When the change in the loss value generated by the deep learning network model in each iteration is less than the given threshold or reaches the maximum number of iterations, terminate the training process of the deep learning network model.

2. The answer selection method based on a knowledge-enhanced graph convolutional network according to claim 1, wherein The specific steps of Step B1 include the following steps: Step B11: Traverse the training dataset DS. After tokenizing the questions and candidate answer texts in it and removing the stop words, each training sample in DS is represented as ds=(q,a,p); where q is the text content of the question, a is the text content of the candidate answer corresponding to the question; p is the label indicating whether the question and the answer are correct, p∈[0,1], 0 indicates that the candidate answer is the wrong answer, and 1 indicates that the candidate answer is the correct answer; the question q is represented as: wherein, is the i-th word in question q, where i = 1, 2, …, m, and m is the number of words in question q; The answer a is represented as: wherein, is the i-th word in answer a, where i = 1, 2, …, n, and n is the number of words in answer a; Step B12: For the question obtained in Step B11 and the answer perform splicing, insert the [CLS] token in front of the question q, and insert the [SEP] tokens before and after the answer a to construct the question-answering input sequence X of the BERT encoding model s ; The Q&A input sequence is represented as: where m and n respectively represent the number of words in the question q and the answer a; Step B13: Input X s into the BERT model to obtain the output sequence of the i-th layer of the model The output sequence E of the last layer of the model s ; According to the positions of the [CLS] and [SEP] tags in the E s sequence, segment the initial representation vectors of the question and the answer, so as to obtain the initial representation vectors E q and E a of the question and the answer respectively; Connect the [CLS] token in to obtain the global semantic feature E cls of the question and the answer; Among them, the output sequence of the i-th layer of the model is expressed as: Among them, the output sequence E of the last layer of the model s is expressed as: Initial feature E of problem q q It is expressed as: where is the word vector corresponding to the 1st word ; is the word vector corresponding to the 2nd word ; is the word vector corresponding to the mth word ; m is the number of words in the question q, and d is the dimension of the word vector Problem a initial feature E a Expressed as: where is the word vector corresponding to the first word, is the word vector corresponding to the second word, is the word vector corresponding to the nth word, n is the number of words in answer a, and d is the dimension of the word vector;​​​ Global semantic feature E of questions and answers cls It is expressed as: Among them is the [CLS] token output by the first-layer model, is the [CLS] token output by the second-layer model, is the [CLS] token output by the l1-th layer model, l1 is the number of encoder layers of BERT; Step B14: Connect the question and the answer to obtain a word sequence Perform syntactic dependency parsing on X qa to generate an undirected syntactic structure dependency graph and encode it into a corresponding (m + n)-order syntactic structure dependency adjacency matrix A; where the representation of A is: Step B15: Perform question text-knowledge matching and multi-hop node expansion for each word in question q and answer a in the knowledge graph ConceptNet; first, for each word in question q perform text-knowledge matching in the knowledge graph to obtain its corresponding knowledge node Similarly, the corresponding knowledge nodes for each word in answer a can be obtained are obtained Secondly, in the process of multi-hop expanding knowledge nodes, multi-hop node selection is performed according to the text-matched knowledge nodes and the relationships between nodes in the knowledge graph; sort the multi-hop selected knowledge nodes according to their initial weights in the knowledge graph, and select the max_n extended knowledge nodes with the largest weights from them; connect the extended nodes with the text-matched knowledge nodes to form a knowledge expansion sequence; use knowledge embedding to map each knowledge node in the knowledge expansion sequence into a continuous low-dimensional vector, and finally form the knowledge expansion sequence feature C of question q and answer a q , C a ; Among them, the knowledge expansion sequence feature C of question q q is expressed as: Among them, l2 = (m + max_n × m) is the length of the problem knowledge expansion sequence, and the dimension of the knowledge word vector is d; is the extended knowledge node of, and max_n is the number of extended nodes; Knowledge extension sequence feature C of answer a a Expressed as: Among them, l3 = (n + max_n × n) is the length of the answer knowledge expansion sequence, and d is the dimension of the knowledge word vector; is the extended knowledge node of, and max_n is the number of extended nodes.

3. The answer selection method based on a knowledge-enhanced graph convolutional network according to claim 2, wherein The specific steps of Step B2 include the following steps: Step B21: Initial features of the question and answer text content Connect them to obtain the text features of the question-answer Among them m + n is the length of the question-answer text sequence, and d is the dimension of the word vector; Step B22: For the text feature E of the question-answer obtained in Step B21 qa perform masked edge weight calculation to obtain the edge weight value matrix M a , and its calculation process is as follows: where m + n is the length of X qa d is the dimension of the E qa vector, W1 and W2 are trainable parameter matrices; Step B23: Take the edge weight matrix M a and perform a dot product operation with the syntactic structure dependency adjacency matrix A obtained in Step B14 to obtain a syntactic structure dependency adjacency matrix with edge weights The calculation process is as follows: Among them, ⊙ represents the element-wise multiplication operation of matrices.

4. The answer selection method based on a knowledge-enhanced graph convolutional network according to claim 3, characterized in that, The specific steps of Step B3 include the following steps: Step B31: Use the text feature E of the question-answer pair qa as the initial representation vector of the graph node, and use a K-layer graph convolutional network to perform graph convolutional operations on the adjacency matrix to update the graph node information; the update process of the hidden state of node i in the k-th layer graph convolutional network is as follows: where \(k\in[1, K]\) represents the number of layers of the graph convolutional network. is the hidden state output by node \(i\) in the \(k\)-th layer network, and Relu() is a non-linear activation function. is a trainable parameter matrix. is a bias vector, \(d\) i represents the dimension of the initial representation vector of node \(i\). Step B32: Connect the hidden states of the K-th layer graph convolutional network to obtain the original structural information features of the question-answer which is expressed as follows: Among them, m + n is the length of the question-answer text sequence, and d is the dimension of the initial representation vector of the node; Step B33: Semantically enhance the text features E of the question-answer qa and the original structural information features of the question-answer in a way of attention calculation to obtain the semantic structural information features of the question-answer The calculation formula is as follows: Among them, m + n is the length of the question-answer text sequence, and d is the dimension of the initial representation vector of the node. W4 and W5 are trainable parameter matrices.

5. The answer selection method based on a knowledge-enhanced graph convolutional network according to claim 4, wherein The specific steps of Step B4 include the following steps: Step B41: Input the initial features E q and E a of the question and answer text content obtained in Step B13, q and C a the knowledge extension sequence features of the question and answer obtained in Step B15 into two attention calculation mechanisms guided by text semantics to obtain the semantic guidance knowledge features of the question q and the answer a Among them The calculation formula is as follows: α q = softmax(tanh(E q W6×(C q W7) T )) Among them, l2 is the length of the knowledge expansion sequence feature C q ; W6 and W7 are trainable parameter matrices; Similarly, the semantic guidance knowledge representation of the answer is obtained Step B42: Semantic-guided knowledge representation of question q and answer a They are respectively input into two different multi-head attention mechanisms to obtain the self-attention knowledge features of the question and the answer Among them, The calculation formula is as follows: Among them, MHA represents the multi-head attention mechanism, num is the number of parallel heads, and Q, k, and V are all semantic-guided question knowledge features. is a trainable parameter matrix, head i represents the output of the i-th attention function, where i ∈ [1, num]; similarly, the self-attention knowledge features of the answer are obtained. Step B43: Input the self-attention knowledge features of the question and answer and the semantic guidance knowledge features into two linear feed-forward layer networks for fusion to obtain the context features H q 、H a ; Among which H q has the following calculation formula: Among them, is a trainable parameter matrix, is a bias vector; Step B45: Input the knowledge context features H of the question and answer q , H a into a gating mechanism for filtering and fusion to obtain the knowledge context feature H of the question-answer qa ; Among which H qa has the following calculation formula: g = sigmoid(H q W 15 :H a W 16 ) H qa = (1 - g) ⊙ H q + g ⊙ H a Among them l2 is the C q length, l3 is the C a length; is a trainable parameter, and ":" is a concatenation operation.

6. The method for answer selection based on a knowledge-enhanced graph convolutional network according to claim 5, wherein The specific steps of Step B5 include the following steps: Step B51: Perform knowledge enhancement on the knowledge context feature H of the question-answer pair qa and the semantic structure information feature of the question-answer pair in the way of attention calculation to obtain the semantic structure information feature of the knowledge-enhanced question-answer pair The calculation formula is as follows: Among them, m + n is the length of the text sequence X of the question-answer pair qa , which is a trainable parameter; Step B52: Input the global semantic feature E obtained in Step B1 cls , into a convolutional neural network with multiple scales to obtain global semantic features with multiple granularities It is expressed as: where MCNN() represents the multi-size CNN.

7. The answer selection method based on a knowledge-enhanced graph convolutional network according to claim 6, characterized in that The specific steps of Step B6 include the following steps: Step B61: Input the semantic structure information features of the knowledge-enhanced question-answer pairs into the forward layer and the backward layer of a bidirectional GRU network, respectively obtaining the state features of the forward hidden layer and the state features of the backward hidden layer where Step B62: Concatenate and , and through a linear layer, obtain the output feature E of the BiGRU for the question-answer pair gru ; perform average pooling on E gru to obtain the aggregated feature of the question-answer The calculation formula is as follows: Among them, are trainable parameters, meanpool() is an average pooling function; Step B63: Aggregate the problem-answer features and the multi-granularity global semantic features to obtain the final question-answer feature representation E final ; E final is represented as follows: Step B64: Input the final question-answer feature E final into a linear classification layer and normalize it using the softmax function to generate the correlation score f(q,a) ∈ [0,1] between the question and the answer. The calculation formula is as follows: f(q,a) = softamx(E final W 19 + b4) Among them, is a trainable parameter matrix, is a bias vector; Step B65: Use the cross-entropy as the loss function to calculate the loss value, update the learning rate through the gradient optimization algorithm Adam, and use backpropagation to iteratively update the model parameters to train the model by minimizing the loss function; the calculation formula for minimizing the loss function L is as follows: where f(q,a) i ∈[0,1] is the relevance score of the question-answer pair calculated by the softmax classifier, and y i ∈[0,1] is the binary classification label.

8. An answer selection system based on a knowledge-enhanced graph convolutional network, characterized in that Adopt an answer selection method based on the knowledge-enhanced graph convolutional network described in any one of the above claims 1 to 7, including: A data collection module that collects the questions and answer records of users in the Q&A platform, and labels the true labels of each question-answer pair to construct the training dataset DS; A text preprocessing module for preprocessing the training samples in the training set, including tokenization and stop word removal; A text encoding module that performs initial encoding on all training samples in the training dataset DS to obtain the initial features of the question and answer text content, the global semantic feature sequence of the question-answer pair, the syntactic structure dependency adjacency matrix of the question-answer, and at the same time performs text-knowledge matching and multi-hop knowledge node extension queries on the question and answer texts from the knowledge graph ConceptNet to obtain the knowledge extension sequence features of the question and answer; The network model training module is used to input the initial features of questions and answer texts, the global semantic features of question-answer pairs, the syntactic structure dependency adjacency matrix of question-answer pairs, and the knowledge expansion sequence features of questions and answers into a deep learning network to obtain the final features of question-answer pairs. The probability of answer correctness is predicted using the final features of the question-answer pairs, and the loss is calculated by comparing with the true class labels in the training set. The entire deep learning network is trained with the goal of minimizing the loss to obtain a deep learning network model based on a knowledge-enhanced graph convolutional network; The answer selection module selects a correct answer for a given question. It uses the deep learning network model based on the knowledge-enhanced graph convolutional network to analyze and process the input question, and outputs the candidate answer with the highest relevance score for the question-answer pair, representing the correct answer selected for the question.

Citation Information

Patent Citations

  • Intelligent question answering method based on XLNet-BiGRU-CRF

    CN113641809A

  • Machine reading understanding method based on BERT and gating attention enhancement network

    CN114398976A