A Chinese automatic question answering method based on radical features and multi-layer attention mechanism

By introducing the radical characteristics of Chinese characters and multi-layer attention mechanism in the Chinese automatic question and answer method, the problem of insufficient accuracy of Chinese automatic question and answer in the prior art is solved, and the accuracy and performance of the model are significantly improved.

CN114118099BActive Publication Date: 2025-05-23ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111325158.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-05-23
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

The existing Chinese automatic question and answer method uses Chinese word vectors as model input, and the accuracy rate is difficult to meet user needs.

Method used

The method based on radical characteristics and multi-layer attention mechanism is adopted, and the model is optimized by adding radical characteristics of Chinese characters and using multi-layer attention mechanisms to improve the accuracy of automatic question-and-answer questions and answers.

Benefits of technology

It effectively improves the accuracy of automatic question-and-answer and improves the performance of the model in Chinese reading comprehension tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003346690760000021
    Figure BDA0003346690760000021
  • Figure BDA0003346690760000033
    Figure BDA0003346690760000033
  • Figure BDA0003346690760000043
    Figure BDA0003346690760000043
Patent Text Reader

Abstract

A Chinese automatic question answering method based on radical features and multi-layer attention mechanism comprises the following steps: step 1, preprocessing a data set; step 2, obtaining a word embedding matrix, and obtaining a radical embedding matrix by random initialization; step 3, converting words into vector representations respectively by word embedding and radical embedding, and appending linguistic features after the word vector; step 4, inputting a document vector sequence and a question vector sequence into different bidirectional RNN networks for encoding; step 5, according to the document vector sequence and the question vector, sequentially calculating the probability of the start and end boundaries of the answer, generating a target probability distribution, step 6, using the data set to train the model for N rounds, calculating the loss and updating the parameters, using the mini-batch strategy to train the model, using the model to process a given document and related questions, and predicting the answer. The present invention improves the accuracy of automatic question answering.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A Chinese automatic question answering method based on radical features and multi-layer attention mechanism, It is characterized in that The method comprises the following steps: Step 1: Preprocess the data set S. The data set S is represented as {S i |S i =(Passage, Question, Answer), 1≤i≤n}, where S i represents the i-th data, which consists of three parts: document Passage, question Question and answer Answer. Answer is a substring of passage. n is the size of the data set S. The data preprocessing steps include: Chinese word segmentation, linguistic feature annotation and word frequency statistics. Step 2: Load the pre-trained word embedding to get the word embedding matrix WE l×d , the radical embedding matrix RE is obtained by random initialization k×r , where l is the number of words in word embedding, d is the dimension of word vector, k is the number of radicals in radical dictionary, r is the dimension of radical embedding, and radical embedding matrix RE is the model training parameter; Step 3: Convert the words in PWList and QWList into vector representations through word embedding and radical embedding, and then append linguistic features to the word vectors to obtain the vector sequence representations vPWList and vQWList of PWList and QWList. The process is as follows: (3.1) Convert the words pwordi in PWList into vector representation Where WE(word) represents the word vector corresponding to the word word, radical(word) = [radicalDict(w 1 ),…,radicalDict(wwcnt (word) )] represents the radical list of word word, wcnt(word) represents the number of Chinese characters in word word, RE(radical(word)) represents the matrix composed of radical vectors corresponding to the Chinese character radicals in word word, represents vector concatenation, c represents the number of convolution output channels, and the final vpword i The dimension is 1×(2d+c+4), and the function CNN_RE(), f match (), f token (), f align ()The results returned are all vectors; (3.2) The word qword in QWList i Convert to vector representation Final vpword i The dimension is 1×(d+c); The vector sequence of PWList is represented as vPWList = [vpword 1 ,vpword 2 ,…,vpwordlen (PWList) ], the vector sequence of QWList is represented as vQWList = [vqword 1 ,vqword 2 ,…,vqwordlen (QWList) ]; Step 4: Input the document vector sequence vPWList and the question vector sequence vQWList into different bidirectional RNN networks for encoding, and obtain the document vector sequence representation PWC and the question vector representation Q containing the question information. The process is as follows; (4.1) Input vPWList into RNN 1 Encoding to obtain the vector sequence Pl = [pl 1 ,pl 2 ,…,pl len(PWList) ]=RNN 1 (vPWList), where RNN 1 Network output result pl i The dimension is 1×h; (4.2) Input vQWList into RNN 2 Encode to obtain the vector sequence Ql = [ql 1 ,ql 2 ,…,qllen (QWList) ]=RNN 2 (vQWList), where RNN 2 Network output result ql i The dimension is 1×h; (4.3) Compress the encoded vector sequence Ql into a vector in w is a trainable parameter vector; (4.4) The encoded vector sequence Pl is further processed based on the attention mechanism to obtain: in in For pl i With ql j The attention weight of (4.5) Input Ph into RNN 3 The document vector sequence containing the question information is encoded in PWC = [pwc 1 ,pwc 2 ,…,pwclen (PWList) ]=RNN 3 (Ph), RNN 3 Network output result pwc i The dimension is 1×h′; Step 5: Based on the document vector sequence representation PWC and the question vector representation Q, calculate each word PWList in PWList in turn i The probability of starting the boundary as an answer and the probability of the answer ending at the boundary Where W s , W e It is a trainable parameter. According to the answer AWList, the target probability distribution PTS is generated at the left boundary l and the right boundary r of PWList. i =Θ(i==l)|1≤i≤len(PWList)] and PTE=[pte i =Θ(i==r)|1≤i≤len(PWList)], where the function Θ(x) returns 1 if x is true and 0 if x is false; Step 6: Divide the dataset S into a training dataset T and a test dataset V. Use the dataset T to train the model for N rounds. start , P end , PTS, PTE calculate the loss and update the parameters, use the mini-batch strategy to train the model, use the test data set V to evaluate the model after each round of training, and take the best performing parameters in N rounds as the model parameters, including RNN network, CNN network and f align The parameters of the fully connected layer α in the () function and RE, W s , W e , w parameters, where the loss calculation method is loss(P start ,PTS)+loss(P send ,PTE); Step 7: Load the trained model parameters, use the model to process a given document p and a related question q, and predict the answer ans; In step 1, the data preprocessing process is as follows: (1.1) Use a Chinese word segmentation tool to perform word segmentation on the dataset S to obtain the word list PWList of Passage = [pword 1 , pword 2 , …, pwordlen (Passage) , QWList = [qword 1 , qword 2 , …, qwordlen (Question) , AWList = [aword 1 , aword 2 , …, awordlen (Answer) , where len(x) represents the number of words in the string x; (1.2) Map the Chinese part-of-speech tagging features and named entity recognition features into numbers to obtain the part-of-speech feature mapping POSMap = { <pos 1 :1>, <pos 2 :2>,…, <pos k :k>}、Named entity feature map NerMap={ <ner 1 :1>, <ner 2 :2>,…, <ner l :l>}, where k is the number of part-of-speech feature categories, l is the number of named entity recognition feature categories, pos i andner j Represent the part-of-speech tagging features and named entity recognition features respectively; (1.3) Use linguistic tools to perform part-of-speech tagging and named entity recognition on PWList, and save the results. Define POS(word, Passage) to represent the part-of-speech features of word in Passage, and Ner(word, Passage) to represent the named entity features of word in Passage; (1.4) Count the frequency of words pwordi appearing in PWList Where count(word,PWList) indicates the number of times word appears in PWList; (1.5) Obtain the radical dictionary of Chinese characters through manual labeling radicalDict = { <w 1 :r 1 >, <w 2 :r 2 >,…, <w m :r m }, where w i For Chinese characters, r i w i The radical of , m is the size of the radical dictionary radicalDict.

2. The Chinese automatic question answering method based on Chinese character radical features and multi-layer attention mechanism as claimed in claim 1, Features In step 3, the functions CNN_RE() and f match (), f token () and f align The expression of () is as follows: Among them, POS(word) represents the part-of-speech tagging feature of word word; NER(word) represents the named entity recognition feature of word; TF(word) represents the frequency of word; α(vector) represents a fully connected layer using ReLU activation function; Conv2D represents a two-dimensional convolution operation with kernel size kernelsize and output channels outputChannels, and QWList(i) represents the i-th word in QWList.

3. A Chinese automatic question answering method based on Chinese character radical features and multi-layer attention mechanism as described in claim 1 or 2, Features , the process of step 7 is as follows: (7.1) Load the trained model parameters; (7.2) Preprocess the document p and question q using the operation in step 1 to obtain PWList and QWList; (7.3) Use the operations of steps 3, 4, and 5 to perform calculations and obtain the value of each word PWList in the document PWList. i The probability P of the interval starting as the answer start (i) and the probability P of ending as the answer end (i); (7.4) Calculate the answer interval The final predicted answer is ans=PWList[l,r], where PWList[l,r] represents the substring in PWList starting from l and ending with r.

Citation Information

Patent Citations

  • Reading understanding method based on attention pooling mechanism

    CN109977199A

  • Construction method of medical intelligent question-answering system based on attention mechanism

    CN110543557A