A Chinese automatic question answering method based on radical features and multi-layer attention mechanism
By introducing the radical characteristics of Chinese characters and multi-layer attention mechanism in the Chinese automatic question and answer method, the problem of insufficient accuracy of Chinese automatic question and answer in the prior art is solved, and the accuracy and performance of the model are significantly improved.
Patent Information
- Application Number
- CN202111325158.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-10
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-11-10
AI Technical Summary
The existing Chinese automatic question and answer method uses Chinese word vectors as model input, and the accuracy rate is difficult to meet user needs.
The method based on radical characteristics and multi-layer attention mechanism is adopted, and the model is optimized by adding radical characteristics of Chinese characters and using multi-layer attention mechanisms to improve the accuracy of automatic question-and-answer questions and answers.
It effectively improves the accuracy of automatic question-and-answer and improves the performance of the model in Chinese reading comprehension tasks.
Smart Images

Figure BDA0003346690760000021 
Figure BDA0003346690760000033 
Figure BDA0003346690760000043
Abstract
Claims
1. A Chinese automatic question answering method based on radical features and multi-layer attention mechanism, It is characterized in that The method comprises the following steps: Step 1: Preprocess the data set S. The data set S is represented as {S i |S i =(Passage, Question, Answer), 1≤i≤n}, where S i represents the i-th data, which consists of three parts: document Passage, question Question and answer Answer. Answer is a substring of passage. n is the size of the data set S. The data preprocessing steps include: Chinese word segmentation, linguistic feature annotation and word frequency statistics. Step 2: Load the pre-trained word embedding to get the word embedding matrix WE l×d , the radical embedding matrix RE is obtained by random initialization k×r , where l is the number of words in word embedding, d is the dimension of word vector, k is the number of radicals in radical dictionary, r is the dimension of radical embedding, and radical embedding matrix RE is the model training parameter; Step 3: Convert the words in PWList and QWList into vector representations through word embedding and radical embedding, and then append linguistic features to the word vectors to obtain the vector sequence representations vPWList and vQWList of PWList and QWList. The process is as follows: (3.1) Convert the words pwordi in PWList into vector representation Where WE(word) represents the word vector corresponding to the word word, radical(word) = [radicalDict(w 1 ),…,radicalDict(wwcnt (word) )] represents the radical list of word word, wcnt(word) represents the number of Chinese characters in word word, RE(radical(word)) represents the matrix composed of radical vectors corresponding to the Chinese character radicals in word word, represents vector concatenation, c represents the number of convolution output channels, and the final vpword i The dimension is 1×(2d+c+4), and the function CNN_RE(), f match (), f token (), f align ()The results returned are all vectors; (3.2) The word qword in QWList i Convert to vector representation Final vpword i The dimension is 1×(d+c); The vector sequence of PWList is represented as vPWList = [vpword 1 ,vpword 2 ,…,vpwordlen (PWList) ], the vector sequence of QWList is represented as vQWList = [vqword 1 ,vqword 2 ,…,vqwordlen (QWList) ]; Step 4: Input the document vector sequence vPWList and the question vector sequence vQWList into different bidirectional RNN networks for encoding, and obtain the document vector sequence representation PWC and the question vector representation Q containing the question information. The process is as follows; (4.1) Input vPWList into RNN 1 Encoding to obtain the vector sequence Pl = [pl 1 ,pl 2 ,…,pl len(PWList) ]=RNN 1 (vPWList), where RNN 1 Network output result pl i The dimension is 1×h; (4.2) Input vQWList into RNN 2 Encode to obtain the vector sequence Ql = [ql 1 ,ql 2 ,…,qllen (QWList) ]=RNN 2 (vQWList), where RNN 2 Network output result ql i The dimension is 1×h; (4.3) Compress the encoded vector sequence Ql into a vector in w is a trainable parameter vector; (4.4) The encoded vector sequence Pl is further processed based on the attention mechanism to obtain: in in For pl i With ql j The attention weight of (4.5) Input Ph into RNN 3 The document vector sequence containing the question information is encoded in PWC = [pwc 1 ,pwc 2 ,…,pwclen (PWList) ]=RNN 3 (Ph), RNN 3 Network output result pwc i The dimension is 1×h′; Step 5: Based on the document vector sequence representation PWC and the question vector representation Q, calculate each word PWList in PWList in turn i The probability of starting the boundary as an answer and the probability of the answer ending at the boundary Where W s , W e It is a trainable parameter. According to the answer AWList, the target probability distribution PTS is generated at the left boundary l and the right boundary r of PWList. i =Θ(i==l)|1≤i≤len(PWList)] and PTE=[pte i =Θ(i==r)|1≤i≤len(PWList)], where the function Θ(x) returns 1 if x is true and 0 if x is false; Step 6: Divide the dataset S into a training dataset T and a test dataset V. Use the dataset T to train the model for N rounds. start , P end , PTS, PTE calculate the loss and update the parameters, use the mini-batch strategy to train the model, use the test data set V to evaluate the model after each round of training, and take the best performing parameters in N rounds as the model parameters, including RNN network, CNN network and f align The parameters of the fully connected layer α in the () function and RE, W s , W e , w parameters, where the loss calculation method is loss(P start ,PTS)+loss(P send ,PTE); Step 7: Load the trained model parameters, use the model to process a given document p and a related question q, and predict the answer ans; In step 1, the data preprocessing process is as follows: (1.1) Use a Chinese word segmentation tool to perform word segmentation on the dataset S to obtain the word list PWList of Passage = [pword 1 , pword 2 , …, pwordlen (Passage) , QWList = [qword 1 , qword 2 , …, qwordlen (Question) , AWList = [aword 1 , aword 2 , …, awordlen (Answer) , where len(x) represents the number of words in the string x; (1.2) Map the Chinese part-of-speech tagging features and named entity recognition features into numbers to obtain the part-of-speech feature mapping POSMap = { <pos 1 :1>, <pos 2 :2>,…, <pos k :k>}、Named entity feature map NerMap={ <ner 1 :1>, <ner 2 :2>,…, <ner l :l>}, where k is the number of part-of-speech feature categories, l is the number of named entity recognition feature categories, pos i andner j Represent the part-of-speech tagging features and named entity recognition features respectively; (1.3) Use linguistic tools to perform part-of-speech tagging and named entity recognition on PWList, and save the results. Define POS(word, Passage) to represent the part-of-speech features of word in Passage, and Ner(word, Passage) to represent the named entity features of word in Passage; (1.4) Count the frequency of words pwordi appearing in PWList Where count(word,PWList) indicates the number of times word appears in PWList; (1.5) Obtain the radical dictionary of Chinese characters through manual labeling radicalDict = { <w 1 :r 1 >, <w 2 :r 2 >,…, <w m :r m }, where w i For Chinese characters, r i w i The radical of , m is the size of the radical dictionary radicalDict.
2. The Chinese automatic question answering method based on Chinese character radical features and multi-layer attention mechanism as claimed in claim 1, Features In step 3, the functions CNN_RE() and f match (), f token () and f align The expression of () is as follows: Among them, POS(word) represents the part-of-speech tagging feature of word word; NER(word) represents the named entity recognition feature of word; TF(word) represents the frequency of word; α(vector) represents a fully connected layer using ReLU activation function; Conv2D represents a two-dimensional convolution operation with kernel size kernelsize and output channels outputChannels, and QWList(i) represents the i-th word in QWList.
3. A Chinese automatic question answering method based on Chinese character radical features and multi-layer attention mechanism as described in claim 1 or 2, Features , the process of step 7 is as follows: (7.1) Load the trained model parameters; (7.2) Preprocess the document p and question q using the operation in step 1 to obtain PWList and QWList; (7.3) Use the operations of steps 3, 4, and 5 to perform calculations and obtain the value of each word PWList in the document PWList. i The probability P of the interval starting as the answer start (i) and the probability P of ending as the answer end (i); (7.4) Calculate the answer interval The final predicted answer is ans=PWList[l,r], where PWList[l,r] represents the substring in PWList starting from l and ending with r.
Citation Information
Patent Citations
Reading understanding method based on attention pooling mechanism
CN109977199A
Construction method of medical intelligent question-answering system based on attention mechanism
CN110543557A