Implementation Method and Device for Machine Reading Comprehension Assisted by Generation Model
Through encoder decoder structure and interactive fusion training, and combining the correct options to generate an expanded vector, the problem of lack of explicit knowledge in multiple-choice reading comprehension and difficulty in training multi-format data sets is solved, and efficient common sense reasoning and accuracy improvement is achieved.
Patent Information
- Application Number
- CN202210285465.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-03-23
AI Technical Summary
The reading comprehension scheme of multiple-choice questions in the prior art has problems such as lack of explicit knowledge and difficulty in training multi-format data sets, as well as the problem that the decoder of pre-trained generation models is not fully utilized.
Based on the reading comprehension method of generative model, through the encoder decoder structure, using interactive fusion training of problems and options, combining correct options to generate expansion vectors, and using teacher-forcing and cross-entropy loss optimization model to achieve common sense reasoning in a single data set.
The reading comprehension accuracy of multiple-choice questions is significantly improved, and is better than existing methods, especially on the CommonsenseQA dataset, which outperforms T5 and UnifiedQA-T5-base.
Smart Images

Figure CN114611510B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, relates to natural language reading comprehension, especially single-choice questions based on common sense, and provides a method and device for realizing machine reading comprehension assisted by a generative model. Background Art
[0002] The reading comprehension ability is an important tool for evaluating whether a computer can understand human language and perform logical reasoning on text. Given a natural language question, the computer needs to rely on its own common sense knowledge and language understanding ability to obtain the correct answer. The current reading comprehension dataset formats are generally divided into the following four types: extractive reading comprehension, represented by SQUAD; generative reading comprehension, represented by NarrativeQA; yes / no questions, represented by BoolQ; and multiple-choice questions, represented by CommonsenseQA. Among them, multiple-choice questions are often more difficult because they generally require combining human common sense and complex multi-hop reasoning to solve, so they can better reflect the computer's ability to understand human language and thus become an important evaluation benchmark.
[0003] The existing methods for solving the natural language reading comprehension problem of multiple-choice questions are generally divided into the following two types: the method of using explicit additional external knowledge to assist in answering questions, and the method of using a generative model to fine-tune multiple-format datasets simultaneously.
[0004] For the first method of using explicit additional external knowledge to assist in answering questions, entities appearing in the question and options are usually extracted, and then the knowledge base of external resources such as Conceptnet is used to extract the relationships connecting the two entities, and then linearization or graph neural network methods are used for modeling. In particular, some methods will also use dictionary information such as Wiktionary to find the description information of the defined words and their lemmas in the question and options from the dictionary, and then splice the original question and options and input them into the pre-trained language model. Typical models include ALBERT+HGN, ALBERT+DESC+KCR, and ALBERT+PathGenerator, etc.
[0005] Another type of method that uses a generative model to fine-tune multiple-format datasets simultaneously. The main idea is to unify multiple-format reading comprehension datasets, such as extractive, generative, single-choice, and yes / no questions, into a text-to-text framework, and then use large-scale seq2seq pre-trained models, such as Google's T5 and Facebook's Bart, etc., to fine-tune a large number of multiple-format datasets simultaneously, so that common sense information can be learned mutually among multiple-format tasks, thereby assisting the answering effect on a single dataset. The representative of this method is UnifiedQA.
[0006] Both of the above two technical methods have achieved good results in multiple-choice questions based on common sense. However, they also have obvious disadvantages. For example, in the first method, using explicit additional external knowledge can indeed provide effective information for the computer to answer questions. However, there may still be problems with the lack of explicit knowledge, such as incomplete knowledge bases and dictionary information, entity association failures, etc. These problems will have a great impact on the effectiveness of this method. The second method models multiple dataset formats into a unified text-to-text format. The problem is that it consumes too much training resources. In fact, the best model of UnifiedQA uses T5-11B, which has as many as 11 billion parameters, bringing great difficulties in training and deployment for organizations with insufficient resources. In addition, when the UnifiedQA method faces specific dataset usage requirements, there may be a large number of other datasets that cannot provide effective knowledge transfer effects or even reduce the effects, resulting in low resource utilization. Summary of the Invention
[0007] The problem to be solved by the present invention is that in the prior art, there are problems of lack of explicit knowledge, difficult and inefficient training of multi-format datasets in the reading comprehension solution for multiple-choice questions, and the underutilization of the decoder in the existing methods of using pre-trained generative models to process multiple-choice questions.
[0008] The technical solution of the present invention is: an implementation method for machine reading comprehension assisted by a generative model. For natural language reading comprehension of multiple-choice questions, a reading comprehension model is constructed based on the encoder-decoder of the sequence-to-sequence model, and is trained using a question set q, a corresponding option set o, and a correct option set a. The reading comprehension model includes two workflows. One is the generation workflow. The question input encoder obtains the question encoding representation Q, and inputs the question encoding representation Q into the decoder to obtain the answer decoding representation Ag. During training, the teacher-forcing loss is calculated according to the correct option. The other is the reading comprehension workflow. The question encoding representation Q is separately input into the decoder to generate a decoding representation as the vector representation Au for question expansion. At the same time, the question is concatenated with each corresponding option and then input into the encoder to obtain the question-option representation QO. The QO and the expanded vector representation Au are interactively fused through a bidirectional matching layer to obtain a fused representation. After that, the fused representation obtains the logit corresponding to each option through a linear layer. During training, the cross-entropy loss is calculated between these logits and the correct option, and the reading comprehension model is trained and optimized by combining the teacher-forcing loss and the cross-entropy loss to obtain a generative reading comprehension model.
[0009] The present invention provides an implementation method for generative reading comprehension. During training, the correct option is used as an auxiliary, enabling the decoder to generate some augmented vectors beneficial for answering questions. These vectors are combined with the representation of the encoder and jointly trained and optimized. The resulting reading comprehension model is used to predict the correct option based on the question of a multiple-choice question, which can significantly improve the accuracy of reading comprehension.
[0010] Furthermore, although large language models can capture a large amount of knowledge during pre-training, their effectiveness is usually based on integrating external knowledge bases, especially in common sense reasoning tasks, such as the understanding of multiple-choice questions. The present invention uses a sequence-to-sequence model (seq2seq model), which can use only the correct options provided within the specified dataset as supervision, without the need for additional common sense knowledge, such as explicit knowledge provided by external resources like ConceptNet, Wiktionary, etc., nor the need for other datasets to assist in learning common sense information. The present invention inputs the questions in the dataset into the Encoder, and at the Decoder end, combines the correct options to output some implicit vector representations beneficial for answering questions, and interacts the questions and options. By evaluating the losses of two workflows, the correct option is finally determined, thereby making full use of the common sense reasoning ability existing in the pre-trained model without using explicit knowledge provided by an additional knowledge base or other formatted datasets, effectively solving the problems of the lack of explicit knowledge and the difficulties and inefficiencies in training with multi-formatted datasets in the prior art, as well as the problem that the decoder cannot be fully utilized in the existing methods for processing multiple-choice questions using pre-trained generative models.
[0011] Based on the encoder-decoder structure, the present invention proposes two new workflows, enabling the encoder-decoder to learn the common sense reasoning relationship among questions-options-correct options. The training of pre-trained language models in the prior art either requires using external resources other than the training dataset to provide common sense knowledge supplementation, or requires using other datasets other than this training dataset for joint training to improve the learning effect of common sense information. The answering effect of the present invention based on only using a single dataset is better than that of existing models. Under the condition of only using the internal answers in the dataset for supervision, the answering metrics of the present invention significantly exceed the existing baseline models such as T5 and UnifiedQA-T5-base that use external resources for assistance. Using the official validation set of the CommonsenseQA dataset as the test set, with the validation set being 10% of the training set divided, the following are the answering metric results: Results based on the T5 base model: T5: 60.93, UnifiedQA: 62.35, and the model of the present invention reaches 63.45. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] The present invention provides an implementation method and device for assisting reading comprehension by expanding vectors with the help of a generative model. A reading comprehension model is constructed based on an encoder-decoder model, and the network structure includes an encoder-decoder module, a bidirectional matching layer module, a linear mapping layer module, and a teacher-forcing and cross-entropy loss module. The present invention proposes two workflows. In pre-training, on the one hand, the correct option is used as an assistant, and at the same time, the decoder generates some expanded vectors that are beneficial to answering questions, which are combined with the question-option representation of the encoder to improve the machine's reading comprehension ability. The encoder is used to encode the question and options. On the one hand, the decoder optimizes itself through the teacher-forcing loss according to the correct option. On the other hand, according to the question-option representation output by the encoder, without using the correct option, the decoder directly generates an expanded vector of the question according to the greedy strategy, combines it with the question-option representation output by the encoder, interacts through the bidirectional matching layer, inputs the interaction result into the linear layer to obtain the logits corresponding to the options, and then optimizes the answering effect through the cross-entropy loss.
[0014] Multiple-choice reading comprehension based on common sense includes questions and several options as inputs, and the understanding task is to infer the correct option among them, that is, the answer. This problem requires the understanding model to perform natural language reasoning on the questions and options in combination with common sense, and finally select the given correct option. For example, in CommonsenseQA, a question is Where would I not want a fox? (Where can't I get a fox?), and the given options are hen house (chicken coop), england (England), mountains (mountains), english hunt (English hunting ground), california (California), and the correct option for this question is hen house (chicken coop). The following definitions are made here:
[0015] Question set: q = [q1, q2,..., q nq
[0016] Option set: o = [o1, o2,..., o no
[0017] Correct option set: a = [a1, a2,..., a na
[0018] The first part of the reading comprehension model is to encode the question alone to obtain the representation Q of the question, as shown in the following formula:
[0019] Q = Encoder(q)
[0020] After obtaining the encoded representation of the question, input it into the decoder Decoder to get the decoded representation of the answer $A_g$, as shown in the following formula:
[0021] A g = Decoder(Q)
[0022] Here, the decoder used is the transformer decoder, and usually the training method is teacher-forcing. Teacher-forcing is a training method for sequence-to-sequence models. It is assumed that the previous outputs at each step are correct, and the distribution of the next word is predicted at each position to fit the true distribution of the next word at that position. Here, the word refers to the token. To achieve parallelization of training, usually the mask matrix of the transformer decoder is set to a lower triangular matrix, so as to ensure that each position can only see the words before that position and not the words after. In addition, when the Decoder is trained according to the correct option $a$, a BOS tag will be added in front of it, representing the start tag of the sentence, and an EOS tag will be added at the back, representing the end tag of the sentence, so that the start and end of the generation can be known during prediction.
[0023] After the decoder gets the decoded representation of the answer $A_g$, it is mapped to the distribution on the vocabulary through a linear layer and a softmax operation, as shown in the following formula:
[0024]
[0025] represents the probability of the $t$-th word in the vocabulary corresponding to the prediction at the $i$-th token position.
[0026] The generation loss, that is, the teacher-forcing loss is:
[0027]
[0028] $n_a$ represents the total number of tokens in the correct option, and $a$ i represents the $i$-th token. For example, if the correct option is "a dog", then $n_a = 2$, $a_0 = a$, $a_1 = dog$.
[0029] On the other hand, in the reading comprehension stream, the present invention uses the same Decoder to decode the question encoding representation Q in an autoregressive manner, selects the next word according to the greedy strategy, defines BOS as the first input at the start of the decoder, and EOS as the last output of the decoder. First, input BOS into the model, map the corresponding representation to the distribution of the vocabulary after obtaining it, then greedily select the word corresponding to the maximum probability, and concatenate all the words that have been generated before and input them into the Decoder again. Finally, iterate the above process until EOS is selected. Thus, we can obtain the representation for assisting in answering questions, and the formula is as follows:
[0030] A u , tokens = Decoder(Q)
[0031] Here, tokens refers to all the tokens obtained at each decoding step using the greedy strategy, and Au is the vector representation of the question expansion.
[0032] To obtain the representation of the interaction between each question and option, concatenate the question with each option respectively and input them into the encoder Encoder simultaneously to obtain the question-option representation QO that integrates question information:
[0033] QO = Encoder(q, o)
[0034] Next, perform the Co-Match operation on the question-option representation QO and the vector representation Au obtained from the previous Decoder, so as to enable these two representations to interact and learn the integrated representation that integrates the information for assisting in answering questions The formula is as follows:
[0035]
[0036] The Co-Match operation here realizes the interactive integration through a bidirectional matching layer, and is defined as follows: Let the two input vectors be respectively:
[0037]
[0038] m, n, and h respectively represent the dimensions of the vectors. Use the method of matrix multiplication to obtain the similarity matrix S:
[0039]
[0040] Among them, the element at the x-th row and y-th column of S represents the similarity between the x-th word in A and the y-th word in B, and is defined as the inner product of the representations of these two words.
[0041] After obtaining the similarity matrix, use the softmax operation to get the attention weights for each word in B corresponding to each word in A, which are defined as follows:
[0042]
[0043] Similarly, we can get the attention weights for each word in A corresponding to each word in B, which are defined as follows:
[0044]
[0045] According to S b and A, we can get the representation of B updated using A, which is defined as follows:
[0046]
[0047] Based on the representation of B updated using A and combined with B itself, we concatenate these two representations and multiply them with Sa to get the representation of A that incorporates the information of B, which is defined as follows:
[0048]
[0049] Using the same method, we can get the representation of A updated using B, which is defined as follows:
[0050]
[0051] Concatenate it with A itself and multiply it with S b to get the representation of B that incorporates the information of A, which is defined as follows:
[0052]
[0053] Finally, combine A and the representation of A that incorporates the information of B, and use the transformation matrix W A to get the final output representation of A, which is defined as follows:
[0054]
[0055] Similarly, combine B and the representation of B that incorporates the information of A, and use another transformation matrix W B to get the final representation of B:
[0056]
[0057] Among them, the two transformation matrices W A and W B are model parameters learned during training, and their dimensions are:
[0058]
[0059]
[0060] According to the obtained fused representation, all options are mapped to corresponding logits through a linear layer, which is defined as follows:
[0061]
[0062] Use the softmax operation to map to the probability that each option is selected as the answer, and use the cross-entropy loss function to obtain the reading comprehension loss, which is defined as follows:
[0063]
[0064] where logit answer is the logit corresponding to the correct option. The training objective of the present invention hopes that the model predicts to make the logit of the correct option as large as possible compared to other wrong options, so as to select the correct option.
[0065] The understanding of multiple-choice questions in the present invention is applicable to single-choice or multiple-choice. For single-choice questions, directly map the correct option to obtain the corresponding logit; for multiple-choice questions, due to the appearance of option combinations, it is no longer possible to simply process according to the option order. The present invention concatenates the representations of the T options of the multiple-choice question in order and uses a linear layer to map it to a 2 T -1-dimensional vector, so that various possible combinations of options are mapped to a new option sorting, converting the multiple-choice question into the form of a single-choice question. The combined sorting of multiple correct options is a number between 1 and 2 T -1. The mapping of the correct option combination is:
[0066]
[0067] I(f) represents whether the f-th option is the correct option, 1 for yes and 0 for no.
[0068] For example, for options A, B, C, and D, the option numbers are in ascending order of characters, and their orders are 0, 1, 2, and 3 respectively. Map each possible multiple-choice combination to a new option sorting to obtain a 15-dimensional vector, representing 15 possible combinations of options. If both A and B are correct options, the label of the correct option combination after mapping is: answer = 1*1 + 2*1 + 4*0 + 8*0 = 3. Thus, the multiple-choice question with four options is converted into a single-choice question with 15 options, and the loss u .
[0069] Finally, combining the generation loss and the reading comprehension loss, the following multi-task optimization loss is obtained:
[0070] L(θ) = λ × loss u+(1 - λ) × loss q
[0071] Here, θ is the model parameter, and λ is defined as:
[0072] λ = rouqe(tokens, a)
[0073] That is, the rouge value between the result of greedy strategy decoding and the correct option. Its meaning is that if all the generated tokens have a high similarity with the correct option, it indicates that the generated result is good, thus relatively reducing the weight of the generation loss and increasing the weight of the reading comprehension loss. Conversely, if the rouge value between the generated tokens and the correct option is small, it means that the generation effect is poor, so the weight of the generation loss is increased accordingly, enabling the model to prioritize improving the generation effect and at the same time preventing the reading Co-Match module from being affected by generation noise and reducing the training effect. The level of similarity here can be judged by setting a threshold.
[0074] Finally, the gradient descent and error backpropagation algorithms are used to optimize the model. Figure 1 In it, SG is the abbreviation of stopgradient, indicating that the gradient here will not be backpropagated. Preferably, the Adam optimizer is adopted. The Adam optimizer uses both the first-order momentum and the second-order momentum to guide the model optimization, which can effectively improve the convergence speed and relieve the model from falling into local optima.
[0075] Next, a specific embodiment is combined to illustrate the implementation of the present invention. The question is the question in CommonsenseQA: Where would I not want a fox? (Where can't I get a fox?), and the given options are hen house, england, mountains, english hunt, california, among which the first option hen house is the correct option. Taking this as an embodiment, the present invention is further described in detail so that those skilled in the art can implement it with reference to the text of the specification.
[0076] Step 101: First, it is necessary to load the pre-trained model required for the experiment. In this embodiment, it is implemented using the transformers library of the Huggingface organization based on pytorch. And it is preferably configured with anaconda to ensure that there are matching pytorch and transformers libraries in the environment. The encoder-decoder structure of the T5 model is adopted and downloaded from the official website https: / / huggingface.co / models. First, tokenize the input question. For the training of the generation process, it is necessary to tokenize the input question "Where would I not want a fox?" and the correct option "hen house" respectively, using the T5 tokenizer for tokenization. The T5 tokenizer uses the sentencepiece algorithm for tokenization, so a word may be divided into multiple tokens. Then, use the tokenization result of the question as input_ids to input into the encoder, and use the tokenization result of the correct option "hen house" as labels. Specifically, the token positions filled in the correct option need to be set to -100, so as to ignore these tokens when calculating the loss. After the question and the correct option are input into the model, the model will automatically add BOS and EOS symbols without manual processing. At this time, the generation loss loss can be obtained from the generation stream output of the model g . Then, use the same method to input the tokenization result of the question as input_ids into the encoder, and let the model decoder generate some tokens autoregressively according to the greedy strategy, and at the same time obtain the extended representation A for reading comprehension u , specifically, call the generate method of the T5 model
[0077] Step 102: Concatenate the question with each of the 5 options respectively, tokenize the result using the T5 tokenizer to obtain the tokenization result, and input it into the encoder to obtain 5 question-option representations QO that integrate question information. Respectively, perform the Co-Match operation on each option representation and the extended representation A for reading comprehension u to obtain the question-option representation Map the question-option representation to the score corresponding to each option through a linear layer, use the softmax operation to map to the distribution of the corresponding selected answer, and finally combine the correct option and use the cross-entropy loss function to obtain the final reading comprehension loss loss u .
[0078] Step 103: According to the tokens generated in Step 101, combined with the correct option, that is, hen house, calculate the Rouge value, which is the weight λ in the loss, using the formula:
[0079] L(θ) = λ × loss u +(1 - λ) × loss g
[0080] Obtain the final loss for model update. Use the torch.optim.Adam optimizer to optimize the reading comprehension model.
[0081] In this embodiment, the maximum length of the Encoder input sequence used by the reading comprehension model is set to 32. The part exceeding the length will be removed, and the part shorter than the maximum length will use <pad>The filling operation is performed. The maximum length of the model Decoder is 16, and the batchsize is set to 1. The learning rate is set to 0.00005, the dropout is set to 0.1, the number of training epochs is 20, and the Adam optimizer uses default parameters. The validation set metric used is accuracy. Finally, the model with the highest validation set accuracy is selected to be tested on the test set, and the option with the highest output probability is taken as the model prediction option during testing. Compared with several other existing understanding models based on the T5 encoder-decoder, the present invention has more excellent answering metrics, as shown in Table 1.
[0082] Table 1
[0083] Based on T5-base csqa test set obqa test set T5 60.93 57.53 UnifiedQA 62.35 58.47 The present invention 63.45 61.67 < / pad>
Claims
1. Implementation method for machine reading comprehension assisted by a generative model, characterized in that For natural language reading comprehension of multiple-choice questions, an encoder-decoder based on the sequence-to-sequence model is used to construct a reading comprehension model, which is trained using a question set q, a corresponding option set o, and a correct option set a. The reading comprehension model includes two workflows. One is the generation workflow. The question input encoder obtains the question encoding representation Q, and inputs the question encoding representation Q into the decoder to obtain the answer decoding representation Ag. During training, the teacher-forcing loss is calculated according to the correct option. The other is the reading comprehension workflow. The question encoding representation Q is separately input into the decoder to generate a decoding representation as the vector representation Au for question expansion. At the same time, the question is concatenated with each corresponding option and then input into the encoder to obtain the question-option representation QO. The QO and the expanded vector representation Au are interactively fused through a bidirectional matching layer to obtain a fused representation. After that, the fused representation passes through a linear layer to obtain the logit corresponding to each option. During training, the cross-entropy loss is calculated between these logits and the correct option, and the teacher-forcing loss and the cross-entropy loss are combined to train and optimize the reading comprehension model to obtain a generative reading comprehension model. When training the reading comprehension model, the loss function of the generation stream and the loss function of the reading comprehension stream are combined to obtain the multi-task optimization loss: L(θ) = λ × loss u + (1 - λ) × loss g loss g For generating the flow loss, loss u is the reading comprehension flow loss, θ is the model parameter, and λ is defined as: λ = rouge(tokens, α) λ is the rouge value between the decoded output of the reading comprehension stream and the correct option. Its meaning is that if the tokens generated by the decoded output of the reading comprehension stream are highly similar to the correct option, it indicates that the generation result is good, thus relatively reducing the weight of the generation loss and increasing the weight of the reading comprehension loss. Conversely, if the rouge value of the generated tokens and the correct option is low, it means that the generation effect is poor, so the weight of the generation loss is correspondingly increased, enabling the model to prioritize improving the generation effect and avoiding being affected by generation noise during the interaction and fusion of the bidirectional matching layer; Finally, the gradient descent and error backpropagation algorithms are used to optimize the model, and the Adam optimizer is adopted.
2. The implementation method of machine reading comprehension assisted by a generative model according to claim 1, characterized in that The question set q, the corresponding option set o, and the correct option set a are derived from a single reading comprehension dataset without using external resources.
3. The implementation method of machine reading comprehension assisted by a generative model according to claim 1, characterized in that The teacher-forcing training method is adopted for the generation stream: assuming that the outputs of the previous steps are all correct, the distribution of the next token is predicted at each position to fit the distribution of the actual next token at that position. The mask matrix of the decoder is set to a lower triangular matrix to ensure that each position can only see the tokens before that position and not the tokens after it; among them, when the decoder is trained according to the correct option, the BOS label and the EOS label are added before and after the correct option respectively to mark the start and end of the correct option. The answer decoded representation Ag is mapped to the distribution on the vocabulary through a linear layer and a softmax operation, as shown in the following formula: represents the probability of the t-th word in the vocabulary corresponding to the prediction of the i-th token position; The generation loss, that is, the teacher-forcing loss, is: Let \(n_a\) denote the total number of tokens in the correct option, and \(a\) i denotes the \(i\)-th token.
4. The implementation method of machine reading comprehension assisted by a generative model according to claim 1, characterized in that In the reading comprehension stream, the question encoded representation Q is separately input into the decoder for decoding. Using the autoregressive method, the auxiliary representation for reading comprehension is obtained according to the greedy strategy as follows: A u , tokens = Decoder(Q) tokens refers to all the tokens obtained at each decoding step using the greedy strategy, and Au is the vector representation extended by the question.
5. The implementation method of machine reading comprehension assisted by a generative model according to claim 1, characterized in that In the reading comprehension stream, the question is concatenated with each option respectively and input into the encoder Encoder simultaneously to obtain the question-option representation QO: QO = Encoder(q, o) The Co-Match fusion operation is performed on the question-option representation QO and the vector representation Au extended by the question for interactive fusion: The Co-Match fusion operation realizes interactive fusion through a bidirectional matching layer. Let the two input vectors be: m, n, and h represent the dimensions of the vectors respectively. The similarity matrix S is obtained by using the matrix multiplication method: Among them, the element at the x-th row and y-th column of the similarity matrix S represents the similarity between the x-th word in A and the y-th word in B, which is defined as the inner product of the representations of these two words. After obtaining the similarity matrix, the softmax operation is used to obtain the attention magnitude of each word in B corresponding to each word in A: Similarly, the attention magnitude of each word in A corresponding to each word in B is obtained: According to S b and A, obtain the representation of B updated using A: Splicing and B, with S a perform matrix multiplication to obtain the representation of A that incorporates the information of B: The same method is used to obtain the representation of A updated using B: Concatenate it with A itself and S b Perform matrix multiplication to obtain the representation of B that incorporates the information of A: Finally, combine A and Use the transformation matrix W A Obtain the representation of the final output of A: Similarly, combining B and using the transformation matrix W B obtain the final representation of B: Among them, two transformation matrices W A and W B are model parameters, learned during training, and their dimensions are:
6. The implementation method of machine reading comprehension assisted by a generative model according to claim 1, characterized in that Fusion representation All options are mapped to corresponding logits through a linear layer, defined as follows: The softmax operation is used to map to the probability that each option is selected as the correct option, and the cross-entropy loss function is used to obtain the reading comprehension loss: where the logit in the fraction answer is the logit corresponding to the correct option mapped by the linear layer. The training objective is to make the logit of the correct option as large as possible compared to the logits of other incorrect options, so as to select the correct option.
7. The implementation method of machine reading comprehension assisted by a generative model according to claim 6, characterized in that For single-choice questions, directly map the correct option to obtain the corresponding logit; for multiple-choice questions, concatenate the representations of the T options of the multiple-choice question in order, and use a linear layer to map them to a 2 T -dimensional vector, so that the combinations of various options are mapped to a new option ranking, and the combination ranking formed by multiple correct options is a number between 1 and 2 T -1. The mapping of the correct option combination is as follows: I(f) represents whether the f-th option is the correct option, where it is 1 if yes and 0 if no.
8. An apparatus for machine reading comprehension assisted by a generative model, characterized in that There is a computer-readable storage medium, in which a computer program is configured. When the computer program is executed by a processor, it implements the reading comprehension model according to any one of claims 1-7.
Citation Information
Patent Citations
A shape filling type reading understanding analysis model and method based on reinforcement learning
CN109840322A
Evaluation method for generative questions and answers
CN112818106A