An unsupervised machine reading comprehension method based on large-scale question self-learning
By adopting an unsupervised method of large-scale problem self-learning in the field of machine reading comprehension, using synthetic data sets and problem generators, the problem lack of labeled data in the new field is solved, and the ability to quickly adapt to the new field is achieved.
Patent Information
- Application Number
- CN202111151305.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-08
- Filing Date
- 2021-09-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-09-29
AI Technical Summary
In the absence of labeled data, machine reading comprehension models are difficult to achieve high accuracy in the new field, and existing methods rely on large amounts of labeled data for pre-training, resulting in model overfitting.
The unsupervised machine reading comprehension method based on large-scale problem self-learning is adopted. By dividing the data into four types (unnoted general data, marked general data, unmarked in-domain data, and marked in-domain data), a synthetic data set is generated using a standard pre-trained model and a problem generator, filtering and training, and finally building the final machine reading comprehension model.
The accuracy of the model is significantly improved without any labeled data and quickly reaches high performance levels after injecting a small amount of labeled data, reducing the cost of data annotation.
Smart Images

Figure CN113836895B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine reading comprehension, and in particular to an unsupervised machine reading comprehension method based on large-scale question self-learning. Background Art
[0002] Many state-of-the-art algorithms for Natural Language Processing (NLP) tasks require manually annotated data. In the early days, we usually did not have any domain-specific labeled datasets, and annotating sufficient amounts of such data was often expensive and laborious. As a result, for many NLP applications, even for resource-rich languages like English, there are only labeled data in a few domains.
[0003] In many NLP applications, it is difficult to obtain a large amount of labeled data. Therefore, in many cases, we train models from a small amount of data. However, the trained model is often overfitted and needs to be generalized to unseen data. Therefore, researchers take advantage of large unlabeled datasets by pre-training language models, which often alleviates the problem of random initialization of network weights, thereby finding better local optimal values and improving the robustness of the agent in unseen environments.
[0004] Recent major advances in machine reading comprehension (MRC) have been achieved by pre-training Transformer language models on large amounts of unlabeled text data and fine-tuning the pre-trained models on manually annotated QA datasets. In the context of pre-trained language models, Gururangan shows the importance of additional pre-training with in-domain data to improve downstream task-specific performance. Summary of the invention
[0005] The present invention mainly provides an unsupervised machine reading comprehension method based on large-scale question self-learning, thereby achieving cold start in a completely new field.
[0006] The present invention solves the above technical problems mainly through the following technical solutions: firstly, the data is divided into four types: unlabeled general data, labeled general data, unlabeled in-domain data, and labeled in-domain data, and then the following steps are performed:
[0007] S1. For unlabeled general data, use the standard pre-trained model for training, and obtain the Transformer-based pre-trained language model as the bottom layer of the architecture;
[0008] S2. For the annotated general data, use the pre-trained language model obtained in step S1 to train a question generator, and use the annotated general data to generate a general domain model for a specific task;
[0009] S3. For the unlabeled in-domain data, use the question generator constructed in step S2 to generate synthetic in-domain data, and then use the general domain model for a specific task to filter it. After filtering, a high-quality synthetic in-domain data set and a low-quality synthetic data set are obtained, and then the high-quality synthetic in-domain data set is trained to obtain a new pre-trained model.
[0010] S4. For the labeled in-domain data, the low-quality synthetic dataset obtained by filtering is mixed and the answers are marked, and then the new pre-trained model is used for training to obtain the final machine reading comprehension model;
[0011] Based on the final machine reading comprehension model, the input data obtains the result of machine reading comprehension.
[0012] Preferably, in step S1, a GPT-2 model or a T5 model is used for model learning.
[0013] Preferably, the question generation based on the trained T5 model is specifically as follows: extracting answers; generating questions based on the extracted answers; accepting the questions and generating an answer; comparing the extracted answers with the generated answers to determine whether the generated questions are correct;
[0014] The question generation based on the trained GPT-2 model is as follows: given the natural order of the language, the sequence s = (s1,…,s n ) is decomposed into the product of conditional forms:
[0015]
[0016] After the GPT-2 model is trained, for each new word, the model calculates the probability of the next word based on all existing characters. Then, based on the probability, the top K high-probability words are selected and randomly sampled from these K candidate words. This process is repeated until a special symbol or a sentence-ending symbol appears.
[0017] Generate this scenario for the question, use special symbols to mark the location of potential answers in the source text, for a paragraph C = [c1,...c n ] and one of the potential answers A=[a1,..,a n ], will be represented as:
[0018] X=([CLS],C,[SEP],A)
[0019] Given the above X, we input it into the trained GPT-2 model or the trained T5 to obtain the hidden vector:
[0020] H=Model(x)
[0021]
[0022] X is the input length, h is the size of the hidden vector; finally H will be input into a layer of fully connected network to get the final result:
[0023]
[0024]
[0025] In the formula, w is a word, W is a matrix, b is a coefficient, and the final result is the best word output by argmax. Both W and b are obtained through learning.
[0026] Preferably, in step S3, active learning is performed on the generated data with round-trip consistency, so as to actively screen out the weak links in the distribution of training data according to the advantages and disadvantages of the existing model in different dimensions, and recommend the next batch of data to be marked.
[0027] Preferably, in step S3, data filtering is performed through round-trip consistency, and learning efficiency is improved through active learning.
[0028] The substantial effect brought by the present invention is that it is applicable to situations without any labels or very small labeled data, and significantly improves the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0030] The technical solution of the present invention is further specifically described below through embodiments and in conjunction with the accompanying drawings.
[0031] Example: We use a variety of pre-trained language models (such as GPT-2 and T5) to generate a large amount of potential questions and answers from unlabeled passages of in-domain text, which allows us to achieve a cold start in a completely new domain. We then pre-train the model based on these generated samples and finally fine-tune it on a specific labeled dataset.
[0032] While the domain-specific model trained on the SQuAD1.1 training dataset achieves state-of-the-art performance on the SQuAD1.1 Dev dataset (EM score of 85%), it is completely unable to perform the same level of reasoning on a completely new domain, NewQA (EM score of 32%). We found that when using synthetic datasets to pre-train models, it is critical to prevent overfitting on the synthetic dataset, as it usually contains many noisy samples. However, these synthetic datasets are very useful when there is no or very little in-domain training data early on, as we can automatically generate "machine" annotated training data in a completely new domain through this method.
[0033] With this approach, we achieve 80% final performance without any labeled data. Moreover, when we inject a small amount of labeled data (10% of the original data), the pre-trained model can quickly reach a final performance level equivalent to 94%. Finally, we evaluate DataDream using the NLP Checklist testing framework, which is used to rigorously test NLP models. Our approach reduces errors by 18% in general language proficiency test items in the NLP Checklist (such as synonyms, problem spelling, time changes, etc.).
[0034] Question generation is a long-standing research topic, and using generated question-answer pairs to improve quality assurance systems has shown huge improvements in low-resource settings with only a small number of samples. However, validating and improving the accuracy of these generated QA pairs is relatively unexplored.
[0035] In machine translation, modeling consistency in both translation directions through dual learning or back-translation can improve the quality of translation models. Back-translation adds synthetically generated parallel data as training examples, which is the inspiration of this work and has achieved state-of-the-art performance in both supervised and unsupervised settings. It is possible to model the joint distribution of questions and answers given the context and use this model directly, while our work uses a generative model to generate synthetic data for pre-training. Combining these two approaches may be a fruitful area of future work.
[0036] QG is used to augment training data for question answering and focuses on text-based quality inspection tasks, aiming to select one or more answer sentences from the text of a given input question. The weight of each data point during training is configured by comparing the generated questions with the original questions when ranking the sentences.
[0037] Translation-based data augmentation mechanisms can be introduced to answer questions. However, these methods are highly dependent on the availability and quality of the translation system. Although we can add more data to the training using MT, it still does not improve significantly due to the difficulty in finding domain-specific data in other languages.
[0038] Using Synthetic QA Corpora Generation can improve the overall MRC task through round-trip consistency. In order to achieve round-trip consistency, the model should have been trained. The main difference from our work is that we assume that our dataset is small and it is difficult to build an initial model. However, they assume that they already have a model and when they want to further improve the model. Therefore, it is difficult to show improvements on new domain datasets and it is difficult to consistently improve performance across domains.
[0039] The main contributions of our proposed Data Dream are in four aspects:
[0040] 1. I proposed four steps to build an NLP system for small sample cases.
[0041] 2. We build Synthetic QA Corpora using multiple different heterogeneous pre-trained language models and show performance improvements on new domains.
[0042] 3. We tested it on the NLP Checklist, which is a rigorous test of NLP models. Our proposed method exceeded the accuracy of the baseline and found that the error rate of common language functions was significantly reduced.
[0043] 4. If the predicted answer is different, we further improve the performance by performing active learning on the generated questions.
[0044] Our overall process is divided into four stages according to different datasets. First, for any NLP field or task, we can divide the dataset into four types:
[0045] 1. Unlabeled general data (e.g., BookCorpus, Wikipedia, etc.).
[0046] 2. Annotated general data (or out-of-domain data) (e.g., SquAD, TriviaQA, HotpotQA, etc.).
[0047] 3. Unlabeled in-domain data (e.g., judicial precedents, insurance clauses, technical specifications, etc.).
[0048] 4. Annotated in-domain data (e.g. manually annotated legal files).
[0049] We take different approaches to these four steps based on the size of the dataset:
[0050] Step 1 (Unlabeled General Data): Research on unlabeled general domain datasets has been actively conducted for 3 years. A large amount of text data is used to build transformer-based pre-trained language, and pre-trained models such as BERT, GPT-2, T5 have become standard NLP processing. We use a transformer-based language model as the lowest level of our architecture.
[0051] Step 2 (Annotated General Data): Our goal is to build machine reading comprehension models, so there are many publicly available datasets. We use this dataset to make a synthetic data generator to make a large-scale in-domain dataset. In addition, we use labeled domain-general datasets to make task-specific (MRC tasks in this work) domain-general models.
[0052] Step 3 (Unlabeled Industry Data): We generate a lot of synthetic in-domain data using the question generator built in step 2. After generating a lot of data, we filter it using the domain model model, which borrows the idea of round-trip consistent question generation. High-quality samples will be used to build a pre-trained model, and we use further filtering methods to improve performance. And when we manually annotate this data, we can use the pre-trained model as an annotation assistant.
[0053] Step 4 (Annotated Industry Data): In the last step, we apply active learning, which uses negative synthetic data from the general model to send to human annotators to label the answers. If the generated questions are ungrammatical and difficult to understand, we ask the annotators to modify the generated questions as much as possible and annotate the answers. Finally, we train the final model using the in-domain labeled dataset.
[0054] In the following sections, we will explain in detail how to implement each step.
[0055] Step 1: Self-learning with unlabeled general data
[0056] In this step, we adopted two different strategies to learn models for unlabeled general data. The first method was proposed by GPT-2. GPT-2 is a large-scale transformer-based language model released by OpenAI in February 2019. It contains 1.5 billion parameters and is trained on a dataset of 8 million web pages. This model is a direct extension of the GPT model, trained on 10 times more data and 10 times more parameters. In terms of performance, the model is able to produce coherent text paragraphs and has achieved SOTA performance on many language modeling benchmarks. Moreover, the model is able to achieve preliminary reading comprehension, machine translation, question answering, and automatic summarization without task-specific training.
[0057] The second strategy is proposed by T5. T5's training data includes the Colossal Clean Crawled Corpus (C4 Corpus), which crawls hundreds of gigabytes of clean English text from the Common Crawl website. The T5 model is a standard Transformer-based Encoder-Decoder model with 11 billion model parameters.
[0058] Step 2: Generate a model by training labeled general data.
[0059] Question generation is the task of automatically generating questions based on a text passage. The simplest approach is question answering. In question generation where the answer is known, the model is provided with the answer and the passage and asked to generate a question for that answer by taking into account the passage context. One reason is that most of the older papers use complex models / processing pipelines and there are no pre-trained models available. As a result, machine-generated questions are often ungrammatical and difficult to understand, making it difficult to use the generated data in real applications. However, recent advances in text generation techniques powered by pre-trained transformer models allow us to generate plausible synthetic data. We used the most powerful generation methods available off the shelf: T5-based generation and GPT-2-based generation.
[0060] Question Generation Based on T5: T5 is a very large new neural network model that was trained on a mixture of unlabeled text and labeled data from popular natural language processing tasks and then fine-tuned individually for each task its authors wanted to solve. T5 is an encoder-decoder model pre-trained on a multi-task mixture of unsupervised and supervised tasks, for which each task was converted to a text-to-text format. In order to generate questions where we know the answer, we usually need 3 models, the first will extract the answer like a span, the second model will generate questions on that answer, and the third will be a QA model that will take the question and produce an answer, and then we can compare the two answers to see if the generated question is correct. Having 3 models for a single task is very complex, so the goal is to create a multi-task model that can do all 3 tasks at the same time.
[0061] GPT-2-based question generation: For question generation using GPT-2, we follow the original standard text generation strategy. Given the natural order of the language model, the joint probability of the sequence s = (sl, .., sn) can be decomposed into a product of conditional
[0062]
[0063] After the above probability model training is completed, the question generation part can be implemented through a variety of random sampling strategies, including sequence top-k. For each new word, the model calculates the probability of the next word based on all existing characters. Then, based on the probability, the top K high-probability words are selected and randomly sampled from these K candidate words. This process is repeated until special symbols, including "?" or sentence end symbols appear.
[0064] In addition, for the question generation scenario, we use special symbols to mark the location of potential answers in the source text. For example, for a paragraph C = [C1, ...C n ] and one of the potential answers A=[a1....A n ], will be represented as:
[0065] X=([CLS],C,[SEP],A)
[0066] Given the above X, we can input it into GPT-2 or T5 to get the hidden vector:
[0067] H=Model(x)
[0068]
[0069] X is the input length, h is the size of the hidden vector. Finally, H will be input into a layer of fully connected network to get the final result:
[0070]
[0071]
[0072] Step 3: Label industry data using the model from step 2.
[0073] A lot of manual annotation is required for AI model training. But the process of manual annotation is costly. In addition, it is difficult for annotators to decide what to ask in machine learning understanding, and human annotators have a lot of duplicates. If the active learner produces the most disagreement in prediction, it decides to query the oracle to label the data sample. This can be measured by entropy and KL-divergence. High variance in the output predictions indicates the most informative data samples. In this paper, we perform active learning on generated data with round-trip consistency, so as to actively screen out the weak links in the distribution of training data according to the strengths and weaknesses of existing models in different dimensions, and recommend the next batch of data that should be labeled, thereby reducing the cost of data labeling and increasing the value of each manually labeled data point.
[0074] After obtaining the above question generation model, we can use the question generation model based on T5 or GPT-2 to annotate any unlabeled industry data, automatically generate potential related questions, and use them to train the question-answering model in the fourth step. However, directly using all generated questions to achieve training results is not ideal because the generated questions contain a lot of noise. Therefore, we invented a round-trip consistency method to control data quality.
[0075] Data filtering by round-trip consistency: Round-trip consistency can be used to filter data. If the model cannot answer the generated questions, examples can be filtered. We also adopted this method to filter data. However, there are some differences between existing works and our work:
[0076] We assume that there is no Xunlei Chain data, so our MRC is trained directly on the generated data.
[0077] Their approach assumes the availability of training data and aims to improve performance given the availability of training data.
[0078] During training, we use the indicator function I(q):
[0079]
[0080] in It is a generated problem. is the given answer. And I(q) is used to filter whether a data point is used.
[0081] Improving learning efficiency through active learning: First, we generate questions about named entities or noun phrases, and then run the trained MRC model from the general domain. If the model cannot predict the answer, we save all samples for active learning. We select the data that the model is least confident about for active learning through the following strategy.
[0082]
[0083] in It is a generated problem. and is the given answer and context. And I(q) is used to filter whether a data point is used.
[0084] Step 4: Fine-tune with labeled industry data
[0085] Training Details: When we have a large amount of in-domain unlabeled data in most cases in the real world, our method can perform task-specific pre-training on those large unlabeled datasets. The whole training process follows the following steps;
[0086] 1. Build multiple question generators from multiple pre-trained language models (e.g., GPT-2 and T5) on publicly available QA datasets (e.g., SQUAD, NQ, and MARCO).
[0087] 2. Use the question generator to generate a large number of questions.
[0088] 3. Use the generated dataset for pre-training.
[0089] 4. Fine-tune the model from the previous step on the labeled dataset.
[0090] We used a large generated QA dataset for pre-training. We used the Span Bert architecture for pre-training and fine-tuning. The target function of the fine-tuning process is to reduce the training error using only labeled data. The main purpose of the fine-tuning step is to re-adjust the weights, which may have been incorrectly trained due to generation errors.
[0091] The final model is evaluated using SQuAD and NewsQA. SQuAD is used to explore the effect of in-domain QG pre-training, which means using the same dataset for question generation and span prediction models. In order to validate a new domain that is completely different from the source of the QG model, we assume that the NewsQA dataset is a new domain dataset and does not contain any training, neither generating questions nor pre-training. The evaluation metrics include standard MRC metrics: EM and F1 score.
[0092] Exact Match (EM): The range of top-1 answers exactly matches the correct answer.
[0093] Fl-Score: We compute the word overlap between the returned span and the ground truth answer at the word level.
[0094] In-Domain vs. Out-of-Domain: Recent natural language processing models have achieved impressive performance when trained and tested on examples from the same dataset, but often perform poorly on out-of-domain (OOD) examples because many unseen events occur in testing.
[0095] We use the SpanBERT architecture, which focuses on pre-training span representations to achieve current state-of-the-art results, to show how the performance differs between in-domain and out-of-domain datasets. We assume
[0096] The SQuADl.l training dataset is an in-domain dataset generated and pre-trained using training questions. We use the NewsQA dataset as the out-of-domain corpus, which does not contain any training samples. We found that the EM score decreased by 78.5% (80.40% -> 17.26%) when testing on out-of-domain data0 However, with the help of the question generator, we are able to generate in-domain QA data on unlabeled samples. On SQuAD1.1, without any labeled data, we can achieve 75% of the final performance, and on NewsQA, without any labeled data, we can achieve 60% of the final performance. Because we included the SQuAD1.1 training data when constructing the question generation.
[0097] Checklist Evaluation: Although measuring the accuracy of retention has been the main method for evaluating generalization, it often overestimates the performance of NLP models, while alternative methods of evaluating models focus on a single task or specific behavior. Inspired by the principles of behavioral testing in software engineering, we introduce CheckList, a task-agnostic approach for testing NLP models. Checklist includes a matrix of general language capabilities and test types that facilitates comprehensive test conception. They illustrate the practicality of Checklist by testing three tasks, identifying critical failures in commercial models and state-of-the-art models. The proposed method, based on pre-training on question generation, achieves an 18% reduction in failure rate, especially in Animals vs Vehicles v2 (39% reduction), Fairness (44% reduction), and Time (93% reduction).
[0098] Effect of labeled data size: To explore the effectiveness of pre-training on data size, we tested QG pre-training on 10% of the dataset and 100% of the dataset. The results show that when we have enough dataset, the model converges faster, but there is not much difference in the final score. This suggests that QG pre-training is more useful in the early stage than in the later stage.
[0099] Effect of generated data size We found that pre-trained models using T5-based generation performed better than GPT-2-based generation. However, when we added both data together, performance improved significantly. Generated questions are often longer than humans, and using both GPT and T5 generation, we were able to add more diverse questions and answers for training. On the same answer “Moninder Singh Pandher”, are T5, GPT, and humans fully questioned? (T5: Who was sentenced to death by the lower court?”, GPT: The craftsman who killed the teenager?, Human: Who was acquitted?”) There was only a small overlap of words between each model. Therefore, model diversity improves the generalization of subsequent MRC models.
[0100] The specific embodiments described herein are merely examples of the spirit of the present invention. Those skilled in the art may make various modifications or additions to the specific embodiments described or replace them in similar ways, but they will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
[0101] Although the term "marking" and "domain" are used more frequently in this article, the possibility of using other terms is not excluded. These terms are used only to more conveniently describe and explain the essence of the present invention; interpreting them as any additional restrictions is contrary to the spirit of the present invention.
Claims
1. An unsupervised machine reading comprehension method based on large-scale question self-learning, characterized in that: First, divide the data into four types: unlabeled general data, labeled general data, unlabeled in-domain data, and labeled in-domain data, and then follow the steps below: S1. For unlabeled general data, use the standard pre-trained model for training, and obtain the Transformer-based pre-trained language model as the bottom layer of the architecture; S2. For the annotated general data, use the pre-trained language model obtained in step S1 to train a question generator, and use the annotated general data to generate a general domain model for a specific task; S3. For the unlabeled in-domain data, use the question generator constructed in step S2 to generate synthetic in-domain data, and then use the general domain model for a specific task to filter it. After filtering, a high-quality synthetic in-domain data set and a low-quality synthetic data set are obtained, and then the high-quality synthetic in-domain data set is trained to obtain a new pre-trained model. S4. For the labeled in-domain data, the low-quality synthetic dataset obtained by filtering is mixed and the answers are marked, and then the new pre-trained model is used for training to obtain the final machine reading comprehension model; Based on the final machine reading comprehension model, input data to obtain the result of machine reading comprehension; In step S1, the standard pre-trained model is the GPT-2 model or the T5 model; The question generation based on the trained T5 model is as follows: extract the answer; generate a question based on the extracted answer; accept the question and generate an answer; compare the extracted answer with the generated answer to determine whether the generated question is correct; The question generation based on the trained GPT-2 model is as follows: given the natural order of the language, the sequence s = (s1,…,s n ) is decomposed into the product of conditional forms: After the GPT-2 model training is completed, for each new word, the model calculates the probability of the next word based on all existing characters; then, based on the probability, the top K high-probability words are selected and randomly sampled from these K candidate words; This process is repeated until a special symbol or a sentence end symbol appears; Generate this scenario for the question, use special symbols to mark the location of potential answers in the source text, for a paragraph C = [c1,...c n ] and one of the potential answers A=[a1,..,a n ], will be represented as: X=([CLS],C,[SEP],A) Given the above X, we input it into the trained GPT-2 model or the trained T5 to obtain the hidden vector: X is the input length, h is the size of the hidden vector; finally H will be input into a layer of fully connected network to get the final result: In the formula, w is a word, W is a matrix, b is a coefficient, and the final result is the best word output by argmax.
2. The unsupervised machine reading comprehension method based on large-scale question self-learning according to claim 1, characterized in that: In step S3, active learning is performed on the generated data with round-trip consistency, so as to actively screen out the weak links in the distribution of training data based on the advantages and disadvantages of the existing model in different dimensions, and recommend the next batch of data to be labeled.
3. The unsupervised machine reading comprehension method based on large-scale question self-learning according to claim 2, characterized in that: In step S3, data filtering is performed through round-trip consistency, and learning efficiency is improved through active learning.
Citation Information
Patent Citations
Method for improving English-Chinese machine translation quality on basis of data selection
CN106844356A
Joint modeling method for spoken language understanding model and language model and dialogue method
CN108962224A