Answer span correction
By constructing a new training set and AI corrector model, detecting and correcting the prediction error of the MRC system, the problems that existing MRC systems produce partially correct answers are solved, improving the accuracy and quality of the answer span, especially in multilingual question-and-answer tasks.
Patent Information
- Application Number
- CN202180071355.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-05
- Filing Date
- 2021-10-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-10-21
AI Technical Summary
When facing answerable questions, existing machine reading comprehension (MRC) systems are prone to produce partially correct answers and lack effective answer span correction mechanisms.
By building a new training set, using the AI corrector model to detect and correct the prediction error of the MRC model, using neural networks to generate improved answer spans, combined with the BERT model for pre-training and fine-tuning, and building a reader plus corrector pipeline to improve the accuracy of the answer.
The accuracy and quality of the MRC system in response span correction is significantly improved, especially in multilingual Q&A tasks, and enhanced the detection and correction capabilities of different error categories.
Smart Images

Figure CN116324929B_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] Embodiments of the present invention relate to answer span correction for machine reading comprehension (MRC) models and systems.
[0002] Answer validation in machine reading comprehension (MRC) involves verifying the extracted answer based on the input context and question pair. Traditional systems address the re-evaluation of the "answerability" of a question given the extracted answer. In the face of answerable questions, traditional MRC systems tend to produce partially correct answers. SUMMARY OF THE INVENTION
[0003] Embodiments relate to answer span correction for machine reading comprehension (MRC) models and systems. One embodiment provides a method for using a computing device to improve answers generated by a natural language question answering system. The method includes receiving, by the computing device, a plurality of questions in the natural language question answering system. The computing device also generates a plurality of answers to the plurality of questions. The computing device also constructs a new training set using the generated plurality of answers, where each answer is compared to the corresponding question among the plurality of questions. The computing device additionally augments the new training set using one or more markers that delimit the span of one or more of the generated plurality of answers. The computing device also trains a new natural language question answering system using the augmented new training set. Embodiments significantly improve the predictions of prior art English readers in different error categories via correction. For MRC systems, some features contribute to the advantage of answer correction, as existing MRC systems tend to produce partially correct answers in the face of answerable questions. Some features contribute to the advantage of detecting and correcting errors in the predictions of the MRC model. Additionally, these features contribute to the advantage of producing answer spans that better match the ground truth (GT), and thus improve the quality of the MRC output answers.
[0004] One or more of the following features may be included. In some embodiments, the augmented new training set for the new natural language question answering system is used to correct the answer span of the reader model of the natural language question answering system.
[0005] In some embodiments, the new natural language question answering system that corrects the answer span is cascaded after the natural language question answering system.
[0006] In one or more embodiments, the method may further include determining, by the new natural language question answering system that corrects the answer span, whether the answer span should be corrected. Additionally, the corrector model of the new natural language question answering system corrects the answer span, producing an improved answer span.
[0007] In some embodiments, the method may additionally include the corrector model using a neural network to generate the improved answer span.
[0008] In one or more embodiments, the method may include determining that an answer span should not be corrected, indicating no need for correction based on defining the GT answer as an input to a new natural language question answering system, and creating new example answers from each original answer among multiple answers.
[0009] In some embodiments, the method may further include using a plurality of the top-k incorrect answer predictions to create example answers for each incorrect answer prediction, where the input is the predicted answer span of a reader model and the target answer is the GT answer.
[0010] In one or more embodiments, the method may additionally include that a plurality of generated answers includes predicted answers, and one or more tokens tokenize the predicted plurality of answers in context for a corrector model to predict new answers.
[0011] These and other features, aspects, and advantages of the present embodiments will be understood with reference to the following specification, the appended claims, and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A cloud computing environment according to an embodiment is depicted;
[0013] Figure 2 A set of abstract model layers according to an embodiment is depicted;
[0014] Figure 3 is a network architecture of a system for answer span correction for performance improvement in machine reading comprehension (MRC) according to an embodiment;
[0015] Figure 4 Shows an example that can be associated with Figure 1 a server and / or client of
[0016] Figure 5 is a block diagram of a distributed system for answer span correction for performance improvement in MRC according to an embodiment;
[0017] Figure 6 Shows an example of a multi-layer bidirectional Transformer encoder (BERT) MRC system;
[0018] Figure 7 Shows an example of a single answer result from a reader of a traditional MRC model or system (given a question and context) and an answer result from a reader plus corrector pipeline (given a question and context) according to an embodiment;
[0019] Figure 8AShows a representative example of training data partitioned into folds for a traditional MRC model or system;
[0020] Figure 8B Shows a representative example of how n-1 individual folds are grouped according to an embodiment to train a separate MRC model for generating predictions for the remaining folds;
[0021] Figure 9 Illustrates a block diagram of the flow of a reader-plus-corrector pipeline of a modified MRC answer span corrector model according to an embodiment;
[0022] Figure 10A Shows according to an embodiment including for Figure 9 Table showing results for answerable questions on the Natural Questions (NQ) MRC benchmark for the Robustly Optimized BERT Approach (RoBERTa) shown, an ensemble method of two readers, and a method using a reader-plus-corrector pipeline (or modified MRC answer span corrector model);
[0023] Figure 10B Shows according to an embodiment including using a reader-plus-corrector pipeline (the modified MRC answer span corrector model shown in Figure 9 and results on the Multilingual Question Answering (MLQA) MRC benchmark dataset for the General Cross-Lingual Transfer task (G-XLT) when the passage is in English;
[0024] Figure 10C Shows according to an embodiment a table including the differences in exact match scores for all 49 MLQA language pair combinations from using a reader-plus-corrector pipeline (modified MRC answer span corrector model); and
[0025] Figure 11 Illustrates a block diagram of a process for answer span correction for performance improvement in MRC according to an embodiment. Detailed Description
[0026] The description of the various embodiments has been presented for purposes of illustration but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments, practical application, or technical improvement found in the marketplace, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.
[0027] Embodiments relate to answer span correction for machine reading comprehension (MRC) models and systems. One embodiment provides a method for improving answers generated by a natural language question answering system using a computing device, including receiving, by the computing device, a plurality of questions in the natural language question answering system. The computing device also generates a plurality of answers to the plurality of questions. The computing device also uses the generated plurality of answers to construct a new training set, where each answer is compared to the corresponding question in the plurality of questions. The computing device additionally uses one or more markers that delimit the span of one or more of the generated plurality of answers to augment the new training set. The computing device also uses the augmented new training set to train a new natural language question answering system.
[0028] One or more embodiments include a corrector that utilizes an artificial intelligence (AI) model (e.g., corrector 960( Figure 9 ))). The AI model can include a trained ML model (e.g., models such as NN, convolutional NN (CNN), recurrent NN (RNN), long short-term memory (LSTM)-based NN, gated recurrent unit (GRU)-based RNN, tree-based CNN, self-attention network (e.g., an NN that uses an attention mechanism as a basic building block; self-attention networks have shown to be effective for sequence modeling tasks without recursion or convolution), BiLSTM (bidirectional LSTM), etc.). An artificial NN is a set of interconnected nodes or neurons.
[0029] It is understood in advance that although the present disclosure includes a detailed description of cloud computing, the implementation of the teachings described herein is not limited to a cloud computing environment. Instead, the embodiments in this example can be implemented in conjunction with any other type of computing environment now known or later developed.
[0030] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processors, memory, storage devices, applications, virtual machines (VMs), and services), which can be rapidly provisioned and released with minimal management effort or interaction with the service provider. The cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0031] The characteristics are as follows:
[0032] On-demand self-service: Cloud consumers can unilaterally and automatically provision computing capabilities, such as server time and network storage, without human interaction with the service provider.
[0033] Broad network access: The capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0034] Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically assigned and reassigned according to demand. There is a sense of location independence in that consumers generally do not control or know the exact location of the provided resources, but can specify location at a higher level of abstraction (e.g., country, state, or data center).
[0035] Rapid elasticity: In some cases, capabilities can be provided quickly and elastically, scaled out rapidly, and released quickly to scale in rapidly. To the consumer, the capabilities available for provisioning generally appear unlimited and can be purchased in any quantity at any time.
[0036] Measured service: Cloud systems automatically control and optimize resource use by leveraging metering capabilities appropriate to the service type at some level of abstraction (e.g., storage, processing, bandwidth, and active consumer accounts). Resource use can be monitored, controlled, and reported, providing transparency for both the provider and consumer of the utilized service.
[0037] The service models are as follows:
[0038] Software as a Service (SaaS): The capabilities provided to the consumer are the ability to use the provider's applications running on a cloud infrastructure. The applications can be accessed from various client devices through a thin client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even the individual application capabilities, with the possible exception of limited user-specific application configuration settings.
[0039] Platform as a Service (PaaS): The capabilities provided to the consumer are the ability to deploy onto the cloud infrastructure consumer-created or acquired applications that are created using programming languages and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but has control over the deployed applications and possibly the application hosting environment configuration.
[0040] Infrastructure as a Service (IaaS): The ability provided to the consumer is the ability to provide processing, storage devices, networks, and other basic computing resources on which the consumer can deploy and run any software, which may include operating systems and application programs. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating system, storage devices, deployed application programs, and limited control over the networking components that may be selected (e.g., host firewall).
[0041] The deployment models are as follows:
[0042] Private cloud: The cloud infrastructure is only used for organizational operations. It can be managed by the organization or a third party and can exist on-premises or off-premises.
[0043] Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by the organization or a third party and can exist on-premises or off-premises.
[0044] Public cloud: The cloud infrastructure is available to the general public or a large industrial group and is owned by the organization selling the cloud services.
[0045] Hybrid cloud: The cloud infrastructure is a combination of two or more clouds (private, community, or public), which remains a single entity but is bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).
[0046] The cloud computing environment is service-oriented, with a focus on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is the infrastructure of a network consisting of interconnected nodes.
[0047] Now referring to Figure 1 , an illustrative cloud computing environment 50 is depicted. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices (such as, for example, a personal digital assistant (PDA) or cellular phone 54A, desktop computer 54B, laptop computer 54C, and / or in-vehicle computer system 54N) used by cloud consumers can communicate. The nodes 10 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as a private cloud, community cloud, public cloud, or hybrid cloud or a combination thereof as described above. This allows the cloud computing environment 50 to provide infrastructure, platform, and / or software as a service, for which the cloud consumer does not need to maintain resources on local computing devices. It should be understood that Figure 1The types of computing devices 54A - 54N shown are only illustrative, and the computing nodes 10 and the cloud computing environment 50 can communicate with any type of computerized device via any type of network and / or network - addressable connection (e.g., using a web browser).
[0048] Now referring to Figure 2 , a set of functional abstraction layers provided by the cloud computing environment 50 ( Figure 1 ) is shown. It should be understood upfront that Figure 2 the components, layers, and functions shown are only illustrative, and the embodiments are not limited thereto. As depicted, the following layers and corresponding functions are provided:
[0049] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: host 61; servers 62 based on RISC (Reduced Instruction Set Computer) architecture; server 63; blade server 64; storage device 65; and network and network components 66. In some embodiments, software components include web application server software 67 and database software 68.
[0050] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 71; virtual storage device 72; virtual network 73 including virtual private networks; virtual applications and operating systems 74; and virtual clients 75.
[0051] In one example, the management layer 80 can provide the functions described below. Resource provisioning 81 provides for the dynamic procurement of computing resources and other resources for tasks to be executed within the cloud computing environment. Metering and pricing 82 provides cost tracking when resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources can include application software licenses. Security provides authentication for cloud consumers and tasks and protection for data and other resources. The user portal 83 provides access to the cloud computing environment for consumers and system administrators. Service - level management 84 provides cloud computing resource allocation and management such that the required service levels are met. Service - level agreement (SLA) planning and fulfillment 85 provides pre - arrangement and procurement of cloud computing resources in anticipation of future demands according to the SLA.
[0052] The workload layer 90 provides examples of functions that can utilize the cloud computing environment. Examples of workloads and functions that can be provided from this layer include: graphics and navigation 91; software development and lifecycle management 92; virtual classroom education delivery 93; data analysis processing 94; transaction processing 95; and answer span correction for performance improvement for MRC processing 96 (see, for example, Figure 5 system 500, reader plus corrector pipeline (modified MRC answer corrector model 900) andFigure 11 Process 1100). As mentioned above, regarding Figure 2 All of the foregoing examples described are merely illustrative, and the embodiments are not limited to these examples.
[0053] It is reiterated that although this disclosure includes a detailed description of cloud computing, the implementation of the teachings set forth herein is not limited to a cloud computing environment. Instead, embodiments can be implemented using any type of cluster computing environment now known or later developed.
[0054] Figure 3 is a network architecture of a system 300 for answer span correction for performance improvement of an MRC model according to an embodiment. As Figure 3 shown, a plurality of remote networks 302 are provided, including a first remote network 304 and a second remote network 306. A gateway 301 can be coupled between the remote networks 302 and a neighboring network 308. In the context of this network architecture 300, the networks 304, 306 can each take any form, including but not limited to a LAN, a WAN such as the Internet, a public switched telephone network (PSTN), an internal telephone network, etc.
[0055] In use, the gateway 301 serves as an entry point from the remote networks 302 to the neighboring network 308. Thus, the gateway 301 can serve as a router capable of directing a given data packet arriving at the gateway 301 and a switch providing an actual path for the given packet in and out of the gateway 301.
[0056] Also included is at least one data server 314 coupled to the neighboring network 308, and the at least one data server 314 can be accessed from the remote networks 302 via the gateway 301. It should be noted that the (one or more) data servers 314 can include any type of computing device / component. A plurality of user devices 316 are coupled to each data server 314. Such user devices 316 can include desktop computers, laptop computers, handheld computers, printers, and / or any other type of device incorporating logic. It should be noted that in some embodiments, the user devices 316 can also be directly coupled to any network.
[0057] Peripheral devices 320 or a series of peripheral devices 320, such as fax machines, printers, scanners, hard disk drives, networking and / or local storage units or systems, etc., can be coupled to one or more of the networks 304, 306, 308. It should be noted that databases and / or additional components can be used with or integrated into any type of network element coupled to the networks 304, 306, 308. In the context of this specification, a network element can refer to any component of a network.
[0058] According to some methods, the methods and systems described herein can be utilized and / or implemented on virtual systems and / or systems that emulate one or more other systems, such as emulating the system of the environment, virtually hosting the system of the environment, emulating the
[0059] Figure 4 FIG. shows a representative hardware system 400 environment associated with Figure 3 a user device 316 and / or a server 314 according to one embodiment. In one example, the hardware configuration includes a workstation having a central processing unit 410 such as a microprocessor, and a plurality of other units interconnected via a system bus 412. Figure 4 The workstation shown may include a random access memory (RAM) 414, a read-only memory (ROM) 416, and an I / O adapter 418 for connecting peripheral devices such as a disk storage unit 420 to the bus 412, a user interface adapter 422 for connecting a keyboard 424, a mouse 426, speakers 428, a microphone 432, and / or other user interface devices such as a touch screen, a digital camera (not shown) to the bus 412, a communication adapter 434 for connecting the workstation to a communication network 435 (e.g., a data processing network), and a display adapter 436 for connecting the bus 412 to a display device 438.
[0060] In one example, the workstation may have an operating system resident thereon, such as Operating System (OS), MA OS, etc. In one embodiment, the system 400 employs a -based file system. It will be understood that other examples can be implemented on platforms and operating systems in addition to those mentioned. Such other examples may include operating systems written in XML, C, and / or C++ languages or other programming languages, as well as object-oriented programming methods. Object-oriented programming (OOP), which has become increasingly used for developing complex applications, can also be used.
[0061] Figure 5FIG. 0 is a block diagram illustrating a distributed system 500 for answer span correction for performance improvement of an MRC model according to one embodiment. In one embodiment, the system 500 includes a client device 510 (e.g., a mobile device, a smart device, a computing system, etc.), a cloud or resource sharing environment 520 (e.g., a public cloud computing environment, a private cloud computing environment, a data center, etc.), and a server 530. In one embodiment, cloud services from the server 530 are provided to the client device 510 through the cloud or resource sharing environment 520.
[0062] Instead of improving the predictability of answerability of a question given an extracted answer solved by a traditional system, one or more embodiments address the problem that existing MRC systems tend to produce partially correct answers when faced with answerable questions. One embodiment provides an AI correction model that rechecks the extracted answer in context to suggest corrections. One embodiment uses the same labeled data used to train the MRC model to build training data for training such an AI correction model. According to one embodiment, the corrector detects errors in the predictions of the MRC model and also corrects the detected errors.
[0063] Figure 6 FIG. 7 shows an example of an MRC system 600 with a multi-layer bidirectional Transformer encoder (BERT) model 630. The BERT model 630 is a multi-layer bidirectional Transformer encoder. Traditional neural machine translation mainly uses RNN or CNN as the model basis for the encoder-decoder architecture. The attention-based Transformer model abandons the traditional RNN and CNN formulas. The attention mechanism is in the form of a fuzzy memory that includes the hidden state of the model. The model selects content to retrieve from the memory. The attention mechanism reduces this problem by allowing the decoder to look back at the source sequence hidden states and then providing their weighted average as an additional input to the decoder. Using attention, the model selects the context most suitable for the current node as the input during the decoding phase. The Transformer model uses an encoder-decoder architecture. The BERT model 630 is a deep bidirectional DNN model. The BERT model 630 applies bidirectional training of the Transformer to language modeling. The Transformer includes an encoder that reads the text input and a decoder that produces predictions for the task. There are two stages for using the BERT model 630: pre-training and fine-tuning. During pre-training, the BERT model 630 is trained on unlabeled data on different pre-training tasks. For fine-tuning, the BERT model 630 is first initialized with pre-trained parameters, and all parameters are fine-tuned using labeled data from downstream tasks. Each downstream task has a separate Transformer (fine-tuned) model 625, even though they are initialized with the same pre-trained parameters.
[0064] The pre-training stage of the BERT model 630 includes masked language model and next sentence prediction. For the masked language model resulting from bidirectionality and the effect of the multi-layer self-attention mechanism used in the BERT model 630, to train a deep bidirectional representation, a percentage (e.g., 15%) of the input tokens are randomly masked, and then the masked tokens are predicted. Similar to a standard language model, the final hidden vector corresponding to the masked token is fed into an output softmax function on the vocabulary (the softmax function turns a vector of K real values into a vector of K real values whose sum is 1). The masked language model objective allows the representation to fuse the context on the left and right sides, which makes it possible to pre-train a deep bidirectional transformer. The BERT model 630 loss function only considers the prediction of masked values and ignores the prediction of unmasked words. For next sentence prediction, the BERT model 630 is also pre-trained for a binary next sentence prediction task, which can be very easily generated from any text corpus. To help the BERT model 630 distinguish between two sentences during training, the input is processed as follows before entering the BERT model 630. The classification [CLS] token 605 is inserted at the beginning of the question 610 (i.e., the first sentence or sentence A), while the separator [SEP] token 615 is inserted at the end of the question 610 and the context (the second sentence or sentence B) 620. The sentence embedding (E) indicating the question 610 or the context 620 is added to each token (e.g., E [CLS] , E [SEP ). Position embeddings (e.g., E1-E N , E'1-E' M )) are added to each token to indicate its position in the sequence.
[0065] To predict whether the context 620 is connected to the question 610, the entire input sequence is passed through the transformer model 625. The output of the [CLS] token 605 is transformed into a vector of shape 2×1 using a classification layer (a learned matrix of weights and biases). The probability of IsNextSequence is determined using the softmax function. For each downstream natural language processing (NLP) task, task-specific inputs and outputs are fed into the BERT model 630, and all parameters are fine-tuned end-to-end. At the input, the pre-trained question 610 and context 620 can be similar to sentence pairs in paraphrasing, hypothesis-premise pairs in entailment, question-passage pairs in question answering, etc. At the output, the token representations are fed into an output layer for token-level tasks, such as sequence tagging or question answering, and the [CLS] representation is fed into an output layer for classification (e.g., output class label C 635). The output layer includes the transformer outputs T1-T N 640 and T [SEP] , T'1-T'M and T [SEP] , which are the answer "start and end" span position classifiers 645.
[0066] Figure 7 shows an example of a reader from a traditional MRC system (e.g., Figure 9 reader 930; given question and context) and an answer result from a reader plus corrector pipeline (MRC answer span corrector model 900( Figure 9 )). The first example includes question 710, result 715 in the context, answer result (R) 716 from the reader using the context, and answer result (R+C) 717 from the reader plus corrector pipeline using the context. The second example includes question 720, result 725 in the context, R 726, and R+C 727.
[0067] Figure 8A shows a representative example 800 of training data divided into multiple folds for a traditional MRC model or system. The training data is divided or parsed into n folds, fold1 810 to fold n 820.
[0068] Figure 8B shows a representative example 830 of how n-1 individual folds 835 are grouped to train separate MRC models (n MRC answer span corrector models 900( Figure 9 )) for making predictions about the remaining folds 840. The n MRC answer span corrector models 900 are each trained on n-1 (where n is an integer greater than or equal to 2) different folds and use them to make predictions about the remaining folds 840. The results from the n different models on the omitted folds are combined to produce an example of the system output for each example in the training set. These [question-context-ground truth answer-system answer] tuples are the basis for constructing the training set of the answer span corrector model 900.
[0069] Figure 9The figure shows a block diagram of the flow of the reader-plus-corrector pipeline of a modified MRC answer span corrector model 900 according to one embodiment. In one embodiment, the output of the reader model 930 is input to the corrector (or corrector model) 960. In one embodiment, the MRC model (reader model 930) includes a transformer-like encoder with two additional classification heads that respectively select the start and end of the answer span. In this embodiment, the answer span corrector (corrector 960) also has a similar architecture. The corrector 960 is trained with data different from that of the reader 930. The modified MRC answer span corrector model 900 rechecks the reader answer 940 (extracted answer) in context to suggest corrections to address issues related to improving the answer span and outputs a corrected answer 970. In one embodiment, the reader answer 940 is delimited with special boundary tokens [T d 950 and [T d 951, and the trained corrector 960 (having an architecture similar to that of the original reader 930) is employed to produce new accurate predictions.
[0070] In one embodiment, the reader 930 is a baseline reader for the standard MRC task of answer extraction from a passage 920 given a question 910. The reader 930 uses two classification heads on top of a pre-trained transformer-based language model, pointing to the start and end positions of the answer span. Subsequently, the entire network is fine-tuned on the target MRC training data. In one embodiment, the input to the corrector 960 contains the boundary tokens [T d 950 and [T d 951 that mark the predictions of the reader (reader answer 940), and the rest of the architecture is similar to the input of the reader 930. In one embodiment, it is desired that the modified MRC answer span corrector model 900 keeps intact the answers that already match the ground truth (GT) spans and corrects the rest of the answers.
[0071] In one embodiment, to generate the training data for the corrector 960, the training set requires the predictions of the reader 930. To obtain the reader 930 predictions, one embodiment partitions or parses the training set into five folds (see, for example Figure 8BExample 830), the reader 930 is trained on four (i.e., n - 1) of these folds, and predictions are obtained on the remaining fold 840. This process is repeated five times to produce reader predictions (reader answers 940) for all (question, answer) pairs in the training set. These reader predictions (reader answers 940) and the original GT annotations are used to generate training examples for the corrector 960. To create examples that do not require correction, new examples 921 are created from each original example (passage 920), where the GT answer itself is delimited in the input, indicating that no correction is required. For examples that require correction, the top k incorrect predictions of the reader 930 (where k is a hyperparameter) are used to create examples for each of them, where the input 925 is the span predicted by the reader 930 and the target is the GT. The presence of both GT (correct) and incorrect predictions in the input data ensures that the corrector 960 learns to detect errors in the reader 930 predictions and correct both of them.
[0072] Figure 10A Table 1000 shows the results 1005 for answerable questions for the Natural Questions (NQ) MRC benchmark including the BERT method for robust optimization (RoBERTa), the results 1006 for an ensemble method for two readers, and the results 1007 for a method using a reader plus corrector pipeline (or Figure 9 the modified MRC answer span corrector model 900 shown in ). In one example embodiment, the modified MRC answer span corrector model 900 is evaluated on answerable questions in the development (dev) and test sets. To calculate the exact match for answerable test set questions, a system that always outputs an answer and obtains the recall value from the leaderboard is used. MLQA (Multi-Lingual Question Answering) includes instances in seven (7) languages: English (en), Arabic (ar), German (de), Spanish (es), Hindi (hi), Vietnamese (vi), and Simplified Chinese (zh).
[0073] The NQ and MLQA readers fine-tune the RoBERTa large and mBERT (packed, 104 languages) language models respectively. The RoBERTa model is first fine-tuned on SQuAD 2.0 and then on NQ. Results show that training on both answerable and unanswerable questions yields a stronger and more robust reader, even when it is evaluated on answerable questions only. The modified MRC answer span corrector model 900 uses the same underlying Transformer language model as the corresponding RoBERTa reader. When creating the training data for the modified MRC answer span corrector model 900, to generate examples that need correction, the two (k = 2) highest-scoring incorrect reader predictions are used (the value of k is tuned on the dev set). Since the goal is to fully correct any inaccuracies in the RoBERTA reader's predictions, exact match (EM) is used as the evaluation metric. In one embodiment, the modified MRC answer span corrector model 900 uses a common architecture for the reader 930 and the corrector 960, but their parameters are separate and learned independently. To compare with a baseline of equal size, the ensemble system for NQ averages the output logits (un-normalized predictions) of two different RoBERTa readers. The results in Table 1000 are obtained by averaging over three seeds. In the dev test, the result 1007 is 0.7 better than the overall performance of the reader result 1006. These results confirm that the correction objective complements the reader's extraction objective well and is fundamental to the overall performance gain of the modified MRC answer span corrector model 900. The results on the answerable questions of NQ show that the result 1007 of the modified MRC answer span corrector model 900 improves by 1.6 points on the dev set and by 1.3 points on the blind test set compared to the result 1005 of the RoBERTa reader method.
[0074] Figure 10B illustrates a reader-plus-corrector pipeline according to one embodiment when the passage is in English Figure 9Table 1020 showing the modified MRC answer span corrector model 900) and the results on the MLQA MRC benchmark dataset for the General Cross-Lingual Transfer (G-XLT) task. Performance is compared in two settings: one where the passage is in English and the question is in any of seven languages (En-Context results 1025), and the other is the G-XLT results 1030, where performance is averaged over all forty-nine (49) (question, passage) language pairs involving seven languages (English, Arabic, German, Spanish, Hindi, Vietnamese, and Simplified Chinese). For MLQA, the Fisher randomization test is used on the number of exact matches in the 158k example test set to verify the statistical significance of the results. As can be seen in Table 1020, at p < 0.01, the reader plus corrector pipeline (modified MRC answer span corrector model 900) performs significantly better than the baseline reader.
[0075] Figure 10C Table 1040 showing the differences in exact match scores for all 49 MLQA language pair combinations including the use of a reader plus corrector pipeline (modified MRC answer span corrector model 900) with a corrector 960 ( Figure 9 ) The results in Table 1040 show the variation in exact matches for all language pair combinations in the MLQA test set. The last row of Table 1040 shows the gain for each passage language averaged over questions in different languages. On average, the corrector 960 gives a performance gain for passages in all languages (last row). The highest gain is observed in the English context, which is expected as the corrector 960 model is trained to correct English answers in context. However, it was also found that the method of the reader plus corrector pipeline (modified MRC answer span corrector model 900) generalizes well to other languages in the zero-shot setting, as the exact match improved in 40 out of 49 language pairs.
[0076] In one embodiment, the changes made by the corrector 960 of the reader plus corrector pipeline (modified MRC answer span corrector model 900) to the predictions of the reader on the NQ dev set represent a total of 13% of the reader model predictions. Of all the changes, 24% result in the correction of incorrect or partially correct answers to the GT answer, and 10% replace the original correct answer with a new correct answer (due to multiple GT annotations in NQ). In 57% of the cases, these changes do not correct the error. However, looking closer, it is observed that the F1 score (a measure of test accuracy) increases in more of these cases (30%) compared to when it decreases (15%). Finally, 9% of the changes introduce an error in the correct reader prediction.
[0077] In one embodiment, the percentage of errors corrected in each of the three error categories: partial coverage, verbosity, and overlap are 9%, 38%, and 22% corrected, respectively. The correction is performed in all categories, but more so in verbosity and overlap than in partial coverage, indicating that corrector 960( Figure 9 ) is more likely to learn the concepts of minimality and syntactic structure better than sufficiency. In one embodiment, the processing using a reader plus corrector pipeline (modified MRC answer span corrector model 900( Figure 9 )) corrects the predictions of the prior art English reader 930 in different error categories. In experiments using one embodiment, the method also generalizes well to multilingual and cross-lingual MRC in seven languages.
[0078] Figure 11 FIG. illustrates a block diagram of a process 1100 for answer span correction for performance improvement of MRC according to one embodiment. In one embodiment, at block 1110, process 1100 receives, by a computing device (from Figure 1 compute node 10, Figure 2 hardware and software layer 60, Figure 3 processing system 300, Figure 4 system 400, Figure 5 system 500, reader plus corrector pipeline( Figure 9 modified MRC answer span corrector model 900), etc.), multiple questions in a natural language question answering system (e.g., Figure 9 MRC model including reader 930 therein). At block 1120, process 1100 also generates, by the computing device, multiple answers to the multiple questions. At block 1130, process 1100 also constructs, using the generated multiple answers, a new training set, where each answer is compared to the corresponding question among the multiple questions. At block 1140, process 1100 additionally augments, by the computing device, the new training set with one or more tags that delimit the span of one or more of the generated multiple answers. At block 1150, process 1100 additionally trains, with the augmented new training set, a new natural language question answering system (e.g., Figure 9 reader plus corrector pipeline in or modified MRC answer span corrector model 900 with corrector 960).
[0079] In one embodiment, process 1100 may also include using the augmented new training set for the new natural language question answering system to correct the answer span features of the reader model of the natural language question answering system (e.g., Figure 9 reader 930 of).
[0080] In one embodiment, process 1100 may additionally include the feature that a new natural language question answering system that corrects answer spans is cascaded after the natural language question answering system.
[0081] In one embodiment, process 1100 may also additionally include the feature that a new natural language question answering system that corrects answer spans determines whether the answer span should be corrected. Additionally, the corrector model (e.g., Figure 9 corrector 960) of the new natural language question answering system corrects the answer span, generating an improved answer span.
[0082] In one embodiment, process 1100 may further additionally include the feature that the corrector model uses a neural network to generate an improved answer span.
[0083] In one embodiment, process 1100 may also include the features of determining that the answer span should not be corrected, indicating no correction needed based on defining the GT answer as the input to the new natural language question answering system, and creating new example answers from each original answer of multiple answers.
[0084] In one embodiment, process 1100 may further include the feature of using multiple of the top-k incorrect answer predictions to create example answers for each incorrect answer prediction, where the input is the predicted answer span of the reader model and the target answer is the GT answer.
[0085] In one embodiment, process 1100 may include the features that the multiple generated answers include predicted answers, and one or more tags tag the multiple predicted answers in context for the corrector model to predict new answers.
[0086] In some embodiments, the features described above contribute to the advantage of significantly improving the predictions of existing natural language readers in different error categories via correction. For MRC systems, some features contribute to the advantage of answer correction because existing MRC systems tend to produce partially correct answers when faced with answerable questions. Some features also contribute to the advantage of detecting errors in the predictions of the MRC model and correcting the detected errors. Additionally, these features contribute to the advantage of generating answer spans that better match the GT and thus improving the quality of the MRC output answers.
[0087] One or more embodiments may be a system, method, and / or computer program product at any possible level of integration of technical details. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present embodiment.
[0088] A computer-readable storage medium can be a tangible device that is capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card having instructions recorded thereon or a raised structure in a groove, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0089] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0090] The computer-readable program instructions for performing the operations of the embodiments may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, executed as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit to perform aspects of this embodiment.
[0091] Aspects of the embodiments are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products. It will be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0092] These computer-readable program instructions can be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can direct a computer, a programmable data processing apparatus, and / or other devices to work in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture including instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0093] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0094] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of the possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified (one or more) logical functions. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be implemented as one step, executed simultaneously, substantially simultaneously, in partial or total time overlap, or these blocks may sometimes be executed in the reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a dedicated hardware-based system that performs the specified functions or acts or a combination of dedicated hardware and computer instructions.
[0095] Unless expressly stated otherwise, a reference in a claim to a single element is not intended to mean "one and only one" but rather "one or more." All structural and functional equivalents of the elements of the above-described exemplary embodiments that are currently known or later become known to those of ordinary skill in the art are intended to be encompassed by the claims herein. A claim element herein shall not be construed as falling under 35 U.S.C. § 112, ¶ 6, unless the element is expressly recited using the phrase "means for" or "step for."
[0096] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the embodiments. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0097] All corresponding structures, materials, acts, and equivalents of the means or step plus function elements in the following claims are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of the embodiments has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the embodiments in the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the embodiments. The embodiments were chosen and described in order to best explain the principles of the embodiments and the practical application, and to enable others of ordinary skill in the art to understand the embodiments with various modifications as are suited to the particular use contemplated.
Claims
1. A method for improving answers generated by a natural language question-answering system using a computing device, the method comprising: Receiving, by the computing device, a plurality of questions in the natural language question-answering system; Generating, by the computing device, a plurality of answers to the plurality of questions using two classification headers of a pre-trained Transformer-based language model by a reader model, wherein the reader model is trained by the following steps: Parsing a training set of the plurality of questions into n folds; Training the reader model on n - 1 of the n folds; Generating predictions for the questions on the remaining fold; And Repeating the parsing, training, and generating n times to produce predictions for all question-answer pairs in the training set; Constructing, by the computing device, a new training set using the generated plurality of answers, with each answer compared to the corresponding question in the plurality of questions; Augmenting, by the computing device, the new training set using one or more markers that define a span of one or more of the generated plurality of answers, wherein the span marks the boundaries of the plurality of answers within the training data; And Training, by the computing device, a new natural language question-answering system using the augmented new training set, wherein the new natural language question-answering system corrects the answer span of the reader model using the augmented new training set, wherein a corrector model of the new natural language question-answering system corrects the answer span to produce an improved answer span, and wherein the corrector model uses a neural network to generate the improved answer span.
2. The method according to claim 1, further comprising: Cascading the new natural language question-answering system that corrects the answer span after the natural language question-answering system.
3. The method according to claim 1, further comprising: For a determination that the answer span should not be corrected, based on defining a ground truth (GT) answer as an input to the new natural language question-answering system and indicating no correction needed, creating new example answers from each original answer in the plurality of answers.
4. The method according to claim 3, further comprising: Using a plurality of the top k incorrect answer predictions to create example answers for each incorrect answer prediction, wherein the input is the predicted answer span of the reader model and the target answer is the GT answer.
5. The method according to claim 1, wherein The generated plurality of answers includes predicted answers, and the one or more markers mark the predicted plurality of answers in context for the corrector model to predict new answers.
6. A computer program product for improving answers generated by a natural language question-answering system, the computer program product comprising program instructions that can be executed by a processor to cause the processor to: Receive, by the processor, a plurality of questions in the natural language question-answering system; Generate, by the processor, a plurality of answers to the plurality of questions using two classification headers of a pre-trained Transformer-based language model by a reader model, wherein the reader model is trained by the following steps: Parsing a training set of the plurality of questions into n folds; Train the reader model on n - 1 of the n folds; Generate predictions for the questions on the remaining folds; And Repeat the parsing, training, and generation n times to produce predictions for all question - answer pairs in the training set; The processor constructs a new training set using the generated multiple answers, with each answer compared to the corresponding question among the multiple questions; The processor augments the new training set using one or more markers that delimit the span of one or more of the generated multiple answers, where the span marks the boundaries of the multiple answers within the training data; And The processor trains a new natural language question - answering system using the augmented new training set, where the new natural language question - answering system corrects the answer span of the reader model using the augmented new training set, where a corrector model of the new natural language question - answering system corrects the answer span to produce an improved answer span, and where the corrector model uses a neural network to generate the improved answer span.
7. The computer program product according to claim 6, wherein, The program instructions executable by the processor also cause the processor to: The processor cascades the new natural language question - answering system that corrects the answer span after the natural language question - answering system.
8. The computer program product according to claim 6, wherein, The program instructions executable by the processor also cause the processor to: For the determination that the answer span should not be corrected, based on defining the ground truth (GT) answer as the input to the new natural language question - answering system, indicating no correction is needed, create new example answers from each original answer among the multiple answers.
9. The computer program product according to claim 8, wherein, The program instructions executable by the processor also cause the processor to: The processor uses multiple of the top k incorrect answer predictions to create example answers for each incorrect answer prediction, where the input is the predicted answer span of the reader model and the target answer is the GT answer.
10. The computer program product according to claim 6, wherein, The generated multiple answers include predicted answers, and the one or more markers mark the predicted multiple answers in context for the corrector model to predict new answers.
11. An apparatus for improving answers generated by a natural language question - answering system, comprising: A memory configured to store instructions; And A processor configured to execute the instructions to: Receive multiple questions in a natural language question - answering system; Generate multiple answers to the multiple questions using a reader model with two classification headers of a pre - trained transformer - based language model, where the reader model is trained by: Parse a training set of multiple questions into n folds; Train the reader model on n - 1 of the n folds; Generate predictions for the questions on the remaining folds; And Repeat the parsing, training, and generation n times to produce predictions for all question - answer pairs in the training set; Construct a new training set using the generated multiple answers, with each answer compared to the corresponding question among the multiple questions; Augmenting the new training set with one or more markers that delimit a span of one or more of the generated answers, where the span marks the boundaries of the multiple answers within the training data; and Training a new natural language question answering system with the augmented new training set, where the new natural language question answering system uses the augmented new training set to correct the answer span of the reader model, where a corrector model of the new natural language question answering system corrects the answer span to produce an improved answer span, and where the corrector model uses a neural network to generate the improved answer span.
12. The apparatus according to claim 11, wherein, The processor is further configured to execute the instructions to: Cascade a new natural language question answering system that corrects the answer span after the natural language question answering system.
13. The device according to claim 11, wherein, The processor is further configured to execute the instructions to: For a determination that the answer span should not be corrected, based on defining a ground truth (GT) answer as the input to the new natural language question answering system, indicating no correction is needed, creating a new example answer from each original answer among the multiple answers; and Using multiple of the top-k incorrect answer predictions to create example answers for each incorrect answer prediction, where the input is the predicted answer span of the reader model and the target answer is the GT answer.
14. The device according to claim 11, wherein, The multiple generated answers include predicted answers, and the one or more markers mark the predicted multiple answers in context for the corrector model to predict a new answer.
Citation Information
Patent Citations
Video question answering method and system based on cross-modal prompt learning
CN114996513A
Constructing imaginary discourse trees to improve answering convergent questions
US20190347297A1