Small-Sample Machine Reading Comprehension Method and Apparatus, and Computer-Readable Storage Medium

Through meta-learning methods and pre-training Bert models, the problem that machine reading comprehension model needs to be retrained in different fields is solved, and the reading comprehension task of quickly adapting to new fields under small sample data is realized.

CN114065728BActive Publication Date: 2025-07-25CHINA TELECOM CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010750691.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-30
Publication Date
2025-07-25
Estimated Expiration
2040-07-30

AI Technical Summary

Technical Problem

Existing machine reading comprehension models need to be retrained for different fields, resulting in inefficiency in small sample data and the inability to quickly adapt to new field tasks.

Method used

Using the meta-learning method, the model parameters are trained on the reading comprehension tasks in related fields, and the parameters of the meta-learning model are iteratively updated. The pre-trained Bert model is used to extract features and quickly generate a new field reading comprehension model.

Benefits of technology

In the new field, only a small amount of sample data is required to quickly converge, improving the model's adaptability and efficiency in small samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114065728B_ABST
    Figure CN114065728B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a small-sample machine reading comprehension method and apparatus, and a computer-readable storage medium. The small-sample machine reading comprehension method includes: for reading comprehension tasks in a related field, training model parameters through sampling extraction tasks respectively, and using the model parameters to iteratively update the model parameters of a meta-learning model; training starting from the meta-learning model for reading comprehension tasks in a new field to generate a reading comprehension model for the new field. The present disclosure can use the method of meta-learning to learn machine reading comprehension tasks in existing related fields, so as to obtain a meta-learning model to accelerate machine reading comprehension tasks with a small sample size in a new field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of natural language processing, and particularly to a few-shot machine reading comprehension method and apparatus, and a computer-readable storage medium. Background Art

[0002] MRC (Machine Reading Comprehension) is essential in scenarios such as intelligent search, online consultation, recommendation, question answering, and dialogue. Machine reading comprehension models usually attempt to understand the knowledge in articles of a specific and segmented field, and the number of articles in these fields may be small, and models need to be retrained from scratch for different fields, even though some articles in different fields are very similar. Summary of the Invention

[0003] In view of at least one of the above technical problems, the present disclosure provides a few-shot machine reading comprehension method and apparatus, and a computer-readable storage medium, which can use the method of meta-learning to learn the machine reading comprehension tasks in existing related fields, so as to obtain a meta-learning model to accelerate the few-shot machine reading comprehension tasks in new fields.

[0004] According to one aspect of the present disclosure, there is provided a few-shot machine reading comprehension method, including:

[0005] For the reading comprehension tasks in related fields, sample and extract tasks to train the model parameters respectively, and use the model parameters to iteratively update the model parameters of the meta-learning model;

[0006] Train the reading comprehension tasks in the new field starting from the meta-learning model to generate a reading comprehension model for the new field.

[0007] In some embodiments of the present disclosure, the training of the reading comprehension tasks in the new field starting from the meta-learning model to generate a reading comprehension model for the new field includes:

[0008] Use the model parameters of the meta-learning model as the initial parameters for the reading comprehension tasks in the new field, and use a small number of samples to train the reading comprehension model for the new field;

[0009] The reading comprehension model for the new field converges quickly, generating a reading comprehension model for the reading comprehension tasks in the new field.

[0010] In some embodiments of the present disclosure, the few-shot machine reading comprehension method further includes:

[0011] Use a pre-trained new-domain reading comprehension model to extract question and document features, and predict the start and end positions of the answer. Among them, the new-domain reading comprehension model is a bidirectional encoder representation model based on Transformer.

[0012] In some embodiments of the present disclosure, for the reading comprehension task in the relevant domain, training the model parameters through sampling extraction tasks respectively and using the model parameters to iteratively update the model parameters of the meta-learning model includes:

[0013] Randomly initialize the model parameters;

[0014] Sample and extract multiple reading comprehension tasks in the relevant domain from the relevant domain task distribution set;

[0015] For each task sampled and extracted, train the meta-learning model to obtain the optimized model parameters for each task;

[0016] Perform meta-update, and update the meta-model parameters according to the optimized model parameters of each task.

[0017] In some embodiments of the present disclosure, obtaining the optimized model parameters for each task includes:

[0018] Calculate the loss function;

[0019] Use gradient descent to find the optimized model parameters that minimize the loss function.

[0020] In some embodiments of the present disclosure, after obtaining the optimized model parameters for each task, the few-shot machine reading comprehension method further includes:

[0021] Judge whether each task sampled and extracted has been trained;

[0022] In the case where each task sampled and extracted has been trained, execute the step of meta-update;

[0023] In the case where each task sampled and extracted has not been trained, execute the step of training the meta-learning model for each task sampled and extracted to obtain the optimized model parameters for each task.

[0024] In some embodiments of the present disclosure, updating the model parameters according to the optimized model parameters of each task includes:

[0025] Update the meta-model parameters to the average value of the loss gradients of each task.

[0026] In some embodiments of the present disclosure, after updating the model parameters according to the optimized model parameters of each task, the few-shot machine reading comprehension method further includes:

[0027] Determine whether the predetermined number of generations of training for the meta-model is completed;

[0028] In the case where the predetermined number of generations of training for the meta-model is completed, perform the step of training the reading comprehension task for the new domain starting from the meta-learning model;

[0029] In the case where the predetermined number of generations of training for the meta-model is not completed, perform the step of sampling and extracting multiple reading comprehension tasks for related domains from the set of related domain task distributions.

[0030] According to another aspect of the present disclosure, there is provided a few-shot machine reading comprehension device, including:

[0031] A related task training module, configured to train the model parameters respectively by sampling and extracting tasks for the reading comprehension tasks of related domains, and use the model parameters to iteratively update the model parameters of the meta-learning model;

[0032] A new task training module, configured to train the reading comprehension task for the new domain starting from the meta-learning model to generate a reading comprehension model for the new domain.

[0033] In some embodiments of the present disclosure, the few-shot machine reading comprehension device is configured to perform operations for implementing the few-shot machine reading comprehension method as described in any one of the above embodiments.

[0034] According to another aspect of the present disclosure, there is provided a few-shot machine reading comprehension device, including a memory and a processor, wherein:

[0035] The memory is configured to store instructions;

[0036] The processor is configured to execute the instructions, so that the few-shot machine reading comprehension device performs operations for implementing the few-shot machine reading comprehension method as described in any one of the above embodiments.

[0037] According to another aspect of the present disclosure, there is provided a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the few-shot machine reading comprehension method as described in any one of the above embodiments is implemented.

[0038] The present disclosure can use the method of meta-learning to learn the machine reading comprehension tasks of existing related domains, so as to obtain a meta-learning model to accelerate the few-shot machine reading comprehension tasks of new domains. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 It is a schematic diagram of some embodiments of the small-sample machine reading comprehension method of the present disclosure.

[0041] Figure 2 It is a schematic diagram of some other embodiments of the small-sample machine reading comprehension method of the present disclosure.

[0042] Figure 3 It is a schematic diagram of some yet other embodiments of the small-sample machine reading comprehension method of the present disclosure.

[0043] Figure 4 It is a schematic diagram of some embodiments of the small-sample machine reading comprehension device of the present disclosure.

[0044] Figure 5 It is a schematic diagram of some other embodiments of the small-sample machine reading comprehension device of the present disclosure. Detailed implementation manners

[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present disclosure in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some embodiments of the present disclosure, rather than all embodiments. The description of at least one exemplary embodiment below is actually only illustrative and in no way limits the present disclosure and its application or use. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.

[0046] Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0047] At the same time, it should be understood that for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0048] For technologies, methods, and devices known to those of ordinary skill in the relevant fields, detailed discussions may not be made, but where appropriate, the said technologies, methods, and devices should be regarded as part of the authorization specification.

[0049] In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0050] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0051] Figure 1 Schematic diagrams of some embodiments of the small-sample machine reading comprehension method of the present disclosure are shown. Preferably, this embodiment can be executed by the small-sample machine reading comprehension device of the present disclosure. The method may include the following steps:

[0052] Step 11: For reading comprehension tasks in related fields, train model parameters through sampling extraction tasks respectively, and use the model parameters to iteratively update the model parameters of the meta-learning model.

[0053] In some embodiments of the present disclosure, step 11 may include steps 111-114, where:

[0054] Step 111: Randomly initialize model parameters.

[0055] Step 112: Sample and extract multiple reading comprehension tasks in related fields from the set of related field task distributions.

[0056] Step 113: For each task sampled and extracted, train the meta-learning model to obtain the optimized model parameters for each task.

[0057] In some embodiments of the present disclosure, in step 113, the step of obtaining the optimized model parameters for each task may include: calculating a loss function; using gradient descent to find the optimized model parameters that minimize the loss function.

[0058] In some embodiments of the present disclosure, after the step of obtaining the optimized model parameters for each task in step 113, the small-sample machine reading comprehension method may further include: determining whether each task sampled and extracted has been trained; in the case where each task sampled and extracted has been trained, execute step 114; in the case where not all tasks sampled and extracted have been trained, repeat step 113, that is, repeat the step of separately training to obtain model parameters.

[0059] Step 114: Perform meta-update, and update the meta-model parameters according to the optimized model parameters of each task.

[0060] In some embodiments of the present disclosure, after the step of updating the model parameters according to the optimized model parameters of each task in step 114, the few-shot machine reading comprehension method may further include: determining whether the training of the meta-model for a predetermined number of generations is completed; in the case where the training of the meta-model for a predetermined number of generations is completed, performing step 12; in the case where the training of the meta-model for a predetermined number of generations is not completed, performing step 112.

[0061] In some embodiments of the present disclosure, step 114 may include: updating the meta-model parameters to the average value of the loss gradients of each task.

[0062] Step 12: Training the reading comprehension task in the new domain starting from the meta-learning model to generate a reading comprehension model for the new domain.

[0063] In some embodiments of the present disclosure, step 12 may include: using the model parameters of the meta-learning model as the initialization parameters for the reading comprehension task in the new domain, and training the reading comprehension model for the new domain with a small number of samples; the reading comprehension model for the new domain converges quickly, generating a reading comprehension model for the reading comprehension task in the new domain.

[0064] In some embodiments of the present disclosure, the few-shot machine reading comprehension method may further include: using the pre-trained reading comprehension model for the new domain to extract question and document features, and predicting the start and end positions of the answer, where the reading comprehension model for the new domain is a Bert (Bidirectional Encoder Representation from Transformers) model.

[0065] The present disclosure can use the method of meta-learning to learn the machine reading comprehension tasks in the existing related domains, so as to obtain a meta-learning model to accelerate the few-shot machine reading comprehension tasks in the new domain.

[0066] In some embodiments of the present disclosure, meta-learning (Meta Learning) of the present disclosure, also known as learning to learn, forms a meta-learning model (Meta-Learning Model) after learning various tasks, so that when facing a new task, the existing meta-learning model can be used to accelerate learning. That is, the model can handle multiple tasks without having to train from scratch for each task.

[0067] Based on the few-shot machine reading comprehension method provided in the above embodiments of the present disclosure, for the reading comprehension tasks in the existing fields, the model parameters are trained respectively through sampling extraction tasks, and these parameters are used to iteratively update the parameters of the meta-learning model. The reading comprehension tasks in the new fields of the above embodiments of the present disclosure are trained starting from this meta-model. In this way, with only a small amount of sample data, the new reading comprehension model can quickly achieve good results. In the machine reading comprehension model part of the above embodiments of the present disclosure, the pre-trained Bert model is used to extract the features of the questions and documents, thereby improving the efficiency of the model in extracting features.

[0068] Figure 2 It is a schematic diagram of some other embodiments of the few-shot machine reading comprehension method of the present disclosure. Preferably, this embodiment can be executed by the few-shot machine reading comprehension device of the present disclosure. The present disclosure uses a meta-learning method based on MAML (Model Agnostic MetaLearning) to solve the reading comprehension problem in natural language processing. The method of the present disclosure may include the following steps:

[0069] Step 20, assume that the reading comprehension model of the present disclosure is f, its parameter is θ, and the set of relevant domain task distributions is P(T).

[0070] Step 21, randomly initialize the parameter θ.

[0071] Step 22, sample and extract multiple relevant domain reading comprehension tasks from the set of relevant domain task distributions P(T).

[0072] In some embodiments of the present disclosure, step 22 may include: selecting 3 relevant domain reading comprehension tasks T = {T1, T2, T3} ∼ p(T) from the set of relevant domain task distributions P(T).

[0073] Step 23, for each task T i , train the meta-learning model to obtain the optimized model parameters for each task.

[0074] In some embodiments of the present disclosure, step 23 may include: for each task T i train the model and calculate the loss function L, and find the parameter θ′ that minimizes the loss function L through gradient descent according to formula (1) i , where α represents the learning rate in formula (1).

[0075]

[0076] By performing step 23 for each task T i optimized parameter θ′ can be obtained, as shown in formula (2).

[0077] θ′ = {θ′1, θ′2, θ′3} (2)

[0078] In some embodiments of the present disclosure, after step 23, the few-shot machine reading comprehension method may further include: determining whether each task sampled and extracted has been trained; in the case where each task sampled and extracted has been trained, performing step 24; otherwise, in the case where each task sampled and extracted has not been trained, repeating step 23 until each task in T has been trained, and then performing step 24.

[0079] Step 24, perform a meta-update, and update the meta-model parameters according to the optimized model parameters of each task.

[0080] In some embodiments of the present disclosure, step 24 may include: performing a meta-update according to formula (3), and updating the randomly initialized parameters to the average of the loss gradients of each task. In formula (3), β represents the average learning rate.

[0081]

[0082] In some embodiments of the present disclosure, after step 24, the few-shot machine reading comprehension method may further include: determining whether the training of the predetermined number of generations N of the meta-model is completed (i.e., whether the training of N Epochs is completed); in the case where the training of the predetermined number of generations N of the meta-model is completed, performing step 25; otherwise, in the case where the training of the predetermined number of generations N of the meta-model is not completed, repeating step 22.

[0083] As Figure 2 shown, the algorithm of the above embodiments of the present disclosure includes two-layer loops. The inner loop (the loop of step 23) calculates the optimized parameters of each task, and the outer loop (the loop of steps 22 - 24) performs a meta-update.

[0084] Step 25, transmit the model f generated in the above steps as the result of meta-learning to a new related task.

[0085] In some embodiments of the present disclosure, as Figure 2 shown, f in the above embodiments of the present disclosure is a reading comprehension model based on Bert, which takes a question + passage as input and predicts the start and end positions of the answer.

[0086] Step 26, input a small number of training samples of the new task T0.

[0087] Step 27, use the model parameters θ trained on the related task (Prior task) fFor the initialization parameters, make fine-tuning.

[0088] In some embodiments of the present disclosure, step 27 may include: For the training of the new task T0, it only needs to be initialized with a small number of samples with θ as the initialization parameter on the basis.

[0089] In step 28, the model converges quickly, generating a new domain reading comprehension model for the reading comprehension task T0 in the new domain.

[0090] The method proposed in the above embodiments of the present disclosure can be directly used in the machine question answering module of the operator's R & D cloud platform project. Specifically, the above embodiments of the present disclosure adopt a meta-learning method to address the problem that different domain natural language reading comprehension tasks require a large amount of sample data to train the model from scratch. The above embodiments of the present disclosure enable the model under the new task to quickly iterate with a small amount of sample data by making full use of the sample data of the existing tasks, providing a good solution for the small sample reading comprehension task.

[0091] Figure 3 It is a schematic diagram of some embodiments of the small sample machine reading comprehension method of the present disclosure. As Figure 3 shown, the small sample machine reading comprehension method of the present disclosure includes: the machine reading comprehension processes MRC1, MRC2,... MRC n for the reading comprehension tasks in the relevant fields, that is, for the reading comprehension tasks in the relevant fields, the model parameters are respectively trained by sampling and extracting tasks (small samples S1, S2,... S n ), and the model parameters are used to iteratively update the model parameters of the meta-learning model M0; and for the reading comprehension task in the new domain, training is performed starting from the meta-learning model using the small sample S to generate a new domain reading comprehension model (i.e., the basic model M).

[0092] The above embodiments of the present disclosure propose a method and device for small sample reading comprehension based on meta-learning. For the reading comprehension tasks in the existing fields, the above embodiments of the present disclosure can respectively train the model parameters by sampling and extracting tasks, and use these parameters to iteratively update the parameters of the meta-learning model. The reading comprehension task in the new domain of the above embodiments of the present disclosure is trained starting from this meta-model, so that only a small amount of sample data is required to quickly make the new reading comprehension model achieve a good effect. The above embodiments of the present disclosure adopt the pre-trained Bert model in the machine reading comprehension model part to extract the features of the question and the document, improving the efficiency of the model in extracting features.

[0093] Figure 4 It is a schematic diagram of some embodiments of the small sample machine reading comprehension device of the present disclosure. As Figure 4As shown in the figure, the small-sample machine reading comprehension device of the present disclosure may include a related task training module 41 and a new task training module 42, where:

[0094] The related task training module 41 is used to train the model parameters for the reading comprehension tasks in the related field by sampling extraction tasks respectively, and use the model parameters to iteratively update the model parameters of the meta-learning model.

[0095] In some embodiments of the present disclosure, the related task training module 41 may be used to randomly initialize the model parameters; sample and extract multiple reading comprehension tasks in the related field from the set of related field task distributions; for each sampled and extracted task, train the meta-learning model to obtain the optimized model parameters for each task; perform meta-update, and update the meta-model parameters according to the optimized model parameters of each task.

[0096] In some embodiments of the present disclosure, when the related task training module 41 obtains the optimized model parameters for each task, it may be used to calculate the loss function; use gradient descent to find the optimized model parameters that minimize the loss function.

[0097] In some embodiments of the present disclosure, after the related task training module 41 obtains the optimized model parameters for each task, it may be used to determine whether each sampled and extracted task has been trained; in the case where each sampled and extracted task has been trained, perform the operation of meta-update; in the case where each sampled and extracted task has not been trained, repeat the operation of training the meta-learning model for each sampled and extracted task to obtain the optimized model parameters for each task.

[0098] In some embodiments of the present disclosure, when the related task training module 41 updates the model parameters according to the optimized model parameters of each task, it may be used to update the meta-model parameters to the average value of the loss gradients of each task.

[0099] In some embodiments of the present disclosure, after the related task training module 41 updates the model parameters according to the optimized model parameters of each task, it may be used to determine whether the predetermined number of generations of training for the meta-model is completed; in the case where the predetermined number of generations of training for the meta-model is completed, perform the operation of training the reading comprehension tasks in the new field starting from the meta-learning model; in the case where the predetermined number of generations of training for the meta-model is not completed, perform the operation of sampling and extracting multiple reading comprehension tasks in the related field from the set of related field task distributions.

[0100] The new task training module 42 is used to train the reading comprehension tasks in the new field starting from the meta-learning model to generate a reading comprehension model for the new field.

[0101] In some embodiments of the present disclosure, the new task training module 42 can be used to train a reading comprehension model for a new domain with a small number of samples, using the model parameters of the meta-learning model as the initial parameters; the reading comprehension model for the new domain converges quickly, generating a reading comprehension model for the reading comprehension task of the new domain.

[0102] In some embodiments of the present disclosure, the new task training module 42 can be used to extract question and document features using a pre-trained reading comprehension model for the new domain, and predict the start and end positions of the answer, where the reading comprehension model for the new domain is a bidirectional encoder representation model based on Transformer.

[0103] In some embodiments of the present disclosure, the few-shot machine reading comprehension device can be used to perform operations to implement the few-shot machine reading comprehension method as described in any of the above embodiments (e.g., Figures 1-3 any of the embodiments).

[0104] Based on the few-shot machine reading comprehension device provided in the above embodiments of the present disclosure, for the reading comprehension tasks in the existing domain, the model parameters can be trained separately by sampling and extracting tasks, and these parameters can be used to iteratively update the parameters of the meta-learning model. The reading comprehension tasks in the new domain of the above embodiments of the present disclosure start training with this meta-model, so that with only a small amount of sample data, the new reading comprehension model can quickly achieve good results. The above embodiments of the present disclosure use a pre-trained Bert model in the machine reading comprehension model part to extract the features of questions and documents, thereby improving the efficiency of the model in extracting features.

[0105] Figure 5 It is a schematic diagram of some other embodiments of the few-shot machine reading comprehension device of the present disclosure. As Figure 5 shown, the few-shot machine reading comprehension device of the present disclosure can include a memory 51 and a processor 52, where:

[0106] The memory 51 is used to store instructions.

[0107] The processor 52 is used to execute the instructions, so that the few-shot machine reading comprehension device performs operations to implement the few-shot machine reading comprehension method as described in any of the above embodiments (e.g., Figures 1-3 any of the embodiments).

[0108] The above embodiments of the present disclosure can use the method of meta-learning to learn the machine reading comprehension tasks in the existing related domains, so as to obtain a meta-learning model to accelerate the few-shot machine reading comprehension tasks in the new domain.

[0109] The method proposed in the above embodiments of the present disclosure can be directly used in the machine question-answering module of the operator's R & D cloud platform project. The above embodiments of the present disclosure adopt a meta-learning method to address the problem that a large amount of sample data is required to train the model from scratch for natural language reading comprehension tasks in different fields. By making full use of the sample data of existing tasks, the above embodiments of the present disclosure enable the model under the new task to quickly iterate with a small amount of sample data, providing a good solution for the reading comprehension task with few samples.

[0110] According to another aspect of the present disclosure, there is provided a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the small-sample machine reading comprehension method as described in any of the above embodiments (for example Figures 1-3 any of the embodiments) is implemented.

[0111] The above embodiments of the present disclosure adopt a meta-learning method based on MAML to solve the reading comprehension problem in natural language processing.

[0112] For the reading comprehension tasks in the existing fields, the above embodiments of the present disclosure can separately train the model parameters by sampling extraction tasks, and use these parameters to iteratively update the parameters of the meta-learning model. The reading comprehension tasks in the new fields of the above embodiments of the present disclosure start training with this meta-model, so that with only a small amount of sample data, the new reading comprehension model can quickly achieve good results. The above embodiments of the present disclosure adopt a pre-trained Bert model in the machine reading comprehension model part to extract the features of questions and documents, thereby improving the efficiency of the model in extracting features.

[0113] The small-sample machine reading comprehension device described above can be implemented as a general-purpose processor, a programmable logic controller (PLC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components or any suitable combination thereof for performing the functions described in this application.

[0114] So far, the present disclosure has been described in detail. To avoid obscuring the concept of the present disclosure, some details well known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.

[0115] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disk, etc.

[0116] The description of the present disclosure has been presented for purposes of illustration and description, and is not intended to be exhaustive or to limit the present disclosure to the forms disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments were chosen and described in order to best explain the principles of the present disclosure and its practical application, and to enable others of ordinary skill in the art to understand the present disclosure and design various embodiments with various modifications suited to particular uses.

Claims

1. A small-sample machine reading comprehension method, characterized in that, include: For reading comprehension tasks in related fields, the model parameters are trained separately through sampling tasks, and the model parameters are used to iteratively update the model parameters of the meta-learning model; The reading comprehension task in the new domain is trained with the meta-learning model as the starting point to generate a reading comprehension model in the new domain; The reading comprehension tasks in the related fields are respectively trained by sampling tasks, and the model parameters are used to iteratively update the model parameters of the meta-learning model, including: Randomly initialize model parameters; From the distribution set of tasks in related fields, sample multiple reading comprehension tasks in related fields; For each sampled task, train the meta-learning model to obtain the optimized model parameters for each task; Determine whether each sampled task has been trained; In the case that each sampled task has not been trained, the steps of training the meta-learning model for each sampled task and obtaining the optimized model parameters for each task are repeated; When each sampled task has been trained, a meta-update is performed to update the meta-model parameters according to the optimized model parameters of each task; The reading comprehension task for the new domain is trained with the meta-learning model as the starting point, and generating the new domain reading comprehension model includes: For the reading comprehension task in the new domain, the model parameters of the meta-learning model are used as initialization parameters, and a small number of samples are used to train the reading comprehension model in the new domain; The new domain reading comprehension model converges quickly and generates a new domain reading comprehension model for reading comprehension tasks in new domains.

2. The small-sample machine reading comprehension method according to claim 1, wherein Also includes: A pre-trained new domain reading comprehension model is used to extract question and document features and predict the start and end positions of the answer. The new domain reading comprehension model is a bidirectional encoder representation model based on the transformer.

3. The small-sample machine reading comprehension method according to claim 1 or 2, characterized in that Obtaining the optimized model parameters for each task includes: Calculate the loss function; Gradient descent is used to find the optimal model parameters that minimize the loss function.

4. The small-sample machine reading comprehension method according to claim 1 or 2, wherein Updating the model parameters according to the optimized model parameters of each task includes: Update the meta-model parameters as the average of the gradients of the losses across tasks.

5. The small-sample machine reading comprehension method according to claim 1 or 2, characterized in that After the model parameters are updated according to the optimized model parameters of each task, it also includes: Determining whether the predetermined number of trainings for the metamodel are completed; After completing the predetermined number of generations of training for the meta-model, executing the step of training a reading comprehension task in a new domain using the meta-learning model as a starting point; In the case where the predetermined number of generations of training for the meta-model is not completed, a step of sampling and extracting a plurality of related-domain reading comprehension tasks from a distribution set of related-domain tasks is performed.

6. A small-sample machine reading comprehension device, characterized in that include: Related task training module, which is used for reading comprehension tasks in related fields. It trains model parameters by sampling tasks and uses the model parameters to iteratively update the model parameters of the meta-learning model. A new task training module, used to train a reading comprehension task in a new domain using the meta-learning model as a starting point to generate a new domain reading comprehension model; Among them, the relevant task training module is used to randomly initialize the model parameters; sample and extract multiple relevant field reading comprehension tasks from the relevant field task distribution set; for each sampled task, train the meta-learning model to obtain the optimized model parameters for each task; determine whether each sampled task has been trained; in the case where each sampled task has not been trained, repeat the operation of training the meta-learning model for each sampled task to obtain the optimized model parameters for each task; in the case where each sampled task has been trained, perform a meta-update to update the meta-model parameters according to the optimized model parameters of each task. Among them, the new task training module is used to train the reading comprehension model for the new field with the model parameters of the meta-learning model as the initial parameters and using a small number of samples; the reading comprehension model for the new field converges quickly to generate a reading comprehension model for the new field for the reading comprehension tasks in the new field.

7. The small-sample machine reading comprehension device according to claim 6, wherein The few-shot machine reading comprehension device is used to perform the operations of implementing the few-shot machine reading comprehension method according to any one of claims 2-5.

8. A small-sample machine reading comprehension device, characterized in that, It includes a memory and a processor, where: The memory is used to store instructions. The processor is used to execute the instructions, so that the few-shot machine reading comprehension device performs the operations of implementing the few-shot machine reading comprehension method according to any one of claims 1-5.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the instructions are executed by the processor, the few-shot machine reading comprehension method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Training sample generation method and device, electronic equipment and storage medium

    CN110795552A

  • Data processing model training method and device, data processing method and device and electronic equipment

    CN111401558A