Model training method, model training device and storage medium

The model training method for question generation systems improves the relevance of generated questions to selected answers by using interrogative words and their interpretations in the training process, resulting in enhanced performance of question-answer pairs.

JP7679902B2Active Publication Date: 2025-05-20RICOH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024060102
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-04-03
Filing Date
2024-04-03
Publication Date
2025-05-20
Estimated Expiration
2044-04-03

Smart Images

  • Figure 0007679902000005
    Figure 0007679902000005
  • Figure 0007679902000006
    Figure 0007679902000006
  • Figure 0007679902000007
    Figure 0007679902000007
Patent Text Reader

Abstract

To provide a model training method and device, and a storage medium.SOLUTION: In a model training process of an example of the present invention, a first presentation text corresponding to an interrogative and a text sample are used as input for a response selection model, and the response selection model is caused to output a response candidate for a specific type. In addition, a second presentation text corresponding to the interrogative, the text sample, and a response sample are used as input for a question generation model, and the question generation model is caused to generate a question related to the selected response candidate by associating the response selection model with the question generation model using the same interrogative and a presentation text including its interpretation. Thereby, performance of a question-response pair generated by the model is improved.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to machine learning and natural language processing (NLP) technologies, and more particularly to a model training method, a model training device, and a storage medium. [Background technology]

[0002] Question generation technology is one of the important technologies in the field of natural language processing. The purpose of question generation is to generate several questions related to a single text specified by a user and teach the answers to these questions in the text. Question generation technology is widely used in Q&A systems and search engines to automatically generate combinations of questions and answers, i.e., question-answer pairs (sometimes simply called Q&A pairs in this paper). Q&A systems and search engines require a large number of question-answer pairs. One of the automatic question-answering methods in Q&A systems is to use a similarity algorithm to match a user's question with a question-answer pair created in advance in a database to obtain an answer. On the other hand, search engines use a machine reading comprehension model that finds the correct answer from the text searched for the user's question. However, the machine reading comprehension model needs to be trained with a large number of Q&A pairs. Generating these Q&A pairs requires a lot of time and effort, and in some cases, taggers are required to have a certain level of expertise. However, question generation technology can significantly reduce the construction cost of Q&A systems and search engines by automatically generating Q&A pairs.

[0003] Currently, mainstream question generation systems usually include one answer selection model and one question generation model. Among them, the answer selection model selects answer candidates from the text, and the question generation model generates related questions based on the text and the answer candidates selected by the answer selection model. One of the conventional question generation methods generates multiple key phrases (answer candidates) by training one answer selection model (sequence-to-sequence model). This answer selection model learned to assign higher probabilities to answer candidates from answers manually selected from a large-scale Q&A data collection (SQuAD). This method also provides a question generation model that generates questions based on the selected answer candidates. This method has the disadvantage that the answer selection model learns to simultaneously select all types of answer candidates in the text during training, making it difficult for the model to capture the features of question value. In addition, the question generation model may generate questions that are not related to the selected answer because additional information (such as the type of answer) cannot be obtained from the answer selection model to assist question generation after selecting an answer.

[0004] Therefore, there is a need for question generation techniques that improve the relevance of generated questions to selected answers. Summary of the Invention [Problem to be solved by the invention]

[0005] At least one embodiment of the present invention provides a model training method, apparatus, and storage medium for improving the relevance of answers selected by an answer selection model and questions generated by a question generation model.

[0006] Improve accuracy in word matching tasks. [Means for solving the problem]

[0007] In order to solve the above technical problems, the present invention provides the following techniques.

[0008] First, an embodiment of the present invention provides a model training method, comprising: generating a first presentation text presenting a selection of answer candidates for a question including the interrogative word and a second presentation text presenting a generation of a question including the interrogative word based on an interrogative word in a target language and its interpretation; obtaining a plurality of original training samples including a text sample, a question sample including the interrogative word, and an answer sample corresponding to the question sample; generating a first training sample for each original training sample using the text sample, the answer sample, and a first presentation text corresponding to the interrogative word in the question sample of the original training sample to obtain a first training set including the plurality of first training samples, and generating a second training sample using the text sample, the question sample, the answer sample, and a second presentation text corresponding to the interrogative word in the question sample of the original training sample to obtain a second training set including the plurality of second training samples; training an answer selection model using the first training set, and training a question generation model using the second training set.

[0009] Optionally, the first training set further comprises at least one third training sample, the third training sample being formed using a first submitted text corresponding to the first interrogative word, a blank answer and the first text sample, determining that for each first text sample, none of all corresponding question samples includes the first interrogative word.

[0010] Optionally, training an answer selection model using the first training set and training a question generation model using the second training set includes inputting a first submitted text and text sample for each training sample in the first training set into an answer selection model, training the answer selection model using corresponding answer samples output by the answer selection model as targets, and generating a trained answer selection model; inputting a second submitted text, text sample and answer sample for each training sample in the second training set into a question generation model, training the question generation model using corresponding question samples output by the question generation model as targets, and generating a trained question generation model.

[0011] Optionally, the above model training method also includes: for a target text, selecting a second interrogative from interrogatives in the target language; inputting a first submission text corresponding to the second interrogative and the target text into the answer selection model to obtain a target answer output by the answer selection model; inputting a second submission text corresponding to the second interrogative, the target text and the target answer into the question generation model to obtain a target question output by the question generation model.

[0012] Optionally, in the model training method, the first and second submitted texts corresponding to the interrogative word include the interrogative word and an interpretation of the interrogative word, respectively.

[0013] Also optionally, the answer selection model is a natural language understanding model and the question generation model is a natural language generation model.

[0014] An embodiment of the present invention provides a model training device including: a first generation module that generates, based on an interrogative word in a target language and its interpretation, a first presentation text that presents a selection of answer candidates for a question including the interrogative word and a second presentation text that presents a generation of a question including the interrogative word; a first acquisition module that acquires a plurality of original training samples including a text sample, a question sample including the interrogative word, and an answer sample corresponding to the question sample; a second generation module that generates, for each original training sample, a first training sample using the text sample, the answer sample, and a first presentation text corresponding to the interrogative word in the question sample of the original training sample to obtain a first training set including the plurality of first training samples; and a training module that trains an answer selection model using the first training set and trains a question generation model using the second training set.

[0015] Optionally, the first training set further comprises at least one third training sample, the third training sample being formed using a first submitted text corresponding to the first interrogative word, a blank answer and the first text sample, determining that for each first text sample, none of all corresponding question samples includes the first interrogative word.

[0016] Optionally, the training module inputs a first submitted text and text sample for each training sample in the first training set into an answer selection model, trains the answer selection model using corresponding answer samples output by the answer selection model as targets, and generates a trained answer selection model; inputs a second submitted text, text sample, and answer sample for each training sample in the second training set into a question generation model, trains the question generation model using corresponding question samples output by the question generation model as targets, and generates a trained question generation model.

[0017] Optionally, the model training device further includes a model application module for selecting a second interrogative word from interrogative words in the target language for a target text; inputting a first submission text corresponding to the second interrogative word and the target text into the answer selection model to obtain a target answer output by the answer selection model; and inputting a second submission text corresponding to the second interrogative word, the target text and the target answer into the question generation model to obtain a target question output by the question generation model.

[0018] Optionally, the first and second presentation texts corresponding to the interrogative word also include the interrogative word and an interpretation of the interrogative word, respectively.

[0019] Also optionally, the answer selection model is a natural language understanding model and the question generation model is a natural language generation model.

[0020] Furthermore, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps of the model training method described above. Effect of the Invention

[0021] Compared with the prior art, according to the model training method and apparatus provided in the embodiments of the present invention, in model training, a first submitted text and text samples corresponding to an interrogative word are used as inputs to an answer selection model, and the answer selection model outputs a certain type of answer candidate; a second submitted text, text samples and answer samples corresponding to the interrogative word are used as inputs to a question generation model, and the answer selection model is associated with the question generation model by the submitted texts containing the same interrogative word and its interpretation, thereby improving the relevance between the selected answer and the generated question, and further improving the performance of the question-answer pairs generated by the model. [Brief description of the drawings]

[0022] Various other benefits and advantages will become apparent to those skilled in the art from the detailed description of the preferred embodiment below. The accompanying drawings are used only to illustrate the preferred embodiment and are not intended to limit the present invention. The same reference numerals are used throughout the drawings to refer to the same parts. [Figure 1] FIG. 1 is a flow chart illustrating a model training method according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram showing an example of creating a presentation text according to an embodiment of the present invention. [Diagram 3] FIG. 3 is a diagram showing an example in which a presentation template is used. [Figure 4] FIG. 4 is a diagram showing an example of creating a presentation text for a question sample according to an embodiment of the present invention. [Diagram 5] FIG. 5 is a diagram illustrating an example of training an answer selection model using a first training sample according to an embodiment of the present invention. [Figure 6] FIG. 6 is a diagram illustrating an example of training a question generation module using a second training sample according to an embodiment of the present invention. [Figure 7] FIG. 7 is a diagram showing a configuration of a model training device according to an embodiment of the present invention. [Figure 8] FIG. 8 is a diagram showing another example of the structure of the model training device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0023] In order to clarify the problems, configurations and effects that the present invention is intended to solve, the following detailed description will be given in conjunction with drawings and specific embodiments. In the following description, specific details of specific arrangements and components are provided only to assist in a full understanding of the embodiments of the present invention. Therefore, it is obvious to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. For the sake of clarity and conciseness, descriptions of known functions and configurations are omitted.

[0024] References throughout the specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present application. Thus, the appearances of "in one embodiment" or "in one embodiment" in various places in the specification do not necessarily refer to the same embodiment. Alternatively, these particular features, structures, or characteristics may be combined in any suitable manner with one or more embodiments. The terms "first," "second," and the like in the specification and claims of the present application are used to distinguish between similar objects, and are not necessarily used to describe a particular sequence or order. It should be understood that such used data are interchangeable in appropriate circumstances, whereby the embodiments of the present invention described herein may, for example, be practiced in sequences other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to the explicitly recited steps or units, but may include other steps or units not explicitly recited or inherent to such process, method, product, or apparatus. In the specification and claims, "and / or" means at least one of the objects connected.

[0025] In each embodiment of the present invention, the magnitude of the numbers of the following processes does not mean the order of execution. The execution order of each process is determined by its function and inherent logic, and does not limit the implementation process of the embodiment of the present invention.

[0026] Typically, a question generation system includes one answer selection model and one question generation model. Previous related research has attempted to improve the performance of question generation in two ways. One is whether the answer selected by the answer selection model has a higher value than the question, and the other is whether the question generated by the question generation model is related to the selected answer. For example, one association technique uses an external parser to analyze expressions in a text (e.g., time, person, etc.) and the relationships between them, and selects expressions with the highest relevance coefficients with other expressions as answer candidates. These answer candidates are then input to the question generation model along with their expression classification to generate questions. The above techniques use expression classification so that the question generation model generates questions that are more related to the answer expressions, but they are highly dependent on the performance of the external parser and have the problem that phrases outside the range of named entities cannot be selected as answer candidates. For example, in the case of English, verb phrases that are not named entities, such as "take a nap," cannot be selected as answer candidates. There is also an approach to select answer candidates by learning a neural network-based answer selection model. Although this technique can select phrases as answers, the answer selection model does not pass additional information to the subsequent question generation model. Furthermore, it is difficult for the question generation model to capture the characteristics of question value from various types of answer candidates.

[0027] In order to improve the relevance between the selected answer and the generated question and the performance of the generated question-answer pair, an embodiment of the present invention provides the following model training method. As shown in Fig. 1, the model training method includes the step 11 of generating a first presentation text that corresponds to the interrogative word and presents a selection of answer candidates to a question including the interrogative word, and a second presentation text that presents a generation of a question including the interrogative word, based on an interrogative word in a target language and its interpretation.

[0028] Here, the target language can be a natural language such as Chinese, English, or Japanese, but this description mainly uses English as an example. In general, natural languages ​​have interrogative words. In English, interrogative words include "who", "when", "where", "howmuch", "howlong", etc. In Chinese, interrogative words include: (outside 1) TIFF0007679902000001.tif14127(outside 2) TIFF0007679902000002.tif15127(outside 3) TIFF0007679902000003.tif14127(outside 4) TIFF0007679902000004.tif14127, etc. The interpretation of each interrogative word can be obtained by searching a related dictionary (eg, Oxford English Dictionary) or encyclopedia (eg, Wikipedia), but the embodiment of the present application is not specifically limited thereto.

[0029] Specifically, we obtain all interrogatives from the target language that have been collected manually in advance, perform string matching between all interrogatives and problem samples in a training set that contains multiple original training samples, and retain the interrogatives that match the problem samples in the training set. For each of the retained interrogatives, we look up their interpretations in dictionaries and encyclopedias. For example, the interpretation of "when" is "the time at something happen."

[0030] For each interrogative word, a first presentation text (answer selection presentation text) and a second presentation text (question generation presentation text) corresponding to the interrogative word are generated based on the interrogative word and its interpretation. The first presentation text presents answer candidates to be selected for a question including the interrogative word. Specifically, the first presentation text corresponding to the interrogative word includes the interrogative word and an interpretation of the interrogative word. The second presentation text presents question generation including the interrogative word. The second presentation text corresponding to the interrogative word also includes the interrogative word and an interpretation of the interrogative word. In this way, the embodiment of the present invention can generate a first presentation text and a second presentation text corresponding to the same interrogative word for each interrogative word.

[0031] The following is a concrete example of creating presentation text using English as an example.

[0032] In this example, a first presentation template and a second presentation template are created in advance. In this description, the first presentation template may be called an answer selection prompt template, and the second presentation template may be called a question generation prompt template.

[0033] The first presentation template is a presentation text. The presentation text has two blank positions, which correspond to an interrogative word and an interpretation of the interrogative word. The presentation text presents a selection of answer candidates for a question containing the interrogative word. When generating a first presentation text corresponding to a certain interrogative word, the interrogative word and its interpretation are respectively filled in the corresponding blank positions in the first presentation template to obtain the first presentation text.

[0034] The second presentation template is also a presentation text. This presentation text also has two blank positions, which correspond to an interrogative word and an interpretation of the interrogative word. This presentation text suggests the generation of a question containing the interrogative word. When generating a second presentation text corresponding to a certain interrogative word, the interrogative word and its interpretation are filled in the corresponding blank positions in the second presentation template to obtain the second presentation text.

[0035] For example, the first presentation template (answer selection prompt template) looks like this:

[0036] Find a text to answer a question which asks about. Here, the blank position corresponding to the first underline is used to fill in the interrogative word, and the blank position corresponding to the second underline is used to fill in the interpretation of the interrogative word.

[0037] The second presentation template (question generation prompt template) looks like this:

[0038] Ask a question which asks about. Taking the interrogative word "When" as an example, the interpretation of "When" is "the time at something happen," and by filling in "When" and its interpretation into the above template, the first presentation text and the second presentation text are obtained as shown in Figure 2. Both of these presentation texts correspond to the interrogative word "When," so there is a one-to-one correspondence between these two presentation texts.

[0039] It should be noted that the above is merely an example of the presentation template / presentation text applied to the embodiment of the present invention, and the embodiment of the present invention may use other forms of presentation template / presentation text. For example, the first presentation template may be any of the following:

[0040] Select an answer to question which is about. Choose an answer to question which is about. The second presentation template may be one of the following:

[0041] Generate a question which is about. Provide a question which is about. The model training method also includes obtaining 12 a plurality of original training samples, including a text sample, a question sample including an interrogative word, and an answer sample corresponding to the question sample.

[0042] Here, an original data set is obtained that includes a number of original training samples. Typically, the training samples include a text sample, one or more questions, and an answer to each question. For ease of processing, the training samples that include a number of questions and answers may be split into a number of original training samples, such that each original training sample includes one question and its answer. Thus, each original training sample includes one text sample, one question sample, and an answer sample corresponding to the question sample. The text sample may be a paragraph of text, the question sample may be a question provided for the text, and the answer sample may be an answer to the question. Typically, the answer sample is a portion of characters of the text, i.e., a substring of the text of the sentence.

[0043] In the embodiment of the present invention, the original training sample refers to any machine-readable comprehension dataset, such as the SQuAD1.1 dataset published by Stanford University. In the dataset, each training data is manually tagged, and each training data includes one sentence, multiple questions related to the sentence, and corresponding answers (each answer is a partial string of a sentence). One sentence, one question related to the sentence, and the answer corresponding to the question can be used as training samples for the answer selection model and the question generation model. These sentences, questions, and answers are obtained in the training process of the answer selection model and the question generation model, but are input to the model in a different manner, that is, the answer is the training target for the answer selection model and the sentence and question are the input, whereas the question is the training target for the question generation model and the sentence and answer are the input.

[0044] The model training method also includes a step 13 of generating a first training sample for each original training sample using a first presented text corresponding to the text sample, the answer sample, and the interrogative word in the question sample of the original training sample, to obtain a first training set including a plurality of the first training samples; and generating a second training sample using a second presented text corresponding to the text sample, the question sample, the answer sample, and the interrogative word in the question sample of the original training sample, to obtain a second training set including a plurality of the second training samples.

[0045] Many pre-trained language models, such as the Bidirectional Encoder Representation from Transformer (BERT) model and the T5 model, have good performance for most natural language processing tasks, and can receive a sentence as a prompt and guide the model to achieve the specified task. As shown in FIG. 3, when the prompt sentence is connected to the target English sentence to be translated, "translate English to German," and input into the T5 model, the T5 model can output the German translation of the target English sentence. Based on this, an embodiment of the present invention provides a method for training a prompt-driven question generation related model, in which both the answer selection model and the question generation model use a prompt mechanism, and specifically, a new training sample is reconstructed based on the original training sample by introducing a prompt text, and then the related model is trained.

[0046] Specifically, the embodiment of the present invention includes in the presented text an interrogative word that reflects the type information of the answer candidate and information on its interpretation. For example, for the interrogative word "when", its answer candidate is a time type. Also, for the interrogative word "where", its answer candidate is a location type. By including the type information of the answer candidate in the presented text, the presented text inputs the answer type information as a suggestion to the answer selection model, and the answer selection model outputs only a specific type of answer candidate each time, so that the model can better capture the characteristics of "question value". Meanwhile, the presented text also conveys similar answer type information to the question generation model, so that the two models can be associated and the question generation model can generate a question that is more related to the selected answer candidate.

[0047] In the embodiment of the present invention, when generating training samples, a first training sample corresponding to each original training sample is generated. Specifically, the first training sample is generated using a first presented text corresponding to the text sample, the answer sample, and the interrogative word in the question sample of the original training sample. In this way, a first training set is obtained by obtaining a plurality of first training samples for a plurality of original training samples. Similarly, a second training sample corresponding to each original training sample is generated. Specifically, the second training sample is generated using a second presented text corresponding to the text sample, the question sample, the answer sample, and the interrogative word in the question sample of the original training sample. In this way, a second training set is obtained by obtaining a plurality of second training samples for a plurality of original training samples.

[0048] Suppose a question sample in a certain original training sample is "When were the Normans in Normandy?" and the interrogative word in it is "When", then for the interrogative word "When", two presentation texts are generated as shown in Fig. 4 to obtain a first training sample and a second training sample. Among them, the first training sample includes the first presentation text, the text sample of the original training sample, and the answer sample shown in Fig. 4, and the second training sample includes the second presentation text, the text sample of the original training sample, the question sample, and the answer sample shown in Fig. 4.

[0049] The model training method also includes a step 14 of training an answer selection model using the first training set and training a question generation model using the second training set.

[0050] In an embodiment of the present invention, the answer selection model is a natural language understanding model. Specifically, it is one of models such as BERT, Roberta, ALBERT, ERNIE, ELECTRA, etc. The answer selection model uses a pre-trained language model such as BERT as infrastructure. For example, a pre-trained language model such as BERT is obtained by pre-training with a large corpus based on the Transformer model architecture, and can realize natural language processing tasks such as answer selection. In addition, the question generation model is a natural language generation model. Specifically, it is one of the T5 model, GPT, BART models, etc. For example, the question generation model uses the T5 model as infrastructure. The T5 model is obtained by pre-training based on the Transformer encoder / decoder architecture, and has a text generation function.

[0051] Here, when training the answer selection model, the first submitted text and the text sample of each first training sample in the first training set are input to the answer selection model, the first submitted text includes the interrogative word in the question sample, and the answer selection model is trained using the corresponding answer sample output by the answer selection model as the target, to obtain the trained answer selection model.

[0052] FIG. 5 shows an example of training an answer selection model using the first training sample. In this case, the answer selection model receives the first presented text and text sample of the first training sample as input, and trains the answer sample of the first training sample output by the model as the training target. The answer generation model shown in FIG. 5 is a pre-trained language model such as BERT. The first presented text and the text sample are paired and input in the format shown in FIG. 5 to the answer selection model. That is, the special delimiter " <sep>" is used to concatenate the first presented text and the text sample as the input for the answer selection model. This input method is determined by the input method used during pre-training with the BERT model.

[0053] For example, Figure 5 Presentation Text 1: Find a text to answer a "when" question which asks about the time at something happens. Text sample:The Normans (Norman:Nourmands;French:Normands;Latin:Normanni) were the people who in the 10 th and 11 th Centuries gave their name to Normandy, a region in France…The distinct cultural and ethnic identity of the Normans emerged initially in the first half of the 10 th century,and it continued to evolve over the succeeding centuries. Sample Answer: in the first half of the 10 th century or in the 10 th and 11 th centuries. That is, in the process of training the answer selection model, the first presented text and the text sample in the first training sample are input to the answer selection model, and the answer selection model selects one text from the text samples and outputs it as an answer text.Then, the similarity between the answer text output by the answer selection model and the answer sample in the first training sample is calculated, and the model parameters of the answer selection model are optimized based on the calculated similarity, so that the answer text output by the answer selection model is made closer to the answer sample.By the above optimization process, a trained answer selection model is finally obtained.

[0054] When training the question generation model, the second presentation text, the text sample, and the answer sample of each second training sample of the second training concentration are input to the question generation model. The question generation model is trained using the corresponding question sample output by the question generation model as a target, and a trained question generation model is obtained.

[0055] Figure 6 shows an example of training the question generation model using the second training sample. In Figure 6, a triplet is formed using an answer sample, a second suggested text, and a text sample, and is input to the question generation model using the input method shown in Figure 6. This input method is determined by the T5 model, which is the infrastructure of the question generation model. Here, the second suggested text is prefixed with "Prompt:" and the sample text is prefixed with "Paragraph:", and then the two are concatenated. Meanwhile, the answer sample is stored in a hidden format, i.e., in XML format ( <answer> and< / answer> ) and enter the text in the format shown in the text sample (underlined part 6 in the figure).

[0056] For example, in Figure 6, Presentation Text 2: Ask a "when" question which asks about the time at something happens. Sample text:The Normans (Norman:Nourmands;French:Normands;Latin:Normanni) were the people who in the 10 th and 11 th Centuries gave their name to Normandy, a region in France…The distinct cultural and ethnic identity of the Normans emerged initially in the first half of the 10 th century,and it continued to evolve over the succeeding centuries. Answer sample:in the 10 th and 11 th centuries. Sample Question:When were the Normans in Normandy? That is, in the training process of the question generation model, the second presentation text, the text sample, and the question sample in the second training sample are input to the question generation model, and the question generation model generates one text and outputs it as a question text. Then, the similarity between the question text output by the question generation model and the question sample in the second training sample is calculated, and the model parameters of the question generation model are optimized based on the calculated similarity, thereby making the question text output by the question generation model closer to the question sample. Through the above optimization process, a finally trained question generation model is obtained.

[0057] Through the above steps, in the model training process, the embodiment of the present invention takes the first submitted text and text samples corresponding to the interrogative word as input to the answer selection model, and makes the answer selection model output a certain type of answer candidate; also takes the second submitted text, text samples and answer samples corresponding to the interrogative word as input to the question generation model, and associates the answer selection model and the question generation model with the submitted text containing the same interrogative word and its interpretation, so as to make the question generation model generate questions related to the selected answer candidates, thereby improving the performance of the question and answer pairs generated by the model.

[0058] The first training sample generated in step 13 above is a positive training sample, i.e., the answer to the question exists in the text sample. Considering that for a question having an interrogative word, the answer to the question may not exist in the text sample, the embodiment of the present invention further adds at least one third training sample (negative training sample) to the first training set when generating the first training set. Specifically, for a certain text sample (for convenience of explanation, referred to as the first text sample here) in the original training samples, a first interrogative word that is not included in all question samples corresponding to the first text sample is identified. Here, the first interrogative word exists in all interrogative words in the target language, and does not exist in all question samples corresponding to the first text sample. In addition, the first interrogative word exists in the question samples of the original training set, and does not exist in all problem samples corresponding to the first text sample. Then, a third training sample is generated using the first presented text corresponding to the first interrogative word, the blank answer, and the first text sample. The blank answer is, for example, "None" (blank as answer).

[0059] For example, suppose the following three original training samples, all of which are shared by all of them, contain the same text sample x:

[0060] in particular: Original training sample a: text sample x, question 1 (including interrogative word 1), answer 1; Original training sample b: text sample x, question 2 (including interrogative word 2), answer 2; Original training sample c: text sample x, question 3 (including interrogative word 3), answer 3; If there are five interrogatives, interrogatives 1 to 5, it can be seen that the interrogatives that are not included in any question sample corresponding to the same text sample x are interrogatives 4 and 5. Therefore, two third training samples (negative training samples) are generated as follows, i.e. 3rd training sample 1: text sample x, 1st presented text corresponding to interrogative word 4, answer (None); 3rd training sample 2: Text sample x, 1st presented text for interrogative word 5, answer (None).

[0061] By training multiple answer selection models with different proportions of positive and negative samples in the first training set, the proportion used by the answer selection model with the best performance index can be selected as the final proportion depending on the performance of the answer selection model.

[0062] Through step 14, the embodiment of the present invention obtains the trained answer selection model and question generation model. Then, the embodiment of the present invention uses the model as a target text to generate a question-answer pair. The target text may be a text input by a user. Specifically, the embodiment of the present invention selects one interrogative word from interrogative words in the target language (for ease of explanation, it is called a second interrogative word). Then, the first input text and the target text corresponding to the second interrogative word are input into the answer selection model to obtain an answer (for ease of explanation, it is called a target answer here) output by the answer selection model. Then, the second input text corresponding to the second interrogative word, the target text and the target answer are input into the question generation model to obtain a target question output by the question generation model. Thus, a pair of question-answer pairs consisting of a target question and a target answer is obtained. In the above process, the manner of generating the first input text and the second input text is the same as that performed in the model training process. The above process may refer to FIG. 5 and FIG. 6. In this case, the text samples in FIG. 5 and FIG. 6 are replaced with the above target text. The method allows one or more interrogative words to be used to obtain one or more question-answer pairs of the target text.

[0063] Based on the above method, an embodiment of the present invention further provides an apparatus for implementing the method. As shown in FIG. 7 , an embodiment of the present invention provides a model training device, including: a first generation module 71 for generating a first presentation text for presenting a selection of answer candidates to a question including the interrogative word and corresponding to the interrogative word based on an interrogative word in a target language and its interpretation; a first acquisition module 72 for acquiring a plurality of original training samples including a text sample, a question sample including the interrogative word, and an answer sample corresponding to the question sample; a second generation module 73 for generating a first training sample for each original training sample using a first presentation text corresponding to the text sample, the answer sample, and the interrogative word in the question sample of the original training sample to obtain a first training set including a plurality of the first training samples; and a training module 74 for training an answer selection model using the first training set and training a question generation model using the second training set.

[0064] In an embodiment of the present invention, the above modules improve the relevance of the trained model to the selected answers and the generated questions.

[0065] Optionally, the first training set includes at least a third training sample, and the model training apparatus includes a third generation module for identifying, for a first text sample, a first interrogative word that is not included in any question sample corresponding to the first text sample; and forming the third training sample using a first submitted text corresponding to the first interrogative word, a blank answer, and the first text sample.

[0066] Optionally, a first submitted text and text sample for each training sample in the first training set is input to an answer selection model, and the answer selection model is trained using the corresponding answer sample output by the answer selection model as a target to generate a trained answer selection model; a second submitted text, text sample and answer sample for each training sample in the second training set is input to a question generation model, and the question generation model is trained using the corresponding question sample output by the question generation model as a target to generate a trained question generation model.

[0067] Optionally, the model training device further includes a model application module for selecting a second interrogative word from interrogative words in the target language for a target text; inputting a first submission text corresponding to the second interrogative word and the target text into the answer selection model to obtain a target answer output by the answer selection model; and inputting a second submission text corresponding to the second interrogative word, the target text and the target answer into the question generation model to obtain a target question output by the question generation model.

[0068] Optionally, the first and second presentation texts corresponding to the interrogative word also include the interrogative word and an interpretation of the interrogative word, respectively.

[0069] Also optionally, the answer selection model is a natural language understanding model and the question generation model is a natural language generation model.

[0070] FIG. 8 is a block diagram showing a hardware configuration of a model training device according to an embodiment of the present invention. As shown in FIG. 8 , the model training apparatus 800 includes a processor 802 and a memory 804 for storing computer program instructions. When the computer program instructions are executed by the processor, the processor 802 executes the following processes: generate a first presentation text presenting a selection of answer candidates for a question including the interrogative word and a second presentation text presenting a generation of a question including the interrogative word based on an interrogative word in a target language and its interpretation; obtain a plurality of original training samples including a text sample, a question sample including an interrogative word, and an answer sample corresponding to the question sample; generate a first training sample for each original training sample using the text sample, the answer sample, and the first presentation text corresponding to the interrogative word in the question sample of the original training sample to obtain a first training set including the plurality of first training samples; generate a second training sample using the text sample, the question sample, the answer sample, and the second presentation text corresponding to the interrogative word in the question sample of the original training sample to obtain a second training set including the plurality of second training samples; train an answer selection model using the first training set, and train a question generation model using the second training set.

[0071] Moreover, as shown in FIG. 8, the model training device 800 also includes a network interface 801 , an input device 803 , a hard disk 805 , and a display device 806 .

[0072] These network interfaces and devices are connected to each other via a bus architecture. The bus architecture may include any number of interconnected buses and bridges. Specifically, various circuits such as one or more central processing units (CPUs) and / or graphics processors (GPUs) represented by processor 802, and one or more memories represented by memory 804 are connected. The bus architecture may also connect various other circuits such as peripheral devices, voltage regulators, power management circuits, etc. The bus architecture allows communication between these components. In addition to a data bus, the bus architecture may include a power bus, a control bus, and a status signal bus. All of these are well known in the art and will not be described in detail.

[0073] The network interface 801 is connected to a network (eg, the Internet, a local area network, etc.), receives data such as original training samples from the network, and stores the received data in the hard disk 805 .

[0074] The input device 803 receives various commands input by an operator and transmits them to the processor 802 for execution. The input device 803 includes a keyboard or a pointing device (e.g., a mouse, a trackball, a touch board, or a touch screen, etc.).

[0075] Display device 806 displays the results of processor 802 executing the instructions, such as displaying model training progress.

[0076] Memory 804 stores programs and data necessary for the execution of an operating system (OS), as well as data such as intermediate results during calculations by processor 802. Memory 804 in embodiments of the present invention may be volatile or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM) used as an external cache. Memory 804 in the apparatus and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0077] In some embodiments, memory 804 stores an operating system 8041 and application programs 8042, which may be executable modules or data structures or submodules or extension modules thereof.

[0078] The operating system 8041 includes various system programs, such as a framework layer, a core library layer, and a driving layer, and is used to realize various core operations and hardware-based tasks. The application program 8042 includes various application programs, such as a web browser, and is used to realize various application operations. The program for executing the method according to this embodiment is included in the application program 8042.

[0079] The method according to the embodiment of the present invention is applied to or implemented by the processor 802. The processor 802 is an integrated circuit board capable of processing signals. Each step of the method is implemented by an integrated logic circuit in hardware or instructions in software form in the processor 802. The processor 802 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-flash programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, capable of implementing or executing each method, step, and logic box disclosed in the embodiment of the present invention. The general-purpose processor may be a microprocessor or any general processor. Each step of the method according to the embodiment of the present invention may be implemented by being executed by a decoder in hardware, or may be implemented by a combination of hardware and software that can be implemented in the decoder. The software module is stored in a storage medium mature in the field, such as a random memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, or the like. The processor 802 reads information from the memory 804, which includes a storage medium on which the software is stored, and configures the hardware to implement the steps of the above-mentioned method.

[0080] The above-described embodiments may be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof, whereby, for hardware implementation, the processing unit may be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processors (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general purpose processors, controllers, microcontrollers, microprocessors, other electronic units performing the functions of the present invention, or a combination thereof.

[0081] Regarding software implementation, the above technology is realized by modules (e.g., processes, functions, etc.) that realize the functions described above. The software code is stored in a memory and executed by a processor. The memory may be implemented inside or outside the processor.

[0082] Specifically, the first training set further includes at least one third training sample, and the computer program is executed by the processor 802 to perform the steps of: identifying, for a first text sample, a first interrogative word that is not included in any question sample corresponding to the first text sample; and forming the third training sample using a first submitted text corresponding to the first interrogative word, a blank answer, and the first text sample.

[0083] In particular, when the computer program is executed by the processor 802, the steps of inputting a first presented text and text sample for each training sample in the first training set into an answer selection model, training the answer selection model using the corresponding answer sample output by the answer selection model as a target, and generating the trained answer selection model; inputting a second presented text, text sample, and answer sample for each training sample in the second training set into a question generation model, training the question generation model using the corresponding question sample output by the question generation model as a target, and generating the trained question generation model are realized.

[0084] In particular, by executing the computer program on the processor 802, the steps of selecting a second interrogative from among interrogatives in the target language for a target text; inputting a first input text corresponding to the second interrogative and the target text into the answer selection model to obtain a target answer output by the answer selection model; and inputting a second input text corresponding to the second interrogative, the target text and the target answer into the question generation model to obtain a target question output by the question generation model are realized.

[0085] Optionally, the first and second presentation texts corresponding to said interrogative word respectively include said interrogative word and an interpretation of said interrogative word.

[0086] Optionally, the answer selection model is a natural language understanding model and the question generation model is a natural language generation model.

[0087] Those skilled in the art of the present invention can easily imagine that the units and algorithm steps of each example described in the above disclosed embodiments can be realized by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed by either hardware or software depends on the specific application and design constraints of the invention. Those skilled in the art can realize the above functions in a way that suits the specific application, but it should not go beyond the scope of the present invention.

[0088] In addition, for convenience and conciseness of explanation, detailed descriptions of the specific operating processes of the above systems, devices and units will be omitted, since it is clear to those skilled in the art that reference can be made to the corresponding processes in the above embodiments.

[0089] It is easily conceivable that the method and apparatus disclosed in the multiple embodiments of the present invention can be realized in other forms. For example, the above-described apparatus is merely schematic. For example, the division of the units is merely one example of logical function allocation, and a different division method may be adopted when actually realizing the present invention. For example, a plurality of units or modules may be combined or integrated into another system, or some functions may be omitted or not executed. Note that the mutual connections, direct connections, or communicative connections shown or disclosed above are connections via a network interface. Indirect connections or communicative connections between devices or units may be electrical, mechanical, or other forms of connections.

[0090] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, i.e., they may be in the same location or distributed across multiple network units. Depending on actual needs, some or all of the units may be selected to achieve the objectives of the embodiments of the present invention.

[0091] Note that each functional unit according to the embodiment of the present invention may be integrated into one processing unit, may be physically independent, or may be integrated into one unit consisting of two or more functional units.

[0092] The functions can be realized in the form of a software functional unit and stored in a computer-readable storage medium when sold or used as an independent product. In this case, the technical solution of the present invention is essentially or contributes to the prior art or the technical solution is expressed in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (such as a personal computer, a server, or a network device) to execute all or part of the steps of the method according to each embodiment of the present invention. The storage medium includes various media capable of storing program code, such as a USB memory, a removable disk, a ROM, a RAM, a magnetic disk, or an optical disk.

[0093] The above description is a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto, and any modifications or replacements that can be easily conceived by any person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of the present invention should be based on the claims.< / sep>

Claims

1. 1. A computer implemented method for training a model, comprising: generating a first presentation text that presents a selection of answer candidates to a question including the interrogative word and a second presentation text that presents a generation of a question including the interrogative word, the first presentation text corresponding to the interrogative word and the second presentation text based on the interrogative word and its interpretation in the target language; obtaining a plurality of original training samples including a text sample, a question sample including an interrogative word, and an answer sample corresponding to the question sample; For each original training sample, generate a first training sample using a first provided text corresponding to the text sample, the answer sample, and the interrogative word in the question sample of the original training sample to obtain a first training set including a plurality of the first training samples; generate a second training sample using a second provided text corresponding to the text sample, the question sample, the answer sample, and the interrogative word in the question sample of the original training sample to obtain a second training set including a plurality of the second training samples; and training an answer selection model using the first training set and training a question generation model using the second training set; A method for training a model.

2. the first training set further comprises at least one third training sample; For a first text sample, a first interrogative word that is not included in any question sample corresponding to the first text sample is identified, and a third training sample is formed using a first submitted text corresponding to the first interrogative word, a blank answer, and the first text sample.

2. The method of claim 1, wherein the model training step is

3. Training an answer selection model using the first training set and training a question generation model using the second training set includes: inputting the first submitted text and text samples for each training sample in the first training set into an answer selection model, training the answer selection model using the corresponding answer samples output by the answer selection model as targets, and generating a trained answer selection model; and inputting the second presentation text, the text sample, and the answer sample for each training sample in the second training set into a question generation model, training the question generation model using the corresponding problem sample output by the question generation model as a target, and generating a trained question generation model; 2. The method of claim 1, wherein the model training step is

4. selecting a second interrogative word from the interrogative words in the target language for the target text; inputting the first input text corresponding to the second interrogative and the target text into the answer selection model to obtain a target answer output by the answer selection model; and inputting a second presentation text corresponding to the second interrogative, the target text, and the target answer into the question generation model to obtain a target question output by the question generation model; 2. The method of claim 1, wherein the model training step is

5. a first presentation text and a second presentation text corresponding to the interrogative word each including the interrogative word and an interpretation of the interrogative word; 2. The method of claim 1, wherein the model training step is

6. The answer selection model is a natural language understanding model, and the question generation model is a natural language generation model.

2. The method of claim 1, wherein the model training step is

7. a first generation module that generates a first presentation text that presents a selection of answer candidates to a question including the interrogative word and a second presentation text that presents a generation of a question including the interrogative word, the first presentation text corresponding to the interrogative word and the selection of answer candidates to a question including the interrogative word, based on an interrogative word in a target language and its interpretation; a first acquisition module for acquiring a plurality of original training samples, the training samples including a text sample, a question sample including an interrogative word, and an answer sample corresponding to the question sample; a second generation module for generating, for each original training sample, a first training sample using a first presentation text corresponding to the text sample, the answer sample, and the interrogative word in the question sample of the original training sample to obtain a first training set including a plurality of the first training samples, and generating a second training sample using a second presentation text corresponding to the text sample, the question sample, the answer sample, and the interrogative word in the question sample of the original training sample to obtain a second training set including a plurality of the second training samples; a training module that uses the first training set to train an answer selection model and uses the second training set to train a question generation model; A model training device comprising:

8. the first training set further comprises at least one third training sample; a third generation module for identifying, for a first text sample, a first interrogative word that is not included in any question sample corresponding to the first text sample; and forming the third training sample using a first submitted text corresponding to the first interrogative word, a blank answer, and the first text sample; 8. The model training device according to claim 7.

9. The training module includes: inputting the first submitted text and text samples for each training sample in the first training set into an answer selection model, training the answer selection model using the corresponding answer samples output by the answer selection model as targets, and generating a trained answer selection model; and inputting the second presentation text, the text sample, and the answer sample for each training sample in the second training set into a question generation model, training the question generation model using the corresponding question sample output by the question generation model as a target, and generating a trained question generation model; 9. A model training device according to claim 7 or 8.

10. a model application module for selecting a second interrogative word from interrogative words in the target language for a target text, inputting a first input text corresponding to the second interrogative word and the target text into the answer selection model to obtain a target answer output by the answer selection model, and inputting a second input text corresponding to the second interrogative word, the target text and the target answer into the question generation model to obtain a target question output by the question generation model; 8. The model training device according to claim 7.

11. a first presentation text and a second presentation text corresponding to the interrogative word each including the interrogative word and an interpretation of the interrogative word; 8. The model training device according to claim 7.

12. The answer selection model is a natural language understanding model, and the question generation model is a natural language generation model.

8. The model training device according to claim 7.

13. A program for causing a computer to execute the model training method according to any one of claims 1 to 6.

14. A computer-readable storage medium storing the program according to claim 13.

Citation Information

Patent Citations

  • Neural network question generation method and system based on interrogative classifier

    CN113094489A

  • Intelligent dialogue method and device based on financial knowledge graph, and electronic equipment

    CN113988071A

  • Question and answer method applied to mineral field knowledge graph, electronic device and storage medium

    CN115269806A

  • Dialog data generating apparatus, dialog data generating method, and program

    JP2019197363A

  • Method for generating triple sample, device, electronic device, and storage medium

    JP2022008207A