A training method and device for an answer extraction model
By filtering relevant questions and answer labels to be queried from the original corpus in machine reading comprehension model training and optimizing the loss function, the problem of low generalization performance in the Chinese answer extraction task is solved, and higher answer extraction accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202010825792.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-08-17
AI Technical Summary
Existing machine reading comprehension models cannot generate query problems that match certain argument types in the Chinese answer extraction task, and the losses considered during the training process are insufficient, resulting in low generalization performance of the model and low accuracy of predicting answers.
By determining the sample text from the original corpus and filtering the questions to be queried and the corresponding answer labels associated with the sample text in the pre-constructed questions, input the questions to be queried and the sample text to be queried and extracted the model, generate the target loss value and optimize the model to improve the accuracy and generalization performance of the training results of the answer extraction model.
Improve the accuracy and efficiency of the answer extraction model in Chinese tasks, improve the generalization performance of the model, and make the generated predicted answers more accurate.
Smart Images

Figure CN114077655B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and particularly to a method and apparatus for training an answer extraction model, a computing device, and a computer-readable storage medium. Background Art
[0002] With the rapid development of the Internet, more and more information is presented to users in the form of electronic texts. To help users quickly find the information they need in the vast amount of information, the concept of information extraction is proposed. Information extraction refers to extracting factual information from natural language texts and describing the information in a structured form; while machine reading comprehension is a research dedicated to teaching machines to read human languages and understand their connotations. Machine reading comprehension tasks focus more on the understanding of passage texts. Machines must learn relevant information from the passages by themselves, rather than using pre-set world knowledge and common sense to answer questions.
[0003] Currently, an important implementation method for training machines to understand human languages is to establish a machine reading comprehension model, and further train the established machine reading comprehension model to obtain the desired machine reading comprehension model, so as to find the answers to questions in text segments based on the trained machine reading comprehension model. However, in the current training process of machine reading comprehension models, for the Chinese answer extraction task, query questions that match certain argument types cannot be generated; in addition, the losses considered in the model training process are not sufficient, and the losses of predicted answers cannot be fully reflected. The generalization performance of the trained model is low, and the accuracy of the generated predicted answers is relatively low. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a method and apparatus for training an answer extraction model, a computing device, and a computer-readable storage medium to solve the technical defects existing in the prior art.
[0005] According to a first aspect of embodiments of the present application, a method for training an answer extraction model is provided, including:
[0006] Determine sample texts from the original corpus, and screen at least one query question associated with the sample texts and the corresponding answer tags in a pre-constructed question set;
[0007] Input any one of the query questions and the sample texts into a pre-trained answer extraction model to determine the answer extraction result of the query question;
[0008] Generate a target loss value of the answer extraction model based on the answer extraction result and the answer tags, and optimize the answer extraction model based on the target loss value to obtain a target answer extraction model.
[0009] Optionally, the question set is constructed in the following manner:
[0010] Extract the event type label and answer type label of the text from the original corpus;
[0011] Integrate the event type label and the answer type label to generate a question label;
[0012] Generate a query question that matches the question label according to the category to which the answer type label included in the question label belongs, and construct a question set based on the query question.
[0013] Optionally, generating a query question that matches the question label according to the category to which the answer type label included in the question label belongs includes:
[0014] If the answer type label included in the question label is the first category, obtain a predefined question template, and construct a query question that matches the question label based on the question label and the question template;
[0015] If the answer type label included in the question label is the second category, perform statistical analysis on the event sentences in the original corpus related to the answer type label of the second category, and construct a query question that matches the question label according to the analysis result.
[0016] Optionally, inputting any one of the questions to be queried and the sample text into a pre-trained answer extraction model to determine the answer extraction result of the question to be queried includes:
[0017] Input any one of the questions to be queried and the sample text as an input set into the answer extraction model. The vector encoding module of the answer extraction model sums the character vector, text vector, and position vector corresponding to each word unit in the input set to generate an encoding vector corresponding to each word unit;
[0018] Calculate the probability distribution of the start position and end position of each word unit as the predicted answer corresponding to the question to be queried based on the encoding vector;
[0019] Determine the answer extraction result corresponding to the question to be queried according to the probability distribution of the start position and end position.
[0020] Optionally, determining the answer extraction result corresponding to the question to be queried according to the probability distribution of the start position and end position includes:
[0021] Use the position of the word unit with the highest probability in the probability distribution of the start position as the start position of the answer in the sample text;
[0022] Use the position of the word unit with the highest probability in the probability distribution of the end position in the sample text as the end position of the answer; and,
[0023] Use the word units between the start position and the end position as the answer extraction result.
[0024] Optionally, generating the target loss value of the answer extraction model based on the answer extraction result and the answer label includes:
[0025] Determine the start position loss of the start position of the answer extraction result in the sample text based on the probability distribution of the start position and the probability of the target start position in the answer label;
[0026] Determine the end position loss of the end position of the answer extraction result in the sample text based on the probability distribution of the end position and the probability of the target end position in the answer label;
[0027] Determine the length loss of the answer extraction result based on the start position and the end position;
[0028] Calculate the target loss value based on the start position loss, the end position loss, and the length loss.
[0029] Optionally, calculating the target loss value based on the start position loss, the end position loss, and the length loss includes:
[0030] Calculate the weighted sum of the start position loss, the end position loss, and the length loss as the target loss value.
[0031] Optionally, the vector encoding module includes an embedding layer and n stack layers;
[0032] Correspondingly, generating the encoding vector corresponding to each word unit includes:
[0033] S11. Input the query question and the sample text as an input set into the embedding layer to obtain a corresponding input vector;
[0034] S12. Input the input vector into the first stack layer to obtain the output vector of the first stack layer;
[0035] S13. Input the output vector of the i-th stack layer into the (i + 1)-th stack layer to obtain the output vector of the (i + 1)-th stack layer, where i ∈ [1, n], and i starts from 1;
[0036] S14. Determine whether i is equal to n - 1. If so, execute step S15; if not, execute step S13.
[0037] S15. Output the output vector of the nth stack layer as the encoding vector of each word unit in the input set.
[0038] According to the second aspect of the embodiments of the present application, there is provided a training device for an answer extraction model, including:
[0039] A screening module, configured to determine a sample text from an original corpus, and screen at least one query question associated with the sample text and a corresponding answer label in a pre-constructed question set;
[0040] A determination module, configured to input any one of the query questions and the sample text into a pre-trained answer extraction model to determine the answer extraction result of the query question;
[0041] A calculation module, configured to generate a target loss value of the answer extraction model based on the answer extraction result and the answer label, and optimize the answer extraction model based on the target loss value to obtain a target answer extraction model.
[0042] According to the third aspect of the embodiments of the present application, there is provided a computing device, including a memory, a processor, and computer instructions stored on the memory and executable on the processor. When the processor executes the instructions, the steps of the training method of the answer extraction model are implemented.
[0043] According to the fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium storing computer instructions, and when the instructions are executed by a processor, the steps of the training method of the answer extraction model are implemented.
[0044] In the embodiments of the present application, after determining the sample text, by screening the query questions associated with the sample text and the corresponding answer labels in the pre-constructed question set, and inputting the query questions and the sample text into the answer extraction model for model training, it is beneficial to improve the accuracy of the training result of the answer extraction model and improve the efficiency of model training; in addition, by calculating the target loss value between the answer extraction result output by the model and the answer label, and optimizing the answer extraction model based on the target loss value, it is beneficial to improve the generalization performance of the answer extraction model. Description of the Drawings
[0045] Figure 1 is a structural block diagram of a computing device provided by an embodiment of the present application;
[0046] Figure 2It is a flowchart of a method for training an answer extraction model provided by an embodiment of the present application;
[0047] Figure 3 It is a schematic diagram of the architecture of a BERT model provided by an embodiment of the present application;
[0048] Figure 4 It is a flowchart of the generation process of encoding vectors provided by an embodiment of the present application;
[0049] Figure 5 It is a schematic diagram of the generation of input vectors of an embedding layer provided by an embodiment of the present application;
[0050] Figure 6 It is a schematic diagram of a method for training an answer extraction model provided by an embodiment of the present application;
[0051] Figure 7 It is a schematic diagram of the structure of a training device for an answer extraction model provided by an embodiment of the present application. Detailed implementation manners
[0052] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.
[0053] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "said", and "the" used in one or more embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more of the associated listed items.
[0054] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first.
[0055] First, the noun terms related to one or more embodiments of the present invention are explained.
[0056] Event extraction: According to the type of event occurrence, extract information such as the trigger word that causes the event to occur, the argument roles participating in the event, and the event type to which they belong.
[0057] Hard-Loss: Harder loss. Here it refers to the triplet loss. During training, a triplet is formed by given an anchor, a positive sample and a negative sample, and a margin parameter is set to make the model strive to reduce the distance between positive sample pairs while pushing away the distance between negative sample pairs.
[0058] MRC: Machine Reading Comprehension. The goal of this task is to extract the answer range from an article through a given question.
[0059] Token: Before any actual processing of the input text, it needs to be segmented into language units such as words, punctuation marks, numbers or letters, and these units are called tokens. For English text, a token can be a word, a punctuation mark, a number, etc. For Chinese text, the smallest token can be a phrase, a character, a punctuation mark, a number, etc.
[0060] BERT model: A bidirectional attention neural network model. The BERT model can predict the current word through the context on both the left and right sides and predict the next sentence through the current sentence. The goal of the BERT model is to use large-scale unlabeled corpora for training to obtain the semantic representation of text containing rich semantic information, and then fine-tune the semantic representation of the text in a specific NLP task and finally apply it to this NLP task.
[0061] Sequence labeling: Simply put, it is to given a sequence and use a related model to make a label for each element in the sequence, which can be entity annotation, part-of-speech annotation, etc.
[0062] CE: Cross-entropy loss function. It is commonly used in classification problems. The weighted sum of the predicted probability vector and the true label vector is used, and the loss is reduced through backpropagation to make it tend to the true label.
[0063] In this application, a training method and device for an answer extraction model, a computing device, and a computer-readable storage medium are provided, and will be described in detail one by one in the following embodiments.
[0064] Figure 1 FIG. shows a structural block diagram of a computing device 100 according to an embodiment of the present application. The components of the computing device 100 include but are not limited to a memory 110 and a processor 120. The processor 120 is connected to the memory 110 through a bus 130, and a database 150 is used to store data.
[0065] The computing device 100 further includes an access device 140, which enables the computing device 100 to communicate via one or more networks 160. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 140 may include one or more of any type of wired or wireless network interface (e.g., Network Interface Card (NIC)), such as an IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0066] In one embodiment of the present application, the above components of the computing device 100 and Figure 1 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 1 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present application. Those skilled in the art may add or replace other components as needed.
[0067] The computing device 100 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 100 can also be a mobile or stationary server.
[0068] Wherein, the processor 120 can execute Figure 2 the steps in the training method of the answer extraction model shown. Figure 2 A flowchart of the training method of the answer extraction model according to an embodiment of the present application is shown, including steps 202 to 206.
[0069] Step 202, determine sample texts from the original corpus, and screen at least one query question and the corresponding answer label associated with the sample texts in a pre-constructed question set.
[0070] Currently, for the event extraction task, one processing method is to divide it into two subtasks: trigger word extraction and argument extraction for processing. Trigger word extraction mostly uses sequence labeling to obtain the entity labels of trigger words, but this processing method is not applicable to the case of argument entity overlap.
[0071] Based on this, a training method for an answer extraction model provided in the embodiments of this specification treats the answer extraction task as an MRC task, determines sample texts from the original corpus, and screens at least one query question associated with the sample texts and corresponding answer labels in a pre-constructed question set. Input any one of the query questions and the sample texts into a pre-trained answer extraction model, determine the answer extraction result of the query question, generate a target loss value for the answer extraction model based on the answer extraction result and the answer label, and optimize the answer extraction model based on the target loss value to obtain a target answer extraction model.
[0072] Specifically, the answer extraction model in the embodiments of this specification is composed of a vector encoding module (BERT model), a start index prediction model, and an end index prediction model. Among them, the schematic architecture diagram of the BERT model is as Figure 3 shown, including 12 stacked layers, and these 12 stacked layers are connected in sequence. Each stacked layer also includes: a self-attention layer, a first normalization layer, a feed-forward layer, and a second normalization layer. The text composed of the article and the question is used as an input set and input into the embedding layer to obtain a text vector, and then the text vector is input into the first stacked layer, and the output vector of the first stacked layer is input into the second stacked layer... and so on. Finally, the output vector of the last stacked layer is obtained. The output vector of the last stacked layer is used as the representation vector of each word unit and input into the feed-forward layer for processing to obtain the encoded vector of the input set.
[0073] In practical applications, the original corpus obtained contains a lot of texts, and the texts are written texts containing certain information content, which can be texts of various lengths such as a sentence, a paragraph, multiple paragraphs, an article, or multiple articles, and this application does not limit this.
[0074] Specifically, when implementing, the construction of the question set can be specifically achieved through the following method:
[0075] Extract the event type label and answer type label of the text from the original corpus;
[0076] Integrate the event type label and the answer type label to generate a question label;
[0077] Generate a query question that matches the question label according to the category to which the answer type label included in the question label belongs, and construct a question set based on the query question.
[0078] Specifically, the event type label and the argument type label are respectively extracted from the text of the original corpus, and then these two labels are integrated to generate a question label (actually a process of traversal and combination), and then a question that matches each question label is constructed to generate a question set.
[0079] For example, the event type tags extracted from the text of the original corpus include: marriage, promotion, and judgment; the argument type tags extracted include: time and location; then, these two types of tags are integrated, and the generated question tags include: promotion - time, marriage - time, judgment - time, promotion - location, marriage - location, and judgment - location.
[0080] After generating the question tags, it is necessary to construct the MRC questions for the use of the question tags. The construction of the actual MRC questions is equivalent to a description of the question tags. A targeted description is equivalent to giving some more semantic information. For example, if the question tag is: promotion - time, the MRC question constructed for it can be: Find out the time when the promotion event occurred;
[0081] However, since some of the tags in the question tags belong to tags that cannot be given appropriate questions, it is necessary to classify the question tags and construct the MRC questions that match them according to the classification results.
[0082] Specifically, when implementing, query questions that match the question tags are generated according to the category to which the answer type tags included in the question tags belong. Specifically, it can be achieved through the following methods:
[0083] If the answer type tag included in the question tag is of the first category, a predefined question template is obtained, and a query question that matches the question tag is constructed based on the question tag and the question template;
[0084] If the answer type tag included in the question tag is of the second category, statistical analysis is performed on the event sentences in the original corpus that are related to the answer type tag of the second category, and a query question that matches the question tag is constructed according to the analysis results.
[0085] Specifically, the first category means that the arguments in the answer type tags have generality. For example, arguments such as time, person, number of people, and organization in the event sentence have generality, and the expressed meanings are basically the same. Therefore, for question tags containing answer type tags of the first category, only a string of the event type needs to be added before each question for distinction. For example, "promotion - time" and "marriage - time" belong to general tags. For general tags, the question template (find out the time when the XX event occurred) can be used, and based on the question template and the question tag, the MRC question can be synthesized with code. Therefore, the MRC question corresponding to "promotion - time" is: Find out the time when the promotion event occurred; the MRC question corresponding to "marriage - time" is: Find out the time when the marriage event occurred.
[0086] The second category indicates that the arguments in the answer type label are not universal. For example: limit up - stocks. Problem labels containing answer type labels of the second category cannot generate appropriate questions. For such labels, event sentence statistical analysis can be used to set a more universal and detailed question description. For example: limit up - stocks, that is, search for events of the type "stocks - limit up" in the text of the original corpus, then determine the universal language description of this type of event, analyze these universal language descriptions, and determine the universal and detailed MRC questions that match them. For example, for "stocks - limit up", the corresponding MRC question is: Find the amplitude of the limit up in the stock limit up event, including the percentage increase, halt, decline, etc.
[0087] By classifying the problem labels and generating MRC questions that match the classification results, a question set can be constructed. Later, after determining the sample text, the query questions associated with the sample text can be screened from the question set, and the query questions and the sample text are input into the answer extraction model for model training, which is beneficial to improving the accuracy of the training results of the answer extraction model and the efficiency of model training.
[0088] Step 204, input any one of the query questions and the sample text into a pre-trained answer extraction model to determine the answer extraction result of the query question.
[0089] Specifically, after constructing the question set, determine the sample text from the original corpus, and screen at least one query question and the corresponding answer label associated with the sample text from the constructed question set, and input the query question and the sample text into the pre-trained answer extraction model to obtain the answer extraction result corresponding to the query question.
[0090] In practical applications, first select a sentence or a paragraph of sample text from the original corpus as the given event sequence, and then determine the query questions according to the entities in the event sequence. In the embodiments of this specification, the query questions are the questions to be answered, and the query questions can be questions associated with the information content in the sample text.
[0091] For example, if the given event sequence is "Zhang San goes to Beijing", the sequence length n is 5, and the entities in this event sequence are "Zhang San" and "Beijing", then the query questions that can be determined according to these two entities can be "Where to go?" and "Who is the person in the event?", and their corresponding answers are "Beijing" and "Zhang San";
[0092] After determining the query questions, the query questions and the event sequence can be used as the input set and input into the BERT model (vector encoding module) in the form of a string to obtain the encoded vector output by the model.
[0093] In specific implementation, any one of the to-be-query questions and the sample text are input into a pre-trained answer extraction model to determine the answer extraction result of the to-be-query question, which can be specifically implemented in the following manner:
[0094] Any one of the to-be-query questions and the sample text are used as an input set and input into the answer extraction model. The vector encoding module of the answer extraction model sums the word vectors, text vectors, and position vectors corresponding to each word unit in the input set to generate an encoded vector corresponding to each word unit;
[0095] Based on the encoded vector, calculate the probability distribution of the start position and the end position of each word unit as the predicted answer corresponding to the to-be-query question;
[0096] According to the probability distribution of the start position and the end position, determine the answer extraction result corresponding to the to-be-query question.
[0097] Furthermore, the vector encoding module includes an embedding layer and n stack layers; Figure 4 is a flowchart of the generation process of the encoded vector provided by an embodiment of this specification. Refer to Figure 4 to generate the encoded vector corresponding to each word unit, which can be specifically implemented in the manner shown in steps 402 to 410:
[0098] Step 402: Input the to-be-query question and the sample text as an input set into the embedding layer to obtain a corresponding input vector.
[0099] Refer to Figure 5 Figure 5 is a schematic diagram of the generation of the input vector. Among them, the input set includes two sentences, "Where to go?" and "Zhang San goes to Beijing". Among them, "Zhang San goes to Beijing" is used as the target text, and "Where to go?" is used as the question.
[0100] Among them, the input vector generated by the embedding layer is formed by summing the following 3 types of vectors:
[0101] Word unit vector - the vector corresponding to each word unit;
[0102] Sentence vector - the sentence vector to which each word unit belongs;
[0103] Position vector - the vector generated by the position corresponding to each word unit.
[0104] Step 404: Input the input vector into the first stack layer to obtain the output vector of the first stack layer;
[0105] Step 406: Input the output vector of the i-th stack layer into the (i + 1)-th stack layer to obtain the output vector of the (i + 1)-th stack layer, where i ∈ [1, n] and i starts from 1;
[0106] Step 408: Determine whether i is equal to n - 1. If so, execute Step 405; if not, execute Step 403;
[0107] Step 410: Output the output vector of the n-th stack layer as the encoding vector of each word unit in the input set.
[0108] Specifically, the input set can adopt the following format: [[cls], question, [sep], sample text, [sep]].
[0109] If it is determined that the question to be queried is "Where to go?" and "Who is the person in the event?", and the event sequence is "Zhang San goes to Beijing", then the input set can be two sentences: "Where to go?" and "Zhang San goes to Beijing". Among them, "Zhang San goes to Beijing" is used as the sample text, and "Where to go?" is used as the question. The input format is: [[cls], go, where,?, [sep], Zhang, San, goes, to, Beijing, [sep]]; The specific schematic diagram is as Figure 4 shown.
[0110] Or, the input set can be two sentences: "Who is the person in the event?" and "Zhang San goes to Beijing". Among them, "Zhang San goes to Beijing" is used as the sample text, and "Who is the person in the event?" is used as the question. The input format is: [[cls], person, in, the, event, is, who,?, [sep], Zhang, San, goes, to, Beijing, [sep]];
[0111] For example, the sample text includes "Zhang San goes to Beijing", and the query question includes "Where to go?". The above sample text and query question are tokenized to obtain the word unit set [[cls], go, where,?, [sep], Zhang, San, goes, to, Beijing, [sep]]. Among them, CLS is the sentence start flag symbol, and SEP is the sentence separation flag symbol. After embedding the above word unit set and inputting it into the BERT model, the output vector of the last stack layer of the model is used as the representation vector of each word unit and input into the feed-forward layer for processing to obtain the encoding vector of the input set as [A1, A2,..., A 10 、A 11 .
[0112] In the embodiments of this specification, by pre - constructing a question set, after determining the sample text, the query questions associated with the sample text and the corresponding answer tags are screened in the question set, and the query questions and the sample text are input into an answer extraction model for model training. The query questions are equivalent to a priori information of the model, and the model outputs an answer extraction result according to the a priori information, which is beneficial to improving the accuracy of the output result.
[0113] After the BERT model outputs the encoded vector, two binary classification strategies can be used for each token in the encoded vector to predict whether it is the start or end position index of an entity (answer), which are specifically obtained by the classification functions shown in Equation (1) and Equation (2):
[0114] P start = softmax eachrow (E·T start )∈R n×2 Equation (1)
[0115] P end = softmax eachrow (E·T end )∈R n×2 Equation (2)
[0116] Among them, Equation (1) is used to predict the probability that each token in the encoded vector is the start position index of an entity; Equation (2) is used to predict the probability that each token in the encoded vector is the end position index of an entity; T in Equation (1) start and T in Equation (2) end are preset model parameters, which are equivalent to the initial weights of the tokens;
[0117] After calculating the probability that each token is the start or end position index of an entity, the probability results can be screened, and an index set is constructed based on the start or end position index with the largest probability in the screening results.
[0118] In specific implementation, according to the probability distributions of the start position and the end position, the answer extraction result corresponding to the query question is determined, which can be specifically implemented in the following ways:
[0119] Take the position of the token with the largest probability in the probability distribution of the start position in the sample text as the start position of the answer;
[0120] Take the position of the token with the largest probability in the probability distribution of the end position in the sample text as the end position of the answer; and,
[0121] Take the tokens between the start position and the end position as the answer extraction result.
[0122] Specifically, after calculating the probabilities of each word unit being the start or end position index of an entity, the probability results can be filtered, and an index set can be constructed based on the start or end position index with the largest probability in the filtered results;
[0123] In practical applications, filtering the probabilities of each word unit being the start position index of an entity, the formula for constructing the start position index set based on the start position index with the largest probability in the filtered results is as shown in Equation (3); filtering the probabilities of each word unit being the end position index of an entity, the formula for constructing the end position index set based on the end position index with the largest probability in the filtered results is as shown in Equation (4):
[0124]
[0125]
[0126] Among them, i in Equation (3) and j in Equation (4) respectively represent the i-th or j-th word unit in the sample text, represents the probability that the i-th word unit is the start position index of an entity, represents the probability that the j-th word unit is the end position index of an entity.
[0127] For example, assuming the sample text is "Zhang San goes to Beijing", the first word unit in the sample text is "Zhang", the second word unit is "San", and so on, the fifth word unit is "Jing"; if the probabilities of each word unit in the sample text being the start position of the answer are [x1, x2, x3, x4, x5] respectively, and the probabilities of each word unit being the end position of the answer are [y1, y2, y3, y4, y5] respectively, among which, among the probabilities of the answer start position, x4 has the largest probability value, then among the probabilities of the answer end position, y5 has the largest probability value, then It can be seen from this that the entity (answer) corresponding to the query question is the 4th and 5th word units in the sample text, and it can be determined that the answer corresponding to the query question is "Beijing".
[0128] In the embodiments of this specification, by calculating the probabilities of each word unit being the start position index or end position index of the answer, and filtering according to the calculation results, taking the word unit with the largest probability as the start or end position in the answer extraction result, and taking the word units between the start position and the end position as the answer extraction result, it is beneficial to ensure the accuracy of the answer extraction result;
[0129] Step 206: Generate the target loss value of the answer extraction model based on the answer extraction result and the answer label, and optimize the answer extraction model based on the target loss value to obtain the target answer extraction model.
[0130] Specifically, after obtaining the answer extraction result of the question to be queried output by the model, the target loss value of the answer extraction model can be calculated based on the answer extraction result and the answer label, and the answer extraction model can be optimized based on the target loss value to obtain the target answer extraction model.
[0131] In specific implementation, to generate the target loss value of the answer extraction model based on the answer extraction result and the answer label, it can be specifically implemented in the following ways:
[0132] Determine the start position loss of the answer extraction result in the sample text based on the probability distribution of the start position and the probability of the target start position in the answer label;
[0133] Determine the end position loss of the answer extraction result in the sample text based on the probability distribution of the end position and the probability of the target end position in the answer label;
[0134] Determine the length loss of the answer extraction result based on the start position and the end position;
[0135] Calculate the target loss value based on the start position loss, the end position loss, and the length loss.
[0136] Furthermore, calculate the target loss value based on the start position loss, the end position loss, and the length loss, that is, calculate the weighted sum of the start position loss, the end position loss, and the length loss as the target loss value.
[0137] In the embodiments of this specification, the triplet loss function (Hard-Loss loss function) is used to calculate the target loss value, and the specific calculation formula is shown in Equations (5), (6), and (7):
[0138]
[0139] L start = CE(P start , Y start ) Equation (6)
[0140] L end = CE(P end , Y end ) Equation (7)
[0141] Among them, k in formula (5) represents the number of elements in the set, E istart is an element in the set; and is an element in the set, represents the ending index of the positive sample that matches the i-th starting index; represents the ending index of the negative sample; α is the boundary parameter;
[0142] Y in formula (6) start is the probability of the starting position index of the true answer label corresponding to the query problem; Y in formula (7) end is the probability of the ending position index of the true answer label corresponding to the query problem;
[0143] When using the Hard-Loss loss function, by giving an anchor (i.e., the target sample), a positive sample and a negative sample to form a triple, and additionally setting a boundary parameter, the model is made to shorten the distance between the target sample and the positive sample, while increasing the distance between the target sample and the negative sample.
[0144] The final target loss value is the weighted sum of the above three losses, that is, the target loss value L = ω1·L hard + ω2·L start + ω3·L end ;
[0145] Among them, ω1, ω2 and ω3 are the weights corresponding to the three loss values respectively. In practical applications, the weights corresponding to the three loss values, the boundary parameter, T start and T end can be set according to actual needs and are not restricted here.
[0146] If it is determined that the probabilities of the 1st word unit and the 2nd word unit as the starting position indices are the largest according to the calculation results of the probabilities of each word unit in the sample text as the starting position of the answer, then If it is determined that the probabilities of the 3rd word unit and the 4th word unit as the ending position indices are the largest according to the calculation results of the probabilities of each word unit in the sample text as the ending position of the answer, then If the ending index of the positive sample that matches the 1st starting index (E 1start ) is E 3end , then the ending index of the negative sample that matches the 1st starting index (E 1start ) is E 4end , similarly, if the ending index of the positive sample that matches the 2nd starting index (E 2start ) is E4end , the end index of the negative sample that matches the second start index (E 2start ) is E 3end ;
[0147] From this, it can be obtained that
[0148] L start and L end are respectively calculated according to the cross-entropy loss function, that is, the probability corresponding to the predicted word unit as the start or end position index of the answer is weighted and calculated with the probability corresponding to the start or end position index of the true answer;
[0149] L hard 、L start and L end After all are calculated, the three losses are weighted and calculated according to their respective corresponding weights to obtain the target loss value.
[0150] After calculating the target loss value, the parameters of the model can be adjusted according to the target loss value to achieve model optimization.
[0151] In addition, after obtaining the answer extraction result, the answer extraction result can be compared with the answer label. If the accuracy of the answer extraction result does not meet the preset condition, the to-be-query question in the question set can be optimized, and the specific optimization method can be determined according to actual needs and is not limited here.
[0152] In the embodiments of this specification, by comparing the answer extraction result with the answer label and optimizing the to-be-query question corresponding to the answer extraction result whose accuracy does not meet the preset condition, it is beneficial to reduce the problem of argument entity overlap to a certain extent; in addition, after obtaining the answer extraction result, the triple loss function is used to calculate the target loss value, and the model parameters are adjusted based on the target loss value, which increases the difficulty of the model training process, enables the model to better distinguish positive and negative samples matching the argument entity, and is beneficial to enhancing the generalization performance of the model.
[0153] The following further illustrates this embodiment with specific examples.
[0154] The schematic diagram of the training method of the answer extraction model provided by the embodiments of this specification is as Figure 6 shown. First is the generation of MRC questions, that is, the event type label and the argument type label are respectively extracted from the text of the original corpus, and then these two types of labels are integrated to generate question labels (actually a process of traversal and combination), and then for each question label, an MRC question matching it is constructed;
[0155] After the MRC question construction is completed, the question q to be queried y is combined with the event sequence X into a string and input into BERT to output a context representation matrix E;
[0156] If the sample text corresponding to the event sequence X is "Zhang San goes to Beijing", the question to be queried can be "Where to go?", and the string is [[cls], go, where,?, [sep], Zhang, San, go, to, Beijing, [sep]]. Input the string into the BERT model to obtain the encoded vector output by the model. Input the encoded vector into the start index prediction model to obtain the probability that each token in the encoded vector is the start position index of the answer, and input the encoded vector into the end index prediction model to obtain the probability that each token in the encoded vector is the end position index of the answer. After obtaining the probability that each token is the start or end position index of the answer, the probability results can be filtered, and the answer extraction result can be determined based on the start or end position index with the highest probability in the filtered results;
[0157] Further, after obtaining the answer extraction result, use the Hard-Loss loss function to calculate the target loss value, and adjust the parameters of the model according to the target loss value to achieve model optimization.
[0158] The training method of the answer extraction model provided by the embodiments of the present application, after determining the sample text, by screening the question to be queried associated with the sample text and the corresponding answer label in the pre-constructed question set, and inputting the question to be queried and the sample text into the answer extraction model for model training, is beneficial to improving the accuracy of the training result of the answer extraction model and improving the efficiency of model training; in addition, by calculating the target loss value between the answer extraction result output by the model and the answer label, and optimizing the answer extraction model based on the target loss value, is beneficial to improving the generalization performance of the answer extraction model.
[0159] Corresponding to the above method embodiments, the present application also provides an embodiment of a training device for an answer extraction model, Figure 7 showing a schematic structural diagram of a training device for an answer extraction model according to an embodiment of the present application. As Figure 7 shown, the device 700 includes:
[0160] A screening module 702, configured to determine a sample text from the original corpus, and screen at least one question to be queried associated with the sample text and the corresponding answer label in a pre-constructed question set;
[0161] A determination module 704, configured to input any one of the questions to be queried and the sample text into a pre-trained answer extraction model to determine the answer extraction result of the question to be queried;
[0162] A calculation module 706, configured to generate a target loss value of the answer extraction model based on the answer extraction result and the answer label, and optimize the answer extraction model based on the target loss value to obtain a target answer extraction model.
[0163] Optionally, the training device of the answer extraction model further includes:
[0164] A label extraction module, configured to extract an event type label and an answer type label of the text from the original corpus;
[0165] A label generation module, configured to integrate the event type label and the answer type label to generate a question label;
[0166] A question set construction module, configured to generate a query question matching the question label according to the category to which the answer type label included in the question label belongs, and construct a question set based on the query question.
[0167] Optionally, the question set construction module includes:
[0168] A first question generation module, configured to, if the answer type label included in the question label is of the first category, obtain a predefined question template, and generate a query question matching the question label based on the question label and the question template;
[0169] A second question generation module, configured to, if the answer type label included in the question label is of the second category, perform statistical analysis on the event sentences related to the answer type label of the second category in the original corpus, and generate a query question matching the question label according to the analysis result.
[0170] Optionally, the determination module 704 includes:
[0171] An encoding vector generation sub-module, configured to input any one of the query questions and the sample text as an input set into the answer extraction model, and the vector encoding module of the answer extraction model sums the word vectors, text vectors, and position vectors corresponding to each word unit in the input set to generate an encoding vector corresponding to each word unit;
[0172] A calculation sub-module, configured to calculate a probability distribution of the start position and the end position of each word unit as the predicted answer corresponding to the query question based on the encoding vector;
[0173] An answer extraction result determination sub-module, configured to determine the answer extraction result corresponding to the query question according to the probability distribution of the start position and the end position.
[0174] Optionally, the answer extraction result determination sub-module includes:
[0175] A start position determination unit, configured to use the position of the word unit with the highest probability in the probability distribution of the start position in the sample text as the start position of the answer;
[0176] An end position determination unit, configured to use the position of the word unit with the highest probability in the probability distribution of the end position in the sample text as the end position of the answer;
[0177] An answer extraction result determination unit, configured to use the word units between the start position and the end position as the answer extraction result.
[0178] Optionally, the calculation module 706 includes:
[0179] A first loss calculation sub-module, configured to determine the start position loss of the answer extraction result at the start position in the sample text based on the probability distribution of the start position and the probability of the target start position in the answer label;
[0180] A second loss calculation sub-module, configured to determine the end position loss of the answer extraction result at the end position in the sample text based on the probability distribution of the end position and the probability of the target end position in the answer label;
[0181] A third loss calculation sub-module, configured to determine the length loss of the answer extraction result based on the start position and the end position;
[0182] A target loss value calculation sub-module, configured to calculate the target loss value based on the start position loss, the end position loss, and the length loss.
[0183] Optionally, the target loss value calculation sub-module is further configured to:
[0184] Calculate the weighted sum of the start position loss, the end position loss, and the length loss as the target loss value.
[0185] Optionally, the vector encoding module includes an embedding layer and n stack layers;
[0186] The encoded vector generation sub-module includes:
[0187] A first input sub-unit, configured to input the query question and the sample text as an input set into the embedding layer to obtain a corresponding input vector;
[0188] A second input subunit, configured to input the input vector into the first stack layer to obtain an output vector of the first stack layer;
[0189] A third input subunit, configured to input the output vector of the i-th stack layer into the (i + 1)-th stack layer to obtain an output vector of the (i + 1)-th stack layer, where i ∈ [1, n], and i starts from 1;
[0190] A judgment subunit, configured to judge whether i is equal to n - 1. If so, run the output sub-module; if not, run the third input subunit;
[0191] The output subunit is configured to output the output vector of the n-th stack layer as the encoding vector of each word unit in the input set.
[0192] It should be noted that each component in the apparatus claim should be understood as a functional module that must be established to implement each step of the program flow or each step of the method. Each functional module is not an actual functional division or separation limitation. The apparatus claim defined by such a set of functional modules should be understood as mainly implementing the functional module framework of the solution through the computer program recorded in the specification, rather than mainly implementing the physical apparatus of the solution through hardware means.
[0193] In an embodiment of the present application, a computing device is further provided, including a memory, a processor, and computer instructions stored on the memory and executable on the processor. When the processor executes the instructions, the steps of the training method of the answer extraction model are implemented.
[0194] An embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions. When the instructions are executed by a processor, the steps of the training method of the answer extraction model as described above are implemented.
[0195] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above training method of the answer extraction model belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above training method of the answer extraction model.
[0196] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0197] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, external hard drives, magnetic disks, optical discs, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0198] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present application is not limited by the described action sequence, because according to the present application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0199] In the above embodiments, the descriptions of the respective embodiments each have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0200] The preferred embodiments of the present application disclosed above are only used to help explain the present application. The optional embodiments do not elaborate on all the details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the present application. The present application selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present application, so that those skilled in the art can well understand and utilize the present application. The present application is only limited by the claims and their full scope and equivalents.
Claims
1. A training method for an answer extraction model, characterized in that, Including: Determine sample text from the original corpus, and screen at least one query question associated with the sample text and the corresponding answer label in a pre-constructed question set, where the question set is constructed in the following manner: Extract the event type label and answer type label of the text from the original corpus; Integrate the event type label and the answer type label to generate a question label; For each question label, construct a question describing the question label to generate a question set; Input any one of the query questions and the sample text into a pre-trained answer extraction model to determine the answer extraction result of the query question; Generate the target loss value of the answer extraction model based on the answer extraction result and the answer label, and optimize the answer extraction model based on the target loss value to obtain a target answer extraction model.
2. The training method of the answer extraction model according to claim 1, characterized in that The constructing a question describing the question label for each question label to generate a question set includes: Generate a query question matching the question label according to the category to which the answer type label included in the question label belongs, and construct a question set based on the query question.
3. The training method of the answer extraction model according to claim 2, wherein The generating a query question matching the question label according to the category to which the answer type label included in the question label belongs includes: If the answer type label included in the question label is the first category, obtain a predefined question template, and construct a query question matching the question label based on the question label and the question template; If the answer type label included in the question label is the second category, perform statistical analysis on the event sentences in the original corpus related to the answer type label of the second category, and construct a query question matching the question label according to the analysis result.
4. The training method of the answer extraction model according to claim 1, wherein, The inputting any one of the query questions and the sample text into a pre-trained answer extraction model to determine the answer extraction result of the query question includes: Input any one of the query questions and the sample text as an input set into the answer extraction model, and the vector encoding module of the answer extraction model sums the word vectors, text vectors, and position vectors corresponding to each word unit in the input set to generate an encoding vector corresponding to each word unit; Calculate the probability distribution of the start position and end position of each word unit as the predicted answer corresponding to the query question based on the encoding vector; Determine the answer extraction result corresponding to the query question according to the probability distribution of the start position and end position.
5. The training method of the answer extraction model according to claim 4, wherein, The determining the answer extraction result corresponding to the query question according to the probability distribution of the start position and end position includes: Take the position of the word unit with the highest probability in the probability distribution of the start position in the sample text as the start position of the answer; Take the position of the word unit with the highest probability in the probability distribution of the end position in the sample text as the end position of the answer; and Take the word units between the start position and the end position as the answer extraction result.
6. The training method of the answer extraction model according to claim 5, characterized in that Generating a target loss value of the answer extraction model based on the answer extraction result and the answer label includes: Determining a start position loss of the answer extraction result at the start position in the sample text based on the probability distribution of the start position and the probability of the target start position in the answer label; Determining an end position loss of the answer extraction result at the end position in the sample text based on the probability distribution of the end position and the probability of the target end position in the answer label; Determining a length loss of the answer extraction result based on the start position and the end position; Calculating the target loss value based on the start position loss, the end position loss, and the length loss.
7. The training method of the answer extraction model according to claim 6, wherein, The calculating the target loss value based on the start position loss, the end position loss, and the length loss includes: Calculating a weighted sum of the start position loss, the end position loss, and the length loss as the target loss value.
8. The training method of the answer extraction model according to claim 4, wherein The vector encoding module includes an embedding layer and n stack layers; Correspondingly, generating the encoding vector corresponding to each word unit includes: S11. Inputting the query question and the sample text as an input set into the embedding layer to obtain a corresponding input vector; S12. Inputting the input vector into the first stack layer to obtain an output vector of the first stack layer; S13. Inputting the output vector of the i-th stack layer into the (i + 1)-th stack layer to obtain an output vector of the (i + 1)-th stack layer, where i ∈ [1, n] and i starts from 1; S14. Judging whether i is equal to n - 1. If so, execute step S15; if not, execute step S13; S15. Outputting the output vector of the n-th stack layer as the encoding vector of each word unit in the input set.
9. A training device for an answer extraction model, characterized in that, Including: A screening module configured to determine a sample text from the original corpus and screen at least one query question associated with the sample text and a corresponding answer label in a pre-constructed question set, where the question set is constructed in the following manner: extracting an event type label and an answer type label of the text from the original corpus; integrating the event type label and the answer type label to generate a question label; constructing a question describing the question label for each question label to generate a question set; A determining module configured to input any one of the query questions and the sample text into a pre-trained answer extraction model to determine an answer extraction result of the query question; A calculating module configured to generate a target loss value of the answer extraction model based on the answer extraction result and the answer label, and optimize the answer extraction model based on the target loss value to obtain a target answer extraction model.
10. A computing device, comprising a memory, a processor, and computer instructions stored on the memory and executable on the processor, characterized in that, When the processor executes the instructions, the steps of the method according to any one of claims 1-8 are implemented.
11. A computer-readable storage medium storing computer instructions, characterized in that, When the instructions are executed by the processor, the steps of the method according to any one of claims 1-8 are implemented.
Citation Information
Patent Citations
A reading understanding model training method and device
CN109816111A
Answer acquisition method and device
CN109977428A
Text analysis model training method and device and text analysis method and device
CN110781663A
Method and device for outputting information
CN111382228A