Accuracy Loss Determination Method and Device

By determining the accuracy loss and optimizing the loss function in machine reading comprehension model training, the problem of insufficient loss in the prior art is solved, and the prediction accuracy of the model is improved.

CN114254750BActive Publication Date: 2025-05-30BEIJING KINGSOFT DIGITAL ENTERTAINMENT CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111632110.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-01-29
Publication Date
2025-05-30
Estimated Expiration
2039-01-29

AI Technical Summary

Technical Problem

The losses considered during the training process of existing machine reading comprehension models are not sufficient to fully reflect the losses of predicting the answers, resulting in a low accuracy rate of predicting the answers.

Method used

A method for determining accuracy loss is proposed. By obtaining training samples, generating predicted answers, and determining the accuracy loss of predicted answers relative to the target answer, determining the loss function based on this loss, and optimizing the reading comprehension model.

Benefits of technology

By fully reflecting the loss of predicted answers, the training efficiency of the reading comprehension model is improved, making the prediction accuracy of the trained model higher.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114254750B_ABST
    Figure CN114254750B_ABST
Patent Text Reader

Abstract

The present application provides a method and an apparatus for determining accuracy loss. Among them, the method for determining accuracy loss includes: determining the position loss of the predicted starting position and the predicted ending position of the predicted answer in the sample article; comparing the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer; comparing the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer; and determining the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss. The method for determining accuracy loss provided by the present application not only improves the accuracy of the accuracy loss, but also fully reflects the loss of the predicted answer, thereby guiding the training process of the reading comprehension model based on the accuracy loss, improving the training efficiency of the reading comprehension model, and making the prediction accuracy of the trained reading comprehension model higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and particularly to a method for determining accuracy loss. This application also relates to an apparatus for determining accuracy loss, a computing device, and a computer-readable storage medium. Background Art

[0002] Natural language processing is the study of various theories and methods for realizing effective communication between humans and computers using natural language. With the rapid development of natural language processing, machine reading comprehension, a popular direction in the field of natural language processing, has also received extensive attention. Machine reading comprehension is a study dedicated to teaching machines to read human language and understand its meaning. Machine reading comprehension tasks focus more on the understanding of passage texts. Machines must learn relevant information from the passage by themselves, rather than using pre-set world knowledge and common sense to answer questions, so it is more challenging.

[0003] Currently, an important implementation method for training machines to understand and read human language is to establish a machine reading comprehension model, and then further train the established machine reading comprehension model to obtain the desired machine reading comprehension model, so as to find the answer to the question in the text segment based on the trained machine reading comprehension model. However, the losses considered in the current training process of machine reading comprehension models are not sufficient, and cannot fully reflect the loss of the predicted answer. Eventually, the accuracy of the predicted answer is relatively low. Summary of the Invention

[0004] In view of this, the embodiments of this application provide a method for determining accuracy loss and a method for training a reading comprehension model to solve the technical defects existing in the prior art. The embodiments of this application also provide an apparatus for determining accuracy loss, an apparatus for training a reading comprehension model, a computing device, and a computer-readable storage medium.

[0005] This application provides a method for training a reading comprehension model, including:

[0006] Obtaining a training sample including a sample question and its corresponding target answer in a sample article;

[0007] Generating a predicted answer to the sample question by inputting the training sample into a reading comprehension model;

[0008] Determining the accuracy loss of the predicted answer relative to the target answer;

[0009] Determining a loss function based on the accuracy loss, and optimizing the reading comprehension model using the loss function.

[0010] Optionally, determining the accuracy loss of the predicted answer relative to the target answer includes:

[0011] Determining the start position loss of the predicted start position of the predicted answer in the sample article, and the end position loss of the predicted end position of the predicted answer in the sample article;

[0012] Comparing the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer;

[0013] Based on the start position loss, the end position loss, and the length loss, determining the accuracy loss of the predicted answer.

[0014] Optionally, determining the start position loss of the predicted start position of the predicted answer in the sample article, and the end position loss of the predicted end position of the predicted answer in the sample article includes:

[0015] Calculating the start probability distribution of the word unit included in the sample article being the start word of the predicted answer, and the end probability distribution of the word unit being the end word of the predicted answer;

[0016] Based on the start probability distribution and the end probability distribution, determining the predicted start position and the predicted end position of the predicted answer in the sample article;

[0017] Based on the probability value corresponding to the predicted start position included in the start probability distribution, determining the start position loss of the predicted start position, and based on the probability value corresponding to the predicted end position included in the end probability distribution, determining the end position loss of the predicted end position.

[0018] Optionally, the predicted start position includes: the position of the word unit with the largest probability value included in the start probability distribution in the sample article;

[0019] The predicted end position includes: the position of the word unit with the largest probability value included in the end probability distribution in the sample article.

[0020] Optionally, the start position loss includes: the difference between the probability value corresponding to the predicted start position and the probability value corresponding to the start position of the target answer;

[0021] The end position loss includes: the difference between the probability value corresponding to the predicted end position and the probability value corresponding to the end position of the target answer.

[0022] Optionally, the comparison of the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer includes:

[0023] Determine the article matrix corresponding to the sample article; the word units in the sample article correspond one-to-one with the elements in the article matrix;

[0024] Determine the predicted start element and the predicted end element corresponding to the predicted start position and the predicted end position of the predicted answer in the article matrix, and the target start element and the target end element corresponding to the start position and the end position of the target answer in the article matrix;

[0025] Determine the predicted answer vector from the predicted start element to the predicted end element, and the target answer vector from the target start element to the target end element;

[0026] Calculate the distance between the predicted answer vector and the target answer vector as the length loss of the predicted answer.

[0027] Optionally, the comparison of the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer includes:

[0028] Determine the predicted start position and the predicted end position of the predicted answer in the sample article;

[0029] Calculate the byte length from the predicted start position to the predicted end position as the byte length of the predicted answer;

[0030] Determine the byte length difference between the byte length of the predicted answer and the byte length of the target answer as the length loss of the predicted answer.

[0031] Optionally, the determination of the accuracy loss of the predicted answer based on the start position loss, the end position loss, and the length loss includes:

[0032] Calculate the weighted sum of the start position loss, the end position loss, and the length loss as the accuracy loss of the predicted answer.

[0033] Optionally, the determination of the accuracy loss of the predicted answer relative to the target answer includes:

[0034] Determine the position loss of the predicted start position and the predicted end position of the predicted answer in the sample article;

[0035] Compare the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer;

[0036] Compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer;

[0037] Determine the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss.

[0038] Optionally, comparing the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer includes:

[0039] Calculate the semantic similarity between each word unit included in the predicted answer and the corresponding word unit in the target answer;

[0040] Calculate the semantic loss between each word unit included in the predicted answer and the corresponding word unit in the target answer based on the semantic similarity and sum them to obtain the semantic loss of the predicted answer.

[0041] This application provides a method for determining accuracy loss, including:

[0042] Determine the position loss of the predicted start position and the predicted end position of the predicted answer in the sample article;

[0043] Compare the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer;

[0044] Compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer;

[0045] Determine the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss.

[0046] Optionally, before determining the position loss of the predicted start position and the predicted end position of the predicted answer in the sample article, it further includes:

[0047] Obtain a training sample including a sample question and its corresponding target answer in the sample article;

[0048] Generate a predicted answer to the sample question by inputting the training sample into a reading comprehension model.

[0049] Optionally, after determining the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss, it further includes:

[0050] Determine a loss function based on the accuracy loss, and use the loss function to optimize the reading comprehension model.

[0051] Optionally, the reading comprehension model is any one of Attentive Reader, Attention Sum Reader, Stanford Attentive Reader, and Gated Attention Reader.

[0052] Optionally, the comparison of the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer includes:

[0053] Calculating the semantic similarity between each word unit included in the predicted answer and the corresponding word unit in the target answer;

[0054] Based on the semantic similarity, calculating the semantic loss between each word unit included in the predicted answer and the corresponding word unit in the target answer and summing them to obtain the semantic loss of the predicted answer.

[0055] Optionally, the determination of the position loss of the predicted start position and the predicted end position of the predicted answer in the sample article includes:

[0056] Determining the start position loss of the predicted start position of the predicted answer in the sample article and the end position loss of the predicted end position of the predicted answer in the sample article;

[0057] Taking the sum of the start position loss and the end position loss as the position loss.

[0058] Optionally, the determination of the start position loss of the predicted start position of the predicted answer in the sample article and the end position loss of the predicted end position of the predicted answer in the sample article includes:

[0059] Calculating the start probability distribution of the word unit included in the sample article being the start word of the predicted answer and the end probability distribution of the word unit being the end word of the predicted answer;

[0060] Based on the start probability distribution and the end probability distribution, determining the predicted start position and the predicted end position of the predicted answer in the sample article;

[0061] Based on the probability value corresponding to the predicted start position included in the start probability distribution, determining the start position loss of the predicted start position, and based on the probability value corresponding to the predicted end position included in the end probability distribution, determining the end position loss of the predicted end position.

[0062] Optionally, the predicted start position includes: the position of the word unit with the largest probability value included in the start probability distribution in the sample article;

[0063] The predicted ending position includes: the position of the word unit with the largest probability value included in the ending probability distribution in the sample article.

[0064] Optionally, the starting position loss includes: the difference between the probability value corresponding to the predicted starting position and the probability value corresponding to the starting position of the target answer;

[0065] The ending position loss includes: the difference between the probability value corresponding to the predicted ending position and the probability value corresponding to the ending position of the target answer.

[0066] Optionally, calculating the starting probability distribution of the word unit included in the sample article as the starting word of the predicted answer and the ending probability distribution of the word unit as the ending word of the predicted answer includes:

[0067] Inputting the sample article and the sample question into a pre-configured classifier, and calculating in the classifier the starting probability distribution of the word unit included in the sample article as the starting word of the predicted answer and the ending probability distribution of the word unit as the ending word of the predicted answer;

[0068] After the calculation, the classifier outputs the starting probability distribution and the ending probability distribution.

[0069] Optionally, comparing the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer includes:

[0070] Determining the article matrix corresponding to the sample article; the word units in the sample article correspond one by one to the elements in the article matrix;

[0071] Determining the predicted starting element and the predicted ending element corresponding to the predicted starting position and the predicted ending position of the predicted answer in the article matrix, and the target starting element and the target ending element corresponding to the starting position and the ending position of the target answer in the article matrix;

[0072] Determining the predicted answer vector from the predicted starting element to the predicted ending element and the target answer vector from the target starting element to the target ending element;

[0073] Calculating the distance between the predicted answer vector and the target answer vector as the length loss of the predicted answer.

[0074] Optionally, comparing the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer includes:

[0075] Determine the predicted start position and the predicted end position of the predicted answer in the sample article;

[0076] Calculate the byte length from the predicted start position to the predicted end position as the byte length of the predicted answer;

[0077] Determine the byte length difference between the byte length of the predicted answer and the byte length of the target answer as the length loss of the predicted answer.

[0078] The present application provides a reading comprehension model training device, including:

[0079] A training sample acquisition module configured to acquire a training sample including a sample question and its corresponding target answer in a sample article;

[0080] A predicted answer generation module configured to generate a predicted answer to the sample question by inputting the training sample into a reading comprehension model;

[0081] An accuracy loss determination module configured to determine the accuracy loss of the predicted answer relative to the target answer;

[0082] A model optimization module configured to determine a loss function based on the accuracy loss and optimize the reading comprehension model using the loss function.

[0083] Optionally, the accuracy loss determination module includes:

[0084] A position loss determination sub-module configured to determine the start position loss of the predicted start position of the predicted answer in the sample article and the end position loss of the predicted end position of the predicted answer in the sample article;

[0085] A length loss determination sub-module configured to compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer;

[0086] An accuracy loss determination sub-module configured to determine the accuracy loss of the predicted answer based on the start position loss, the end position loss, and the length loss.

[0087] The present application provides an accuracy loss determination device, including:

[0088] A second position loss determination sub-module configured to determine the position loss of the predicted start position and the predicted end position of the predicted answer in the sample article;

[0089] A semantic loss determination sub-module, configured to compare the word units included in the predicted answer with the word units included in the target answer, and determine the semantic loss of the predicted answer;

[0090] A second length loss determination sub-module, configured to compare the predicted answer with the target answer in the sample article, and determine the length loss of the predicted answer;

[0091] A second accuracy loss determination sub-module, configured to determine the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss.

[0092] This application provides a computing device, including:

[0093] A memory and a processor;

[0094] The memory is used to store computer-executable instructions, and when the processor executes the computer-executable instructions, the steps of the reading comprehension model training method or the accuracy loss determination method are implemented.

[0095] This application provides a computer-readable storage medium, which stores computer instructions, and when the instructions are executed by a processor, the steps of the reading comprehension model training method or the accuracy loss determination method are implemented.

[0096] Compared with the prior art, this application has the following advantages:

[0097] This application provides a reading comprehension model training method, including: obtaining a training sample including a sample question and its corresponding target answer in a sample article; generating a predicted answer to the sample question by inputting the training sample into a reading comprehension model; determining the accuracy loss of the predicted answer relative to the target answer; determining a loss function based on the accuracy loss, and optimizing the reading comprehension model by using the loss function.

[0098] In the training process of the reading comprehension model provided by this application, by inputting a training sample into the reading comprehension model to generate a predicted answer of the reading comprehension model to the sample question, and comparing the predicted answer of the sample question with the actual target answer to determine the loss of the predicted answer relative to the actual target answer, thereby guiding the training process of the reading comprehension model on the basis of determining the loss, improving the training efficiency of the reading comprehension model, and making the prediction accuracy of the trained reading comprehension model higher.

[0099] The present application provides a method for determining accuracy loss, including: determining the position loss of the predicted starting position and the predicted ending position of the predicted answer in the sample article; comparing the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer; comparing the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer; and determining the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss.

[0100] In summary, in the method for determining accuracy loss provided by the present application, during the process of determining the accuracy loss, by determining the position loss of the predicted starting position and the predicted ending position of the predicted answer in the sample article, and comparing the predicted answer of the sample question with the actual target answer to determine the semantic loss and the length loss of the predicted answer relative to the actual target answer, thereby determining the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss, that is, determining the accuracy loss of the predicted answer from three aspects of position, semantics, and length, not only improves the accuracy of the accuracy loss, but also fully reflects the loss of the predicted answer, and thus guides the training process of the reading comprehension model based on this accuracy loss, can improve the training efficiency of the reading comprehension model, and make the prediction accuracy of the trained reading comprehension model higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0101] Figure 1 is a processing flowchart of a method for training a reading comprehension model provided by an embodiment of the present application;

[0102] Figure 2 is a schematic structural diagram of a device for training a reading comprehension model provided by an embodiment of the present application;

[0103] Figure 3 is a block diagram of the structure of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0104] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0105] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0106] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0107] This application provides a method for training a reading comprehension model, and this application also provides a device for training a reading comprehension model, a computing device, and a computer-readable storage medium. The following will be described in detail one by one in combination with the accompanying drawings of the embodiments provided in this application, and each step of the method will be described.

[0108] An embodiment of the method for training a reading comprehension model provided in this application is as follows:

[0109] Referring to the attached Figure 1 , which shows a processing flow chart of a method for training a reading comprehension model provided in this embodiment.

[0110] Step S102, obtain a training sample including a sample question and its corresponding target answer in a sample article.

[0111] The life cycle of the model mainly includes three main stages: the construction stage, the training stage, and the application stage; the method for training a reading comprehension model provided in this application is to train the already constructed reading comprehension model during the model construction stage so that the trained reading comprehension model can predict more accurate answers when applied.

[0112] In addition, the reading comprehension model training method provided by this application can also train the reading comprehension model during its application process. For example, every time a question and an article are input into the reading comprehension model to predict the answer to the question in the article, by using the predicted question, article, and the predicted answer to the question in the article this time as training samples to optimize the reading comprehension model, it can not only make the prediction accuracy during the application process of the reading comprehension model higher, but also make the optimization and adjustment of the reading comprehension model closer to the actual business of applying this reading comprehension model.

[0113] It should be noted that the reading comprehension model described in the embodiments of this application refers to a machine reading comprehension model. There are many specific models in the field of machine reading comprehension research. For example, common machine reading comprehension models include: AttentiveReader, Attention Sum Reader (AS Reader), Stanford Attentive Reader (Stanford AR), and Gated Attention Reader (GA Reader), etc.

[0114] In the embodiments of this application, a training sample consists of three parts: an article, a question, and the true answer to the question in the article. For the convenience of description, the article will be hereinafter referred to as the sample article, the question as the sample question, the true answer to the sample question in the sample article as the target answer, and the answer obtained by inputting the sample question and the sample article into the reading comprehension model for prediction as the predicted answer.

[0115] Step S104, generating a predicted answer to the sample question by inputting the training sample into the reading comprehension model.

[0116] In specific implementation, in order to evaluate the gap between the predicted answer obtained by the reading comprehension model and the target answer, it is necessary to input the training sample into the reading comprehension model to obtain the predicted answer predicted by this reading comprehension model. Specifically, the sample article and the sample question included in the training sample are input into the reading comprehension model, and the reading comprehension model performs prediction calculations on the sample question in the sample article, and finally outputs the predicted answer to the sample question in the sample article predicted by it.

[0117] Step S106, determining the accuracy loss of the predicted answer relative to the target answer.

[0118] In a preferred implementation manner provided by the embodiments of this application, to determine the accuracy loss of the predicted answer relative to the target answer, the following specific method is adopted:

[0119] 1) Determine the start position loss of the predicted answer at the predicted start position in the sample article, and the end position loss of the predicted answer at the predicted end position in the sample article;

[0120] In the embodiments of the present application, for the predicted start position of the predicted answer in the sample article, preferably, calculate the start probability distribution of the word units included in the sample article being the start word of the predicted answer, and determine the predicted start position of the predicted answer in the sample article based on the start probability distribution.

[0121] Preferably, the predicted start position of the predicted answer in the sample article refers to the position of the word unit with the largest probability value included in the start probability distribution in the sample article.

[0122] It can be seen that by calculating the probability of each word unit in the sample article being the start word of the predicted answer, and taking the word unit with the largest probability in the sample article as the start word of the predicted answer, the prediction accuracy of the start word of the predicted answer is improved.

[0123] Similar to the above-provided method for predicting the start position of the predicted answer in the sample article, for the predicted end position of the predicted answer in the sample article, it is also calculated by calculating the end probability distribution of the word units included in the sample article being the end word of the predicted answer, and determining the predicted end position of the predicted answer in the sample article based on the end probability distribution.

[0124] Preferably, the predicted end position of the predicted answer in the sample article refers to the position of the word unit with the largest probability value included in the end probability distribution in the sample article.

[0125] It can be seen that by calculating the probability of each word unit in the sample article being the end word of the predicted answer, and taking the word unit with the largest probability in the sample article as the end word of the predicted answer, the prediction accuracy of the end word of the predicted answer can also be improved.

[0126] Specifically, when the reading comprehension model predicts the corresponding answer of the predicted answer in the sample article, if the reading comprehension model needs to calculate the probability of each word unit in the sample article being the start word of the predicted answer and calculate the probability of each word unit in the sample article being the end word of the predicted answer during the prediction process, then the start probability distribution of the word units included in the sample article being the start word of the predicted answer and the end probability distribution of the word units included in the sample article being the end word of the predicted answer can be read in the reading comprehension model.

[0127] In addition, the starting probability distribution of the starting word and the ending probability distribution can also be obtained by inputting the sample article and the sample question into a pre-configured classifier, and calculating the probabilities that the word units included in the sample article are the starting word or the ending word of the predicted answer in the classifier. After the calculation is completed, the classifier outputs the starting probability distribution and the ending probability distribution.

[0128] On the basis of determining the predicted starting position and the predicted ending position of the predicted answer in the sample article as described above, further calculate the loss of the predicted starting position compared to the starting position of the target answer in the sample article. This loss is the starting position loss of the predicted starting position. Preferably, the starting position loss of the predicted starting position refers to the difference between the probability value corresponding to the predicted starting position and the probability value corresponding to the starting position of the target answer.

[0129] For example, if the sample article consists of 100 word units, calculate the probability that each word unit is the starting word of the predicted answer, and then determine the position of the word unit with the highest probability (probability value of 85%) in the sample article as the predicted starting position of the predicted answer. If this predicted starting position is also the starting position of the target answer in the sample article, then the probability value corresponding to the starting position of the target answer in the sample article is 1. The loss of this predicted starting position is equal to the probability value 1 corresponding to the starting position in the sample article minus the probability value 85% corresponding to the predicted starting position. Finally, the starting position loss of this predicted starting position is 1 - 85% = 0.15.

[0130] Similar to the determination process of the starting position loss of the predicted starting position, on the basis of determining the predicted ending position and the predicted ending position of the predicted answer in the sample article as described above, further calculate the loss of the predicted ending position compared to the ending position of the target answer in the sample article. This loss is the ending position loss of the predicted ending position. The ending position loss of the predicted ending position refers to the difference between the probability value corresponding to the predicted ending position and the probability value corresponding to the ending position of the target answer.

[0131] For example, if the sample article consists of 100 word units, calculate the probability that each word unit is the ending word of the predicted answer, and then determine the position of the word unit with the highest probability (probability value of 70%) in the sample article as the predicted ending position of the predicted answer. If this predicted ending position is also the ending position of the target answer in the sample article, then the probability value corresponding to the ending position of the target answer in the sample article is 1. The loss of this predicted ending position is equal to the probability value 1 corresponding to the ending position in the sample article minus the probability value 70% corresponding to the predicted ending position. Finally, the ending position loss of this predicted ending position is 1 - 70% = 0.3.

[0132] 2) Compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer.

[0133] In a preferred embodiment provided by the embodiments of the present application, the length loss of the predicted answer is specifically determined in the following manner:

[0134] (a) Determine the article matrix corresponding to the sample article;

[0135] There is a one-to-one correspondence between the word units in the sample article and the elements in the article matrix, and each word unit corresponds to an element in the article matrix;

[0136] (b) Determine the predicted start element and the predicted end element corresponding to the predicted start position and the predicted end position of the predicted answer in the article matrix, and the target start element and the target end element corresponding to the start position and the end position of the target answer in the article matrix;

[0137] (c) Determine the predicted answer vector from the predicted start element to the predicted end element, and the target answer vector from the target start element to the target end element;

[0138] (d) Calculate the distance between the predicted answer vector and the target answer vector as the length loss of the predicted answer.

[0139] For example, the sample article consists of 100 word units, specifically shown as 5 rows in the sample article, with 20 words in each row. By establishing a mapping relationship between the rows in the sample article and the rows in the matrix, and a mapping relationship between the columns in the sample article and the columns in the matrix, a corresponding matrix is constructed for the sample article, and each element in the matrix corresponds to a word unit in the sample article;

[0140] Then determine the predicted start element and the predicted end element corresponding to the predicted start position and the predicted end position in the matrix, and further determine the predicted answer vector from the predicted start element to the predicted end element; similarly, determine the target start element and the target end element corresponding to the start position and the end position of the target answer in the matrix, and further determine the target answer vector from the target start element to the target end element;

[0141] Finally, calculate the Euclidean distance between the predicted answer vector and the target answer vector as the length loss of the predicted answer relative to the target answer.

[0142] In addition to the method for determining the length loss of the predicted answer provided above, the length loss of the predicted answer can also be determined in other ways. For example, it is preferably determined in the following way: First, determine the predicted start position and the predicted end position of the predicted answer in the sample article; then, calculate the byte length from the predicted start position to the predicted end position as the byte length of the predicted answer; finally, determine the byte length difference between the byte length of the predicted answer and the byte length of the target answer as the length loss of the predicted answer.

[0143] 3) Based on the start position loss, the end position loss, and the length loss, determine the accuracy loss of the predicted answer.

[0144] The accuracy loss of the predicted answer is preferably determined by calculating the weighted sum of the start position loss, the end position loss, and the length loss.

[0145] For example, the accuracy loss Loss of the predicted answer is:

[0146] Loss = Loss_start + Loss_end + Loss_length

[0147] where Loss_start is the start position loss, Loss_end is the end position loss, and Loss_length is the length loss.

[0148] Step S108, determine a loss function based on the accuracy loss, and use the loss function to optimize the reading comprehension model.

[0149] According to the determined accuracy loss of the predicted answer relative to the target answer, determine the loss function (evaluation function) for training the reading comprehension model, and then use the loss function to optimize the reading comprehension model. For example, adjust the parameters or weight coefficients of the reading comprehension model. Finally, after the reading comprehension model is trained, the obtained reading comprehension model has a higher prediction accuracy for the predicted answer.

[0150] In the process of determining the accuracy loss of the predicted answer relative to the target answer in this embodiment, it is preferably to determine the final accuracy loss of the predicted answer relative to the target answer according to the start position loss, the end position loss, and the length loss. In addition, other losses related to accuracy can also be used in the process of determining the accuracy loss. For example, the following provides the determination of the accuracy using position loss, semantic loss, and length loss:

[0151] 1) Determine the positional loss of the predicted starting position and the predicted ending position of the predicted answer in the sample article;

[0152] wherein, the positional loss is equal to the sum of the starting position loss of the predicted starting position and the ending position loss of the predicted ending position;

[0153] 2) Compare the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer;

[0154] Specifically, the semantic loss is preferably obtained by calculating the semantic similarity between each word unit included in the predicted answer and the corresponding word unit in the target answer, and calculating and summing the semantic losses between each word unit included in the predicted answer and the corresponding word unit in the target answer based on the semantic similarity, and finally obtaining the semantic loss of the predicted answer;

[0155] 3) Compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer;

[0156] 4) Determine the accuracy loss of the predicted answer based on the positional loss, the semantic loss, and the length loss.

[0157] Based on the above, the positional loss, the semantic loss, and the length loss are used to determine the accuracy loss, and further determine the loss function (evaluation function) for training the reading comprehension model, and then use the loss function to optimize the reading comprehension model, so as to obtain a reading comprehension model with higher prediction accuracy.

[0158] In summary, the reading comprehension model training method provided by the present application, in the process of training the reading comprehension model, inputs the training samples into the reading comprehension model to generate the predicted answer of the reading comprehension model to the sample questions, and compares the predicted answer of the sample questions with the actual target answer to determine the loss of the predicted answer relative to the actual target answer, so as to guide the training process of the reading comprehension model based on the determined loss, improve the training efficiency of the reading comprehension model, and make the prediction accuracy of the trained reading comprehension model higher.

[0159] An embodiment of a reading comprehension model training device provided by the present application is as follows:

[0160] In the above embodiment, a reading comprehension model training method is provided. Correspondingly, the present application also provides a reading comprehension model training device, which will be described below with reference to the accompanying drawings.

[0161] Refer to the appendix Figure 2 , which shows a schematic structural diagram of a reading comprehension model training device provided by an embodiment of the present application.

[0162] Since the device embodiments are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The device embodiments described below are merely illustrative.

[0163] This application provides a reading comprehension model training device, including:

[0164] A training sample acquisition module 202, configured to acquire training samples including sample questions and their corresponding target answers in sample articles;

[0165] A predicted answer generation module 204, configured to generate a predicted answer to the sample question by inputting the training sample into the reading comprehension model;

[0166] An accuracy loss determination module 206, configured to determine the accuracy loss of the predicted answer relative to the target answer;

[0167] A model optimization module 208, configured to determine a loss function based on the accuracy loss and optimize the reading comprehension model using the loss function.

[0168] Optionally, the accuracy loss determination module 206 includes:

[0169] A position loss determination sub-module, configured to determine a start position loss of the predicted start position of the predicted answer in the sample article and an end position loss of the predicted end position of the predicted answer in the sample article;

[0170] A length loss determination sub-module, configured to compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer;

[0171] An accuracy loss determination sub-module, configured to determine the accuracy loss of the predicted answer based on the start position loss, the end position loss, and the length loss.

[0172] Optionally, the position loss determination sub-module includes:

[0173] A probability distribution calculation sub-unit, configured to calculate a start probability distribution that a word unit included in the sample article is the start word of the predicted answer and an end probability distribution that the word unit is the end word of the predicted answer;

[0174] A position determination sub-unit, configured to determine the predicted start position and the predicted end position of the predicted answer in the sample article based on the start probability distribution and the end probability distribution;

[0175] A loss determination subunit, configured to determine a start position loss of the predicted start position based on a probability value corresponding to the predicted start position included in the start probability distribution, and determine an end position loss of the predicted end position based on a probability value corresponding to the predicted end position included in the end probability distribution.

[0176] Optionally, the predicted start position includes: the position of the word unit with the largest probability value included in the start probability distribution in the sample article;

[0177] The predicted end position includes: the position of the word unit with the largest probability value included in the end probability distribution in the sample article.

[0178] Optionally, the start position loss includes: the difference between the probability value corresponding to the predicted start position and the probability value corresponding to the start position of the target answer;

[0179] The end position loss includes: the difference between the probability value corresponding to the predicted end position and the probability value corresponding to the end position of the target answer.

[0180] Optionally, the length loss determination sub-module includes:

[0181] A matrix determination subunit, configured to determine an article matrix corresponding to the sample article; the word units in the sample article correspond one-to-one with the elements in the article matrix;

[0182] An element determination subunit, configured to determine a predicted start element and a predicted end element corresponding to the predicted start position and the predicted end position of the predicted answer in the article matrix, and a target start element and a target end element corresponding to the start position and the end position of the target answer in the article matrix;

[0183] A vector determination subunit, configured to determine a predicted answer vector from the predicted start element to the predicted end element, and a target answer vector from the target start element to the target end element;

[0184] A first length loss determination subunit, configured to calculate the distance between the predicted answer vector and the target answer vector as the length loss of the predicted answer.

[0185] Optionally, the length loss determination sub-module includes:

[0186] A predicted position determination subunit, configured to determine a predicted start position and a predicted end position of the predicted answer in the sample article;

[0187] A byte length determination subunit, configured to calculate the byte length from the predicted start position to the predicted end position as the byte length of the predicted answer;

[0188] A second length loss determination subunit, configured to determine the byte length difference between the byte length of the predicted answer and the byte length of the target answer as the length loss of the predicted answer.

[0189] Optionally, the accuracy loss determination sub-module is specifically configured to calculate the weighted sum of the start position loss, the end position loss, and the length loss as the accuracy loss of the predicted answer.

[0190] Optionally, the accuracy loss determination module 206 includes:

[0191] A second position loss determination sub-module, configured to determine the position loss of the predicted start position and the predicted end position of the predicted answer in the sample article;

[0192] A semantic loss determination sub-module, configured to compare the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer;

[0193] A second length loss determination sub-module, configured to compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer;

[0194] A second accuracy loss determination sub-module, configured to determine the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss.

[0195] Optionally, the semantic loss determination sub-module includes:

[0196] A semantic similarity calculation subunit, configured to calculate the semantic similarity between each word unit included in the predicted answer and the corresponding word unit in the target answer;

[0197] A semantic loss determination subunit, configured to calculate and sum the semantic loss between each word unit included in the predicted answer and the corresponding word unit in the target answer based on the semantic similarity to obtain the semantic loss of the predicted answer.

[0198] An embodiment of a computing device provided by the present application is as follows:

[0199] Figure 3FIG. 0 is a structural block diagram showing a computing device 300 according to an embodiment of the present specification. Components of the computing device 300 include, but are not limited to, a memory 310 and a processor 320. The processor 320 is connected to the memory 310 via a bus 330, and a database 350 is used to store data.

[0200] The computing device 300 further includes an access device 340, which enables the computing device 300 to communicate via one or more networks 360. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 340 may include one or more of any type of wired or wireless network interface (e.g., Network Interface Card (NIC)), such as an IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC) interface, and so on.

[0201] In an embodiment of the present specification, the above components of the computing device 300 and Figure 3 other components not shown therein may also be connected to each other, for example, via a bus. It should be understood that Figure 3 the shown structural block diagram of the computing device is only for illustrative purposes and is not a limitation on the scope of the present specification. Those skilled in the art may add or replace other components as needed.

[0202] The computing device 300 may be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 300 may also be a mobile or stationary server.

[0203] The present application provides a computing device, including a memory 310, a processor 320, and computer instructions stored on the memory and executable on the processor. The processor 320 is used to execute the following computer-executable instructions:

[0204] Obtain a training sample including a sample question and its corresponding target answer in a sample article;

[0205] Generate a predicted answer to the sample question by inputting the training sample into a reading comprehension model;

[0206] Determine the accuracy loss of the predicted answer relative to the target answer;

[0207] Determine a loss function based on the accuracy loss, and use the loss function to optimize the reading comprehension model.

[0208] Optionally, the determining the accuracy loss of the predicted answer relative to the target answer includes:

[0209] Determine the start position loss of the predicted start position of the predicted answer in the sample article, and the end position loss of the predicted end position of the predicted answer in the sample article;

[0210] Compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer;

[0211] Based on the start position loss, the end position loss, and the length loss, determine the accuracy loss of the predicted answer.

[0212] Optionally, the determining the start position loss of the predicted start position of the predicted answer in the sample article, and the end position loss of the predicted end position of the predicted answer in the sample article includes:

[0213] Calculate the start probability distribution of the word unit included in the sample article as the start word of the predicted answer, and the end probability distribution of the word unit as the end word of the predicted answer;

[0214] Based on the start probability distribution and the end probability distribution, determine the predicted start position and the predicted end position of the predicted answer in the sample article;

[0215] Based on the probability value corresponding to the predicted start position included in the start probability distribution, determine the start position loss of the predicted start position, and based on the probability value corresponding to the predicted end position included in the end probability distribution, determine the end position loss of the predicted end position.

[0216] Optionally, the predicted start position includes: the position of the word unit with the largest probability value included in the start probability distribution in the sample article;

[0217] The predicted end position includes: the position of the word unit with the largest probability value included in the end probability distribution in the sample article.

[0218] Optionally, the start position loss includes: the difference between the probability value corresponding to the predicted start position and the probability value corresponding to the start position of the target answer;

[0219] The ending position loss includes the difference between the probability value corresponding to the predicted ending position and the probability value corresponding to the ending position of the target answer.

[0220] Optionally, the comparing the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer includes:

[0221] Determine the article matrix corresponding to the sample article; the word units in the sample article correspond one-to-one with the elements in the article matrix;

[0222] Determine the predicted starting element and the predicted ending element corresponding to the predicted starting position and the predicted ending position of the predicted answer in the article matrix, and the target starting element and the target ending element corresponding to the starting position and the ending position of the target answer in the article matrix;

[0223] Determine the predicted answer vector from the predicted starting element to the predicted ending element, and the target answer vector from the target starting element to the target ending element;

[0224] Calculate the distance between the predicted answer vector and the target answer vector as the length loss of the predicted answer.

[0225] Optionally, the comparing the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer includes:

[0226] Determine the predicted starting position and the predicted ending position of the predicted answer in the sample article;

[0227] Calculate the byte length from the predicted starting position to the predicted ending position as the byte length of the predicted answer;

[0228] Determine the byte length difference between the byte length of the predicted answer and the byte length of the target answer as the length loss of the predicted answer.

[0229] Optionally, the determining the accuracy loss of the predicted answer based on the starting position loss, the ending position loss, and the length loss includes:

[0230] Calculate the weighted sum of the starting position loss, the ending position loss, and the length loss as the accuracy loss of the predicted answer.

[0231] Optionally, the determining the accuracy loss of the predicted answer relative to the target answer includes:

[0232] Determine the position loss of the predicted starting position and the predicted ending position of the predicted answer in the sample article;

[0233] Compare the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer;

[0234] Compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer;

[0235] Determine the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss.

[0236] Optionally, the comparing the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer includes:

[0237] Calculate the semantic similarity between each word unit included in the predicted answer and the corresponding word unit in the target answer;

[0238] Calculate the semantic loss between each word unit included in the predicted answer and the corresponding word unit in the target answer based on the semantic similarity and sum them to obtain the semantic loss of the predicted answer.

[0239] An embodiment of a computer-readable storage medium provided by this application is as follows:

[0240] An embodiment of this application further provides a computer-readable storage medium, which stores computer instructions that are executed by a processor for:

[0241] Obtain a training sample including a sample question and its corresponding target answer in a sample article;

[0242] Generate a predicted answer to the sample question by inputting the training sample into a reading comprehension model;

[0243] Determine the accuracy loss of the predicted answer relative to the target answer;

[0244] Determine a loss function based on the accuracy loss, and use the loss function to optimize the reading comprehension model.

[0245] Optionally, the determining the accuracy loss of the predicted answer relative to the target answer includes:

[0246] Determine the start position loss of the predicted start position of the predicted answer in the sample article, and the end position loss of the predicted end position of the predicted answer in the sample article;

[0247] Compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer;

[0248] Determine the accuracy loss of the predicted answer based on the start position loss, the end position loss, and the length loss.

[0249] Optionally, determining the start position loss of the predicted start position of the predicted answer in the sample article, and the end position loss of the predicted end position of the predicted answer in the sample article includes:

[0250] Calculate the start probability distribution where the word unit included in the sample article is the start word of the predicted answer, and the end probability distribution where the word unit is the end word of the predicted answer;

[0251] Determine the predicted start position and the predicted end position of the predicted answer in the sample article based on the start probability distribution and the end probability distribution;

[0252] Determine the start position loss of the predicted start position based on the probability value corresponding to the predicted start position included in the start probability distribution, and determine the end position loss of the predicted end position based on the probability value corresponding to the predicted end position included in the end probability distribution.

[0253] Optionally, the predicted start position includes: the position of the word unit with the largest probability value included in the start probability distribution in the sample article;

[0254] The predicted end position includes: the position of the word unit with the largest probability value included in the end probability distribution in the sample article.

[0255] Optionally, the start position loss includes: the difference between the probability value corresponding to the predicted start position and the probability value corresponding to the start position of the target answer;

[0256] The end position loss includes: the difference between the probability value corresponding to the predicted end position and the probability value corresponding to the end position of the target answer.

[0257] Optionally, comparing the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer includes:

[0258] Determine the article matrix corresponding to the sample article; the word units in the sample article correspond one-to-one with the elements in the article matrix;

[0259] Determine the predicted start element and the predicted end element corresponding to the predicted start position and the predicted end position of the predicted answer in the article matrix, and the target start element and the target end element corresponding to the start position and the end position of the target answer in the article matrix;

[0260] Determine a predicted answer vector from the predicted start element to the predicted end element, and a target answer vector from the target start element to the target end element;

[0261] Calculate the distance between the predicted answer vector and the target answer vector as the length loss of the predicted answer.

[0262] Optionally, comparing the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer includes:

[0263] Determine the predicted start position and the predicted end position of the predicted answer in the sample article;

[0264] Calculate the byte length from the predicted start position to the predicted end position as the byte length of the predicted answer;

[0265] Determine the byte length difference between the byte length of the predicted answer and the byte length of the target answer as the length loss of the predicted answer.

[0266] Optionally, determining the accuracy loss of the predicted answer based on the start position loss, the end position loss, and the length loss includes:

[0267] Calculate the weighted sum of the start position loss, the end position loss, and the length loss as the accuracy loss of the predicted answer.

[0268] Optionally, determining the accuracy loss of the predicted answer relative to the target answer includes:

[0269] Determine the position loss of the predicted start position and the predicted end position of the predicted answer in the sample article;

[0270] Compare the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer;

[0271] Compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer;

[0272] Determine the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss.

[0273] Optionally, comparing the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer includes:

[0274] Calculate the semantic similarity between each word unit included in the predicted answer and the corresponding word unit in the target answer;

[0275] Based on the semantic similarity, calculate the semantic loss between each word unit included in the predicted answer and the corresponding word unit in the target answer and sum them up to obtain the semantic loss of the predicted answer.

[0276] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above reading comprehension model training method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above reading comprehension model training method.

[0277] The computer instructions include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0278] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0279] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0280] The preferred embodiments of the present application disclosed above are only used to help illustrate the present application. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present application, so that those skilled in the art can well understand and utilize the present application. The present application is only limited by the claims and their full scope and equivalents.

Claims

1. A method for determining accuracy loss, characterized in that, it includes: Determine the position loss of the predicted answer at the predicted start position and the predicted end position in the sample article. Among them, the determination of the position loss of the predicted answer at the predicted start position and the predicted end position in the sample article includes: determining the start position loss of the predicted start position of the predicted answer in the sample article, and the end position loss of the predicted end position of the predicted answer in the sample article; determining the sum of the start position loss and the end position loss as the position loss; Compare the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer; Compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer; Determine the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss.

2. The method for determining accuracy loss according to claim 1, characterized in that, before determining the position loss of the predicted answer at the predicted start position and the predicted end position in the sample article, it further includes: Obtain a training sample including a sample question and its corresponding target answer in the sample article; Generate a predicted answer to the sample question by inputting the training sample into a reading comprehension model.

3. The method for determining accuracy loss according to claim 1 or 2, characterized in that, after determining the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss, it further includes: Determine a loss function based on the accuracy loss, and use the loss function to optimize the reading comprehension model.

4. The method for determining accuracy loss according to claim 2, characterized in that, the reading comprehension model is any one of Attentive Reader, Attention Sum Reader, Stanford Attentive Reader, and GatedAttention Reader.

5. The method for determining accuracy loss according to claim 1, characterized in that, the comparison of the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer includes: Calculate the semantic similarity of each word unit included in the predicted answer and the corresponding word unit in the target answer; Calculate the semantic loss of each word unit included in the predicted answer and the corresponding word unit in the target answer based on the semantic similarity and sum them to obtain the semantic loss of the predicted answer.

6. The method for determining accuracy loss according to claim 1, characterized in that, the determination of the start position loss of the predicted start position of the predicted answer in the sample article, and the end position loss of the predicted end position of the predicted answer in the sample article includes: Calculate the start probability distribution of the word unit included in the sample article being the start word of the predicted answer, and the end probability distribution of the word unit being the end word of the predicted answer; Determine the predicted start position and the predicted end position of the predicted answer in the sample article based on the starting probability distribution and the ending probability distribution; Determine the start position loss of the predicted start position based on the probability value corresponding to the predicted start position included in the starting probability distribution, and determine the end position loss of the predicted end position based on the probability value corresponding to the predicted end position included in the ending probability distribution.

7. The method for determining the accuracy loss according to claim 6, wherein, the predicted start position includes: the position of the word unit with the largest probability value included in the starting probability distribution in the sample article; the predicted end position includes: the position of the word unit with the largest probability value included in the ending probability distribution in the sample article.

8. The method for determining the accuracy loss according to claim 7, wherein, the start position loss includes: the difference between the probability value corresponding to the predicted start position and the probability value corresponding to the start position of the target answer; the end position loss includes: the difference between the probability value corresponding to the predicted end position and the probability value corresponding to the end position of the target answer.

9. The method for determining the accuracy loss according to claim 6, wherein, the calculating the starting probability distribution of the word unit in the sample article being the starting word of the predicted answer and the ending probability distribution of the word unit being the ending word of the predicted answer includes: Input the sample article and the sample question into a pre-configured classifier, and calculate in the classifier the starting probability distribution of the word unit in the sample article being the starting word of the predicted answer and the ending probability distribution of the word unit being the ending word of the predicted answer; After the calculation, the classifier outputs the starting probability distribution and the ending probability distribution.

10. The method for determining the accuracy loss according to claim 1, wherein, the comparing the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer includes: Determine the article matrix corresponding to the sample article; the word units in the sample article correspond one by one to the elements in the article matrix; Determine the predicted start element and the predicted end element corresponding to the predicted start position and the predicted end position of the predicted answer in the article matrix, and the target start element and the target end element corresponding to the start position and the end position of the target answer in the article matrix; Determine the predicted answer vector from the predicted start element to the predicted end element and the target answer vector from the target start element to the target end element; Calculate the distance between the predicted answer vector and the target answer vector as the length loss of the predicted answer.

11. The method for determining the accuracy loss according to claim 1, wherein, the comparing the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer includes: Determine the predicted start position and the predicted end position of the predicted answer in the sample article; Calculate the byte length from the predicted start position to the predicted end position as the byte length of the predicted answer; Determine the byte length difference between the byte length of the predicted answer and the byte length of the target answer as the length loss of the predicted answer.

12. An accuracy loss determination device, characterized in that, it includes: A second position loss determination sub-module, configured to determine the position loss of the predicted start position and the predicted end position of the predicted answer in the sample article, wherein the second position loss determination sub-module is further configured to: determine the start position loss of the predicted start position of the predicted answer in the sample article, and the end position loss of the predicted end position of the predicted answer in the sample article; determine the sum of the start position loss and the end position loss as the position loss; A semantic loss determination sub-module, configured to compare the word units included in the predicted answer with the word units included in the target answer to determine the semantic loss of the predicted answer; A second length loss determination sub-module, configured to compare the predicted answer with the target answer in the sample article to determine the length loss of the predicted answer; A second accuracy loss determination sub-module, configured to determine the accuracy loss of the predicted answer based on the position loss, the semantic loss, and the length loss.

13. A computing device, characterized in that, it includes: A memory and a processor; The memory is used to store computer-executable instructions, and when the processor executes the computer-executable instructions, it implements the steps of the accuracy loss determination method according to any one of claims 1 to 11.

14. A computer-readable storage medium, which stores computer instructions, characterized in that, when the instructions are executed by a processor, it implements the steps of the accuracy loss determination method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Neural network training and construction method and device, and object detection method and device

    CN106295678A

  • Recommendation method for technical labels in software question and answer community

    CN107798624A