An English multiple-choice question answer prediction method, device, equipment and storage medium

By judging whether the pending sentence contains dialogue in the English multiple-choice questions, and implementing the corresponding option prediction strategy, the answers to the English multiple-choice questions are directly predicted, and the problem of poor answer matching in the photo search scenario of English multiple-choice questions is solved, achieving high-accurate answer prediction.

CN114741485BActive Publication Date: 2025-06-10GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110025511.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-08
Publication Date
2025-06-10
Estimated Expiration
2041-01-08

AI Technical Summary

Technical Problem

In the scene of taking photos and searching questions, English multiple-choice questions cannot match satisfactory answers due to insufficient coverage of the question bank, and the existing technology is difficult to effectively solve this problem.

Method used

By obtaining the English multiple-choice question stem and multiple options, substituting each option into the question stem, determining whether the pending statement contains dialogue, and executing different option prediction strategies based on this, determining the target option from multiple options as the answer.

Benefits of technology

There is no need to rely on a huge exercise bank, which saves data storage space and improves answer prediction accuracy, solving the problem that the exercise bank cannot get the correct answer because of insufficient coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114741485B_ABST
    Figure CN114741485B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment and storage medium for predicting answers to English multiple-choice questions. After obtaining the stem and multiple options of an English multiple-choice question, each option is substituted into the stem to obtain multiple sentences to be processed. It is determined whether the sentences to be processed contain conversations. Based on whether the sentences to be processed contain conversations, different option prediction strategies are executed to determine a target option from multiple options as the answer to the English multiple-choice question. By classifying English multiple-choice questions into conversation-based and non-conversation-based types, different option prediction strategies are executed for conversation-based and non-conversation-based multiple-choice questions, and a target option is determined from multiple options as the answer to the English multiple-choice question. In a human-like way of solving problems, the answer to the English multiple-choice question is directly predicted without relying on a huge question bank, which not only saves data storage space but also solves the problem that the correct answer cannot be obtained due to the insufficient coverage of the question bank, and improves the accuracy of answer prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of natural language processing, and in particular, to a method, device, equipment and storage medium for predicting answers to English multiple-choice questions. Background Art

[0002] In recent years, with the booming development of online education, there have emerged more and more online question bank websites, APPs, etc. Among them, the function of taking pictures to search for questions is one of the functions that students use very frequently and is very popular. Students can take pictures of the exercises they don't know, and then view the specific answers and analysis in the APP to help them understand the corresponding exercises and knowledge points. However, in this scenario of taking pictures to search for questions, in addition to relying on the accuracy of the algorithm itself, it also relies on a very large question bank. If the exercise taken by the student is not in the question bank, even if the algorithm is accurate, it is impossible to return a satisfactory answer to the student. However, no matter how large this question bank is, it is impossible to cover all the exercises in the world, and there will always be some exercises that cannot be retrieved. For this scenario, it is generally very difficult to handle at present.

[0003] Among many disciplines, the problem of relying on a question bank for English exercises is particularly prominent. The main reason is that many English exercises come from some newspapers, magazines, etc. in life, and even teachers can make sentences by themselves. For some grammar multiple-choice questions, just randomly dig a few blanks, and different exercises can be generated. Especially for the same sentence, different blanks are dug, the knowledge points examined are different, and different exercises are corresponding. Therefore, the speed of generating new English questions is much higher than that of other disciplines. Therefore, in the scenario of taking pictures to search for questions, English is also the discipline that most often fails to match satisfactory results. Summary of the Invention

[0004] The present invention provides a method, device, equipment and storage medium for predicting answers to English multiple-choice questions, which does not need to rely on a large question bank, not only saves data storage space, but also improves the accuracy of answer prediction.

[0005] In a first aspect, an embodiment of the present invention provides a method for predicting answers to English multiple-choice questions, including:

[0006] Obtain the stem and multiple options of an English multiple-choice question;

[0007] Substitute each of the options into the stem to obtain multiple statements to be processed;

[0008] Determine whether the statement to be processed contains a dialogue;

[0009] Based on whether the statement to be processed contains a dialogue, execute different option prediction strategies to determine a target option from multiple options as the answer to the English multiple-choice question.

[0010] Optionally, determining whether the statement to be processed contains a dialogue includes:

[0011] Inputting each of the statements to be processed into a preset dialogue determination model for processing to determine whether the statement to be processed contains a dialogue.

[0012] Optionally, inputting each of the statements to be processed into a preset dialogue determination model for processing to determine whether the statement to be processed contains a dialogue includes:

[0013] Performing word embedding processing on the characters in the statement to be processed to obtain a representation vector of the statement to be processed;

[0014] Extracting a first feature vector from the representation vector for characterizing whether the statement to be processed contains a dialogue;

[0015] Mapping the first feature vector to a probability value that the statement to be processed contains a dialogue;

[0016] Based on the probability value, determining whether the statement to be processed contains a dialogue.

[0017] Optionally, based on whether the statement to be processed contains a dialogue, performing different option prediction strategies to determine a target option from multiple options, including:

[0018] When the statement to be processed contains a dialogue, calculating the text matching degree between the previous sentence and the next sentence of the statement to be processed;

[0019] Based on the text matching degrees corresponding to the statements to be processed, determining the target option.

[0020] Optionally, calculating the text matching degree between the previous sentence and the next sentence of the statement to be processed includes:

[0021] Inputting the statement to be processed into the input layer of the Roberta model for embedding processing to obtain an embedding matrix;

[0022] Inputting the embedding matrix into the encoding layer of the Roberta model for processing to obtain an encoding matrix;

[0023] Performing a linear transformation on the encoding matrix to obtain a second feature vector;

[0024] Mapping the second feature vector to the text matching degree between the previous sentence and the next sentence of the statement to be processed.

[0025] Optionally, the input layer of the Roberta model includes a word embedding layer, a position embedding layer, and a segment embedding layer. Inputting the statement to be processed into the input layer of the Roberta model for embedding processing to obtain an embedding matrix includes:

[0026] Perform word embedding operation on the characters of the to-be-processed statement in the word embedding layer to obtain a word embedding matrix;

[0027] Perform position embedding operation on the characters of the to-be-processed statement in the position embedding layer to obtain a position embedding matrix;

[0028] Perform segment embedding operation on the to-be-processed statement in the segment embedding layer to obtain a segment embedding matrix;

[0029] Sum up the word embedding matrix, the position embedding matrix and the segment embedding matrix to obtain an embedding matrix.

[0030] Optionally, the encoding layer of the Roberta model includes M stacked multi-head attention layers in sequence. The multi-head attention layer has an input matrix and an output matrix, where M is a positive integer greater than or equal to 2. Input the embedding matrix into the encoding layer of the Roberta model for processing to obtain an encoding matrix, including:

[0031] Input the embedding matrix as the input matrix of the first multi-head attention layer into the first multi-head attention layer for processing;

[0032] Use the output matrix of the previous multi-head attention layer as the input matrix of the next multi-head attention layer until the output matrix of the last multi-head attention layer is obtained as the encoding matrix.

[0033] Optionally, the processing process of each multi-head attention layer includes:

[0034] Process the input matrix of the multi-head attention layer based on the multi-head attention mechanism to obtain an attention matrix;

[0035] Sum up the attention matrix and the input matrix of the multi-head attention layer to obtain a fusion matrix;

[0036] Input the fusion matrix into a fully connected feed-forward layer for processing to obtain a fully connected matrix;

[0037] Sum up the fully connected matrix and the fusion matrix to obtain the output matrix of the multi-head attention layer.

[0038] Optionally, based on whether the to-be-processed statement contains a dialogue, execute different option prediction strategies to determine a target option from multiple options, including:

[0039] When the to-be-processed statement does not contain a dialogue, input the to-be-processed statement into a preset grammar error correction model for processing to obtain a corrected target statement;

[0040] Use the option corresponding to the to-be-processed statement that is the same as the corrected target statement as the intermediate option;

[0041] When the number of the intermediate options is 1, use the intermediate option as the target option;

[0042] When the number of the intermediate options is greater than 1, calculate the perplexity of the to-be-processed statement corresponding to the intermediate option;

[0043] Use the intermediate option corresponding to the to-be-processed statement with the minimum perplexity as the target option.

[0044] Optionally, calculating the perplexity of the to-be-processed statement corresponding to the intermediate option includes:

[0045] Input the to-be-processed statements corresponding to the intermediate options into a preset perplexity calculation model for processing to obtain the probability distribution of each character in the to-be-generated statement in the context semantic environment;

[0046] Calculate the perplexity of the to-be-processed statement based on the probability distribution.

[0047] Optionally, calculating the perplexity of the to-be-processed statement based on the probability distribution includes:

[0048] Calculate the product of all probability values in the probability distribution to obtain a first value;

[0049] Calculate the reciprocal of the first value to obtain a second value;

[0050] Calculate the Nth root of the second value as the perplexity, where N is the number of characters in the to-be-processed statement.

[0051] In a second aspect, an embodiment of the present invention further provides an English multiple-choice question answer prediction device, including:

[0052] A question acquisition module, configured to acquire the stem and multiple options of an English multiple-choice question;

[0053] A to-be-processed statement determination module, configured to substitute each of the options into the stem to obtain multiple to-be-processed statements;

[0054] A dialogue judgment module, configured to judge whether the to-be-processed statement contains a dialogue;

[0055] A target option determination module, configured to execute different option prediction strategies based on whether the to-be-processed statement contains a dialogue, and determine a target option from multiple options as the answer to the English multiple-choice question.

[0056] In a third aspect, an embodiment of the present invention further provides a computer device, including:

[0057] One or more processors;

[0058] A storage device for storing one or more programs;

[0059] When the one or more programs are executed by the one or more processors, the one or more processors implement the English multiple-choice question answer prediction method provided in the first aspect of the present invention.

[0060] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the English multiple-choice question answer prediction method provided in the first aspect of the present invention.

[0061] For the English multiple-choice question answer prediction method provided in the embodiment of the present invention, after obtaining the stem and multiple options of an English multiple-choice question, each option is respectively substituted into the stem to obtain multiple statements to be processed, and it is determined whether the statements to be processed contain dialogues. Based on whether the statements to be processed contain dialogues, different option prediction strategies are executed to determine a target option from multiple options as the answer to the English multiple-choice question. By classifying English multiple-choice questions into dialogue-based and non-dialogue-based types, different option prediction strategies are executed for dialogue-based and non-dialogue-based multiple-choice questions, and a target option is determined from multiple options as the answer to the English multiple-choice question. In a way similar to human problem-solving, the answer to the English multiple-choice question is directly predicted without relying on a huge question bank. This not only saves data storage space but also solves the problem that the coverage of the question bank is insufficient to obtain the correct answer, improving the accuracy of answer prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a flowchart of an English multiple-choice question answer prediction method provided in Embodiment 1 of the present invention;

[0063] Figure 2A It is a flowchart of an English multiple-choice question answer prediction method provided in Embodiment 2 of the present invention;

[0064] Figure 2B It is a schematic structural diagram of the input layer of the Roberta model provided in the embodiment of the present invention;

[0065] Figure 2C It is a schematic structural diagram of the encoding layer of the Roberta model provided in the embodiment of the present invention;

[0066] Figure 3 It is a schematic structural diagram of an English multiple-choice question answer prediction device provided in Embodiment 3 of the present invention;

[0067] Figure 4 It is a schematic structural diagram of a computer device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only the parts related to the present invention rather than all the structures are shown in the accompanying drawings.

[0069] Embodiment 1

[0070] Figure 1 The figure is a flowchart of a method for predicting the answers to English multiple-choice questions provided in Embodiment 1 of the present invention. This embodiment is applicable to predicting the answers to English multiple-choice questions. This method can be executed by the device for predicting the answers to English multiple-choice questions provided in the embodiments of the present invention. The device can be implemented in software and / or hardware and integrated into the computer device provided in the embodiments of the present invention, such as Figure 1 As shown, the method specifically may include the following steps:

[0071] S101. Obtain the stem and multiple options of the English multiple-choice question.

[0072] The English multiple-choice questions in the embodiments of the present invention may include English multiple-choice questions in the pre-school education stage to the high school education stage (or the K12 education stage). Of course, in other embodiments of the present invention, they may also include English multiple-choice questions in the higher education stage, which are not limited herein in the embodiments of the present invention.

[0073] In the embodiments of the present invention, when a user encounters an English multiple-choice question that cannot be answered, the user (such as a student) can take a photo of the area where the stem and multiple options of the English multiple-choice question are located through a user terminal (such as a smart phone, a tablet computer, a personal computer) to obtain a picture including the stem and multiple options of the multiple-choice question. The text content in the picture is recognized by the optical character recognition (OCR) technology on the local user terminal or the server, and then the stem and multiple options of the English multiple-choice question are obtained. In some other embodiments of the present invention, the user can input the stem and multiple options of the multiple-choice question in a preset format on the user terminal. For example, input the stem in the stem column and input multiple options in the option column, and then obtain the stem and multiple options of the English multiple-choice question.

[0074] It should be noted that the above ways of obtaining the stem and multiple options of the English multiple-choice question are exemplary descriptions rather than limitations of the embodiments of the present invention. In other embodiments of the present invention, the stem and multiple options of the English multiple-choice question can also be obtained through other means, which are not limited herein in the embodiments of the present invention.

[0075] S102. Substitute each option into the stem respectively to obtain multiple to-be-processed statements.

[0076] Specifically, each option is substituted into the question stem respectively to obtain multiple statements to be processed corresponding to each option. Exemplarily, for English multiple-choice questions, there are usually 4 options, namely Option A, Option B, Option C, and Option D. Substituting the 4 options into the question stem respectively, 4 statements to be processed are obtained.

[0077] S103. Determine whether the statement to be processed contains a dialogue.

[0078] Specifically, determine whether each statement to be processed contains a dialogue, and further determine whether the type of the multiple-choice question being examined is a dialogue-based type. In the embodiments of the present invention, according to the different contents examined by English multiple-choice questions, English multiple-choice questions are divided into dialogue-based and non-dialogue-based types.

[0079] Among them, for dialogue-based multiple-choice questions, what is examined is whether the connection between the upper sentence and the lower sentence is smooth and conforms to the context after substituting the option into the question stem. The question stem of dialogue-based multiple-choice questions usually includes 2 sentences, which we call the upper sentence and the lower sentence, and there is an obvious separation between the upper sentence and the lower sentence to distinguish different dialogue characters. Dialogue-based multiple-choice questions have 2 dialogue characters, and the issuers of the upper sentence and the lower sentence are these two dialogue characters respectively. For non-dialogue-based multiple-choice questions, the question stem usually includes 1 sentence, or includes multiple sentences, but there is no separation for distinguishing dialogue characters between two adjacent sentences. Based on the above characteristics of dialogue-based multiple-choice questions, it can be determined whether the statement to be processed contains a dialogue.

[0080] S104. Based on whether the statement to be processed contains a dialogue, execute different option prediction strategies to determine the target option from multiple options as the answer to the English multiple-choice question.

[0081] Execute different option prediction strategies based on whether the statement to be processed contains a dialogue, and then determine the target option from multiple options as the answer to the English multiple-choice question. Exemplarily, in some embodiments of the present invention, when the statement to be processed contains a dialogue, the first prediction strategy can be executed; when the statement to be processed does not contain a dialogue, the second prediction strategy can be executed. Among them, the first prediction strategy can be: calculate the text matching degree between the previous sentence and the next sentence of each statement to be processed, and then use the option corresponding to the statement to be processed with the highest text matching degree as the target option. Among them, the text matching degree is used to characterize whether the connection between the previous sentence and the next sentence is smooth and conforms to the context. The second prediction strategy can be: first determine whether there is a grammar error in the statement to be processed, then calculate the perplexity of the statement to be processed without grammar errors, and then use the option corresponding to the statement to be processed with the smallest perplexity as the target option. Among them, the perplexity is used to measure the goodness of a probability distribution or probability model in predicting samples. More generally speaking, it is to describe the probability that the statement to be processed is a normal sentence, and whether the statement obtained by substituting this option into the question stem conforms to people's normal expression habits. It should be noted that the first prediction strategy and the second prediction strategy in the above embodiments are exemplary descriptions rather than limitations of the embodiments of the present invention.

[0082] The method for predicting the answer to an English multiple-choice question provided by the embodiment of the present invention, after obtaining the question stem and multiple options of the English multiple-choice question, substitutes each option into the question stem respectively to obtain multiple statements to be processed, determines whether the statement to be processed contains a dialogue, and based on whether the statement to be processed contains a dialogue, executes different option prediction strategies to determine the target option from multiple options as the answer to the English multiple-choice question. By classifying English multiple-choice questions into dialogue types and non-dialogue types, and executing different option prediction strategies for dialogue-type and non-dialogue-type multiple-choice questions, the target option is determined from multiple options as the answer to the English multiple-choice question. Using a human-like problem-solving method, the answer to the English multiple-choice question is directly predicted without relying on a large question bank, which not only saves data storage space, but also solves the problem that the coverage of the question bank is insufficient to obtain the correct answer, and improves the accuracy of answer prediction.

[0083] Embodiment 2

[0084] Figure 2A It is a flowchart of a method for predicting the answer to an English multiple-choice question provided by Embodiment 2 of the present invention. This embodiment is refined on the basis of Embodiment 1 above, and details the detailed processes of different prediction strategies, as Figure 2A shown. The method includes:

[0085] S201. Obtain the question stem and multiple options of the English multiple-choice question.

[0086] As described in the foregoing embodiments, the English multiple-choice questions in the embodiments of the present invention may include English multiple-choice questions in the pre-school education stage to the high school education stage (or the K12 education stage). Of course, in other embodiments of the present invention, English multiple-choice questions in the higher education stage may also be included, and the embodiments of the present invention do not make limitations here. The stem and multiple options of the English multiple-choice questions can be obtained by taking pictures or by manual input by the user, and the embodiments of the present invention do not make limitations here.

[0087] S202. Substitute each option into the stem respectively to obtain a plurality of statements to be processed.

[0088] Specifically, substitute each option into the stem respectively to obtain a plurality of statements to be processed corresponding to each option.

[0089] S203. Determine whether the statement to be processed contains a dialogue.

[0090] Specifically, in the embodiments of the present invention, according to the different contents examined by the English multiple-choice questions, the English multiple-choice questions are divided into dialogue types and non-dialogue types. Among them, the non-dialogue types are divided into grammar correction types and expression types. The grammar correction type multiple-choice questions are used to examine English grammar, and the expression type multiple-choice questions are used to examine normal expression habits.

[0091] Determining whether the statement to be processed contains a dialogue is essentially a binary classification problem. A dialogue judgment model is pre-trained, and then each statement to be processed is input into the trained dialogue judgment model for processing, and it can be judged whether the statement to be processed contains a dialogue. Specifically, the dialogue judgment model can be a common convolutional neural network model. The processing process of the dialogue judgment model for the statement to be processed is as follows:

[0092] 1. Perform word embedding processing on the characters in the statement to be processed to obtain a representation vector of the statement to be processed.

[0093] Characters can be understood as placeholders in the statement to be processed, and each word or punctuation is a character. In the embodiments of the present invention, perform word embedding processing on the characters in the statement to be processed to obtain a representation vector of the statement to be processed. Word Embedding refers to converting a word into a vector representation. Specifically, the method of word embedding can be implemented by one-hot encoding or Word2Vec algorithm, etc., and the embodiments of the present invention will not elaborate here.

[0094] 2. Extract a first feature vector from the representation vector for characterizing whether the statement to be processed contains a dialogue.

[0095] Exemplarily, the characterization vector can be input into a convolutional network for processes such as convolution and pooling to obtain a first feature vector characterizing whether the statement to be processed contains a dialogue. Specifically, the convolutional network may include one or more convolutional layers, which are not limited in the embodiments of the present invention.

[0096] 3. Map the first feature vector to a probability value that the statement to be processed contains a dialogue.

[0097] Specifically, the first feature vector can be mapped to a value between [0, 1] via a classification function to obtain a probability value that the statement to be processed contains a dialogue. Specifically, the classification function can be the sigmoid function.

[0098] 4. Determine whether the statement to be processed contains a dialogue based on the probability value.

[0099] Specifically, compare the obtained probability value with a preset value. When the probability value is greater than or equal to the preset value (the preset value can be set to 0.8), it is determined that the statement to be processed contains a dialogue. When the probability value is less than the preset value, it is determined that the statement to be processed does not contain a dialogue.

[0100] When the statement to be processed contains a dialogue, execute step S204 - step S205. When the statement to be processed does not contain a dialogue, execute step S206 - step S211.

[0101] S204. Calculate the text matching degree between the previous sentence and the next sentence of the statement to be processed.

[0102] Specifically, when the statement to be processed contains a dialogue, it indicates that the English multiple-choice question is a dialogue-type multiple-choice question, and what is examined is whether the connection between the previous sentence and the next sentence is fluent and conforms to the context after substituting the options into the question stem. Therefore, the text matching degree between the previous sentence and the next sentence of each statement to be processed can be calculated, and then the target option can be determined according to the text matching degree.

[0103] In some embodiments of the present invention, the process of calculating the text matching degree between the previous sentence and the next sentence of the statement to be processed is as follows:

[0104] 1. Input the statement to be processed into the input layer of the Roberta model for embedding processing to obtain an embedding matrix.

[0105] Specifically, the input layer of the Roberta (Robustly Optimized BERT Pretraining Approach) model is used to perform word embedding (Token Embedding) operations on the characters in the statement to be processed, obtaining a word embedding matrix, perform position embedding (Position Embedding) operations on the characters in the statement to be processed, obtaining a position embedding matrix, and perform segment embedding (Segment Embedding) operations on the statement to be processed, obtaining a segment embedding matrix. Among them, the word embedding operation refers to representing each character with the encoding in the dictionary library and converting it into a word embedding vector with a fixed dimension. The matrix composed of all word embedding vectors is the word embedding matrix; the position embedding operation refers to numbering the positions of each character in the statement to be processed, and then each number corresponds to a vector. By combining the position vector and the word vector, certain position information is introduced for each word. The segment embedding operation refers to encoding the statement to be processed to distinguish different sentences in the statement to be processed. The distinguishing method is that all the words in one of the adjacent two sentences are assigned 0, and all the words in the other sentence are assigned 1, and the starting position of each sentence is represented by the identifier "CLS", and the ending position of each sentence is represented by the identifier "SEP".

[0106] Figure 2B FIG. is a schematic structural diagram of the input layer of the Roberta model provided by the embodiment of the present invention. Exemplarily, as Figure 2B shown, the input layer of the Roberta model includes a word embedding layer, a position embedding layer, and a segment embedding layer. Inputting the statement to be processed into the input layer of the Roberta model for embedding processing, the process of obtaining the embedding matrix is as follows:

[0107] Perform word embedding operations on the characters of the statement to be processed in the word embedding layer to obtain a word embedding matrix. Perform position embedding operations on the characters of the statement to be processed in the position embedding layer to obtain a position embedding matrix. Perform segment embedding operations on the statement to be processed in the segment embedding layer to obtain a segment embedding matrix. Add the word embedding matrix, the position embedding matrix, and the segment embedding matrix to obtain the embedding matrix E.

[0108] 2. Input the embedding matrix into the encoding layer of the Roberta model for processing to obtain a coding matrix.

[0109] The encoding layer of the Roberta model processes the embedding matrix E based on the multi-head self-attention mechanism to obtain the coding matrix T. Figure 2C FIG. is a schematic structural diagram of the encoding layer of the Roberta model provided by the embodiment of the present invention, as Figure 2CAs shown, specifically, the encoding layer of the Roberta model includes M stacked multi-head attention layers in sequence. The multi-head attention layer has an input matrix and an output matrix, where M is a positive integer greater than or equal to 2. Usually, the multi-head attention layer is also called a transformer. Therefore, the encoding layer of the Roberta model is composed of M stacked transformers. Exemplarily, in the embodiment of the present invention, M = 12. The embedding matrix E is input into the encoding layer of the Roberta model for processing to obtain the encoding matrix T, including:

[0110] The embedding matrix E is used as the input matrix of the first multi-head attention layer and input into the first multi-head attention layer for processing.

[0111] The output matrix of the previous multi-head attention layer is used as the input matrix of the next multi-head attention layer, and so on, until the output matrix of the last multi-head attention layer is used as the encoding matrix T.

[0112] Among them, each multi-head attention layer processes the input matrix based on the multi-head attention mechanism. Exemplarily, the processing process of the multi-head attention layer is as follows:

[0113] 1) Process the input matrix of the multi-head attention layer based on the multi-head attention mechanism to obtain the attention matrix.

[0114] Specifically, taking the first multi-head attention layer as an example, the processing process of the multi-head attention mechanism (Multi-Head Attention) is as follows:

[0115] First, perform three linear transformations on the embedding matrix E using three different linear transformation coefficients to obtain matrix Q, matrix K, and matrix V respectively.

[0116] Q = EW i Q

[0117] K = EW i K

[0118] V = EW i V

[0119] Among them, W i Q is the linear transformation coefficient of matrix Q of the i-th multi-head attention layer, W i K is the linear transformation coefficient of matrix K of the i-th multi-head attention layer, W i V is the linear transformation coefficient of matrix V of the i-th multi-head attention layer, where i = 1 here.

[0120] Next, perform m linear transformations on matrices Q, K, and V respectively to obtain matrix Q i , matrix K i and matrix V i , where i ∈ m and m is the number of attention heads in the multi-head attention layer. Exemplarily, taking m = 2 as an example, the processing process of the multi-head attention layer is described exemplarily.

[0121] Next, calculate the dot product of matrix Q i and matrix K i to obtain the first sub-matrix, and calculate the quotient of the first sub-matrix and the square root of the dimension of matrix K i to obtain the second sub-matrix. Then, perform normalization processing on the second sub-matrix to obtain the third sub-matrix. Next, calculate the dot product of the third sub-matrix and matrix V i to obtain the fourth sub-matrix (i.e., head i ).

[0122]

[0123] Among them, is the dot product of matrix Q i and matrix Ki, is the transposed matrix of K i , d 1 is the dimension of matrix K i , and softmax is the normalization processing.

[0124] Finally, splice the m fourth sub-matrices head i to obtain the spliced matrix, and perform a linear transformation on the first spliced matrix to obtain the attention matrix M.

[0125] M = Multihead(Q, K, V) = concat(head 1 ,..., head m )W 0

[0126] Among them, concat is matrix splicing, and W 0 is the linear transformation coefficient for performing a linear transformation on the spliced matrix.

[0127] 2) Add the attention matrix and the input matrix of the multi-head attention layer to obtain the fusion matrix.

[0128] Add the attention matrix output by the multi-head attention mechanism (Multi-Head Attention) and the input matrix of the multi-head attention layer to obtain the fusion matrix. As Figure 2CAs shown, taking the first - layer multi - head attention layer as an example, the attention matrix M output by the multi - head attention mechanism is added to the input matrix (embedding matrix E) of the multi - head attention layer to obtain a fusion matrix.

[0129] In a specific embodiment of the present invention, in order to accelerate the convergence speed of the network, the attention matrix M can be normalized (Normalization, denoted as norm in the figure). In order to reduce the over - fitting phenomenon of the network, the matrix after normalization is input into a dropout layer for random dropout operation to obtain matrix M 1 . Then, a residual connection is made between the output of the dropout layer and the input of the multi - head attention layer, that is, matrix M 1 is added to the embedding matrix E to obtain the fusion matrix M 2 .

[0130] 3) The fusion matrix is input into a fully - connected feed - forward layer for processing to obtain a fully - connected matrix.

[0131] Specifically, in the embodiment of the present invention, in order to accelerate the convergence speed of the network, the fusion matrix M 2 can be pre - processed by layer normalization to obtain matrix M 3 . The layer normalization process is as follows:

[0132]

[0133] where t i represents normalizing each row of the fusion matrix M 2 , u L and σ L represent the mean and variance of each sample respectively, α and β represent the parameter vectors for scaling and translation, ε is a bias parameter to avoid the denominator being zero, and after normalizing each row, matrix M 3 is obtained.

[0134] Then, matrix M 3 is input into a fully - connected feed - forward layer (Fully Connected Feed Forward Network, FFN) for processing to obtain the fully - connected matrix M 4 . Specifically, the processing process of the fully - connected feed - forward layer is shown as follows:

[0135] M 4 = FFN(M 3 ) = Max(0, M 3 W 1 + b 1 )W 2 + b 2

[0136] Specifically, the fully connected feed-forward layer first performs a non-linear transformation on the matrix M 3 with transformation parameters (W 1 , b 1 ) to obtain the matrix M 3 W 1 +b 1 . Then, the non-linear activation function Max(0, a) is used to perform non-linear activation on M 3 W 1 +b 1 . Next, a linear transformation is performed again on the matrix obtained after non-linear activation, with transformation parameters (W 2 , b 2 ). The non-linear activation function Max(0, a) is used to replace the negative elements in the matrix M 3 W 1 +b 1 with 0.

[0137] 4) Add the fully connected matrix and the fusion matrix to obtain the output matrix of the multi-head attention layer.

[0138] Specifically, in an embodiment of the present invention, in order to accelerate the convergence speed of the network and reduce the overfitting phenomenon of the network, the fully connected matrix M 4 can be pre-normalized and randomly dropped. Then, the output of the random drop is connected in residual to the input of the fully connected feed-forward layer, that is, the output of the random drop is added to the fusion matrix M 2 to obtain the output matrix of the multi-head attention layer.

[0139] And so on, the output matrix of the previous second multi-head attention layer is used as the input matrix of the next second multi-head attention layer until the output matrix of the last second multi-head attention layer is obtained as the encoding matrix T.

[0140] 3. Perform a linear transformation on the encoding matrix to obtain the second feature vector.

[0141] Specifically, perform a linear transformation on the encoding matrix T again to convert the encoding matrix T into the second feature vector.

[0142] 4. Map the second feature vector to the text matching degree between the previous sentence and the next sentence of the sentence to be processed.

[0143] Specifically, map the second feature vector through a classification function to the range [0, 1] to obtain the text matching degree between the previous sentence and the next sentence of the sentence to be processed. The text matching degree is used to characterize whether the connection between the previous sentence and the next sentence is smooth and conforms to the context. Specifically, the classification function can be the sigmoid function or the softmax function.

[0144] S205. Determine the target option based on the text matching degree corresponding to each statement to be processed.

[0145] The text matching degree is used to characterize whether the connection between the previous sentence and the next sentence is smooth and conforms to the context. The higher the text matching degree, the smoother the connection between the previous sentence and the next sentence of the statement to be processed and the more it conforms to the context. In the embodiments of the present invention, the option corresponding to the statement to be processed with the highest text matching degree is taken as the target option, and this target option is the answer to the predicted English multiple-choice question.

[0146] S206. Input the statement to be processed into a preset grammar error correction model for processing to obtain the corrected target statement.

[0147] Specifically, when the statement to be processed does not contain a dialogue, input the statement to be processed into a preset grammar error correction model for processing to obtain the corrected target statement.

[0148] For grammar error correction multiple-choice questions, it is essentially a seq2seq problem, and the grammar error correction model is a seq2seq model. That is, the input of the grammar error correction model is a statement with grammar errors, and the output is a statement without grammar errors after grammar correction. The seq2seq model belongs to a type of encoder-decoder structure. In the common encoder-decoder structure, the basic idea is to use two RNNs, one RNN as the encoder and the other RNN as the decoder. The encoder is responsible for compressing the input sequence into a vector of a specified length, and this vector can be regarded as the semantics of this sequence. This process is called encoding. The simplest way to obtain the semantic vector is to directly use the hidden state of the last input as the semantic vector C. It is also possible to perform a transformation on the last hidden state to obtain the semantic vector, or perform a transformation on all the hidden states of the input sequence to obtain the semantic variable.

[0149] The decoder is responsible for generating a specified sequence based on the semantic vector, and this process is also called decoding. The simplest way is to use the semantic variable obtained by the encoder as the initial state and input it into the RNN of the decoder to obtain the output sequence. It can be seen that the output of the previous moment will be used as the input of the current moment, and among them, the semantic vector C only participates in the operation as the initial state, and the subsequent operations have nothing to do with the semantic vector C. There is another way of the decoder processing method, that is, the semantic vector C participates in the operations of all moments of the sequence. The output of the previous moment is still used as the input of the current moment, but the semantic vector C will participate in the operations of all moments.

[0150] S207. Take the option corresponding to the statement to be processed that is the same as the corrected target statement as the intermediate option.

[0151] Specifically, compare the corrected target statement with the input statement to be processed to see if they are the same. If they are the same, it means that the statement to be processed obtained by substituting this option into the question stem has no grammar errors. If they are not the same, it means that the statement to be processed obtained by substituting this option into the question stem has grammar errors. Select the option corresponding to the statement to be processed that is the same as the corrected target statement as the intermediate option, that is, filter out the options corresponding to the statements to be processed that have no grammar errors as the intermediate options.

[0152] S208. Determine whether the number of intermediate options is greater than 1.

[0153] Specifically, determine whether the number of intermediate options is greater than 1. If the number of intermediate options is equal to 1, then execute step S209. If the number of intermediate options is greater than 1, then execute steps S210 - S211.

[0154] S209. Take the intermediate option as the target option.

[0155] Specifically, when the number of filtered intermediate options is 1, that is, only one option is substituted into the question stem to obtain a statement to be processed without grammar errors, and the others have grammar errors, then it can be considered that this unique intermediate option is the target option, that is, the answer to the English multiple-choice question.

[0156] S210. Calculate the perplexity of the statements to be processed corresponding to the intermediate options.

[0157] Specifically, the statements to be processed without grammar errors may conform to the grammar rules but do not conform to people's normal expression habits. For example, for the question stem "I___Natural Language Processing", for the option "hit", although the statement obtained by substituting this option into the question stem has no grammar errors, it obviously does not conform to the normal expression habit. When the number of intermediate options is greater than 1, that is, there are multiple options that are substituted into the question stem to obtain statements to be processed without grammar errors. At this time, it is necessary to calculate the perplexity of the statements to be processed corresponding to the intermediate options, and then determine the target option according to the perplexity of each statement to be processed. Specifically, the calculation process of the perplexity of the statement to be processed is as follows:

[0158] 1. Input the statements to be processed corresponding to each intermediate option into the pre-set perplexity calculation model for processing to obtain the probability distribution of each character in the statement to be generated in the context semantic environment.

[0159] Secondly, for expression multiple-choice questions, more consideration is given to semantic content, that is, whether it is appropriate to fill the blank with the options and whether it conforms to people's expression habits. And to judge whether such an expression is appropriate, it is often based on whether such an expression appears frequently. Therefore, we can use a large amount of English language materials in life to train such a model to measure whether the word in the blank is appropriate and, if so, to what extent.

[0160] In the embodiments of the present invention, a generative pre-training model (Generative Pre-Training, GPT), specifically the GPT2.0 model, is introduced as the perplexity calculation model. The language model of GPT2.0 is a pre-trained language model open-sourced by the OpenAI team. It uses approximately 10 million articles and nearly 40G of text content. Therefore, it can learn a large number of associations between texts and has very strong semantic expression capabilities. At the same time, it is a rare unidirectional language model among pre-trained language models. The calculation method of the unidirectional language model is closer to the calculation method of the normal language model, improving the calculation accuracy of perplexity.

[0161] The GPT2.0 model uses pre-training and downstream fine-tuning methods to process NLP tasks and solve dynamic semantic problems. After being processed by the GPT2.0 model, the probability distribution of each character in the generated sentence corresponding to the sentence to be processed corresponding to each intermediate option in the context semantic environment is obtained.

[0162] 2. Calculate the perplexity of the sentence to be processed based on the probability distribution.

[0163] Next, calculate the perplexity of a corresponding sentence to be processed based on the probability distribution as the perplexity PP(S). Specifically, the calculation method is as follows:

[0164]

[0165] Specifically, first, calculate the product of all probability values in the probability distribution to obtain the first value P(w 1 w 2 ...w N ), where w 1 w 2 ...w N are the characters in the sentence to be processed. Here, N is the number of characters in the sentence to be processed, and S is the input sentence to be processed with a length of N. Next, calculate the reciprocal of the first value to obtain the second value, and then calculate the Nth root of the second value as the first perplexity. To simplify the calculation process, the above product calculation process can be converted into a logarithmic connection process, and P(w i |w 1:i-1 ) represents the probability value of the i-th symbol based on the context semantic environment.

[0166] S211. Take the intermediate option corresponding to the statement to be processed with the minimum perplexity as the target option.

[0167] After calculating the perplexity of the statements to be processed corresponding to each intermediate option, take the intermediate option corresponding to the statement to be processed with the minimum perplexity as the target option, and this target option is the answer to the predicted English multiple-choice question.

[0168] The method for predicting the answer to an English multiple-choice question provided by the embodiment of the present invention divides the English multiple-choice question into a dialogue type and a non-dialogue type, executes different option prediction strategies for the dialogue-type and non-dialogue-type multiple-choice questions, determines the target option from multiple options as the answer to the English multiple-choice question, and directly predicts the answer to the English multiple-choice question in a human-like problem-solving manner without relying on a huge question bank. This not only saves data storage space but also can solve the problem that the correct answer cannot be obtained due to the insufficient coverage of the question bank, improving the accuracy of answer prediction. In addition, multiple perspectives such as context, semantics, grammar, and expression habits are considered simultaneously for answer prediction, enhancing the accuracy of the model. The embodiment of the present invention introduces external data through a perplexity calculation model, obtains powerful semantic information from a large amount of data to assist in the prediction of the final question answer, and enhances the credibility, accuracy, and generalization ability of the model.

[0169] Embodiment III

[0170] Figure 3 As shown in the structural schematic diagram of an apparatus for predicting the answer to an English multiple-choice question provided by Embodiment III of the present invention, Figure 3 as shown, the apparatus includes:

[0171] A question acquisition module 301, configured to acquire the stem and multiple options of an English multiple-choice question;

[0172] A statement to be processed determination module 302, configured to substitute each of the options into the stem to obtain multiple statements to be processed;

[0173] A dialogue judgment module 303, configured to judge whether the statement to be processed contains a dialogue;

[0174] A target option determination module 304, configured to execute different option prediction strategies based on whether the statement to be processed contains a dialogue, and determine the target option from multiple options as the answer to the English multiple-choice question.

[0175] In some embodiments of the present invention, the dialogue judgment module 303 is further configured to:

[0176] Input each of the statements to be processed into a preset dialogue judgment model for processing to judge whether the statement to be processed contains a dialogue.

[0177] In some embodiments of the present invention, the dialogue judgment module 303 includes:

[0178] A word embedding sub-module, configured to perform word embedding processing on the characters in the to-be-processed statement to obtain a representation vector of the to-be-processed statement;

[0179] A first feature vector extraction sub-module, configured to extract a first feature vector from the representation vector for characterizing whether the to-be-processed statement contains a dialogue;

[0180] A probability value obtaining sub-module, configured to map the first feature vector to a probability value that the to-be-processed statement contains a dialogue;

[0181] A dialogue judgment sub-module, configured to judge whether the to-be-processed statement contains a dialogue based on the probability value.

[0182] In some embodiments of the present invention, the target option determination module 304 includes:

[0183] A text matching degree calculation sub-module, configured to calculate the text matching degree between the previous sentence and the next sentence of the to-be-processed statement when the to-be-processed statement contains a dialogue;

[0184] A first target option determination sub-module, configured to determine a target option based on the text matching degree corresponding to each to-be-processed statement.

[0185] In some embodiments of the present invention, the text matching degree calculation sub-module includes:

[0186] An embedding processing unit, configured to input the to-be-processed statement into the input layer of the Roberta model for embedding processing to obtain an embedding matrix;

[0187] An encoding matrix obtaining unit, configured to input the embedding matrix into the encoding layer of the Roberta model for processing to obtain an encoding matrix;

[0188] A linear transformation unit, configured to perform a linear transformation on the encoding matrix to obtain a second feature vector;

[0189] A text matching degree obtaining unit, configured to map the second feature vector to the text matching degree between the previous sentence and the next sentence of the to-be-processed statement.

[0190] In some embodiments of the present invention, the input layer of the Roberta model includes a word embedding layer, a position embedding layer, and a segment embedding layer, and the embedding processing unit includes:

[0191] A word embedding sub-unit, configured to perform a word embedding operation on the characters of the to-be-processed statement in the word embedding layer to obtain a word embedding matrix;

[0192] A position embedding sub - unit, configured to perform a position embedding operation on the characters of the to - be - processed statement in the position embedding layer to obtain a position embedding matrix;

[0193] A segment embedding sub - unit, configured to perform a segment embedding operation on the to - be - processed statement in the segment embedding layer to obtain a segment embedding matrix;

[0194] An addition sub - unit, configured to add the word embedding matrix, the position embedding matrix, and the segment embedding matrix to obtain an embedding matrix.

[0195] In some embodiments of the present invention, the encoding layer of the Roberta model includes M stacked multi - head attention layers in sequence. The multi - head attention layer has an input matrix and an output matrix, where M is a positive integer greater than or equal to 2. The encoding matrix acquisition unit includes:

[0196] A first processing sub - unit, configured to input the embedding matrix as the input matrix of the first - layer multi - head attention layer into the first - layer multi - head attention layer for processing;

[0197] A second processing sub - unit, configured to use the output matrix of the previous - layer multi - head attention layer as the input matrix of the next - layer multi - head attention layer until the output matrix of the last - layer multi - head attention layer is obtained as the encoding matrix.

[0198] In some embodiments of the present invention, the second processing sub - unit is configured to:

[0199] Process the input matrix of the multi - head attention layer based on the multi - head attention mechanism to obtain an attention matrix;

[0200] Add the attention matrix to the input matrix of the multi - head attention layer to obtain a fusion matrix;

[0201] Input the fusion matrix into a fully - connected feed - forward layer for processing to obtain a fully - connected matrix;

[0202] Add the fully - connected matrix to the fusion matrix to obtain the output matrix of the multi - head attention layer.

[0203] In some embodiments of the present invention, the target option determination module 304 further includes:

[0204] An error - correction processing sub - module, configured to input the to - be - processed statement into a preset grammar error - correction model for processing to obtain a corrected target statement when the to - be - processed statement does not contain a dialogue;

[0205] An intermediate option acquisition sub - module, configured to use the option corresponding to the to - be - processed statement that is the same as the corrected target statement as the intermediate option;

[0206] A second target option determination sub-module, configured to use the intermediate option as the target option when the number of the intermediate options is 1;

[0207] A perplexity calculation sub-module, configured to calculate the perplexity of the to-be-processed statement corresponding to the intermediate option when the number of the intermediate options is greater than 1;

[0208] A third target option determination sub-module, configured to use the intermediate option corresponding to the to-be-processed statement with the minimum perplexity as the target option.

[0209] In some embodiments of the present invention, the perplexity calculation sub-module includes:

[0210] A probability distribution acquisition unit, configured to respectively input the to-be-processed statements corresponding to the intermediate options into a preset perplexity calculation model for processing, to obtain the probability distribution of each character in the to-be-generated statement in the context semantic environment;

[0211] A perplexity calculation unit, configured to calculate the perplexity of the to-be-processed statement based on the probability distribution.

[0212] In some embodiments of the present invention, the perplexity calculation unit includes:

[0213] A first numerical calculation sub-unit, configured to calculate the product of all probability values in the probability distribution to obtain a first numerical value;

[0214] A second numerical calculation sub-unit, configured to calculate the reciprocal of the first numerical value to obtain a second numerical value;

[0215] A perplexity calculation sub-unit, configured to calculate the Nth root of the second numerical value as the perplexity, where N is the number of characters in the to-be-processed statement.

[0216] The above English multiple-choice question answer prediction device can execute the method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0217] Embodiment 4

[0218] Embodiment 4 of the present invention provides a computer device. Figure 4 As shown in the structural schematic diagram of the computer device provided in Embodiment 4 of the present invention, Figure 4 as shown, the computer device includes a processor 401, a memory 402, a communication module 403, an input device 404, and an output device 405; the number of processors 401 in the computer device can be one or more. Figure 4 Taking one processor 401 as an example; the processor 401, the memory 402, the communication module 403, the input device 404, and the output device 405 in the computer device can be connected through a bus or other means.Figure 4 Take the bus connection as an example. The above-mentioned processor 401, memory 402, communication module 403, input device 404, and output device 405 can be integrated on the control main board of the computer device.

[0219] As a computer-readable storage medium, the memory 402 can be used to store software programs, computer-executable programs, and modules, such as the module corresponding to the English multiple-choice question answer prediction method in this embodiment. The processor 401 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory 402, that is, implements the English multiple-choice question answer prediction method provided in the above embodiment.

[0220] The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device. In addition, the memory 402 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 402 can further include a memory remotely set relative to the processor 401, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.

[0221] The communication module 403 is used to establish a connection with external devices (such as smart terminals) and implement data interaction with external devices. The input device 404 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function control of the computer device.

[0222] A computer device provided in this embodiment can execute the English multiple-choice question answer prediction method provided in any of the above embodiments of the present invention, and specifically has corresponding functions and beneficial effects.

[0223] Embodiment Five

[0224] Embodiment Five of the present invention provides a storage medium containing computer-executable instructions, on which a computer program is stored. When the program is executed by a processor, it implements the English multiple-choice question answer prediction method provided in any of the above embodiments of the present invention. The method includes:

[0225] Obtain the stem and multiple options of the English multiple-choice question;

[0226] Substitute each of the options into the stem to obtain multiple statements to be processed;

[0227] Determine whether the statement to be processed contains a dialogue;

[0228] Execute different option prediction strategies based on whether the statement to be processed contains a dialogue, and determine a target option from multiple options as the answer to the English multiple-choice question.

[0229] Of course, for the storage medium containing computer-executable instructions provided by the embodiments of the present invention, the computer-executable instructions are not limited to the method operations described above, and can also execute relevant operations in the method for predicting the answer to the English multiple-choice question provided by the embodiments of the present invention.

[0230] It should be noted that for the embodiments of the device, equipment, and storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0231] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk, or optical disc of a computer, etc., including several instructions for causing a computer device (which can be a robot, personal computer, server, or network device, etc.) to execute the method for predicting the answer to the English multiple-choice question described in any embodiment of the present invention.

[0232] It should be noted that in the above device, the included modules, sub-modules, units, and sub-units are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0233] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution device. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0234] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0235] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. An English multiple-choice question answer prediction method, characterized in that, it includes: Obtain the stem and multiple options of an English multiple-choice question; Substitute each of the options into the stem to obtain multiple statements to be processed; Determine whether the statement to be processed contains a dialogue; Based on whether the statement to be processed contains a dialogue, execute different option prediction strategies to determine a target option from multiple options as the answer to the English multiple-choice question, including: When the statement to be processed contains a dialogue, calculate the text matching degree between the upper sentence and the lower sentence of the statement to be processed; determine the target option based on the text matching degree corresponding to each statement to be processed; When the statement to be processed does not contain a dialogue, determine the statement to be processed without grammar errors, and determine the target option based on the perplexity of the statement to be processed without grammar errors.

2. The English multiple-choice question answer prediction method according to claim 1, characterized in that, Determining whether the statement to be processed contains a dialogue includes: Input each of the statements to be processed into a preset dialogue judgment model for processing, and determine whether the statement to be processed contains a dialogue.

3. The English multiple-choice question answer prediction method according to claim 2, characterized in that, Inputting each of the statements to be processed into a preset dialogue judgment model for processing, and determining whether the statement to be processed contains a dialogue includes: Performing word embedding processing on the characters in the statement to be processed to obtain a representation vector of the statement to be processed; Extract a first feature vector from the representation vector for characterizing whether the statement to be processed contains a dialogue; Map the first feature vector to a probability value that the statement to be processed contains a dialogue; Based on the probability value, determine whether the statement to be processed contains a dialogue.

4. The English multiple-choice question answer prediction method according to claim 1, characterized in that, Calculating the text matching degree between the upper sentence and the lower sentence of the statement to be processed includes: Input the statement to be processed into the input layer of the Roberta model for embedding processing to obtain an embedding matrix; Input the embedding matrix into the encoding layer of the Roberta model for processing to obtain an encoding matrix; Perform a linear transformation on the encoding matrix to obtain a second feature vector; Map the second feature vector to the text matching degree between the upper sentence and the lower sentence of the statement to be processed.

5. The English multiple-choice question answer prediction method according to claim 4, characterized in that, The input layer of the Roberta model includes a word embedding layer, a position embedding layer, and a segment embedding layer. Inputting the statement to be processed into the input layer of the Roberta model for embedding processing to obtain an embedding matrix includes: Perform word embedding operations on the characters of the statement to be processed in the word embedding layer to obtain a word embedding matrix; Perform position embedding operations on the characters of the statement to be processed in the position embedding layer to obtain a position embedding matrix; Perform segment embedding operations on the statement to be processed in the segment embedding layer to obtain a segment embedding matrix; Add the word embedding matrix, the position embedding matrix, and the segment embedding matrix to obtain an embedding matrix.

6. The method for predicting the answers of English multiple-choice questions according to claim 4, characterized in that, the encoding layer of the Roberta model includes M stacked multi-head attention layers in sequence, the multi-head attention layer has an input matrix and an output matrix, M is a positive integer greater than or equal to 2, and the embedding matrix is input into the encoding layer of the Roberta model for processing to obtain an encoding matrix, including: taking the embedding matrix as the input matrix of the first multi-head attention layer and inputting it into the first multi-head attention layer for processing; taking the output matrix of the previous multi-head attention layer as the input matrix of the next multi-head attention layer until the output matrix of the last multi-head attention layer is used as the encoding matrix.

7. The method for predicting the answers of English multiple-choice questions according to claim 6, characterized in that, the processing process of each multi-head attention layer includes: processing the input matrix of the multi-head attention layer based on the multi-head attention mechanism to obtain an attention matrix; adding the attention matrix and the input matrix of the multi-head attention layer to obtain a fusion matrix; inputting the fusion matrix into a fully connected feed-forward layer for processing to obtain a fully connected matrix; adding the fully connected matrix and the fusion matrix to obtain the output matrix of the multi-head attention layer.

8. The method for predicting the answers of English multiple-choice questions according to any one of claims 1-7, characterized in that, when the statement to be processed does not contain a dialogue, determining a statement to be processed without grammar errors, and determining a target option based on the perplexity of the statement to be processed without grammar errors, including: when the statement to be processed does not contain a dialogue, inputting the statement to be processed into a preset grammar error correction model for processing to obtain a corrected target statement; taking the option corresponding to the statement to be processed that is the same as the corrected target statement as an intermediate option; when the number of intermediate options is 1, taking the intermediate option as the target option; when the number of intermediate options is greater than 1, calculating the perplexity of the statement to be processed corresponding to the intermediate option; taking the intermediate option corresponding to the statement to be processed with the smallest perplexity as the target option.

9. The method for predicting the answers of English multiple-choice questions according to claim 8, characterized in that, calculating the perplexity of the statement to be processed corresponding to the intermediate option includes: inputting the statements to be processed corresponding to each intermediate option into a preset perplexity calculation model for processing to obtain the probability distribution of each character in the statement to be generated in the context semantic environment; calculating the perplexity of the statement to be processed based on the probability distribution.

10. The method for predicting the answers of English multiple-choice questions according to claim 9, characterized in that, calculating the perplexity of the statement to be processed based on the probability distribution includes: calculating the product of all probability values in the probability distribution to obtain a first value; calculating the reciprocal of the first value to obtain a second value; calculating the Nth root of the second value as the perplexity, where N is the number of characters in the statement to be processed.

11. An apparatus for predicting the answers of English multiple-choice questions, characterized in that, comprising: A question acquisition module for acquiring the stem and multiple options of an English multiple-choice question; A to-be-processed statement determination module for substituting each of the options into the stem to obtain a plurality of to-be-processed statements; A dialogue judgment module for judging whether the to-be-processed statement contains a dialogue; A target option determination module for executing different option prediction strategies based on whether the to-be-processed statement contains a dialogue, and determining a target option from the multiple options as the answer to the English multiple-choice question; The target option determination module includes: A text matching degree calculation sub-module for calculating the text matching degree between the upper sentence and the lower sentence of the to-be-processed statement when the to-be-processed statement contains a dialogue; A first target option determination sub-module for determining a target option based on the text matching degree corresponding to each to-be-processed statement; The target option determination module is specifically configured to, when the to-be-processed statement does not contain a dialogue, determine a to-be-processed statement without a grammar error, and determine a target option based on the perplexity of the to-be-processed statement without a grammar error.

12. A computer device characterized in that it includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the English multiple-choice question answer prediction method according to any one of claims 1-10.

13. A computer-readable storage medium, on which a computer program is stored characterized in that when the program is executed by a processor, it implements the English multiple-choice question answer prediction method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Multi-view fused answer selection method for reading understanding choice questions

    CN111339281A

  • Answer determination method and device based on AI recognition, electronic equipment and medium

    CN111680515A