Scoring method, training method, computer device and storage medium for oral Q&A
By combining speech recognition, semantic extraction and acoustic model's feature vectors for scoring, the problems of accuracy and fairness of oral question-and-answer scores in the prior art are solved, and more accurate and fair scoring results are achieved.
Patent Information
- Application Number
- CN202111618214.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-12-27
AI Technical Summary
The existing oral Q&A scoring schemes have limitations in accuracy, especially due to the impact of speech recognition errors on the scoring effect, which leads to unfair and stable scores.
By obtaining the question text and the audio of students' answers, using the preset speech recognition model, semantic extraction model and acoustic model, the speech feature vector, text feature vector and acoustic feature vector were extracted respectively, and the fusion score was performed based on the scoring model to improve the accuracy and fairness of the scoring.
This method can alleviate the impact of speech recognition errors on scores, improve the accuracy and stability of scores, and make the scores more fair and just.
Smart Images

Figure CN114360537B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method for scoring oral answers, a training method, a computer device, and a storage medium. Background Art
[0002] With the deep reform of the education system, more and more attention is paid to the improvement of students' oral language level, such as their English oral language level. Oral examinations have been introduced into the high school entrance examination and college entrance examination in many regions. Among them, question-and-answer, as the most representative mode of daily communication and language exchange, is frequently examined. At the same time, with the increasing maturity of artificial intelligence-related technologies, machine grading assistance has gradually become the mainstream. However, the question-and-answer questions are relatively open, and not only pronunciation information but also semantic understanding need to be considered, making it challenging for machines to score fairly and objectively.
[0003] The existing oral answer scoring scheme is as follows: a speech recognition model is used to recognize the student's answer audio, and then the similarity between the oral recognition text and the artificially preset standard answer is calculated. In addition, the speech is segmented according to the oral recognition text, and phoneme pronunciation features such as GOP (goodness of pronunciation) are calculated based on the segmentation boundaries. Then, the extracted speech and semantic features are regressed to obtain the corresponding score for the student. However, this scheme has certain limitations and low accuracy. Summary of the Invention
[0004] The embodiments of the present application provide a method for scoring oral answers, a training method, a computer device, and a storage medium, which can more accurately implement the scoring of oral answers.
[0005] In a first aspect, the present application provides a method for scoring oral answers, the method comprising:
[0006] Obtain a question text;
[0007] Obtain the answer audio of the student's answer;
[0008] Input the answer audio into a preset speech recognition model to obtain the speech feature vector corresponding to the answer audio and the oral recognition text;
[0009] Input the oral recognition text and the question text into a preset semantic extraction model to obtain the text feature vector corresponding to the answer audio;
[0010] Input the answer audio into a preset acoustic model to obtain the acoustic feature vector corresponding to the answer audio;
[0011] Based on a preset scoring model, obtain the predicted score corresponding to the answer audio according to the speech feature vector, the text feature vector, and the acoustic feature vector corresponding to the answer audio.
[0012] In a second aspect, the present application provides a method for training a spoken Q&A scoring model, where the spoken Q&A scoring model includes a speech recognition model, a semantic extraction model, an acoustic model, and a scoring model;
[0013] The training method includes:
[0014] Obtain the question text;
[0015] Obtain the response audio of the student's answer and the corresponding annotated score;
[0016] Input the response audio into a pre-trained speech recognition model to obtain the speech feature vector and the spoken recognition text corresponding to the response audio;
[0017] Input the spoken recognition text and the question text into a pre-trained semantic extraction model to obtain the text feature vector corresponding to the response audio;
[0018] Input the response audio into a pre-trained acoustic model to obtain the acoustic feature vector corresponding to the response audio;
[0019] Based on a preset scoring model, obtain the predicted score corresponding to the response audio according to the speech feature vector, text feature vector, and acoustic feature vector corresponding to the response audio;
[0020] Adjust the model parameters of at least one of the speech recognition model, semantic extraction model, acoustic model, and scoring model according to the predicted score corresponding to the response audio and the annotated score.
[0021] In a third aspect, the present application provides a computer device, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the steps of any one of the above methods when executing the computer program.
[0022] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. If the computer program is executed by a processor, the steps of any one of the above methods are implemented.
[0023] The present application discloses a scoring method, a training method, a computer device, and a storage medium for oral Q&A. The scoring method includes: inputting the answering audio into a preset speech recognition model to obtain the speech feature vector and the oral recognition text corresponding to the answering audio; inputting the oral recognition text and the question text into a preset semantic extraction model to obtain the text feature vector corresponding to the answering audio; inputting the answering audio into a preset acoustic model to obtain the acoustic feature vector corresponding to the answering audio; and based on a preset scoring model, obtaining the predicted score corresponding to the answering audio according to the speech feature vector, the text feature vector, and the acoustic feature vector corresponding to the answering audio. By extracting the speech feature vector through the speech recognition model, the text feature vector through the semantic extraction model, the acoustic feature vector through the acoustic model, and finally scoring according to the extracted multiple vectors, the influence of speech recognition errors on the scoring effect in the existing oral Q&A scoring scheme can be alleviated or avoided, thereby making the scoring more fair and stable. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 It is a schematic flowchart of the scoring method for oral Q&A according to an embodiment of the present application;
[0026] Figure 2 It is a schematic diagram of the application scenario of the scoring method in an embodiment;
[0027] Figure 3 It is a schematic diagram of scoring by the oral Q&A scoring model in an embodiment;
[0028] Figure 4 It is a schematic diagram of extracting the speech feature vector in an embodiment;
[0029] Figure 5 It is a schematic diagram of extracting the text feature vector in an embodiment;
[0030] Figure 6 It is a schematic diagram of extracting the acoustic feature vector in an embodiment;
[0031] Figure 7 It is a schematic flowchart of the training method of the oral Q&A scoring model according to another embodiment of the present application;
[0032] Figure 8 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0034] The flowchart shown in the accompanying drawings is only an example, and does not necessarily include all contents and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined or partially merged. Therefore, the actual execution order may be changed according to the actual situation.
[0035] Embodiments of the present application provide a method for scoring oral answers, a training method, a computer device, and a storage medium to achieve scoring based on the question text to be answered and the audio of the student's answer, that is, scoring oral answers.
[0036] Currently, for scoring oral answers, a speech recognition model is usually used to recognize the audio of the student's answer, and then the similarity between the recognized text of the speech and the artificially preset standard answer is calculated. Among the measurement indicators related to similarity, there are the longest common substring (lcs), BLEU (Bilingual Evaluation Understudy), METEOR (Metric for Evaluation of Translation with Explicit Ordering), etc. Then, the speech is segmented according to the recognized text of the speech, and phoneme pronunciation features such as GOP (goodness of pronunciation) are calculated according to the segmentation boundaries. Then, the extracted speech and semantic features are regressed to obtain the corresponding score of the student.
[0037] Currently, the scoring of oral answers has at least one of the following defects:
[0038] First, this solution is divided into three steps. First, speech recognition is performed to obtain the text, then the similarity is calculated using the obtained text, and finally all features are regressed for scoring. This strategy will produce cascading errors. The errors in speech recognition have a greater impact on text feature extraction, and the errors in feature extraction also have a greater impact on scoring.
[0039] Second, it is difficult to meet the requirements in text similarity calculation. Many slight changes in the students' answers will cause semantic changes, resulting in score differences. The existing text similarity is difficult to reflect the impact brought by such slight perturbations.
[0040] III. For questions with open answers, it is difficult to enumerate standard answers by manpower and rules. Therefore, existing solutions are difficult to give fair scores to some open answers, and excellent divergent answers are often judged as low scores.
[0041] IV. The speech recognition training data lacks domain adaptation for oral scenarios, which has certain limitations in recognizing the audio of middle school students speaking a second foreign language, and it is difficult to accurately recognize common grammar errors and pronunciation flaws.
[0042] Based on this, the inventors of the present application improved the scoring method for oral Q&A to solve at least one of the above defects.
[0043] The scoring method for oral Q&A provided in the embodiments of the present application can be applied to a terminal or a server. The terminal can be an electronic device such as a mobile phone, a tablet computer, a laptop computer, a desktop computer, a personal digital assistant, etc.; the server can be an independent server or a server cluster. For the convenience of understanding, the following embodiments will introduce the method applied to the server in detail.
[0044] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0045] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a scoring method for oral Q&A provided by an embodiment of the present application.
[0046] In some embodiments, as Figure 2 shown, the server obtains the question text and the answer audio of the student's answer from the terminal, generates a predicted score corresponding to the answer audio according to the scoring method for oral Q&A, and sends the predicted score to the terminal.
[0047] In some optional embodiments, the question text is the text stored locally by the device for implementing the scoring method for oral Q&A, the text obtained by the device from the network, the text obtained by the device from the input device connected thereto, the text obtained by the device from other electronic devices, the text converted by the device according to voice information, etc. The answer audio is the audio stored locally by the device for implementing the scoring method for oral Q&A, the audio obtained by the device from the network, the audio obtained by the device from the input device connected thereto, the audio obtained by the device from other electronic devices, etc.
[0048] Please refer to Figure 1 and Figure 3 , the scoring method for oral Q&A includes the following steps S110 to step S160.
[0049] Step S110: Obtain the question text.
[0050] For example, the problem text is represented as a = {a 1, a 2, …a m,}, where m is the number of words in the problem text.
[0051] Step S120: Obtain the response audio of the student's answer.
[0052] Exemplarily, when starting to answer, the student triggers the start of recording through operations such as pressing a key, and triggers the end of recording after answering, to obtain the response audio. Of course, it is not limited to this. For example, it can start and end recording according to the voice commands of the student during answering, such as "start" and "answer completed", to obtain the response audio.
[0053] Step S130: Input the response audio into a preset speech recognition model to obtain the speech feature vector corresponding to the response audio and the spoken language recognition text.
[0054] Exemplarily, speech recognition is also known as Automatic Speech Recognition (ASR), which can convert human speech into text.
[0055] In some embodiments, the speech recognition model can extract the speech feature vector from the response audio, and recognize the spoken language recognition text according to the speech feature vector. Exemplarily, perform speech recognition on the response audio to obtain the spoken language recognition text, and take the intermediate vector in the speech recognition process as the speech feature vector.
[0056] Exemplarily, the speech recognition model includes an Encoder sub-model and a Decoder sub-model, and the speech recognition model can be an Encoder-Decoder model. The Encoder sub-model is used to encode the response audio to obtain the speech feature vector, and the speech feature vector is a hidden layer vector; the Decoder sub-model is used to decode the speech feature vector to obtain the spoken language recognition text.
[0057] Optionally, as Figure 4 shown, the vector obtained by the Encoder sub-model encoding the response audio can also be processed by a preset neural network, such as a fully connected process, and the vector after the neural network processing is used as the speech feature vector. Processing the vector obtained by the Encoder sub-model encoding the response audio through a neural network can extract more useful information and improve the accuracy of scoring.
[0058] In some embodiments, please refer to Figure 4, step S130 inputs the response audio into a preset speech recognition model to obtain a speech feature vector and a spoken language recognition text corresponding to the response audio, including the following steps S131 to S133.
[0059] Step S131: Input the response audio into the first feature extraction model of the speech recognition model to obtain the first speech feature of the response audio; Step S132: Input the first speech feature into the encoder sub-model of the speech recognition model to obtain a speech feature vector corresponding to the response audio; Step S133: Input the speech feature vector into the decoder sub-model of the speech recognition model to obtain a spoken language recognition text corresponding to the response audio.
[0060] Exemplarily, the first feature extraction model is used to extract the fbank feature of the response audio, that is, the first speech feature can be the fbank feature. The steps for obtaining the fbank feature of the speech signal generally include the following steps: pre-emphasis, framing, windowing, short-time Fourier transform (STFT), mel filtering, mean removal, etc. For example, when extracting the fbank feature, the feature window length is 25 ms, the frame shift is 10 ms, and the dimension of the first speech feature is T×40 dimensions, where T is the number of frames of the response audio.
[0061] Exemplarily, the encoder sub-model includes multiple layers, such as 12 layers of transformers (multi-head self-attention networks). The output of the upper layer of the Transformer will be used as the input of the lower layer of the Transformer. Each layer of the Transformer uses the multi-head attention mechanism to process the vector.
[0062] The input of the encoder sub-model is the fbank feature of T×40 dimensions, and the output of the encoder sub-model is a vector of T×S dimensions. S is the hidden layer vector dimension of the transformer layer. The output of the last layer of the transformer of the encoder sub-model can be used as the speech feature vector corresponding to the response audio to participate in the subsequent scoring task.
[0063] Exemplarily, the decoder sub-model includes multiple layers, such as 6 layers of transformers, and decodes the speech feature vector to obtain the spoken language recognition text.
[0064] In some embodiments, the speech recognition model is trained based on pre-training data containing a large amount of high-score domestic student spoken language reading data, and thus has a strong modeling ability for the characteristics of student spoken language.
[0065] Exemplarily, for a speech recognition model in the field of spoken language, in the pre-training stage, a large number of high-score reading questions are first selected for relatively rough bottom training, and then a large amount of finely annotated spoken language transcription data is selected for training. In the training stage, the speech recognition model is first used to recognize the text, and then the output of the last layer of the encoder is used as the speech feature vector to participate in the subsequent scoring task.
[0066] For example, the spoken language recognition text corresponding to the response audio is represented as g = {g 1, g 2, …g l,}, where l is the number of words in the spoken language recognition text.
[0067] Step S140: Input the spoken language recognition text and the question text into a preset semantic extraction model to obtain the text feature vector corresponding to the response audio.
[0068] The text feature vector is used to indicate the semantic features of the spoken language recognition text and the question text. In some embodiments, the semantic extraction model can be referred to as a text feature extraction module, which inputs the spoken language recognition text and the question text and outputs the extracted text feature vector, that is, the semantic feature.
[0069] In some embodiments, step S140 inputs the spoken language recognition text and the question text into a preset semantic extraction model to obtain the text feature vector corresponding to the response audio, including the following steps S141 to step S142.
[0070] Step S141: Input a preset start character (such as [CLS]), the spoken language recognition text, the question text, and a preset interval character (such as [SEP]) between the spoken language recognition text and the question text into the embedding sub-model of the semantic extraction model to obtain an embedding vector.
[0071] By splicing the spoken language recognition text and the question text to obtain a target text, the spoken language recognition text is used as the first part of the target text, and the start character, such as [CLS], is added at the beginning of the sentence; the question text is used as the second part of the target text, and an interval character, such as [SEP], is used to fill between the first part and the second part.
[0072] Exemplarily, please refer to Figure 5 , the question text is represented as a = {a 1, a 2, …a m,}, and the spoken language recognition text is represented as g = {g 1, g 2, …g l,, plus the start character and the separator character, the number of characters in the target text is l + m + 2. First, each character in the target text is converted into a character vector. The embedding vector includes l + m + 2 character vectors, for example, denoted as H 0 =[e 1 , e 2 , …, e l+m+2 , where e i represents the character vector of the i-th character in the target text, where the first character vector e 1 is the character vector of the start character, and the (l + 2)-th character vector e l+2 is the character vector of the separator character. Optionally, an end character can also be added after the question text. The end character can also be [SEP] for example. Then the embedding vector includes l + m + 3 character vectors. Optionally, the target text can be represented as [a; q].
[0073] Step S142: Input the embedding vector into the multi-head self-attention sub-model of the semantic extraction model to obtain the corresponding text feature vector.
[0074] In some embodiments, the multi-head self-attention sub-model is a BERT model (Bidirectional Encoder Representations from Transformers). Optionally, the multi-head self-attention sub-model is a pre-trained BERT model. Of course, the multi-head self-attention sub-model is not limited to the BERT model. For example, it can also be a recurrent neural network model, a convolutional neural network model, or a combination of multiple models / networks. The BERT model is a bidirectional language representation model. Using the Transformer network (a neural network based on self-attention) as the unit module, it is pre-trained on a large-scale corpus using two upstream tasks: masked language model (MLM) and next sentence prediction (NSP). Compared with recurrent neural networks and convolutional neural networks, it has stronger semantic modeling capabilities and only needs simple fine-tuning on downstream tasks to achieve good results. Using pre-trained multi-head self-attention sub-models such as the BERT model can perform deeper semantic modeling.
[0075] Please refer to Figure 5, where the coherence classification of the upper and lower sentences is one of the upstream tasks of the BERT model, which is used to predict whether the second sentence is a follow-up of the first sentence. In common QA (Question Answering) tasks, the question Q and the answer A are often concatenated as the input of the BERT model, and the BERT model predicts whether the answer A can answer the question Q. In the embodiments of the present application, the BERT model is used to perform semantic modeling on the target text to obtain the text feature vectors corresponding to the speech recognition text and the question text.
[0076] Exemplarily, a pre-trained Bert model is used to extract text features to characterize the coherence between the question text and the speech recognition text.
[0077] The BERT model is a deep language model with 12 layers of Transformer networks (multi-head self-attention networks) and without shared parameters. The output of the upper layer of the Transformer network will be used as the input of the lower layer of the Transformer network. The output of the last layer of the Transformer network can be expressed as follows:
[0078]
[0079] Where is the vector corresponding to the starting character [CLS], is the vector corresponding to the first character in the speech recognition text is the vector corresponding to the separator character [SEP], is the vector corresponding to the last character in the question text
[0080] The output output of the last layer of the Transformer network 12 contains the hidden layer representations of all characters in the target text. Exemplarily, after pre-extracting the speech recognition text and the question text through the 12 layers of Transformer networks of the pre-trained Bert model, the last layer of the Transformer network outputs a vector with a dimension of (l + m + 2) × 768, and this vector can be used as the text feature vector corresponding to the speech recognition text and the question text, of course, it is not limited to this.
[0081] Exemplarily, please refer to Figure 5 , the hidden layer representation of the starting character [CLS] can be selected as the text content representation (or can be called semantic representation) corresponding to all characters in the starting character, the key point text, the separator character, and the candidate's answer text, that is, the hidden layer representation of the text vector.
[0082] Optionally, please refer to Figure 5, the text feature vectors corresponding to the spoken language recognition text and the question text at least include output 12 and the vector corresponding to the starting character [CLS] in 12 . Optionally, it may also include output 12 and the average value of the vectors corresponding to the other characters except the starting character in 12 : MEAN(output 12 [1:]). That is, the text feature vectors corresponding to the spoken language recognition text and the question text can be expressed as follows:
[0083] Hidden = [output 12 [0]; MEAN(output 12 [1:])]
[0084] In some embodiments, the semantic extraction model is pre-trained in two stages. In the first stage, a large number of adversarial samples are generated in an unsupervised manner, and the semantic extraction model is pre-trained with texts with slight changes, so that the semantic extraction model can take into account both sensitivity and robustness for spoken language data. Exemplarily, a first training sample is obtained, and the first training sample includes the spoken language recognition text; based on a preset data augmentation rule, the spoken language recognition text is augmented to obtain the augmented first training sample; the semantic extraction model is pre-trained according to the first training sample. Among them, the training data includes the spoken language recognition texts of students, and the spoken language recognition texts are augmented. The data augmentation method is as follows: the part-of-speech of the corpus is recognized, and specific words are replaced with synonyms and antonyms, and the replacement ratio is, for example, 40%, so as to generate samples with the same semantics. The pre-training method of the semantic extraction model is, for example, consistent with the pre-training method of the original Bert model.
[0085] The training of the second stage of the semantic extraction model can use a large number of dialogue samples to model the dialogue coherence evaluation, so as to not rely on manual answer formulation for the questions to be answered, and achieve the effect of fair scoring. Exemplarily, a second training sample is obtained, and the second training sample includes the question text and the corresponding answer text; according to the second training sample, the semantic extraction model pre-trained based on the first training sample is pre-trained to obtain the pre-trained semantic extraction model. The second stage is pre-training for question-answer coherence. The data corpus is question-answer data, the related question-answer pairs are positive samples, and randomly sampled question-answer pairs are negative samples, and then training is carried out.
[0086] Exemplarily, through the above two-stage pre-training of the semantic extraction model, the robustness and sensitivity of the model to the semantic understanding of the spoken language recognition text are increased. At the same time, through a large number of spoken language question-answer data pre-training, it has better modeling for short discourse coherence.
[0087] Step S150: Input the response audio into a preset acoustic model to obtain an acoustic feature vector corresponding to the response audio.
[0088] The speech recognition model in step S130 pays more attention to the speech content, and the identified features lack information related to scoring dimensions such as prosody and timbre standardization. By using the acoustic model to extract acoustic features from the response audio, the extracted acoustic features can indicate the original acoustic information of the response audio, making the scoring information more diverse.
[0089] In some embodiments, refer to Figure 6 , step S150 inputting the response audio into a preset acoustic model to obtain an acoustic feature vector corresponding to the response audio includes the following steps S151 to S152.
[0090] Step S151: Based on the second feature extraction model of the acoustic model, obtain the second speech feature of the response audio; step S152: Input the second speech feature into the acoustic information extraction sub-model of the acoustic model to obtain an acoustic feature vector corresponding to the response audio.
[0091] In some embodiments, the second feature extraction model is used to extract MFCC (Mel-Frequency Cepstral Coefficients) features of the response audio, that is, the second speech feature can be MFCC features. Exemplarily, the second feature extraction model can perform DCT (discrete cosine transform) cepstral processing on the fbank features obtained by the first feature extraction model to obtain the MFCC features. The essence of DCT is to remove the correlation between each dimension of the signal and map the signal to a low-dimensional space, making the MFCC features have better discriminability.
[0092] Exemplarily, refer to Figure 6 , the acoustic information extraction sub-model of the acoustic model includes 3 convolutional modules and 3 bidirectional LSTM (Long short-term memory) modules to extract features from the second speech feature, and finally output an acoustic feature vector. The acoustic feature vector is, for example, a vector with a dimension of T×S. Optionally, the vector dimension of the acoustic feature vector output by the acoustic model is the same as that of the speech feature vector obtained in step S130.
[0093] In some embodiments, the speech recognition model and the acoustic model can form a speech processing module, which inputs the student audio and obtains the recognized text, text feature vector, and speech feature vector.
[0094] Step S160: Based on a preset scoring model, obtain the predicted score corresponding to the response audio according to the speech feature vector, text feature vector, and acoustic feature vector corresponding to the response audio.
[0095] The scoring model can also be referred to as a feature fusion and scoring module. This module takes the speech feature vector, text feature vector, and acoustic feature vector as inputs and outputs the final score. By fusing the speech feature vector, text feature vector, and acoustic feature vector, the scoring process fully considers the speech information, text semantic information, and acoustic features, integrating speech recognition, semantic understanding, and acoustic features, making the scoring accuracy higher.
[0096] In some embodiments, the speech recognition model, semantic extraction model, acoustic model, and scoring model can form an oral Q&A scoring model.
[0097] The scoring method of the embodiments of the present application extracts the speech feature vector through the speech recognition model, the text feature vector through the semantic extraction model, the acoustic feature vector through the acoustic model, and finally scores according to the extracted multiple vectors, which can alleviate or avoid the influence of speech recognition errors on the scoring effect in the existing oral Q&A scoring scheme, making the scoring more fair and stable.
[0098] In some embodiments, the acoustic model, speech recognition model, and semantic extraction model are pre-trained using a large amount of student oral data, enabling each model to adapt to the student oral domain, thereby making the scoring more fair and stable.
[0099] In some embodiments, calibration data can be used to fine-tune the entire oral Q&A scoring model during the formal exam, thereby realizing the joint training of the speech recognition model, semantic extraction model, and scoring model, alleviating the cascading error, and reducing the influence of speech recognition on text feature extraction.
[0100] Exemplarily, the training process of the oral Q&A scoring model is divided into two stages. The first stage is modular pre-training, using a large amount of data in the oral Q&A field to pre-train the speech recognition model and semantic extraction model. The second stage is fine-tuning in the calibration scenario. Given a small amount of student responses and scores, the semantic extraction model, the encoder sub-model of the speech recognition model, and the acoustic model are fine-tuned using the backpropagation algorithm to obtain the trained oral Q&A scoring model. By jointly training each model in the oral Q&A scoring model, the cascading error can be alleviated, such as alleviating the influence of speech recognition errors on the scoring effect, resulting in a better scoring effect. It can be understood that the training of the semantic extraction model, in addition to pre-training, can also include joint training with the speech recognition model and acoustic model during the calibration fine-tuning stage.
[0101] In some embodiments, the method for training a spoken Q&A scoring model includes calibration and fine-tuning of the spoken Q&A scoring model, which may include the following steps: obtaining question text; obtaining the response audio of the student's answer and the corresponding annotated score; inputting the response audio into a pre-trained speech recognition model to obtain the speech feature vector and the spoken recognition text corresponding to the response audio; inputting the spoken recognition text and the question text into a pre-trained semantic extraction model to obtain the text feature vector corresponding to the response audio; inputting the response audio into a pre-trained acoustic model to obtain the acoustic feature vector corresponding to the response audio; based on a preset scoring model, obtaining the predicted score corresponding to the response audio according to the speech feature vector, text feature vector, and acoustic feature vector corresponding to the response audio; and adjusting the model parameters of at least one of the speech recognition model, semantic extraction model, acoustic model, and scoring model according to the predicted score and the annotated score corresponding to the response audio.
[0102] Exemplarily, based on a preset loss function, determining a model loss value according to the predicted score and the annotated score corresponding to the response audio, and adjusting the model parameters of at least one of the speech recognition model, semantic extraction model, acoustic model, and scoring model according to the model loss value.
[0103] The scoring method for spoken Q&A provided by the embodiments of the present application includes: inputting the response audio into a preset speech recognition model to obtain the speech feature vector and the spoken recognition text corresponding to the response audio; inputting the spoken recognition text and the question text into a preset semantic extraction model to obtain the text feature vector corresponding to the response audio; inputting the response audio into a preset acoustic model to obtain the acoustic feature vector corresponding to the response audio; and based on a preset scoring model, obtaining the predicted score corresponding to the response audio according to the speech feature vector, text feature vector, and acoustic feature vector corresponding to the response audio. By respectively extracting the speech feature vector through the speech recognition model, the text feature vector through the semantic extraction model, the acoustic feature vector through the acoustic model, and finally scoring according to the extracted multiple vectors, the influence of speech recognition errors on the scoring effect in the existing spoken Q&A scoring scheme can be alleviated or avoided, thereby making the scoring more fair and stable.
[0104] In some embodiments, in the face of open-ended questions with high divergence, it is difficult to ensure the fairness of scoring by using the traditional method of manually making answers and matching the similarity with the student's answers. However, the scoring method for spoken Q&A in the embodiments of the present application can well model the coherence of the conversation and improve the accuracy of scoring by extracting the text feature vectors of the spoken recognition text and the question text based on the semantic extraction model.
[0105] In some embodiments, in view of the characteristics of high text sensitivity, such as the phenomenon that a slight change in words leads to a huge change in the meaning of a sentence, the existing text representation is difficult to solve well. The scoring method for oral Q&A in the embodiments of the present application uses a pre-trained Bert model to extract the text feature vectors of the oral recognition text and the question text, and adds a variety of interference data in the pre-training stage, thereby increasing the sensitivity of the pre-trained Bert model to language changes and improving the accuracy of scoring.
[0106] Please refer to the foregoing embodiments Figure 7 , and the embodiments of the present application further provide a training method for an oral Q&A scoring model. Among them, as Figure 3 shown, the oral Q&A scoring model includes a speech recognition model, a semantic extraction model, an acoustic model, and a scoring model.
[0107] Please refer to Figure 7 , and the training method includes steps S210 to S250.
[0108] Step S210: Obtain the question text;
[0109] Step S220: Obtain the response audio of the student's answer and the corresponding marked score;
[0110] Step S230: Input the response audio into the pre-trained speech recognition model to obtain the speech feature vector corresponding to the response audio and the oral recognition text;
[0111] Step S240: Input the oral recognition text and the question text into the pre-trained semantic extraction model to obtain the text feature vector corresponding to the response audio;
[0112] Step S250: Input the response audio into the pre-trained acoustic model to obtain the acoustic feature vector corresponding to the response audio;
[0113] Step S260: Based on a preset scoring model, obtain the predicted score corresponding to the response audio according to the speech feature vector, text feature vector, and acoustic feature vector corresponding to the response audio;
[0114] Step S270: Adjust the model parameters of at least one of the speech recognition model, semantic extraction model, acoustic model, and scoring model according to the predicted score corresponding to the response audio and the marked score.
[0115] In some embodiments, the training method further includes:
[0116] Obtain a first training sample, where the first training sample includes an oral recognition text;
[0117] Based on preset data augmentation rules, perform data augmentation on the speech recognition text to obtain a first training sample after data augmentation;
[0118] Pre-train the semantic extraction model according to the first training sample;
[0119] Obtain a second training sample, where the second training sample includes a question text and a corresponding answer text;
[0120] According to the second training sample, pre-train the semantic extraction model pre-trained based on the first training sample to obtain the pre-trained semantic extraction model.
[0121] Exemplarily, the semantic extraction model undergoes two-stage pre-training. In the first stage, a large number of adversarial samples are generated in an unsupervised manner, and the semantic extraction model is pre-trained with texts with slight changes, so that the semantic extraction model can take into account both sensitivity and robustness to spoken language data. The training in the second stage can use a large number of dialogue samples for dialogue coherence evaluation modeling, so as to not rely on manual answer formulation for the questions to be answered, achieving the effect of fair scoring.
[0122] In some embodiments, the step of inputting the response audio into the pre-trained speech recognition model to obtain the speech feature vector and the speech recognition text corresponding to the response audio includes:
[0123] Input the response audio into the first feature extraction model of the speech recognition model to obtain the first speech feature of the response audio;
[0124] Input the first speech feature into the encoder sub-model of the speech recognition model to obtain the speech feature vector corresponding to the response audio;
[0125] Input the speech feature vector into the decoder sub-model of the speech recognition model to obtain the speech recognition text corresponding to the response audio.
[0126] The specific principle and implementation manner of the training method of the spoken language question-and-answer scoring model provided by the embodiments of the present application are similar to those of the scoring method of the spoken language question-and-answer in the foregoing embodiments, and will not be elaborated here.
[0127] The method of the present application can be used in many general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0128] Exemplarily, the above method can be implemented in the form of a computer program, which can run on a computer device as shown in Figure 8 below.
[0129] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device can be a server or a terminal.
[0130] Referring to Figure 8 , the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.
[0131] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which when executed, can cause the processor to execute the steps of any of the foregoing methods.
[0132] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0133] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can be caused to execute the steps of any of the foregoing methods.
[0134] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that the structure of the computer device is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0135] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0136] Among them, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:
[0137] Obtain the question text;
[0138] Obtain the response audio of the student's answer;
[0139] Input the response audio into a preset speech recognition model to obtain the speech feature vector and the spoken language recognition text corresponding to the response audio;
[0140] Input the spoken language recognition text and the question text into a preset semantic extraction model to obtain the text feature vector corresponding to the response audio;
[0141] Input the response audio into a preset acoustic model to obtain the acoustic feature vector corresponding to the response audio;
[0142] Based on a preset scoring model, obtain the predicted score corresponding to the response audio according to the speech feature vector, text feature vector, and acoustic feature vector corresponding to the response audio.
[0143] Among them, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:
[0144] Obtain the question text;
[0145] Obtain the response audio of the student's answer and the corresponding marked score;
[0146] Input the response audio into a pre-trained speech recognition model to obtain the speech feature vector and the spoken language recognition text corresponding to the response audio;
[0147] Input the spoken language recognition text and the question text into a pre-trained semantic extraction model to obtain the text feature vector corresponding to the response audio;
[0148] Input the response audio into a pre-trained acoustic model to obtain the acoustic feature vector corresponding to the response audio;
[0149] Based on a preset scoring model, obtain the predicted score corresponding to the response audio according to the speech feature vector, text feature vector, and acoustic feature vector corresponding to the response audio;
[0150] According to the predicted score corresponding to the response audio and the marked score, adjust the model parameters of at least one of the speech recognition model, semantic extraction model, acoustic model, and scoring model.
[0151] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application, such as:
[0152] A computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement the steps of any one of the oral answer scoring methods provided in the embodiments of this application.
[0153] Among them, the computer-readable storage medium can be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device.
[0154] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
Claims
1. A method for scoring oral Q&A, characterized in that, it includes: Obtain the question text; Obtain the response audio of the student's answer; Input the response audio into a preset speech recognition model to obtain the speech feature vector and the oral recognition text corresponding to the response audio; Input the oral recognition text and the question text into a preset semantic extraction model to obtain the text feature vector corresponding to the response audio; Input the response audio into a preset acoustic model to obtain the acoustic feature vector corresponding to the response audio; Based on a preset scoring model, according to the speech feature vector, text feature vector, and acoustic feature vector corresponding to the response audio, obtain the predicted score corresponding to the response audio; Among them, the step of inputting the oral recognition text and the question text into a preset semantic extraction model to obtain the text feature vector corresponding to the response audio includes: Input the target text obtained by splicing the oral recognition text and the question text into the embedding sub-model of the semantic extraction model to obtain an embedding vector; Input the embedding vector into the multi-head self-attention sub-model of the semantic extraction model to obtain the corresponding text feature vector; the text feature vector is a semantic feature used to represent the coherence between the question text and the oral recognition text.
2. The scoring method according to claim 1, characterized in that, the step of inputting the response audio into a preset speech recognition model to obtain the speech feature vector and the oral recognition text corresponding to the response audio includes: Input the response audio into the first feature extraction model of the speech recognition model to obtain the first speech feature of the response audio; Input the first speech feature into the encoder sub-model of the speech recognition model to obtain the speech feature vector corresponding to the response audio; Input the speech feature vector into the decoder sub-model of the speech recognition model to obtain the oral recognition text corresponding to the response audio.
3. The scoring method according to claim 2, characterized in that, the step of inputting the response audio into a preset acoustic model to obtain the acoustic feature vector corresponding to the response audio includes: Based on the second feature extraction model of the acoustic model, obtain the second speech feature of the response audio; Input the second speech feature into the acoustic information extraction sub-model of the acoustic model to obtain the acoustic feature vector corresponding to the response audio.
4. The scoring method according to claim 3, characterized in that, the first speech feature is the fbank feature, and the second speech feature is the MFCC feature.
5. The scoring method according to any one of claims 1-4, characterized in that, the step of inputting the target text obtained by splicing the oral recognition text and the question text into the embedding sub-model of the semantic extraction model to obtain an embedding vector includes: Input a preset start character, the oral recognition text, the question text, and a preset interval character between the oral recognition text and the question text into the embedding sub-model of the semantic extraction model to obtain an embedding vector.
6. A method for training an oral Q&A scoring model, characterized in that, The spoken Q&A scoring model includes a speech recognition model, a semantic extraction model, an acoustic model, and a scoring model; The training method includes: Obtain question texts; Obtain the response audio of the student's answer and the corresponding annotated score; Input the response audio into a pre-trained speech recognition model to obtain the speech feature vector and the spoken recognition text corresponding to the response audio; Input the spoken recognition text and the question text into a pre-trained semantic extraction model to obtain the text feature vector corresponding to the response audio; Input the response audio into a pre-trained acoustic model to obtain the acoustic feature vector corresponding to the response audio; Based on a preset scoring model, obtain the predicted score corresponding to the response audio according to the speech feature vector, text feature vector, and acoustic feature vector corresponding to the response audio; Adjust the model parameters of at least one of the speech recognition model, semantic extraction model, acoustic model, and scoring model according to the predicted score corresponding to the response audio and the annotated score; Wherein, the step of inputting the spoken recognition text and the question text into a preset semantic extraction model to obtain the text feature vector corresponding to the response audio includes: Input the target text obtained by concatenating the spoken recognition text and the question text into the embedding sub-model of the semantic extraction model to obtain an embedding vector; Input the embedding vector into the multi-head self-attention sub-model of the semantic extraction model to obtain the corresponding text feature vector; the text feature vector is a semantic feature used to characterize the coherence between the question text and the spoken recognition text.
7. The training method according to claim 6, wherein, The training method further includes: Obtain a first training sample, where the first training sample includes a spoken recognition text; Based on a preset data augmentation rule, perform data augmentation on the spoken recognition text to obtain a data-augmented first training sample; Pre-train the semantic extraction model according to the first training sample; Obtain a second training sample, where the second training sample includes a question text and the corresponding answer text; According to the second training sample, pre-train the semantic extraction model pre-trained based on the first training sample to obtain the pre-trained semantic extraction model.
8. The training method according to claim 6 or 7, wherein, The step of inputting the response audio into a pre-trained speech recognition model to obtain the speech feature vector and the spoken recognition text corresponding to the response audio includes: Input the response audio into the first feature extraction model of the speech recognition model to obtain the first speech feature of the response audio; Input the first speech feature into the encoder sub-model of the speech recognition model to obtain the speech feature vector corresponding to the response audio; Input the speech feature vector into the decoder sub-model of the speech recognition model to obtain the spoken recognition text corresponding to the response audio.
9. A computer device, wherein, The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and, when executing the computer program, implement: the steps of the method for scoring a spoken language Q&A according to any one of claims 1-5; and / or the steps of the method for training a spoken language Q&A scoring model according to any one of claims 6-8.
10. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements: the steps of the method for scoring a spoken language Q&A according to any one of claims 1-5; and / or the steps of the method for training a spoken language Q&A scoring model according to any one of claims 6-8.
Citation Information
Patent Citations
Question-answer scoring method, device and equipment based on artificial intelligence, and storage medium
CN110797010A
Voice processing method and device, electronic equipment and computer readable storage medium
CN111833853A
Semantic recognition method and device, computer equipment and storage medium
CN112232086A