Method and apparatus for training a spoken language evaluation model
By employing a training method for oral assessment models that incorporates multi-task learning and loss function optimization, the challenge of assessment without reference text is addressed. This approach achieves efficient and accurate oral scoring and grade probability prediction, thereby enhancing the comprehensiveness and precision of the assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING YUANLI WEILAI SCI & TECH CO LTD
- Filing Date
- 2021-11-01
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies cannot achieve text-free assessment in oral language evaluation, and ordered regression methods are computationally intensive, affecting the accuracy and efficiency of scoring.
A multi-task learning approach is adopted, which trains the oral assessment model by acquiring sample audio features and combining mean square loss and cross-entropy loss functions. The model is then optimized to perform oral scoring without reference text, thereby achieving multi-dimensional evaluation and grade probability prediction.
It improves the accuracy and interpretability of oral assessment, enables probabilistic assessment at multiple levels, and enhances the model's prediction accuracy and the comprehensiveness of the scoring.
Smart Images

Figure CN116072150B_ABST
Abstract
Description
Technical Field
[0001] This manual relates to the field of audio processing technology, and in particular to training methods and devices for spoken language assessment models. Background Technology
[0002] With the development of internet technology, speech recognition technology is being applied in more and more fields. In spoken language assessment scenarios, computers typically automatically evaluate the overall pronunciation of a sentence read aloud by a user. However, the content read by the user often requires corresponding reference text, making it impossible to achieve text-free spoken language scoring. Furthermore, while scoring fluency based on ordered regression can achieve text-free assessment, this method requires pairing the scoring audio with all anchor audio during inference to determine fluency and obtain the final score. This process not only requires pre-pairing but also involves significant computational costs. Therefore, an effective solution is urgently needed to address these issues. Summary of the Invention
[0003] In view of this, embodiments of this specification provide a training method for an oral language assessment model. This specification also relates to a training apparatus for an oral language assessment model, an oral language assessment method, an oral language assessment device, a computing device, and a computer-readable storage medium, to address the technical deficiencies existing in the prior art.
[0004] According to a first aspect of the embodiments of this specification, a method for training a spoken language assessment model is provided, comprising:
[0005] Obtain the sample audio and the corresponding spoken language score of the sample audio;
[0006] The audio features are determined based on the sample audio, and the audio features are input into the oral assessment model for processing to obtain the oral assessment score and the probability distribution of oral level.
[0007] A first loss value is calculated based on the assessed oral communication score and the sample oral communication score, and a second loss value is calculated based on the oral communication level probability distribution and the sample oral communication score.
[0008] The oral assessment model is tuned based on the first loss value and the second loss value, and training continues until the training stop condition is met.
[0009] Optionally, obtaining the sample audio and the corresponding spoken language score includes:
[0010] Obtain the sample audio;
[0011] Determine the fluency score for the fluency dimension, the pronunciation score for the pronunciation dimension, and the semantic score for the semantic dimension of the sample audio.
[0012] Calculate the average of the fluency score, the pronunciation score, and the speech score, and determine the sample spoken language score based on the calculation results.
[0013] Optionally, determining the audio features based on the sample audio includes:
[0014] Acoustic features are constructed based on the sample audio;
[0015] The acoustic features are input into a pre-trained acoustic model for processing to obtain the audio features.
[0016] Optionally, calculating the first loss value based on the assessed spoken language score and the sample spoken language score includes:
[0017] The mean squared loss function is read based on the evaluation oral score and the sample oral score;
[0018] The first loss value is obtained by calculating the evaluation oral score and the sample oral score using the mean square loss function.
[0019] Optionally, calculating the second loss value based on the spoken language level probability distribution and the sample spoken language score includes:
[0020] The cross-entropy loss function is read based on the spoken language level probability distribution and the sample spoken language scores;
[0021] The cross-entropy loss function is used to calculate the spoken language level probability distribution and the sample spoken language score to obtain a second loss value.
[0022] Optionally, tuning the oral assessment model based on the first loss value and the second loss value includes:
[0023] Determine the loss weight corresponding to the second loss value, and calculate the third loss value based on the loss weight and the second loss value;
[0024] The target loss value is obtained by summing the first loss value and the third loss value;
[0025] The oral assessment model is tuned based on the target loss value.
[0026] Optionally, after the step of obtaining the sample audio and the corresponding sample spoken score is performed, the method further includes:
[0027] The oral scores of the samples are divided to obtain oral proficiency level information;
[0028] Accordingly, the step of inputting the audio features into the spoken language assessment model for processing to obtain the spoken language level probability distribution includes:
[0029] The audio features are input into the spoken language assessment model for processing to obtain the sample spoken language level probability corresponding to the spoken language level information.
[0030] The spoken language level probability distribution corresponding to the sample audio is constructed based on the sample spoken language level probability.
[0031] According to a second aspect of the embodiments of this specification, a training apparatus for a spoken language assessment model is provided, comprising:
[0032] The acquisition module is configured to acquire sample audio and the corresponding spoken language score of the sample audio;
[0033] The processing module is configured to determine audio features based on the sample audio, and input the audio features into the oral assessment model for processing to obtain the oral assessment score and the oral level probability distribution;
[0034] The calculation module is configured to calculate a first loss value based on the assessed spoken language score and the sample spoken language score, and to calculate a second loss value based on the spoken language level probability distribution and the sample spoken language score;
[0035] The training module is configured to tune the oral assessment model based on the first loss value and the second loss value, and continue training until the training stop condition is met.
[0036] According to a third aspect of the embodiments of this specification, a spoken language assessment method is provided, comprising:
[0037] Acquire spoken audio and determine spoken audio features based on the spoken audio;
[0038] The spoken audio features are input into the spoken language assessment model in the training method of the spoken language assessment model for processing to obtain the spoken language level probability distribution;
[0039] The target spoken language score of the spoken language audio is calculated based on the spoken language level probability distribution.
[0040] Optionally, calculating the target spoken language score of the spoken language audio based on the spoken language level probability distribution includes:
[0041] Based on the spoken language level probability distribution, extract the positive probability value of each spoken language level corresponding to the spoken language audio;
[0042] The target spoken language score of the spoken language audio is obtained by summing the positive probability values for each spoken language level.
[0043] According to a fourth aspect of the embodiments of this specification, a spoken language assessment device is provided, comprising:
[0044] The audio acquisition module is configured to acquire spoken audio and determine spoken audio features based on the spoken audio.
[0045] The spoken language assessment module is configured to input the spoken language audio features into the spoken language assessment model in the training method of the spoken language assessment model for processing, and obtain the spoken language level probability distribution.
[0046] The score calculation module is configured to calculate the target spoken language score of the spoken language audio based on the spoken language level probability distribution.
[0047] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising:
[0048] Memory and processor;
[0049] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the training method or steps of the oral assessment model.
[0050] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of a training method or a speaking assessment method for a spoken language assessment model.
[0051] The oral language assessment model training method provided in this embodiment aims to score any spoken language without reference text. After obtaining sample audio and its corresponding sample spoken language score, audio features are determined based on the sample audio and input into the oral language assessment model for processing to obtain the assessed spoken language score and the probability distribution of spoken language levels. Then, a first loss value is calculated based on the assessed spoken language score and the sample spoken language score, and a second loss value is calculated based on the probability distribution of spoken language levels and the sample spoken language score. This combines the two loss values to optimize the oral language assessment model until a model that meets the stopping condition is trained for use in oral language assessment scenarios. This method employs multi-task learning to train the model, effectively improving the model's prediction accuracy. It also enables the model to learn the ability to predict spoken language level probabilities, allowing for probability assessment at multiple levels. This effectively improves the interpretability of spoken language predictions, thereby enhancing the accuracy of oral language assessment. Attached Figure Description
[0052] Figure 1 This is a flowchart of a training method for an oral assessment model provided in one embodiment of this specification;
[0053] Figure 2This is a schematic diagram of a training method for an oral assessment model provided in one embodiment of this specification;
[0054] Figure 3 This is a schematic diagram of the structure of a training device for an oral assessment model provided in one embodiment of this specification;
[0055] Figure 4 This is a flowchart of an oral assessment method provided in one embodiment of this specification;
[0056] Figure 5 This is a flowchart illustrating a spoken language assessment method applied in a spoken language tutoring scenario, provided by one embodiment of this specification.
[0057] Figure 6 This is a schematic diagram of the structure of an oral language assessment device provided in one embodiment of this specification;
[0058] Figure 7 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0059] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0060] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0061] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0062] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0063] A loss function is a function that maps the values of a random event or its related random variables to non-negative real numbers to represent the "risk" or "loss" of that random event. In applications, the loss function is often used as a learning criterion in relation to optimization problems; that is, the model is solved and evaluated by minimizing the loss function.
[0064] TDNN (Time-Delay Neural Network) is a convolutional neural network applied to speech recognition problems. It uses FFT-preprocessed speech signals as input, and its hidden layers consist of two one-dimensional convolutional kernels to extract translation-invariant features in the frequency domain.
[0065] ASR (Automatic Speech Recognition) is a technology that converts human speech into text.
[0066] This specification provides a training method for an oral assessment model, and also relates to a training device for an oral assessment model, an oral assessment method, an oral assessment device, a computing device, and a computer-readable storage medium, which will be described in detail in the following embodiments.
[0067] In practical applications, labeling an audio segment with a score is quite difficult, heavily reliant on the subjective standards of the annotators. Different annotators may label the same audio segment differently, and even the same annotator's labeling standards may change at different times. This results in annotations that are difficult to use for supervised learning and lack precision. Existing technologies mostly employ ordered regression methods, but these require selecting appropriate anchor samples for each score level, and sample selection affects the final score accuracy. This not only consumes a significant amount of computation, but even with offline storage of anchor sample features, inference time remains high. Therefore, an effective solution is urgently needed to address these issues.
[0068] The oral language assessment model training method provided in this embodiment aims to score any spoken language without reference text. After obtaining sample audio and its corresponding sample spoken language score, audio features are determined based on the sample audio and input into the oral language assessment model for processing to obtain the assessed spoken language score and the probability distribution of spoken language levels. Then, a first loss value is calculated based on the assessed spoken language score and the sample spoken language score, and a second loss value is calculated based on the probability distribution of spoken language levels and the sample spoken language score. This combines the two loss values to optimize the oral language assessment model until a model that meets the stopping condition is trained for use in oral language assessment scenarios. This method employs multi-task learning to train the model, effectively improving the model's prediction accuracy. It also enables the model to learn the ability to predict spoken language level probabilities, allowing for probability assessment at multiple levels. This effectively improves the interpretability of spoken language predictions, thereby enhancing the accuracy of oral language assessment.
[0069] Figure 1 A flowchart illustrating a training method for a spoken language assessment model according to an embodiment of this specification is shown, specifically including the following steps:
[0070] Step S102: Obtain the sample audio and the corresponding spoken language score of the sample audio.
[0071] Specifically, the sample audio refers to the samples used to train the spoken language assessment model capable of scoring spoken audio. The length and content of the sample audio can be set according to actual needs, and this embodiment does not impose any limitations. Correspondingly, the sample spoken language score specifically refers to the score characterizing the quality of the sample audio. A higher score indicates that the speaker of the sample audio has more standard pronunciation, and vice versa. The sample spoken language score will serve as the label corresponding to the sample audio and will be used to train the spoken language assessment model.
[0072] Therefore, in order to enable the trained oral assessment model to score any spoken audio without reference text, the model training method provided in this embodiment will adopt a multi-task learning approach to train the oral assessment model. During the training phase, various loss functions will be used to tune parameters to train a model that can predict the probability of spoken audio levels, thus ensuring the scoring accuracy during the usage phase.
[0073] Furthermore, when obtaining the spoken language scores corresponding to the sample audio, in order to consider more dimensions and ensure the comprehensiveness and accuracy of the scoring when scoring the spoken audio in the future, the scores can be evaluated from multiple dimensions and the labels corresponding to multiple samples. In this embodiment, the specific implementation method is as follows:
[0074] Obtain the sample audio; determine the fluency score, pronunciation score, and semantic score of the sample audio; calculate the average of the fluency score, pronunciation score, and speech score, and determine the sample spoken language score based on the calculation results.
[0075] Specifically, the fluency score in the fluency dimension refers to the score corresponding to the fluency of the sample audio, the pronunciation score in the pronunciation dimension refers to the score corresponding to the pronunciation accuracy of the sample audio, and the semantic score in the semantic dimension refers to the score corresponding to the sentence content of the sample audio.
[0076] Based on this, after obtaining the sample audio, in order to ensure that the scoring criteria are more comprehensive when scoring spoken audio in the future, we can determine the fluency score corresponding to the fluency dimension, the pronunciation score corresponding to the pronunciation dimension, and the semantic score corresponding to the semantic dimension of the sample audio. Then, we can determine the sample spoken score corresponding to the sample audio by calculating the average of the three, and use this as a label for subsequent model training. This can effectively improve the prediction accuracy of the model and enable the model to be evaluated by comprehensively considering the prediction capabilities of multiple dimensions.
[0077] This embodiment uses the process of training the model with a single sample audio file as an example. Each iteration of the model training process can be described in the same or corresponding description in this embodiment, and will not be elaborated further here.
[0078] For example, the sample audio is "The weather is so nice today". The fluency score S1, pronunciation score S2, and semantic score S3 of the sample audio are obtained. Then, the fluency score S1, pronunciation score S2, and semantic score S3 are summed and averaged to obtain the sample spoken language score (S1+S2+S3) / 3 = S of the sample audio "The weather is so nice today". This score is used as a label for the sample audio to train the spoken language assessment model.
[0079] In summary, by simultaneously scoring sample audio from the dimensions of fluency, semantics, and pronunciation, and calculating the spoken score of the sample audio by taking the average, we can evaluate the sample audio from multiple dimensions. This ensures that the labels can map the pronunciation of the sample audio from various angles, enabling accurate training of the model in the future.
[0080] Step S104: Determine audio features based on the sample audio, and input the audio features into the oral assessment model for processing to obtain the oral assessment score and the probability distribution of oral level.
[0081] Specifically, after obtaining the sample audio and its corresponding spoken scores, a multi-task learning approach will be used to train the model to improve its predictive ability. That is, the model will be trained using multiple methods during the training phase, enabling the spoken assessment model to learn different spoken scoring abilities from multiple perspectives for use in subsequent spoken assessment processing.
[0082] Specifically, audio features refer to the feature expressions that can be input into the oral assessment model. Correspondingly, the oral assessment score refers to the score given by the oral assessment model to the sample audio during the training phase. The oral level probability distribution refers to the distribution corresponding to the probability that the sample audio is better than each level and worse than each level after the oral assessment model evaluates the oral level of the sample audio during the training phase. This avoids the problem of directly determining the oral audio score at once during the oral assessment phase, which is too rigid and has low interpretability. It realizes the comprehensive evaluation of oral audio from multiple levels, improving the accuracy and comprehensiveness of the evaluation.
[0083] Based on this, once the sample audio is obtained, audio features will be constructed based on the sample audio to obtain input that can be input into the oral assessment model to be trained. Then, it will be input into the oral assessment model to be trained for processing to obtain the oral assessment score corresponding to the sample audio and the oral level probability distribution corresponding to the sample audio. This will facilitate subsequent parameter tuning of the model from multiple perspectives and improve the prediction accuracy of the model.
[0084] Furthermore, in order to reduce the interference of other factors on the spoken language assessment model when determining audio features, the sample audio can be converted into text, and the features corresponding to the decoded text can be extracted as the input of the spoken language assessment model. In this embodiment, the specific implementation method is as follows:
[0085] Acoustic features are constructed based on the sample audio; the acoustic features are input into a pre-trained acoustic model for processing to obtain the audio features.
[0086] Specifically, acoustic features refer to the vector representation of the sample audio in the time domain, used to characterize the features of the sample audio; correspondingly, an acoustic model refers to a neural network capable of converting audio into text. During the conversion process, the decoder in the neural network can generate text. In order to train the spoken language assessment model, the vector representation before decoding into text can be selected as audio features for subsequent processing; that is, audio features are the vector representation before decoding into text during the process of converting sample audio into text by the acoustic model.
[0087] In practical applications, in order to convert sample audio into audio features that conform to the input of the spoken language assessment model, a TDNN model pre-trained based on ASR can be selected as the acoustic model to construct the audio features corresponding to the sample audio, which is convenient for subsequent processing by the spoken language assessment model.
[0088] In summary, by using a pre-trained acoustic model to process sample audio to obtain audio features, we can fully integrate the features involved in audio-to-text conversion to influence the prediction accuracy of the spoken language assessment model, thereby training a spoken language assessment model that meets the usage requirements.
[0089] Furthermore, when predicting the probability distribution of spoken language proficiency levels using a spoken language assessment model, since the level information predicted by the model is pre-learned, it is necessary to segment the spoken language proficiency information during the sample preprocessing stage. This facilitates the spoken language assessment model in predicting the probability distribution of spoken language proficiency levels based on this information. In this embodiment, the specific implementation method is as follows:
[0090] The spoken language scores of the samples are divided to obtain spoken language level information; the audio features are input into the spoken language assessment model for processing to obtain the sample spoken language level probability corresponding to the spoken language level information; and the spoken language level probability distribution corresponding to the sample audio is constructed based on the sample spoken language level probability.
[0091] Specifically, spoken language proficiency information refers to the information corresponding to the levels divided according to the spoken language scores of the samples. Different scores correspond to different levels, so as to represent the quality of the audio through different levels.
[0092] Based on this, once the sample audio scores are obtained, the sample audio can be divided to obtain spoken language level information. When the model makes predictions on the sample audio, the model outputs the sample spoken language level probability corresponding to the sample audio. In other words, the spoken language assessment model will predict the probability values of the sample audio being better than or worse than the level in each spoken language level, so as to construct a spoken language level probability distribution based on the sample spoken language level probability corresponding to each level, and optimize the model in combination with the sample spoken language scores.
[0093] For example, see Figure 2As shown, based on the sample audio "The weather is really nice today", further, after determining the sample spoken language score S of the sample audio, it is divided into levels according to the requirements, resulting in k levels. At this time, the acoustic features corresponding to the sample audio are constructed, and then the acoustic features are input into the TDNN model pre-trained based on ASR for processing to obtain the audio features corresponding to the sample audio. At this time, a multi-task training method is used to train the spoken language assessment model, that is, the audio features are input into the spoken language assessment model to be trained for processing to obtain the predicted spoken language score Sn of the model, as well as the probability values of the sample audio being better than the level and worse than the level in each of the k levels. That is, when the sample audio corresponds to level k=1, the probability values of being better than level k=1 and worse than level k=1 are obtained. And so on, the probability values corresponding to the k levels are obtained through the spoken language assessment model for subsequent model optimization.
[0094] In summary, by pre-classifying spoken language level information based on sample spoken language scores, the spoken language assessment model learns the ability to classify spoken language level probabilities according to the level information. This enables the model to predict the spoken language level probability value corresponding to the sample audio at each level, effectively improving the interpretability of the model.
[0095] Step S106: Calculate a first loss value based on the assessed spoken language score and the sample spoken language score, and calculate a second loss value based on the spoken language level probability distribution and the sample spoken language score.
[0096] Specifically, based on the oral assessment scores and oral proficiency distributions output by the oral assessment model, and further, since a multi-task learning approach is used to train the oral assessment model, different loss functions need to be combined when tuning the model parameters. That is, a first loss value can be calculated based on the oral assessment scores and sample oral scores, and a second loss value can be calculated based on the oral proficiency distribution and sample oral scores, so that multiple models can be optimized with two different loss values in the subsequent process.
[0097] It should be noted that the first loss value and the second loss value will be calculated using two different loss functions. The loss function corresponding to the first loss value is used to assist in training the oral assessment model, that is, to optimize the model from the perspective of assessment scores, while the loss function corresponding to the second loss value is used to primarily train the oral assessment model.
[0098] Furthermore, when calculating the first loss value, considering that the first loss value is used to assist in training the oral assessment model, the mean squared loss function can be selected. In this embodiment, the specific implementation method is as follows:
[0099] The mean squared loss function is read based on the assessed spoken language score and the sample spoken language score; the mean squared loss function is used to calculate the assessed spoken language score and the sample spoken language score to obtain the first loss value.
[0100] Specifically, the formula for calculating the mean squared loss function is shown in equation (1) below:
[0101]
[0102] Among them, y i This represents the predicted value, specifically the predicted speaking score; s i This represents the labeled value, i.e., the oral score of the sample.
[0103] Based on this, once the assessment speaking score and the sample speaking score are obtained, the mean squared loss function (MSE) can be read based on the assessment speaking score and the sample speaking score. Then, the mean squared loss function is used to calculate the assessment speaking score and the sample speaking score to obtain the first loss value.
[0104] Furthermore, in order to improve the model's predictive ability and enable the model to predict spoken language level probabilities at multiple levels, a second loss value can be calculated using the cross-entropy loss function. In this embodiment, the specific implementation method is as follows:
[0105] The cross-entropy loss function is read based on the spoken language level probability distribution and the sample spoken language score; the cross-entropy loss function is used to calculate the spoken language level probability distribution and the sample spoken language score to obtain the second loss value.
[0106] Specifically, the calculation formula for the cross-entropy loss function is shown in equation (2) below:
[0107]
[0108] Where k represents the oral assessment level, t represents the output of the t-th group in the corresponding oral assessment level, i represents the number of samples, and p i t y represents the probability that the i-th sample is greater than or equal to rank t. i t The label represents the value of a sample, and is 1 if the i-th sample is greater than the rank t, otherwise it is 0; s i This represents the labeled value, i.e., the spoken score of the sample. It should be noted that the loss function is multiplied by 1 / |(s i -t)| is used to characterize that when the score is close to the level t, the oral assessment model has difficulty learning the corresponding ability. In this case, a larger weight can be assigned to assist the model in learning.
[0109] For example, after obtaining the predicted sample score S and the probability distribution of spoken language level corresponding to the sample audio, the predicted sample score S and the sample spoken language score Sn can be substituted into the MSE loss function as shown in Equation (1) to calculate the first loss value MSELoss. At the same time, the spoken language level probability of the k levels corresponding to the sample audio and the sample spoken language score Sn can be substituted into the cross-entropy loss function as shown in Equation (2) to calculate the second loss value CELoss, so as to facilitate the subsequent optimization of the spoken language assessment model by combining the two loss values.
[0110] In summary, by using both mean squared loss and cross-entropy loss functions to tune the model's parameters, we can not only improve the model's learning ability but also effectively enhance its predictive ability, thereby obtaining a spoken language assessment model that meets the usage requirements.
[0111] Step S108: Adjust the parameters of the spoken language assessment model based on the first loss value and the second loss value, and continue training until the training stop condition is met.
[0112] Specifically, based on the first and second loss values obtained above, the model can then be tuned using both values. After tuning, the model is trained continuously to improve its predictive ability through optimization, until a spoken language assessment model that meets the training stopping condition is obtained. The training stopping condition can be the number of iterations or a comparison of loss values; this embodiment does not impose any limitations on this.
[0113] Furthermore, when tuning the oral assessment model using the first and second loss values, since the two loss values control the model to learn different predictive abilities, in order to enable the model to accurately assess spoken audio, the total loss value can be calculated before tuning the parameters. In this embodiment, the specific implementation is as follows:
[0114] Determine the loss weight corresponding to the second loss value, and calculate the third loss value based on the loss weight and the second loss value; sum the first loss value and the third loss value to obtain the target loss value; and tune the oral assessment model based on the target loss value.
[0115] Specifically, the calculation of the third loss value can be completed using the following formula (3):
[0116] Loss=CELoss+λMSELoss (3)
[0117] Where Loss represents the third loss value, and λ represents the weight value corresponding to the mean square loss function.
[0118] Based on this, after obtaining the first and second loss values, a weighted sum can be performed on the first and second loss values, and the sum can be used as the third loss value. Then, the parameters of the oral assessment model can be tuned based on the third loss value. This process can be repeated until an oral assessment model that meets the training stopping condition is obtained and stored.
[0119] Following the example above, after obtaining the first loss value MSELoss and the second loss value CELoss, the third loss value can be calculated using Equation (3). Then, the oral assessment model is tuned based on the third loss value. If the tuned model does not meet the training stopping condition, the above processing method will continue to be used to train the model until an oral assessment model that meets the training stopping condition is obtained, which can then be used for subsequent oral audio evaluation.
[0120] The oral language assessment model training method provided in this embodiment aims to score any spoken language without reference text. After obtaining sample audio and its corresponding sample spoken language score, audio features are determined based on the sample audio and input into the oral language assessment model for processing to obtain the assessed spoken language score and the probability distribution of spoken language levels. Then, a first loss value is calculated based on the assessed spoken language score and the sample spoken language score, and a second loss value is calculated based on the probability distribution of spoken language levels and the sample spoken language score. This combines the two loss values to optimize the oral language assessment model until a model that meets the stopping condition is trained for use in oral language assessment scenarios. This method employs multi-task learning to train the model, effectively improving the model's prediction accuracy. It also enables the model to learn the ability to predict spoken language level probabilities, allowing for probability assessment at multiple levels. This effectively improves the interpretability of spoken language predictions, thereby enhancing the accuracy of oral language assessment.
[0121] Corresponding to the above method embodiments, this specification also provides embodiments of a training device for an oral assessment model. Figure 3 This specification shows a schematic diagram of the structure of a training device for a spoken language assessment model according to an embodiment. Figure 3 As shown, the device includes:
[0122] The acquisition module 302 is configured to acquire sample audio and the corresponding sample spoken score of the sample audio;
[0123] Processing module 304 is configured to determine audio features based on the sample audio, and input the audio features into the oral assessment model for processing to obtain the oral assessment score and the oral level probability distribution;
[0124] The calculation module 306 is configured to calculate a first loss value based on the evaluation oral score and the sample oral score, and to calculate a second loss value based on the oral level probability distribution and the sample oral score;
[0125] Training module 308 is configured to tune the oral assessment model based on the first loss value and the second loss value, and continue training until the training stop condition is met.
[0126] In an optional embodiment, the acquisition module 302 is further configured to:
[0127] Obtain the sample audio; determine the fluency score, pronunciation score, and semantic score of the sample audio; calculate the average of the fluency score, pronunciation score, and speech score, and determine the sample spoken language score based on the calculation results.
[0128] In an optional embodiment, the processing module 304 is further configured to:
[0129] Acoustic features are constructed based on the sample audio; the acoustic features are input into a pre-trained acoustic model for processing to obtain the audio features.
[0130] In an optional embodiment, the computing module 306 is further configured to:
[0131] The mean squared loss function is read based on the assessed spoken language score and the sample spoken language score; the mean squared loss function is used to calculate the assessed spoken language score and the sample spoken language score to obtain the first loss value.
[0132] In an optional embodiment, the computing module 306 is further configured to:
[0133] The cross-entropy loss function is read based on the spoken language level probability distribution and the sample spoken language score; the cross-entropy loss function is used to calculate the spoken language level probability distribution and the sample spoken language score to obtain the second loss value.
[0134] In an optional embodiment, the training module 308 is further configured to:
[0135] Determine the loss weight corresponding to the second loss value, and calculate the third loss value based on the loss weight and the second loss value; sum the first loss value and the third loss value to obtain the target loss value; and tune the oral assessment model based on the target loss value.
[0136] In an optional embodiment, the training device for the spoken language assessment model further includes:
[0137] The segmentation module is configured to segment the spoken language scores of the samples to obtain spoken language level information;
[0138] Accordingly, the processing module 304 is further configured to:
[0139] The audio features are input into the spoken language assessment model for processing to obtain the sample spoken language level probability corresponding to the spoken language level information; based on the sample spoken language level probability, the spoken language level probability distribution corresponding to the sample audio is constructed.
[0140] The training device for the spoken language assessment model provided in this embodiment, in order to score any spoken language without reference text, after obtaining sample audio and its corresponding sample spoken language score, determines audio features based on the sample audio and inputs them into the spoken language assessment model for processing to obtain the assessed spoken language score and the spoken language level probability distribution. At this time, a first loss value is calculated based on the assessed spoken language score and the sample spoken language score, and a second loss value is calculated based on the spoken language level probability distribution and the sample spoken language score. This achieves optimization of the spoken language assessment model by combining the two loss values until a spoken language assessment model that meets the stopping condition is trained for use in spoken language assessment scenarios. It realizes the use of multi-task learning to train the model, which can effectively improve the prediction accuracy of the model, and at the same time enable the model to learn the ability to predict the probability of spoken language level, so as to enable subsequent probability assessment at multiple levels, effectively improving the interpretability of spoken language prediction, thereby improving the accuracy of spoken language assessment.
[0141] The above is a schematic scheme of a training device for an oral assessment model according to this embodiment. It should be noted that the technical solution of this oral assessment model training device and the technical solution of the oral assessment model training method described above belong to the same concept. For details not described in detail in the technical solution of the oral assessment model training device, please refer to the description of the technical solution of the oral assessment model training method described above.
[0142] This implementation also provides an example of an oral assessment method, as follows:
[0143] Figure 4 A flowchart of a spoken language assessment method according to an embodiment of this specification is shown, which specifically includes the following steps:
[0144] Step S402: Obtain spoken audio and determine spoken audio features based on the spoken audio.
[0145] Specifically, spoken audio refers to the audio to be evaluated, which can be uploaded by the user or collected through a terminal device. This embodiment does not make any limitations here. Correspondingly, spoken audio features refer to feature expressions constructed based on spoken audio.
[0146] It should be noted that the process of determining spoken audio features based on spoken audio can be found in the above embodiments, where the description of audio features is determined based on sample audio. For ease of description, this embodiment will not elaborate further.
[0147] Step S404: Input the spoken audio features into the spoken language assessment model in the training method of the spoken language assessment model for processing to obtain the spoken language level probability distribution.
[0148] Specifically, after obtaining the spoken audio features corresponding to the spoken audio, they can be input into a trained spoken language assessment model for processing to obtain the probability of the spoken audio at each spoken language assessment level. This facilitates the subsequent calculation of the spoken audio score by combining the probability corresponding to each spoken language assessment level.
[0149] Step S406: Calculate the target spoken language score of the spoken language audio based on the spoken language level probability distribution.
[0150] Specifically, after predicting the probability of spoken audio levels in each spoken language assessment dimension using the spoken language assessment model, the target spoken language score can then be calculated by combining the probabilities of each level. This allows for a comprehensive scoring of spoken audio by combining the probability distributions of each level, thereby improving the interpretability of the scoring.
[0151] Furthermore, in order to improve the accuracy of the calculation when calculating the target oral score based on the probability distribution of oral proficiency levels, this embodiment is implemented as follows:
[0152] Based on the spoken language level probability distribution, the positive probability value corresponding to each spoken language level of the spoken language audio is extracted; the positive probability values of each spoken language level are summed to obtain the target spoken language score of the spoken language audio.
[0153] Specifically, the positive probability value refers to the probability that the spoken audio is better than that level in each spoken language level. The formula for calculating the target spoken language score is shown in equation (4) below:
[0154]
[0155] Where, p i t This represents the probability value that the spoken audio is greater than level t.
[0156] Based on this, after determining the positive probability values of spoken audio at each spoken level, the positive probability values can be summed to determine the target spoken score corresponding to the spoken audio based on the summation result, thus enabling the spoken audio to be scored.
[0157] For example, the process involves acquiring user-uploaded spoken language audio for evaluation, constructing acoustic features corresponding to the audio, and then inputting these features into a pre-trained TDNN model for processing to obtain spoken language audio features. Further, these features are input into a trained spoken language evaluation model for processing to obtain the probability of the audio being better than each spoken language level. Assuming there are four spoken language levels, the model predicts that the probability of the audio being better than the first level is 0.9, the probability of it being better than the second level is 0.5, the probability of it being better than the third level is 0.2, and the probability of it being better than the fourth level is 0.15.
[0158] Furthermore, after obtaining the probability that the oral proficiency level to be evaluated is better than each oral proficiency level, the probability value corresponding to the oral proficiency level is input into formula (4). The oral proficiency score corresponding to the oral audio to be evaluated is determined to be 0.9+0.5+0.2+0.15=1.75. Finally, the obtained oral proficiency score is fed back to the user to represent the user's pronunciation accuracy so that the user can make corrections.
[0159] In summary, to accurately evaluate spoken audio, we can construct spoken audio features after obtaining the audio, then input these features into a trained spoken language evaluation model for processing. This yields the probability distribution of spoken language levels for the audio to be evaluated. Finally, based on this probability distribution, we can calculate the target spoken language score for the audio. This approach effectively improves the accuracy of spoken language evaluation, enabling evaluation even without reference text, thus broadening the scope of evaluation scenarios and enhancing the user experience.
[0160] The following is in conjunction with the appendix Figure 5 Taking the application of the oral assessment method provided in this manual in an oral tutoring scenario as an example, the oral assessment method will be further explained. Among other things, Figure 5 This specification illustrates a flowchart of a spoken language assessment method applied in a spoken language tutoring scenario, according to an embodiment of this specification. The method specifically includes the following steps:
[0161] Step S502: Obtain sample audio.
[0162] Step S504: Determine the fluency score for the fluency dimension, the pronunciation score for the pronunciation dimension, and the semantic score for the semantic dimension of the sample audio.
[0163] Step S506: Calculate the average of the fluency score, pronunciation score, and speech score, and determine the sample spoken language score based on the calculation results.
[0164] Step S508: Construct acoustic features based on sample audio, and input the acoustic features into a pre-trained acoustic model for processing to obtain audio features.
[0165] Step S510: Input the audio features into the oral assessment model for processing to obtain the oral assessment score and the probability distribution of oral level.
[0166] Step S512: Calculate the evaluation oral score and the sample oral score using the mean square loss function to obtain the first loss value.
[0167] Step S514: Calculate the spoken language level probability distribution and sample spoken language scores using the cross-entropy loss function to obtain the second loss value.
[0168] Step S516: Calculate the target loss value based on the first loss value and the second loss value.
[0169] Step S518: Adjust the parameters of the oral assessment model based on the target loss value, and continue training until the training stopping condition is met.
[0170] Step S520: Receive the spoken audio uploaded by the user to be evaluated.
[0171] Step S522: Construct spoken audio features based on the spoken audio to be evaluated, and input them into the spoken language evaluation model for processing to obtain the spoken language level probability distribution.
[0172] Step S524: Based on the probability distribution of spoken language levels, extract the positive probability value of each spoken language level corresponding to the spoken language audio to be evaluated.
[0173] Step S526: Sum the positive probability values for each spoken language level to obtain the target spoken language score for the spoken language audio to be evaluated, and then provide feedback to the user.
[0174] In summary, to accurately evaluate spoken audio, we can construct spoken audio features after obtaining the audio, then input these features into a trained spoken language evaluation model for processing. This yields the probability distribution of spoken language levels for the audio to be evaluated. Finally, based on this probability distribution, we can calculate the target spoken language score for the audio. This approach effectively improves the accuracy of spoken language evaluation, enabling evaluation even without reference text, thus broadening the scope of evaluation scenarios and enhancing the user experience.
[0175] Corresponding to the above method embodiments, this specification also provides embodiments of oral language assessment devices. Figure 6A schematic diagram of the structure of a spoken language assessment device provided in one embodiment of this specification is shown. Figure 6 As shown, the device includes:
[0176] The audio acquisition module 602 is configured to acquire spoken audio and determine spoken audio features based on the spoken audio.
[0177] The oral assessment module 604 is configured to input the oral audio features into the oral assessment model in the training method of the oral assessment model for processing, and obtain the oral level probability distribution.
[0178] The score calculation module 606 is configured to calculate the target spoken language score of the spoken language audio based on the spoken language level probability distribution.
[0179] In an optional embodiment, the score calculation module 606 is further configured to:
[0180] Based on the spoken language level probability distribution, the positive probability value corresponding to each spoken language level of the spoken language audio is extracted; the positive probability values of each spoken language level are summed to obtain the target spoken language score of the spoken language audio.
[0181] The spoken language assessment device provided in this embodiment, in order to accurately assess spoken language audio, can construct spoken language audio features corresponding to the spoken language audio after obtaining the spoken language audio, and then input them into a trained spoken language assessment model for processing to obtain the spoken language level probability distribution corresponding to the spoken language audio to be assessed. Finally, the target spoken language score corresponding to the spoken language audio is calculated based on the spoken language level probability distribution. This can effectively improve the accuracy of spoken language assessment, realize spoken language assessment even without reference text, effectively expand the scope of assessment scenarios, and improve user experience.
[0182] The above is a schematic scheme of an oral language assessment device according to this embodiment. It should be noted that the technical solution of this oral language assessment device and the technical solution of the oral language assessment method described above belong to the same concept. For details not described in detail in the technical solution of the oral language assessment device, please refer to the description of the technical solution of the oral language assessment method described above.
[0183] Figure 7 A structural block diagram of a computing device 700 according to an embodiment of this specification is shown. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.
[0184] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., a Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) interface, a Wi-MAX interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0185] In one embodiment of this specification, the above-described components of the computing device 700 and Figure 7 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 7 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0186] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 700 can also be a mobile or stationary server.
[0187] The processor 720 is used to execute computer-executable instructions to implement training methods or oral assessment methods for oral assessment models.
[0188] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the technical solution of the above-mentioned training method for the oral assessment model or the oral assessment method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-mentioned training method for the oral assessment model or the oral assessment method.
[0189] An embodiment of this specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, are used for a training method or a speaking assessment method for a speaking assessment model.
[0190] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the above-mentioned training method for the oral assessment model or the oral assessment method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the above-mentioned training method for the oral assessment model or the oral assessment method.
[0191] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0192] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0193] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this specification is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this specification. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this specification.
[0194] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0195] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. These embodiments have been selected and specifically described in this specification to better explain the principles and practical applications of this specification, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A training method for an oral assessment model, characterized in that, include: Obtain the sample audio and the corresponding spoken language score of the sample audio; Audio features are determined based on the sample audio, and the audio features are input into the oral assessment model for processing to obtain the oral assessment score and the oral level probability distribution. The oral level probability distribution includes the distribution of the probability corresponding to the sample audio being higher or lower than each level. A first loss value is calculated based on the assessed oral communication score and the sample oral communication score, and a second loss value is calculated based on the oral communication level probability distribution and the sample oral communication score. The oral assessment model is tuned based on the first loss value and the second loss value, and training continues until the training stop condition is met.
2. The method according to claim 1, characterized in that, The acquisition of sample audio and the corresponding spoken language score includes: Obtain the sample audio; Determine the fluency score for the fluency dimension, the pronunciation score for the pronunciation dimension, and the semantic score for the semantic dimension of the sample audio. Calculate the average of the fluency score, the pronunciation score, and the semantic score, and determine the sample spoken language score based on the calculation results.
3. The method according to claim 1, characterized in that, The step of determining audio features based on the sample audio includes: Acoustic features are constructed based on the sample audio; The acoustic features are input into a pre-trained acoustic model for processing to obtain the audio features.
4. The method according to claim 1, characterized in that, The calculation of the first loss value based on the assessed spoken language score and the sample spoken language score includes: The mean squared loss function is read based on the evaluation oral score and the sample oral score; The first loss value is obtained by calculating the evaluation oral score and the sample oral score using the mean square loss function.
5. The method according to claim 1, characterized in that, The step of calculating the second loss value based on the spoken language level probability distribution and the sample spoken language score includes: The cross-entropy loss function is read based on the spoken language level probability distribution and the sample spoken language scores; The cross-entropy loss function is used to calculate the spoken language level probability distribution and the sample spoken language score to obtain a second loss value.
6. The method according to any one of claims 1-5, characterized in that, The parameter tuning of the spoken language assessment model based on the first loss value and the second loss value includes: Determine the loss weight corresponding to the second loss value, and calculate the third loss value based on the loss weight and the second loss value; The target loss value is obtained by summing the first loss value and the third loss value; The oral assessment model is tuned based on the target loss value.
7. The method according to any one of claims 1-5, characterized in that, After the steps of obtaining the sample audio and the corresponding spoken language score are executed, the method further includes: The oral scores of the samples are divided to obtain oral proficiency level information; Accordingly, the step of inputting the audio features into the spoken language assessment model for processing to obtain the spoken language level probability distribution includes: The audio features are input into the spoken language assessment model for processing to obtain the sample spoken language level probability corresponding to the spoken language level information. The spoken language level probability distribution corresponding to the sample audio is constructed based on the sample spoken language level probability.
8. A training device for an oral assessment model, characterized in that, include: The acquisition module is configured to acquire sample audio and the corresponding spoken language score of the sample audio; The processing module is configured to determine audio features based on the sample audio, and input the audio features into the oral assessment model for processing to obtain the oral assessment score and the oral level probability distribution, wherein the oral level probability distribution includes the distribution corresponding to the probability of the sample audio being higher or lower than each level; The calculation module is configured to calculate a first loss value based on the assessed spoken language score and the sample spoken language score, and to calculate a second loss value based on the spoken language level probability distribution and the sample spoken language score; The training module is configured to tune the oral assessment model based on the first loss value and the second loss value, and continue training until the training stop condition is met.
9. A method for oral language assessment, characterized in that, include: Acquire spoken audio and determine spoken audio features based on the spoken audio; The spoken audio features are input into the spoken language assessment model in any one of claims 1-7 for processing to obtain a spoken language level probability distribution, wherein the spoken language level probability distribution includes the distribution corresponding to the probability of the spoken audio being higher or lower than each level; The target spoken language score of the spoken language audio is calculated based on the spoken language level probability distribution.
10. The method according to claim 9, characterized in that, The calculation of the target spoken language score of the spoken language audio based on the spoken language level probability distribution includes: Based on the spoken language level probability distribution, extract the positive probability value of each spoken language level corresponding to the spoken language audio; The target spoken language score of the spoken language audio is obtained by summing the positive probability values for each spoken language level.
11. A spoken language assessment device, characterized in that, include: The audio acquisition module is configured to acquire spoken audio and determine spoken audio features based on the spoken audio. The spoken language assessment module is configured to input the spoken language audio features into the spoken language assessment model in any one of the methods of claims 1-7 for processing to obtain a spoken language level probability distribution, wherein the spoken language level probability distribution includes the distribution corresponding to the probability of the spoken language audio being higher than and lower than each level; The score calculation module is configured to calculate the target spoken language score of the spoken language audio based on the spoken language level probability distribution.
12. A computing device, characterized in that, It includes a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 7 or 9 to 10.
13. A computer-readable storage medium storing computer instructions, characterized in that, When executed by a processor, this instruction implements the steps of the method according to any one of claims 1 to 7 or 9 to 10.
Citation Information
Patent Citations
Oral test and evaluation method, device and device for generating oral test and evaluation model
CN109272992A
Voice fluency detection method and device and electronic equipment
CN112951270A