Spoken language evaluation method and device, related equipment and computer program product

By combining the inference scoring model and calibration model, the problem of oral evaluation systems in the existing technology relying too much on reference answers is solved, and more accurate and fair oral evaluation results are achieved.

CN119942858AActive Publication Date: 2025-05-06IFLYTEK CO LTD

Patent Information

Application Number
CN202510029197.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The existing oral evaluation system relies too much on reference answers when rating, resulting in the reduction of the accuracy of the score when the candidate gives the correct answer but is not within the scope of the reference answer.

Method used

By obtaining the candidate's answer data, identifying the answer audio and converting it into text, and combining the reasoning scoring model and the calibration model, comprehensive scoring is performed. The reasoning scoring model includes pronunciation scores and completeness scores. The calibration model is based on pre-training for expert scoring, which can more accurately evaluate candidates' oral skills.

Benefits of technology

It improves the accuracy of oral evaluation and can more fairly evaluate candidates' comprehensive abilities, including language ability, logical thinking ability and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942858A_ABST
    Figure CN119942858A_ABST
Patent Text Reader

Abstract

The invention discloses a spoken language evaluation method and device, related equipment and a computer program product, and the method comprises the steps: obtaining the answer data of an examinee, the answer data comprises a question, the answer audio of the examinee, and a reference answer; identifying an answer text corresponding to the answer audio, and combining the answer text and the answer data to obtain a reasoning score of the examinee through a configured reasoning scoring model; a configured calibration model is obtained, the calibration model is obtained through pre-training based on answer texts of calibrated examinees, reasoning scores of the calibrated examinees and expert scores, and the calibrated examinees are part of examinees extracted from all the examinees participating in the oral test; and according to the answer text and the reasoning score of each examinee, scoring by using the calibration model to obtain a final score of each examinee. Compared with the mode of determining the score by purely calculating the similarity between the answer text and the reference answer in the prior art, the spoken language evaluation result obtained by the scheme of the invention is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of oral language assessment, and more specifically, to an oral language assessment method, apparatus, related equipment and computer program product. Background Art

[0002] In the speaking test scenario, candidates are usually required to demonstrate their speaking ability through a series of speaking tasks. These tasks may include introducing themselves, describing pictures or diagrams, answering questions, participating in conversations, etc.

[0003] With the development of artificial intelligence technology, many oral assessment systems can already realize automatic scoring functions. These systems convert the candidates' oral expressions into text through speech recognition technology, and use natural language processing algorithms to analyze various aspects of the text. However, in the examination scenario, in addition to demonstrating language ability, candidates also need to demonstrate confidence, logical thinking ability, and adaptability. Therefore, the oral assessment test is not only a test of language ability, but also a test of comprehensive ability. However, due to the unique performance of candidates, the standard reference answers cannot cover the range of all suitable answers. The current oral assessment system's scoring relies heavily on the content of the reference answers, but the number of reference answers is limited, and all correct answers cannot be included in the test paper making stage, resulting in the candidate giving the correct answer, but because the answer is not included in the reference answer of the test paper in advance, the text similarity between the candidate's answer and the reference answer is not high, resulting in a lower score, which reduces the accuracy of the oral assessment. Summary of the invention

[0004] In view of the above problems, this application is proposed to provide a spoken language assessment method, apparatus, related equipment and computer program product to improve the accuracy of spoken language assessment. The specific solution is as follows:

[0005] In a first aspect, a spoken language assessment method is provided, comprising:

[0006] Obtaining the examinee's answer data, wherein the answer data includes the question, the examinee's answer audio and reference answers;

[0007] Identify the answer text corresponding to the answer audio, combine the answer text and the answer data, and obtain the reasoning score of the examinee through the configured reasoning scoring model;

[0008] Obtaining a configured calibration model, wherein the calibration model is obtained by pre-training based on the answer texts of the calibration candidates, the reasoning scores of the calibration candidates, and the expert scores, wherein the calibration candidates are selected from all the candidates participating in this oral test;

[0009] The calibration model is used to score each examinee's answer text and reasoning score to obtain a final score for each examinee.

[0010] In one possible design, in another implementation of the first aspect of the embodiment of the present application, the inference scoring model includes a pronunciation scoring model and a completeness scoring model;

[0011] The process of obtaining the reasoning score of the examinee by combining the answer text and the answer data through the configured reasoning score model includes:

[0012] Processing the answer audio using the pronunciation scoring model to obtain a pronunciation score output by the model;

[0013] Processing the question, the reference answer and the answer text using the completeness scoring model to obtain a completeness score output by the model;

[0014] The examinee's reasoning score is composed of the pronunciation score and the completeness score.

[0015] In a possible design, in another implementation of the first aspect of the embodiments of the present application, when the reference answers include more than two, the process of using the completeness scoring model to process the question, the reference answer, and the answer text to obtain a completeness score output by the model includes:

[0016] Calculate the similarity between the answer text and each reference answer, and select the target reference answer with the highest similarity;

[0017] The completeness scoring model is used to process the question, the target reference answer and the answer text to obtain a completeness score output by the model.

[0018] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the process of extracting calibration candidates from all candidates participating in the oral test includes:

[0019] All candidates who participated in the oral test were sorted according to their reasoning scores;

[0020] Some calibration candidates are selected in each set scoring range to form all the calibration candidates.

[0021] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the calibration model includes: a feature extraction model corresponding to each examinee, and a regression model;

[0022] The feature extraction model corresponding to each examinee is a model trained by using the examinee's answer text in an unsupervised training manner, and the feature extraction model is used to extract text features of the corresponding examinee's answer text;

[0023] The regression model is trained by taking the text features and reasoning scores of the test-taking examinees' answer texts as inputs and taking the expert scores of the test-taking examinees as outputs.

[0024] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the regression model includes: a first regression model and a second regression model;

[0025] The first regression model is trained by taking the text features of the test-taking examinee's answer text as input and taking the expert score of the test-taking examinee as output;

[0026] The second regression model is trained by taking the expert scoring result predicted by the first regression model for the calibration examinee and the reasoning score of the calibration examinee as input, and taking the expert score of the calibration examinee as output.

[0027] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the process of scoring each examinee using the calibration model according to the answer text and the reasoning score of each examinee to obtain the final score of each examinee includes:

[0028] For each candidate:

[0029] Sending the answer text of the examinee to a corresponding feature extraction model to extract text features of the answer text of the examinee;

[0030] Sending the text features of the answer text of the examinee into the first regression model to obtain a first score as an output;

[0031] The first score and the examinee's reasoning score are sent to the second regression model to obtain a final score as output.

[0032] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the process of identifying the answer text corresponding to the answer audio includes:

[0033] Extracting acoustic features of the answer audio;

[0034] The acoustic features are input into a configured speech recognition model, which includes an encoder, a decoder and a keyword excitation module. The keyword excitation module encodes the features of each keyword in the reference answer, predicts the sub-word score of each keyword, performs weighted fusion on the sub-word scores predicted by the decoder and the sub-word scores predicted by the keyword excitation module, and determines the recognition result based on the fused sub-word scores as the answer text.

[0035] In a possible design, in another implementation of the first aspect of the embodiment of the present application, when there are multiple examinees to be evaluated, a recognition core is configured for each examinee, and a reasoning core corresponding to the multiple recognition cores is configured;

[0036] The process of identifying the answer text corresponding to each candidate's answer audio includes:

[0037] Construct the input features of the speech recognition model through the recognition core corresponding to each examinee, mark the inference sequence number of the recognition core and send it to the inference core, receive the inference result returned by the inference core, and output the final recognized answer text based on the inference result;

[0038] The inference core obtains the input features carrying the inference sequence number sent by each recognition core, and splices the multiple input features according to the inference sequence number, and sends them to the speech recognition model for inference, and returns the inference result to the corresponding recognition core according to the inference sequence number.

[0039] In a second aspect, a spoken language evaluation device is provided, comprising:

[0040] The answer data acquisition unit is used to acquire the examinee's answer data, wherein the answer data includes the question, the examinee's answer audio and the reference answer;

[0041] An audio recognition unit, used to recognize the answer text corresponding to the answer audio;

[0042] A reasoning scoring unit, used to combine the answer text and the answer data to obtain the reasoning score of the examinee through a configured reasoning scoring model;

[0043] A calibration model acquisition unit, used to acquire a configured calibration model, wherein the calibration model is obtained by pre-training based on the answer texts of the calibration examinees, the reasoning scores of the calibration examinees and the expert scores, and the calibration examinees are some examinees selected from all examinees participating in this oral test;

[0044] The final score determination unit is used to score each candidate based on the answer text and the reasoning score using the calibration model to obtain the final score of each candidate.

[0045] In a third aspect, an electronic device is provided, comprising: a memory and a processor;

[0046] The memory is used to store programs;

[0047] The processor is used to execute the program to implement each step of the oral evaluation method described in any one of the first aspects of the present application.

[0048] In a fourth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the oral language assessment method described in any one of the first aspects of the present application are implemented.

[0049] In a fifth aspect, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements the various steps of the oral language assessment method described in any one of the first aspects of the present application.

[0050] By means of the above technical scheme, the present application determines the reasoning score of the examinee through the reasoning scoring model by identifying the answer text corresponding to the answer audio of the examinee, combining the answer text and the reference answer and answer audio in the answer data. On this basis, in order to improve the accuracy of oral evaluation, a calibration model is further introduced, that is, some calibration examinees can be extracted from all examinees participating in this oral test, and the experts score the calibration examinees to obtain the expert score, which is more accurate. On this basis, the calibration model is trained using the answer text, reasoning score and expert score of the calibration examinee, and the calibration model can predict more accurate expert scores based on the examinee's answer text and reasoning score. Therefore, for each examinee, the calibration model can be used to score according to its answer text and reasoning score to obtain the final score of each examinee. In the process of determining the final score of the examinee, the present application combines the reference answer on the one hand, and combines the expert's calibration score on the other hand, giving the calibration model a certain degree of freedom in scoring based on the reference answer. Compared with the existing technology of simply calculating the similarity between the answer text and the reference answer to determine the score, the oral evaluation result obtained by the present application scheme is more accurate. In addition, since the calibration model is trained based on the expert scores of the calibration candidates, it can be understood as customized training and learning for the scoring scale and student answers of each oral test. It is therefore more suitable for evaluating the oral performance of each candidate in this oral test, and further improves the accuracy of the oral evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0052] Figure 1 A schematic diagram of an implementation system architecture of the spoken language assessment method provided in an embodiment of the present application;

[0053] Figure 2 A flow chart of a spoken language assessment method provided in an embodiment of the present application;

[0054] Figure 3 A structural diagram of a pronunciation scoring model provided in an embodiment of the present application;

[0055] Figure 4 A schematic diagram of the structure of a completeness scoring model provided in an embodiment of the present application;

[0056] Figure 5 A schematic diagram of a speech recognition model structure provided in an embodiment of the present application;

[0057] Figure 6 A flowchart of a method for recognizing audio answers of multiple examinees provided in an embodiment of the present application;

[0058] Figure 7 A flowchart of another spoken language assessment method provided in an embodiment of the present application;

[0059] Figure 8 A schematic diagram of the structure of a spoken language evaluation device provided in an embodiment of the present application;

[0060] Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0062] This application provides a spoken language assessment method that can be applied to Figure 1 The system architecture shown in FIG. 1 may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1A server is included as an example for explanation).

[0063] The terminal 100 in this embodiment can specifically be a device with data processing capabilities, such as a tablet computer, a learning machine, a teaching terminal, or a laptop computer.

[0064] The terminal 100 or the server 200 can be used alone to execute the spoken language evaluation method provided in the embodiment of the present application. In addition, the terminal 100 and the server 200 can also be used together to execute the spoken language evaluation method provided in the embodiment of the present application.

[0065] Exemplarily, the terminal 100 obtains the test-taking data of the examinee, completes the oral assessment of the examinee, and obtains the oral assessment result.

[0066] In another example, the terminal 100 obtains the test-taker's answer data and uploads the answer data to the server 200 . The server 200 completes the oral assessment of the test-taker and returns the oral assessment result to the terminal 100 after obtaining it.

[0067] In another example, the terminal 100 and the server 200 can cooperate with each other to complete the oral assessment of the examinee, that is, some steps of the oral assessment are implemented on the terminal 100, and other steps are implemented on the server 200, and the two coordinate and cooperate to complete the oral assessment.

[0068] The oral assessment solution provided in the embodiment of the present application can perform oral assessment on the candidate's answer data (such as answer audio) transmitted offline, and can also perform oral assessment on the candidate's answer data transmitted online.

[0069] The oral assessment solution of the present application can be applied to oral examination scenarios of various subjects, such as English oral examination assessment, Chinese oral examination assessment, etc.

[0070] The present application embodiment provides a spoken language evaluation method, which is illustrated by applying the method to a computer device. The computer device may be Figure 1 The terminal 100 or the system consisting of the terminal 100 and the server 200. Figure 2 , the oral evaluation method specifically comprises the following steps:

[0071] Step S100, obtaining the examinee's answer data, wherein the answer data includes questions, the examinee's answer audio and reference answers.

[0072] Specifically, for all candidates participating in this oral test, the answer data of each candidate includes test paper information and corresponding answer audio. The test paper information here is a file made according to the test paper making rules, which may include questions and corresponding reference answers.

[0073] Depending on the type of oral test questions, the corresponding reference answers may be one or more. For example, questions such as sentence translation and paraphrase may have more than one reference answer.

[0074] Step S110, identifying the answer text corresponding to the answer audio, combining the answer text and the answer data, and obtaining the candidate's reasoning score through the configured reasoning scoring model.

[0075] Specifically, for each candidate's answer audio, the recognition result, i.e. the answer text, is obtained through the speech recognition model.

[0076] Furthermore, in this step, the answer text and the answer data can be combined to use a pre-trained reasoning scoring model to perform score prediction, and obtain the candidate's reasoning score output by the model.

[0077] In this embodiment, the reasoning scoring model is configured to predict the corresponding reasoning score based on the answer text and the answer data.

[0078] The inference scoring model can at least refer to the reference answers in the answer data, and determine the inference score based on the matching between the answer text and the reference answer. On this basis, the inference scoring model can further analyze and process the answer audio in the answer data to score the candidate's pronunciation, and determine the final inference score based on the pronunciation, the matching between the answer text and the reference answer.

[0079] Step S120, obtaining a configured calibration model, wherein the calibration model is pre-trained based on the calibration examinee's answer text, the calibration examinee's reasoning score and the expert score.

[0080] The calibration candidates are selected from all the candidates who take the oral test. The calibration candidates may be selected according to a set selection rule, for example, a set number of calibration candidates are randomly selected from all the candidates.

[0081] This embodiment provides an optional extraction method:

[0082] First, all candidates participating in this oral test are sorted according to their reasoning scores.

[0083] Furthermore, some calibration candidates are selected in each set scoring interval to form all the calibration candidates.

[0084] By adopting the extraction method of this embodiment, the extracted calibration examinees can cover various scores as much as possible, which is convenient for comprehensively and fairly evaluating the level of the entire examinees.

[0085] After the calibration candidates are extracted, experts can be arranged to score the test data of the calibration candidates to obtain the expert scores of the calibration candidates. It is understandable that the results of the expert scores are more accurate and can be used as standard scores.

[0086] On this basis, the calibration model is trained based on the test-taking text, reasoning score and expert score of the calibration test-taker. The calibration model is configured to predict the expert score of the test-taker based on the test-taking text and reasoning score.

[0087] Step S130: Scoring is performed using the calibration model according to the answer text and the reasoning score of each examinee to obtain a final score for each examinee.

[0088] Specifically, the above steps can be used to obtain the answer text and reasoning score of each candidate. On this basis, the calibration model can be called to predict the final score to obtain the final score of each candidate.

[0089] The oral evaluation method provided in the embodiment of the present application, by identifying the answer text corresponding to the answer audio of the examinee, combining the answer text and the reference answer and answer audio in the answer data, and determining the examinee's reasoning score through the reasoning scoring model. On this basis, in order to improve the accuracy of oral evaluation, a calibration model is further introduced, that is, some calibration examinees can be extracted from all examinees participating in this oral test, and experts score the calibration examinees to obtain expert scores, which are more accurate. On this basis, the calibration model is trained using the answer text, reasoning score and expert score of the calibration examinee, and the calibration model can predict more accurate expert scores based on the examinee's answer text and reasoning score. Therefore, for each examinee, the calibration model can be used to score according to its answer text and reasoning score to obtain the final score of each examinee. In the process of determining the examinee's final score, the present application combines the reference answer on the one hand, and combines the expert's calibration score on the other hand, giving the calibration model a certain degree of freedom in scoring based on the reference answer. Compared with the existing technology of simply calculating the similarity between the answer text and the reference answer to determine the score, the oral evaluation result obtained by the present application scheme is more accurate. In addition, since the calibration model is trained based on the expert scores of the calibration candidates, it can be understood as customized training and learning for the scoring scale and student answers of each oral test. It is therefore more suitable for evaluating the oral performance of each candidate in this oral test, and further improves the accuracy of the oral evaluation results.

[0090] In some embodiments of the present application, the process of obtaining the reasoning score of the examinee through the reasoning scoring model is described.

[0091] The reasoning scoring model can measure the completeness of the answer text from the text level. In addition, it can further measure the pronunciation of the answer audio from the answer audio level to obtain the pronunciation score. The completeness score and pronunciation score are comprehensively considered to determine the candidate's reasoning score.

[0092] Take the example of an inference score that includes both a pronunciation score and a completeness score:

[0093] The inference scoring model may include a pronunciation scoring model and a completeness scoring model.

[0094] For the pronunciation scoring model, an end-to-end pronunciation scoring model is trained in a data-driven manner.

[0095] The pronunciation scoring model is used to process the test takers’ answer audio and obtain the pronunciation score output by the model.

[0096] Combination Figure 3 An optional structure of a pronunciation scoring model is shown in the example. The pronunciation scoring model may include an audio encoder Encoder and a fully connected layer FC. The Encoder may adopt an encoder based on a Conformer structure.

[0097] The input of the pronunciation scoring model may include the acoustic features of the answer audio, such as Fbank features, etc. In addition, it may further include the frame number mask corresponding to the acoustic features in the answer audio.

[0098] like Figure 3 As shown, a "CLS" tag can be added before the input acoustic features. On this basis, the output of the encoder at the "CLS" position can be fed into the fully connected layer to predict the pronunciation score.

[0099] The pronunciation scoring model uses audio samples marked with pronunciation scoring labels as training data during training, and is trained according to a set loss function. For example, the loss function may be a Huber loss function. The specific calculation process is as follows:

[0100] For each audio sample i, calculate the difference between the predicted value and the true value:

[0101] According to the absolute value of the difference and the size of δ, the loss is calculated in two cases:

[0102] 1) If The loss is half the square of the difference, that is:

[0103]

[0104] 2) If The loss is:

[0105]

[0106] 3) Average the loss of all audio samples:

[0107]

[0108] Where n represents the number of audio samples, yi represents the true value of the pronunciation score of the i-th audio sample (i.e., the pronunciation score label), Represents the value predicted by the model, δ is the parameter of Huber loss, which is used to control the trade-off between the mean absolute error MAE and the mean square error MSE of the loss function. A larger δ makes the loss function closer to the mean square error MSE at large errors, and a smaller δ makes the loss function closer to the mean absolute error MAE.

[0109] Of course, the above only illustrates the case of the Huber loss function. The pronunciation scoring model may also use other types of loss functions during the training phase, which will not be listed one by one in this embodiment.

[0110] In this embodiment, by training an end-to-end pronunciation scoring model, the trained pronunciation scoring model can be used to predict the pronunciation score corresponding to the candidate's answer audio, thereby measuring the candidate's level from the perspective of pronunciation.

[0111] For the completeness scoring model, it can also be trained end-to-end in a data-driven manner.

[0112] The completeness scoring model is used to process the questions, reference answers and answer texts in the examinees' answer data to obtain the completeness score output by the model.

[0113] The completeness scoring model can use BERT or other backbone networks as the main structure. Figure 4 An example of using BERT as a completeness scoring model for the main structure is illustrated.

[0114] The input text sequence of the completeness scoring model may include question information, reference answers, and answer text.

[0115] Among them, delimiters can be set between different types of input. For example, the input text sequence can be in the following format:

[0116] [CLS]Question sequence [SEP]Ref sequence [SEP]ASR sequence [SEP].

[0117] Among them, Question represents the question, Ref represents the reference answer, and ASR represents the answer text.

[0118] It should be noted that if there are more than two reference answers corresponding to the current question in the answer data, the similarity between the candidate's answer text and each reference answer can be calculated, and the target reference answer with the highest similarity can be selected.

[0119] When calculating the similarity between the candidate's answer text and each reference answer, the Levenshtein distance between the answer text and each reference answer can be calculated. The smaller the distance, the higher the similarity between the two. Therefore, the reference answer with the smallest distance can be selected as the target reference answer Ref.

[0120] When constructing the input text sequence of the completeness scoring model, the input text sequence can be composed of questions, target reference answers, and answer texts.

[0121] After obtaining the input text sequence, the token sequence corresponding to the input text sequence can be determined. Exemplarily, a tokenizer can be used to process the input text sequence to obtain a token sequence, namely token_id.

[0122] Furthermore, the input of BERT can also include token_id_type and token_mask.

[0123] Among them, token_id_type can be an all-zero sequence with the same length as token_id, and the position value of the token belonging to the answer text in token_id is set to 1, that is, the answer text in token_id is marked.

[0124] token_mask represents a mask, which is used to retain valid data in multiple batches. token_mask can be a sequence of all 1s with the same length as token_id.

[0125] The token_id, token_id_type and token_mask constructed above are added together to get the input of BERT. BERT processes it and predicts the output completeness score.

[0126] In this embodiment, by training an end-to-end completeness scoring model, the trained completeness scoring model can be used to predict the corresponding completeness scores for the questions, reference answers and answer texts in the examinee's answer data, thereby measuring the completeness level of the answer text from the text level.

[0127] After the pronunciation score and the completeness score of the examinee are obtained in the above embodiment, the two scores can be used together to form the examinee's reasoning score.

[0128] In some embodiments of the present application, the process of identifying the answer text corresponding to the examinee's answer audio in the aforementioned step S110 is explained.

[0129] In one possible implementation, a traditional speech recognition model can be used to recognize the candidate's answer audio to obtain the answer text.

[0130] In some other possible implementations, in order to improve the accuracy of speech recognition results, this embodiment introduces a composition structure of a speech recognition model, which adopts an encoder-decoder structure and is provided with a keyword excitation module.

[0131] When building a speech recognition model, subwords can be used as modeling units. Compared with using whole words as modeling units, using subwords as modeling units can reduce false triggers.

[0132] For the test taker’s answer audio, the acoustic features are extracted and sent to the acoustic encoder for encoding. The encoded features are sent to the decoder. The decoder predicts the subword score P based on the encoded features. ed .

[0133] The keyword excitation module encodes the features of each keyword in the reference answer and predicts the subword score P of each keyword. hot .

[0134] In a possible implementation, the keyword excitation module may further refer to the hidden features extracted by the decoder when predicting the subword score of each keyword, that is, the keyword excitation module may predict the subword score P of each keyword based on the feature encoding result of each keyword in the reference answer and the hidden features extracted by the decoder. hot .

[0135] Finally, the subword score P predicted by the decoder is ed and the subword score P predicted by the keyword excitation module hot Perform weighted fusion:

[0136] P total =P ed ×(1-coeff)+P hot ×coeff

[0137] Among them, P total represents the fused sub-word score, coeff represents the score weight of the keyword excitation module, and for example, coeff can be equal to 0.9.

[0138] Based on the fused subword score P total Determine the recognition result as the answer text.

[0139] The structure of the speech recognition model provided in this embodiment can incentivize the sub-word scores of the keywords in the reference answers by adding a keyword incentive module, thereby improving the probability that the recognition results hit the keywords in the reference answers, and avoiding the impact of homophone recognition differences on subsequent evaluation results.

[0140] The input of the keyword excitation module includes the feature coding of each keyword in the reference answer. Considering that the conventional speech recognition model is easily confused with homophones, mainly because the pronunciation information is not considered. For this reason, the feature coding of the keyword in this embodiment may include the feature coding of the subword hotword_sub and the phoneme level coding of the subword hotword_phn.

[0141] In one possible implementation, firstly, duplicate words in the reference answer are removed, and then the following processing steps are performed for each duplicated word:

[0142] 1) Get the subword sequence of a word;

[0143] 2) Record the length and position of each subword:

[0144] 3) Obtain the pronunciation sequence of the word and its corresponding subword pronunciation sequence;

[0145] 4) Record the length of each subword’s pronunciation sequence and the position of the subword’s pronunciation sequence.

[0146] The above subword sequence, subword length, and subword position are concatenated together to form the feature code of the subword, hotword_sub. The above subword pronunciation sequence, subword pronunciation sequence length, and subword pronunciation sequence position are concatenated together to form the phoneme-level code of the subword, hotword_phn.

[0147] It is understandable that in the above process of concatenating different data, a set value can be used for padding, for example, -1 is used for padding to maintain the consistency of dimensions when concatenating data.

[0148] like Figure 5 As shown, an optional composition structure of a speech recognition model is provided, in which the optional composition structure of a decoder Encoder and a keyword excitation module is exemplified.

[0149] For the keyword excitation module, when performing feature processing on the feature encoding hotword_sub of the subword and the phoneme-level encoding hotword_phn of the subword, Conformer can be used for feature encoding. Compared with LSTM, Conformer has better encoding effect.

[0150] In addition, a RAM (Real Additive Margin)-softmax layer can also be set in the keyword excitation module. The traditional softmax method only considers the category information of data points during training, but ignores the distance relationship between different data points. RAM-softmax introduces the concept of marginal distance on this basis, and increases the classification robustness of data points by adding margins. In simple terms, it increases the distance between data points of the same type and reduces the distance between data points of different types, thereby improving classification accuracy.

[0151] In some embodiments of the present application, considering the actual oral assessment scenario, it is necessary to assess all candidates participating in the oral test, that is, the number of candidates to be assessed is multiple. According to the traditional method, the answer audio of each candidate needs to be recognized in sequence, which seriously reduces the efficiency of the overall oral assessment. This embodiment provides a solution:

[0152] In this embodiment, a recognition core is configured for each examinee, and an inference core corresponding to the multiple recognition cores is configured. A speech recognition model is deployed in the inference core, and the speech recognition model may include an encoder and a decoder.

[0153] The process of identifying the answer text corresponding to each candidate's answer audio may include:

[0154] Through the recognition core corresponding to each candidate, the input features of the speech recognition model are constructed, and the inference sequence number of the recognition core is marked and sent to the inference core. The inference result returned by the inference core is received, and the final recognized answer text is output based on the inference result.

[0155] The inference core obtains the input features carrying the inference sequence number sent by each recognition core, and concatenates the input features of multiple channels according to the inference sequence number, and sends them to the speech recognition model for inference, and returns the inference results to the corresponding recognition core according to the inference sequence number.

[0156] Combination Figure 6 , which illustrates the process of completing the recognition of the answer audio of each candidate through multiple recognition cores and one reasoning core.

[0157] The input features of the speech recognition model are constructed through the recognition core corresponding to each candidate, and the inference sequence number of the recognition core is marked and sent to the inference core. After detecting the output of the decoder returned by the inference core (the probability score of the recognition result at the current moment), the set search algorithm is executed ( Figure 6Taking the BeamSearch algorithm as an example), that is, searching the recognition result path according to the set search algorithm based on the recognition result probability score at the time, and judging whether the set search algorithm stops searching, that is, judging whether the search stop condition is reached. If not, constructing the output of the search algorithm and sending it to the inference core for the decoder of the speech recognition model to infer again. If so, outputting the final recognized answer text.

[0158] The inference core obtains the input features carrying the inference sequence number sent by each recognition core, and splices multiple data according to the inference sequence number, that is, splices the multiple input features and sends them to the encoder in the speech recognition model for inference, records the output of the encoder according to the inference sequence number, splices the output of the encoder again and sends it to the decoder for inference, records the recognition result probability score of the decoder output at the current moment according to the inference sequence number, and returns the output of the decoder to the corresponding recognition core according to the inference sequence number.

[0159] By adopting the system configuration of multiple recognition cores and one reasoning core provided in this embodiment, the recognition process of audio data of multiple different candidates can be executed in parallel by the multiple recognition cores, and the reasoning core performs reasoning after the multiple input features are spliced. In view of the characteristics of deep learning matrix operations, the number of batches can be increased to improve the overall computing efficiency. The reasoning core grasps the original bits of the data according to the reasoning sequence number, performs multi-batch reasoning, and returns the reasoning results to the recognition core corresponding to the reasoning sequence number, thereby improving the overall computing efficiency.

[0160] In another possible implementation, if the speech recognition model deployed in the inference core adopts Figure 5 The architecture shown in the figure includes an encoder, a decoder and a keyword excitation module. The input features of the speech recognition model constructed in the recognition core may include: the acoustic features of the answer audio (as the input features of the encoder), and the features of each keyword in the reference answer (as the input features of the keyword excitation module).

[0161] Correspondingly, in the inference core, multiple acoustic features are spliced ​​according to the inference sequence number, multiple keyword features are spliced, the acoustic feature splicing results are sent to the encoder for inference, the keyword feature splicing results are sent to the keyword excitation module for inference, the outputs of the encoder and the keyword recording module are recorded respectively according to the inference sequence number, the output of the encoder is spliced ​​again and sent to the decoder for inference, the output of the decoder is recorded according to the inference sequence number, and the output of the decoder is weightedly fused with the output of the keyword recording module according to the inference sequence number to obtain the recognition result probability score at the current moment, and return it to the corresponding recognition core according to the inference sequence number, so that the recognition core can search for the recognition result path through the search algorithm.

[0162] In some embodiments of the present application, the calibration model obtained in the aforementioned step S120 is introduced.

[0163] In a possible implementation, the calibration model may include: a feature extraction model corresponding to each examinee, and a regression model.

[0164] Among them, the feature extraction model corresponding to each examinee is a model trained by using the examinee's answer text in an unsupervised training method, and the feature extraction model is used to extract text features of the corresponding examinee's answer text.

[0165] Exemplarily, the feature extraction model can adopt a restricted Boltzmann machine model RBM, an autoencoder or other types of neural network models. Take the RBM model as an example:

[0166] The uni / bi / tri ternary modeling unit can be used to train the RBM model corresponding to each candidate's answer text in an unsupervised training manner.

[0167] The regression model in the calibration model is trained by taking the text features and reasoning scores of the calibration candidates' answer texts as input and taking the expert scores of the calibration candidates as output.

[0168] That is, for the answer text of the calibration candidate, the text features of the answer text are first extracted through the feature extraction model corresponding to the candidate, and the text features of the calibration candidate and its reasoning score are used as input. The expert score of the calibration candidate is predicted through the regression model to obtain the final score result.

[0169] In a possible implementation, the number of regression models may be one, and the regression model may be trained using the text features and reasoning scores of the calibration examinees' answer texts as training samples and the expert scores of the calibration examinees as sample labels.

[0170] In another possible implementation, the regression model may include a first regression model and a second regression model.

[0171] The first regression model is trained with the text features of the test takers' answer texts as input and the expert scores of the test takers as output.

[0172] The second regression model is trained by taking the expert scoring results predicted by the first regression model for the calibration candidates and the reasoning scores of the calibration candidates as inputs, and taking the expert scores of the calibration candidates as outputs.

[0173] By setting up the first and second regression models, the first regression model predicts a rough range of expert scores based on the text features of the examinee's answer text, and the second regression model further predicts a more accurate expert score based on the rough range of expert scores and inference scores predicted by the first regression model, thereby improving the accuracy of the scoring results.

[0174] The first regression model may adopt various types of regression models, such as a ridge regression model.

[0175] The number of the second regression model can be one or more, and the type can also adopt multiple types of regression models. Exemplarily, the second regression model includes three different types of regression models, namely: linear regression model, decision tree model and support vector machine model.

[0176] The scoring result is predicted by each type of second regression model respectively, and the prediction results of the second regression models of different types are integrated, such as averaging, to obtain the final scoring result.

[0177] In a possible implementation, taking the calibration model including the feature extraction model, the first regression model and the second regression model as an example, step S130, scoring by using the calibration model according to the answer text and the reasoning score of each examinee to obtain the final score of each examinee, may include:

[0178] For each candidate:

[0179] The examinee's answer text is sent to the corresponding feature extraction model to extract the text features of the examinee's answer text.

[0180] The text features of the examinee's answer text are sent to the first regression model to obtain the first output score.

[0181] The first score and the examinee's reasoning score are fed into the second regression model to obtain the final score as output.

[0182] Reference Figure 7 , Figure 7 A flow chart of a spoken language assessment method provided in an embodiment of the present application.

[0183] The oral assessment method can be specifically divided into three stages: pre-calibration processing, calibration process and post-calibration processing.

[0184] Processing before calibration:

[0185] Collect all the test-taking data of the examinees, including questions, reference answers and audio of the answers. Use the test paper analysis module to analyze the answering data and obtain the questions and reference answers. Use the silence segment detection module to detect the non-silent segments in the answering audio, and use the audio feature extraction module to extract the acoustic features of the non-silent segments and send them to the recognition module.

[0186] The recognition module can utilize the acoustic features of the input, and in combination with the key words in the reference answer, adopt the speech recognition model for recognition to obtain the recognized answer text.

[0187] Combined with the candidate's answer text and answer data, the reasoning score of the candidate is obtained through the reasoning scoring module. The reasoning scores and answer texts of all candidates constitute the initial evaluation results of all candidates.

[0188] During the calibration process:

[0189] The calibration candidates for expert scoring are selected from all the candidates, and the experts score the test answers of the calibration candidates to obtain the expert scores of the calibration candidates. The final evaluation results of the calibration candidates are composed of the expert scores, test answers and reasoning scores of the calibration candidates.

[0190] From the initial assessment results of all candidates, the answer text of each candidate is extracted respectively, and the RBM model corresponding to the answer text of each candidate is trained using an unsupervised training method.

[0191] For the calibration candidates, the expert scores of the calibration candidates are extracted, and the text features of the calibration candidates' answer texts are extracted through the RBM model corresponding to the calibration candidates' answer texts. The text features and expert scores are used to train the ridge regression model in a supervised manner.

[0192] Furthermore, for the calibration examinees, the scores and reasoning scores output by the ridge regression model are used as input, and the expert scores are used as output to supervise the training of the three regression models. Here, the three regression models can include three different types of regression models, such as linear regression models, decision tree models, and support vector machine models.

[0193] The RBM model, ridge regression model and triple regression model corresponding to each examinee after the above training are saved as calibration models.

[0194] In the post-calibration processing:

[0195] For each candidate to be evaluated (since the calibration candidates have been scored by experts, all candidates except the calibration candidates can be used as candidates to be evaluated), the RBM features are extracted through the RBM model and sent to the ridge regression model to predict the first score. Furthermore, the first score and the candidate's inference score are sent to the three regression models to obtain the output scores of the three regression models, and the average is taken as the final score of the candidate.

[0196] The following is a description of a spoken language evaluation device provided in an embodiment of the present application. The spoken language evaluation device described below and the spoken language evaluation method described above can be referenced to each other.

[0197] See also Figure 8 , Figure 8 This is a schematic diagram of the structure of a spoken language assessment device disclosed in an embodiment of the present application.

[0198] like Figure 8 As shown, the device may include:

[0199] The answer data acquisition unit 11 is used to acquire the examinee's answer data, wherein the answer data includes the question, the examinee's answer audio and the reference answer;

[0200] An audio recognition unit 12 is used to recognize the answer text corresponding to the answer audio;

[0201] The reasoning scoring unit 13 is used to combine the answer text and the answer data to obtain the reasoning score of the examinee through the configured reasoning scoring model;

[0202] The calibration model acquisition unit 14 is used to acquire the configured calibration model, wherein the calibration model is obtained by pre-training based on the answer text of the calibration examinee, the reasoning score of the calibration examinee and the expert score, and the calibration examinee is a part of the examinees selected from all the examinees participating in this oral test;

[0203] The final score determination unit 15 is used to score each examinee based on the answer text and the reasoning score using the calibration model to obtain the final score of each examinee.

[0204] In a possible implementation, the inference scoring model includes a pronunciation scoring model and a completeness scoring model, and the process in which the inference scoring unit combines the answer text and the answer data to obtain the inference scoring of the examinee through the configured inference scoring model includes:

[0205] Processing the answer audio using the pronunciation scoring model to obtain a pronunciation score output by the model;

[0206] Processing the question, the reference answer and the answer text using the completeness scoring model to obtain a completeness score output by the model;

[0207] The examinee's reasoning score is composed of the pronunciation score and the completeness score.

[0208] In a possible implementation, when the reference answers include more than two, the reasoning scoring unit uses the completeness scoring model to process the question, the reference answer, and the answer text to obtain a completeness score output by the model, including:

[0209] Calculate the similarity between the answer text and each reference answer, and select the target reference answer with the highest similarity;

[0210] The completeness scoring model is used to process the question, the target reference answer and the answer text to obtain a completeness score output by the model.

[0211] In a possible implementation, the apparatus of the present application may further include:

[0212] The calibration candidate extraction unit is used to extract calibration candidates from all candidates participating in this oral test. The process includes:

[0213] All candidates who participated in the oral test were sorted according to their reasoning scores;

[0214] Some calibration candidates are selected in each set scoring range to form all the calibration candidates.

[0215] In a possible implementation, the calibration model includes: a feature extraction model corresponding to each examinee, and a regression model;

[0216] The feature extraction model corresponding to each examinee is a model trained by using the examinee's answer text in an unsupervised training manner, and the feature extraction model is used to extract text features of the corresponding examinee's answer text;

[0217] The regression model is trained by taking the text features and reasoning scores of the test-taking examinees' answer texts as inputs and taking the expert scores of the test-taking examinees as outputs.

[0218] In a possible implementation, the regression model includes: a first regression model and a second regression model;

[0219] The first regression model is trained by taking the text features of the test-taking examinee's answer text as input and taking the expert score of the test-taking examinee as output;

[0220] The second regression model is trained by taking the expert scoring result predicted by the first regression model for the calibration examinee and the reasoning score of the calibration examinee as input, and taking the expert score of the calibration examinee as output.

[0221] On this basis, the final score determination unit scores each examinee's answer text and reasoning score using the calibration model to obtain the final score of each examinee, including:

[0222] For each candidate:

[0223] Sending the answer text of the examinee to a corresponding feature extraction model to extract text features of the answer text of the examinee;

[0224] Sending the text features of the answer text of the examinee into the first regression model to obtain a first score as an output;

[0225] The first score and the examinee's reasoning score are sent to the second regression model to obtain a final score as output.

[0226] In a possible implementation, the process of the audio recognition unit recognizing the answer text corresponding to the answer audio includes:

[0227] Extracting acoustic features of the answer audio;

[0228] The acoustic features are input into a configured speech recognition model, which includes an encoder, a decoder and a keyword excitation module. The keyword excitation module encodes the features of each keyword in the reference answer, predicts the sub-word score of each keyword, performs weighted fusion on the sub-word scores predicted by the decoder and the sub-word scores predicted by the keyword excitation module, and determines the recognition result based on the fused sub-word scores as the answer text.

[0229] In a possible implementation, when there are multiple examinees to be evaluated, a recognition core is configured for each examinee, and a reasoning core corresponding to the multiple recognition cores is configured. The process of the audio recognition unit recognizing the answer text corresponding to the answer audio of each examinee includes:

[0230] Construct the input features of the speech recognition model through the recognition core corresponding to each examinee, mark the inference sequence number of the recognition core and send it to the inference core, receive the inference result returned by the inference core, and output the final recognized answer text based on the inference result;

[0231] The inference core obtains the input features carrying the inference sequence number sent by each recognition core, and splices the multiple input features according to the inference sequence number, and sends them to the speech recognition model for inference, and returns the inference result to the corresponding recognition core according to the inference sequence number.

[0232] The present application also provides an electronic device in an embodiment. Fig. 9 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as tablet computers, learning machines, teaching terminals, laptop computers, etc. Fig. 9 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0233] like Fig. 9 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603 to implement the oral evaluation method of the aforementioned embodiment of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0234] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Fig. 9 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0235] An embodiment of the present application also provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any oral assessment method provided in the embodiment of the present application.

[0236] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any oral assessment method provided in the embodiment of the present application.

[0237] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0238] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0239] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0240] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

[0241] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.

Claims

1. A spoken language assessment method, characterized in that: include: Obtaining the examinee's answer data, wherein the answer data includes the question, the examinee's answer audio and reference answers; Identify the answer text corresponding to the answer audio, combine the answer text and the answer data, and obtain the reasoning score of the examinee through the configured reasoning scoring model; Obtaining a configured calibration model, wherein the calibration model is obtained by pre-training based on the answer texts of the calibration candidates, the reasoning scores of the calibration candidates, and the expert scores, wherein the calibration candidates are selected from all the candidates participating in this oral test; The calibration model is used to perform scoring based on the answer text and the reasoning score of each examinee to obtain a final score for each examinee.

2. The method according to claim 1, characterized in that The inference scoring model includes a pronunciation scoring model and a completeness scoring model; The process of obtaining the reasoning score of the examinee by combining the answer text and the answer data through the configured reasoning score model includes: Processing the answer audio using the pronunciation scoring model to obtain a pronunciation score output by the model; Processing the question, the reference answer and the answer text using the completeness scoring model to obtain a completeness score output by the model; The examinee's reasoning score is composed of the pronunciation score and the completeness score.

3. The method according to claim 2, characterized in that When the reference answers include more than two, the process of using the completeness scoring model to process the question, the reference answer and the answer text to obtain a completeness score output by the model includes: Calculate the similarity between the answer text and each reference answer, and select the target reference answer with the highest similarity; The completeness scoring model is used to process the question, the target reference answer and the answer text to obtain a completeness score output by the model.

4. The method according to claim 1, characterized in that: The process of selecting the candidates for calibration from all the candidates taking the oral test includes: All candidates who participated in the oral test were sorted according to their reasoning scores; Some calibration candidates are selected in each set scoring range to form all the calibration candidates.

5. The method according to claim 1, characterized in that: The calibration model includes: a feature extraction model corresponding to each examinee, and a regression model; The feature extraction model corresponding to each examinee is a model trained by using the examinee's answer text in an unsupervised training manner, and the feature extraction model is used to extract text features of the corresponding examinee's answer text; The regression model is trained by taking the text features and reasoning scores of the test-taking examinees' answer texts as inputs and taking the expert scores of the test-taking examinees as outputs.

6. The method according to claim 5, characterized in that The regression model includes: a first regression model and a second regression model; The first regression model is trained by taking the text features of the test-taking examinee's answer text as input and taking the expert score of the test-taking examinee as output; The second regression model is trained by taking the expert scoring result predicted by the first regression model for the calibration examinee and the reasoning score of the calibration examinee as input, and taking the expert score of the calibration examinee as output.

7. The method according to claim 6, characterized in that The process of scoring each examinee's answer text and reasoning score using the calibration model to obtain a final score for each examinee includes: For each candidate: Sending the answer text of the examinee to a corresponding feature extraction model to extract text features of the answer text of the examinee; Sending the text features of the answer text of the examinee into the first regression model to obtain a first score as an output; The first score and the examinee's reasoning score are sent to the second regression model to obtain a final score as output.

8. The method according to claim 1, characterized in that The process of identifying the answer text corresponding to the answer audio includes: Extracting acoustic features of the answer audio; The acoustic features are input into a configured speech recognition model, which includes an encoder, a decoder and a keyword excitation module. The keyword excitation module encodes the features of each keyword in the reference answer, predicts the sub-word score of each keyword, performs weighted fusion on the sub-word scores predicted by the decoder and the sub-word scores predicted by the keyword excitation module, and determines the recognition result based on the fused sub-word scores as the answer text.

9. The method according to claim 1, characterized in that: When there are multiple examinees to be evaluated, a recognition core is configured for each examinee, and a reasoning core corresponding to the multiple recognition cores is configured; The process of identifying the answer text corresponding to each candidate's answer audio includes: Construct the input features of the speech recognition model through the recognition core corresponding to each examinee, mark the inference sequence number of the recognition core and send it to the inference core, receive the inference result returned by the inference core, and output the final recognized answer text based on the inference result; The inference core obtains the input features carrying the inference sequence number sent by each recognition core, and splices the multiple input features according to the inference sequence number, and sends them to the speech recognition model for inference, and returns the inference result to the corresponding recognition core according to the inference sequence number.

10. A spoken language evaluation device, characterized in that: include: The answer data acquisition unit is used to acquire the examinee's answer data, wherein the answer data includes the question, the examinee's answer audio and the reference answer; An audio recognition unit, used to recognize the answer text corresponding to the answer audio; A reasoning scoring unit, used to combine the answer text and the answer data to obtain the reasoning score of the examinee through a configured reasoning scoring model; A calibration model acquisition unit, used to acquire a configured calibration model, wherein the calibration model is obtained by pre-training based on the answer texts of the calibration examinees, the reasoning scores of the calibration examinees and the expert scores, and the calibration examinees are some examinees selected from all examinees participating in this oral test; The final score determination unit is used to score each candidate based on the answer text and the reasoning score using the calibration model to obtain the final score of each candidate.

11. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the spoken language evaluation method according to any one of claims 1 to 9.

12. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the spoken language evaluation method according to any one of claims 1 to 9 is implemented.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, each step of the spoken language evaluation method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Calibration optimization method and system for speaking test evaluation

    CN104464423A

  • Scoring method, device and equipment for oral test, storage medium and program product

    CN114333787A

  • Subjective item scoring method, model training method, computer equipment and storage medium

    CN114357964A

  • Scoring method and training method for spoken language questions and answers, computer equipment and storage medium

    CN114360537A

  • Model training method and device, speaking topic evaluation method and device and computer storage medium

    CN118335111A

Cited By

  • Oral answer detection method, device and equipment and storage medium

    CN121054030A