An education large model evaluation method for a spoken language practice scene

The educational large-scale model evaluation method, which integrates multimodal feature fusion and dynamic penalty mechanism, solves the problems of single evaluation dimensions and manual dependence in existing technologies. It realizes multi-dimensional automated evaluation of oral practice of large models, improving evaluation efficiency and accuracy.

CN120164492BActive Publication Date: 2025-11-28BEIJING NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510482906.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-11-28
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Existing oral assessment methods mainly focus on pronunciation accuracy and fluency, neglecting the multi-dimensional performance of large models in oral dialogues, and relying on manual evaluation leads to inefficiency and strong subjectivity.

Method used

Employing multimodal feature fusion technology and a dynamic penalty mechanism, this system utilizes a multi-dimensional evaluation framework encompassing speech recognition, speech generation, and content generation. By combining deep learning models such as MultiPA, HuBERT, RoBERTa, and the prompt large model, it automatically assesses multiple dimensions of speech recognition, including accuracy, fluency, and prosody generation quality, providing objective evaluation results.

Benefits of technology

It enables multi-dimensional evaluation of large-scale oral practice, reduces manual costs, improves evaluation efficiency, provides intuitive evaluation results, facilitates users in selecting suitable oral practice products, and supports personalized learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164492B_ABST
    Figure CN120164492B_ABST
Patent Text Reader

Abstract

The application discloses a kind of education big model evaluation methods for oral practice scene, comprising: obtaining speech data and corresponding standard text, by speech data and test model are interacted, obtain the recognition text, audio data and answer text provided by test model;The standard text is compared and evaluated with the recognition text, and the speech recognition accuracy is obtained;Based on multimodal feature fusion, the pronunciation accuracy, fluency and prosody of audio data are evaluated, and the audio score is obtained;Based on dynamic penalty mechanism and the prompt big model constructed, the grammar accuracy of answer text is evaluated, based on the prompt big model constructed, the on-topic degree of answer text is evaluated, and the answer text score is obtained;The speech recognition accuracy, audio score and answer text score are comprehensively summarized, and the final evaluation result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of education large model evaluation and oral evaluation, and particularly relates to an education large model evaluation method for an oral practice scenario. BACKGROUND

[0002] Oral practice refers to an activity of improving individual language oral expression ability through dialogue, imitation, role-playing and the like, and it plays an important role in language learning. With the continuous improvement of the capabilities of large-scale language models in speech recognition and speech generation, they are gradually applied to oral dialogue interaction between users and systems, and a series of artificial intelligence products in oral practice scenarios are derived. These products not only can understand the user's voice input, but also can generate natural and fluent oral answers, thereby helping users learn foreign languages and improve oral skills. At present, there are many oral practice products for primary and secondary school students, and how to evaluate the quality of their answers has become a problem to be considered. However, the related technical solutions mainly have the following deficiencies:

[0003] 1. Evaluation angle is not applicable: the existing oral evaluation methods mainly focus on providing pronunciation feedback to users, and most of them evaluate the pronunciation accuracy and fluency of people, but are not applicable to evaluating the answers of large models. There is still a gap in the evaluation method for the generated content of large model products.

[0004] 2. Evaluation dimension is single: the current oral evaluation method mainly evaluates the quality of speech generation, such as the GOP (Goodness of Pronunciation) method based on Hidden Markov Model (HMM), which mainly evaluates pronunciation accuracy. Oral practice is completed through dialogue between users and large model products, and more attention is paid to interactivity with users. Most products and methods only focus on the accuracy and fluency of pronunciation, and ignore whether the model answers the question well, whether the grammar is used accurately, and whether the content is rich. The evaluation dimension is single, which leads to the fact that the existing evaluation system cannot fully reflect the performance of large models in actual oral dialogue, and therefore there is still a lack of a more comprehensive oral evaluation method for large models.

[0005] 3. Subjective factors have a great influence: the MOS (Mean Opinion Score) method used to evaluate the naturalness of synthesized speech or oral pronunciation is a subjective evaluation method based on human, which can be used to evaluate whether the speech generated by the model sounds close to human pronunciation. However, this method relies on human subjective feelings and directly reflects the subjective feelings of people on the naturalness of speech, and is greatly influenced by individual differences and subjective factors of evaluators, and the results will have a certain degree of subjectivity and uncertainty.

[0006] 4. High labor cost: The existing method for evaluating the naturalness of pronunciation, the subject matter of the content, and other subjective evaluations all require human involvement. Participants need to spend time listening and scoring, which is inefficient in large-scale speech database quality evaluation. At the same time, subjective evaluation requires strict operation procedures and standardized test guides to reduce the impact of operation errors, which also increases the labor cost and complexity. SUMMARY

[0007] To solve the above technical problems, the present application provides a large education model evaluation method for oral practice scenarios to solve the problems existing in the prior art.

[0008] To achieve the above purpose, the present application provides a large education model evaluation method for oral practice scenarios, comprising:

[0009] Obtain speech data and corresponding standard text, interact with the test model through the speech data, and obtain the recognized text, audio data and answer text provided by the test model;

[0010] Compare and evaluate the standard text and the recognized text to obtain the speech recognition accuracy;

[0011] Evaluate the pronunciation accuracy, fluency and prosody of the audio data based on multi-modal feature fusion to obtain the audio score;

[0012] Evaluate the grammatical accuracy of the answer text based on a dynamic penalty mechanism and a constructed prompt large model, and evaluate the on-topic degree of the answer text based on the constructed prompt large model to obtain the answer text score;

[0013] Comprehensively summarize the speech recognition accuracy, audio score and answer text score to obtain the final evaluation result.

[0014] Optionally, the process of comparing and evaluating the standard text and the recognized text comprises:

[0015] Error evaluation is performed on the recognized text and the standard text to obtain the word error rate, character error rate and word information loss rate; the word error rate, character error rate and word information loss rate are averaged and integrated by dynamic weighting to obtain the speech recognition accuracy:

[0016] Accuracy=w wer *max(0,1-WER)+w cer *max(0,1-CER)+w wil *max(0,1-WIL)

[0017] Wherein, Accuracy represents the speech recognition accuracy, w wer is the weight of the word error rate WER, wcer is the weight of the character error rate CER, w wil is the weight of the word information loss rate WIL;

[0018] Among them, the weight in dynamic weighting is adjusted according to the evaluation requirements.

[0019] Optionally, the audio data is evaluated by a MultiPA model, and a data processing process of the MultiPA model includes:

[0020] The audio data is feature-extracted by HuBERT to obtain frame-level features;

[0021] The audio data is transcribed by an ASR model, wherein the ASR model provides a recognized transcription text when transcribing, a Whisper base.en generates an alternative target transcription text in an inference stage, and text features are obtained;

[0022] The recognized transcription text, the target transcription text and the audio data are aligned by a Charsiu model to obtain alignment information, and the alignment information includes word-level alignment features, phoneme-level alignment features and phoneme vectors;

[0023] The target transcription text and the recognized transcription text are extracted by a RoBERTa model to obtain corresponding embedding vectors, and are spliced into word-level semantic embeddings;

[0024] The frame-level features are processed by average pooling according to the alignment information to obtain word-level features; and the word-level features are aggregated to obtain sentence-level features;

[0025] The phoneme features and the word-level features are aligned and fused according to the alignment information to obtain phoneme-level features;

[0026] The audio features, the text features, the word-level semantic embeddings and the features after multi-granularity alignment are fused by a Transformer encoder, different numbers of Transformer encoders are used to process different types of features respectively to realize deep fusion, the fusion results are spliced to obtain a final unified feature representation, and the final unified feature representation is processed by a linear layer to obtain sentence-level pronunciation accuracy, fluency and prosody results, wherein the features after multi-granularity alignment include word-level features, sentence-level features and phoneme-level features.

[0027] Optionally, the process of evaluating pronunciation accuracy of the audio data includes:

[0028] The first alignment feature includes the time proportion of the target phoneme and the recognized phoneme, the time interval between the phoneme and the adjacent phoneme, the time alignment error of the target phoneme and the recognized phoneme, and the matching probability of the target phoneme and the recognized phoneme, i.e., the alignment probability.

[0029] The factor accuracy, i.e., the pronunciation accuracy score, is obtained by calculating the matching degree of the target phoneme sequence and the recognized phoneme sequence through the Levenshtein distance.

[0030] Optionally, the fluency evaluation process of the audio data includes:

[0031] The word duration and interval features in the alignment feature are obtained, and the fluency score of the speech is calculated according to the word duration and interval features; according to the interval feature, a pause analysis result is obtained; the syntactic and semantic coherence of the target transcription text and the recognized transcription text are compared and evaluated by embedding analysis of the target transcription text and the recognized transcription text, to obtain a semantic consistency score; and the fluency score, the pause analysis result and the semantic consistency score are input into a linear regression layer to obtain a fluency score.

[0032] Optionally, the process of prosody evaluation of the audio data includes:

[0033] The fundamental frequency feature of the audio data is extracted through HuBERT, and the energy level and alignment degree are obtained according to the fundamental frequency feature; the rhythm score is calculated according to the alignment degree; and whether the change of the fundamental frequency conforms to the intonation rule of the language is analyzed according to the fundamental frequency feature, the energy level and the rhythm score, and a prosody score is obtained through a linear regression layer.

[0034] Optionally, the process of rhythm score is:

[0035]

[0036] Wherein, RMSE is the root mean square error of the rhythm, and Max Timing Difference is the maximum time error of the alignment.

[0037] Optionally, the process of syntax accuracy evaluation of the answer text includes:

[0038] The answer text is checked for syntax by the constructed prompt large model and statistics are obtained, to obtain the number and type of syntax errors, and a syntax accuracy score is calculated according to the number and type of syntax errors:

[0039] accuracy=max(0,(1-error_count / word_count)*length_penalty)

[0040] wherein error_count is the number of syntax errors, word_count is the total number of words, and length_penalty is a length penalty coefficient adjusted according to the sentence length of the answer text.

[0041] Optionally, the process of evaluating the on-topic degree of the answer text comprises:

[0042] The answer text is scored according to the designed rules by the constructed prompt large model to obtain the on-topic degree score.

[0043] Optionally, the speech recognition accuracy, audio score, and answer text score are comprehensively summarized by averaging or weighted calculation.

[0044] Compared with the prior art, the present application has the following advantages and technical effects:

[0045] The evaluation method proposed in the present application, after experimental testing, exhibits a series of significant and beneficial effects compared to the prior art.

[0046] 1. Reducing labor costs and improving efficiency

[0047] The present application significantly reduces the dependence on manual evaluation and reduces labor costs through automated oral evaluation technology. The MultiPA model combines self-supervised learning, multi-task evaluation, and deep learning technology, which does not require a large number of human interventions compared to traditional methods, thereby improving the efficiency and scalability of the evaluation.

[0048] 2. Multi-dimensional comprehensive evaluation

[0049] The present application proposes a more comprehensive evaluation method for oral practice large model products, which not only evaluates pronunciation accuracy, fluency, and rhythm, but also covers multiple dimensions such as grammar use and content on-topic degree. This method can comprehensively reflect the performance of the large model in actual oral dialogue, filling the gap in the evaluation method of large model generated content in the prior art. At the same time, the proposed oral practice large model evaluation method reduces the interference of subjective factors, provides more objective and accurate evaluation results, and makes the evaluation results more secure and reliable.

[0050] 3. Intuitive comparison of evaluation results

[0051] The evaluation result of the present application is intuitive and easy to understand, which can clearly show the performance of the oral practice product in various aspects, including pronunciation, grammar, content, etc., facilitating the user to understand and use. And it can generate clear and quantitative evaluation report, enhancing the intuitiveness and comparability of the evaluation result, so that the user and the developer can identify the advantages and disadvantages of each oral practice product at a glance. This comparison provides a comprehensive perspective, helping users to choose the most suitable oral practice product according to specific learning needs and goals, so as to realize personalized and efficient language learning. The present application also provides an effective evaluation tool for oral practice products in the field of intelligent education, and provides strong support for the further development and application of the education large model oral practice evaluation technology. BRIEF DESCRIPTION OF DRAWINGS

[0052] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments of this application and their description together with the drawings serve to explain the application. In the drawings:

[0053] Figure 1 The evaluation framework diagram of the embodiment of the present application. DETAILED DESCRIPTION

[0054] It should be noted that the embodiments and features in the present application can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0055] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0056] In view of the above problems, how to comprehensively evaluate the model in the oral dialogue between the user and the large model, design a new evaluation method, which can objectively and comprehensively evaluate the performance of the large model in the oral practice scene, especially from the aspects of pronunciation accuracy, fluency, rhythm, content quality, etc. Comprehensive analysis from multiple dimensions has become an important direction of current technical development. The present application aims to solve the problems in the prior art, and proposes a comprehensive and multi-dimensional large model oral evaluation method by using advanced deep learning technology, which helps personalized learning and intelligent evaluation in the field of education.

[0057] The present application proposes a multi-dimensional education large model evaluation method for oral practice scene, which relates to a method for evaluating the multi-dimensional capabilities of a large model in an oral practice scene, such as speech recognition, pronunciation quality, and content generation.

[0058] The method constructs an evaluation framework including speech recognition, speech generation quality, and content generation quality by quantization algorithm and multi-modal feature fusion technology, and introduces a dynamic weight distribution mechanism and a deep feature alignment method, which is significantly different from the single dimension evaluation mode of the prior art, and aims to comprehensively reflect the performance of the large model in the actual spoken dialogue. The technical innovation points of the scheme are embodied in the following structured and process design:

[0059] Firstly, in order to construct this efficient and easy-to-implement quantitative evaluation method, the present application designs an evaluation framework in which each module can integrate the evaluation results of the large model in multiple dimensions such as speech recognition, speech generation and content generation, and finally output a comprehensive score. Figure 1 As shown in the figure, the main process and data flow of each link are shown.

[0060] 1. Speech input: the user interacts with the large model through speech input.

[0061] 2. Data preparation: the input speech signal is recognized by the test model and converted into text form as the text data of the large model speech recognition, and the standard text data is prepared for comparison, the audio of the large model answer is used as the test data of the speech generation quality, and the text of the answer is used as the test data of the content generation quality.

[0062] 3. Speech recognition evaluation: first, compare the text of the large model speech recognition with the standard text to evaluate the speech recognition accuracy.

[0063] 4. Speech generation quality evaluation: comprehensive score of pronunciation, fluency and prosody of the audio speech generated by the large model.

[0064] 5. Content generation quality evaluation: evaluate the answer text content generated by the model in the dimensions of grammatical accuracy and topic relevance.

[0065] 6. Comprehensive score: through the average algorithm of each dimension score, the scores of each dimension are summarized into a comprehensive score to generate the final evaluation result.

[0066] The present application also defines the evaluation dimensions, evaluation content and score standard, and table 1 is a spoken dialogue evaluation dimension table, the specific content is shown in table 1 as follows:

[0067] Table 1

[0068]

[0069]

[0070]

[0071] Meanwhile, the application also innovatively constructs a multi-dimensional evaluation method. Unlike traditional single-index evaluation methods, the scheme adopts a dynamic weight distribution mechanism and multi-modal feature fusion technology, combined with a speech fluency quantification algorithm, a dynamic penalty mechanism for grammatical accuracy, etc., to realize full-link objective evaluation from acoustic features to grammatical content.

[0072] 1. Speech recognition accuracy evaluation

[0073] Speech recognition accuracy is one of the key indicators for evaluating large models in spoken language practice scenarios. The application uses a comprehensive scoring algorithm based on error rate dynamic weighting to break through the limitations of traditional single indicators (such as WER). The following formula is used to realize the normalization processing of multi-dimensional errors:

[0074] Accuracy=w wer *max(0,1-WER)+w cer *max(0,1-CER)+w wil *max(0,1-WIL)

[0075] Wherein, the weight distribution is w wer =0.4, w cer =0.3, w wil =0.3, the influence of word error rate (WER), character error rate (CER) and word information loss rate (WIL) on speech recognition performance can be balanced through dynamic (adjustable at any time) weighted average, forming a more objective quantitative score.

[0076] The evaluation algorithm can accurately evaluate the understanding ability of the large model to the user's speech input, and ensure the accuracy of the model in the speech recognition link.

[0077] 2. Speech generation quality evaluation

[0078] Speech generation quality is an important aspect of evaluating the performance of large models in spoken language interaction. Based on the open-source multi-task speech pronunciation evaluation model MultiPA, the application extracts multi-modal features and fuses them to realize joint evaluation of sentence-level pronunciation accuracy, fluency, and prosody. The core algorithm includes:

[0079] 2.1 Fluency quantification algorithm

[0080] Through speech speed consistency analysis and pause detection model, the speech naturalness is calculated:

[0081]

[0082] Wherein, t iEnd time of the ith word, Expected Timing is the ideal word interval time, combined with the interval features extracted by the alignment tool, dynamically detect unnatural pauses and repetition.

[0083] 2.2 Prosody scoring model

[0084] Introduce the joint analysis of fundamental frequency (F0), energy (Energy) and rhythm (Rhythm), and define the rhythm scoring formula:

[0085]

[0086] Where RMSE is the root mean square error of rhythm, and Max Timing Difference is the maximum time error of alignment. By smoothing the fundamental frequency curve and aligning the target audio, the naturalness of intonation is quantified, and the objective scoring consistency with manual scoring is more than 90%.

[0087] Using the MultiPA model can comprehensively evaluate the pronunciation quality at the sentence level (accuracy, fluency, prosody) and word level (accuracy, stress), and provide multi-dimensional feedback. The model can consider multiple aspects of speech to provide comprehensive pronunciation quality evaluation. In the present application, only the sentence-level score is used as the pronunciation quality evaluation score.

[0088] 3. Content generation quality evaluation

[0089] Content generation quality evaluation focuses on evaluating whether the content expressed by the large model in spoken language generation is accurate and reasonable, clear, relevant to the context, and consistent with the theme. The evaluation dimensions include grammatical accuracy and topic relevance.

[0090] 3.1 Grammatical accuracy dynamic penalty mechanism

[0091] To solve the problem of long text grammatical accuracy evaluation bias, a penalty coefficient based on sentence length, length_penalty, is proposed, and the accuracy is calculated based on the following formula:

[0092] accuracy=max(0,(1-error_count / word_count)*length_penalty)

[0093] When word_count>10, length_penalty dynamically decays with length (lower limit 0.5), which can effectively balance the grammatical error tolerance rate of long and short sentences.

[0094] 3.2 Topic relevance

[0095] Using a large model, give a prompt containing the definition and example of the topic, based on the semantic correlation between the large model generated content and the input question, design a five-level scoring rule (1-5 points), and output the evaluation score and reason according to the format, get the topic degree score, and ensure that the evaluation results meet objectivity and logicality at the same time.

[0096] Finally, the evaluation scores of speech recognition, speech generation quality, and content generation quality are integrated, and the evaluation results are obtained by calculating the average of each dimension.

[0097] The above structured design realizes the triple breakthrough of evaluation dimension, algorithm accuracy and application scenario, provides a reusable technical framework for oral interaction ability evaluation of education large model, and provides a new technical scheme for intelligent evaluation in education field.

[0098] The above technical solutions are described in detail:

[0099] 1. Data set construction idea

[0100] The speech data and corresponding text data need to be collected as the input audio of the model to ensure that the data set is representative and covers various oral situations. The construction process of the data set is as follows:

[0101] 1.1 Data collection: Collect primary and secondary school English teaching materials, public oral data sets (such as Speechocean) and speech data recorded by language learners as input audio. The input data should include various speech situations, such as daily conversation, academic discussion, and emotional expression. The input speech signal is recognized and converted into text form by the tested model as the test data of speech recognition, and the standard text data is prepared for comparison. The audio of the large model answer is used as the test data of speech generation quality, and the text data of the answer is used as the test data of content generation quality.

[0102] 1.2 Standard text: Collect the text obtained when collecting questions or the transcription text of each speech obtained by the automatic speech recognition (ASR) model as the standard text. Ensure that each speech data has accurate standard text for subsequent evaluation and comparison.

[0103] 1.3 Data preprocessing: Denoise and segment the collected speech data to ensure that the data is clear and meets the evaluation requirements. Then, classify these speech data into corresponding folders for evaluation of different models.

[0104] 2. Evaluation method

[0105] The evaluation method of the present application is divided into multiple dimensions, and the specific evaluation process and calculation method are as follows:

[0106] 2.1 Speech recognition accuracy evaluation

[0107] Speech recognition accuracy is used to evaluate the performance of large models in speech recognition. Standard evaluation metrics, including word error rate (WER), character error rate (CER), and word loss rate (WIL), are used, and the specific steps are as follows:

[0108] (1) Data input: Input the text for speech transcription recognition of the test model and the standard text.

[0109] (2) Error calculation: The error between the recognized text and the standard text is evaluated using the Jiwer library of Python, and the WER, CER and WIL error rates are calculated.

[0110] (3) Weighted average score: Each error rate is first converted into a corresponding accuracy score, and then the accuracy of each speech sample is weighted and averaged according to the set weights. The initial weight allocation is WER=0.4, CER=0.3, WIL=0.3, to obtain the overall speech recognition accuracy score.

[0111] Calculation formula:

[0112] Accuracy = w wer *max(0,1-WER)+w cer *max(0,1-CER)+w wil *max(0,1-WIL) where:

[0113] w wer : This is the weight of the word error rate (WER), set to 0.4.

[0114] w cer : This is the weight of the character error rate (CER), set to 0.3.

[0115] w wil : This is the weight of the word information loss rate (WIL), set to 0.3.

[0116] This formula ensures that each accuracy value is not negative, and it performs a weighted average of different error rates according to preset weights. The weights can also be dynamically adjusted to balance the impact of the three error rates on speech recognition performance in order to obtain a comprehensive accuracy score.

[0117] 2.2 Speech Generation Quality Evaluation

[0118] The evaluation of speech generation quality mainly includes three aspects: pronunciation accuracy, fluency, and prosody. This invention uses the open-source MultiPA model for evaluation, and the detailed implementation mainly includes the following steps:

[0119] 2.2.1 Data Preparation and Loading

[0120] (1) Audio file: Load the input audio file of the test model as a tensor format, and the audio format needs to be converted to.wav in advance. The audio sampling rate is usually set to 16000 Hz.

[0121] (2) ASR transcription: Generate the transcription text of the audio through the ASR model Whisper.

[0122] 2.2.2 Multimodal feature extraction

[0123] Feature extraction is divided into audio features and text features, as well as their alignment information.

[0124] (1) Audio feature extraction: Use the self-supervised learning model HuBERT to extract audio features from the input audio, output frame-level features, capture the timing information and acoustic features of the speech, and these features can capture the details of the speech.

[0125] (2) Text feature extraction: When ASR transcribes, use two versions of the Whisper model to transcribe, Whisperbase.en provides the recognized transcript as the initial recognition result, and Whispermedium.en generates an alternative target transcript in the inference stage to replace the real target text in an open scenario. (Since the open dialogue scenario cannot obtain the real target text), the model uses two different sizes of Whisper versions to obtain text transcriptions of different qualities, so as to better evaluate the pronunciation quality in an open scenario.

[0126] (3) Alignment information: Process the recognized transcript, target transcript, and audio signal through the alignment tool Charsiu, output word-level alignment features: including duration, interval, time difference, Levenshtein distance, etc., phoneme-level alignment features: phoneme duration, phoneme interval, phoneme probability, etc., and phoneme vector: generate a hot encoding phoneme representation for further fusion.

[0127] (4) Semantic embedding feature extraction: Use RoBERTa to extract the semantic embedding of the target transcript and the recognized transcript, output the embedding vector of the target text and the recognized text, and concatenate it as a word-level semantic embedding.

[0128] 2.2.3 Feature fusion

[0129] The core of feature fusion is to align, aggregate, and process features of different modalities (audio and text) and different granularities (frame level, word level, and phoneme level), achieving multi-scale feature fusion.

[0130] (1) Multi-modal and multi-granularity alignment

[0131] Frame-to-word alignment: Using the alignment information provided by Charsiu, the frame-level features extracted by HuBERT are averaged-pooled (frame→word) to generate word-level features.

[0132] Word-to-sentence alignment: The word-level features are further aggregated into sentence-level features (word→sen) through average-pooling.

[0133] Phone-level alignment: Phone features (including phone probabilities and a hot encoding vector) are aligned and fused with word-level features through Charsiu's alignment results to obtain phone-level features.

[0134] The features after multi-granularity alignment include word-level features, sentence-level features, and phone-level features.

[0135] The core purpose of the alignment process is to ensure that information from different modalities (audio and text) can be matched and fused at the same granularity level.

[0136] (2) Feature deep fusion

[0137] The audio features, text features, word-level semantic embeddings, and features after multi-granularity alignment are input into the TransformerEncoder. Different numbers of heads (h = 2, 3, 5, 7) of the Transformer encoder are used to process different types of features, achieving deep fusion of features.

[0138] (3) Feature concatenation (Concat)

[0139] The fused features are concatenated by the "Concat" operation to connect these features in a specific dimension, forming a unified feature representation. This unified representation contains comprehensive information such as the acoustic characteristics, semantic content, and timing features of the speech, which can be used as the basis data for scoring.

[0140] (4) Linear and convolutional layers

[0141] Linear layer (Linear): Processes the final unified sentence-level feature representation after concatenation, outputting sentence-level evaluation scores (such as accuracy, fluency, prosody, etc.).

[0142] Convolutional layer (Conv1d): Processes the final word-level feature representation, outputting word-level evaluation scores (such as word accuracy, stress, etc.).

[0143] 2.2.4 Pronunciation scoring

[0144] The MultiPA model outputs multi-dimensional pronunciation scores, with the sentence level as follows:

[0145] (1) Accuracy: Evaluates the accuracy of pronunciation, whether it matches the standard pronunciation.

[0146] (2) Fluency: Evaluates the coherence and fluency of the speech.

[0147] (3) Prosody: Evaluates the intonation, stress, and rhythm in the speech.

[0148] At the word level, there is accuracy and stress, and in this invention, only the sentence level score is used as the pronunciation quality evaluation score. After the model output, the system can provide feedback based on these scores to evaluate the speech generation quality.

[0149] The specific implementation ideas of the three dimensions of speech generation quality evaluation in the MultiPA model are as follows:

[0150] Pronunciation accuracy evaluation

[0151] (1) Phoneme-level alignment: Use Charsiu tool to align the audio data with the recognized transcription text at the phoneme level. Extract alignment features, including the time proportion (Duration) of the target phoneme and the recognized phoneme, the time interval (Interval) between phonemes and adjacent phonemes, the time alignment error (Time Difference) of the target phoneme and the recognized phoneme, and the matching probability of the target phoneme and the recognized phoneme - alignment probability (Phone Probability):

[0152] (2) Feature calculation: Calculate the matching degree (dislocation, replacement, deletion) of the target phoneme sequence and the recognized phoneme sequence through Levenshtein distance. Generate phoneme accuracy, calculation formula:

[0153]

[0154] The derived pronunciation accuracy score is based on the accuracy of phoneme recognition, the higher the accuracy, the better the evaluation result.

[0155] (3) Model fusion: Combine speech embedding (extracted from HuBERT), alignment features, and phoneme-level features, and fuse features through a multi-layer Transformer encoder. Use a linear layer to regress to generate an accuracy score.

[0156] Fluency evaluation

[0157] (1) Speech rate measurement: Use the word duration (Duration) and interval (Interval) in the alignment features to calculate the consistency of the speech rate.

[0158] (2) Speech Rate Calculation: Assess the fluency of generated speech by calculating the time interval between adjacent syllables (i.e., speech rate). Calculation formula:

[0159]

[0160] t i : End time of the i-th word.

[0161] Expected Timing: Ideal word interval time.

[0162] (3) Pause Analysis: Detect unnatural pause positions and durations (too long or too short) by extracting interval features from alignment features. Mark excessive repeated words or speech delays (e.g., "uh," "um") specifically.

[0163] (4) Semantic Fluency: Use RoBERTa embedding analysis to identify the syntactic and semantic coherence of the text. Compare the target text with the recognized text to assess semantic consistency.

[0164] (5) Model Generation: Integrate speech rate, pause, and semantic fluency features through a Transformer encoder process. Use a linear layer regression to generate a fluency score, with higher fluency resulting in higher scores.

[0165] Prosody Evaluation

[0166] (1) Feature Extraction:

[0167] Fundamental Frequency (F0): Extract the fundamental frequency features of audio frames using HuBERT. Detect the smoothness of the fundamental frequency curve and its alignment with the target audio.

[0168] Energy: Calculate the energy level of speech frames to assess the strength of speech accents.

[0169] Rhythm: Analyze the difference between the rhythm distribution of target words or phonemes and the actual pronunciation.

[0170] (2) Rhythm Score: Calculation formula:

[0171]

[0172] RMSE: Root Mean Square Error of Rhythm.

[0173] Max Timing Difference: Maximum time error of alignment.

[0174] (3) Pitch Naturalness: Compare the target fundamental frequency variation with the actual pronunciation fundamental frequency variation to measure whether the fundamental frequency variation conforms to the pitch rules of the language.

[0175] (4) Model generation: Integrate fundamental frequency, energy and rhythm features, analyze intonation naturalness, and fuse multi-dimensional prosody information through Transformer encoder. Use linear layer to generate prosody score.

[0176] The derived prosody score reflects the naturalness and expressiveness of the speech, and the better the prosody, the higher the score.

[0177] Through the above steps, the MultiPA model can comprehensively evaluate the pronunciation accuracy, fluency and prosody of the model's answers, helping to optimize the quality of the model's spoken language generation.

[0178] 2.3 Content generation quality evaluation

[0179] 2.3.1 Grammar accuracy evaluation

[0180] Grammar accuracy is used to measure the grammatical correctness of the model-generated speech text. The present invention uses a large model for grammar evaluation and introduces a length penalty coefficient for reasonable calculation. Since the total number of words in a sentence has different effects on accuracy, longer sentences may have calculation bias. The more words in a sentence, the higher the accuracy may be even with a small number of errors, which may cause longer sentences to be evaluated as "more accurate". This is unfair, so a length penalty coefficient is added to the accuracy formula to make the calculation formula more accurate. The specific steps are as follows:

[0181] (1) Construct prompt: The prompt gives detailed evaluation rules and example content, and inputs the text generated by the test model into the large model for grammar checking.

[0182] (2) Error statistics: The large model identifies grammatical errors in the text and counts the number and type of errors, such as tense, subject-verb agreement, etc.

[0183] (3) Penalty coefficient principle: For sentences less than 10 words, length_penalty is 1.0 (no penalty). When word_count>10, length_penalty gradually decreases, but not less than 0.5. This method is suitable for sentences of different lengths, while avoiding high penalties for shorter sentences and low accuracy scores for longer sentences.

[0184] (4) Accuracy calculation: Calculate the grammar accuracy through the formula accuracy=max(0,(1-error_count / word_count)*length_penalty).

[0185] error_count is the number of syntax errors, word_count is the total number of words, and length_penalty is the length penalty coefficient, which is used to adjust the accuracy score in long texts.

[0186] Specifically, this sentence length-based penalty coefficient allows longer sentences to have more reasonable accuracy scores. In the calculate_accuracy function, the penalty coefficient is added to make the accuracy calculation more accurate. Only sentences containing errors will be affected by the length penalty, which more smoothly controls the penalty, thereby avoiding the situation where the accuracy of error-free sentences is undesirably reduced.

[0187] 2.3.2 Topic Relevance Evaluation

[0188] Topic relevance evaluation is used to determine whether the content generated by the large model is related to the user's input question or context. The present invention uses a suitable prompt and a large model to evaluate the topic relevance, and the specific steps are as follows:

[0189] (1) Construct prompt: design the prompt according to the rules, detail the scoring criteria, define the topic relevance and relevance, and specify the requirements of each dimension.

[0190] -- Topic Relevance Definition: Topic relevance refers to whether the answer directly addresses the question asked, and whether it closely revolves around the theme or core issue of the conversation. A high topic relevance answer will directly respond to the core of the question and not deviate from the main focus of the question. Example: If the question is "What is the weather like today?", a relevant answer might be "Today the weather is sunny, with a temperature of about 25 degrees." This answer directly answers the question about the weather and does not deviate from the theme.

[0191] -- Relevance Definition: Relevance refers to the degree to which the answer is related to the question asked, even if the answer may not directly answer the question, but provides information related to the question. A high relevance answer may not directly answer the question, but provides background information or indirect answers that help understand the question. Example: For the same question "What is the weather like today?", a relevant but not direct answer might be "It rained all day yesterday, but the weather forecast says it will clear up today." This answer provides information about yesterday's weather and the weather forecast, although it does not directly describe today's weather, but is related to the question and helps understand the weather conditions today.

[0192] In general, an ideal answer should have both high topic relevance and high relevance, i.e. directly answering the question while also providing sufficient background information or relevant details to help the questioner fully understand the question.

[0193] The prompt requires the model to give brief scoring reasons to enhance the transparency and logic of the score. At the same time, the dialogue content is integrated into the prompt so that the model can directly quote and evaluate it.

[0194] (2) Five-level scoring rules: Design five-level scoring rules (1-5 points, 1 point for completely off-topic, 5 points for highly on-topic), according to the relevance and on-topic degree of the model-generated content to the target content, give the corresponding score.

[0195] (3) Large model evaluation: Use a large model to determine whether the generated content is on topic, give a grade score and specific scoring reasons, and give scores according to the rules.

[0196] 3. Evaluation results

[0197] During the evaluation process, the application integrates the evaluation results of each dimension and calculates the final comprehensive evaluation score by averaging. The specific steps are as follows:

[0198] (1) Dimensional scoring: Give a score for each evaluation dimension (such as speech recognition accuracy, pronunciation accuracy, fluency, prosody, grammatical accuracy, on-topic degree, etc.).

[0199] (2) Average synthesis: According to the scores of each dimension, the total score is obtained by calculating the average. Weights can also be set for weighted averaging, for example, speech recognition accuracy may account for 30%, pronunciation accuracy for 25%, fluency and prosody for 20%, and content evaluation for 25%.

[0200] (3) Feedback output: According to the comprehensive score results, generate an evaluation report and provide it to the user or developer, pointing out the strengths and weaknesses of the model in various aspects to help improve it.

[0201] The above is only the preferred specific embodiment of the present application, but the protection scope of the present application is not limited to this, any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An educational large model evaluation method for a spoken language practice scenario, characterized in that, The method comprises the following steps: acquire voice data and corresponding standard text, interact with the test model through the voice data, and acquire the recognized text, audio data and answer text provided by the test model; compare and evaluate the standard text and the recognized text to obtain the voice recognition accuracy; based on multi-modal feature fusion, evaluate the audio data for pronunciation accuracy, fluency and prosody to obtain an audio score; the process of prosody evaluation on the audio data comprises: extract the fundamental frequency features of the audio data through HuBERT, and acquire the energy level and alignment degree according to the fundamental frequency features, calculate the rhythm score according to the alignment degree, analyze whether the change of the fundamental frequency conforms to the intonation rules of the language according to the fundamental frequency features, the energy level and the rhythm score, and obtain the prosody score through the linear regression layer; based on the dynamic penalty mechanism and the constructed prompt large model, evaluate the grammatical accuracy of the answer text, and based on the constructed prompt large model, evaluate the on-topic degree of the answer text to obtain the answer text score; comprehensively summarize the voice recognition accuracy, the audio score and the answer text score to obtain the final evaluation result.

2. The method of claim 1, wherein the process of comparing and evaluating the standard text and the recognized text comprises: error evaluation is performed on the recognized text and the standard text to obtain the word error rate, the character error rate and the word information loss rate; the word error rate, the character error rate and the word information loss rate are averaged and integrated by dynamic weighting to obtain the voice recognition accuracy: wherein the weights in the dynamic weighting are adjusted according to the evaluation requirements. wherein Accuracy denotes speech recognition accuracy, w wer is a weight of a word error rate WER, w cer is a weight of a character error rate CER, w wil is a weight of a word information loss rate WIL; 3. The method of claim 1, wherein the audio data is evaluated by a MultiPA model, and the data processing process of the MultiPA model comprises: extracting features of the audio data by HuBERT to obtain audio features; transcribing the audio data by an ASR model, wherein the ASR model provides a recognized transcription text when transcribing, and a Whisper medium.en generates an alternative target transcription text in the inference stage to obtain text features; aligning the recognized transcription text, the target transcription text and the audio data by a Charsiu model to obtain alignment information, which includes word-level alignment features, phoneme-level alignment features and phoneme vectors; extracting semantic embedding of the target transcription text and the recognized transcription text by a RoBERTa model to obtain corresponding embedding vectors and concatenate them into word-level semantic embedding; performing average pooling on the frame-level features according to the alignment information to obtain word-level features; and aggregating the word-level features to obtain sentence-level features; aligning and fusing the phoneme features and the word-level features according to the alignment information to obtain phoneme-level features. ​ ​ The audio features, text features, word-level semantic embeddings, and multi-granularity aligned features are fused by a Transformer encoder, different types of features are processed by different head number of Transformer encoders to realize deep fusion, the fusion results are spliced to obtain the final unified feature representation, and the final unified feature representation is processed by a linear layer to obtain the sentence-level pronunciation accuracy, fluency, and rhythm results, wherein the multi-granularity aligned features include word-level features, sentence-level features, and phoneme-level features.

4. The method of claim 1, wherein, The process of evaluating the pronunciation accuracy of the audio data includes; The Charsiu model is used to perform phoneme-level alignment on the audio data and the recognized transcription text to obtain first alignment features, and the first alignment features include the time proportion of the target phoneme and the recognized phoneme, the time interval between phonemes and adjacent phonemes, the time alignment error of the target phoneme and the recognized phoneme, and the matching probability of the target phoneme and the recognized phoneme. The matching degree of the target phoneme sequence and the recognized phoneme sequence is calculated by the Levenshtein distance to obtain the phoneme accuracy, i.e., the pronunciation accuracy score.

5. The method of claim 1, wherein, The process of evaluating the fluency of the audio data includes: The word duration and interval features in the alignment features are obtained, the fluency score of the speech is calculated according to the word duration and interval features, the pause analysis result is obtained according to the interval features, the syntactic and semantic coherence of the target transcription text and the recognized transcription text is analyzed by embedding the target transcription text and the recognized transcription text, the syntactic and semantic coherence of the target transcription text and the recognized transcription text is compared and evaluated to obtain a semantic consistency score, and the fluency score, the pause analysis result, and the semantic consistency score are input into a linear regression layer to obtain the fluency score.

6. The method of claim 1, wherein, The process of evaluating the rhythm score includes: wherein RMSE is the root mean square error of the rhythm, and Max Timing Difference is the maximum time error of the alignment.

7. The method of claim 1, wherein, The process of evaluating the grammatical accuracy of the answer text includes: The answer text is checked for grammatical errors by the constructed prompt large model, and the number and types of grammatical errors are counted to calculate the grammatical accuracy score: accuracy = max(0, (1 - error_count / word_count) * length_penalty) wherein error_count is the number of grammatical errors, word_count is the total number of words, and length_penalty is a length penalty coefficient adjusted according to the sentence length of the answer text.

8. The method of claim 1, wherein, The process of evaluating the on-topic degree of the answer text includes: The answer text is scored according to the designed rules by the constructed prompt large model to obtain a topic cutting degree score.

9. The method of claim 1, wherein, The speech recognition accuracy, the audio score and the answer text score are comprehensively summarized by means of average synthesis or weighted calculation.

Citation Information

Patent Citations

  • Intelligent evaluation method, system and device for spoken English test voice

    CN115497455A

  • Method and device for automatically evaluating spoken language fluency

    CN118335086A