A speech expression evaluation method and device for video and electronic equipment

CN122511292APending Publication Date: 2026-08-04HUAZHONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG NORMAL UNIV
Filing Date
2026-04-23
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0005]针对现有技术的缺陷,本申请的目的在于提供一种面向视频的话语表达评估方法、装置及电子设备,旨在解决现有研究缺乏一套科学普适的针对课堂实录视频的教师话语表达评估方法,导致教师话语表达的评估准确性较低的问题

Benefits of technology

本申请构建一套完整的、从视频输入到报告输出的端到端数智化评估系统,通过将课堂实录视频预处理、多维度特征自动化提取(语速值、语调变化、话语清晰度)与大语言模型驱动的智能报告生成有机串联,实现了从原始教学视频到结构化评估报告的全自动化流程,通过选定对话语表达评估影响较大的语速、语调和话语清晰度作为评估教师话语表达的言语指标,提升教师话语表达的评估准确性,且大语言模型还基于具体教学背景参数,包括学生年龄段和学科类型,从而使大语言模型动态调整评估基准,提升教师话语表达的评估准确性和普适性,突破了传统研究只能局限于特定场景的制约。该方案为一线教师提供了直观、可操作的话语表达特征反馈,有效解决了现有研究停留在学术层面、缺乏面向实践应用的系统性评估方法与工具的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122511292A_ABST
    Figure CN122511292A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of education evaluation, and specifically discloses a speech expression evaluation method and device for a video and electronic equipment, and the method comprises the following steps: extracting audio from a classroom recording video, and segmenting and screening valid audio segments containing only teacher speech from the audio; obtaining the intelligibility level, tone variation and speech speed value of the valid audio segments; and generating an evaluation report by using a large language model based on a teaching background parameter, the speech speed value, the tone variation, the intelligibility level and a preset evaluation index weight. The method can improve the evaluation accuracy and universality of teacher speech expression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of educational assessment technology, and more specifically, relates to a method, device and electronic device for assessing discourse expression in video. Background Technology

[0002] Teachers' verbal expression is one of the most crucial teaching behaviors in the classroom, and its quality directly impacts teaching effectiveness. Many scholars have dedicated themselves to studying the characteristics of teaching language under different teaching models and content themes, thereby decoding the linguistic code of effective classrooms.

[0003] However, current research on the feature extraction of teachers' discourse in classroom recordings remains primarily focused on academic exploration. This research typically involves measuring and analyzing specific linguistic features to explore differences in these features across different subjects and teaching formats. At the practical level, a systematic assessment method and tool that is easily understood and applied by frontline teachers has yet to be found. Furthermore, the diverse classroom settings across different subjects and student age groups make establishing a universal, objective assessment standard extremely difficult, and existing research is often limited to specific scenarios.

[0004] In summary, existing research lacks a scientifically universally applicable method for assessing teachers' verbal expression in recorded classroom videos, resulting in low accuracy in assessing teachers' verbal expression. Summary of the Invention

[0005] In view of the shortcomings of the existing technology, the purpose of this application is to provide a method, device and electronic device for evaluating discourse expression in video, which aims to solve the problem that the existing research lacks a scientific and universally applicable method for evaluating teachers' discourse expression in classroom recording videos, resulting in low accuracy of evaluation of teachers' discourse expression.

[0006] To achieve the above objectives, firstly, this application provides a method for evaluating discourse expression in video, comprising: Audio is extracted from classroom recording videos, and valid audio segments containing only the teacher's voice are segmented and filtered from the audio. Obtain the clarity level, intonation variation, and speech rate of the effective audio segment; An evaluation report is generated using a large language model based on teaching background parameters, speech rate value, intonation variation, clarity level, and preset evaluation index weights.

[0007] This application constructs a complete end-to-end digital intelligent assessment system, from video input to report output. By organically linking classroom video preprocessing, automated extraction of multi-dimensional features (speech rate, intonation variation, and speech clarity), and intelligent report generation driven by a large language model, it achieves a fully automated process from raw teaching videos to structured assessment reports. By selecting speech rate, intonation, and speech clarity—which have a significant impact on the assessment of teachers' speech expression—as speech indicators, the accuracy of teacher speech expression assessment is improved. Furthermore, the large language model is based on specific teaching background parameters, including student age group and subject type, enabling it to dynamically adjust the assessment benchmark, thereby improving the accuracy and universality of teacher speech expression assessment and breaking through the limitations of traditional research that is confined to specific scenarios. This solution provides frontline teachers with intuitive and operable feedback on speech expression characteristics, effectively addressing the problem that existing research remains at the academic level and lacks systematic assessment methods and tools for practical application.

[0008] According to the video-based discourse expression assessment method provided in this application, the step of segmenting and filtering out valid audio segments containing only the teacher's speech from the audio includes: Detecting valid speech segments based on speech activity detection technology; Speaker recognition was performed using the Gaussian Mixture Model-Universal Background Model (GMM-UBM) algorithm. By setting a likelihood ratio score threshold, the teacher's speech was distinguished and extracted to obtain effective audio segments containing the teacher's speech.

[0009] This application employs the GMM-UBM algorithm for speaker recognition and sets a likelihood ratio score threshold, enabling precise differentiation of teacher speech, student speech, and silent segments from complex classroom recordings. It effectively eliminates interference from student responses, classroom discussions, and other non-teacher speech. This ensures that subsequent feature extraction and evaluation analysis are strictly focused on the teacher's teaching performance, significantly improving the accuracy and relevance of the evaluation results and avoiding evaluation bias caused by the inclusion of irrelevant speech.

[0010] According to the video-oriented discourse expression evaluation method provided in this application, the process of obtaining the clarity level of the effective audio segment includes: Train an Emphasized Channel Attention, Propagation and Aggregation in Time DelayNeural Network (ECAPA-TDNN) model on a speaker recognition dataset; Fine-tuning the ECAPA-TDNN model to adapt it to the speech clarity level classification task; The effective audio segments are classified using a finely tuned ECAPA-TDNN model to obtain a clarity level.

[0011] This application uses the ECAPA-TDNN model pre-trained on a large speaker recognition dataset as its foundation and fine-tunes it on a specially constructed teacher speech dataset. It fully leverages the advantages of transfer learning, successfully transferring general speaker recognition capabilities to the specific task of speech clarity assessment. It can achieve high-precision and automated clarity level classification on a smaller scale of domain data, which not only reduces the dependence on large-scale labeled data, but also ensures the professionalism and robustness of the model in the specific acoustic scenario of teacher speech.

[0012] According to the video-oriented discourse expression evaluation method provided in this application, the process of obtaining the intonation changes of the effective audio segment includes: Analyze the fundamental frequency and energy changes of the effective audio segment, and calculate the first-order difference between the smoothed fundamental frequency sequence and the energy sequence; The dynamic threshold is determined based on the first-order difference between the smoothed fundamental frequency sequence and the energy sequence; Based on the relationship between the fundamental frequency change and the energy change relative to the dynamic threshold, the intonation is determined to be rising, falling, or neutral.

[0013] This application achieves objective and adaptive determination of intonation type by comprehensively analyzing the first-order difference between fundamental frequency and energy, and adopting a dynamic threshold mechanism based on standard deviation for intonation classification. The dynamic threshold can be automatically adjusted according to the intonation variation range of different speakers (such as teachers of different genders and ages), effectively overcoming the subjectivity and scenario limitations brought about by using fixed thresholds, making the intonation evaluation results more objective and reliable, and providing a highly versatile technical means for quantitatively analyzing intonation characteristics in different teaching situations.

[0014] According to the video-oriented discourse expression evaluation method provided in this application, the process of obtaining the speech rate value of the effective audio segment includes: The effective audio segments are converted into text using a speech recognition model; The language type is detected based on the speech recognition results. For languages ​​that use characters as the basic unit, the number of valid characters in the transcribed text is counted. For languages ​​that use spaces to separate words, the number of valid words in the transcribed text is counted. Based on the number of valid characters or valid words and the duration of valid speech, the speech rate value per unit time is calculated.

[0015] This application automatically detects the language type based on speech recognition results and adaptively selects the number of characters or words as the unit of measurement, making the speech rate assessment module compatible with classroom scenarios in multiple languages ​​such as Chinese and English. This significantly improves the system's versatility and applicability, eliminating the need to develop separate assessment modules for different languages, reducing the system's complexity, and enabling it to be flexibly applied to teaching assessment needs in different language environments.

[0016] Secondly, this application provides a device for evaluating discourse expression in video, comprising: The preprocessing module is used to extract audio from classroom recording videos and segment and filter out valid audio segments containing only the teacher's voice from the audio. The acquisition module is used to acquire the clarity level, intonation variation, and speech rate value of the effective audio segment; The assessment module is used to generate an assessment report based on teaching background parameters, the speech rate value, the intonation variation, the clarity level, and the preset assessment index weights, using a large language model.

[0017] Thirdly, this application provides an electronic device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the video-oriented discourse expression evaluation method described in the first aspect or any possible implementation thereof.

[0018] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the video-oriented discourse expression evaluation method described in the first aspect or any possible implementation thereof.

[0019] Fifthly, this application provides a computer program product that, when run on a processor, causes the processor to execute the video-oriented discourse expression evaluation method described in the first aspect or any possible implementation of the first aspect.

[0020] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0021] Overall, the technical solutions conceived in this application have the following beneficial effects compared with the prior art: This application constructs a complete end-to-end digital intelligent assessment system, from video input to report output. By organically linking classroom video preprocessing, automated extraction of multi-dimensional features (speech rate, intonation variation, and speech clarity), and intelligent report generation driven by a large language model, it achieves a fully automated process from raw teaching videos to structured assessment reports. By selecting speech rate, intonation, and speech clarity—which have a significant impact on the assessment of teachers' speech expression—as speech indicators, the accuracy of teacher speech expression assessment is improved. Furthermore, the large language model is based on specific teaching background parameters, including student age group and subject type, enabling it to dynamically adjust the assessment benchmark, thereby improving the accuracy and universality of teacher speech expression assessment and breaking through the limitations of traditional research that is confined to specific scenarios. This solution provides frontline teachers with intuitive and operable feedback on speech expression characteristics, effectively addressing the problem that existing research remains at the academic level and lacks systematic assessment methods and tools for practical application. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart illustrating the video-oriented discourse expression evaluation method provided in the embodiments of this application; Figure 2 This is a schematic diagram of the framework of the video-oriented discourse expression evaluation method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the video-oriented discourse expression evaluation device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0025] In this article, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The symbol " / " in this article indicates that the related objects are in an "or" relationship; for example, A / B means A or B.

[0026] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0027] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, for example, multiple processing units means two or more processing units, multiple elements means two or more elements, etc.

[0028] Next, combined Figures 1-2 The video-oriented discourse expression evaluation method provided in the embodiments of this application is introduced.

[0029] Figure 1 This is a flowchart illustrating the video-oriented discourse expression evaluation method provided in the embodiments of this application, as shown below. Figure 1 As shown, the method includes the following steps: Step S1: Extract audio from the classroom recording video, and segment and filter out valid audio segments that contain only the teacher's voice. To achieve digital assessment of teachers' speech in classroom recordings, the first step is to extract effective audio from the recordings.

[0030] Step S2: Obtain the clarity level, intonation variation, and speech rate of the valid audio segments; This application selects speech rate, tone, and speech clarity as speech indicators for evaluating teachers' speech expression.

[0031] Speech rate refers to the speed at which a teacher speaks during a lecture. In Chinese, the statistical unit is usually "words per minute," while in other languages ​​such as English, the statistical unit is "words per minute."

[0032] In classroom teaching, intonation refers to the speaking tone created by teachers through configuring the pitch, volume, and speed of their voices. It can be mainly divided into three categories: rising intonation, falling intonation, and level intonation.

[0033] Speech clarity refers to the accuracy and intelligibility of a teacher's pronunciation when speaking.

[0034] Figure 2 This is a schematic diagram of the framework of the video-oriented discourse expression evaluation method provided in the embodiments of this application, such as... Figure 2As shown, in one embodiment of this application, the method includes three stages: classroom recording video preprocessing, teacher speech expression evaluation, and generation of teacher speech expression evaluation report. In the teacher speech expression evaluation stage, after obtaining an effective audio segment containing the teacher's voice, the clarity level features, tone change features, and speech rate value features are extracted through the intonation evaluation module, speech rate evaluation module, and speech clarity evaluation module.

[0035] Step S3: Based on teaching background parameters, speech rate, intonation variation, clarity level, and preset evaluation index weights, an evaluation report is generated using a large language model.

[0036] Optionally, the large language model used can be qwen3-max, deepseek, or other large language models, and this application does not limit this.

[0037] like Figure 2 As shown in one embodiment of this application, in the stage of generating a teacher's discourse expression assessment report, in order to achieve the relevance and contextualization of the assessment report, the teaching information teacher_info input by the user is first received and processed. This information contains key teaching background parameters, including the student's age group (such as primary school students, middle school students, university students, etc.) and subject type (such as humanities, science, etc.). These contextual parameters are input into a large language model, such as qwen3-max, as important contextual basis, to ensure that the subsequent generated report analysis and assessment criteria can be closely matched with the specific teaching scenario, thereby improving the application value and guiding significance of the assessment results.

[0038] To construct a scientific and reasonable multi-dimensional evaluation system, the Analytic Hierarchy Process (AHP) was used to determine the weights of each evaluation indicator. The Delphi expert questionnaire was used to collect the relative importance judgments of experts in the field on each indicator. The judgment matrix was constructed based on the AHP 1-9 scaling method. The expert team was composed of experts in the education field and experts in the technology field in a balanced ratio. After multiple rounds of iterative consultation, the consensus of the experts' opinions met the statistical test standard of CR < 0.1.

[0039] The evaluation dimensions for teachers' verbal expression comprehensively consider core teaching expression elements such as appropriate speaking speed, intonation variation characteristics, and speech clarity. This evaluation index system, as shown in Table 1, provides a scientific and reasonable quantitative basis for evaluating teaching verbal expression.

[0040] Table 1 Evaluation Index System for Teachers' Verbal Expression

[0041] The aforementioned multi-dimensional evaluation indicators and their weighting system are embedded into the report generation process of the large language model through structured prompt word design.

[0042] Specifically, the system comprehensively calculates a structured assessment data summary by integrating the raw evaluation values ​​of teacher discourse features (such as speech rate, intonation variation rate, and clarity level) obtained from the previous three modules into effective teacher audio, based on the indicator weights. This data summary, along with the teaching background parameter `teacher_info`, constitutes the input for this module. Through specific instruction prompts, the large language model is guided to automatically generate a well-formatted, in-depth, and specific suggestion-based written assessment report of teacher discourse expression based on the input structured data and contextual parameters, thus completing the intelligent conversion from quantitative data to a qualitative analysis report.

[0043] This application presents a video-based discourse expression assessment method that constructs a complete end-to-end digital intelligent assessment system from video input to report output. By organically linking classroom video preprocessing, automated extraction of multi-dimensional features (speech rate, intonation variation, and speech clarity), and intelligent report generation driven by a large language model, it achieves a fully automated process from raw teaching videos to structured assessment reports. By selecting speech rate, intonation, and speech clarity—which have a significant impact on discourse expression assessment—as speech indicators, it improves the accuracy of teacher discourse expression assessment. Furthermore, the large language model dynamically adjusts the assessment benchmark based on specific teaching background parameters, including student age group and subject type, thereby enhancing the accuracy and universality of teacher discourse expression assessment and overcoming the limitations of traditional research confined to specific scenarios. This solution provides frontline teachers with intuitive and operable feedback on discourse expression characteristics, effectively addressing the problem that existing research remains at the academic level and lacks systematic assessment methods and tools for practical application.

[0044] In some embodiments, step S1 specifically includes: Detecting valid speech segments based on speech activity detection technology; The GMM-UBM algorithm is used for speaker recognition. By setting a likelihood ratio score threshold, the teacher's speech is distinguished and extracted to obtain effective audio segments containing the teacher's speech.

[0045] like Figure 2 As shown in one embodiment of this application, the classroom recording video preprocessing stage consists of the following three steps: The first step is audio track extraction: using the moviepy library, the complete audio track is extracted from the classroom recording video file to obtain the original audio data containing the entire classroom teaching process; The second step is teacher speech segmentation: the raw audio data is processed using the ST automated segmentation tool. This tool first detects valid speech segments in the audio based on WebRTC VAD technology, and then uses the GMM-UBM algorithm for speaker recognition. By setting a likelihood ratio score threshold of 0.2, it accurately distinguishes between teacher speech, student speech, and silent segments, thereby extracting segments containing only teacher speech from the complete classroom recording. This method effectively eliminates interference from student answers, classroom discussions, and other non-teacher speech, ensuring that subsequent analysis focuses on the teacher's teaching performance. The third step is to filter effective segments: the segmented teacher audio segments are filtered by duration, and only audio segments with a duration of more than one second are retained as effective teacher audio data for the final digital assessment.

[0046] The above steps together constitute a complete preprocessing workflow for classroom recording videos, providing a guarantee for accurate evaluation of teachers' speech and expression in the subsequent process.

[0047] In some embodiments, the process of obtaining the clarity level of the valid audio segment in step S2 includes: Step S21: Train the ECAPA-TDNN model on the speaker recognition dataset; Step S22: Fine-tune the ECAPA-TDNN model to adapt it to the speech clarity level classification task. Step S23: Classify the effective audio segments using the finely tuned ECAPA-TDNN model to obtain the clarity level.

[0048] Specifically, we first constructed a Teacher Audio dataset, which includes approximately 57 lessons covering different grade levels, regions, and subject types, ensuring data diversity and representativeness. The specific construction process of the dataset is as follows: First, the audio track was separated from each classroom recording video. The complete course audio was evenly divided into segments of 8 seconds each, and all audio segments were converted to WAV format to ensure data format consistency. Then, valid segments containing only the teacher's voice were selected from each 8-second audio segment, excluding segments containing student speeches, group conversations, or background noise, to ensure that subsequent analysis focused on the teacher's independent voice. Finally, the selected segments were further divided into three clarity levels: "high," "medium," and "low," based on the clarity of the speech.

[0049] Optionally, after annotation is completed, a consistency check algorithm (such as the Kappa statistic) is used to evaluate the consistency of the results from multiple annotators to ensure the reliability of annotation quality.

[0050] Then, a speech clarity assessment model is trained. To achieve efficient and accurate assessment of teachers' speech clarity, this application adopts the core idea of ​​transfer learning. Specifically, the ECAPA-TDNN model, pre-trained on the large speaker recognition dataset VoxCeleb, is selected as the basic feature extractor. Through systematic data preprocessing and targeted fine-tuning training, its general speaker recognition capabilities are transferred to the specific task of speech clarity assessment. The specific steps include: 1a. Data preprocessing: Standardize, clean, and enhance the audio samples in the Teacher Audio dataset to ensure that the speech data input to the model is of high quality and consistent.

[0051] First, data loading and cleaning were performed: the sampling rate of all audio samples was standardized to 44.1kHz, and multi-channel audio was mixed into a single channel to eliminate noise introduced by differences in vocal channels. Second, length normalization was performed: all audio samples were standardized to a fixed duration (approximately 8 seconds) through truncation or end padding to ensure compatibility of input dimensions during batch training. Finally, data augmentation was applied to the training set during the training phase: various acoustic transformations were applied with random probabilities, including adding Gaussian background noise of random intensity to simulate real acoustic environments, applying random gain changes to enhance the model's robustness to volume fluctuations, applying small random time shifts to weaken the influence of absolute time position, and applying slight time stretching with low probability to simulate speech rate changes. This series of augmentation strategies effectively improved the model's generalization ability to different acoustic conditions and pronunciation habits.

[0052] 2a. Fine-tuning Training: Fine-tuning training aims to adapt the pre-trained model to the speech clarity level classification task. First, the model structure is adaptively modified: the ECAPA-TDNN backbone network is retained as the feature extractor, and its original classification head is replaced with a task-specific classification module. Specifically, a normalized cosine classifier is introduced, the calculation process of which can be formalized as follows:

[0053] in, The 192-dimensional acoustic feature vector output by the model. The classifier weight matrix is... Represents the L2 norm. Through the analysis of... and Perform L2 normalization on each object, map it to a unit spherical space, and apply a scaling factor. (In this embodiment, it is set to 30) to enhance the discriminativeness of the category decision boundary, thereby achieving accurate mapping and differentiation of the three clarity levels of "high", "medium" and "low".

[0054] Optionally, a hierarchical parameter update strategy is adopted during training: the general feature extraction layer at the bottom of the model is frozen to maintain its pre-trained knowledge; the mid-to-high-level network components (including Bottle2neck module, channel attention mechanism SEBlocks, feature aggregation layer, statistical pooling layer, fully connected layer and cosine classifier weights) are unfrozen and targeted gradient updates are performed using Teacher Audio data, so that the model learns high-level semantic features related to clarity.

[0055] Optionally, in terms of optimization strategy, a group learning rate configuration is adopted: differentiated learning rates (1e-5 and 5e-5) are set for the unfrozen backbone network and the classification head respectively, and the learning rate is dynamically adjusted in combination with Warmup and cosine annealing scheduler, while gradient pruning (threshold) is implemented. Setting it to 1.0 ensures the stability of the training process.

[0056] Optionally, for the loss function, a labeled smoothing and class-weighted cross-entropy loss is used, the expression of which is:

[0057] Where C is the total number of categories (C=3 here). For category c, smooth label and + ( 0.05 is the label smoothing factor. (This is the original tag) The model predicts the probability of class c. The weights are for class c. The class weights are calculated based on the training set distribution, and additional weights are applied to the "medium" and "low" levels that the model is prone to confusion, in order to effectively alleviate the class imbalance problem and improve the model's generalization performance and classification robustness in complex acoustic scenarios.

[0058] In some embodiments, the process of obtaining the intonation changes of the valid audio segments in step S2 includes: Step S24: Analyze the fundamental frequency and energy changes of the effective audio segments, and calculate the first-order difference between the smoothed fundamental frequency sequence and the energy sequence; Step S25: Determine the dynamic threshold based on the first-order difference between the smoothed fundamental frequency sequence and the energy sequence; Step S26: Based on the relationship between fundamental frequency change and energy change relative to the dynamic threshold, the intonation is determined to be rising, falling, or neutral.

[0059] By analyzing the fundamental frequency and energy variations of audio signals, multidimensional prosodic features are extracted, thereby enabling automatic classification of intonation types. Specifically, the implementation process includes the following three stages: 1b. Preprocessing Stage: In the preprocessing stage, a series of normalization operations are performed on the input teacher's audio to improve the robustness and accuracy of subsequent feature extraction. The specific steps are as follows: First, robust noise suppression is performed on the audio using a spectral subtraction-based noise reduction technique. A 200-millisecond noise sample from the beginning of the audio is extracted for estimation, and an 80% noise reduction ratio is set to effectively suppress background noise interference. Subsequently, the denoised signal undergoes pre-emphasis filtering. A transfer function of H(z) = 1- z - ¹( A digital filter is used to enhance the high-frequency components of the speech signal, thereby improving the accuracy of subsequent fundamental frequency detection. Finally, the filtered continuous speech signal is analyzed and windowed to convert it into a short-time stationary signal suitable for time-frequency analysis. Specifically, a Hanning window with a frame length of 25 milliseconds and a frame shift of 10 milliseconds is used for framing.

[0060] 2b. Feature Extraction Stage: This stage employs a dual feature extraction strategy combining fundamental frequency and energy to comprehensively capture the prosodic characteristics of intonation. For fundamental frequency features, an improved pYIN algorithm is used for estimation within the 80-400Hz fundamental frequency range. This algorithm integrates the robustness of the YIN method with probabilistic interpolation techniques, effectively addressing fundamental frequency breaks and noise interference to obtain a continuous fundamental frequency trajectory. For energy features, the root mean square value of each frame of speech signal is calculated to obtain the energy trajectory reflecting changes in speech intensity. To further improve the reliability of the features, the extracted fundamental frequency sequence... Perform smoothing processing. For Each point in If the point is valid (i.e. Then, for that fundamental frequency point Perform a smoothing operation. Use a window size of [size missing]. A sliding window of 5 frames, for the fundamental frequency point The smoothing calculation formula is:

[0061] This allows for the suppression of abnormal fluctuations while maintaining the natural continuity of the intonation trajectory.

[0062] 3b. Intonation Classification Stage: This stage, based on a dynamic threshold mechanism, comprehensively analyzes the extracted fundamental frequency and energy features to automatically classify rising, falling, and neutral intonations. The specific implementation process is as follows: First, the first-order difference of the smoothed fundamental frequency trajectory is calculated to obtain the fundamental frequency variation sequence. The energy change sequence is obtained by the first-order difference of the initial energy trajectory. Then, a dynamic classification threshold is established, primarily excluding fundamental frequency variation sequences. Extreme outliers exceeding 1.5 times the interquartile range (IQR) were identified. The fundamental frequency classification threshold was calculated based on the standard deviation of the remaining data. Similarly, the energy classification threshold is obtained. This mechanism can adapt to the range of intonation variations of different speakers; finally, a comprehensive intonation determination is performed. The determination logic comprehensively considers the fundamental frequency change trend and the energy change trend. and When, mark the audio frame as "rising tone"; when and When the intonation is falling, it is marked as "falling intonation"; otherwise, it is marked as "neutral intonation".

[0063] Through the above processing steps, key prosodic features can be automatically and objectively extracted from the original teacher's speech, and intonation classification can be completed. This allows for the calculation of the intonation variation rate, providing a quantitative indicator for subsequent comprehensive evaluation of discourse performance. Intonation Variation Rate ,in The frame number marked as pitch up. The number of frames marked as down-tone is N, where N is the total number of frames.

[0064] In some embodiments, the process of obtaining the speech rate value of the valid audio segment in step S2 includes: Step S27: Use a speech recognition model to convert valid audio segments into text; Step S28: Detect the language type based on the speech recognition results. For languages ​​that use characters as the basic unit, count the number of valid characters in the transcribed text. For languages ​​that use spaces to separate words, count the number of valid words in the transcribed text. Step S29: Calculate the speech rate value per unit time based on the number of valid characters or valid words and the effective speech duration.

[0065] The speech rate assessment module aims to extract objective and accurate speech rate metrics from input audio through an automated process. Its core lies in using an advanced speech recognition model to convert audio into text and combining it with a language-adaptive statistical strategy to calculate the number of effective language units per unit time. The implementation of this module mainly includes the following steps: First, the input teacher's audio is preprocessed and subjected to high-precision text conversion. A framework based on the Whisper large-scale speech recognition model is used to convert the audio signal into corresponding text transcription. During this process, speech activity detection technology is simultaneously applied to automatically identify and exclude silent or non-speech segments in the audio, ensuring that subsequent analysis is based only on valid speech intervals. Second, language-adaptive text unit statistics are performed. The system automatically detects the language type of the original audio based on the speech recognition results and adaptively selects the appropriate measurement strategy accordingly: for languages ​​such as Chinese, Japanese, and Korean, which use characters as the basic unit, the number of valid characters in the transcribed text is counted; for languages ​​such as English and French, which use spaces to separate words, the number of valid words is counted. Finally, the speech rate index is calculated based on the above statistical results. Where N is the number of valid characters or words counted during the effective speech time period, and T is the effective speech duration (in minutes) after being filtered by speech activity detection, thus obtaining a speech rate quantification value in "characters per minute" or "words per minute".

[0066] The speech rate characteristic index output by this module provides key data for assessing the fluency and rhythm of teachers' speech.

[0067] The teacher discourse expression assessment system constructed in this application adheres to the principles of scientific rigor, systematic approach, and applicability. Based on classroom recordings, it effectively extracts and quantifies the teaching language features contained in teachers' original, valid speech signals. This system realizes the transformation from theoretical discussion to practical application, providing frontline teachers with intuitive and operable feature feedback.

[0068] This application employs targeted feature extraction and analysis strategies to address the context-dependent and subjective challenges of assessment criteria. For speech clarity feature extraction, a pre-trained ECAPA-TDNN model is fine-tuned and trained on a specific dataset (Teacher Audio) to achieve high-precision, automated speech clarity classification. For intonation feature extraction, rising, falling, and level tones are labeled by comprehensively analyzing the synergistic trends of fundamental frequency and energy. In terms of comprehensive analysis, for features highly dependent on the teaching context, such as speech rate and intonation, teaching background parameters are innovatively introduced and combined with a large-scale model. By identifying key variables such as subject matter and student age group, the assessment benchmark is dynamically adjusted, achieving a balance between universality and accuracy, thus overcoming the limitations of traditional research that is confined to specific scenarios.

[0069] The video-oriented discourse evaluation device provided in this application is described below. The video-oriented discourse evaluation device described below can be referred to in correspondence with the video-oriented discourse evaluation method described above.

[0070] Figure 3 This is a schematic diagram of the structure of a video-oriented discourse expression evaluation device provided in an embodiment of this application, as shown below. Figure 3 As shown, the device 300 includes: The preprocessing module 310 is used to extract audio from classroom recording videos and segment and filter out valid audio segments that contain only the teacher's voice. The acquisition module 320 is used to acquire the clarity level, intonation variation, and speech rate value of valid audio segments; The assessment module 330 is used to generate an assessment report based on teaching background parameters, speech rate, intonation variation, clarity level, and preset assessment indicator weights using a large language model.

[0071] It should be understood that the above-described device is used to execute the methods in the above embodiments. The implementation principle and technical effect of the corresponding program modules in the device are similar to those described in the above methods. The working process of the device can be referred to the corresponding process in the above methods, and will not be repeated here.

[0072] Based on the methods in the above embodiments, Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown in the illustration, this application provides an electronic device that may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440. The processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions stored in the memory 430 to execute the video-oriented speech expression evaluation method described in the above embodiment.

[0073] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the video-oriented speech expression evaluation method described in the various embodiments of this application.

[0074] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the video-oriented speech expression evaluation method in the above embodiments.

[0075] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the video-oriented speech expression evaluation method in the above embodiments.

[0076] It is understood that the processor in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.

[0077] The method steps in this application embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0078] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0079] It is understood that the various numerical designations used in the embodiments of this application are merely for the convenience of description and are not intended to limit the scope of the embodiments of this application.

[0080] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for evaluating discourse expression in video, characterized in that, include: Audio is extracted from classroom recording videos, and valid audio segments containing only the teacher's voice are segmented and filtered from the audio. Obtain the clarity level, intonation variation, and speech rate of the effective audio segment; An evaluation report is generated using a large language model based on teaching background parameters, speech rate value, intonation variation, clarity level, and preset evaluation index weights.

2. The video-oriented discourse expression evaluation method according to claim 1, characterized in that, The step of segmenting and filtering out valid audio segments containing only the teacher's voice from the audio includes: Detecting valid speech segments based on speech activity detection technology; The GMM-UBM algorithm is used for speaker recognition. By setting a likelihood ratio score threshold, the teacher's speech is distinguished and extracted to obtain effective audio segments containing the teacher's speech.

3. The video-oriented discourse expression evaluation method according to claim 1, characterized in that, The process of obtaining the clarity level of the valid audio segment includes: Train the ECAPA-TDNN model on a speaker recognition dataset; Fine-tuning the ECAPA-TDNN model to adapt it to the speech clarity level classification task; The effective audio segments are classified using a finely tuned ECAPA-TDNN model to obtain a clarity level.

4. The video-oriented discourse expression evaluation method according to claim 1, characterized in that, The process of obtaining the intonation changes of the effective audio segments includes: Analyze the fundamental frequency and energy changes of the effective audio segment, and calculate the first-order difference between the smoothed fundamental frequency sequence and the energy sequence; The dynamic threshold is determined based on the first-order difference between the smoothed fundamental frequency sequence and the energy sequence; Based on the relationship between the fundamental frequency change and the energy change relative to the dynamic threshold, the intonation is determined to be rising, falling, or neutral.

5. The video-oriented discourse expression evaluation method according to claim 1, characterized in that, The process of obtaining the speech rate value of the effective audio segment includes: The effective audio segments are converted into text using a speech recognition model; The language type is detected based on the speech recognition results. For languages ​​that use characters as the basic unit, the number of valid characters in the transcribed text is counted. For languages ​​that use spaces to separate words, the number of valid words in the transcribed text is counted. Based on the number of valid characters or valid words and the duration of valid speech, the speech rate value per unit time is calculated.

6. A device for evaluating discourse expression in video, characterized in that, include: The preprocessing module is used to extract audio from classroom recording videos and segment and filter out valid audio segments containing only the teacher's voice from the audio. The acquisition module is used to acquire the clarity level, intonation variation, and speech rate value of the effective audio segment; The assessment module is used to generate an assessment report based on teaching background parameters, the speech rate value, the intonation variation, the clarity level, and the preset assessment index weights, using a large language model.

7. An electronic device, characterized in that, include: At least one memory for storing computer programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform a video-oriented speech expression evaluation method as described in any one of claims 1-5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is run on a processor, the processor performs the video-oriented speech expression evaluation method as described in any one of claims 1-5.

9. A computer program product, characterized in that, When the computer program product is run on a processor, the processor performs the video-oriented speech expression evaluation method as described in any one of claims 1-5.