A deep learning-based teaching quality evaluation method and system

By using a deep learning-based teaching quality assessment method, screening and comparing teaching speech with voiceprints, constructing a training dataset, and distinguishing between background noise and background noise, the problem of insufficient assessment accuracy in existing technologies is solved, and efficient and accurate speech quality assessment and teaching quality feedback are achieved.

CN122392572APending Publication Date: 2026-07-14CNSCI SOFT EDUCATIONAL TECH (BEIJING) CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CNSCI SOFT EDUCATIONAL TECH (BEIJING) CORP
Filing Date
2026-06-15
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing speech quality assessment methods rely on manual listening or fixed thresholds, which are subjective and inefficient, making them difficult to adapt to large-scale online teaching. Furthermore, the single dimension of feature extraction leads to insufficient assessment accuracy and a tendency to make misjudgments or omissions.

Method used

A deep learning-based teaching quality assessment method is adopted. By screening abnormal speech segments in teaching speech recognition, constructing a training dataset using voiceprint comparison and feature extraction, training a classification model to distinguish between noise and background noise, calculating scores using differentiated evaluation weights, adaptively determining the unit time, and generating quality levels.

Benefits of technology

It enables precise quality assessment of teaching audio, distinguishes between background noise and background noise, improves the accuracy and adaptability of the assessment, reflects the actual impact of teaching content, and provides specific guidance to improve teaching quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122392572A_ABST
    Figure CN122392572A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of intelligent teaching, and particularly relates to a teaching quality evaluation method and system based on deep learning. The method comprises extracting feature data to be evaluated from normal speech, evaluating the feature data to be evaluated by using a speech evaluation model to generate an evaluation result, and determining the quality grade of the normal speech according to the evaluation result, wherein the feature extraction from noise speech and noise speech comprises extracting amplitude information and frequency information from a sound production section. The present application performs screening on teaching speech to identify abnormal sound sections; through voiceprint comparison, the teaching speech in which the abnormal sound sections that can match the pre-stored voiceprint are classified as noise speech, and the teaching speech that cannot be matched is classified as noise speech, so as to distinguish the noise speech originating from the background environment from the noise speech originating from the teaching subject, overcome the evaluation error problem caused by regarding the two as noise without distinction, and lay a data foundation for subsequent evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent teaching technology, specifically relating to a teaching quality assessment method and system based on deep learning. Background Technology

[0002] Distance learning and smart education have become integral parts of modern education. The clarity and fluency of the audio used in teaching, as the main medium for knowledge transfer and interaction between teachers and students, are crucial to the quality of teaching activities. Whether learners can hear clearly is one of the criteria for judging the effectiveness of teaching.

[0003] Currently, speech quality assessment relies heavily on manual listening or rule-based systems with fixed thresholds. This approach is inherently subjective and inefficient, making it unsuitable for large-scale online teaching scenarios. While some methods now incorporate machine learning or signal processing techniques to enhance the intelligence of the assessment, their feature extraction dimensions are relatively limited. They typically focus only on macroscopic acoustic parameters such as frequency and amplitude, failing to effectively distinguish between the teacher's effective vocalization and complex acoustic interferences like background noise and electrical current. This results in inaccurate assessment results, prone to misjudgments or omissions.

[0004] To address the aforementioned problems, this invention provides a teaching quality assessment method and system based on deep learning. Summary of the Invention

[0005] The purpose of this invention is to provide a teaching quality assessment method and system based on deep learning, which can more accurately assess the quality of teaching audio, distinguish between background noise and noise in teaching audio, avoid treating effective background noise as noise, and ensure the accuracy of the assessment.

[0006] This invention discloses a teaching quality assessment method based on deep learning, comprising: When the speech evaluation model and the feature data to be evaluated corresponding to normal speech are obtained, the following steps are performed: evaluate the feature data to be evaluated using the speech evaluation model to generate evaluation results; and determine the quality level of normal speech based on the evaluation results. The process of obtaining the speech evaluation model includes: extracting features from noisy speech and slurred speech to construct a training dataset; and training a classification model using the training dataset to obtain the speech evaluation model; wherein, when constructing the training dataset for noisy speech, amplitude information is specified as the input feature and frequency information is specified as the output feature; and when constructing the training dataset for slurred speech, frequency information is specified as the input feature and amplitude information is specified as the output feature. Among them, feature extraction from noisy and slurred speech includes: extracting amplitude and frequency information from the vocal segment.

[0007] Preferably, determining the difference between noise and background noise includes: Screening is performed on the teaching speech to identify abnormal speech segments; voiceprint comparison is performed on the abnormal speech segments to match the abnormal speech segments with pre-stored voiceprints in the voiceprint database; If an abnormal sound segment matches a pre-stored voiceprint, the teaching speech containing the abnormal sound segment will be marked as noisy speech; if an abnormal sound segment does not match a pre-stored voiceprint, the teaching speech containing the abnormal sound segment will be marked as noisy speech. Among them, teaching voices that were not marked as noise or slurred speech were identified as normal speech.

[0008] Preferably, screening the teaching voice to identify abnormal audio segments includes: inputting the teaching voice into a discrimination model to obtain an amplitude curve; and in the amplitude curve, if a segment is identified whose maximum amplitude is greater than or equal to a preset reference amplitude upper limit, then the audio segment corresponding to the segment is identified as an abnormal audio segment.

[0009] Preferably, determining the quality level of normal speech based on the evaluation results includes: The evaluation results are analyzed as the number of occurrences of a specific audio event within a unit of time; a score is calculated based on the number of occurrences and the corresponding preset evaluation weights. The system compares the scores with preset standard evaluation values ​​to generate quality judgment information, and determines the quality level based on the quality judgment information.

[0010] Preferably, the preset evaluation weights include: The evaluation weight for noise-related events is calculated based on the cumulative duration of noise-related events within a unit of time; and the evaluation weight for background noise events is a preset fixed value.

[0011] Preferably, determining the unit time includes: Based on the initial sampling period, the sampling period is iteratively adjusted and the average fluctuation value of the preset event count within the period is calculated until the average fluctuation value is greater than or equal to the preset standard fluctuation value. The final sampling period is then determined as a unit time.

[0012] This invention also discloses a teaching quality assessment system based on deep learning, comprising: The speech classification module is used to process teaching speech and classify it into noise speech, noisy speech, or normal speech. The evaluation model generation module is used to respond to the classification results of the speech classification module. It constructs a training dataset using features extracted from noisy and slurred speech, and trains a classification model based on the training dataset to generate a speech evaluation model. And a quality judgment module, which responds to the speech evaluation model generated by the evaluation model generation module, and uses the speech evaluation model to evaluate the feature data to be evaluated extracted from normal speech in order to determine the quality level of normal speech.

[0013] Preferably, the speech classification module is used to: screen the teaching speech to identify abnormal speech segments; and perform voiceprint comparison on the abnormal speech segments to classify the teaching speech into noise speech or noisy speech based on whether the abnormal speech segments match the pre-stored voiceprints in the voiceprint database.

[0014] Preferably, the quality judgment module is used for: The evaluation results are analyzed as the number of times a specific audio event occurs within a unit of time; a score is calculated based on the number of occurrences and the corresponding preset evaluation weights. The score is compared with the preset standard evaluation value to determine the quality level.

[0015] Preferably, the quality determination module is further used for: The evaluation weight of noise events is calculated based on the cumulative duration of noise events within a unit of time. And a preset fixed value is used as the evaluation weight for noise-related events.

[0016] Beneficial effects 1. This invention performs screening on teaching speech to identify abnormal sound segments; through voiceprint comparison, teaching speech containing abnormal sound segments that can match pre-stored voiceprints is classified as noise speech, and teaching speech that cannot be matched is classified as noise speech, thereby distinguishing noise speech originating from the background environment from noise speech originating from the teaching subject, overcoming the problem of inaccurate evaluation caused by indiscriminately treating the two as noise, and laying a data foundation for subsequent evaluation.

[0017] 2. This invention adaptively determines the unit time for counting the number of occurrences of specific audio events based on the fluctuation of those events; and adopts differentiated evaluation weights for different types of events when calculating scores, so that the evaluation criteria can fit the content rhythm and interference type of the teaching audio, and the final score can reflect the actual impact of different specific audio events on teaching quality.

[0018] 3. This invention utilizes classified noise and slurred speech to construct a training dataset and employs a targeted feature strategy. For noise and slurred speech, amplitude information is used as the input feature and frequency information as the output feature; for slurred speech, frequency information is used as the input feature and amplitude information as the output feature. The speech evaluation model trained based on this training dataset is ultimately used to evaluate the quality of normal speech. By enabling the speech evaluation model to learn the acoustic feature mapping relationship specific to different types of interference, it can identify potential quality problems when evaluating normal speech. Attached Figure Description

[0019] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a system module diagram of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are merely for explaining the invention and are not intended to limit the scope of protection of the invention.

[0021] Example 1 Please see Figure 1 As shown, this embodiment provides a teaching quality assessment method based on deep learning, including the following steps: S1. Acquisition and preprocessing of teaching audio; The data acquisition unit acquires the original teaching audio generated in the teaching scenario. To facilitate subsequent processing, the acquired teaching audio is digitized, and preliminary signal processing can be selectively performed, such as noise reduction, silence removal, and signal amplitude normalization. The processed audio data is saved as a standard format audio file for use in subsequent steps.

[0022] Furthermore, the data acquisition unit is preferably a built-in microphone integrated into the terminal device or an external recording device.

[0023] S2. Screening and classification of teaching voice recordings; The audio file is input into a preset abnormal audio segment recognition mechanism for processing. This mechanism specifically includes: Calculate the energy or amplitude of the audio file over time to generate an amplitude curve; in the amplitude curve, identify whether there are segments whose duration exceeds the preset minimum duration, and the maximum amplitude in the segment is greater than or equal to the preset reference amplitude limit. It should be noted that the preset standard amplitude of normal teaching voice is usually around -20dBFS, and the upper limit of the reference amplitude can be set to -10dBFS; If the aforementioned segment exists, the audio segment corresponding to that segment will be identified as an abnormal audio segment, and the start time, duration, and maximum amplitude of the abnormal audio segment in the teaching speech will be recorded as a ratio to the preset standard amplitude of the normal teaching speech. Voiceprint comparison is performed on abnormal sound segments. The voiceprint features extracted from the abnormal sound segments are matched with a voiceprint database containing pre-stored voiceprint features of the instructor and other known human voices. The voiceprint features include the speaker's fundamental frequency, formants and other unique acoustic characteristics. If the voiceprint features of the abnormal sound segment successfully match other pre-stored voiceprints in the database besides the instructor, the teaching speech containing the abnormal sound segment is marked as noisy speech. Such speech usually corresponds to the conversation of other people in the teaching environment. If the voiceprint features of the abnormal sound segment cannot be matched with any human voice in the voiceprint database, the teaching speech containing the abnormal sound segment will be marked as noise speech. Such speech usually corresponds to non-human voice interference in the environment, such as electrical noise, object collision sound or outdoor noise. If no abnormal audio segments are detected, or if the voiceprint of an abnormal audio segment matches the teaching voice of the instructor, it is determined to be normal speech.

[0024] S3. Feature extraction and dataset construction; Deep feature extraction is performed on the obtained noisy speech, slurred speech, and normal speech respectively. The specific steps include: After the digital signal of the teaching voice is processed by framing, the vocal segments containing valid voice signals are identified and extracted from each frame by methods such as voice activity detection. Two key acoustic feature information types, amplitude information and frequency information, are extracted from the sound-producing section. Amplitude information includes short-time energy, short-time average amplitude, etc., which are uniformly quantized and defined as the first region feature value. Frequency information includes fundamental frequency, formant, Mel frequency cepstral coefficient, etc., which are uniformly quantized and defined as the second region feature value.

[0025] When constructing the training dataset for subsequent processing, a specific mapping configuration method is adopted, as follows: For noisy speech, the first region feature value is designated as the input feature, and the second region feature value is designated as the corresponding output feature; for noisy speech, the second region feature value is designated as the input feature, and the first region feature value is designated as the corresponding output feature.

[0026] When constructing each data point, a category label is generated simultaneously, which is a label attached to each data pair to indicate whether the data pair originates from noisy or background noise. This category label, along with the input and output features, constitutes a complete training dataset. The training dataset constructed in this way can contain acoustic feature mapping relationships under different types of interference. Feature data extracted from normal speech is saved separately as feature data to be evaluated.

[0027] S4. Establishment of evaluation rules and speech quality assessment; The training dataset is processed according to a predetermined process to establish evaluation rules and parameter systems. Specifically: Based on the correspondence between a large number of input features and output features in the training dataset, as well as their respective category labels, a computational structure is automatically constructed through iterative calculation and parameter optimization. This structure can deduce the quantitative index of the corresponding audio event based on the input feature data to be evaluated. This computational structure is the speech evaluation model. The feature data to be evaluated is input into the speech evaluation model for processing to generate quantitative evaluation results. The evaluation results are further analyzed into the number of occurrences of specific audio events such as discontinuity, sudden volume changes, and specific frequency noise within a dynamically determined unit of time.

[0028] Furthermore, the method for determining the unit time includes: setting an initial sampling period; using the start time of the teaching voice as a reference, counting a preset specific audio event within each sampling period to obtain a series of count items; The fluctuation value is calculated based on the difference between two adjacent count items. This value is used to measure the stability of the frequency of events. All fluctuation values ​​are recorded in a list. The average fluctuation value is obtained by calculating the average of all fluctuation values ​​in the list. The average fluctuation value is compared with the preset standard fluctuation value. If the average fluctuation value is less than the preset standard fluctuation value, the sampling period is gradually adjusted and the above counting and calculation steps are repeated. If the average fluctuation value is greater than or equal to the preset standard fluctuation value, the current sampling period is determined as the final unit time.

[0029] After obtaining the number of occurrences, a score is calculated based on the number of occurrences and the corresponding preset evaluation weight. The preset evaluation weight is set differently based on the event type. For noise-type events parsed in the evaluation results, the corresponding evaluation weight is a value dynamically calculated based on the cumulative occurrence duration of such events within a unit of time; while for background noise events, the corresponding evaluation weight is a preset fixed value.

[0030] The calculated score is compared with the preset standard evaluation value to generate quality judgment information. If the score is higher than the preset standard evaluation value, the quality judgment information is determined to be normal; if the score is lower than the preset standard evaluation value, it is determined to be a warning.

[0031] Furthermore, this embodiment also includes a verification step to ensure the reliability of the evaluation results, specifically: The current feature data to be evaluated is compared with all the data features in the training dataset to calculate the overall similarity or distance metric. If the metric is lower than the preset confidence threshold, it indicates that the speech features to be evaluated differ too much from the existing training samples. In this case, a prompt message can be generated to indicate that the confidence of the current evaluation result is low, or a backup evaluation scheme based on simplified rules can be used for supplementary evaluation.

[0032] A speech evaluation model is a computational model used to infer quantitative indicators of potential audio events based on acoustic features extracted from normal speech. Its specific definition is as follows:

[0033] In the formula, This represents a quantitative indicator of speech quality, indicating the comprehensive deviation between the acoustic features of the speech to be evaluated and the interference feature patterns learned by the model. This represents a special conversion function for noise-speech, which is obtained through training on noise-speech and is used to predict frequency features from amplitude features. This represents a dedicated transformation function for noisy speech, obtained through training on noisy speech, used to predict amplitude features from frequency features; The amplitude feature vector at time t represents the vector extracted from the speech to be evaluated at time frame t, which is composed of the feature values ​​of the first region. The frequency feature vector at time t represents the vector extracted from the speech to be evaluated at time frame t, which is composed of the feature values ​​of the second region. This indicates a noise speech category identifier, meaning an identifier representing a noise speech category; This indicates a category identifier for noisy speech, meaning an identifier representing a category of noisy speech. This represents the weighting coefficient, which is a preset hyperparameter used to adjust and balance the weight of the two different types (noise and noise) in the final quantification index. This represents the total number of frames, which means the total number of speech frames contained within the time window used to calculate the quantification index. denoted as L2 norm squared, which means the square of the Euclidean distance between two vectors, used to measure the difference between the predicted feature vector and the actual feature vector.

[0034] S5. Output of evaluation results; By combining comprehensive quality assessment information with the calculated scores, the final quality level of normal speech is determined, specifically divided into multiple levels such as excellent, good, average, and poor. Based on this final quality level, corresponding processing opinions are generated and output.

[0035] For voices rated as poor, it is recommended to check the microphone equipment or adjust the recording environment; for voices rated as average, it is recommended to pay attention to controlling the speaking speed and volume, thus providing specific guidance for improving teaching quality.

[0036] Example 2 Please see Figure 2 As shown, this embodiment provides a teaching quality assessment system based on deep learning, including the following modules: The speech classification module processes the raw teaching speech and classifies it into noise speech, noisy speech, or normal speech, providing filtered and labeled data for subsequent model training and quality evaluation.

[0037] In the specific execution process, the received teaching audio is screened to identify abnormal audio segments. Specifically: The teaching voice is input into a preset discrimination model to obtain the amplitude curve of the voice. The discrimination model is preferably an anomaly detection model based on energy or amplitude features. During the scanning of the amplitude curve, if a segment is identified whose maximum amplitude is greater than or equal to the preset upper limit of the reference amplitude, the audio segment corresponding to that segment in time is identified as an abnormal sound segment, thereby quickly locating any sudden high-energy abnormal sounds that may exist in the audio.

[0038] After identifying abnormal audio segments, a voiceprint comparison is performed on the abnormal audio segments. Specifically, the voiceprint features of the abnormal audio segments are matched with pre-stored voiceprints in the voiceprint database, which stores the voiceprint features of the teacher and other known human voices. If an abnormal sound segment can be successfully matched with any pre-stored voiceprint, it means that the sound may come from other people's conversations. Therefore, the teaching speech containing the abnormal sound segment is marked as noisy speech as a whole. If no pre-stored voiceprint is matched for the abnormal sound segment, the sound is considered to be non-human interference, such as electrical noise, object collision sound, or outdoor noise. In this case, the teaching speech containing the abnormal sound segment is marked as noise speech.

[0039] All teaching audio that was not marked as background noise or noisy audio was identified as normal audio and will be used for subsequent quality assessments.

[0040] The evaluation model generation module automatically builds and trains a speech evaluation model for quality assessment in response to the classification results of the speech classification module.

[0041] The training dataset is constructed using the noise and slurred speech output by the speech classification module. Key acoustic features, including amplitude and frequency information, are extracted from the vocal segments of these speech segments. In constructing the training dataset, a specific mapping configuration method is adopted, specifically: When processing noisy speech, the extracted amplitude information is designated as the input feature of the model, and the corresponding frequency information is designated as the output feature. When processing noisy speech, the extracted frequency information is designated as the input feature, and the corresponding amplitude information is designated as the output feature. In this way, the model is trained to learn the specific acoustic feature mapping relationship between amplitude and frequency features for noise and interference.

[0042] After the training dataset is constructed, it is used to train a classification model. The classification model can be a support vector machine, a deep neural network, or any other model suitable for handling such acoustic feature classification tasks. The specific model type here is only a preferred implementation of the present invention and does not constitute a limitation of the present invention.

[0043] After sufficient training, the resulting stable model is the speech evaluation model required by this system and is passed to the quality judgment module for use.

[0044] The quality assessment module is used to evaluate the quality of normal speech using the speech evaluation model generated by the evaluation model generation module, and finally determine its final quality level.

[0045] Once a speech evaluation model and the feature data to be evaluated extracted from normal speech are obtained, the speech evaluation model is used to evaluate the feature data to generate an evaluation result. This evaluation result typically manifests as the identification and labeling of specific audio events present in normal speech. Specific audio events include suspected background human voice segments or non-speech noise segments.

[0046] Based on the evaluation results, the final quality level of normal speech is determined. This process includes the following steps: The evaluation results are analyzed as the number of occurrences of various specific audio events, such as noise events and background noise events, within a specific dynamically determined unit of time. The unit of time is adaptive, specifically: Based on the initial sampling period, this module iteratively adjusts the sampling period and calculates the average fluctuation value of the preset event count within each period. When the average fluctuation value is greater than or equal to the preset standard fluctuation value for the first time, it means that a time scale that can stably reflect the frequency of event occurrence has been found. At this point, the final sampling period is determined as a dynamically determined unit time.

[0047] A comprehensive score is calculated based on the frequency of occurrence of the event and its corresponding preset evaluation weight. The preset evaluation weight is set differently, and the specific setting method is as follows: For noise-related events, the evaluation weight is dynamically calculated, depending on the cumulative duration of the event within a dynamically determined unit of time. The longer the duration, the higher the weight, and the heavier the deduction. For background noise events, the evaluation weight is set to a preset fixed value, reflecting that such events have a fixed negative impact regardless of their duration.

[0048] The calculated score is compared with the preset standard evaluation value to generate quality judgment information. Based on this quality judgment information, the system finally determines the final quality level of the normal speech segment.

Claims

1. A teaching quality assessment method based on deep learning, characterized in that, include: When the speech evaluation model and the feature data to be evaluated corresponding to normal speech are obtained, the following steps are performed: use the speech evaluation model to evaluate the feature data to be evaluated to generate evaluation results; And based on the assessment results, determine the quality level of normal speech; The process of obtaining the speech evaluation model includes: extracting features from noisy and smog speech to construct a training dataset; and using the training dataset to train a classification model to obtain the speech evaluation model. Among them, feature extraction from noisy and slurred speech includes: extracting amplitude and frequency information from the vocal segment; Specifically, when constructing a training dataset for noisy speech, amplitude information is designated as the input feature and frequency information is designated as the output feature; when constructing a training dataset for noisy speech, frequency information is designated as the input feature and amplitude information is designated as the output feature.

2. The teaching quality assessment method based on deep learning according to claim 1, characterized in that, Identifying noise and slurred speech includes: Screening of teaching speech to identify abnormal segments; Perform voiceprint comparison on abnormal sound segments to match the abnormal sound segments with pre-stored voiceprints in the voiceprint database; if the abnormal sound segment matches a pre-stored voiceprint, the teaching speech containing the abnormal sound segment is marked as noisy speech; if the abnormal sound segment does not match a pre-stored voiceprint, the teaching speech containing the abnormal sound segment is marked as noisy speech. Among them, teaching voices that were not marked as noise or slurred speech were identified as normal speech.

3. The teaching quality assessment method based on deep learning according to claim 2, characterized in that, Screening instructional speech to identify abnormal segments includes: The teaching audio is input into the discrimination model to obtain the amplitude curve; Furthermore, if the maximum amplitude of a segment is found to be greater than or equal to the preset upper limit of the reference amplitude in the amplitude curve, the audio segment corresponding to the segment will be identified as an abnormal audio segment.

4. The teaching quality assessment method based on deep learning according to claim 1, characterized in that, Based on the assessment results, the quality level of normal speech is determined as follows: The evaluation results are analyzed as the number of audio events occurring within a unit of time; a score is calculated based on the number of occurrences and the corresponding preset evaluation weights. The system compares the scores with preset standard evaluation values ​​to generate quality judgment information, and determines the quality level based on the quality judgment information.

5. The teaching quality assessment method based on deep learning according to claim 4, characterized in that, The preset evaluation weights include: The evaluation weight for noise-related events is calculated based on the cumulative duration of the noise-related event within a unit of time. And the evaluation weight for noise-related events, which is a preset fixed value.

6. The teaching quality assessment method based on deep learning according to claim 5, characterized in that, Determining the unit of time includes: Based on the initial sampling period, the sampling period is iteratively adjusted and the average fluctuation value of the preset event count within the period is calculated until the average fluctuation value is greater than or equal to the preset standard fluctuation value. The final sampling period is then determined as a unit time.

7. A teaching quality assessment system based on deep learning, characterized in that, include: The speech classification module is used to process teaching speech and classify it into noise speech, noisy speech, or normal speech. The evaluation model generation module is used to respond to the classification results of the speech classification module. It constructs a training dataset using features extracted from noisy and slurred speech, and trains a classification model based on the training dataset to generate a speech evaluation model. And a quality judgment module, which responds to the speech evaluation model generated by the evaluation model generation module, and uses the speech evaluation model to evaluate the feature data to be evaluated extracted from normal speech in order to determine the quality level of normal speech.

8. A teaching quality evaluation system based on deep learning according to claim 7, characterized in that, The voice classification module is used for: The teaching speech is screened to identify abnormal speech segments; and voiceprint comparison is performed on the abnormal speech segments to classify the teaching speech into noise speech or noisy speech based on whether the abnormal speech segments match the pre-stored voiceprints in the voiceprint database.

9. A teaching quality evaluation system based on deep learning according to claim 7, characterized in that, The quality assessment module is used for: The evaluation results are analyzed as the number of audio events occurring per unit of time. The score is calculated based on the frequency of occurrence and the corresponding preset evaluation weights; And compare the score with the preset standard evaluation value to determine the quality level.

10. A teaching quality evaluation system based on deep learning according to claim 9, characterized in that, The quality assessment module is also used for: The evaluation weight of noise events is calculated based on the cumulative duration of noise events within a unit of time; and a preset fixed value is used as the evaluation weight of noise events.