An artificial intelligence-based teaching effect evaluation system and a teaching speech emotion recognition method

By performing feature segment selection and gain compensation on the teacher's speech signal, combined with multi-path convolution and convolutional attention processing, the problem of emotion recognition under the high coupling of teaching style and emotional expression is solved, and the accuracy and robustness of emotion recognition are improved.

CN120766726BActive Publication Date: 2025-11-04SICHUAN VOCATIONAL & TECHN COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511295138.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-11-04
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

In speech scenarios where teachers' teaching styles are highly coupled with their emotional expressions, traditional emotion recognition methods struggle to accurately distinguish and capture emotional information, resulting in insufficient recognition accuracy and robustness.

Method used

By acquiring teachers' speech signals, filtering feature segments, using Mel frequency cepstral coefficient features and time-frequency entropy for gain compensation, and combining multi-path convolution and convolutional attention to perform gradient enhancement on the spectrogram, the probability distribution of teaching style and emotional fluctuations can be identified.

Benefits of technology

In speech scenarios where teaching style and emotion are highly coupled, the accuracy and robustness of emotion recognition are improved, the sensitivity to emotional changes and feature discrimination are enhanced, and the separation and analysis of style and emotion are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766726B_ABST
    Figure CN120766726B_ABST
Patent Text Reader

Abstract

The application provides a teaching effect evaluation system based on artificial intelligence and a teaching voice emotion recognition method. The voice signal of a teacher is subjected to feature segment screening to obtain different emotion fluctuation segments and teaching style segments. Gain compensation is performed on the prosody features in the teaching style segments to obtain gain compensation coefficients of the prosody features in the teaching style segments. The emotion fluctuation segments are converted into spectrograms, convolution attention of the spectrograms in different convolution channels is determined, and then gradient enhancement is performed on the emotion intensity distribution of the spectrograms to obtain gradient enhancement features of the emotion intensity distribution. The probability distribution of the teaching style and the probability distribution of the emotion fluctuation in the voice signal are recognized through the gain compensation coefficients of the prosody features and the gradient enhancement features of the emotion intensity distribution, and then an emotion recognition result of the voice signal is generated. The scheme of the application can realize emotion recognition in a voice scene in which the teaching style and emotional expression of a teacher are highly coupled to cause feature distortion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of teaching effect evaluation, more specifically, the present application relates to a teaching effect evaluation system based on artificial intelligence and a teaching speech emotion recognition method. BACKGROUND

[0002] Teaching effect evaluation measures the teaching quality of teachers and the learning effect of students through multi-dimensional data analysis. Traditional evaluation mainly relies on examination results, which is difficult to reflect emotional factors in the teaching process. In recent years, the development of teaching speech emotion recognition technology provides a new perspective for teaching effect evaluation. This technology uses prosodic features, speech recognition features, and emotion intensity distribution information, combined with probability models or deep learning methods, to accurately capture the emotional changes and teaching style characteristics in the teacher's speech. Through the analysis of emotional fluctuations and teaching style, the influence of teachers' emotional expression on the classroom atmosphere and student participation can be revealed. Integrating emotion recognition into teaching effect evaluation not only enriches the evaluation indicators, but also improves the scientificity and effectiveness of the evaluation, providing strong support for personalized teaching improvement and intelligent education management.

[0003] In the context of teacher teaching speech, teaching style and emotional expression are often closely intertwined. Teachers' speech not only carries information to convey knowledge, but also integrates diverse teaching styles and rich emotional colors. This high coupling leads to complex distortion and nonlinear changes in prosodic, intonation, speech rate, and other features in the speech signal, posing a huge challenge to traditional emotion recognition methods. Since the influence of teaching style on speech features is intertwined with emotional fluctuations, single features or simple models cannot accurately distinguish and capture emotional information, thereby affecting the accuracy and robustness of emotion recognition. Therefore, how to perform emotion recognition in the context of highly coupled teacher teaching style and emotional expression leading to feature distortion in speech has become a problem faced by the industry. SUMMARY

[0004] The present application provides a teaching effect evaluation system based on artificial intelligence and a teaching speech emotion recognition method, which can perform emotion recognition in the context of highly coupled teacher teaching style and emotional expression leading to feature distortion in speech.

[0005] In a first aspect, the present application provides a teaching speech emotion recognition method for performing speech emotion recognition in a teaching effect evaluation system based on artificial intelligence. The method includes the following steps:

[0006] Obtaining the speech signal of the teacher during classroom teaching;

[0007] Performing feature segment screening on the speech signal to obtain different emotional fluctuation segments and teaching style segments;

[0008] The prosodic features in the teaching style segment are gain compensated according to the mel-frequency cepstral coefficient features and the time-frequency entropy of the speech signal, to obtain gain compensation coefficients of the prosodic features in the teaching style segment.

[0009] The different spectrograms are converted from all the emotional fluctuation segments, multi-path convolution is performed on each spectrogram to obtain convolution attention of each spectrogram in different convolution paths, and gradient enhancement is performed on the emotional intensity distribution of each spectrogram through the convolution attention of each spectrogram in different convolution paths to obtain gradient enhancement features of each emotional intensity distribution.

[0010] The probability distribution of the teaching style and the probability distribution of the emotional fluctuation in the speech signal are identified through the gain compensation coefficients of the prosodic features and the gradient enhancement features of all the emotional intensity distributions, and then the emotional recognition result of the speech signal is generated based on all the probability distributions.

[0011] In some embodiments, the feature segment screening is performed on the speech signal to obtain different emotional fluctuation segments and teaching style segments, and specifically includes:

[0012] The speech signal is segmented to obtain different audio segments.

[0013] All the audio segments are clustered based on the radial basis distance between the audio segments to obtain different audio segment clusters.

[0014] The different emotional fluctuation segments are obtained by identifying and classifying all the audio segments according to a preset speech emotion classification model and the emotional label of the audio segment closest to the cluster center in each audio segment cluster.

[0015] All the audio segments except the emotional fluctuation segments are used as teaching style segments.

[0016] In some embodiments, the gain compensation of the prosodic features in the teaching style segment according to the mel-frequency cepstral coefficient features and the time-frequency entropy of the speech signal specifically includes:

[0017] The mel-frequency cepstral coefficient features and the time-frequency entropy of the speech signal are determined.

[0018] The stability of the prosodic features in the teaching style segment is analyzed through the time-frequency entropy to obtain a stability margin of the prosodic features in the teaching style segment.

[0019] The gain compensation coefficients of the prosodic features in the teaching style segment are obtained by superimposing and compensating the prosodic features in the teaching style segment through the stability margin and the mel-frequency cepstral coefficient features.

[0020] In some embodiments, converting all the emotional fluctuation segments into different spectrograms specifically comprises:

[0021] performing short-time Fourier transform on each emotional fluctuation segment to extract frequency spectrum information of each emotional fluctuation segment at different time points;

[0022] generating a spectrogram of each emotional fluctuation segment according to the frequency spectrum information of the emotional fluctuation segment at different time points.

[0023] In some embodiments, performing multi-path convolution on each spectrogram to obtain convolution attention of each spectrogram in different convolution paths specifically comprises:

[0024] performing parallel convolution on each spectrogram through convolution kernels of different scales to obtain scale feature maps of each spectrogram in different convolution paths;

[0025] performing channel attention enhancement on the scale feature maps of each spectrogram in different convolution paths to obtain channel attention of the scale feature maps of the spectrogram in different convolution paths;

[0026] performing spatial attention enhancement on the scale feature maps of each spectrogram in different convolution paths to obtain spatial attention of the scale feature maps of the spectrogram in different convolution paths;

[0027] performing feature fusion on the channel attention of the scale feature maps of each spectrogram in different convolution paths and the spatial attention of the scale feature maps of the spectrogram in different convolution paths to obtain the convolution attention of the spectrogram in different convolution paths.

[0028] In some embodiments, performing gradient enhancement on the emotional intensity distribution of each spectrogram through the convolution attention of the spectrogram in different convolution paths to obtain gradient enhancement features of each emotional intensity distribution specifically comprises:

[0029] determining the emotional intensity distribution of each spectrogram;

[0030] performing gradient mutation identification on the emotional intensity distribution of each spectrogram to obtain a gradient mutation region in the emotional intensity distribution of the spectrogram;

[0031] performing feature enhancement on the emotional intensity distribution of each spectrogram through the convolution attention of the spectrogram in different convolution paths and the gradient mutation region in the emotional intensity distribution of the spectrogram to obtain gradient enhancement features of each emotional intensity distribution.

[0032] In some embodiments, identifying the probability distribution of teaching style and the probability distribution of emotional fluctuation in the speech signal through the gain compensation coefficient of prosodic features and the gradient enhancement features of all emotional intensity distributions specifically comprises:

[0033] extracting speech recognition features of the speech signal through a preset speech recognition model;

[0034] probability mapping teaching style in the speech signal according to the speech recognition features and the gain compensation coefficient of all prosodic features, to obtain the probability distribution of teaching style in the speech signal;

[0035] probability mapping emotional fluctuation in the speech signal through the speech recognition features and the gradient enhancement features of all emotional intensity distributions, to obtain the probability distribution of emotional fluctuation in the speech signal.

[0036] In some embodiments, generating the emotional recognition result of the speech signal based on all probability distributions specifically comprises:

[0037] determining the emotional recognition confidence of the speech signal according to the probability distribution of emotional fluctuation and the probability distribution of teaching style;

[0038] taking the emotional recognition confidence as the emotional recognition result of the speech signal.

[0039] In some embodiments, the speech signal of the teacher in the classroom teaching is collected through an array microphone with directional sound pickup function in a digital teaching device.

[0040] In a second aspect, the present application provides a teaching effect evaluation system based on artificial intelligence, which comprises a teaching speech emotion recognition unit, and the teaching speech emotion recognition unit comprises:

[0041] an acquisition module configured to acquire a speech signal of a teacher in classroom teaching;

[0042] a processing module configured to perform feature segment screening on the speech signal to obtain different emotional fluctuation segments and teaching style segments;

[0043] The processing module is further configured to perform gain compensation on prosodic features in the teaching style segment according to the mel-frequency cepstral coefficient features and time-frequency entropy of the speech signal, to obtain the gain compensation coefficient of the prosodic features in the teaching style segment.

[0044] The processing module is further configured to convert all the emotional fluctuation segments into different spectrograms, perform multi-path convolution on each spectrogram to obtain convolution attention of each spectrogram in different convolution paths, perform gradient enhancement on the emotional intensity distribution of each spectrogram through the convolution attention of each spectrogram in different convolution paths, and obtain gradient enhancement features of each emotional intensity distribution.

[0045] The execution module is configured to identify the probability distribution of the teaching style and the probability distribution of the emotional fluctuation in the speech signal through the gain compensation coefficient of the prosodic feature and the gradient enhancement features of all the emotional intensity distributions, and further generate the emotional recognition result of the speech signal based on all the probability distributions.

[0046] The technical scheme provided by the embodiments disclosed in the present application has the following beneficial effects:

[0047] In the teaching effect evaluation system and the teaching speech emotion recognition method based on artificial intelligence provided by the present application, a speech signal of a teacher during classroom teaching is acquired through a digital teaching device; different emotional fluctuation segments and teaching style segments are obtained by performing feature segment screening on the speech signal; gain compensation is performed on prosodic features in the teaching style segments according to a mel-frequency cepstral coefficient feature and a time-frequency entropy of the speech signal, so as to obtain a gain compensation coefficient of the prosodic features in the teaching style segments; all the emotional fluctuation segments are converted into different spectrograms, multi-path convolution is performed on each spectrogram to obtain convolution attention of each spectrogram in different convolution paths, gradient enhancement is performed on an emotional intensity distribution of each spectrogram through the convolution attention of each spectrogram in different convolution paths, and gradient enhancement features of each emotional intensity distribution are obtained; the probability distribution of the teaching style and the probability distribution of the emotional fluctuation in the speech signal are identified through the gain compensation coefficient of the prosodic feature and the gradient enhancement features of all the emotional intensity distributions, and further the emotional recognition result of the speech signal is generated based on all the probability distributions.

[0048] Therefore, this application can identify the probability distribution of teaching style and the probability distribution of emotional fluctuation in the speech signal through the gain compensation coefficient of prosodic features and the gradient enhancement features of all emotional intensity distributions. Firstly, by filtering feature segments of the speech signal to distinguish emotional fluctuation segments from teaching style segments, a preliminary decoupling of style and emotional information can be achieved structurally, alleviating feature confusion caused by their mutual interference. Simultaneously, this step allows subsequent emotion analysis to focus more on the actual emotional change regions, improving the targeting of feature extraction. Secondly, by applying gain compensation to the prosodic features within the teaching style segments, the influence of teaching style factors on prosodic characteristics can be effectively weakened. The suppression effect of prosodic features restores their intrinsic state, and the gain compensation coefficient of prosodic features effectively reduces the recognition error caused by teaching style interference, thereby improving the accuracy of emotion recognition in teaching speech scenarios where teaching style and emotion are highly coupled. Next, converting emotional fluctuation segments into spectrograms helps transform time-series signals into a visualized time-frequency structure, thus more comprehensively showcasing emotion-related features such as pitch changes, energy distribution, and frequency dynamics in speech. Based on this, a multi-path convolutional neural network is used for processing, allowing for multi-scale and multi-angle extraction of emotional features from different convolutional kernel sizes, directions, and depths, enhancing the model's understanding of complex emotional patterns. The convolutional attention mechanism further weights and evaluates each feature channel, automatically focusing on the most emotionally discriminative regions and suppressing invalid feature interference from teaching styles or background noise, thus improving the efficiency of emotional information expression. Furthermore, convolutional attention enhances the emotional intensity distribution of the spectrogram through gradient enhancement, strengthening the expressive effect of high emotional intensity regions while suppressing redundant or interfering components. This ensures that emotional features remain clear and recognizable even under the influence of style. The gradient-enhanced features not only improve the ability to perceive subtle emotional fluctuations but also enhance the overall discriminative power of the features, thereby enabling performance in speech scenarios where style and emotion are highly coupled and feature distortion is severe. The present invention provides more accurate and robust emotion recognition. Gain compensation coefficients effectively restore the true impact of teaching style factors on prosodic features, reducing the misleading influence of style on emotion judgment. Gradient enhancement features strengthen the distribution characteristics of emotion signals in the spectrogram, improving sensitivity to emotion changes. Through gain compensation coefficients and gradient enhancement features, the probability distributions of teaching style and emotion fluctuations in the speech signal are identified, achieving separation and analysis of style and emotion. Finally, the emotion recognition result of the speech signal is generated based on all probability distributions. In summary, the solution of this application can achieve emotion recognition in speech scenarios where the high coupling between teachers' teaching style and emotional expression causes feature distortion. Attached Figure Description

[0049] Figure 1 This is an exemplary flowchart of a teaching speech emotion recognition method according to some embodiments of this application;

[0050] Figure 2 is a flowchart of determining a gain compensation coefficient according to some embodiments of the present application;

[0051] Figure 3 is a flowchart of determining convolution attention according to some embodiments of the present application;

[0052] Figure 4 is a structural diagram of a teaching speech emotion recognition unit according to some embodiments of the present application;

[0053] Figure 5 is a structural diagram of a computer device for implementing a teaching speech emotion recognition method according to some embodiments of the present application.

[0054] BRIEF DESCRIPTION OF THE DRAWINGS

[0055] 400, a teaching speech emotion recognition unit;

[0056] 401, an acquisition module;

[0057] 402, a processing module;

[0058] 403, an execution module;

[0059] 500, a computer device;

[0060] 501, a processor;

[0061] 502, a communication bus;

[0062] 503, a memory;

[0063] 504, a communication interface. DETAILED DESCRIPTION

[0064] In order to better understand the technical solutions of the present application, the technical solutions of the present application will be described in detail below in combination with the drawings in the specification and specific embodiments.

[0065] Referring to Figure 1 The figure is an exemplary flowchart of a teaching speech emotion recognition method according to some embodiments of the present application, which mainly includes the following steps:

[0066] In step 101, the speech signal of the teacher in classroom teaching is acquired.

[0067] In specific implementation, the speech signal of the teacher in classroom teaching is acquired through the database in the digital teaching device.

[0068] It should be noted that the digital teaching equipment in the present application represents a modern teaching hardware for assisting classroom teaching, collecting teaching data and realizing intelligent interaction; the voice signal represents the sound wave data of the teacher in the classroom teaching, and the voice signal of the teacher in the classroom teaching is collected through the array microphone with directional sound pickup function in the digital teaching equipment.

[0069] In step 102, the feature segment screening is performed on the voice signal to obtain different emotional fluctuation segments and teaching style segments.

[0070] In some embodiments, the feature segment screening on the voice signal to obtain different emotional fluctuation segments and teaching style segments can be realized by the following steps:

[0071] Segmenting the voice signal to obtain different audio segments;

[0072] Clustering all the audio segments based on the radial basis distance between the audio segments to obtain different audio segment clusters.

[0073] According to the preset voice emotion classification model and the emotional label of the audio segment closest to the cluster center in each audio segment cluster, all the audio segments are identified and classified to obtain different emotional fluctuation segments;

[0074] All the audio segments except the emotional fluctuation segments are used as teaching style segments.

[0075] It should be noted that the voice emotion classification model in the present application represents a deep learning model trained based on a large amount of labeled emotional data, which adopts convolutional neural network and recurrent neural network structures and can effectively capture the emotional change rule in the voice; the emotional label represents a category identification vector of the emotional state expressed in the audio segment; the emotional fluctuation segment represents an audio segment in which the emotional state in the voice signal fluctuates obviously; and the teaching style segment represents an audio segment reflecting the personal teaching expression characteristics of the teacher in the voice signal.

[0076] In specific implementation, firstly, the speech signal is segmented to obtain different audio segments. This can be achieved by using existing Silo VAD speech activity detection technology to segment the speech signal, thereby obtaining different audio segments. Secondly, all audio segments are clustered based on the radial basis function (RBF) distance between them to obtain different audio segment clusters. This can be achieved by using the Librosa library in Python to extract the chroma features of each audio segment, calculating the Euclidean distance between the chroma features of any two audio segments, inputting the Euclidean distance into the radial basis function, and using the output value of the radial basis function as the radial basis distance between the two audio segments. Then, a radial basis distance matrix is ​​constructed based on all the obtained radial basis distances. Based on the radial basis distance matrix, a spectral clustering algorithm is used to cluster all audio segments, and the resulting clusters are all taken as audio segment clusters. Finally, based on the preset speech emotion classification model and the values ​​in each audio segment cluster, the data are further analyzed. The emotional labels of the audio segments closest to the cluster center are used to identify and classify all audio segments to obtain different emotional fluctuation segments. This can be achieved in the following way: an emotional label of all audio segments is extracted using a preset speech emotion classification model. For each audio segment cluster, the cosine similarity between the emotional label of the audio segment closest to the cluster center and the emotional labels of all audio segments in the cluster is calculated. The mean of all cosine similarities is then calculated, and the mean is used as a similarity threshold. All audio segments with cosine similarities greater than the similarity threshold are considered as emotional fluctuation segments, thus obtaining different emotional fluctuation segments. Finally, the identified emotional fluctuation segments are removed from all audio segments, and the audio segments after removing the emotional fluctuation segments are considered as teaching style segments.

[0077] In step 103, the prosodic features within the teaching style segment are gain-compensated based on the Mel frequency cepstral coefficients and time-frequency entropy of the speech signal, thereby obtaining the gain compensation coefficients of the prosodic features within the teaching style segment.

[0078] In some embodiments, reference Figure 2 As shown in the figure, this is a flowchart illustrating the process of determining the gain compensation coefficient in some embodiments of this application. In this embodiment, gain compensation is performed on the prosodic features within the teaching style segment based on the Mel-frequency cepstral coefficient characteristics and time-frequency entropy of the speech signal. The gain compensation coefficient for the prosodic features within the teaching style segment can be obtained by the following steps:

[0079] First, in step 1031, the Mel frequency cepstral coefficient characteristics and time-frequency entropy of the speech signal are determined;

[0080] Secondly, in step 1032, the stability of the prosodic features in the teaching style segment is analyzed by the time-frequency entropy, and a stability margin of the prosodic features in the teaching style segment is obtained;

[0081] Then, in step 1033, the prosodic features in the teaching style segment are superimposed and compensated by the stability margin and the mel-frequency cepstral coefficient feature, and a gain compensation coefficient of the prosodic features in the teaching style segment is obtained.

[0082] It should be noted that the mel-frequency cepstral coefficient feature in the present application represents a feature reflecting the voice timbre and vocal tract characteristics of the speech signal; the time-frequency entropy represents the uncertainty of the speech signal in the time-frequency domain, and the time-frequency entropy can reflect the dynamic changes and prosodic information of the speech signal.

[0083] In specific implementation, the mel-frequency cepstral coefficient feature and the time-frequency entropy of the speech signal can be determined in the following manner: the Librosa speech processing tool is used to extract the cepstral coefficient of each speech frame from the speech signal, and a matrix composed of the cepstral coefficients of all speech frames is taken as the mel-frequency cepstral coefficient feature of the speech signal; the speech signal is subjected to fast Fourier transform to extract the time-frequency energy spectrum, then the time-frequency energy spectrum is subjected to normalization processing, the probability distribution of each time-frequency point in the time-frequency energy spectrum is calculated, finally the information entropy of all time-frequency points is calculated by using the Shannon entropy formula, and the obtained information entropy is taken as the time-frequency entropy.

[0084] It should be noted that the prosodic features in the present application represent the acoustic characteristics of the prosodic features of the teaching speech in the teaching style segment; the stability margin represents the margin reflecting the stability of the prosodic fluctuation in the teaching style segment.

[0085] In a specific implementation, the stability margin of the prosodic features in the teaching style segment can be obtained by performing stability analysis on the prosodic features in the teaching style segment using the time-frequency entropy, which can be achieved in the following manner. First, the pitch module in the Praat speech processing tool is used to extract the fundamental frequency trajectory of the teaching style segment, where the fundamental frequency trajectory reflects the pitch variation of the teaching style segment. The intensity module in the Praat is then used to extract the intensity information, where the intensity information reflects the sound loudness variation. The set composed of the fundamental frequency trajectory and the intensity information is taken as the prosodic features in the teaching style segment. Second, the fundamental frequency trajectory of the prosodic features in the teaching style segment is discretized into intervals, the frequency of occurrence of the fundamental frequency values in each interval is counted, the probability distribution of the fundamental frequency is calculated, the entropy value of the probability distribution is calculated using the Shannon entropy formula, and the obtained entropy value is taken as the fundamental frequency entropy value. Furthermore, the intensity information is discretized into intervals, the frequency of occurrence of the intensity values in each interval is counted, the probability distribution of the intensity is calculated, the entropy value of the probability distribution is calculated using the Shannon entropy formula, and the obtained entropy value is taken as the intensity entropy value. Then, the absolute values of the differences between the time-frequency entropy and the intensity entropy value and the fundamental frequency entropy value are taken, the sum of the two absolute values is calculated, and the obtained sum is taken as the stability margin of the prosodic features in the teaching style segment.

[0086] It should be noted that the gain compensation coefficient in the present application represents a parameter of non-emotional fluctuation caused by the enhancement and compensation of the prosodic features of the teaching style in the speech scenario where the high coupling between the teaching style and the emotional expression causes feature distortion.

[0087] In a specific implementation, the gain compensation coefficient of the prosodic features in the teaching style segment can be obtained by performing superimposed compensation on the prosodic features in the teaching style segment using the stability margin and the mel-frequency cepstral coefficient feature, which can be achieved in the following manner. First, the stability margin is taken as a weight to perform weighted adjustment on all the cepstral coefficients in the mel-frequency cepstral coefficient feature, and the matrix obtained after the weighted adjustment is taken as a superimposed compensation matrix. Each value in the superimposed compensation matrix is taken as a superimposed compensation value. Second, for each superimposed compensation value in the superimposed compensation matrix, the superimposed compensation value is added to each fundamental frequency value on the fundamental frequency trajectory in the prosodic features in the teaching style segment, and the sum is taken as a fundamental frequency compensation value. The superimposed compensation value is also added to each intensity value in the intensity information in the prosodic features in the teaching style segment, and the sum is taken as an intensity compensation value. In this way, different multiple fundamental frequency compensation values and multiple intensity compensation values are obtained. All the fundamental frequency compensation values are normalized, and the vector obtained after the normalization is taken as a fundamental frequency compensation vector. All the intensity compensation values are also normalized, and the vector obtained after the normalization is taken as an intensity compensation vector. The Euclidean distance between the intensity compensation vector and the fundamental frequency compensation vector is calculated, and the obtained Euclidean distance is taken as the gain compensation coefficient of the prosodic features in the teaching style segment.

[0088] In step 104, all the emotional fluctuation segments are converted into different spectrograms, each spectrogram is subjected to multi-path convolution to obtain the convolution attention of each spectrogram in different convolution paths, and the gradient enhancement feature of each emotional intensity distribution is obtained by gradient enhancement of the emotional intensity distribution of each spectrogram through the convolution attention of each spectrogram in different convolution paths.

[0089] In some embodiments, converting all the emotional fluctuation segments into different spectrograms can be achieved by the following steps:

[0090] Performing short-time Fourier transform on each emotional fluctuation segment to extract the frequency spectrum information of each emotional fluctuation segment at different time points;

[0091] Generating the spectrogram of each emotional fluctuation segment according to the frequency spectrum information of each emotional fluctuation segment at different time points.

[0092] It should be noted that the spectrogram in the present application represents the time-frequency intensity distribution diagram of the emotional fluctuation segment over time.

[0093] In specific implementation, first, performing short-time Fourier transform on each emotional fluctuation segment to extract the frequency spectrum information of each emotional fluctuation segment at different time points can be achieved by the following way, i.e., performing short-time Fourier transform on each emotional fluctuation segment to obtain the energy distribution of each emotional fluctuation segment at different time points, and taking the obtained energy distribution as the frequency spectrum information of the corresponding emotional fluctuation segment at different time points; second, generating the spectrogram of each emotional fluctuation segment according to the frequency spectrum information of each emotional fluctuation segment at different time points can be achieved by the following way, i.e., arranging the frequency spectrum information of each emotional fluctuation segment at different time points in time sequence to generate a two-dimensional matrix, wherein the horizontal axis is time, the vertical axis is frequency, and each element in the matrix represents the energy intensity of the corresponding time-frequency point, thereby constituting the time-frequency intensity distribution diagram of each emotional fluctuation segment, and taking the obtained time-frequency intensity distribution diagram as the spectrogram, thereby obtaining the spectrogram of each emotional fluctuation segment.

[0094] In some embodiments, referring to Figure 3 The figure is a flowchart for determining convolution attention in some embodiments of the present application, and in the present embodiment, multi-path convolution is performed on each spectrogram to obtain the convolution attention of each spectrogram in different convolution paths, which can be achieved by the following steps:

[0095] Parallel convolution is performed on each spectrogram by using convolution kernels of different scales to obtain the scale feature map of each spectrogram in different convolution paths;

[0096] channel attention enhancement is performed on the scale feature maps of each spectrogram in different convolution channels to obtain channel attention of the scale feature maps of each spectrogram in different convolution channels.

[0097] spatial attention enhancement is performed on the scale feature maps of each spectrogram in different convolution channels to obtain spatial attention of the scale feature maps of each spectrogram in different convolution channels.

[0098] feature fusion is performed on the channel attention of the scale feature maps of each spectrogram in different convolution channels and the spatial attention of the scale feature maps of each spectrogram in different convolution channels to obtain convolution attention of each spectrogram in different convolution channels.

[0099] It should be noted that the scale feature map in the present application represents a feature map of a spectrogram at different convolution scales.

[0100] In a specific implementation, the scale feature maps of each spectrogram in different convolution channels can be obtained by performing parallel convolution on each spectrogram by using convolution kernels of different scales, that is, by setting multiple convolution channels and using convolution kernels of different scales (for example, 3x3, 5x5, 7x7) to perform convolution on each spectrogram, and taking the feature maps obtained by convolution as the scale feature maps of each spectrogram in different convolution channels.

[0101] It should be noted that the channel attention in the present application represents channel features that contribute more to emotion recognition.

[0102] In a specific implementation, the channel attention of the scale feature maps of each spectrogram in different convolution channels can be obtained by the following method, that is, for the scale feature maps of each spectrogram in different convolution channels, a Squeeze-and-Excitation module is used to model the channel dimension of the scale feature maps, the Squeeze-and-Excitation module compresses the spatial dimension by global average pooling operation, thereby extracting the global information of each channel, then passes through two fully connected networks to generate weight coefficients of each channel, finally, multiplies all weight coefficients with the channels of the scale feature maps one by one, and takes the features obtained by multiplication as the channel attention of the scale feature maps, thereby obtaining the channel attention of the scale feature maps of each spectrogram in different convolution channels.

[0103] It should be noted that the spatial attention in the present application represents key region features that contribute more to emotion recognition.

[0104] In a specific implementation, the spatial attention enhancement of the scale feature map of each spectrogram in different convolution channels can be realized in the following manner: for the scale feature map of each spectrogram in different convolution channels, average pooling and maximum pooling operations are respectively performed on the channel dimension of the scale feature map to obtain two spatial description maps of the scale feature map, then the two spatial description maps of the scale feature map are spliced in the channel dimension, the spliced features are input into a convolution layer, a spatial attention map is generated through the convolution layer, then the spatial attention map is multiplied with the scale feature map element by element, and the multiplied features are taken as the spatial attention of the scale feature map, so as to obtain the spatial attention of the scale feature map of each spectrogram in different convolution channels.

[0105] It should be noted that the convolution attention in the present application represents the feature of enhancing the expression of the emotion feature of the spectrogram.

[0106] In a specific implementation, the feature fusion of the channel attention of the scale feature map of each spectrogram in different convolution channels and the spatial attention of the scale feature map of each spectrogram in different convolution channels can be realized in the following manner: the channel attention of the scale feature map of each spectrogram in different convolution channels and the spatial attention of the scale feature map of each spectrogram in different convolution channels are multiplied element by element, and the multiplied features are taken as the convolution attention of each spectrogram in different convolution channels.

[0107] In some embodiments, the gradient enhancement of the emotion intensity distribution of each spectrogram is performed through the convolution attention of each spectrogram in different convolution channels, and the gradient enhanced feature of each emotion intensity distribution can be realized in the following steps:

[0108] determining the emotion intensity distribution of each spectrogram;

[0109] performing gradient mutation identification on the emotion intensity distribution of each spectrogram to obtain a gradient mutation region in the emotion intensity distribution of each spectrogram;

[0110] performing feature enhancement on the emotion intensity distribution of each spectrogram through the convolution attention of each spectrogram in different convolution channels and the gradient mutation region in the emotion intensity distribution of each spectrogram, to obtain the gradient enhanced feature of each emotion intensity distribution.

[0111] It should be noted that the emotion intensity distribution in the present application represents the spatial distribution of the emotion expression intensity in the spectrogram.

[0112] In a specific implementation, the determination of the emotion intensity distribution of each spectrogram can be achieved by the following method: for each spectrogram, a fixed sliding window is used to scan the spectrogram from left to right and from top to bottom based on the sliding window technology, where the sliding window can be set to 32x32, and the sliding step can be set to 8; the edge intensity of the sliding window at each sliding position is calculated using the Sobel operator; a matrix composed of all edge intensities in the sliding order is obtained; the matrix is up-sampled to the size of the spectrogram using the bilinear interpolation method; and the matrix obtained after up-sampling is used as the emotion intensity distribution of the spectrogram, thereby obtaining the emotion intensity distribution of each spectrogram.

[0113] It should be noted that the gradient mutation region in the present application represents a continuous region in which the emotion intensity changes sharply in space in the emotion intensity distribution of the spectrogram.

[0114] In a specific implementation, the gradient mutation region in the emotion intensity distribution of each spectrogram can be obtained by the following method: for each spectrogram, different local maximum points in the emotion intensity distribution of the spectrogram are detected using a local maximum detection method; and the connectivity of the emotion intensity distribution of the spectrogram is analyzed using a region growing algorithm with each local maximum point as a seed point, thereby obtaining a plurality of different regions, and each obtained region is used as a gradient mutation region in the emotion intensity distribution of the spectrogram, and then the gradient mutation region in the emotion intensity distribution of each spectrogram is obtained.

[0115] It should be noted that the gradient enhancement feature in the present application represents a response feature that highlights the emotion expression in the spectrogram.

[0116] In a specific implementation, the gradient enhancement feature of each emotion intensity distribution can be obtained by the following method: for each spectrogram, the convolution attention of the spectrogram in different convolution channels and the gradient mutation region in the emotion intensity distribution of the spectrogram are weighted and fused respectively, then the features obtained by all the weighted and fused are spliced, and the matrix obtained by the splicing is used as the gradient enhancement feature of the emotion intensity distribution of the spectrogram, thereby obtaining the gradient enhancement feature of each emotion intensity distribution.

[0117] In step 105, the probability distribution of the teaching style and the probability distribution of the emotion fluctuation in the speech signal are identified by the gain compensation coefficient of the prosodic feature and the gradient enhancement feature of all emotion intensity distributions, and then the emotion recognition result of the speech signal is generated based on all the probability distributions.

[0118] In some embodiments, the probability distribution of teaching style and the probability distribution of emotional fluctuation in the speech signal can be identified by the gain compensation coefficient of prosodic features and the gradient enhancement feature of all emotional intensity distributions using the following steps:

[0119] extracting speech recognition features of the speech signal by a preset speech recognition model;

[0120] probability mapping of teaching style in the speech signal according to the speech recognition features and the gain compensation coefficient of all prosodic features, to obtain the probability distribution of teaching style in the speech signal;

[0121] probability mapping of emotional fluctuation in the speech signal by the speech recognition features and the gradient enhancement feature of all emotional intensity distributions, to obtain the probability distribution of emotional fluctuation in the speech signal.

[0122] It should be noted that the preset speech recognition model in the present application is Wav2Vec 2.0, wherein the Wav2Vec 2.0 model is an advanced self-supervised speech modeling method, which can realize efficient speech feature extraction in the absence of a large amount of manually labeled data. The Wav2Vec 2.0 model extracts low-level acoustic features from the original waveform through a convolutional neural network, and then uses a Transformer encoder to capture global context to generate high-dimensional context-aware representations. In the pre-training stage, a contrastive learning method is used to enable the model to distinguish between real and negative sampled feature segments, thereby enhancing semantic modeling capabilities. In the fine-tuning stage, a small amount of labeled data can be used to quickly adapt to specific speech recognition tasks.

[0123] It should also be noted that the speech recognition features in the present application represent the speech features of the signal frames in the speech signal.

[0124] In specific implementation, the speech recognition features of the speech signal can be extracted by the preset speech recognition model using the following method, i.e., extracting features of the speech signal by the preset speech recognition model to obtain the context-aware representation of each signal frame in the speech signal, wherein the context-aware representation is a high-dimensional feature vector reflecting the acoustic features of the signal frame in the context of the speech signal. All context-aware representations are concatenated in time sequence, and the vector obtained by concatenation is taken as the speech recognition features of the speech signal.

[0125] It should be noted that the teaching style segment in the present application is composed of multiple consecutive signal frames; and the probability distribution of teaching style in the speech signal represents the probability distribution of the tendency of the teacher's teaching style in the speech signal.

[0126] In a specific implementation, the probability distribution of the teaching style in the speech signal can be obtained by multiplying the gain compensation coefficient of the prosody feature and all the prosody features in the speech signal, and the probability distribution of the teaching style in the speech signal can be obtained by the following method: for each teaching style segment, the gain compensation coefficient of the prosody feature in the teaching style segment is multiplied by the context-aware representation of all signal frames in the teaching style segment, and then the context-aware representations multiplied by the gain compensation coefficient are averaged and pooled, and the vector obtained after the average pooling is taken as the compensation pooling vector of the teaching style segment, so as to obtain the compensation pooling vector of each teaching style segment, and then all the compensation pooling vectors are averaged and pooled, and the vector obtained after the normalization by the Softmax function is taken as the probability distribution of the teaching style in the speech signal.

[0127] It should be noted that the emotion fluctuation segment in the present application is composed of a plurality of continuous signal frames; each emotion fluctuation segment corresponds to a spectrogram, and each spectrogram corresponds to an emotion intensity distribution, that is, each emotion fluctuation segment corresponds to an emotion intensity distribution.

[0128] In addition, it should be noted that the probability distribution of the emotion fluctuation in the speech signal in the present application represents the probability distribution of the teacher emotion fluctuation in the speech signal.

[0129] In a specific implementation, the probability distribution of the emotion fluctuation in the speech signal can be obtained by multiplying the gain compensation coefficient of the prosody feature and all the prosody features in the speech signal, and the probability distribution of the teaching style in the speech signal can be obtained by the following method: for each teaching style segment, the gain compensation coefficient of the prosody feature in the teaching style segment is multiplied by the context-aware representation of all signal frames in the teaching style segment, and then the context-aware representations multiplied by the gain compensation coefficient are averaged and pooled, and the vector obtained after the average pooling is taken as the compensation pooling vector of the teaching style segment, so as to obtain the compensation pooling vector of each teaching style segment, and then all the compensation pooling vectors are averaged and pooled, and the vector obtained after the normalization by the Softmax function is taken as the probability distribution of the teaching style in the speech signal.

[0130] In some embodiments, the emotion recognition result of the speech signal can be obtained based on all the probability distributions by the following steps:

[0131] According to the probability distribution of the emotion fluctuation and the probability distribution of the teaching style, the emotion recognition confidence of the speech signal is determined.

[0132] The emotion recognition confidence is taken as the emotion recognition result of the speech signal.

[0133] It should be noted that the emotion recognition confidence in the present application represents a quantitative index of the reliability of the emotion recognition result in the speech signal.

[0134] In a specific implementation, the emotion recognition confidence of the speech signal is determined according to the probability distribution of the emotion fluctuation and the probability distribution of the teaching style, which can be implemented in the following manner: aligning the probability distribution of the emotion fluctuation with the probability distribution of the teaching style, selecting the maximum value at the same position, and taking the vector composed of all selected values as the emotion recognition vector, for example: the probability distribution of the emotion fluctuation is [0.25, 0.30, 0.20, 0.25], the probability distribution of the teaching style is [0.18, 0.32, 0.27, 0.23], and the emotion recognition vector is [0.25, 0.32, 0.27, 0.25]. Then, the sum of the values in the emotion recognition vector that belong to the probability distribution of the emotion fluctuation is calculated, and the sum is divided by the sum of all values in the emotion recognition vector, and the result of the division is taken as the emotion recognition confidence of the speech signal, for example: the emotion recognition confidence is (0.25+0.25) / (0.25+0.32+0.27+0.25)=0.46.

[0135] In addition, another aspect of the present application provides an artificial intelligence-based teaching effectiveness evaluation system, which, in some embodiments, includes a teaching speech emotion recognition unit, which is used for Figure 4 The figure is a structural schematic diagram of a teaching speech emotion recognition unit according to some embodiments of the present application. The teaching speech emotion recognition unit 400 includes an acquisition module 401, a processing module 402, and an execution module 403, which are described as follows:

[0136] The acquisition module 401 is mainly used for acquiring the speech signal of the teacher during classroom teaching in the present application.

[0137] The processing module 402 is used for feature segment screening of the speech signal to obtain different emotion fluctuation segments and teaching style segments in the present application.

[0138] It should be noted that the processing module 402 is also used for gain compensation of the prosodic features in the teaching style segment according to the mel-frequency cepstral coefficient features and the time-frequency entropy of the speech signal to obtain the gain compensation coefficient of the prosodic features in the teaching style segment.

[0139] It should be noted that the processing module 402 is also used to convert all the emotional fluctuation segments into different spectrograms, perform multi-path convolution on each spectrogram, obtain the convolution attention of each spectrogram in different convolution paths, perform gradient enhancement on the emotional intensity distribution of each spectrogram through the convolution attention of each spectrogram in different convolution paths, and obtain the gradient enhancement features of each emotional intensity distribution.

[0140] The execution module 403 is mainly used to identify the probability distribution of the teaching style and the probability distribution of the emotional fluctuation in the speech signal through the gain compensation coefficient of the prosodic feature and the gradient enhancement features of all the emotional intensity distributions, and further generate the emotional recognition result of the speech signal based on all the probability distributions.

[0141] In addition, the present application also provides a computer device, which comprises a memory and a processor, the memory stores code, and the processor is configured to acquire the code and execute the above-mentioned teaching speech emotion recognition method.

[0142] In some embodiments, with reference to Figure 5 The figure is a structural schematic diagram of a computer device for implementing the teaching speech emotion recognition method according to some embodiments of the present application. The teaching speech emotion recognition method in the above-mentioned embodiments can be implemented by the computer device shown in the figure, which comprises at least one processor 501, a communication bus 502, a memory 503, and at least one communication interface 504. Figure 5 The processor 501 can be a general central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more circuits for controlling the execution of the teaching speech emotion recognition method in the present application.

[0143] The communication bus 502 can be used to transmit information between the above-mentioned components.

[0144] The communication bus 502 can be used to transmit information between the above-mentioned components.

[0145] The memory 503 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM), or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory 503 can exist independently, and is connected to the processor 501 through the communication bus 502. The memory 503 can also be integrated with the processor 501.

[0146] The memory 503 is configured to store program code for implementing the solutions of the present application, and the processor 501 is configured to control the execution of the program code. The processor 501 is configured to execute the program code stored in the memory 503. The program code can include one or more software modules. The methods described in the above method embodiments can be implemented by the processor 501 and one or more software modules in the program code in the memory 503.

[0147] The communication interface 504 is configured to communicate with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc., using any transceiver-like device.

[0148] In a specific implementation, as an embodiment, the computer device can include a plurality of processors, each of which can be a single-CPU processor or a multi-CPU processor. The processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0149] The computer device described above can be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device can be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of the present application do not limit the type of the computer device.

[0150] In addition, the present application also provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the teaching speech emotion recognition method described above.

[0151] Although the preferred embodiments of the present application have been described, those skilled in the art who are informed of the basic inventive concept can make further changes and modifications to the embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0152] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.

Claims

1. A method for recognizing teaching speech emotions, performing speech emotion recognition in an AI-based teaching effectiveness evaluation system, characterized in that, The method includes the following steps: Acquire the teacher's voice signals during classroom teaching; By filtering feature segments from the speech signal, segments with different emotional fluctuations and teaching styles are obtained; Gain compensation is performed on the prosodic features within the teaching style segment based on the Mel frequency cepstral coefficient characteristics and time-frequency entropy of the speech signal, to obtain the gain compensation coefficient of the prosodic features within the teaching style segment. All emotional fluctuation segments are converted into different spectrograms. Multi-path convolution is performed on each spectrogram to obtain the convolutional attention of each spectrogram in different convolutional paths. The emotional intensity distribution of each spectrogram is gradient-enhanced by the convolutional attention of each spectrogram in different convolutional paths to obtain the gradient-enhanced feature of each emotional intensity distribution. The probability distributions of teaching style and emotional fluctuations in the speech signal are identified by the gain compensation coefficient of prosodic features and the gradient enhancement features of all emotional intensity distributions. Then, the emotion recognition result of the speech signal is generated based on all the probability distributions.

2. The method as described in claim 1, characterized in that, The speech signal is subjected to feature segment filtering to obtain different emotional fluctuation segments and teaching style segments, specifically including: The speech signal is segmented to obtain different audio segments; All audio segments are clustered based on the radial basis distance between them to obtain different audio segment clusters; Based on the preset speech emotion classification model and the emotion label of the audio segment closest to the cluster center in each audio segment cluster, all audio segments are identified and classified to obtain different emotion fluctuation segments; All audio clips, excluding those showing emotional fluctuations, are considered as teaching style clips.

3. The method as described in claim 1, characterized in that, Gain compensation is performed on the prosodic features within the teaching style segment based on the Mel-frequency cepstral coefficients and time-frequency entropy of the speech signal. The specific gain compensation coefficients for the prosodic features within the teaching style segment include: Determine the Mel-frequency cepstral coefficient characteristics and time-frequency entropy of the speech signal; The stability of the prosodic features within the teaching style segment is obtained by performing stability analysis on the prosodic features within the teaching style segment using the time-frequency entropy. The prosodic features within the teaching style segment are superimposed and compensated using the stability margin and the Mel frequency cepstral coefficient features to obtain the gain compensation coefficient of the prosodic features within the teaching style segment.

4. The method as described in claim 1, characterized in that, Converting all emotional fluctuations into different spectrograms specifically includes: A short-time Fourier transform is performed on each emotional fluctuation segment to extract the spectral information of each emotional fluctuation segment at different time points; A spectrogram of each emotional fluctuation segment is generated based on the spectral information of each emotional fluctuation segment at different time points.

5. The method as described in claim 1, characterized in that, Performing multi-path convolutions on each spectrogram to obtain the convolutional attention for each spectrogram in different convolutional paths specifically includes: Parallel convolutions are performed on each spectrogram using convolution kernels of different scales to obtain scale feature maps of each spectrogram in different convolutional paths; Channel attention enhancement is performed on the scale feature maps of each spectrogram in different convolutional paths to obtain the channel attention of the scale feature maps of each spectrogram in different convolutional paths; Spatial attention enhancement is performed on the scale feature maps of each spectrogram in different convolutional paths to obtain the spatial attention of the scale feature maps of each spectrogram in different convolutional paths; The channel attention of the scale feature map of each spectrogram in different convolutional paths and the spatial attention of the scale feature map of each spectrogram in different convolutional paths are fused to obtain the convolutional attention of each spectrogram in different convolutional paths.

6. The method as described in claim 1, characterized in that, Gradient enhancement of the sentiment intensity distribution of each spectrogram is achieved by performing convolutional attention on each spectrogram in different convolutional paths, resulting in gradient enhancement features for each sentiment intensity distribution. Specifically, these features include: Determine the distribution of sentiment intensity for each spectrogram; Gradient mutation identification is performed on the sentiment intensity distribution of each spectrogram to obtain the gradient mutation region in the sentiment intensity distribution of each spectrogram; By using convolutional attention in different convolutional paths for each spectrogram and gradient abrupt change regions in the sentiment intensity distribution of each spectrogram, the sentiment intensity distribution of each spectrogram is enhanced to obtain gradient-enhanced features for each sentiment intensity distribution.

7. The method as described in claim 1, characterized in that, The probability distributions of teaching style and emotional fluctuations in the speech signal are identified by using the gain compensation coefficient of prosodic features and the gradient enhancement features of all emotional intensity distributions. Specifically, this includes: The speech recognition features of the speech signal are extracted using a preset speech recognition model; Based on the gain compensation coefficients of the speech recognition features and all prosodic features, the teaching style in the speech signal is probabilistically mapped to obtain the probability distribution of the teaching style in the speech signal; The probability distribution of emotional fluctuations in the speech signal is obtained by probabilistically mapping the speech recognition features and gradient enhancement features of all emotional intensity distributions.

8. The method as described in claim 1, characterized in that, The emotion recognition results generated based on all probability distributions specifically include: The confidence level of emotion recognition of the speech signal is determined based on the probability distribution of emotional fluctuations and the probability distribution of teaching styles; The confidence level of the emotion recognition is used as the emotion recognition result of the speech signal.

9. The method as described in claim 1, characterized in that, The teacher's voice signal during classroom teaching is collected using an array microphone with directional pickup function within the digital teaching equipment.

10. A teaching effectiveness evaluation system based on artificial intelligence, the system comprising a teaching voice emotion recognition unit, characterized in that, The teaching voice emotion recognition unit includes: The acquisition module is used to acquire the teacher's voice signals during classroom teaching; The processing module is used to filter feature segments from the speech signal to obtain segments with different emotional fluctuations and teaching styles. The processing module is also used to perform gain compensation on the prosodic features within the teaching style segment based on the Mel frequency cepstral coefficient characteristics and time-frequency entropy of the speech signal, so as to obtain the gain compensation coefficient of the prosodic features within the teaching style segment. The processing module is also used to convert all emotional fluctuation segments into different spectrograms, perform multi-path convolution on each spectrogram to obtain the convolutional attention of each spectrogram in different convolutional paths, and perform gradient enhancement on the emotional intensity distribution of each spectrogram through the convolutional attention of each spectrogram in different convolutional paths to obtain the gradient enhancement feature of each emotional intensity distribution. The execution module is used to identify the probability distribution of teaching style and the probability distribution of emotional fluctuation in the speech signal through the gain compensation coefficient of prosodic features and the gradient enhancement features of all emotional intensity distributions, and then generate the emotion recognition result of the speech signal based on all probability distributions.

Citation Information

Patent Citations

  • Voice emotion recognition method, device and equipment and readable storage medium

    CN120356488A

  • Emotion recognition apparatus, emotion recognition model learning apparatus, methods and programs for the same

    US20230095088A1