Emotion adaptive evaluation method and device based on acoustic semantic dual-stream emotion representation

CN120766724BActive Publication Date: 2026-09-18HUAZHONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511006872.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2026-09-18
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

[0003]然而,不同教学场景下的环境噪声和场景特征差异会影响语音情感识别的准确性,且教学效果评估的可靠性较低

Benefits of technology

[0015] (1) This invention provides an adaptive emotion evaluation method based on dual-stream acoustic and semantic emotion representation. It extracts speech emotion features through a cross-modal fusion network of multi-level acoustic features and innovatively introduces standardized scene description templates. A teaching scene perception training mechanism guided by a large model-generated text transforms scene information into feature representations and injects them into the acoustic features as a joint representation. This design not only improves the robustness of the speech emotion recognition model in different teaching scenarios but also enhances the model's ability to perceive scene-related emotion features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766724B_ABST
    Figure CN120766724B_ABST
Patent Text Reader

Abstract

The application discloses an emotion adaptive evaluation method and device based on acoustic semantic double-flow emotion representation, which comprises the following steps: extracting and preprocessing the input speech signal to obtain mel-frequency cepstral coefficient features and mel-spectrogram features; converting teaching scene information generated by a large language model into scene features through a preset scene description template; adaptively fusing the mel-frequency cepstral coefficient features and the mel-spectrogram features injected with the scene features respectively to generate joint emotion features; and mapping the recognized joint emotion features to corresponding teaching styles through a first mapping mechanism and a second mapping mechanism, and evaluating the matching degree of the teaching styles and the current teaching scene, wherein the first mapping mechanism is an emotion-style mapping mechanism, and the second mapping mechanism is a style-scene mapping mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent sensing technology, and more specifically, to an adaptive emotion evaluation method and device based on dual-stream emotional representation of sound and semantics. Background Technology

[0002] Currently, artificial intelligence technology is widely used in the education field. Intelligent assessment technology is playing an increasingly important role in classroom teaching quality evaluation and teacher professional development. In actual teaching, teachers need to flexibly adjust their emotional changes according to different teaching scenarios (such as commenting on student answers, explaining difficult concepts, and classroom discussions). This adjustment is reflected in multiple dimensions, including the teacher's vocal characteristics (such as pitch, speaking speed, and volume) and emotional expression characteristics.

[0003] However, environmental noise and scene characteristics in different teaching scenarios can affect the accuracy of speech emotion recognition, and the reliability of teaching effectiveness evaluation is low. Summary of the Invention

[0004] To address at least one deficiency or improvement need in the existing technology, this invention provides an emotion adaptive assessment method and device based on voice-semantic dual-stream emotion representation. This method can accurately identify teachers' emotional expressions, thereby achieving accurate assessment of teaching emotions, further judging the matching degree between the current teaching emotions and the current scenario, and providing reasonable adjustment suggestions, effectively improving teaching effectiveness.

[0005] To achieve the above objectives, according to a first aspect of the present invention, an adaptive emotion assessment method based on dual-stream emotional representation of speech and semantics is provided. The method includes: extracting and preprocessing Mel-frequency cepstral coefficient features and Mel-spectral graph features from the input speech signal; converting teaching scenario information generated with the assistance of a large language model into scenario features using a preset scenario description template; adaptively fusing the Mel-frequency cepstral coefficient features and Mel-spectral graph features, which have received scene feature injections, to generate joint emotion features; and mapping the identified joint emotion features to corresponding teaching styles through a first mapping mechanism and a second mapping mechanism, and evaluating the degree of matching between the teaching style and the current teaching scenario, wherein the first mapping mechanism is an emotion-style mapping mechanism and the second mapping mechanism is a style-scenario mapping mechanism.

[0006] In an exemplary embodiment, the step of converting the teaching scenario information generated by the large language model into scenario features using a preset scenario description template includes: converting the scenario description text into a word sequence using a word segmenter, obtaining initial features through an encoder; performing feature projection and pooling operations on the Mel frequency cepstral coefficients, and performing feature projection and pooling operations on the Mel spectrogram branches; expanding the scenario features into temporal features and concatenating them with the Mel frequency cepstral coefficient features in the time dimension; and expanding the scenario features into spatial features and concatenating them with the image patch features of the Mel spectrogram branches.

[0007] In an exemplary embodiment, the step of adaptively fusing the Mel frequency cepstral coefficient features and Mel spectrogram features, which have respectively received scene feature injection, to generate joint emotion features includes: converting the Mel frequency cepstral coefficient features and Mel spectrogram features to the same dimension through linear projection; calculating the first attention output of the Mel spectrogram features for the Mel frequency cepstral coefficient features through a bidirectional attention mechanism; and calculating the second attention output of the Mel spectrogram features for the Mel frequency cepstral coefficient features.

[0008] In one exemplary embodiment, after the second attention output of the process of calculating the Mel frequency cepstral coefficient features for Mel spectrogram features, the method further includes: obtaining a fixed-dimensional feature vector through multi-head attention processing and global average pooling; and outputting an emotion category probability distribution through feature fusion and a multilayer perceptron.

[0009] In an exemplary embodiment, after outputting the emotion category probability distribution, the method further includes: obtaining the emotion category probability distribution for each time window and calculating a weighted average emotion distribution; calculating teaching style features through an emotion-teaching style mapping matrix; evaluating the scene matching degree based on a preset scene style weight vector through a teaching style-scene mapping matrix, and obtaining a teaching style evaluation vector by combining the conformity of acoustic features.

[0010] In an exemplary embodiment, after obtaining the teaching style evaluation vector, the method further includes: setting an evaluation threshold based on the teaching style evaluation vector; and triggering an adjustment suggestion when any dimension of teaching style, scene style, and acoustic feature conformity is lower than the evaluation threshold.

[0011] According to a second aspect of the present invention, an adaptive emotion assessment device based on dual-stream emotional representation of speech and semantics is also provided, comprising: an extraction unit for extracting and preprocessing Mel frequency cepstral coefficient features and Mel spectrogram features of the input speech signal; a conversion unit for converting teaching scene information generated with the assistance of a large language model into scene features through a preset scene description template; a fusion unit for adaptively fusing the Mel frequency cepstral coefficient features and Mel spectrogram features that have received scene feature injection respectively to generate joint emotion features; and a first evaluation unit for mapping the identified joint emotion features to corresponding teaching styles through a first mapping mechanism and a second mapping mechanism, and evaluating the degree of matching between the teaching style and the current teaching scene, wherein the first mapping mechanism is an emotion-style mapping mechanism and the second mapping mechanism is a style-scene mapping mechanism.

[0012] According to a third aspect of the invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to execute the above-described emotion adaptive evaluation method based on acoustic-semantic dual-stream emotion representation at runtime.

[0013] According to a fourth aspect of the present invention, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described emotion adaptive evaluation method based on voice-semantic dual-stream emotion representation through the computer program.

[0014] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:

[0015] (1) This invention provides an adaptive emotion evaluation method based on dual-stream acoustic and semantic emotion representation. It extracts speech emotion features through a cross-modal fusion network of multi-level acoustic features and innovatively introduces standardized scene description templates. A teaching scene perception training mechanism guided by a large model-generated text transforms scene information into feature representations and injects them into the acoustic features as a joint representation. This design not only improves the robustness of the speech emotion recognition model in different teaching scenarios but also enhances the model's ability to perceive scene-related emotion features.

[0016] (2) Two key mechanisms, emotion-style mapping and style-scenario mapping, were established to map the identified emotional characteristics to the corresponding teaching styles and assess their matching degree with the current teaching scenario. Through multi-dimensional evaluation indicators, real-time suggestions for adjusting teaching emotions can be provided to teachers, helping them better adapt to the needs of different teaching scenarios and improve teaching effectiveness. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating an optional emotion adaptive evaluation method based on dual-stream emotional representation based on acoustic and semantics, provided for an embodiment of this application;

[0019] Figure 2 A flowchart illustrating another optional emotion adaptive evaluation method based on acoustic-semantic dual-stream emotion representation is provided for the embodiments of this application.

[0020] Figure 3 A schematic diagram illustrating an optional audio data acquisition scenario in a teacher's teaching environment, provided as an embodiment of this application;

[0021] Figure 4 A flowchart illustrating an optional pre-trained large language model for injecting teaching scene awareness into an embodiment of this application;

[0022] Figure 5 A schematic diagram illustrating the results of an optional two-stream network model provided in an embodiment of this application;

[0023] Figure 6 A schematic diagram of the structure of an optional emotion adaptive evaluation device based on acoustic-semantic dual-stream emotion representation provided in an embodiment of this application;

[0024] Figure 7 This is a schematic diagram of an optional electronic device provided in an embodiment of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0026] The terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0027] According to one aspect of the embodiments of this application, an adaptive emotion evaluation method based on dual-stream emotional representation of speech and semantics is provided. The following is in conjunction with... Figure 1 This application describes an adaptive emotion assessment method based on dual-stream emotional representation of speech and semantics, provided by an embodiment of the present application.

[0028] Figure 1 This is a flowchart illustrating an optional emotion adaptive evaluation method based on voice-semantic dual-stream emotion representation provided in an embodiment of this application. Figure 1 As shown, the process of this method may include the following steps:

[0029] S102, extract and preprocess the Mel frequency cepstral coefficient features and Mel spectrogram features of the input speech signal;

[0030] S104, transforms the teaching scenario information generated with the assistance of the large language model into scenario features by using a preset scenario description template;

[0031] S106, adaptively fuse the Mel frequency cepstral coefficient features and Mel spectrogram features that have received scene feature injection respectively to generate joint emotion features;

[0032] S108, through the first mapping mechanism and the second mapping mechanism, the identified joint emotional features are mapped to the corresponding teaching style, and the matching degree between the teaching style and the current teaching scenario is evaluated, wherein the first mapping mechanism is an emotion-style mapping mechanism and the second mapping mechanism is a style-scenario mapping mechanism.

[0033] This application provides an adaptive emotion assessment method based on dual-stream emotional representation of speech and semantics, which is applicable to educational scenarios. It provides teachers with targeted suggestions for adjusting teaching emotions for different teaching scenarios, effectively improving teaching quality and effectiveness.

[0034] Specifically, firstly, the input speech signal is processed by framing. A sampling rate of 16kHz, a frame length of 32ms, and a frame shift of 16ms are used. This setup ensures 50% overlap between adjacent frames. The framing process can be represented as: x i(n)=x(M·i+n),n=0,1,…,N-1, where x i (n)∈R N Let x(·) represent the nth sampling point of the i-th frame, x(·) be the original speech signal, i be the frame index, M be the frame shift, and N be the frame length.

[0035] For Mel spectrogram extraction, a window function needs to be applied to each frame of signal to reduce spectral leakage. The process of windowing and Fourier transform can be represented as follows:

[0036]

[0037] Where w(n) is the Hamming window function, x ′ i(n) is the windowed signal, X i (k) represents the spectrum of the i-th frame signal, where k is the frequency index. The power spectrum is then weighted and summed using 40 Mel filters, followed by logarithmic compression, and finally adjusted to a uniform size using bilinear interpolation. This process can be represented as:

[0038]

[0039] Where S i (m) represents the output of the i-th frame after passing through the m-th Mel filter, H m (k) is the frequency response of the m-th Mel filter, |X i (k)| 2 This represents the power spectrum. S ′ i(m) is the filter output after taking the logarithm.

[0040] Mel frequency cepstral coefficients (MFCC) feature extraction. By performing a discrete cosine transform on the output of the Mel filter bank, the Mel frequency cepstral coefficients (MFCC) feature extraction can be expressed as:

[0041]

[0042] Among them MFCC i (n) represents the nth MFCC coefficient in the i-th frame, M is the number of Mel filters (40), and n ranges from 0 to 12, resulting in 13-dimensional basic MFCC coefficients. To better capture the dynamic changes in speech features, the first and second order difference coefficients of the MFCCs were then calculated to better identify prosodic patterns in emotional expression:

[0043]

[0044] Where Δ t Let c be the first-order difference coefficient at time t. t ΔΔ represents the MFCC coefficient at time t. tThese are the second-order difference coefficients. Finally, the static MFCC coefficients are combined with their dynamic features to form a 39-dimensional feature vector: C t =[MFCC t ,Δ t ,ΔΔ t ].

[0045] like Figure 2 As shown, the input teaching speech signal is extracted and preprocessed using Mel frequency cepstral coefficient features and Mel spectrogram features to provide basic feature representations for subsequent dual-stream feature processing. A standardized scene description template is used to transform the teaching scene information generated with the assistance of a large language model into feature representations, providing scene-aware guidance for dual-stream acoustic feature processing. Two independent branches are used to perform deep feature extraction on Mel frequency cepstral coefficient features and Mel spectrogram features respectively, while simultaneously receiving scene feature injections to enhance the scene-awareness of the features. The feature representations obtained from the two branches are adaptively fused to generate a comprehensive emotional feature representation, achieving effective integration of multimodal information. Through two key mechanisms, emotion-style mapping and style-scene mapping, the identified emotional features are mapped to corresponding teaching styles, and their matching degree with the current teaching scene is evaluated, providing teachers with suggestions for adjusting their teaching emotions.

[0046] Through steps S102 to S108, the input speech signal is extracted and preprocessed using Mel frequency cepstral coefficient features and Mel spectrogram features. The teaching scenario information generated by the large language model is transformed into scenario features using a preset scenario description template. The Mel frequency cepstral coefficient features and Mel spectrogram features, which have received scene feature injection, are adaptively fused to generate joint emotion features. Through a first mapping mechanism and a second mapping mechanism, the identified joint emotion features are mapped to corresponding teaching styles, and the matching degree between the teaching style and the current teaching scenario is evaluated. The first mapping mechanism is an emotion-style mapping mechanism, and the second mapping mechanism is a style-scenario mapping mechanism, which can accurately identify the teacher's emotional expression, thereby achieving accurate evaluation of teaching emotions, further judging the matching degree between the current teaching emotions and the current scenario, and providing reasonable adjustment suggestions, effectively improving teaching effectiveness.

[0047] In an exemplary embodiment, the step of converting the teaching scenario information generated with the assistance of a large language model into scenario features through a preset scenario description template includes:

[0048] S11, use a word segmenter to convert the scene description text into a word sequence, and obtain the initial features through an encoder;

[0049] S12, perform characteristic projection and pooling operations on the Mel frequency cepstral coefficients, and perform characteristic projection and pooling operations on the Mel spectrogram branches;

[0050] S13 expands the scene features into temporal features and concatenates them with the Mel frequency cepstral coefficient features in the time dimension; expands the scene features into spatial features and concatenates them with the image patch features of the Mel spectrogram branch.

[0051] In this embodiment, as Figure 3 As shown, a standardized scene description template for the large language model was designed to systematically describe the characteristics of different teaching scenarios. This template encompasses two key dimensions: the teaching environment and the teaching activities, comprehensively depicting the features of each scenario. The basic form of the template is: "This is a {teaching scenario}, and a {teaching activity} is underway." Here, {teaching scenario} describes the physical environment of the teaching scenario, such as "large lecture hall," "multimedia classroom," or "laboratory"; {teaching activity} describes the specific teaching activity, such as "the teacher is explaining a concept" or "students are having a group discussion." This template-based approach facilitates the next step of analyzing the scene text using the large language model, providing high-level prior information about the scene for the emotion recognition model.

[0052] Text Feature Encoding: Based on the designed large language model scene description template, the RoBERTa encoder is used to convert the text description generated by the large language model into a high-dimensional feature representation. First, a word segmenter is used to convert the scene description text into a sequence of words that the model can process. Then, the initial feature representation is obtained through the RoBERTa encoder.

[0053]

[0054] Where L is the sequence length. text prompt d is the scene description text generated based on the template. r This represents the hidden dimension of RoBERTa. This initial feature contains rich semantic information about the teaching scenario, laying the foundation for subsequent feature fusion.

[0055] Considering that the Mel frequency cepstral coefficients (MFCC) feature branch and the Mel spectrogram feature branch have different feature dimensions and processing methods, the text features need to undergo corresponding projection transformations to effectively fuse them with the features of the two branches. First, feature projection and pooling operations are performed on the MFCC branch:

[0056]

[0057] Among them, Linear mfcc For a learnable linear projection layer, AvgPool is an average pooling operation that compresses sequential features into a single scene representation vector. A similar process is then performed on the Mel spectrogram branch.

[0058]

[0059] Among them, Linear mel As another learnable linear projection layer, features are mapped to a 96-dimensional space to match the feature dimensions of the Mel spectrogram branches, ensuring that scene features can match the feature dimensions of both branches.

[0060] After feature projection is completed, the scene features are injected into two feature processing branches. For the MFCC branch, the scene features are expanded into temporal features and concatenated with the Mel frequency cepstral coefficients (MFCC) features in the time dimension. For the Mel spectrogram branch, the scene features are expanded into spatial features and concatenated with image patch features.

[0061]

[0062] Among them MFCC features The original Mel-frequency cepstral coefficient (MFCC) feature sequence is given, where T1 is the time step. The Expand operation expands the pooled scene features into a form that can be concatenated with the MFCC features. features For image patch feature sequences, N is the number of image patches. After the two acoustic feature branches are merged, each branch will have one more scene feature frame.

[0063] In an exemplary embodiment, the adaptive fusion of Mel frequency cepstral coefficient features and Mel spectrogram features, which have respectively received scene feature injection, to generate joint emotion features includes:

[0064] S21, the Mel frequency cepstral coefficient features and Mel spectrogram features are converted to the same dimension by linear projection;

[0065] S22, calculates the first attention output of Mel spectrogram features for Mel frequency cepstral coefficient features through a bidirectional attention mechanism;

[0066] S23, calculate the second attention output of the Mel spectrogram features for the Mel frequency cepstral coefficient features.

[0067] In this embodiment, as Figure 4 As shown, when processing Mel spectrogram features, the two-dimensional image features first need to be converted into a sequence form to adapt to the Transformer's processing mechanism. A patch-based processing strategy is adopted, dividing the input 224×224 Mel spectrogram into several non-overlapping 4×4 image patches, then flattening them into 16-dimensional vectors. These vectors are then mapped to a high-dimensional feature space using a learnable linear projection matrix, and a learnable positional code is added to each image patch. This process can be represented by the following formula: x = M P·b+p, where b∈R 16 M represents the flattened image block. P ∈R C×16 It is a learnable projection matrix, p∈R C Here, C is the position encoding vector, and C is the target feature dimension (set to 96). After this processing, the input features are transformed into a sequence of shape B×N×C, where B is the batch size. P = 4 represents the number of image patches, and P = 4 represents the size of the image patches.

[0068] Multi-stage feature extraction is performed through four progressive processing stages. In each stage, multiple Transformer blocks are used for feature transformation. Each Transformer block contains two main components: a window multi-head self-attention mechanism and a feedforward network. In the window multi-head self-attention calculation, the input feature X∈R is first processed... N×C The query (U), key (P), and value (M) matrices are obtained through three independent linear transformations:

[0069]

[0070] Among them W U W P W m ∈R C×C This is a learnable parameter matrix. Then, attention weights are computed and features are aggregated within each attention head:

[0071]

[0072] The outputs of multiple attention heads are concatenated and linearly transformed to obtain the final multi-head attention output:

[0073] MultiAttn(X)=Concat(head1,…,head h W out

[0074] Among them W out ∈R C×C This is for outputting the projection matrix.

[0075] To break the limitations of fixed windows and enhance cross-window information flow, a cyclic shift window mechanism is introduced. By alternating between regular windows and shift windows between adjacent Transformer layers, features that were not originally in the same window can interact. Specifically, for an input feature map X∈R... H×W×C First, the feature map is divided into regular M×M windows. In the shifted window attention layer, the feature map first undergoes a cyclic shift operation, with the shift distance being half the window size. This shift operation can be represented as: CyclicShift represents a cyclic shift operation. The shift distance is rounded down. The window size M is set to 7, so that each window contains 7×7 image patches. This setting maintains computational efficiency while capturing an appropriate range of local dependencies. On the shifted feature map, the window is re-divided and self-attention is computed. To accurately encode positional information, a relative position encoding matrix B is introduced, enabling the attention computation to be aware of feature positional relationships.

[0076]

[0077] in This is a relative position encoding matrix, containing the relative positional relationships between all position pairs within the window. After calculation, the original permutation X of the features is recovered through a reverse shift operation. out .

[0078] Feature downsampling is introduced between different stages, allowing the network to gradually transition from local detailed features to global features. At each downsampling point, the feature map is reconstructed using a non-overlapping 2×2 window. Specifically, for the input feature X∈R... H×W×C First, feature rearrangement and normalization are performed: X reshaped =Reshape(LayerNorm(X)), where LayerNorm ensures the stability of the feature distribution. Then, the feature dimensions are adjusted through a linear transformation: Y = X reshaped ·W+b, where W∈R 4c×2c As a learnable transformation matrix, output features The spatial dimension is halved, while the number of channels is doubled. Through this hierarchical feature extraction, fine-grained time-frequency features are preserved in the shallow layer, while more abstract high-level information is captured in the deep layer.

[0079] Specifically, to obtain the Mel frequency cepstral coefficients (MFCC) features, Figure 5 A portion of the diagram shows the Voiceformer network model based on the speech signal structure used in this process, such as... Figure 5 As shown:

[0080] Considering that the typical duration of phonemes in speech is between 20-80 ms, feature encoding is first performed at the most basic acoustic unit level. For the input Mel-frequency cepstral coefficient (MFCC) feature matrix, its feature dimension is first expanded by linear projection, and then an overlapping segmentation strategy is adopted, using a sliding window to capture local contextual information. This process can be represented as:

[0081]

[0082] in The input is the Mel-frequency cepstral coefficient (MFCC) feature matrix, where d1 = 256 is the projected feature dimension, and Tw1 is set to the number of frames corresponding to 50ms. Specifically, for a 16kHz sampling rate speech signal, T1 is typically the audio length (in seconds) × 62.5; for example, for a 4-second audio segment, T1 ≈ 250 frames. For each segment, a self-attention mechanism is used to capture local contextual relationships.

[0083]

[0084] Based on the acoustic unit encoding described above, we will next construct higher-level feature representations to capture syllable and word-level speech patterns. Considering that the average duration of a syllable in Chinese is approximately 200-250ms, we will first create learnable word vectors. As a global reference, the features of adjacent acoustic units are then grouped and integrated. This process is achieved through the following steps:

[0085]

[0086] in k is typically set to 8 to cover a complete syllable. For each group of features α j G is obtained by aggregation using an attention-weighted approach. j :

[0087]

[0088] Where ∑ represents the summation of all tokens in the current group j, W q It is a learnable projection matrix.

[0089] Then, word-level encoding is performed on the aggregated feature sequence:

[0090]

[0091] To systematically integrate speech features across different time scales, a multi-level feature aggregation mechanism was designed. This mechanism is based on the prosodic hierarchy discovered in phonetic research, including a phoneme layer (approximately 50ms), a syllable layer (approximately 200ms), and a prosodic phrase layer (approximately 1000ms). Furthermore, features are progressively aggregated to different time scales: Where M i These correspond to the merging window sizes for the three time scales. Specifically, M1 is set to 4 (for the phoneme level), M2 to 16 (for the syllable level), and M3 to 80 (for the prosodic phrase level). These window sizes are based on typical time scales in phonetic research. Finally, these multi-scale features are integrated using residual connectivity and layer normalization. Where γ iThese are learnable scale weights used to balance the contributions of features at different time scales.

[0092] The cross-modal fusion module achieves deep interaction between Mel frequency cepstral coefficient (MFCC) features and Mel spectrogram features through a bidirectional attention mechanism. It adaptively fuses the feature representations obtained from the two branches to generate a comprehensive sentiment feature representation, thus effectively integrating multimodal information. Specifically, it includes the following steps:

[0093] First, the feature dimensions of the two branches are unified to facilitate subsequent cross-modal interactions. Linear projections are then performed on the Mel frequency cepstral coefficients (MFCC) features and the Mel spectrogram features, respectively.

[0094]

[0095] Where M∈R T1×256 Let S be the characteristic sequence of Mel frequency cepstral coefficients (MFCC), where S ∈ R. N×96 For the characteristic sequence of Mel spectrogram, W m ∈R 256×d and W s ∈R 96×d The projection matrix is ​​learnable, and d = 128 represents the uniform hidden dimension. T1 is the length of the MFCC sequence (approximately 250 frames), and N is the number of image patches.

[0096] To fully capture the interrelationship between the two modal features, a bidirectional cross-modal attention mechanism was designed. First, the attention of the Mel frequency cepstral coefficients (MFCC) features to the Mel spectrogram features was calculated:

[0097]

[0098] Among them W qm W ks W vs ∈R d×d The transformation matrix for query, key, and value. The Mel-frequency cepstral coefficients (MFCC) features are represented by the Mel spectrogram after attention. Similarly, the attention A of the Mel spectrogram features to the Mel-frequency cepstral coefficients (MFCC) features is calculated. s2m .

[0099] In one exemplary embodiment, after the second attention output of the process for calculating the Mel frequency cepstral coefficient features for Mel spectrogram features, the method further includes:

[0100] S31, after multi-head attention processing and global average pooling, yields a fixed-dimensional feature vector;

[0101] S32 outputs the probability distribution of sentiment categories through feature fusion and multilayer perceptron.

[0102] For example, the attention-followed feature sequence is subjected to temporal-dimensional global average pooling to obtain a fixed-dimensional feature representation f. m and f s Then the two features are concatenated and fused using a nonlinear transformation: f c =σ(W c [f m ;f s ]+b c ), where [;] represents the feature concatenation operation, W c ∈R d×2d Here, σ represents the parameters of the fusion layer, and σ is the ReLU activation function. To enhance the expressive power of the features, layer normalization is introduced: f n =LayerNorm(f c +f m +f s ), in which residual connections are added to preserve the original feature information.

[0103] Finally, sentiment classification was performed using a multilayer perceptron.

[0104]

[0105] in Here, C = 7 represents the number of sentiment categories, and y represents the final category probability distribution. Through this progressive feature fusion and transformation, the model can fully utilize the complementary information of the two modalities, improving the accuracy of sentiment recognition.

[0106] In one exemplary embodiment, after the output sentiment category probability distribution, the method further includes:

[0107] S41, Obtain the probability distribution of sentiment categories for each time window and calculate the weighted average sentiment distribution;

[0108] S42, teaching style characteristics are calculated using the affect-teaching style mapping matrix;

[0109] S43, the scene matching degree is evaluated based on the preset scene style weight vector through the teaching style-scene mapping matrix, and the teaching style evaluation vector is obtained by combining the conformity of acoustic features.

[0110] Optionally, a temporal analysis of emotional expression during the teaching process is performed. Since teachers' emotional expression dynamically changes with the teaching scenario, the teaching process is divided into 300-second windows (considering the integrity of the teaching content and the continuity of emotional expression). For each time window t, the probability distributions of seven basic emotions (happiness, concern, seriousness, calmness, enthusiasm, attentiveness, and neutrality) are obtained: P t =[p{t,1} ,p {t,2} ,…,p {t,7} ],

[0111] Where p {t,i} Let represent the probability of the i-th basic emotion within the t-th time window. To comprehensively grasp the characteristics of emotional changes during the teaching process, a weighted average emotion distribution is calculated: Here E d This reflects the average distribution of various emotions throughout the teaching process. t It is a time window weight, which will be given higher weight in key teaching stages.

[0112] Furthermore, two key mapping mechanisms are established: 1. Affective-teaching style mapping: defining a matrix. Where n style This represents the number of teaching styles. Each row can be seen as the combined weight of a particular teaching style on the seven emotion categories. The seven-dimensional emotion distribution E obtained from a segment of the teaching process... d Perform a linear mapping: In order to obtain a standardized distribution of teaching styles, we need to analyze the style... raw Normalization is performed: Style current Each component is located in the interval [0,1], and the sum of the elements is 1, thus realizing the mapping of multiple emotions to multiple styles.

[0113] For example, a "positive and encouraging" teaching style can have a higher weighting for emotions such as "happiness, enthusiasm, and concern," while the weighting for emotions such as "seriousness and neutrality" may be lower. At the same time, a "warm and interactive" style is also allowed to have corresponding weightings for "happiness" and "peace." In this way, the same emotion can affect different styles, and different emotions can also influence the same style, thus achieving the flexibility of many-to-many mapping.

[0114] 2. Teaching Style-Scenario Mapping: Each teaching scenario has its ideal combination of teaching styles. For example, the style requirements for a student answering and commenting scenario can be represented as: Where w i This represents the ideal weight of the i-th teaching style in this scenario. It can be set by education experts based on their teaching experience, or it can be learned in actual teaching scenarios.

[0115] Based on these two mapping relationships, the matching degree between the current teaching style and the scenario requirements is calculated: M s =Sim(Style) current Scene required ), where Style current This is the currently detected distribution of teaching styles, Scene requiredThis represents the ideal style distribution required for the current scenario. Similarity is calculated using Sim(S1,S2) = 1 - JSD(S1||S2), where JSD is the Jensen-Shannon divergence. Simultaneously, the acoustic features are evaluated to determine if they meet the requirements of the teaching style needed for the current scenario.

[0116]

[0117] Where f k f represents the actual acoustic characteristic values ​​(pitch, speech rate, volume). {k,std} It is the standard feature value of the dominant teaching style in the current scenario, σ k It is the allowable deviation range, β k It is the feature importance weight.

[0118] In one exemplary embodiment, after obtaining the teaching style evaluation vector, the method further includes:

[0119] S51, based on the teaching style assessment vector, set the assessment threshold;

[0120] S52, when any dimension of teaching style, scene style and acoustic feature conformity is lower than the evaluation threshold, an adjustment suggestion is triggered.

[0121] The evaluation results from each dimension are integrated into a teaching style evaluation vector:

[0122] S=[α1·Style match ;α2·Scene match α3·C a ]

[0123] Style match Accuracy reflecting teaching style (0-1), Scene match Reflects scene adaptability (0-1), C a Reflecting the acoustic feature conformity (0-1), where α1, α2, and α3 are the weights of each dimension, and α1 + α2 + α3 = 1. Based on the evaluation vector S, an evaluation threshold θ is set. When any dimension falls below the threshold, an adjustment suggestion is triggered: 1. Style match <θ: Indicates that the teaching style is not accurate enough. Example: "The current emotional expression is too serious. It is recommended to increase positive and constructive emotional expression." 2. Scene match <θ: Indicates a mismatch with the scene requirements. Example: "The current scene needs more interactive features; it is recommended to add questions and feedback." 3.C a <θ: indicates that the acoustic characteristics need to be adjusted. Example: "It is recommended to increase the speaking speed appropriately when explaining to motivate the audience, and to use an upward tone to avoid a flat tone."

[0124] A threshold-based multi-dimensional assessment mechanism can help teachers precisely adjust their teaching emotions. Table 1 below is an example of such a mechanism:

[0125] Table 1

[0126]

[0127] This embodiment employs a multi-dimensional teaching style assessment mechanism that uses emotion-teaching style mapping and scenario fit evaluation to monitor the rationality of teaching emotions in real time and provide teachers with specific and feasible adjustment suggestions, effectively improving teaching outcomes. The introduction of teaching scenario perception methods makes emotion recognition and style assessment results more aligned with actual teaching needs, enhancing the application value of the method in real teaching environments.

[0128] According to another aspect of the embodiments of this application, an evaluation apparatus is also provided for implementing the above-described emotion adaptive evaluation method based on the dual-stream emotional representation of speech and semantics. Figure 6 This is a schematic diagram of the structure of an optional emotion adaptive evaluation device based on voice-semantic dual-stream emotion representation according to an embodiment of this application, as shown below. Figure 6 As shown, the device may include:

[0129] Extraction unit 602 is used to extract and preprocess Mel frequency cepstral coefficient features and Mel spectrogram features from the input speech signal;

[0130] The conversion unit 604 is used to convert the teaching scenario information generated with the assistance of the large language model into scenario features through a preset scenario description template.

[0131] The fusion unit 606 is used to adaptively fuse the Mel frequency cepstral coefficient features and the Mel spectrogram features that have received scene feature injection respectively to generate joint emotion features;

[0132] The first evaluation unit 608 is used to map the identified joint emotional features to the corresponding teaching style through a first mapping mechanism and a second mapping mechanism, and to evaluate the degree of matching between the teaching style and the current teaching scenario. The first mapping mechanism is an emotion-style mapping mechanism and the second mapping mechanism is a style-scenario mapping mechanism.

[0133] It should be noted that the extraction unit 602 in this embodiment can be used to perform the above step S102, the conversion unit 604 in this embodiment can be used to perform the above step S104, the fusion unit 606 in this embodiment can be used to perform the above step S106, and the first evaluation unit 608 in this embodiment can be used to perform the above step S108.

[0134] Through the aforementioned modules, the input speech signal is extracted and preprocessed using Mel frequency cepstral coefficient features and Mel spectrogram features. The teaching scenario information generated with the assistance of a large language model is transformed into scenario features using a preset scenario description template. The Mel frequency cepstral coefficient features and Mel spectrogram features, which have received scene feature injection respectively, are adaptively fused to generate joint emotion features. Through a first mapping mechanism and a second mapping mechanism, the identified joint emotion features are mapped to corresponding teaching styles, and the matching degree between the teaching style and the current teaching scenario is evaluated. The first mapping mechanism is an emotion-style mapping mechanism, and the second mapping mechanism is a style-scenario mapping mechanism, which can accurately identify the teacher's emotional expression, thereby achieving accurate evaluation of teaching emotions, further judging the matching degree between the current teaching emotions and the current scenario, and providing reasonable adjustment suggestions, effectively improving teaching effectiveness.

[0135] In one exemplary embodiment, the conversion unit includes:

[0136] The acquisition module is used to convert scene description text into a sequence of words using a word segmenter and to acquire initial features through an encoder.

[0137] The pooling module is used to perform feature projection and pooling operations on Mel frequency cepstral coefficients and on Mel spectrogram branches.

[0138] The stitching module is used to expand scene features into temporal features and stitch them with Mel frequency cepstral coefficient features in the time dimension, and to expand scene features into spatial features and stitch them with image patch features of Mel spectrogram branches.

[0139] In one exemplary embodiment, the fusion unit includes:

[0140] The linear projection module is used to convert Mel frequency cepstral coefficient features and Mel spectrogram features to the same dimension through linear projection.

[0141] The first calculation module is used to calculate the first attention output of the Mel spectrogram features for the Mel frequency cepstral coefficient features through a bidirectional attention mechanism.

[0142] The second calculation module is used to calculate the second attention output of the Mel spectrogram features for the Mel frequency cepstral coefficient features.

[0143] In one exemplary embodiment, the apparatus further includes:

[0144] The average pooling unit is used to obtain a fixed-dimensional feature vector after multi-head attention processing and global average pooling.

[0145] The output unit is used to output the probability distribution of sentiment categories through feature fusion and a multilayer perceptron.

[0146] In one exemplary embodiment, the apparatus further includes:

[0147] The first calculation unit is used to obtain the probability distribution of sentiment categories for each time window and calculate the weighted average sentiment distribution.

[0148] The second calculation unit is used to calculate teaching style characteristics through the emotion-teaching style mapping matrix;

[0149] The second assessment unit is used to evaluate the scene matching degree based on the preset scene style weight vector through the teaching style-scene mapping matrix, and to obtain the teaching style evaluation vector by combining the conformity of acoustic features.

[0150] In one exemplary embodiment, the apparatus further includes:

[0151] The unit is set to define the assessment threshold based on the teaching style assessment vector;

[0152] The triggering unit is used to trigger adjustment suggestions when any dimension of teaching style, scene style, or acoustic feature conformity falls below the evaluation threshold.

[0153] It should be noted that the examples and scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run in a hardware environment and can be implemented by software or hardware. The hardware environment includes a network environment.

[0154] According to another aspect of the embodiments of this application, a storage medium is also provided. Optionally, in this embodiment, the storage medium can be used to execute the program code of any of the above-described emotion adaptive evaluation methods based on acoustic-semantic dual-stream emotion representations in the embodiments of this application.

[0155] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps:

[0156] S1, extracts and preprocesses the Mel frequency cepstral coefficient features and Mel spectrogram features of the input speech signal;

[0157] S2 transforms the teaching scenario information generated with the assistance of the large language model into scenario features through a preset scenario description template;

[0158] S3, adaptively fuses the Mel frequency cepstral coefficient features and Mel spectrogram features that have received scene feature injection respectively, to generate joint emotion features;

[0159] S4. Through the first mapping mechanism and the second mapping mechanism, the identified joint emotional features are mapped to the corresponding teaching styles, and the matching degree between the teaching styles and the current teaching scenario is evaluated. The first mapping mechanism is an emotion-style mapping mechanism, and the second mapping mechanism is a style-scenario mapping mechanism.

[0160] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated in this embodiment.

[0161] The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, as well as magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0162] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described emotion adaptive evaluation method based on voice-semantic dual-stream emotion representation is also provided. The electronic device may be a server, a terminal, or a combination thereof.

[0163] Figure 7 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application, such as... Figure 7 As shown, it includes a processor 702, a communication interface 704, a memory 706, and a communication bus 708. The processor 702, communication interface 704, and memory 706 communicate with each other via the communication bus 708.

[0164] Memory 706 is used to store computer programs;

[0165] When processor 702 executes a computer program stored in memory 706, it performs the following steps:

[0166] S1, extracts and preprocesses the Mel frequency cepstral coefficient features and Mel spectrogram features of the input speech signal;

[0167] S2 transforms the teaching scenario information generated with the assistance of the large language model into scenario features through a preset scenario description template;

[0168] S3, adaptively fuses the Mel frequency cepstral coefficient features and Mel spectrogram features that have received scene feature injection respectively, to generate joint emotion features;

[0169] S4. Through the first mapping mechanism and the second mapping mechanism, the identified joint emotional features are mapped to the corresponding teaching styles, and the matching degree between the teaching styles and the current teaching scenario is evaluated. The first mapping mechanism is an emotion-style mapping mechanism, and the second mapping mechanism is a style-scenario mapping mechanism.

[0170] Optionally, the communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The symbol is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned electronic device and other devices.

[0171] The memory may include RAM, or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0172] As an example, the memory 706 described above may include, but is not limited to, the extraction unit 602, the conversion unit 604, the fusion unit 606, and the first evaluation unit 608 of the emotion adaptive evaluation device based on dual-stream emotional representation of speech and semantics. Furthermore, it may include, but is not limited to, other module units of the emotion adaptive evaluation device based on dual-stream emotional representation of speech and semantics, which will not be elaborated upon in this example.

[0173] The processors mentioned above can be general-purpose processors, including but not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; they can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0174] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here.

[0175] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0176] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0177] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between devices or units may be electrical or other forms.

[0178] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0179] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0180] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0181] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0182] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of embodiments of this disclosure upon considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described herein. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

[0183] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0184] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An adaptive emotion assessment method based on dual-stream vocal and semantic emotion representation, characterized in that, include: The input speech signal is extracted and preprocessed using Mel frequency cepstral coefficient features and Mel spectrogram features; The teaching scenario information generated with the assistance of the large language model is transformed into scenario features by using a preset scenario description template; The Mel frequency cepstral coefficient features and Mel spectrogram features, which have received scene feature injection respectively, are adaptively fused to generate joint emotion features; The first mapping mechanism and the second mapping mechanism are used to map the identified joint emotional features to the corresponding teaching styles and evaluate the degree of matching between the teaching styles and the current teaching scenario. The first mapping mechanism is an emotion-style mapping mechanism and the second mapping mechanism is a style-scenario mapping mechanism. The step of converting the teaching scenario information generated with the assistance of the large language model into scenario features through a preset scenario description template includes: The scene description text is converted into a sequence of words using a word segmenter, and the initial features are obtained through an encoder. Characteristic projection and pooling operations are performed on the cepstral coefficients of the Mel frequency, and characteristic projection and pooling operations are performed on the branches of the Mel spectrogram; Scene features are extended into temporal features and concatenated with Mel frequency cepstral coefficient features in the time dimension. Scene features are also extended into spatial features and concatenated with image patch features of the Mel spectrogram branch.

2. The emotion adaptive evaluation method based on dual-stream emotional representation of speech and semantics as described in claim 1, characterized in that, The adaptive fusion of Mel frequency cepstral coefficient features and Mel spectrogram features, which have received scene feature injection respectively, to generate joint emotion features includes: The Mel frequency cepstral coefficient features and Mel spectrogram features are converted to the same dimension by linear projection; The first attention output of Mel spectrogram features for Mel frequency cepstral coefficient features is calculated using a bidirectional attention mechanism. Calculate the second attention output for the Mel spectrogram features with respect to the Mel frequency cepstral coefficient features.

3. The emotion adaptive evaluation method based on dual-stream emotional representation of speech and semantics as described in claim 2, characterized in that, Following the second attention output of the process of calculating the Mel frequency cepstral coefficient features for Mel spectrogram characteristics, the method further includes: After multi-head attention processing and global average pooling, a fixed-dimensional feature vector is obtained; By using feature fusion and a multilayer perceptron, the probability distribution of sentiment categories is output.

4. The emotion adaptive evaluation method based on dual-stream emotional representation of speech and semantics as described in claim 3, characterized in that, Following the output sentiment category probability distribution, the method further includes: Obtain the probability distribution of sentiment categories for each time window and calculate the weighted average sentiment distribution; Teaching style characteristics are obtained by calculating the affect-teaching style mapping matrix; The teaching style-scene mapping matrix is ​​used to evaluate the scene matching degree based on the preset scene style weight vector, and the matching degree of acoustic features is combined to obtain the teaching style evaluation vector.

5. The emotion adaptive evaluation method based on dual-stream emotional representation of speech and semantics as described in claim 4, characterized in that, After obtaining the teaching style assessment vector, the method further includes: Based on the teaching style assessment vector, an assessment threshold is set; An adjustment suggestion is triggered when any dimension of teaching style, scene style, or acoustic feature conformity falls below the evaluation threshold.

6. An emotion adaptive evaluation device based on dual-stream emotional representation of speech and semantics, characterized in that, include: The extraction unit is used to extract and preprocess the Mel frequency cepstral coefficient features and Mel spectrogram features of the input speech signal; The transformation unit is used to transform teaching scenario information generated with the assistance of a large language model into scenario features through a preset scenario description template. The fusion unit is used to adaptively fuse the Mel frequency cepstral coefficient features and Mel spectrogram features that have received scene feature injection respectively to generate joint emotion features; The first evaluation unit is used to map the identified joint emotional features to the corresponding teaching style through a first mapping mechanism and a second mapping mechanism, and to evaluate the degree of matching between the teaching style and the current teaching scenario. The first mapping mechanism is an emotion-style mapping mechanism and the second mapping mechanism is a style-scenario mapping mechanism. The conversion unit includes: The acquisition module is used to convert scene description text into a sequence of words using a word segmenter and to acquire initial features through an encoder. The pooling module is used to perform feature projection and pooling operations on Mel frequency cepstral coefficients and on Mel spectrogram branches. The stitching module is used to expand scene features into temporal features and stitch them with Mel frequency cepstral coefficient features in the time dimension, and to expand scene features into spatial features and stitch them with image patch features of Mel spectrogram branches.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 5.

8. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 5 through the computer program.

Citation Information

Patent Citations

  • Learner online interest detection method fusing self-attention mechanism

    CN117197856A

  • Speech emotion recognition system based on deep learning

    CN118430587A