Voice emotion recognition method based on deep residual network and frequency domain attention mechanism

Through the speech emotion recognition method of deep residual network and frequency domain attention mechanism, the logarithmic Meier spectrogram feature extraction and frequency domain attention module processing are used, combined with the label smooth cross entropy loss function optimization, the problem of low accuracy of speech emotion recognition in existing models is solved, and efficient emotion classification and model simplification is achieved.

CN120581041APending Publication Date: 2025-09-02TIANHE COLLEGE GUANGDONG POLYTECHNIC NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510618474.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing voice emotion recognition model cannot effectively classify voice emotions, resulting in low recognition accuracy.

Method used

The deep residual network and frequency domain attention mechanism are adopted to optimize the generalization capability of the model through logarithmic Meier spectrogram feature extraction, frequency domain attention module processing and label smooth cross-entropy loss function optimization, combined with data enhancement and lightweight classifier.

Benefits of technology

It significantly improves the accuracy of speech emotion recognition, reduces the complexity of the model, and adapts to emotion classification in small sample scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120581041A_ABST
    Figure CN120581041A_ABST
Patent Text Reader

Abstract

The invention provides a voice emotion recognition method based on a deep residual network and a frequency domain attention mechanism, and the method comprises the steps: inputting an original voice signal, and converting the original voice signal into a logarithmic Mel spectrogram; extracting features of the logarithmic Mel spectrogram, retaining a convolutional layer of the logarithmic Mel spectrogram, removing an original classification layer, and outputting a multi-channel feature map; processing the multi-channel feature map through a frequency domain attention module to obtain a modulated feature map; and based on the modulated feature map and the emotion category label, performing model optimization by using a label smooth cross entropy loss function and regularization constraint to obtain an emotion category classification result. Different frequency band weights are dynamically distributed through FAM; and then, in combination with a label smooth cross entropy loss function and regularization constraint, overfitting in a small sample scene is inhibited, the generalization ability of the model is optimized, the emotion classification precision is remarkably improved, and the accuracy of voice emotion recognition by the model is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sound analysis technology, and in particular to a speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism. Background Art

[0002] Emotion, as a fundamental paralinguistic signal, often conveys the speaker's intentions and psychological state. Accurately identifying emotions in speech can help voice interaction systems better understand the speaker's needs and respond appropriately, effectively improving the user experience. Speech emotion recognition technology, by analyzing the emotional characteristics of speech signals, is widely used in fields such as human-computer interaction and mental health monitoring. The primary task of speech emotion recognition is to extract the emotional information contained in speech and identify its category. Traditional methods rely on manually extracted acoustic features (such as fundamental frequency and Mel-frequency cepstral coefficients) combined with shallow classification models. However, these methods have limited feature expression capabilities and struggle to capture emotional relevance in complex scenarios.

[0003] However, existing models are unable to classify emotions, resulting in low accuracy in voice emotion recognition. Summary of the Invention

[0004] The present invention provides a speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism, which is used to solve the defect in the prior art that emotions cannot be classified, resulting in a low accuracy rate of the model in speech emotion recognition.

[0005] The present invention provides a speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism, comprising:

[0006] Input the original speech signal and convert it into a logarithmic Mel spectrum.

[0007] Extract the features of the logarithmic Mel-spectrogram, retain its convolutional layer and remove the original classification layer, and output a multi-channel feature map;

[0008] The multi-channel feature map is processed by the frequency domain attention module to obtain the modulated feature map;

[0009] Based on the modulated feature map and emotion category labels, the label smoothed cross entropy loss function and regularization constraints are used to optimize the model and obtain the emotion category classification results.

[0010] According to a speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism provided by the present invention, the input of the original speech signal is converted into a logarithmic Mel-spectrogram, comprising:

[0011] The original speech signal is converted into a time spectrum through short-time Fourier transform, and a logarithmic Mel spectrum graph is generated through Mel filter bank and nonlinear compression;

[0012] Based on the logarithmic Mel-spectrogram, data enhancement is performed using random cropping, horizontal flipping, color jittering, and random rotation strategies.

[0013] According to a speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism provided by the present invention, the method extracts features of a logarithmic Mel-spectrogram, retains its convolutional layer and removes the original classification layer, including:

[0014] Input the logarithmic Mel-spectrogram, use the ResNet50 model to remove its original classification layer, retain the convolutional layer as the feature extractor, and output a multi-channel feature map.

[0015] According to a speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism provided by the present invention, the multi-channel feature map is processed by the frequency domain attention module to obtain a modulated feature map, including:

[0016] Perform global average pooling on the multi-channel feature map to generate a frequency band energy vector;

[0017] The frequency band energy vector is reduced and restored through the fully connected layer to generate the frequency domain weight vector;

[0018] Multiply the frequency domain weight vector and the multi-channel feature map channel by channel and output the modulated feature map.

[0019] According to a speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism provided by the present invention, the multi-channel feature map is processed by the frequency domain attention module to obtain a modulated feature map, including:

[0020] The modulated feature map is input into the lightweight classifier to output the probability distribution of emotion categories;

[0021] During training, all parameters in the ResNet50 model except Layer 4 and the frequency domain attention module are frozen, and only the weights of Layer 4 and the frequency domain attention module are updated.

[0022] According to a speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism provided by the present invention, the freezing of all parameters in the ResNet50 model except for Layer 4 and the frequency domain attention module includes:

[0023] The pre-trained weights of the shallow network of ResNet50 are retained, its parameters are frozen, and only the parameters of the deep network and the frequency domain attention module are updated;

[0024] Control the total amount of trainable parameters to meet the requirements of small sample learning;

[0025] By freezing shallow parameters, the back-propagation chain gradient calculation is simplified.

[0026] According to the present invention, a speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism is provided. The method uses a label smoothed cross entropy loss function and a regularization constraint to optimize the model based on the modulated feature map and the emotion category label to obtain the emotion category classification result, including:

[0027] The total loss function is:

[0028] L=L CE +λR L2

[0029] L CE is the label smoothed cross entropy loss function, λ is the balance hyperparameter, R L2 is the regularization term;

[0030] The AdamW optimizer is used to update the trainable parameters, and the cosine annealing strategy is used to dynamically adjust the learning rate;

[0031] The label smoothed cross entropy loss function is:

[0032]

[0033] Among them, ∈ = 0.1 is the smoothing coefficient, C = 4 is the number of categories;

[0034] The cosine annealing strategy formula is:

[0035]

[0036] Initial learning rate η max , the minimum learning rate η min , period T max .

[0037] According to a speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism provided by the present invention, a logarithmic Mel-spectrogram is input and normalized and uniformly scaled;

[0038] Calculate the classification accuracy based on the logarithmic Mel-spectrogram to evaluate the generalization performance of the model;

[0039] Use TensorRT tools to quantize the trained model;

[0040] Merge the computational steps of the convolutional layer and the batch normalization layer to generate a fused lightweight inference model;

[0041] Output optimized deployment model for real-time speech emotion classification.

[0042] The present invention further provides an electronic device, comprising:

[0043] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, the steps of the speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism as described above are implemented.

[0044] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described methods for speech emotion recognition based on a deep residual network and a frequency domain attention mechanism.

[0045] The speech emotion recognition method based on deep residual network and frequency domain attention mechanism provided by the present invention converts the original speech signal into a logarithmic Mel-spectrogram, uses pre-trained ResNet to extract deep frequency domain features, and dynamically assigns weights of different frequency bands through FAM; then combines the label smoothed cross entropy loss function with regularization constraints to suppress overfitting in small sample scenarios, optimize the model's generalization ability, significantly improve the emotion classification accuracy, and thus improve the model's accuracy in speech emotion recognition. At the same time, by retaining its convolutional layer and removing the original classification layer, the model complexity is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0047] Figure 1 This is a flow chart of the speech emotion recognition method based on deep residual network and frequency domain attention mechanism provided by the present invention. DETAILED DESCRIPTION

[0048] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0049] like Figure 1As shown, the speech emotion recognition method based on deep residual network and frequency domain attention mechanism of the present invention includes:

[0050] Input the original speech signal and convert it into a logarithmic Mel spectrum.

[0051] Extract the features of the logarithmic Mel-spectrogram, retain its convolutional layer and remove the original classification layer, and output a multi-channel feature map;

[0052] The multi-channel feature map is processed by the frequency domain attention module to obtain the modulated feature map;

[0053] Based on the modulated feature map and emotion category labels, the label smoothed cross entropy loss function and regularization constraints are used to optimize the model and obtain the emotion category classification results.

[0054] The original speech signal is converted into a logarithmic Mel-spectrogram, deep frequency domain features are extracted using pre-trained ResNet, and weights of different frequency bands are dynamically assigned through FAM. Subsequently, the label smoothed cross entropy loss function and regularization constraints are combined to suppress overfitting in small sample scenarios, optimize the model's generalization ability, and significantly improve the accuracy of emotion classification, thereby improving the model's accuracy in speech emotion recognition. At the same time, by retaining its convolutional layer and removing the original classification layer, the model complexity is reduced.

[0055] The input original speech signal is converted into a logarithmic Mel-spectrogram, including:

[0056] The original speech signal is converted into a time spectrum by short-time Fourier transform (STFT), and spectrum leakage is reduced by frame division and windowing.

[0057] First input the original speech signal and convert it into a time spectrum through short-time Fourier transform (STFT):

[0058]

[0059] Where x(τ) is the original speech signal, w(t-τ) is the Hamming window function, which reduces the spectrum leakage at the edge of the frame, and X(t,f) is the time-spectrum, which represents the energy distribution at time t and frequency f.

[0060] A logarithmic Mel-spectrogram M(t,m) is generated through a Mel filter bank and nonlinear compression. The logarithmic Mel-spectrogram further simulates the auditory characteristics of the human ear and performs nonlinear compression on the frequency axis.

[0061]

[0062] Among them H m (f) is the weight of the mth Mel filter bank, simulating the sensitivity of the human ear to different frequency bands.

[0063] Based on the logarithmic Mel-spectrogram, data enhancement is performed using random cropping, horizontal flipping, color jittering, and random rotation strategies.

[0064] Random cropping: Speech emotion features are localized in time and frequency. Cropping forces the model to learn local features rather than relying on global statistics.

[0065] From the original spectrogram M∈R T×F Randomly select a subregion (t0, f0, △t, △f) and scale it to a fixed size of 224×224:

[0066]

[0067] Horizontal flip: The short-term stability of speech signals (10-30ms) makes flipping have minimal impact on frequency domain characteristics;

[0068] Flip along the time axis:

[0069]

[0070] Color dithering: Different microphones have different frequency responses (such as low-frequency enhancement / attenuation). By adjusting α m ,β m distribution, covering the gain range of real devices,

[0071]

[0072] Random rotation: Rotation is equivalent to time-frequency distortion, simulating changes in speech speed (such as fast / slow speaking). Experiments show that a rotation angle |θ|>15° will cause confusion in emotion labels;

[0073] Perform a small rotation on the spectrogram, θ∈[-15°,15°]:

[0074]

[0075] Suppose you need to train a speech emotion recognition model to distinguish between "angry" and "calm." The following is an example of how data augmentation strategies can be used:

[0076] 1. Random cropping:

[0077] For example, from a 256×256 mel-spectrogram generated from a 3-second segment of angry speech, a 1-second subregion with a frequency range of 100–4000 Hz is randomly selected, with the upper left corner at t = 0.5 seconds and f = 100 Hz, and resized to 224×224. This cropped spectrogram forces the model to focus on local features, such as high-frequency energy spikes, rather than relying on the overall spectral distribution.

[0078] 2. Horizontal flip:

[0079] The spectrogram of the same angry speech segment was flipped along the time axis, essentially mirroring it. For example, the "anger" feature in the original spectrogram was concentrated in the second half, but after flipping it, it became concentrated in the first half. Because speech is stable within a short 20ms, the flipped spectrum still retains high-frequency energy characteristics, allowing the model to effectively identify emotions.

[0080] 3. Color dithering:

[0081] Simulate the differences in recording the same "angry" voice using different microphones:

[0082] Add 1.2 times gain to the low frequency band of the spectrum to simulate a low frequency enhancement device, such as 0-500Hz;

[0083] Add a -0.05 offset to the mid-high frequency band to simulate background noise interference, for example 2000-4000Hz.

[0084] By randomly adjusting the gain and offset of each frequency band, the model learns to ignore device differences and focus on emotion-related frequency bands.

[0085] 4. Random rotation:

[0086] The spectrogram of calm speech is rotated 10 degrees clockwise, which is equivalent to slightly stretching the time axis, simulating the effect of slower speech. The rotated spectrum maintains "calm" characteristics (such as uniform energy distribution), while the rotation angle is kept within 15 degrees to avoid excessive distortion that could cause the model to misinterpret the speech as "sad."

[0087] Furthermore, the feature extraction of the logarithmic Mel-spectrogram, retaining its convolution layer and removing the original classification layer, includes:

[0088] The logarithmic Mel-spectrogram is input, and the original classification layer is removed using the ResNet50 model. The convolutional layer is retained as the feature extractor, and a multi-channel feature map is output to improve the accuracy of the dataset and the training efficiency.

[0089] Specifically, improve the pre-trained ResNet50 model:

[0090] Remove the original classification layer: Let the input Mel spectrum map X∈R 224×224×3 , the forward propagation of ResNet50 can be expressed as:

[0091] F=ResNet trunc (X),F∈R B×2048×7×7

[0092] Freeze underlying parameters: Only fine-tune Layer 4, the deep feature extraction layer, and keep the pre-trained weights of the shallow layers Conv1 to Layer 3 unchanged.

[0093] The multi-channel feature map is processed by the frequency domain attention module to obtain a modulated feature map. The importance of different frequency components is dynamically adjusted by the frequency domain weight to obtain physiological characteristics consistent with speech emotions, such as high-frequency energy is positively correlated with anger, including:

[0094] Perform global average pooling on the multi-channel feature map to generate a frequency band energy vector;

[0095] Compress the multi-channel feature map F along the spatial dimension (H×W) and retain the frequency band statistics:

[0096] Global Average Pooling (GAP):

[0097]

[0098] Where F is the Fourier transform, indicating that the GAP result reflects the energy of the frequency band, F c,i,j is the value of channel c of the input feature map at the spatial position (i, j).

[0099] The frequency band energy vector is reduced and restored through the fully connected layer to generate the frequency domain weight vector;

[0100] First, reduce the dimension to reduce the amount of calculation:

[0101] Z=W1f,W1∈R 256×2048 ,Z∈R 256

[0102] The purpose is to compress the 2048-dimensional features to 256 dimensions through the matrix W1, and the number of parameters is reduced from O(C 2 ) down to

[0103] Second, non-linear activation (ReLU):

[0104] z′=ReLU(z)

[0105] Next, restore the original dimensions:

[0106] w′=W2z,′W2∈R 2048×256 ,w′∈R 2048 ,

[0107] The goal is to use the matrix W2 to learn how to combine 256-dimensional features to reconstruct channel weights.

[0108] Multiply the frequency domain weight vector and the multi-channel feature map channel by channel and output the modulated feature map.

[0109] Sigmoid normalization:

[0110] Clamp the weights to the interval [0,1]:

[0111]

[0112] If w c ≈1, indicating that the cth frequency band is critical for emotion recognition. c ≈0, indicating that the frequency band can be suppressed, weight w c ∈[0,1], the larger the value, the more important the frequency band is for emotion classification.

[0113] Feature weighting is used to perform channel-level modulation on the original feature map, perform channel-level weighting on the original feature map, suppress irrelevant frequency bands, such as low-frequency noise, and enhance key frequency bands;

[0114] Y c, i , j=F c,i,j w c

[0115] w of the high-frequency channel (e.g. c = 1500) c When it is larger, it enhances emotional characteristics such as "anger";

[0116] w of the low-frequency channel (e.g. c = 300) c When it is smaller, irrelevant background noise is reduced.

[0117] To solve the problem of overfitting in small samples and retain the generalization ability of pre-trained features, lightweight classification and parameter freezing are performed:

[0118] The purpose of the lightweight classifier is to build a classifier with a small number of parameters but sufficient expressiveness, which is suitable for 2048-dimensional input features and 4 types of output.

[0119] Classifier structure:

[0120]

[0121] Where: W1∈R 512×2048 (reduced dimension matrix); W2∈R 4×512 (Classification Matrix)

[0122] Furthermore, the multi-channel feature map is processed by the frequency domain attention module to obtain a modulated feature map, including:

[0123] The modulated feature map is input into the lightweight classifier to output the probability distribution of emotion categories;

[0124] During training, all parameters in the ResNet50 model except Layer 4 and the frequency domain attention module are frozen, and only the weights of Layer 4 and the frequency domain attention module are updated.

[0125] The freezing of all parameters in the ResNet50 model except Layer4 and the frequency domain attention module includes:

[0126] The pre-trained weights of the shallow network of ResNet50 are retained, its parameters are frozen, and only the parameters of the deep network and the frequency domain attention module are updated;

[0127] Control the total amount of trainable parameters to meet the requirements of small sample learning;

[0128] By freezing shallow parameters, the back-propagation chain gradient calculation is simplified.

[0129] Parameter freezing is divided into three steps:

[0130] 1. Extracting hierarchical features of pre-trained features

[0131] Universality of shallow features: The shallow layers of the convolutional network (Conv1 to Layer3) extract universal features.

[0132] Deep feature specificity: Layer 4 learns task-related features, tuned for speech emotion recognition.

[0133] 2. VC dimension constraints for small sample learning

[0134] According to the Vapnik-Chervonenkis theory, the model complexity C(f) must satisfy:

[0135]

[0136] Where N is the number of samples and ∈ is the target error.

[0137] Model complexity C(f) is usually related to factors such as the number of model parameters and structural complexity. In this model, the parameters of the ResNet50 model and the changes in parameters after freezing some layers are combined to illustrate its impact on model complexity.

[0138] After freezing the shallow layer:

[0139] The number of training parameters was reduced from 25.5M to 5.7M;

[0140] satisfy Emotion Recognition Dataset (IEMOCAP) N≈3k.

[0141] 3. Gradient Propagation Analysis

[0142] After freezing the shallow layers, the back propagation chain rule is simplified, thus preventing the shallow gradient from disappearing and reducing the memory usage:

[0143]

[0144] By freezing the shallow fixed feature extractor, the frequency-domain attention weights w learned by the FAM module directly reflect the contribution of different frequency bands to emotion classification. For example, the weight of the high-frequency channel (1500-2000) is significantly higher than that of the low-frequency band (0-500), which is consistent with the conclusion in psychological research that "anger is often manifested as increased high-frequency energy."

[0145] Based on the modulated feature map and emotion category labels, the model is optimized using the label smoothed cross entropy loss function and regularization constraints to obtain the emotion category classification results, and the frequency domain attention module (FAM) and classifier parameters are optimized simultaneously while maintaining the feature extraction stability of the ResNet50 frozen layer, including:

[0146] The total loss function is:

[0147] L=L CE +λR L2

[0148] It is formalized as follows

[0149]

[0150] L CE is the label smoothed cross entropy loss function, λ is the balance hyperparameter, R L2 is the regularization term;

[0151] The AdamW optimizer is used to update the trainable parameters, and the cosine annealing strategy is used to dynamically adjust the learning rate;

[0152] Label smoothed cross entropy loss:

[0153] The hard labels in the original standard cross entropy caused the model to overfit the training label distribution.

[0154]

[0155] Label smoothing improvements: Label smoothing softens hard labels (0 or 1) to prevent the model from being overconfident in training samples and improves robustness to noisy labels. After smoothing, the validation set accuracy increased by 1.2%;

[0156] The true label y i To switch from a hard label (0 or 1) to a soft label:

[0157]

[0158] Therefore, the label smoothed cross entropy loss function after smoothing is:

[0159]

[0160] Among them, ∈ = 0.1 is the smoothing coefficient, balancing the true label and uniform distribution, and C = 4 is the number of categories;

[0161] For example, when the true label of a speech is "happy" and the hard label is (1,0), through label smoothing with a smoothing coefficient ∈=0.1, the label is softened to 0.9, 0.1. The model no longer requires the predicted probability of "happy" to be close to 100%, but allows a small amount of probability to be assigned to other categories, such as "sadness". This allows the model to avoid misjudgment due to overconfidence when facing ambiguous speech, such as laughter with a tearful tone. The test set accuracy rate is improved from 88% to 89.5%.

[0162] Frequency domain attention regularization

[0163] Frequency-domain attention regularization

[0164] The weight w of the FAM module should avoid over-suppression or over-enhancement of certain frequency bands. L2 regularization constrains the weight amplitude:

[0165]

[0166] Regularization is used to prevent the FAM weight w from becoming extreme (such as all 0 or all 1), ensuring balanced participation of each frequency band.

[0167] For example, when processing sad speech, the frequency attention module (FAM) tends to oversuppress low frequencies (0-200Hz), resulting in the neglect of low-pitched tones. L2 regularization is used to constrain the weights, forcing the weights of high and low frequencies to be balanced, with the high frequency band (2000-4000Hz) being the highest. For example, by reducing the high-frequency weight from 0.95 to 0.7 and increasing the low-frequency weight from 0.05 to 0.3, the model can simultaneously capture the low tones and high-frequency vibrato of "sad" speech, improving classification accuracy by 1.8%.

[0168] Optimizer Selection: Adaptability of AdamW

[0169] The standard Adam optimizer may cause weight decay to fail in fine-tuning scenarios. The update rule is improved:

[0170] θ t =θ t-1 -η

[0171] in:

[0172] is the first-order and second-order moment estimation of the gradient; λ is the decoupled weight attenuation coefficient, which defaults to 10 -4

[0173] More effectively controls parameter amplitudes, preventing weight explosion in the FAM module. On the IEMOCAP dataset, AdamW consistently improves final accuracy by 0.8% over Adam.

[0174] In the early stage of training, the model parameters are updated with a large amplitude. AdamW decouples the weight decay and λ=10 -4 This prevents frequency-domain attention weights from exceeding a reasonable range due to gradient explosion, such as a sudden increase from 1.0 to 5.0. For example, without AdamW, weights can get out of control, causing the model to misclassify background noise as "happy." With AdamW, weights stabilize within the range of 0 to 1.2, making classification results more reliable.

[0175] It is difficult to balance the convergence speed and stability in the fine-tuning phase with a fixed learning rate, so it is necessary to apply the cosine annealing learning rate to schedule the learning rate η. t Decay according to the cosine function;

[0176] The cosine annealing strategy formula is:

[0177]

[0178] Initial learning rate η max =10 -4 , the minimum learning rate η min =10 -6 ,

[0179] Period T max =100, a large learning rate is used for rapid convergence in the early stage, and a small learning rate is used for fine tuning in the later stage.

[0180] At the beginning of training, the learning rate is set to 10 -4 , the model quickly learns basic features, such as the sudden increase in high-frequency energy of the "happy" voice; by the 50th round, the learning rate is reduced to 10 -6 The model began to fine-tune the distribution of frequency-domain attention weights. For example, the previously ignored mid-frequency band (500-1500Hz) was gradually recognized as a key feature of "sadness," such as sobbing. The final validation set accuracy increased from 90.1% to 91.3%.

[0181] Preferably, the model is validated to evaluate its generalization performance in real scenarios and avoid data augmentation interference;

[0182] Input the logarithmic Mel spectrum, perform normalization and uniform size scaling;

[0183] Calculate the classification accuracy based on the logarithmic Mel-spectrogram to evaluate the generalization performance of the model;

[0184] Use TensorRT tools to quantize the trained model;

[0185] Merge the computational steps of the convolutional layer and the batch normalization layer to generate a fused lightweight inference model;

[0186] Output optimized deployment model for real-time speech emotion classification.

[0187] Specifically, the validation set is processed:

[0188] Use only normalization and scaling, disable random augmentation, and standard tests:

[0189]

[0190] in Represents function composition.

[0191] Performance index calculation:

[0192] Classification accuracy:

[0193] Model deployment optimization

[0194] (1) TensorRT quantization

[0195] By reducing numerical precision (FP32→FP16 / INT8), the model size and computational complexity are reduced, and the inference speed is improved. The weights and activation values ​​are converted from 32-bit floating point to 16-bit FP32→FP16 quantization:

[0196] W FP16 =cast FP16 (W FP32 ),X FP16 =cast FP16 (X FP32 )

[0197] FP16: represents 16-bit floating point, with controllable precision loss, suitable for most speech emotion recognition.

[0198] INT8: represents an 8-bit integer. It requires calibration to determine the dynamic range and has a slightly greater loss of precision. Therefore, it is necessary to verify whether the emotion classification accuracy meets the requirements.

[0199] This reduces model size by 50% (FP16) or 75% (INT8), and increases inference speed by 2.3 times (FP16) or 3.1 times (INT8).

[0200] (2) Layer Fusion

[0201] Combine consecutive linear operations (such as Conv+BN+ReLU) into a single kernel function:

[0202]

[0203] in

[0204] μ, σ are the mean and variance of BN; γ, β are scaling and offset parameters, which can reduce the number of memory accesses and reduce latency. After combining Conv+BN+ReLU into a single matrix operation, the latency is reduced by 60%.

[0205] After optimization through TensorRT quantization (FP16 / INT8) and layer fusion technology, the model's inference latency on Jetson Xavier was reduced from 22.1ms to 5.6ms, the volume was reduced by 75%, the power consumption was reduced by 35%, the real-time processing speed reached 150FPS, and the accuracy loss was controlled within 1.3% (IEMOCAP test set). Data augmentation was disabled in the verification phase to ensure unbiased evaluation.

[0206] An electronic device may include: a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory communicate with each other via the communications bus. The processor may call logic instructions in the memory to execute a speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism.

[0207] In addition, the logical instructions in the above-mentioned memory can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0208] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech emotion recognition method based on deep residual network and frequency domain attention mechanism provided by the above methods.

[0209] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the speech emotion recognition method based on deep residual network and frequency domain attention mechanism provided by the above methods.

[0210] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0211] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A speech emotion recognition method based on deep residual network and frequency domain attention mechanism, characterized in that: include: Input the original speech signal and convert it into a logarithmic Mel spectrum. Extract the features of the logarithmic Mel-spectrogram, retain its convolutional layer and remove the original classification layer, and output a multi-channel feature map; The multi-channel feature map is processed by the frequency domain attention module to obtain the modulated feature map; Based on the modulated feature map and emotion category labels, the label smoothed cross entropy loss function and regularization constraints are used to optimize the model and obtain the emotion category classification results.

2. The speech emotion recognition method based on deep residual network and frequency domain attention mechanism according to claim 1 is characterized in that The input original speech signal is converted into a logarithmic Mel-spectrogram, including: The original speech signal is converted into a time spectrum through short-time Fourier transform, and a logarithmic Mel spectrum graph is generated through Mel filter bank and nonlinear compression; Based on the logarithmic Mel-spectrogram, data enhancement is performed using random cropping, horizontal flipping, color jittering, and random rotation strategies.

3. The speech emotion recognition method based on deep residual network and frequency domain attention mechanism according to claim 1 is characterized in that The method of extracting features of the logarithmic Mel-spectrogram, retaining its convolutional layer and removing the original classification layer, includes: Input the logarithmic Mel-spectrogram, use the ResNet50 model to remove its original classification layer, retain the convolution layer as the feature extractor, and output a multi-channel feature map.

4. The speech emotion recognition method based on deep residual network and frequency domain attention mechanism according to claim 3 is characterized in that The multi-channel feature map is processed by the frequency domain attention module to obtain a modulated feature map, including: Perform global average pooling on the multi-channel feature map to generate a frequency band energy vector; The frequency band energy vector is reduced and restored through the fully connected layer to generate the frequency domain weight vector; Multiply the frequency domain weight vector and the multi-channel feature map channel by channel and output the modulated feature map.

5. The speech emotion recognition method based on deep residual network and frequency domain attention mechanism according to claim 3 is characterized in that The multi-channel feature map is processed by the frequency domain attention module to obtain a modulated feature map, including: The modulated feature map is input into the lightweight classifier to output the probability distribution of emotion categories; During training, all parameters in the ResNet50 model except Layer 4 and the frequency domain attention module are frozen, and only the weights of Layer 4 and the frequency domain attention module are updated.

6. The speech emotion recognition method based on deep residual network and frequency domain attention mechanism according to claim 5 is characterized in that The freezing of all parameters in the ResNet50 model except Layer4 and the frequency domain attention module includes: The pre-trained weights of the shallow network of ResNet50 are retained, its parameters are frozen, and only the parameters of the deep network and the frequency domain attention module are updated; Control the total amount of trainable parameters to meet the requirements of small sample learning; By freezing shallow parameters, the back-propagation chain gradient calculation is simplified.

7. The speech emotion recognition method based on deep residual network and frequency domain attention mechanism according to claim 1 is characterized in that The model optimization is performed based on the modulated feature map and the emotion category label using the label smoothing cross entropy loss function and the regularization constraint to obtain the emotion category classification result, including: The total loss function is: L=L CE +λR L2 L CE is the label smoothed cross entropy loss function, λ is the balance hyperparameter, R L2 is the regularization term; The AdamW optimizer is used to update the trainable parameters, and the cosine annealing strategy is used to dynamically adjust the learning rate; The label smoothed cross entropy loss function is: Among them, ∈ = 0.1 is the smoothing coefficient, C = 4 is the number of categories; The cosine annealing strategy formula is: Initial learning rate η max , the minimum learning rate η min , period T max .

8. The speech emotion recognition method based on deep residual network and frequency domain attention mechanism according to claim 1 is characterized in that Input the logarithmic Mel spectrum, perform normalization and uniform size scaling; Calculate the classification accuracy based on the logarithmic Mel-spectrogram to evaluate the generalization performance of the model; Use TensorRT tools to quantize the trained model; Merge the computational steps of the convolutional layer and the batch normalization layer to generate a fused lightweight inference model; Output optimized deployment model for real-time speech emotion classification.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the speech emotion recognition method based on deep residual network and frequency domain attention mechanism are implemented as described in any one of claims 1 to 8.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the speech emotion recognition method based on a deep residual network and a frequency domain attention mechanism are implemented as described in any one of claims 1 to 8.