Speech emotion recognition method and system based on multi-scale adaptive feature fusion
Through the multi-scale adaptive feature fusion method, combined with global information estimation and efficient self-attention branch, the problem of inaccurate feature fusion in traditional methods is solved, and more efficient speech emotion recognition is achieved.
Patent Information
- Application Number
- CN202510426048.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional speech emotion recognition methods are difficult to dynamically adjust the importance of features in feature fusion, and the deep correlation between features is not fully explored, resulting in unsatisfactory recognition accuracy.
The multi-scale adaptive feature fusion method is adopted, and the feature fusion process is optimized through the adaptive mechanism, combined with global information estimation branches and efficient self-attention branches, extract multi-scale features and explore their correlations, and use a fully connected network for classification.
It improves the accuracy of speech emotion recognition, can capture emotional information in speech more accurately, and enhances the classification ability of the model.
Smart Images

Figure CN120356487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech signal processing, and specifically to a speech emotion recognition method and system based on multi-scale adaptive feature fusion. Background Art
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.
[0003] Speech Emotion Recognition (SER) is a technology that automatically identifies the emotional state of a speaker by analyzing acoustic features (such as pitch, rhythm, intensity, spectrum, etc.) in the speech signal. It combines knowledge in fields such as signal processing, machine learning, and psychology, and has broad application prospects in intelligent human-computer interaction, social robots, customer service and call centers, medical health monitoring, and smart homes, and has received high attention from the industrial and academic communities in recent years.
[0004] Generally, SER is modeled as a classification task related to speech signals, and the goal is to classify speech samples into specific discrete emotion categories, such as sadness, anger, or happiness, etc. This task usually includes two key steps: first, extracting the features of the speech samples, and then classifying the samples into the corresponding emotion categories based on these features.
[0005] The speech signal contains various information, such as time-frequency features, pitch features, semantic features, and high-dimensional depth features, and it is often difficult for a single feature to comprehensively describe the emotional state. Therefore, through feature fusion, effectively integrating features at different levels and different modalities can enhance the model's representation ability for complex emotion patterns. For example, combining low-level acoustic features with high-level features extracted by deep learning helps to capture subtle changes in emotional expressions while reducing noise interference.
[0006] Although traditional feature fusion methods have achieved certain results in Speech Emotion Recognition (SER), these methods usually use fixed weights or simple splicing methods for feature integration, and it is difficult to dynamically adjust the importance of features according to the characteristics of different speech samples, resulting in some redundant or irrelevant features affecting the model performance. At the same time, the deep correlation between features is often not fully exploited. Especially in decision-level fusion, the results of different classifiers are usually calculated independently, lacking effective information interaction. Therefore, in the Speech Emotion Recognition (SER) task, the accuracy of the results obtained by using traditional feature fusion methods is still not ideal. Summary of the Invention
[0007] To solve the technical problems existing in the above background art, the present invention provides a speech emotion recognition method and system based on multi-scale adaptive feature fusion, comprehensively utilizes multi-level information, optimizes the feature fusion process through an adaptive mechanism, enables the model to more accurately capture the emotion information in speech, and improves the accuracy of speech emotion recognition.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] The first aspect of the present invention provides a speech emotion recognition method based on multi-scale adaptive feature fusion, including the following steps:
[0010] Obtain a speech signal and preprocess it to obtain the Mel-frequency cepstral coefficients corresponding to the speech signal;
[0011] Extract the time feature and frequency feature of the Mel-frequency cepstral coefficients, fuse the information of the two along the frequency dimension, and use a multi-scale block to fuse convolution blocks of different sizes along the channel dimension to obtain multi-scale features;
[0012] The extracted multi-scale features are expanded in channels and divided into a global information estimation branch and an efficient self-attention branch. The efficient self-attention branch extracts fine-grained local features through self-attention, and the global information estimation branch uses downsampling to extract low-frequency content and captures non-local information by combining global variance modulation; by fusing the global features and local features obtained from the two branches, the correlation between different features is mined and a deep feature representation is obtained to obtain the fused deep time-frequency features;
[0013] The fused deep time-frequency features are classified using a fully connected network to determine the emotion category corresponding to the speech signal.
[0014] Further, the preprocessing includes the following steps:
[0015] Obtain the speech signal to be recognized, perform frame segmentation and remove the leading and trailing silences;
[0016] Perform windowing on each speech segment and obtain the power spectrum using the short-time Fourier transform;
[0017] Based on the power spectrum, obtain the logarithmic Mel spectrogram through logarithmic operation, and obtain the Mel-frequency cepstral coefficients through discrete cosine transform.
[0018] Further, extracting the time feature and frequency feature of the Mel-frequency cepstral coefficients, fusing the information of the two along the frequency dimension, and using a multi-scale block to fuse convolution blocks of different sizes along the channel dimension to obtain multi-scale features includes the following steps:
[0019] Extract frequency-domain features from the Mel-frequency cepstral coefficients of the voice signal lamp, while retaining the temporal dependence between frames, to form a time-frequency feature map. Extract features in both the time domain and the frequency domain, and fuse the information of both along the frequency dimension;
[0020] Apply multi-scale blocks, connect convolutional blocks of different sizes along the channel dimension, and combine residual connections to enhance the feature representation ability;
[0021] Apply batch normalization to obtain multi-scale features \(X\in\mathbb{R}\) C×H×W , where \(C\) represents the number of channels, \(H\) represents the frequency dimension, and \(W\) represents the time dimension.
[0022] Furthermore, the extracted multi-scale features are extended in channels and divided into a global information estimation branch and an efficient self-attention branch. Specifically:
[0023] The input multi-scale feature map \(X\in\mathbb{R}\) C×H×W , use a \(1\times1\) convolution to expand the channels, and divide the channels into a global information estimation branch and an efficient self-attention branch, where \(C\) represents the number of channels, \(H\) represents the frequency dimension, and \(W\) represents the time dimension.
[0024] Furthermore, the global information estimation branch uses downsampling to extract low-frequency content and combines global variance modulation to capture non-local information. Specifically:
[0025] Extract and downsample the low-frequency features, and perform \(3\times3\) depth convolution to obtain non-local feature representations;
[0026] Measure the statistical divergence by calculating the variance of the features;
[0027] Fuse the variance with the non-local feature representation through addition and \(1\times1\) convolution to adaptively extract representative global features \(A\) t .
[0028] Furthermore, the efficient self-attention branch extracts fine-grained local features through self-attention. Specifically:
[0029] Divide the input feature \(B\) into local detail estimates \(B\) l and the original feature \(B\) h along the channel dimension;
[0030] Generate self-attention queries \(Q\), keys \(K\), and values \(V\) through convolution and splitting operations, and flatten the feature dimensions of \(Q\), \(K\), and \(V\);
[0031] Obtain local detail estimates \(B\) i through scaled dot-product attention;
[0032] Reshape the local detail estimates \(B\) i to \(B\) n using the shape of the input features;
[0033] Connect the local detail estimate B n with the original feature B h to generate an enhanced local feature B d .
[0034] Furthermore, add the global feature A l and the local feature B d by element-wise addition, and perform 1×1 convolution to obtain the fused deep time-frequency feature.
[0035] The second aspect of the present invention provides a system for implementing the above method, including:
[0036] A preprocessing module, configured to: obtain a voice signal and preprocess it to obtain the Mel-frequency cepstral coefficients corresponding to the voice signal;
[0037] A multi-scale feature extraction module, configured to: extract the time feature and frequency feature of the Mel-frequency cepstral coefficients, fuse the information of the two along the frequency dimension, and use a multi-scale block to fuse convolution blocks of different sizes along the channel dimension to obtain multi-scale features;
[0038] An adaptive feature fusion module, configured to: the extracted multi-scale features are expanded through channels and divided into a global information estimation branch and an efficient self-attention branch. The efficient self-attention branch extracts fine-grained local features through self-attention, and the global information estimation branch uses downsampling to extract low-frequency content and captures non-local information by combining global variance modulation; by fusing the global features and local features obtained from the two branches, the correlation between different features is mined and a deep feature representation is obtained to obtain the fused deep time-frequency feature;
[0039] An emotion classification module, configured to: classify the fused deep time-frequency feature using a fully connected network to determine the emotion category corresponding to the voice signal.
[0040] The third aspect of the present invention provides a computer-readable storage medium.
[0041] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the above-mentioned voice emotion recognition method based on multi-scale adaptive feature fusion.
[0042] The fourth aspect of the present invention provides a computer device.
[0043] A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the above-mentioned voice emotion recognition method based on multi-scale adaptive feature fusion.
[0044] Compared with the prior art, the above one or more technical solutions have the following beneficial effects:
[0045] In view of the limitations of traditional methods in feature utilization and fusion strategies, a more refined feature extraction and optimization mechanism is constructed. This method is based on dual-scale time-frequency modeling and can simultaneously capture the short-term dynamic changes and spectral distributions of speech signals, making the expression of emotional information more comprehensive and accurate. And the adaptive feature fusion module accurately explores the correlations between different features, thereby enhancing the accuracy of the model. Finally, the fused deep features are input into the classification network to efficiently and accurately identify speech emotion categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation to the invention.
[0047] Figure 1 is a schematic diagram of the speech emotion recognition process provided by one or more embodiments of the invention;
[0048] Figure 2 is a schematic diagram of the internal structure of the time-frequency feature extraction module provided by one or more embodiments of the invention;
[0049] Figure 3 is a schematic diagram of the internal structure of the adaptive feature fusion module provided by one or more embodiments of the invention;
[0050] Figure 4 is a schematic diagram of the structure of the speech emotion recognition system based on multi-scale adaptive feature fusion provided by one or more embodiments of the invention;
[0051] Figure 5 is the experimental data of comparing this solution with the baseline model using the IEMOCAP dataset provided by one or more embodiments of the invention;
[0052] Figure 6 is the confusion matrix diagram constructed using the IEMOCAP dataset provided by one or more embodiments of the invention;
[0053] Figure 7 is the confusion matrix diagram constructed using the RAVDESS dataset provided by one or more embodiments of the invention;
[0054] Figure 8 is the visualization scatter plot of the GLAM baseline model corresponding to the experimental data provided by one or more embodiments of the invention;
[0055] Figure 9It is a visualization scatter plot of the MAFF proposed model corresponding to experimental data provided by one or more embodiments of the present invention. Detailed implementation manners
[0056] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0057] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0058] It should be noted that the terms used in the following embodiments are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0059] Term explanation:
[0060] Mel-Frequency Cepstral Coefficients (MFCC), one of the most commonly used feature extraction methods in speech signal processing, is widely used especially in speech recognition and sentiment recognition (SER). It converts the speech signal into a set of coefficients that can reflect acoustic features by mimicking the human ear's perception characteristics of sound.
[0061] The following embodiments provide a speech emotion recognition method and system based on multi-scale adaptive feature fusion, which comprehensively utilizes multi-level information and optimizes the feature fusion process through an adaptive mechanism, enabling the model to more accurately capture the emotion information in speech and improving the accuracy of speech emotion recognition.
[0062] Embodiment 1:
[0063] As Figure 1 shown, the speech emotion recognition method based on multi-scale adaptive feature fusion in this embodiment includes the following steps:
[0064] Obtain the speech signal to be recognized and preprocess it, and extract Mel-Frequency Cepstral Coefficients (MFCC) to characterize the key acoustic features of the speech;
[0065] Extract the time and frequency features of MFCC to capture the dynamic changes and spectral information of the speech, and model the two-scale time-frequency features in the time domain and frequency domain.
[0066] Input the extracted time-frequency features into the adaptive feature fusion module to fully explore the correlation between different features and obtain deep feature representations;
[0067] Feed the fused deep time-frequency features into a fully connected network for classification to determine the emotion category corresponding to the speech signal.
[0068] As a further implementation, obtain the speech signal to be recognized and preprocess it, and extract Mel Frequency Cepstral Coefficients (MFCCs) to characterize the key acoustic features of the speech. The specific steps include:
[0069] After obtaining the speech signal to be recognized, first perform frame segmentation on it and remove the head and tail silences. The speech signal is divided into 2-second segments with an overlap of 1.6 seconds to enhance the ability to capture context information.
[0070] Window each segment and perform Short-Time Fourier Transform (STFT) to obtain the power spectrum; then, input the power spectrum into the Mel filter bank, perform logarithmic operation to obtain the log Mel spectrogram, and further apply Discrete Cosine Transform (DCT) to extract Mel Frequency Cepstral Coefficients (MFCCs) to provide an efficient feature representation for subsequent emotion recognition.
[0071] As a further implementation, extract the time and frequency features of MFCCs to capture the short-time dynamic changes and spectral information of the speech, and model the two-scale time-frequency features in the time domain and frequency domain. The specific steps include:
[0072] Extract frequency domain features from the MFCCs of the original speech while retaining the time dependence between frames to form a time-frequency feature map. Use 1×3 and 3×1 parallel convolutional filters to extract features in the time domain and frequency domain respectively, and fuse the information of the two along the frequency dimension.
[0073] Apply multi-scale blocks, perform convolution using 1×1 and 3×3 filters, and connect them along the channel dimension. At the same time, combine residual connections to enhance the feature representation ability and use normalization methods to improve the training stability.
[0074] Apply 5×5 convolution and batch normalization to further strengthen the time-frequency feature representation.
[0075] As a further implementation, input the extracted time-frequency features into the adaptive feature fusion module to fully explore the correlation between different features and obtain deep feature representations. The specific steps include:
[0076] For a given input feature map, first expand the channels with 1×1 convolution, and then divide the channels into two parts: the GIE (Global Information Estimation) branch and the ESA (Efficient Self-Attention) branch.
[0077] Collaboratively utilize the mutual coupling of local and non-local features. The Efficient Self-Attention (ESA) branch focuses on extracting fine-grained local features through self-attention; the Global Information Estimation (GIE) branch uses downsampling to extract low-frequency content and captures non-local information by combining global variance modulation.
[0078] Fuse the global features and local features through element-wise addition to form the output of the adaptive feature fusion module.
[0079] As a further implementation, feed the fused deep time-frequency features into a fully-connected network (FC) for classification to determine the emotion category corresponding to the speech signal. The specific steps include:
[0080] The fully-connected network (FC) layer maps the features to the emotion category space through linear transformation to achieve emotion classification.
[0081] In this embodiment, the speech emotion recognition method based on multi-scale adaptive feature fusion includes the following steps:
[0082] S1. Obtain the speech signal to be recognized and preprocess it, and extract Mel Frequency Cepstral Coefficients (MFCCs) to characterize the key acoustic features of the speech.
[0083] S2. Extract the time and frequency features of the MFCCs to capture the short-time dynamic changes and spectral information of the speech, and model the two-scale time-frequency features in the time domain and frequency domain.
[0084] S3. Input the extracted time-frequency features into the adaptive feature fusion module to fully explore the correlation between different features and obtain deep feature representations.
[0085] S4. Feed the fused deep time-frequency features into a fully-connected network for classification to determine the emotion category corresponding to the speech signal.
[0086] In step S1, obtain the speech signal to be recognized and preprocess it, and extract Mel Frequency Cepstral Coefficients (MFCCs) to characterize the key acoustic features of the speech. Specifically, it includes:
[0087] Step S11. Obtain the speech to be recognized, frame the speech to be recognized, remove the leading and trailing silences, and unify the length of the speech to obtain the speech signal;
[0088] Preprocess the.wav format speech files in the dataset. Using the initial sampling frequency of the speech, 16000 Hz, perform noise reduction, removal of leading and trailing silences, etc. on the speech, and then unify the length of each speech to 2 seconds. For speeches longer than 2 seconds, perform cropping, and for speeches shorter than 2 seconds, perform padding. In addition, adopt an overlapping method of 1.6 seconds to capture more context information.
[0089] Step S12: First, window the speech signal to reduce discontinuities between frames and maintain a smooth transition of the signal. The Hamming Window is usually adopted, which can reduce spectral leakage caused by signal truncation and improve the accuracy of spectral estimation.
[0090] Then, perform the Fast Fourier Transform (FFT) on each windowed speech signal frame to convert the time-domain signal into a frequency-domain signal, so as to obtain the energy distribution of different frequency components. Perform modulus square operation on the spectrum obtained by FFT calculation to get the Power Spectral Density (PSD), which is used to measure the energy distribution of the signal at different frequencies, thereby extracting key spectral information.
[0091] Next, to better conform to the auditory characteristics of the human ear, input the power spectrum into the Mel Filter Bank for processing. The Mel Filter Bank consists of a series of triangular filters that divide the frequency axis according to the Mel Scale, making the filters in the low-frequency part more densely distributed and fewer in the high-frequency part. This transformation simulates the characteristics that the human ear is more sensitive to low-frequency signals and has lower resolution for high-frequency signals.
[0092] After being processed by the Mel Filter Bank, perform logarithmic transformation on the energy output by the filter to obtain the Log-Mel Spectrogram. The role of logarithmic transformation is to compress the dynamic range, make low-energy features more prominent, improve the robustness of features, and reduce the influence of noise.
[0093] Finally, apply the Discrete Cosine Transform (DCT) to the logarithmic Mel coefficients to remove the correlation between features and extract the main spectral features. DCT maps the time-frequency features to the cepstral domain, effectively reducing the dimension of the data while retaining the most representative feature components.
[0094] Ultimately, extract the Mel-Frequency Cepstral Coefficients (MFCC) for subsequent speech emotion recognition tasks.
[0095] In step S2, extract the time and frequency features of MFCC to capture the short-term dynamic changes and spectral information of the speech, and model the two-scale time-frequency features in the time domain and frequency domain. The specific extraction steps are as follows:
[0096] Step S21: Two sets of parallel convolutional filters with kernel sizes of 1×3 and 3×1 are used to extract depth features in the time domain and frequency domain respectively, and the two output feature maps are concatenated along the frequency dimension to integrate time and frequency information.
[0097] Step S22: In subsequent layers, a multi-scale block is applied, which performs convolution using 1×1 and 3×3 filters and concatenates along the channel dimension. At the same time, residual connections are combined to enhance the feature representation ability, and a normalization method is used to improve the training stability. The specific process is as Figure 2 shown.
[0098] Step S23: Finally, a convolutional layer with a larger kernel size of 5×5 and batch normalization are used to obtain multi-scale features X ∈ R C ×H×W , where C represents the number of channels, H represents the frequency dimension, and W represents the time dimension.
[0099] In step S3, the extracted time-frequency features are input into the adaptive feature fusion module to fully explore the correlation between different features and obtain deep feature representations, including:
[0100] Step S31: To generate the deep representation of emotional speech, an adaptive feature fusion network is designed, as Figure 3 shown, which uses the global information estimation (GIE) branch to explore non-local information and the efficient self-attention (ESA) branch to capture local information.
[0101] Given the input feature map X ∈ R C×H×W , first, the channels are expanded using a 1×1 convolution, and then the channels are divided into two parts:
[0102] A, B = Split(Conv 1×1 (X), [C, C]) ∈ R C×H×W
[0103] Step S32: The global information estimation branch (GIE) is used to aggregate and estimate global information from multi-scale features. A method of capturing global features using variance is adopted. First, the low-frequency features are downsampled and extracted, and then 3×3 depth convolution is performed to obtain the non-local feature representation:
[0104]
[0105] In the formula, P(·) represents adaptive average pooling, and the scaling factor is 4.
[0106] Step S33: Next, the variance of the features is calculated to measure the statistical divergence:
[0107]
[0108] In the formula, σ2 (A) represents the variance of feature A, which is obtained by calculating the squared deviation of all eigenvalues from the mean μ in each channel.
[0109] Step S34: Then, the variance is fused with through addition and 1×1 convolution to adaptively extract the representative global feature A t :
[0110]
[0111] In the formula, S(·) represents the Swish activation function, U(·) represents the upsampling operation, and ⊙ represents the element-wise multiplication.
[0112] Step S35: Considering that the Global Information Estimation (GIE) module mainly focuses on capturing non-local structural information, an Efficient Self-Attention (ESA) mechanism is introduced to more comprehensively extract local detailed features. First, the input feature B is divided into B l and B h along the channel dimension:
[0113] B l ,B h = Split(B, [C1, C - C1])
[0114] where C1 controls the number of channels in each part.
[0115] Step S36: Next, self-attention queries (Q), keys (K), and values (V) are generated through convolution and splitting operations:
[0116] Q, K, V = Split(Conv2d(B l ), [d, d, C1])
[0117] Then they are split into Q ∈ R d×H×W and K ∈ R d×H×W representing tensors of queries and keys. representing the tensor of values.
[0118] Step S37: For ease of calculation, the feature dimensions of Q, K, and V are flattened, i.e., Q, K ∈ R d×HW , The local detail estimation B i is obtained using scaled dot-product attention:
[0119]
[0120] Step S38: Then, B i is reshaped to using the shape of the input feature, and the local detail estimation Bn Connect with the original feature B h to generate the enhanced local feature B d :
[0121] B d = Proj(Concat(B h + B n ))
[0122] Step S39. Finally, add the global feature A l and the local feature B d by element-wise addition, and perform 1×1 convolution to form the output of the adaptive feature fusion network:
[0123] Y = Conv 1×1 (A t + B d ) ∈ R C×H×W
[0124] This design balances the collaborative modeling capabilities of non-local and local features. Specifically, the adaptive feature fusion network realizes the collaborative modeling of global and local features through the Global Information Estimation (GIE) and Efficient Self-Attention (ESA) branches. The GIE branch extracts global information through downsampling, depth convolution, and variance calculation, enhancing the modeling ability for context and long-range dependencies; the ESA branch uses the self-attention mechanism to capture local details, enhancing the perception ability for short-term emotional fluctuations. Finally, the global and local features are fused by element-wise addition, enabling the model to not only understand the overall emotional trend but also accurately depict local emotional changes.
[0125] In step S4, the fused deep time-frequency features are fed into a fully connected network for classification to determine the emotional category corresponding to the speech signal. The specific steps include:
[0126] The fused deep time-frequency features are first flattened for input into the fully connected (FC) layer for further feature mapping. Subsequently, through the non-linear transformation of the multi-layer fully connected network, the features become more compact and the discrimination ability for different emotional categories is enhanced.
[0127] The probability distribution of each emotional category is calculated through the Softmax layer, and the category with the highest probability is selected as the final emotional prediction result, thus completing the speech emotion recognition task.
[0128] This embodiment presents a speech emotion recognition method based on multi-scale adaptive feature fusion. Aiming at the limitations of traditional methods in feature utilization and fusion strategies, a more refined feature extraction and optimization mechanism is constructed. This method is based on dual-scale time-frequency modeling, which can simultaneously capture the short-term dynamic changes and spectral distributions of speech signals, making the expression of emotion information more comprehensive and accurate. And through the adaptive feature fusion module, the correlation between different features is accurately mined, thereby enhancing the accuracy of the model. Finally, the fused deep features are input into the classification network to efficiently and accurately identify speech emotion categories.
[0129] The adaptive feature fusion module accurately mines the correlation between different features and synergistically utilizes the mutual coupling of local and non-local features. The efficient self-attention (ESA) branch focuses on extracting fine-grained local features through self-attention; the global information estimation (GIE) branch uses downsampling to extract low-frequency content and captures non-local information by combining global variance modulation. This complementary modeling enables the adaptive feature fusion module to achieve more accurate emotion recognition.
[0130] Experimental data:
[0131] Compare with mature models in the industry, such as Figure 5 As shown. Among them, on the IEMOCAP dataset and RAVDESS dataset, the common WA (weighted average recall rate) is used: to measure the accuracy of the model on the overall dataset, that is, the correct classification ratio considering class imbalance. UA (unweighted average recall rate): to measure the average performance of the model on each class, unaffected by class imbalance. F1 score: to balance accuracy and recall, suitable for multi-class classification tasks with class imbalance.
[0132] Make a confusion matrix on the IEMOCAP and RAVDESS datasets, as Figures 6 - 7 shown. The confusion matrix analyzes the classification effects of different emotion categories, where the rows represent the true categories and the columns represent the categories predicted by the model. The values on the diagonal represent the number of samples correctly classified.
[0133] Make a visualization scatter plot with the baseline model on the IEMOCAP dataset, as Figures 8 - 9 shown. Intuitively display the distribution of emotion categories and compare the classification effects of different models. The clearer the distribution of category boundaries, the better.
[0134] Embodiment 2:
[0135] A system for implementing the above method includes:
[0136] A dataset construction module, configured to:
[0137] S1. Obtain the speech signal to be recognized and preprocess it, and extract Mel Frequency Cepstral Coefficients (MFCCs) to characterize the key acoustic features of the speech.
[0138] S2. Extract the time and frequency features of the MFCCs to capture the short-term dynamic changes and spectral information of the speech, and model the two-scale time-frequency features in the time domain and frequency domain.
[0139] S3. Input the extracted time-frequency features into the adaptive feature fusion module to fully explore the correlations between different features and obtain deep feature representations.
[0140] S4. Feed the fused deep time-frequency features into a fully connected network for classification to determine the emotional category corresponding to the speech signal.
[0141] Preprocessing module, configured to: obtain the speech signal to be recognized and preprocess it, and extract Mel Frequency Cepstral Coefficients (MFCCs) to characterize the key acoustic features of the speech.
[0142] Multi-scale feature extraction module, configured to: extract the time and frequency features of the MFCCs to capture the short-term dynamic changes and spectral information of the speech, and model the two-scale time-frequency features in the time domain and frequency domain.
[0143] Adaptive feature fusion module, configured to: input the extracted time-frequency features into the adaptive feature fusion module to fully explore the correlations between different features and obtain deep feature representations.
[0144] Emotion classification module, configured to: feed the fused deep time-frequency features into a fully connected network for classification to determine the emotional category corresponding to the speech signal.
[0145] Embodiment III:
[0146] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the speech emotion recognition method based on multi-scale adaptive feature fusion as described in Embodiment I above.
[0147] Embodiment IV:
[0148] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the speech emotion recognition method based on multi-scale adaptive feature fusion as described in Embodiment I above.
[0149] The steps or modules involved in the second to fourth embodiments above correspond to those in the first embodiment. For specific implementation manners, reference may be made to the relevant description part of the first embodiment. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0150] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for speech emotion recognition based on multi-scale adaptive feature fusion, characterized in that: Including the following steps: Obtain a speech signal and perform preprocessing to obtain the Mel-frequency cepstral coefficients corresponding to the speech signal; Extract the temporal features and frequency features of the Mel-frequency cepstral coefficients, fuse the information of the two along the frequency dimension, and use multi-scale blocks to fuse convolutional blocks of different sizes along the channel dimension to obtain multi-scale features; The extracted multi-scale features are expanded in channels and divided into a global information estimation branch and an efficient self-attention branch. The efficient self-attention branch extracts fine-grained local features through self-attention. The global information estimation branch uses downsampling to extract low-frequency content and captures non-local information by combining global variance modulation; by fusing the global features and local features obtained from the two branches, the correlation between different features is mined and a deep feature representation is obtained to obtain the fused deep time-frequency features; The fused deep time-frequency features are classified using a fully connected network to determine the emotion category corresponding to the speech signal.
2. The method for speech emotion recognition based on multi-scale adaptive feature fusion according to claim 1, characterized in that, The extracted multi-scale features are expanded in channels and divided into a global information estimation branch and an efficient self-attention branch, specifically: The input multi-scale feature map X ∈ R C×H×W , and use 1×1 convolution to expand the channels, dividing the channels into a global information estimation branch and an efficient self-attention branch. Here, C represents the number of channels, H represents the frequency dimension, and W represents the time dimension.
3. The speech emotion recognition method based on multi-scale adaptive feature fusion according to claim 1, wherein The global information estimation branch uses downsampling to extract low-frequency content and captures non-local information by combining global variance modulation, specifically: Perform downsampling extraction on the low-frequency features and perform 3×3 depth convolution processing to obtain non-local feature representations; Measure the statistical divergence by calculating the variance of the features; Fuse the variance with the non-local feature representation through addition and 1×1 convolution, and adaptively extract the representative global feature A t .
4. The method for speech emotion recognition based on multi-scale adaptive feature fusion according to claim 1, characterized in that The efficient self-attention branch extracts fine-grained local features through self-attention, specifically: Divide the input feature B into local detail estimates B l and the original feature B h ; Generate the query Q, key K, and value V of self-attention through convolution and splitting operations, and flatten the feature dimensions of Q, K, and V; Obtain the local detail estimate B through scaled dot-product attention i ; Estimate local detail B using the input characteristic shape i Reshape it into B n ; Connect the local detail estimate B n with the original feature B h to generate the enhanced local feature B d .
5. The speech emotion recognition method based on multi-scale adaptive feature fusion according to claim 1, wherein Fuse the global feature A l and the local feature B d by element addition, and perform 1×1 convolution to obtain the fused deep time-frequency feature.
6. The speech emotion recognition method based on multi-scale adaptive feature fusion according to claim 1, wherein Preprocessing includes the following steps: Obtain the speech signal to be recognized, perform frame segmentation and remove the leading and trailing silences; Apply a window to each speech segment and obtain the power spectrum using the short-time Fourier transform; Based on the power spectrum, obtain the logarithmic Mel spectrogram through logarithmic operation and obtain the Mel-frequency cepstral coefficients through discrete cosine transform.
7. The method for speech emotion recognition based on multi-scale adaptive feature fusion according to claim 1, characterized in that, Extract the temporal features and frequency features of the Mel-frequency cepstral coefficients, fuse the information of the two along the frequency dimension, and use multi-scale blocks to fuse convolutional blocks of different sizes along the channel dimension to obtain multi-scale features, including the following steps: Extract the frequency-domain features from the Mel-frequency cepstral coefficients of the speech signal, while retaining the temporal dependence between frames, form a time-frequency feature map, extract features in both the time domain and the frequency domain, and fuse the information of the two along the frequency dimension; Apply multi-scale blocks, connect convolutional blocks of different sizes along the channel dimension, and combine residual connections to enhance the feature representation ability; Apply batch normalization to obtain multi-scale features \(X\in\mathbb{R}\) C×H×W , where \(C\) represents the number of channels, \(H\) represents the frequency dimension, and \(W\) represents the time dimension.
8. A speech emotion recognition system based on multi-scale adaptive feature fusion, characterized in that, Including: A preprocessing module configured to: obtain a speech signal and perform preprocessing to obtain the Mel-frequency cepstral coefficients corresponding to the speech signal; A multi-scale feature extraction module configured to: extract the temporal features and frequency features of the Mel-frequency cepstral coefficients, fuse the information of the two along the frequency dimension, and use multi-scale blocks to fuse convolutional blocks of different sizes along the channel dimension to obtain multi-scale features; An adaptive feature fusion module is configured to: After the extracted multi-scale features are expanded in channels, they are divided into a global information estimation branch and an efficient self-attention branch. The efficient self-attention branch extracts fine-grained local features through self-attention, and the global information estimation branch uses downsampling to extract low-frequency content and captures non-local information by combining global variance modulation; By fusing the global features and local features obtained from the two branches, the correlation between different features is mined and a deep feature representation is obtained, and the fused deep time-frequency features are obtained; An emotion classification module is configured to: The fused deep time-frequency features are classified using a fully connected network to determine the emotion category corresponding to the speech signal.
9. A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the speech emotion recognition method based on multi-scale adaptive feature fusion described in any one of claims 1-7 above are implemented.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps in the speech emotion recognition method based on multi-scale adaptive feature fusion described in any one of claims 1-7 are implemented.
Citation Information
Cited By
English pronunciation correction method using AI speech recognition
CN120783751A
English pronunciation correction method using AI voice recognition
CN120783751B
Immersive traditional cultural language audio feature extraction method based on artificial intelligence
CN120895025A
Optical fiber distributed voiceprint recognition method based on multi-scale decomposition and hybrid recombination
CN121938375A