A speech emotion recognition method based on feature fusion

By adopting feature fusion method in speech emotion recognition technology, combining convolutional neural network and classification layer feature fusion algorithm, the problem of single feature selection and classifier model selection in the existing technology is solved, and the accuracy of speech emotion recognition is improved.

CN114495990BActive Publication Date: 2025-05-06ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210217251.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-05-06
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

The existing speech emotion recognition technology is relatively single in feature selection and classifier model selection, resulting in a lack of correlation between the extracted features and affecting the recognition accuracy.

Method used

The speech emotion recognition method based on feature fusion is adopted to extract the deep features of acoustic features through convolutional neural networks, and the deep features of MFCC are fused with the other three acoustic features, combined with the classification layer feature fusion algorithm to improve the recognition accuracy.

Benefits of technology

The feature utilization rate of acoustic features is improved, the range of feature utilization of data is expanded, the classifier recognition probability is effectively utilized, and the speech emotion recognition rate is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495990B_ABST
    Figure CN114495990B_ABST
Patent Text Reader

Abstract

A method for speech emotion recognition based on feature fusion, comprising: step 1) data acquisition and preprocessing; step 2) inputting the speech emotion recognition network based on feature fusion designed in the present invention to perform emotion recognition; step 3) obtaining the emotion recognition result. The present invention uses the classification layer feature fusion method to recognize speech emotions, designs and implements a method for fusing the deep features of MFCC (Mel-frequency cepstral coefficients) with traditional acoustic features, uses the classification layer feature fusion algorithm to fuse the MFCC deep features with the zero-crossing rate, Mel frequency, and spectrum centroid, performs fusion calculation on the output recognition results through the specified decision fusion rules, and finally selects the one with the largest probability in the probability distribution as the recognition result. The invention has great application value for speech emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a speech emotion recognition method based on feature fusion. Background Art

[0002] The process of speech emotion recognition is similar to the image classification process. The traditional recognition process is mainly divided into three aspects. The first is to preprocess the speech data, including normalization, averaging, data enhancement and other operations. The second is the extraction and selection of speech features. Commonly used speech features are: gene frequency, formant, Mel cepstral coefficients, etc. Finally, a suitable classifier model is used to analyze the obtained emotion-related features to identify the emotions contained in the speech. The selection of speech features and the selection of classifier models have the greatest impact on the accuracy of emotion recognition. Commonly used classifiers include support vector machines, Gaussian mixture models, random forests, etc. In recent years, with the development of deep learning, these traditional classification models have gradually been replaced. The selection and extraction methods of traditional speech features are relatively simple, the extracted features lack correlation, and the features cannot contain emotion-related information. Since deep learning has shown strong performance in feature extraction, more and more researchers use deep learning methods to extract speech features. Among them, the method that improves the emotion recognition rate more significantly is to use neural networks to obtain more deep features through operations such as convolution. The main process of this method includes: first, select an acoustic feature, then analyze this feature, and finally obtain the classification result. Although this method improves the recognition rate to a certain extent, due to the large number of features in speech, this method only focuses on the deep features of one acoustic feature but ignores the impact of other acoustic features on the accuracy. Summary of the invention

[0003] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provide a speech emotion recognition method based on feature fusion.

[0004] The present invention solves the technical problem by adopting the following technical solution:

[0005] A speech emotion recognition method based on feature fusion comprises the following steps:

[0006] Step 1: Acquire and preprocess user voice data;

[0007] Step 2: Put the data into the feature extractor to extract features;

[0008] Step 3: Input the speech emotion recognition network for emotion recognition;

[0009] Step 4: Obtain speech emotion recognition results.

[0010] Step 1: Use the Python library function librosa.load() to read the speech that needs emotion recognition, save it as a numpy data type, and preprocess it by constructing a pre-emphasis function and a windowing and framing function.

[0011] Step 2 specifically includes:

[0012] 1) Use the Python library function librosa.feature.zero_crossing_rate() to extract the zero crossing rate of the saved data;

[0013] 2) Use the Python library function librosa.feature.melspectrogram() to extract the Mel frequency of the saved data;

[0014] 3) Use the Python library function librosa.feature.spectral_centroid() to extract the spectral centroid of the saved data;

[0015] 4) Use the python library function librosa.feature.mfcc() function to extract the MFCC (Mel Frequency Cepstral Coefficient) of the saved data.

[0016] Step 3: The speech emotion recognition network is divided into three parts: deep feature extraction subnet, classifier and classification layer feature fusion:

[0017] 1) Deep feature extraction subnet;

[0018] The MFCC in step 2 is sent to CNN (convolutional neural network), and the feature is convolved to obtain deep features. The network structure includes four convolution parts, each of which includes a convolution layer, a pooling layer, a normalization layer and a Dropout layer. The deep features are obtained through the convolution operation of the network.

[0019] 2) Classifier;

[0020] The cross entropy loss function is used to measure the difference between the predicted value and the true value distribution, and the emotion of the speech is divided into seven types: neutral, angry, afraid, happy, sad, disgusted, and bored.

[0021] 3) Classification layer feature fusion;

[0022] Feature fusion is performed based on the speech features extracted by the feature extractor and the deep features obtained by the deep feature extraction subnet. This part adopts the classification layer feature fusion algorithm.

[0023] The feature fusion algorithm based on the classification layer first records the feature categories extracted from the speech signal as n categories, inputs these n categories of features into n classifiers for training, and then uses the test data to obtain m categories of classification results. The recognition probability of the test data is expressed as {P ij (K), i = 1Ln, j = 1Lm}, and then these recognition results are fused and calculated according to the specified decision fusion rules, and finally the probability distribution of m different classification results is obtained, which is expressed as {q j (K),j=1Lm}, the new decision probability q obtained by decision fusion j (K) is calculated using formula (1):

[0024]

[0025] Among them, j represents the jth test data, q′ j (K) is calculated as follows:

[0026] q′ j (K) = R(p ij (K)) ⑵

[0027] The R function here represents the rules of classification layer fusion, and the predicted label (3) is finally calculated based on the judgment probability:

[0028] l(K)=argmax(q j (K)) ⑶

[0029]

[0030] The decision rule of the classification layer feature fusion algorithm uses the summation method, as shown in formula (4).

[0031] Step 4: Obtain speech emotion recognition results

[0032] According to the probability distribution provided in step 3, the final classification result is selected from the classification result sets with the largest probability, and the recognition result is matched to the seven emotions of neutral, angry, afraid, happy, sad, disgusted, and boredom.

[0033] The present invention has the following beneficial effects:

[0034] (1) The convolutional neural network is used to extract the deep features of acoustic features, which improves the feature utilization rate of acoustic features.

[0035] (2) The deep features of MFCC are integrated with the other three acoustic features to expand the feature utilization range of the data.

[0036] (3) The classification layer feature fusion method effectively utilizes the classifier recognition probability to improve the speech emotion recognition rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is the overall flow chart of the present invention.

[0038] Figure 2 It is a flow chart of the classification layer feature fusion in the present invention. DETAILED DESCRIPTION

[0039] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0040] A speech emotion recognition method based on feature fusion comprises the following steps:

[0041] Step 1: Acquire and preprocess user voice data;

[0042] Step 2: Put the data into the feature extractor to extract features;

[0043] Step 3: Input the speech emotion recognition network for emotion recognition;

[0044] Step 4: Obtain speech emotion recognition results.

[0045] Step 1: Use the Python library function librosa.load() to read the speech that needs emotion recognition, save it as a numpy data type, and preprocess it by constructing a pre-emphasis function and a windowing and framing function.

[0046] Step 2 specifically includes:

[0047] 1) Use the Python library function librosa.feature.zero_crossing_rate() to extract the zero crossing rate of the saved data;

[0048] 2) Use the Python library function librosa.feature.melspectrogram() to extract the Mel frequency of the saved data;

[0049] 3) Use the Python library function librosa.feature.spectral_centroid() to extract the spectral centroid of the saved data;

[0050] 4) Use the python library function librosa.feature.mfcc() function to extract the MFCC (Mel Frequency Cepstral Coefficient) of the saved data.

[0051] Step 3: The speech emotion recognition network is divided into three parts: deep feature extraction subnet, classifier and classification layer feature fusion:

[0052] 1) Deep feature extraction subnet;

[0053] The MFCC in step 2 is sent to CNN (convolutional neural network), and the feature is convolved to obtain deep features. The network structure includes four convolution parts, each of which includes a convolution layer, a pooling layer, a normalization layer and a Dropout layer. The deep features are obtained through the convolution operation of the network.

[0054] 2) Classifier;

[0055] The cross entropy loss function is used to measure the difference between the predicted value and the true value distribution, and the emotion of the speech is divided into seven types: neutral, angry, afraid, happy, sad, disgusted, and bored.

[0056] 3) Classification layer feature fusion;

[0057] Feature fusion is performed based on the speech features extracted by the feature extractor and the deep features obtained by the deep feature extraction subnet. This part adopts the classification layer feature fusion algorithm.

[0058] The feature fusion algorithm based on the classification layer first records the feature categories extracted from the speech signal as n categories, inputs these n categories of features into n classifiers for training, and then uses the test data to obtain m categories of classification results. The recognition probability of the test data is expressed as {P ij (K), i = 1Ln, j = 1Lm}, and then these recognition results are fused and calculated according to the specified decision fusion rules, and finally the probability distribution of m different classification results is obtained, which is expressed as {q j (K),j=1Lm}, the new decision probability q obtained by decision fusion j (K) is calculated using formula (1):

[0059]

[0060] Among them, j represents the jth test data, q′ j (K) is calculated as follows:

[0061] q′ j (K) = R(p ij (K)) ⑵

[0062] The R function here represents the rules of classification layer fusion, and finally the judgment probability is calculated.

[0063] The predicted label (3) is:

[0064] l(K)=argmax(q j (K)) ⑶

[0065]

[0066] The decision rule of the classification layer feature fusion algorithm uses the summation method, as shown in formula (4).

[0067] Step 4: Obtain speech emotion recognition results

[0068] According to the probability distribution provided in step 3, the final classification result is selected from the classification result sets with the largest probability, and the recognition result is matched to the seven emotions of neutral, angry, afraid, happy, sad, disgusted, and boredom.

[0069] The present invention designs a speech emotion recognition system based on feature fusion. The system mainly includes data acquisition and preprocessing, feature extractor and speech emotion recognition network. Data preprocessing mainly performs pre-emphasis and windowing and framing operations on the acquired data to normalize the sound data and eliminate some invalid data, which is convenient for feature extraction. The feature extractor extracts features from the data through the library function in python. The system extracts four features: zero crossing rate, Mel frequency, spectrum centroid, and MFCC (Mel frequency cepstral coefficient). The speech emotion recognition network consists of three parts: a deep feature extraction subnet, a classifier, and a classification layer feature fusion. The deep feature extraction subnet uses a convolutional neural network to perform a convolution operation on MFCC to extract its deep features. The classifier uses a cross entropy loss function to measure the difference between the distribution of predicted values ​​and true values ​​and completes the classification of speech emotions. The classification layer feature fusion is to fuse the MFCC deep features with the zero-crossing rate, Mel frequency, and spectrum centroid. This part uses a feature fusion algorithm based on the classification layer to correspond the above four features to four classifiers respectively. The model is first trained, and the test data is fused and calculated on the output recognition results through the specified decision fusion rules. Finally, the one with the largest probability in the probability distribution is selected as the recognition result. This invention has great application value for speech emotion recognition.

Claims

1. A speech emotion recognition method based on feature fusion, comprising the following steps: Step 1: Obtain and preprocess user voice data. Use the Python library function librosa.load() to read the voice that needs emotion recognition, save it as a numpy data type, and preprocess it by constructing a pre-emphasis function and a windowing and framing function. Step 2: Put the data into the feature extractor and extract the following four features for subsequent deep extraction and fusion of features. The specific steps are as follows: 1) Use the Python library function librosa.feature.zero_crossing_rate() to extract the zero crossing rate of the saved data; 2) Use the Python library function librosa.feature.melspectrogram() to extract the Mel frequency of the saved data; 3) Use the Python library function librosa.feature.spectral_centroid() to extract the spectral centroid of the saved data; 4) Use the python library function librosa.feature.mfcc() to extract the Mel frequency cepstral coefficients MFCC of the saved data; Step 3: Input the speech emotion recognition network for emotion recognition; the speech emotion recognition network includes a deep feature extraction subnet, a classifier, and a classification layer feature fusion: The deep feature extraction subnet sends the above MFCC into the convolutional neural network CNN, performs convolution operation on the feature to obtain deep features. The network structure includes four convolution parts, each of which includes a convolution layer, a pooling layer, a normalization layer and a Dropout layer. The deep features are obtained through the convolution operation of the network; The loss function used by the classifier is the cross entropy loss function, which is used to measure the difference between the predicted value and the true value distribution, and classify the emotion of the speech into seven emotions: neutral, angry, afraid, happy, sad, disgusted, and bored; The classification layer feature fusion is based on the speech features extracted by the feature extractor and the deep features obtained by the deep feature extraction subnet. This part adopts the classification layer feature fusion algorithm; The feature fusion algorithm based on the classification layer first records the feature categories extracted from the speech signal as n categories, inputs these n categories of features into n classifiers for training, and then uses the test data to obtain m categories of classification results. The recognition probability of the test data is expressed as {P ij (K), i = 1Ln, j = 1Lm}, and then these recognition results are fused and calculated according to the specified decision fusion rules, and finally the probability distribution of m different classification results is obtained, which is expressed as {q j (K), j = 1Lm}, the new decision probability q obtained by decision fusion j (K) is calculated using formula (1): Among them, j represents the jth test data, q′ j (K) is calculated as follows: q′ j (K)=R(p ij (K)) (2) The R function here represents the rules of classification layer fusion, and finally the predicted label (3) is calculated based on the judgment probability: l(K)=argmax(q j (K)) (3) The decision rule of the classification layer feature fusion algorithm uses the summation method, as shown in formula (4); Step 4: Obtain speech emotion recognition results; According to the probability distribution provided in step 3, the final classification result is selected from the classification result sets with the largest probability, and the recognition result is matched to the seven emotions of neutral, angry, afraid, happy, sad, disgusted, and boredom.

Citation Information

Patent Citations

  • Multi-mode emotion recognition and classification method

    CN107808146A

  • Emotion recognition method and device based on multi-feature fusion and storage medium

    CN110110653A