A method of speech emotion recognition based on neural network

By combining the neural network model of CNN and BILSTM, key features of speech signals are extracted and selected, and the problem of long recognition time and poor results in traditional methods is solved, and fast and accurate speech emotion recognition is achieved.

CN114495989BActive Publication Date: 2025-05-06ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210216452.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-05-06
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

Traditional speech emotion recognition methods require a long time and have poor recognition effects, and existing methods cannot effectively utilize the continuity and interpretability of speech signals.

Method used

A neural network-based speech emotion recognition method is adopted, and speech features are extracted using CNN combined with BILSTM, key features are selected through attention mechanism, and emotional classification is performed using the full connection layer and softmax layer.

Benefits of technology

Fast and accurate speech emotion recognition is achieved, and the rich emotional characteristics in speech signals can be adaptively extracted, improving recognition accuracy and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495989B_ABST
    Figure CN114495989B_ABST
Patent Text Reader

Abstract

A speech emotion recognition method using a neural network includes: step 1) acquiring speech feature MFCC; step 2) inputting the speech emotion recognition network based on a neural network designed in the present invention to perform speech emotion recognition; step 3) obtaining a speech emotion recognition result. The present invention uses a neural network to perform emotion recognition on a speech signal, designs and implements a speech emotion recognition network based on a neural network, and when performing speech emotion recognition, only the speech signal to be recognized needs to be input to obtain the recognized speech emotion. For the input speech signal, the speech emotion recognition network can adaptively extract speech features containing rich emotional information, and use the emotional features to perform accurate speech emotion recognition. The invention can accurately perform speech emotion recognition for everyone, and has great application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a speech emotion recognition method based on a neural network. Background Art

[0002] Speech emotion recognition usually analyzes speech signal processing to determine individual emotional information. Traditional speech emotion recognition methods are all based on machine learning, including SVM, HMM, Bayesian and other machine learning methods. These recognition methods require a certain amount of learning time. The recognition of a speech signal often takes a long time and the effect is not very good. Thanks to the development of neural networks, more and more speech emotion recognition methods based on deep learning neural networks have been proposed. These methods can complete the recognition of speech signals in a shorter time, and the accuracy of the recognition results is relatively accurate. Among these methods, some methods directly use convolutional neural network extraction to perform speech emotion recognition on speech signal features. Although such methods reduce the dimensionality of speech signal features, they cannot perform speech emotion recognition based on the continuity of speech signals, and the interpretability of the methods is low. Another long short-term memory network extracts features in time continuity to perform speech emotion recognition. These methods have high interpretability, and through key feature selection, feature information containing richer emotions is selected, which increases the accuracy of speech emotion recognition. Summary of the invention

[0003] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provide a speech emotion recognition method based on a neural network.

[0004] The present invention solves the technical problem by adopting the following technical solution:

[0005] A speech emotion recognition method based on a neural network comprises the following steps:

[0006] Step 1: Obtain voice feature information;

[0007] Step 2: inputting the speech emotion recognition network based on neural network designed in the present invention to perform speech emotion recognition;

[0008] Step 3: Obtain speech emotion recognition results.

[0009] Step 1: Obtain speech feature parameters; specifically including:

[0010] 1) Use the python library function python_speech_features.mfcc() to read the MFCC feature coefficients in the speech signal;

[0011] 2) Use the python library function python_speech_features.delta() to read the first-order difference MFCC feature coefficients and the second-order difference MFCC feature coefficients of the speech signal;

[0012] 3) Use the Python library function numpy.hstack() to concatenate the three speech features and convert the speech features into a size of 199*39. Use the same size for all speech signals to be evaluated.

[0013] 4) Use the Python library function transforms.ToTensor() to convert the feature data into a Tensor vector for use in the subsequent speech emotion recognition network;

[0014] The speech emotion recognition network used in step 2 is divided into three parts: feature extractor, key feature selection subnet and evaluation subnet:

[0015] 1) Feature extractor;

[0016] The feature extractor uses CNN (convolutional neural network) combined with BILSTM (bidirectional long short-term memory network) as a feature extractor to further extract the features of the speech signal; after the feature extractor, the convolutional neural network performs dimensionality reduction processing on the speech features, and the long short-term memory network makes the speech signal features continuous, and finally generates new key features. The output feature is represented as H = {h1, h2, ..., h L};

[0017] 2) Key feature selection subnet;

[0018] The key feature selection subnet further selects features with key information content in the speech features, that is, areas with key features, based on the speech features extracted by the feature extractor;

[0019] The input of the key feature selection subnet is the dimension reduction features generated by the feature extractor. For these features, the ATTENTION attention mechanism network is used to score the importance of speech emotion features, focus on the parts with prominent speech emotion features, and select features with key information, as shown in formula (1):

[0020]

[0021] K, Q, and V represent the key features of the input features, the query value, and the weight value of the current key feature. The first step is to calculate the correlation between each query value and each key feature with the corresponding weight coefficient, expressed as S(Q, K). The second step is to use the Softmax function for normalization, where L xRepresents the length of the feature data, as shown in formula (2).

[0022]

[0023] The third step is to perform weighted summation on the weight coefficient and the corresponding key value to obtain the final attention value, as shown in formula (3).

[0024]

[0025] 3) Evaluate subnet;

[0026] The evaluation subnet performs speech emotion evaluation; for the input speech signal, the feature extractor extracts the features of the speech signal, and the key feature selection subnet selects the key features of the speech signal with outstanding information. The speech features are extracted by the feature extractor, and the final speech emotion classification result is obtained through two fully connected layers and a softmax layer;

[0027] Step 3: Obtain speech emotion recognition results

[0028] When performing speech emotion analysis, speech emotions are divided into seven emotion categories: anger, sadness, happiness, fear, neutral, disgust and boredom.

[0029] After the speech signal to be evaluated is input into the speech emotion evaluation network, its corresponding emotion type will be output, ranging between 7 emotion types.

[0030] The present invention has the following beneficial effects:

[0031] (1) Accurately identify emotion based on speech features.

[0032] (2) Representative emotional features can be selected from speech features, that is, feature areas with key emotional information.

[0033] (3) The invention can accurately identify and evaluate the speaker's voice, making it easier for others to understand the other person's emotional fluctuations, and has great application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is the overall flow chart of the present invention.

[0035] Figure 2 It is a structural diagram of the identification network in the present invention. DETAILED DESCRIPTION

[0036] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0037] A method for speech emotion recognition based on a neural network comprises the following steps:

[0038] Step 1: Obtain voice feature information;

[0039] Step 2: inputting the speech emotion recognition network based on neural network designed in the present invention to perform emotion recognition;

[0040] Step 3: Obtain speech emotion recognition results.

[0041] Step 1 specifically includes:

[0042] 1) Use the python library function python_speech_features.mfcc() to read the MFCC feature coefficients in the speech signal;

[0043] 2) Use the python library function python_speech_features.delta() to read the first-order difference MFCC feature coefficients and the second-order difference MFCC feature coefficients of the speech signal;

[0044] 3) Use the Python library function numpy.hstack() to concatenate the three speech features and convert the speech features into a size of 199*39. Use the same size for all speech signals to be evaluated.

[0045] 4) Use the Python library function transforms.ToTensor() to convert the feature data into a Tensor vector for use in the subsequent speech emotion recognition network;

[0046] The bone age assessment network used in step 2 is divided into four parts, namely feature extractor, region of interest selection subnet, guidance subnet and evaluation subnet:

[0047] The speech emotion recognition network used in step 2 is divided into three parts: feature extractor, key feature selection subnet and evaluation subnet:

[0048] 1) Feature extractor;

[0049] The feature extractor uses CNN (convolutional neural network) combined with BILSTM (bidirectional long short-term memory network) as a feature extractor to further extract the features of the speech signal; after the feature extractor, the convolutional neural network performs dimensionality reduction processing on the speech features, and the long short-term memory network makes the speech signal features continuous, and finally generates new key features. The output feature is represented as H = {h1, h2, ..., h L};

[0050] 2) Key feature selection subnet;

[0051] The key feature selection subnet further selects features with key information content in the speech features, that is, areas with key features, based on the speech features extracted by the feature extractor;

[0052] The input of the key feature selection subnet is the dimension reduction features generated by the feature extractor. For these features, the ATTENTION attention mechanism network is used to score the importance of speech emotion features, focus on the parts with prominent speech emotion features, and select features with key information, as shown in formula (1):

[0053]

[0054] K, Q, and V represent the key features of the input features, the query value, and the weight value of the current key feature. The first step is to calculate the correlation between each query value and each key feature with the corresponding weight coefficient, expressed as S(Q, K). The second step is to use the Softmax function for normalization, where L x Represents the length of the feature data, as shown in formula (2).

[0055]

[0056] The third step is to perform weighted summation on the weight coefficient and the corresponding key value to obtain the final attention value, as shown in formula (3).

[0057]

[0058] 3) Evaluate subnet;

[0059] The evaluation subnet performs speech emotion evaluation; for the input speech signal, the feature extractor extracts the features of the speech signal, and the key feature selection subnet selects the key features of the speech signal with outstanding information. The speech features are extracted by the feature extractor, and the final speech emotion classification result is obtained through two fully connected layers and a softmax layer;

[0060] Step 3: Obtain speech emotion recognition results

[0061] When performing speech emotion analysis, speech emotions are divided into seven emotion categories: anger, sadness, happiness, fear, neutral, disgust and boredom.

[0062] After the speech signal to be evaluated is input into the speech emotion recognition network, its corresponding emotion type will be output, ranging between 7 emotion types.

[0063] The present invention designs a speech emotion recognition model based on a neural network. The model consists of three parts: a feature extractor, a key feature selection subnet and an evaluation subnet. The feature extractor is implemented based on a convolutional neural network and a recurrent neural network, and uses CNN combined with BILSTM to extract speech features. The key feature selection subnet is used to select key areas in the features, and these feature areas contain representative emotional features that can help classification. The evaluation subnet uses the extracted features to perform speech emotion recognition. The speech emotion recognition model proposed in the invention can further extract emotional feature information in the speech, and use these feature information to improve the accuracy of speech emotion recognition. The present invention uses a neural network method to perform emotion recognition on speech features, designs and implements a speech emotion recognition network based on a neural network. When performing speech emotion recognition, only the speech information to be evaluated needs to be input to obtain the recognized speech emotion type. For the input speech signal, the speech emotion recognition network can adaptively extract features containing rich emotions, and use emotional features to perform accurate speech emotion evaluation. The invention can accurately perform speech emotion recognition for everyone, and has great application value.

Claims

1. A speech emotion recognition method based on a neural network, comprising the following steps: Step 1: Obtain speech feature parameters; specifically including: 1) Use the python library function python_speech_features.mfcc() to read the MFCC feature coefficients in the speech signal; 2) Use the python library function python_speech_features.delta() to read the first-order difference MFCC feature coefficients and the second-order difference MFCC feature coefficients of the speech signal; 3) Use the python library function numpy.hstack() to concatenate the three speech feature parameters and convert the speech features into a size of 199*39. Use the same size for all speech signals to be evaluated. 4) Use the Python library function transforms.ToTensor() to convert the feature data into a Tensor vector for use in the subsequent speech emotion recognition network; Step 2: Input speech feature parameters to analyze speech emotion; the speech emotion recognition network includes a feature extractor, a key feature selection subnet and a recognition subnet: The feature extractor uses a convolutional neural network (CNN) combined with a bidirectional long short-term memory (BILSTM) network as a feature extractor to further extract the features of the speech signal; after the feature extractor, the convolutional neural network performs dimensionality reduction processing on the speech features, and the long short-term memory network makes the speech signal features continuous, and finally generates new key features. The output feature is represented as H = {h1, h2, ..., h L }; The key feature selection subnet further selects features with key information content in the speech features, that is, areas with key features, based on the speech features extracted by the feature extractor; The input of the key feature selection subnet is the dimension reduction features generated by the feature extractor. For these features, the ATTENTION attention mechanism network is used to score the importance of speech emotion features, focus on the parts with prominent speech emotion features, and select features with key information, as shown in formula (1): Where K, Q, and V represent the key features of the input features, the query value, and the weight value of the current key feature respectively; the first step is to calculate the correlation between each query value and each key feature to obtain the corresponding weight coefficient, which is expressed as S i ; The second step is to use the Softmax function for normalization, where L x Represents the length of the feature data, as shown in formula (2); The third step is to perform weighted summation on the weight coefficient and the corresponding weight value of the current key feature to obtain the final attention value, as shown in formula (3); The evaluation subnet performs speech emotion evaluation. For the input speech signal, the feature extractor extracts the features of the speech signal, and the key feature selection subnet selects the key features of the speech signal that highlight the amount of emotional information. The speech features are extracted by the feature extractor, and the final speech emotion classification result is obtained after passing through two fully connected layers and a softmax layer. Step 3: Obtain speech emotion evaluation results; When performing voice emotion analysis, voice emotions are divided into 7 types: angry, sad, happy, afraid, neutral, disgusted and bored; After the speech signal to be evaluated is input into the speech emotion recognition network, its corresponding emotion type will be output, and the range of the emotion type is between the above 7 emotion types.

Citation Information

Patent Citations

  • Voice emotion recognition based on direction and self-attention mechanism and bidirectional long-short-term network

    CN110400579A

  • Speech emotion recognition method of parallel convolutional recurrent neural network based on spectrogram characteristics

    CN110534132A