Speech Emotion Recognition Method and Device

By using the dual-channel deep neural network model of Mel cepstral coefficient features and manual features in speech emotion recognition, and combining it with the XGBoost model for fusion, the problems existing in speech emotion recognition of a single feature set and model are solved, achieving higher recognition accuracy.

CN115035917BActive Publication Date: 2025-06-27KONKA GROUP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210848992.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-06-27
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

In the prior art, the speech emotion recognition method has the problem of insufficient expression of emotional characteristics of a single feature set and inaccurate classification of a single model in certain emotional categories, resulting in low recognition accuracy.

Method used

A speech emotion recognition method is adopted to improve the accuracy of speech emotion recognition by using the dual-channel deep neural network model of Mel cepstrum coefficient features and manual features at the same time, and combining it with the XGBoost model.

Benefits of technology

Through this method, the problems of insufficient expression of a single feature set and inaccurate classification of a single model are solved, which significantly improves the accuracy of speech emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035917B_ABST
    Figure CN115035917B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of speech data processing, and in particular, to a speech emotion recognition method and apparatus. The method includes: obtaining speech data to be recognized; extracting Mel cepstral coefficient features, ComParE feature sets, and audio bag-of-words feature sets according to the speech data to be recognized; inputting the Mel cepstral coefficient features, ComParE feature sets, and audio bag-of-words feature sets into a trained speech emotion recognition model, where the speech emotion recognition model includes a dual-channel neural network sub-model and an extreme gradient boosting sub-model; obtaining a first probability matrix according to the Mel cepstral coefficient features, audio bag-of-words feature sets, and dual-channel neural network sub-model, and obtaining a second probability matrix according to the ComParE feature sets and extreme gradient boosting sub-model; fusing the first probability matrix and the second probability matrix into a comprehensive probability matrix, and determining a prediction result of the speech data to be recognized according to the comprehensive probability matrix. By using the present disclosure, the accuracy of speech emotion recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of speech data processing, and particularly to a speech emotion recognition method and apparatus. Background Art

[0002] With the development of artificial intelligence, more and more researchers have started to focus on the direction of speech emotion recognition. Emotion is a phenomenon that synthesizes human behavior, thought, and feeling. A person's emotion can be felt from his language, and speech information plays a crucial role in language. It not only contains the content that the speaker wants to express, but also contains the speaker's emotional information at that time, whether it is anger or joy. This plays a crucial role in certain industries, such as the service industry, where corresponding service adjustments can be made according to the emotions expressed in the customer's words; in psychotherapy, the treatment plan can also be adjusted according to the patient's emotions.

[0003] Of course, a person's emotional expression can be reflected not only from a complete sentence or some words and phrases in the language, but also from some non-verbal information. For example, when a person is sad, there will be some sobbing sounds, and when a person is excited, there will be screaming sounds. These are all non-verbal information, and they can also express a person's emotion.

[0004] Currently, common technologies in speech emotion recognition are as follows:

[0005] 1. Use a single feature set as the model input. For example, traditional machine learning algorithms such as SVM (Support Vector Machine), XGBoost (eXtreme Gradient Boosting), GMM (Gaussian mixture model), KNN (K-Nearest Neighbor), HMM (Hidden Markov Model), etc. Some low-level handcrafted feature sets LLDs (Low Level Descriptors) or HSFs (High level Statistics Functions, features obtained by performing some statistics on the basis of LLDs) can be used as the feature set to be input into the model for training; or, deep learning methods such as spectrogram + CRNN (Convolutional Recurrent Neural Network) or handcrafted features + CRNN can be used. After frame windowing is performed on the original signal, many frames can be obtained. For each frame, FFT (Fast Fourier Transform) is performed. The role of the Fourier transform is to convert the time-domain signal into a frequency-domain signal. Stacking the frequency-domain signals (spectrograms) after FFT for each frame in time can obtain a spectrogram. Handcrafted features include: ComParE feature set (Computational Paralinguistics ChallengE, the feature set used in the ComParE challenge), BoAW (Bag-of-Audio-Words), etc. Since 2013, the ComParE challenge has required the use of a designed feature set, which contains 6,373 static features and is obtained by calculating various functions on the LLDs, called the ComParE feature set. BoAW is a further organizational representation of features and is obtained by calculating the LLDs according to a codebook.

[0006] However, there are the following problems with this method of using a single feature set as the model input: Since the current speech emotion feature representation is not clear, using a single feature set to represent a certain emotion may cause the loss of emotion information, resulting in inaccurate classification.

[0007] 2. Use a single model. For example, use the original speech signal plus a deep learning network model. The original speech signal read from an audio file is usually called raw waveform, which is a one-dimensional array. Its length is determined by the audio length and the sampling rate. For example, if the sampling rate Fs is 16KHz, it means 16,000 points are sampled in one second. At this time, if the audio length is 10 seconds, then there are 160,000 values in the raw waveform, and the magnitude of the values usually represents the amplitude.

[0008] However, this method of using a single model has the following problems: Since there are many types of human emotions and each model has different abilities to distinguish different emotions, it is very difficult to have a model with good classification effects for each type of emotion, which in turn leads to inaccurate classification of some emotion categories. Summary of the Invention

[0009] To solve the problems of the prior art, embodiments of the present invention provide a method and device for speech emotion recognition, which can improve the accuracy of speech emotion recognition. The technical solutions are as follows:

[0010] According to one aspect of the present disclosure, there is provided a method for speech emotion recognition, the method comprising:

[0011] Obtain the speech data to be recognized;

[0012] Extract Mel cepstral coefficients features, ComParE feature set and audio bag-of-words feature set according to the speech data to be recognized;

[0013] Input the Mel cepstral coefficients features, ComParE feature set and audio bag-of-words feature set into a trained speech emotion recognition model, the speech emotion recognition model includes a dual-channel neural network sub-model and an extreme gradient boosting sub-model;

[0014] Obtain a first probability matrix according to the Mel cepstral coefficients features, audio bag-of-words feature set and dual-channel neural network sub-model, and obtain a second probability matrix according to the ComParE feature set and extreme gradient boosting sub-model;

[0015] Fuse the first probability matrix and the second probability matrix into a comprehensive probability matrix, and determine the prediction result of the speech data to be recognized according to the comprehensive probability matrix.

[0016] According to another aspect of the present disclosure, there is provided a speech emotion recognition device, the speech emotion recognition device is used to execute the above speech emotion recognition method, and the device includes:

[0017] An obtaining module, configured to obtain the speech data to be recognized;

[0018] An extraction module, configured to extract Mel cepstral coefficient features, ComParE feature sets, and audio bag-of-words feature sets according to the speech data to be recognized;

[0019] An input module, configured to input the Mel cepstral coefficient features, ComParE feature sets, and audio bag-of-words feature sets into a trained speech emotion recognition model, where the speech emotion recognition model includes a dual-channel neural network sub-model and an extreme gradient boosting sub-model;

[0020] A processing module, configured to obtain a first probability matrix according to the Mel cepstral coefficient features, audio bag-of-words feature sets, and dual-channel neural network sub-model, and obtain a second probability matrix according to the ComParE feature sets and extreme gradient boosting sub-model;

[0021] A determination module, configured to fuse the first probability matrix and the second probability matrix into a comprehensive probability matrix, and determine a prediction result of the speech data to be recognized according to the comprehensive probability matrix.

[0022] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0023] A processor; and

[0024] A memory storing a program,

[0025] wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method described in any one of the above-mentioned speech emotion recognition methods.

[0026] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in any one of the above-mentioned speech emotion recognition methods.

[0027] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the method described in any one of the above-mentioned speech emotion recognition methods.

[0028] One or more technical solutions provided in the embodiments of the present application adopt a dual-channel deep neural network model that simultaneously uses Mel cepstral coefficient features and manual features, and at the same time supplements the XGBoost model as a fusion solution, which solves the problem of insufficient expression of emotion features by a single feature set and the problem of inaccurate classification of a single model in some emotion categories, thereby improving the accuracy of speech emotion recognition. Description of the Drawings

[0029] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features, and advantages of the present disclosure are disclosed. In the drawings:

[0030] Figure 1 A flowchart of a voice emotion recognition method according to an exemplary embodiment of the present disclosure is shown;

[0031] Figure 2 A flowchart of a method for training an initial voice emotion recognition model according to an exemplary embodiment of the present disclosure is shown;

[0032] Figure 3 A schematic diagram of the overall structure of an initial voice emotion recognition model according to an exemplary embodiment of the present disclosure is shown;

[0033] Figure 4 A schematic diagram of the structure of a first channel module according to an exemplary embodiment of the present disclosure is shown;

[0034] Figure 5 A schematic diagram of the structure of a second channel module according to an exemplary embodiment of the present disclosure is shown;

[0035] Figure 6 A flowchart of a method for voice emotion recognition based on a voice emotion recognition model according to an exemplary embodiment of the present disclosure is shown;

[0036] Figure 7 A schematic block diagram of a voice emotion recognition device according to an exemplary embodiment of the present disclosure is shown;

[0037] Figure 8 A schematic block diagram of an exemplary electronic device capable of implementing the embodiments of the present disclosure is shown. Detailed Description of the Embodiments

[0038] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0039] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0040] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0041] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0042] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0043] The following refers to the attached Figure 1 Describe a voice emotion recognition method of this disclosure. This method can be implemented by an electronic device. The electronic device can include a terminal or a server. This method includes the following steps:

[0044] Step 101, obtain the voice data to be recognized;

[0045] Step 102, extract Mel cepstral coefficient features, ComParE feature set and audio bag-of-words feature set according to the voice data to be recognized;

[0046] Step 103, input the Mel cepstral coefficient features, ComParE feature set and audio bag-of-words feature set into the trained voice emotion recognition model. The voice emotion recognition model includes a dual-channel neural network sub-model and an extreme gradient boosting sub-model;

[0047] Step 104, obtain the first probability matrix according to the Mel cepstral coefficient features, audio bag-of-words feature set and dual-channel neural network sub-model, and obtain the second probability matrix according to the ComParE feature set and extreme gradient boosting sub-model;

[0048] Step 105, fuse the first probability matrix and the second probability matrix into a comprehensive probability matrix, and determine the prediction result of the voice data to be recognized according to the comprehensive probability matrix.

[0049] Optionally, the dual-channel neural network sub-model includes a first channel module, a second channel module and a classification module;

[0050] Based on the Mel-frequency cepstral coefficient features, the audio bag-of-words feature set, and the dual-channel neural network sub-model, a first probability matrix is obtained, including:

[0051] Input the Mel-frequency cepstral coefficient features into the first channel module to obtain a 64-dimensional first feature vector;

[0052] Input the audio bag-of-words feature set into the second channel module to obtain a 64-dimensional second feature vector;

[0053] Process the first feature vector and the second feature vector through the classification module to obtain the first probability matrix.

[0054] Optionally, the first channel module includes two convolutional layers, a two-layer LSTM network, and a fully connected layer;

[0055] Input the Mel-frequency cepstral coefficient features into the first channel module to obtain a 64-dimensional first feature vector, including:

[0056] Convolve the input Mel-frequency cepstral coefficient features through two convolutional layers. After each convolution, batch normalization operation and rectified linear unit activation are performed, and max pooling layer is used for downsampling;

[0057] Perform feature extraction on the downsampled feature data through a two-layer LSTM network, input the extracted data into the fully connected layer, and obtain a 64-dimensional first feature vector through the fully connected layer.

[0058] Optionally, the second channel module includes two fully connected layers;

[0059] Input the audio bag-of-words feature set into the second channel module to obtain a 64-dimensional second feature vector, including:

[0060] Input the audio bag-of-words feature set into two fully connected layers. After passing through each fully connected layer, batch normalization operation and rectified linear unit activation are performed, and then dropout processing is carried out to obtain a 64-dimensional second feature vector.

[0061] Optionally, the classification module includes a concatenation sub-module, a fully connected layer, and a Softmax function;

[0062] Process the first feature vector and the second feature vector through the classification module to obtain the first probability matrix, including:

[0063] Concatenate the first feature vector and the second feature vector through the concatenation sub-module to obtain a 128-dimensional feature vector;

[0064] Input the 128-dimensional feature vector into the fully connected layer, and after activating the output of the fully connected layer with the Softmax function, obtain the first probability matrix.

[0065] Optionally, fusing the first probability matrix and the second probability matrix into a comprehensive probability matrix includes:

[0066] Combining two numerical values at corresponding positions of the first probability matrix and the second probability matrix to form a new two-dimensional vector;

[0067] Calculating the L2 norm of the new two-dimensional vector according to the following formula (1):

[0068]

[0069] where a ji represents a numerical value of the first probability matrix, b ji represents a numerical value of the second probability matrix, j represents the number of rows of the matrix, that is, the j-th speech data in the speech data to be recognized, i represents the number of columns of the matrix, and P ji is the predicted probability of the i-th class label of the j-th speech data;

[0070] Forming a comprehensive probability matrix according to the L2 norm (a kind of Euclidean norm) of the new two-dimensional vector.

[0071] Optionally, the training process of the speech emotion recognition model includes:

[0072] Obtaining an untrained initial speech emotion recognition model, where the initial speech emotion recognition model includes an initial dual-channel neural network sub-model and an initial extreme gradient boosting sub-model;

[0073] Obtaining a training set and performing data augmentation on the training set;

[0074] Extracting the Mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set of the training set according to the augmented training set;

[0075] Training the initial dual-channel neural network sub-model according to the Mel cepstral coefficient features and the audio bag-of-words feature set of the training set to obtain a trained dual-channel neural network sub-model;

[0076] Performing grid search on the parameters of the initial extreme gradient boosting sub-model through GridSearchCV (Grid Search Cross Validation) to determine the initial parameters of the initial extreme gradient boosting sub-model, and training the initial extreme gradient boosting sub-model with the determined initial parameters according to the extracted ComParE feature set to obtain a trained extreme gradient boosting sub-model;

[0077] Obtaining a development set and extracting the Mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set of the development set;

[0078] According to the Mel cepstral coefficient features of the development set, the audio bag-of-words feature set, and the trained dual-channel neural network sub-model, obtain the first probability matrix of the development set. According to the ComParE feature set and the extreme gradient boosting sub-model, obtain the second probability matrix of the development set;

[0079] Fuse the first probability matrix of the development set and the second probability matrix of the development set into the comprehensive probability matrix of the development set. According to the comprehensive probability matrix of the development set, determine the prediction result of the development set;

[0080] Evaluate the performance of the trained speech emotion recognition model according to the prediction result of the development set.

[0081] In the embodiments of the present disclosure, a dual-channel deep neural network model that simultaneously uses Mel cepstral coefficient features and hand-crafted features is adopted, and the XGBoost model is supplemented as a fusion solution, which solves the problem of insufficient expression of emotion features by a single feature set and the problem of inaccurate classification of a single model in certain emotion categories, thereby improving the accuracy of speech emotion recognition.

[0082] The following refers to the attached Figure 2 Describe a method for training an initial speech emotion recognition model of the present disclosure. This method can be implemented by an electronic device. The electronic device can include a terminal or a server. The method includes the following steps:

[0083] Step 201, obtain an untrained initial speech emotion recognition model. The initial speech emotion recognition model includes an initial dual-channel neural network sub-model and an initial extreme gradient boosting sub-model.

[0084] In a feasible implementation, the overall structure of the initial speech emotion recognition model can be as Figure 3 shown. The input data is respectively brought into the first channel module and the second channel module, and the results of the two channels are merged and then passed through the classification module to obtain the output of the model prediction category. Among them:

[0085] (1) The first channel module: The design of this module can be as Figure 4 shown. The input Mel cepstral coefficient features pass through two convolutional layers, and the convolutional kernels are both 3*3. After each convolution, batch normalization operation and rectified linear unit activation are performed, and max pooling with a pool size of 2 is used for downsampling. Then, it passes through two layers of LSTM (Long Short-Term Memory) networks. Set the hidden size of this LSTM network to 128 and the dropout rate to 0.5. Concatenate the hidden layer states of the last moment of the two-layer LSTM and send them into the FC fully connected layer, and finally output a 64-dimensional feature vector.

[0086] (2) Second channel module: The design of this module can be as follows Figure 5 shown. The extracted audio bag-of-words feature set is used as the input and sent to two fully connected (FC) layers. The output dimensions of the two fully connected layers are 512 and 64 respectively. And after each pass through the fully connected layer, batch normalization operation and rectified linear unit activation are performed, and then dropout processing is carried out to prevent overfitting. Finally, a 64-dimensional feature vector is output.

[0087] (3) Classification module: The design of this model can refer to the classification module part in Figure 3 . The 64-dimensional results output by the two channels are concatenated to obtain a 128-dimensional feature vector. After passing through a fully connected layer and being activated by the Softmax function, a 6-dimensional output vector is obtained, which respectively corresponds to the prediction probability sizes of 6 labels.

[0088] Step 202: Obtain the training set and the development set.

[0089] In a feasible implementation, there are various ways to obtain the dataset used for training the model. One feasible way is provided by Natalie Holz of MPI in Frankfurt. The characteristics are vocalizations with different emotional intensities such as laughter, crying, groaning or screaming, indicating different emotions. The dataset consists of 625 training sets (which can be called the Train set) and 460 development sets (which can be called the Development set), and all are female voices.

[0090] Optionally, in order to more accurately evaluate the performance of the trained model, the trained model can also be tested using test data. The test data has 276 male voices. These datasets have 6 labels, including achievement, anger, fear, pain, happiness and surprise, and the duration of each sampling is about 1 second.

[0091] It should be noted that the number of each type of label in the database provided by Natalie Holz of MPI in Frankfurt is relatively balanced, so no compensation measures need to be taken.

[0092] Step 203: Perform data augmentation on the training set.

[0093] In a feasible implementation, data augmentation can improve the diversity of data and enhance the robustness of the model. It is generally used for the training set. Neural networks require a large number of parameters, and the parameters of many neural networks are in the millions. And to make these parameters work correctly, a large amount of data is needed for training. However, in many actual projects, it is difficult to find sufficient data to complete the task. To solve this problem, the training samples can be randomly changed instead of preparing a large number of training samples. This can reduce the dependence of the model on certain attributes, thereby improving the generalization ability of the model.

[0094] The present disclosure uses the librosa library (a third-party library for Python speech signal processing) to perform speed and pitch enhancement on the sample data in the training set. Since both the training set and the development set are female voices, while the test set are male voices, in order to reduce the difference between the two, the present disclosure also uses the Parselmouth library (a singing synthesis speech synthesis annotation library) to perform male-female voice conversion on the test speech, converting male voices to female voices, thereby reducing the degree of mismatch between the training and testing processes.

[0095] Step 204: According to the enhanced training set, extract the Mel-frequency cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set of the training set.

[0096] In a feasible implementation manner, the Mel-frequency cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set are described as follows:

[0097] (1) Mel-frequency cepstral coefficient features:

[0098] Mel-frequency cepstral coefficients (MFCC, Mel-Frequency Cepstral Coefficients) are one of the common features in speech signal processing. The process of extracting Mel-frequency cepstral coefficient features can be: first, truncate the input speech data or pad it with zeros to align the data, unifying it to a length of 1 s. Data alignment is for the convenience of batch input into the deep network model for calculation. For example, if one speech is 1.5 s long and another speech is 0.8 s long, the feature lengths extracted from them are also different, and they cannot be divided into the same batch for calculation. After data alignment, use the librosa tool to extract Mel-frequency cepstral coefficient features from the 1-s speech data, where the n_mfcc parameter is set to 40, obtaining a Mel-frequency cepstral coefficient feature of 40 * 88. Librosa is a Python toolkit for audio, music analysis, and processing, including various common functions such as time-frequency processing, feature extraction, and sound graph plotting.

[0099] (2) ComParE feature set:

[0100] Since 2013, the ComParE challenge has required the use of a designed feature set, which contains 6,373 static features calculated by various functions on the LLD and is called the feature set. It can be obtained through the openSmile open-source package. openSmile is a tool that runs in command-line form. By configuring the config file, it can be used as a feature extractor for signal processing and machine learning. It has characteristics such as high modularity and flexibility. The most basic function of openSMILE can be used for the extraction of speech signal features. Of course, it can also analyze other forms of signals, such as visual signals, medical physiological signals, etc. openSMILE is written in C++ and has the characteristics of high speed and efficiency. It has a flexible architecture and can run on major mainstream operating systems.

[0101] (3) Bag of Audio Words feature set (which can be called the BoAW feature set)

[0102] BoAW is a further organizational representation of features. This process is calculated using the ComParE feature set obtained in the previous step. The BoAW representation can be obtained using the openXBOW open-source package, and finally 2,000 features are extracted. openXBOW is an open-source toolkit for generating bag-of-words (BoW) representations from multimodal inputs. In the BoW principle, the histogram of words is first used as a feature for document classification, but its idea is and can be easily adapted, for example, acoustic or visual descriptors, introducing the previous step of vector quantization. The openXBOW toolkit supports any number of input features and text inputs and connects computational sub-packages to the final package. It provides various extensions and options. openXBOW is the first publicly available toolkit for generating cross-modal vocabulary packages, and the functions of this tool have been verified in different scenarios: sentiment analysis in tweets, snore classification, and time-dependent sentiment recognition based on sound, language, and visual information, with results superior to other feature representations.

[0103] Step 205: According to the Mel cepstral coefficient features of the training set and the bag of audio words feature set, train the initial dual-channel neural network sub-model to obtain a trained dual-channel neural network sub-model.

[0104] In a feasible implementation, the following training parameters are used when training the initial dual-channel neural network sub-model:

[0105] (1) Number of training epochs: 80;

[0106] (2) Batch size of data: 32;

[0107] (3) Initial learning rate: 0.001;

[0108] (4) Learning rate decay rate: 97%;

[0109] (5) Loss function: cross-entropy loss;

[0110] (6) Optimizer: Adam;

[0111] Input the Mel-frequency cepstral coefficient (MFCC) features of size [32*40*88] and the Bag of Audio Words (BoAW) features of size [32*2000] into the model; among them, 32 in the MFCC features represents the data batch size, 40 represents the dimension of the input MFCC feature vector, 88 represents the frame length of the input speech; 32 in the BoAW features represents the data batch size, and 2000 represents the dimension of the input BoAW feature vector.

[0112] Use the MFCC features of size [32*40*88] for calculation in the first channel. First, reshape this feature to [32*1*40*88], where the newly added 1 represents the number of channels, for the convenience of subsequent two-dimensional convolution operations. After bringing this feature into the first convolutional layer and passing through the pooling layer, it becomes a tensor data of size [32*16*20*44] (a data format of neural network). Then bring this data into the second convolutional layer, and after passing through the pooling layer, it becomes a tensor data of size [32*32*10*22]. Then transform the dimension of this data to [22*32*320], and bring it into the LSTM layer. The output hidden layer state is [32*2*128], where 32 represents the data batch size, 2 represents the number of hidden layers, and 128 represents the hidden layer feature dimension. Concatenate the two hidden layers to form a data of size [32*256]. Pass it through the fully connected layer to get a tensor data of size [32*64].

[0113] Use the BoAW features of size [32*2000] for calculation in the second channel. Pass this feature through the first fully connected layer to get a data of size [32*512], and then pass it through the second fully connected layer to get a tensor data of size [32*64].

[0114] In the Classification Block, the tensor data of size [32*64] output from the first channel and the tensor data of size [32*64] output from the second channel are concatenated to form a tensor data of size [32*128]. This data is fed into the fully connected layer and after passing through softmax, it becomes the output data of size [32*6]. This data represents a probability matrix of [32*6], where 32 is the data batch size and 6 is the predicted probability for each category. The output matrix and the true training data labels are brought into the cross-entropy loss function to calculate the loss value, and then through the Adam optimization algorithm for backpropagation to reduce the loss, thereby improving the model classification effect.

[0115] Step 206: Perform grid search on the parameters of the initial extreme gradient boosting sub-model through the GridSearchCV grid search algorithm to determine the initial parameters of the initial extreme gradient boosting sub-model. Train the initial extreme gradient boosting sub-model with the determined initial parameters according to the extracted ComParE feature set to obtain the trained extreme gradient boosting sub-model.

[0116] In a feasible implementation, since there are numerous XGBoost parameters, the present disclosure uses GridSearchCV in the sklearn toolkit to perform grid search on the XGBoost parameters. GridSearchCV grid search can achieve automatic parameter tuning and return the best parameter combination. For example, when a training model or fitting strategy is selected and a parameter list is given for selection, through grid search, it can automatically tune to the optimal and return the parameter combination and score. Finally, the parameters of XGBoost are determined as follows:

[0117] (1) learning_rate = 0.3;

[0118] (2) min_child_weight = 1;

[0119] (3) max_depth = 6;

[0120] (4) gamma = 0;

[0121] (5) alpha = 0;

[0122] (6) subsample = 1;

[0123] Among them, learning_rate is the learning rate range [0, 1], with a default value of 0.3. The smaller this parameter, the slower the calculation speed. The larger this parameter, the greater the possibility of non-convergence; min_child_weight is the minimum weight sum in each leaf, with a range of [0, +∞), and the default value is 1. The larger this parameter, the less likely to overfit; max_depth is the maximum depth of each tree, with a range of [0, +∞), and the default value is 6. The larger this parameter, the easier to overfit; gamma is the parameter that controls the number of leaves, with a range of [0, +∞), and the default value is 0. The larger this parameter, the less likely to overfit; alpha is the L1 regularization parameter, with a range of [0, +∞), and the default value is 0. The larger this parameter, the less likely to overfit. subsample is the sample sampling ratio, with a range of (0, 1], and the default value is 1. If it takes 0.5, it means randomly using 50% of the sample set for training.

[0124] After determining the parameters, the 6373-dimensional ComParE manual features of each speech and the labels corresponding to each speech are brought into the model for training to obtain a trained extreme gradient boosting sub-model.

[0125] Step 207: Obtain the development set, and extract the mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set of the development set.

[0126] In a feasible implementation manner, the method for extracting the mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set of the development set can refer to the above step 204, and the present disclosure will not elaborate here.

[0127] Step 208: According to the mel cepstral coefficient features, audio bag-of-words feature set of the development set, and the trained dual-channel neural network sub-model, obtain the first probability matrix of the development set. According to the ComParE feature set and the extreme gradient boosting sub-model, obtain the second probability matrix of the development set.

[0128] In a feasible implementation manner, since the development set includes 460 speech data, therefore, both the first probability matrix and the second probability matrix are matrices with a specification of [460 * 6].

[0129] Step 209: Fuse the first probability matrix of the development set and the second probability matrix of the development set into the comprehensive probability matrix of the development set, and determine the prediction result of the development set according to the comprehensive probability matrix of the development set.

[0130] In a feasible implementation manner, perform an L2 norm fusion on the first probability matrix predicted by the dual-channel neural network model and the second probability matrix predicted by the XGBoost model. The fusion method is as follows:

[0131] Assume is the first probability matrix predicted by the dual-channel neural network model, is the second probability matrix predicted by XGBoost, where j represents the j-th voice data in the development set.

[0132] The numbers at their corresponding positions are combined into a new two-dimensional vector i is the number of label categories. Then, the L2 norm of the new two-dimensional vector is calculated according to the following formula (1):

[0133]

[0134] where a ji represents a numerical value of the first probability matrix, b ji represents a numerical value of the second probability matrix, j represents the number of rows of the matrix, that is, the j-th voice data in the voice data to be recognized, i represents the number of columns of the matrix, and P ji is the predicted probability of the i-th class label of the j-th voice data.

[0135] The L2 norms of the new two-dimensional vectors are combined into a [460*6] matrix, which is the comprehensive probability matrix.

[0136] Step 210: Evaluate the performance of the trained voice emotion recognition model according to the prediction results of the development set.

[0137] In a feasible implementation, the comprehensive probability matrix is a matrix with a [460*6] specification. Among them, 6 numerical values in each row are the probabilities of 6 labels corresponding to a voice data in the development set. The maximum probability is selected from the 6 probabilities, and the label corresponding to the maximum probability is determined as the predicted label corresponding to the voice data. For example, the predicted probability vector of the first voice is The label sequence is [a sense of achievement, angry, afraid, pain, pleasure, surprise], then the predicted label is the position label corresponding to the maximum probability of 0.5, that is, the label predicted by this voice is pain.

[0138] In the embodiments of the present disclosure, a dual-channel deep neural network model that simultaneously uses Mel cepstrum coefficient features and manual features is adopted, and an XGBoost model is supplemented as a fusion solution, which solves the problem of insufficient expression of emotion features by a single feature set, and the problem of inaccurate classification of a single model in some emotion categories, thereby improving the accuracy of voice emotion recognition.

[0139] The following refers to the appendix Figure 6Describe a method for speech emotion recognition based on a speech emotion recognition model according to the present disclosure. This method can be implemented by an electronic device, which can include a terminal or a server. In the embodiments of the present disclosure, when recognizing the speech data to be recognized, the speech data to be recognized can be a single speech data or multiple speech data. When the speech data to be recognized is a single speech data, the probability matrix in the embodiments of the present disclosure is a [1*6] matrix, which can also be regarded as a vector composed of 6 elements. When the speech data to be recognized is multiple speech data, the probability matrix in the embodiments of the present disclosure is an [n*6] matrix, where n represents the number of speech data.

[0140] In the embodiments of the present disclosure, the speech emotion recognition model includes a dual-channel neural network sub-model and an extreme gradient boosting sub-model. The dual-channel neural network sub-model includes a first channel module, a second channel module, and a classification module. The method includes the following steps:

[0141] Step 301, obtain the speech data to be recognized.

[0142] Step 302, extract Mel-frequency cepstral coefficient features, ComParE feature sets, and audio bag-of-words feature sets according to the speech data to be recognized.

[0143] In a feasible implementation manner, the process of extracting features can refer to step 204 in the above embodiments, and the present disclosure will not elaborate here.

[0144] Step 303, input the Mel-frequency cepstral coefficient features into the first channel module of the dual-channel neural network sub-model of the speech emotion recognition model to obtain a 64-dimensional first feature vector.

[0145] Optionally, the first channel module includes two convolutional layers, two layers of LSTM networks, and a fully connected layer.

[0146] In a feasible implementation manner, step 303 can specifically include the following steps 3031-3032:

[0147] Step 3031, perform convolution on the input Mel-frequency cepstral coefficient features through two convolutional layers. After each convolution, batch normalization operation and rectified linear unit activation are performed, and max pooling layer is used for downsampling.

[0148] Step 3032, extract features from the downsampled feature data through two layers of LSTM networks, input the extracted data into the fully connected layer, and obtain a 64-dimensional first feature vector through the fully connected layer.

[0149] Step 304, input the audio bag-of-words feature sets into the second channel module of the dual-channel neural network sub-model of the speech emotion recognition model to obtain a 64-dimensional second feature vector.

[0150] Optionally, the second channel module may include two fully connected layers.

[0151] In a feasible implementation, the process of using two fully connected layers to extract features from the audio bag-of-words feature set may be as follows:

[0152] Input the audio bag-of-words feature set into two fully connected layers. After each pass through the fully connected layer, perform batch normalization operation and rectified linear unit activation, and then perform dropout processing to obtain a 64-dimensional second feature vector.

[0153] Step 305: Process the first feature vector and the second feature vector through the classification module to obtain a first probability matrix.

[0154] Optionally, the classification module may include a splicing sub-module, a fully connected layer, and a Softmax function.

[0155] In a feasible implementation, step 303 may specifically include the following steps 3051 - 3052:

[0156] Step 3051: Splice the first feature vector and the second feature vector through the splicing sub-module to obtain a 128-dimensional feature vector.

[0157] Step 3052: Input the 128-dimensional feature vector into the fully connected layer, and after activating the output of the fully connected layer using the Softmax function, obtain a first probability matrix.

[0158] Step 306: Obtain a second probability matrix according to the ComParE feature set and the extreme gradient boosting sub-model.

[0159] In a feasible implementation, input the Mel cepstral coefficient features of the speech data to be recognized into the audio bag-of-words feature set of the dual-channel neural network sub-model to obtain a first probability matrix, and input the ComParE feature set of the speech data to be recognized into the extreme gradient boosting sub-model to obtain a second probability matrix. The specifications of the first probability matrix and the second probability matrix are exactly the same.

[0160] Step 307: Fuse the first probability matrix and the second probability matrix into a comprehensive probability matrix.

[0161] In a feasible implementation, the first probability matrix and the second probability matrix can be fused with an L2 norm, and the corresponding operations can be as follows in steps 3071 - 3073:

[0162] Step 3071: Combine the two numerical values at the corresponding positions of the first probability matrix and the second probability matrix to form a new two-dimensional vector.

[0163] For example, assume that is the first probability matrix, is the second probability matrix. Then, the values at the corresponding positions of the two probability matrices are combined to form a new two-dimensional vector where j represents the j-th speech data in the speech data to be recognized, and i is the number of label categories.

[0164] Step 3072: Calculate the L2 norm of the new two-dimensional vector according to the following formula (1):

[0165]

[0166] where a ji represents a value of the first probability matrix, b ji represents a value of the second probability matrix, j represents the number of rows of the matrix, that is, the j-th speech data in the speech data to be recognized, i represents the number of columns of the matrix, and P ji is the predicted probability of the i-th class label of the j-th speech data.

[0167] Step 3073: Form a comprehensive probability matrix according to the L2 norm of the new two-dimensional vector.

[0168] In a feasible implementation, P ji is formed into a matrix according to the position of the subscript, which is the comprehensive probability matrix.

[0169] Step 308: Determine the prediction result of the speech data to be recognized according to the comprehensive probability matrix.

[0170] In a feasible implementation, the comprehensive probability matrix is a matrix of [n*6] specification, where n represents the number of speech data in the speech data to be recognized, and 6 values in each row are the probabilities of 6 labels corresponding to one speech data in the speech data to be recognized. The largest probability is selected from the 6 probabilities, and the label corresponding to the largest probability is determined as the predicted label corresponding to the speech data.

[0171] In the embodiments of the present disclosure, a dual-channel deep neural network model that simultaneously uses Mel cepstral coefficient features and handcrafted features is adopted, and the XGBoost model is supplemented as a fusion solution at the same time, which solves the problem of insufficient expression of emotion features by a single feature set and the problem of inaccurate classification of a single model in some emotion categories, thereby improving the accuracy of speech emotion recognition.

[0172] The embodiments of the present disclosure provide a speech emotion recognition device, which is used to implement the above speech emotion recognition method. As Figure 7Schematic block diagram of the voice emotion recognition device shown. The voice emotion recognition device 700 includes: an acquisition module 701, an extraction module 702, an input module 703, a processing module 704, and a determination module 705.

[0173] The acquisition module 701 is used to acquire the voice data to be recognized;

[0174] The extraction module 702 is used to extract Mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set according to the voice data to be recognized;

[0175] The input module 703 is used to input the Mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set into the trained voice emotion recognition model. The voice emotion recognition model includes a dual-channel neural network sub-model and an extreme gradient boosting sub-model;

[0176] The processing module 704 is used to obtain a first probability matrix according to the Mel cepstral coefficient features, audio bag-of-words feature set, and dual-channel neural network sub-model, and obtain a second probability matrix according to the ComParE feature set and extreme gradient boosting sub-model;

[0177] The determination module 705 is used to fuse the first probability matrix and the second probability matrix into a comprehensive probability matrix, and determine the prediction result of the voice data to be recognized according to the comprehensive probability matrix.

[0178] Optionally, the dual-channel neural network sub-model includes a first channel module, a second channel module, and a classification module;

[0179] The processing module 704 is further used for:

[0180] Input the Mel cepstral coefficient features into the first channel module to obtain a 64-dimensional first feature vector;

[0181] Input the audio bag-of-words feature set into the second channel module to obtain a 64-dimensional second feature vector;

[0182] Process the first feature vector and the second feature vector through the classification module to obtain a first probability matrix.

[0183] Optionally, the first channel module includes two convolutional layers, a two-layer LSTM network, and a fully connected layer;

[0184] The processing module 704 is further used for:

[0185] Convolve the input Mel cepstral coefficient features through two convolutional layers. After each convolution, batch normalization operation and rectified linear unit activation are performed, and max pooling layer is used for downsampling;

[0186] The downsampled feature data is subjected to feature extraction through a two - layer LSTM network, and the extracted data is input into a fully - connected layer to obtain a 64 - dimensional first feature vector through the fully - connected layer.

[0187] Optionally, the second channel module includes two fully - connected layers;

[0188] The processing module 704 is further configured to:

[0189] Input the audio bag - of - words feature set into two fully - connected layers. After passing through the fully - connected layer each time, batch normalization operation and rectified linear unit activation are performed, and then dropout processing is carried out to obtain a 64 - dimensional second feature vector.

[0190] Optionally, the classification module includes a splicing sub - module, a fully - connected layer, and a Softmax function;

[0191] The processing module 704 is further configured to:

[0192] Splice the first feature vector and the second feature vector through the splicing sub - module to obtain a 128 - dimensional feature vector;

[0193] Input the 128 - dimensional feature vector into the fully - connected layer, and after activating the output of the fully - connected layer using the Softmax function, obtain a first probability matrix.

[0194] Optionally, the determination module 705 is further configured to:

[0195] Form a new two - dimensional vector from two numerical values at corresponding positions of the first probability matrix and the second probability matrix;

[0196] Calculate the L2 norm of the new two - dimensional vector according to the following formula (1):

[0197]

[0198] where a ji represents a numerical value of the first probability matrix, b ji represents a numerical value of the second probability matrix, j represents the number of rows of the matrix, that is, the j - th speech data in the speech data to be recognized, i represents the number of columns of the matrix, and P ji is the predicted probability of the i - th class label of the j - th speech data;

[0199] Form a comprehensive probability matrix according to the L2 norm of the new two - dimensional vector.

[0200] Optionally, the device further includes a training module 706;

[0201] The training module 706 is used for:

[0202] Obtain an untrained initial speech emotion recognition model, where the initial speech emotion recognition model includes an initial dual-channel neural network sub-model and an initial extreme gradient boosting sub-model;

[0203] Obtain a training set and perform data augmentation on the training set;

[0204] According to the augmented training set, extract the Mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set of the training set;

[0205] According to the Mel cepstral coefficient features and the audio bag-of-words feature set of the training set, train the initial dual-channel neural network sub-model to obtain a trained dual-channel neural network sub-model;

[0206] Perform grid search on the parameters of the initial extreme gradient boosting sub-model through the GridSearchCV grid search algorithm to determine the initial parameters of the initial extreme gradient boosting sub-model, and train the initial extreme gradient boosting sub-model with the determined initial parameters according to the extracted ComParE feature set to obtain a trained extreme gradient boosting sub-model;

[0207] Obtain a development set and extract the Mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set of the development set;

[0208] According to the Mel cepstral coefficient features, audio bag-of-words feature set of the development set, and the trained dual-channel neural network sub-model, obtain the first probability matrix of the development set, and according to the ComParE feature set and the extreme gradient boosting sub-model, obtain the second probability matrix of the development set;

[0209] Fuse the first probability matrix and the second probability matrix of the development set into the comprehensive probability matrix of the development set, and determine the prediction result of the development set according to the comprehensive probability matrix of the development set;

[0210] Evaluate the performance of the trained speech emotion recognition model according to the prediction result of the development set.

[0211] In the embodiments of the present disclosure, a dual-channel deep neural network model that simultaneously uses Mel cepstral coefficient features and handcrafted features is adopted, and an XGBoost model is supplemented as a fusion solution, which solves the problem of insufficient expression of emotion features by a single feature set and the problem of inaccurate classification of a single model in some emotion categories, thereby improving the accuracy of speech emotion recognition.

[0212] Reference Figure 8, the structural block diagram of the electronic device 800 that can be used as a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0213] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 802 or the computer program loaded from the storage unit 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.

[0214] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device that can input information into the electronic device 800. The input unit 806 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 807 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include, but is not limited to, magnetic disks, optical disks. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0215] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above. For example, in some embodiments, the above-described method of voice emotion recognition can be implemented as a computer software program that is tangibly included in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. In some embodiments, the computing unit 801 can be configured to execute the above-described method of voice emotion recognition in any other suitable manner (e.g., by means of firmware).

[0216] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0217] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0218] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device that provides machine instructions and / or data to a programmable processor (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)), including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that provides machine instructions and / or data to a programmable processor.

[0219] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0220] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0221] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other.

Claims

1. A method for speech emotion recognition, characterized in that, Including: Obtain the speech data to be recognized; Extract Mel-frequency cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set according to the speech data to be recognized; Input the Mel-frequency cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set into the trained speech emotion recognition model, where the speech emotion recognition model includes a dual-channel neural network sub-model and an extreme gradient boosting sub-model; Obtain a first probability matrix according to the Mel-frequency cepstral coefficient features, audio bag-of-words feature set, and dual-channel neural network sub-model, and obtain a second probability matrix according to the ComParE feature set and extreme gradient boosting sub-model; Fuse the first probability matrix and the second probability matrix into a comprehensive probability matrix, and determine the prediction result of the speech data to be recognized according to the comprehensive probability matrix.

2. The speech emotion recognition method according to claim 1, characterized in that, The dual-channel neural network sub-model includes a first channel module, a second channel module, and a classification module; The obtaining a first probability matrix according to the Mel-frequency cepstral coefficient features, audio bag-of-words feature set, and dual-channel neural network sub-model includes: Input the Mel-frequency cepstral coefficient features into the first channel module to obtain a 64-dimensional first feature vector; Input the audio bag-of-words feature set into the second channel module to obtain a 64-dimensional second feature vector; Process the first feature vector and the second feature vector through the classification module to obtain a first probability matrix.

3. The speech emotion recognition method according to claim 2, characterized in that, The first channel module includes two convolutional layers, a two-layer LSTM network, and a fully connected layer; The inputting the Mel-frequency cepstral coefficient features into the first channel module to obtain a 64-dimensional first feature vector includes: Perform convolution on the input Mel-frequency cepstral coefficient features through two convolutional layers. After each convolution, batch normalization operation and rectified linear unit activation are performed, and max pooling layer is used for downsampling; Perform feature extraction on the downsampled feature data through a two-layer LSTM network, input the extracted data into the fully connected layer, and obtain a 64-dimensional first feature vector through the fully connected layer.

4. The speech emotion recognition method according to claim 2, wherein The second channel module includes two fully connected layers; The inputting the audio bag-of-words feature set into the second channel module to obtain a 64-dimensional second feature vector includes: Input the audio bag-of-words feature set into two fully connected layers. After passing through each fully connected layer, batch normalization operation and rectified linear unit activation are performed, and then dropout processing is performed to obtain a 64-dimensional second feature vector.

5. The voice emotion recognition method according to claim 2, characterized in that, The classification module includes a splicing sub-module, a fully connected layer, and a Softmax function; The processing the first feature vector and the second feature vector through the classification module to obtain a first probability matrix includes: Splice the first feature vector and the second feature vector through the splicing sub-module to obtain a 128-dimensional feature vector; Input the 128-dimensional feature vector into the fully connected layer, and after activating the output of the fully connected layer with the Softmax function, obtain a first probability matrix.

6. The voice emotion recognition method according to claim 1, characterized in that, The fusing the first probability matrix and the second probability matrix into a comprehensive probability matrix includes: Form a new two-dimensional vector by combining the two numerical values at the corresponding positions of the first probability matrix and the second probability matrix; Calculate the L2 norm of the new two-dimensional vector according to the following formula (1): Among them, a ji represents a value of the first probability matrix, b ji represents a value of the second probability matrix, j represents the number of rows of the matrix, that is, the j-th voice data in the voice data to be recognized, i represents the number of columns of the matrix, and P ji is the predicted probability of the i-th class label of the j-th voice data; Form a comprehensive probability matrix based on the L2 norm of the new two-dimensional vector.

7. The voice emotion recognition method according to claim 1, wherein The training process of the speech emotion recognition model includes: Obtain an untrained initial speech emotion recognition model, where the initial speech emotion recognition model includes an initial dual-channel neural network sub-model and an initial extreme gradient boosting sub-model; Obtain a training set and perform data augmentation on the training set; Extract the Mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set of the training set according to the augmented training set; Train the initial dual-channel neural network sub-model according to the Mel cepstral coefficient features and audio bag-of-words feature set of the training set to obtain a trained dual-channel neural network sub-model; Perform grid search on the parameters of the initial extreme gradient boosting sub-model through the GridSearchCV grid search algorithm to determine the initial parameters of the initial extreme gradient boosting sub-model, and train the initial extreme gradient boosting sub-model with the determined initial parameters according to the extracted ComParE feature set to obtain a trained extreme gradient boosting sub-model; Obtain a development set and extract the Mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set of the development set; Obtain the first probability matrix of the development set according to the Mel cepstral coefficient features, audio bag-of-words feature set, and trained dual-channel neural network sub-model of the development set, and obtain the second probability matrix of the development set according to the ComParE feature set and extreme gradient boosting sub-model; Fuse the first probability matrix and the second probability matrix of the development set into the comprehensive probability matrix of the development set, and determine the prediction result of the development set according to the comprehensive probability matrix of the development set; Evaluate the performance of the trained speech emotion recognition model according to the prediction result of the development set.

8. A speech emotion recognition device, comprising: An acquisition module for acquiring speech data to be recognized; An extraction module for extracting Mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set according to the speech data to be recognized; An input module for inputting the Mel cepstral coefficient features, ComParE feature set, and audio bag-of-words feature set into a trained speech emotion recognition model, where the speech emotion recognition model includes a dual-channel neural network sub-model and an extreme gradient boosting sub-model; A processing module for obtaining a first probability matrix according to the Mel cepstral coefficient features, audio bag-of-words feature set, and dual-channel neural network sub-model, and obtaining a second probability matrix according to the ComParE feature set and extreme gradient boosting sub-model; A determination module for fusing the first probability matrix and the second probability matrix into a comprehensive probability matrix, and determining the prediction result of the speech data to be recognized according to the comprehensive probability matrix.

9. An electronic device, comprising: A processor; And A memory for storing programs Wherein, the program includes instructions which, when executed by the processor, cause the processor to execute the method according to any one of claims 1-7.

10. A computer program product, comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Voice emotion recognition model and method based on joint feature representation

    CN108899051A

  • Voice emotion recognizing method and device, computer device and storage medium

    CN110021308A