A Cross-Database Speech Emotion Recognition Method Based on Multi-Task Learning and Sub-Domain Adaptation

Through multi-task learning and subdomain adaptive methods, the feature distribution of cross-border speech emotion recognition is aligned with the deep autoencoder and the local maximum mean error algorithm, which solves the feature difference problem caused by training and test data from different corpus, and improves the performance of cross-border speech emotion recognition.

CN113870900BActive Publication Date: 2025-07-25HENAN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111125098.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-25
Publication Date
2025-07-25
Estimated Expiration
2041-09-25

AI Technical Summary

Technical Problem

In real-life application scenarios, training data and test data come from different corpuses, resulting in different feature distribution differences, and cross-border speech emotion recognition problems that affect model recognition performance.

Method used

Multitask learning and subdomain adaptation methods are adopted to compress feature redundancy information through a deep autoencoder, and aligned emotional and gender subdomain feature distributions using a local maximum mean error algorithm, combined with cross-entropy optimization training network to achieve unsupervised domain adaptation emotion classification.

Benefits of technology

The performance of cross-border speech emotion recognition is improved, the difference in feature distribution between different corpus is reduced, and the generalization ability of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
  • Figure SMS_4
    Figure SMS_4
Patent Text Reader

Abstract

The present invention proposes a cross-database speech emotion recognition method based on multi-task learning and sub-domain adaptation. The present invention includes the following steps: First, the high-dimensional speech features extracted from the source domain and the target domain are respectively input into a deep autoencoder network to compress the redundant information of the features and obtain low-dimensional emotion features; Then, a sub-domain adaptation algorithm is used to divide the low-dimensional feature space into an emotion sub-domain feature space and a gender sub-domain feature space respectively, so as to reduce the feature distribution distance; Finally, emotion recognition is used as the main task and gender recognition is used as the auxiliary task to learn more common emotion information. The method proposed by the present invention can effectively improve the performance of cross-database speech emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technical Field

[0002] The present invention belongs to the technical field of speech signal processing, and particularly relates to a cross-corpus speech emotion recognition method based on multi-task learning and sub-domain adaptation. Background Art

[0003] Speech emotion recognition is an important part of emotion computing and also an important research direction in the field of artificial intelligence. Speech emotion recognition is to convert the human speech emotion signal into a digital signal through a computer, and through the learning of the computer, enable it to have the ability to recognize human speech emotions. Since in real application scenarios, it is difficult to ensure that the training data and the test data come from the same corpus, this has caused a great difference in the feature distributions of the training and test data, seriously affecting the recognition performance of the model.

[0004] Inspired by the successful application of transfer learning and multi-task learning in the field of speech emotion recognition, sub-domain adaptation is introduced in the research of cross-corpus speech emotion recognition to reduce the difference in feature distributions between different domains, and multi-task learning is used to improve the generalization ability of the model.

[0005] Therefore, the present invention mainly focuses on cross-corpus speech emotion recognition between different corpora. In the low-dimensional emotion feature space, a sub-domain adaptation algorithm is used to reduce the feature distribution distance. It should be noted that the present invention simultaneously reduces the feature distribution distance in both the emotion sub-domain feature space and the gender sub-domain feature space to improve the cross-corpus speech emotion recognition performance. Summary of the Invention

[0006] In order to learn more identical speech emotion information in the source domain and the target domain and achieve unsupervised domain adaptation emotion classification, a cross-corpus speech emotion recognition method based on multi-task learning and sub-domain adaptation is proposed. The specific steps are as follows:

[0007] (1) Feature preprocessing: First, select the data with the same emotion categories in the source domain corpus and the target domain corpus as the training set and the test set respectively, and then extract their acoustic features and perform normalization processing on them;

[0008] (2) Feature processing: Input the source domain and target domain features obtained after normalization in step (1) into a deep autoencoder respectively to compress the redundant feature information and obtain low-dimensional emotion features with strong representativeness. Assume that the input of the deep autoencoder is X and the decoded output is Then the reconstruction loss of the deep autoencoder is as follows:

[0009]

[0010]

[0011] In this way, the emotional representation of the source domain and the target domain in the low-dimensional space is obtained; at the same time, the real emotional label and gender label of the source domain are used as cross entropy to optimize the division of the subdomain space. The cross entropy is calculated as follows:

[0012]

[0013] in is the predicted probability;

[0014] (3) Subdomain feature distribution alignment: The local maximum mean discrepancy (LMMD) is used to divide the subdomain feature space into the emotion subdomain feature space and the gender subdomain feature space. The emotion subdomain feature distribution alignment algorithm is expressed as:

[0015]

[0016] in Encode the low-dimensional features of the source domain output by the deep autoencoder The weight of each feature belonging to sentiment category c, Encode the output low-dimensional features of the target domain for the deep autoencoder The weight of each feature in the emotion category c. At the same time, the attribute feature of emotion, that is, gender, is aligned, and the distribution of gender subdomain features is aligned as follows:

[0017]

[0018] in is the low-dimensional feature of the source domain The weight of each feature belonging to gender category a, The target domain sample The weight of each feature belonging to gender category a;

[0019] (4) Training model: The entire network training is continuously optimized by the Adam optimizer. The cross entropy of the source domain’s sentiment label and gender label is calculated to optimize the accurate division of the subdomain space in step (3). The loss function of the entire network is expressed as:

[0020]

[0021] in and They are the reconstruction loss of the deep autoencoder, and They are the emotional information cross entropy loss and gender information cross entropy loss of the source domain features, and They are the emotion subdomain feature distribution distance and gender subdomain feature distribution distance based on LMMD respectively;

[0022] (6) Repeat steps (2) and (3), and iteratively train the network model by the gradient descent method, continuously reducing the loss function in step (5) until the model is optimal;

[0023] (7) Use the network model trained in step (6) and the sofmatx classifier to identify the target domain features without noise in step (1), and finally achieve the emotion recognition of speech emotion under the condition of cross-corpus. Description of the Drawings

[0024] As shown in the drawings, Figure 1 It is a flowchart of a cross-corpus speech emotion recognition method based on multi-task learning and sub-domain adaptation. Detailed Implementation Manner

[0025] The present invention will be further described below in conjunction with the detailed implementation manner.

[0026] (1) Select two speech emotion corpora, EMO-DB and CASIA, as the source domain database and the target domain database respectively;

[0027] (2) Select the common emotion speech of the above two corpora as the data set. Specifically, select 4 types of emotion speech (anger, fear, happiness, sadness) in the EMO-DB database, with a total of 327 speeches; 4 types of emotion speech (anger, fear, happiness, sadness) in the CASIA Chinese speech emotion database, with a total of 800 speech data. Use the open source toolkit Opensmile to extract the standard feature set of the 2010 International Speech Emotion Recognition Challenge, and the features extracted from each speech are 1582-dimensional;

[0028] (3) Make the emotion labels and gender labels of the source domain speech signals, and the labels are represented as one-hot vectors;

[0029] (4) Based on the deep autoencoder, the number of hidden layers is 3, and the hidden layer neuron nodes are set to 1200, 500, and 1200 respectively. The activation function in the encoding stage uses the LeakyReLU function, and the activation function in the decoding stage uses the ReLU function;

[0030] (5) After preprocessing the features of the source domain and target domain data sets obtained in (2), input them into the deep dynamic encoder to extract low-dimensional emotion features;

[0031] (6) In the low-dimensional emotion space, the distribution distance of the low-dimensional emotion features between the source domain and the target domain is measured based on the emotion sub-domain adaptive algorithm and the gender sub-domain adaptive algorithm of LMMD. When calculating the sub-domain adaptive loss function based on LMMD, the low-dimensional emotion features of the source domain, as well as the low-dimensional gender features and corresponding true labels of the source domain, are required; however, the emotion labels and gender labels of the target domain need to generate pseudo-labels using the probability distribution calculated by softmax;

[0032] (7) The loss function of the entire network is expressed as:

[0033]

[0034] Where and are the reconstruction losses of the deep autoencoder respectively, and are the cross-entropy losses of the emotion information and gender information of the source domain features respectively, and are the distribution distances of the emotion sub-domain features and gender sub-domain features based on LMMD respectively;

[0035] (8) The learning rate and batch size of the model are both set to 0.00001 and 100. The network model is trained using the Adam gradient descent method. The model is iteratively trained 800 times, and the classifier uses softmax. In the emotion sub-domain adaptation and gender sub-domain adaptation, the feature mapping function of LMMD uses the multi-kernel Gaussian function, and the number of Gaussian kernels is set to 5 and 2 respectively;

[0036] (9) Normalize the target domain speech signal to be recognized and input it into the deep autoencoder trained in step (8). Use the softmax classifier to output the category with the highest probability, which is the recognized target domain emotion category;

[0037] The scope of protection requested by the present invention is not limited to the description of this specific embodiment.

Claims

1. A cross-database speech emotion recognition method based on multi-task learning and sub-domain adaptation, characterized in that, It includes the following steps: (1) Feature preprocessing: First, select data with the same sentiment category from the source domain corpus and the target domain corpus as the training set and the test set respectively. Then extract their acoustic features and perform normalization processing on them; (2) Feature processing: The source domain and target domain features obtained after normalization in step (1) are respectively input into a deep autoencoder to compress redundant feature information and obtain low-dimensional sentiment features with strong representational power. Assuming the input of the deep autoencoder is X and the decoded output is Then the reconstruction loss of the deep autoencoder is as follows: Thereby obtaining the sentiment representations of the source domain and the target domain in the low-dimensional space; at the same time, use the true sentiment labels and gender labels of the source domain as cross-entropy to optimize the division of the sub-domain space. The cross-entropy calculation is as follows: wherein is the predicted probability; (3) Sub-domain feature distribution alignment: Use the local maximum mean discrepancy (LMMD) to divide the sub-domain feature space into the sentiment sub-domain feature space and the gender sub-domain feature space respectively. The sentiment sub-domain feature distribution alignment algorithm is expressed as: where is the low-dimensional source domain feature output by the deep autoencoder encoding The weight of each feature belonging to the emotion category c in is the low-dimensional target domain feature output by the deep autoencoder encoding The weight of each feature belonging to the emotion category c in, and at the same time align the attribute features of the emotion, i.e., gender. The gender sub-domain feature distribution alignment is as follows: where is the low-dimensional feature of the source domain and the weight of each feature belonging to gender category a in it, is the target domain sample and the weight of each feature belonging to gender category a in it; (4) Training the model: The entire network training is continuously optimized through the Adam optimizer. Calculate the cross-entropy from the sentiment labels and gender labels of the source domain respectively to optimize the accurate division of the sub-domain space in step (3). The loss function of the entire network is expressed as: L sum = L RS + L RT + L E + L G + L emotion + L gender (6) where L RS and L RT are the reconstruction losses of the deep autoencoder, L E and L G are the cross-entropy losses of the sentiment information and gender information of the source domain features respectively, L emotion and L gender are the distribution distances of the sentiment sub-domain features and gender sub-domain features based on LMMD respectively; (6) Repeat steps (2) and (3), and iteratively train the network model through the gradient descent method, continuously reducing the loss function in step (5) until the model is optimal; (7) Use the network model trained in step (6), and use the sofmatx classifier to identify the target domain features without noise in step (1), and finally achieve the sentiment recognition of speech sentiment under the condition of cross-corpus.

Citation Information

Patent Citations

  • High-performance anti-noise speech emotion recognition method based on deep neural network

    CN112927723A

  • Sub-domain self-adaptive cross-library speech emotion recognition method based on depth auto-encoder

    CN113077823A