A cross-database speech emotion recognition method based on collaborative boundary-aware adversarial learning
By constructing a cross-database speech emotion recognition method based on collaborative boundary-aware adversarial learning and utilizing decision boundary alignment network and feature distribution alignment network, the problem of decision boundary misalignment in cross-database speech emotion recognition is solved, thereby improving the recognition accuracy and generalization ability of the model.
Patent Information
- Application Number
- CN202511032472.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing cross-corpus speech emotion recognition methods face data differences between different corpora, especially the decision boundary misalignment problem caused by differences in age, language and expression, which leads to decreased model performance and difficulty in effectively identifying emotion categories.
A collaborative boundary-aware adversarial learning method is adopted to reduce the feature distribution differences and decision boundary differences between the source and target domains by constructing a decision boundary alignment network and a feature distribution alignment network. A multi-stage training strategy, including sentiment classification loss, domain classification loss and gradient reversal mechanism, is adopted to optimize model parameters to adapt to different corpora.
The accuracy of cross-database speech emotion recognition and the generalization ability of the model are improved, and it can better identify target domain speech samples near the decision boundary, thereby enhancing the stability and recognition performance of the model.
Smart Images

Figure CN120526808B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech emotion recognition, and in particular to a cross-database speech emotion recognition method based on collaborative boundary-aware adversarial learning. Background Art
[0002] Speech is the most direct and natural form of communication in daily life, containing rich semantic and emotional information. By analyzing speech characteristics such as speech rate, pitch, and volume, listeners can more accurately perceive the speaker's emotional state. Speech emotion recognition technology aims to use computers to analyze the emotional information in speech signals, thereby improving the intelligence and effectiveness of human-computer interaction. Currently, speech emotion recognition technology has been applied in various fields, including medical diagnosis, intelligent transportation, and intelligent customer service.
[0003] While speech emotion recognition technology has made significant progress, practical applications still face numerous challenges. Existing methods are typically designed and evaluated based on the same corpus. However, due to differences in acquisition equipment and recording environments, the distribution of training and test data often differs significantly, which can significantly degrade model performance.
[0004] To address this issue, researchers have shifted their focus to cross-corpus speech emotion recognition methods. Unlike traditional speech emotion recognition methods, cross-corpus speech emotion recognition methods use training and test data from different corpora. By integrating data from multiple corpora, cross-corpus speech emotion recognition can effectively address data discrepancies and significantly improve the model's generalization capabilities.
[0005] In recent years, researchers have conducted extensive work to reduce the disparity in feature distributions between corpora, but the problem of decision boundary misalignment caused by differences in age, language, and expression remains unresolved. Specifically, existing methods focus on extracting corpus-invariant features while neglecting the alignment of task-specific decision boundaries between different emotion categories. This results in target domain speech samples near the decision boundary being easily misclassified. Summary of the Invention
[0006] Purpose of the invention: The technical problem to be solved by the present invention is to address the deficiencies of the existing technology and provide a cross-database speech emotion recognition method based on collaborative boundary-aware adversarial learning, comprising the following steps:
[0007] Step 1: Obtain source and target domain speech samples, perform preprocessing, including amplitude normalization, pre-emphasis filtering, frame segmentation and windowing, and convert the preprocessed speech samples into Mel spectrogram features;
[0008] Step 2: Build a cross-database speech emotion recognition model based on collaborative boundary-aware adversarial learning to reduce the differences between source and target domain speech samples from the dual dimensions of feature distribution and decision boundary, thereby achieving emotional knowledge transfer.
[0009] Step 3: adopt a multi-stage training strategy, which includes first-stage training, second-stage training, and third-stage training. A maximum-minimization difference detection loss function is used to obtain all target domain speech samples with incorrect mapping of emotion labels due to decision boundary differences, and adjust the decision boundary between categories.
[0010] Step 1 includes:
[0011] Step 1-1: Randomly select two databases from the public speech emotion database with significant heterogeneity to form the source domain corpus and the target domain corpus , to simulate the actual multi-source heterogeneous data source scenario;
[0012] A label consistency screening method is used to retain only speech samples with the same emotion category, ensuring that the source domain and the target domain share a unified emotion label space.
[0013] Step 1-2: source domain speech samples and target domain speech samples Perform amplitude normalization, pre-emphasis filtering, frame division and windowing processing;
[0014] Steps 1-3: Select a supervised source domain corpus and unsupervised target domain corpus Conduct collaborative training;
[0015] Source domain corpus With emotional labels, for Corresponding sentiment labels; target domain corpus No emotional labels, Simulate actual speech samples without emotion labels.
[0016] Steps 1-2 include: using maximum normalization to scale the sample amplitude to [-1, 1] to reduce the impact of differences in the acquisition environment on the sample amplitude; using a Gaussian filter to enhance the high-frequency information of the speech and reduce background noise; using a Hanning window with a frame length of 25ms and a frame shift of 10ms to divide the frame and add windows to reduce spectral leakage; then, converting to a Mel frequency scale to retain emotional information and remove redundant features to achieve consistency between the source domain and target domain speech sample features, and converting to a Mel frequency scale that conforms to the auditory characteristics of the human ear to ensure the consistency of heterogeneous data in the source and target domains and the stability of emotional features.
[0017] Step 2 includes:
[0018] Step 2-1: The source domain speech sample and target domain speech samples The Mel spectrogram is input into the shared feature extraction network to extract the source domain deep speech features with strong migration ability and discriminativeness. and target domain deep speech features ; The shared feature extraction network uses the VGG11 network as the backbone network;
[0019] The VGG11 network is structured and parameterized, removing the last three fully connected layers and retaining only the convolutional and pooling layers.
[0020] Step 2-2, build a decision boundary alignment network: use two sentiment classifiers with the same structure but different initialization parameters to classify the source domain deep speech features extracted in step 2-1 Predict emotion categories from different perspectives;
[0021] Both sentiment classifiers are multi-layer perceptron structures, denoted as the first sentiment classifier and the second sentiment classifier respectively. The weight matrices and bias vectors of the two sentiment classifiers are initialized through Gaussian distribution, but different random seeds are used to generate parameters, thereby forming a sentiment decision boundary with discriminative diversity in the feature space.
[0022] The first emotion classifier is used to classify the source domain speech samples The sentiment label prediction results are:
[0023] ,
[0024] in, The source domain sentiment category prediction result generated by the first sentiment classifier, is the Softmax function used to convert the sentiment category score into a probability distribution, is the weight matrix of the first sentiment classifier, is the ReLU activation function used to enhance emotional features and suppress irrelevant features. It is a shared feature extraction network composed of multiple convolutional layers. is the bias vector of the first sentiment classifier;
[0025] The second emotion classifier is used to classify the source domain speech samples The sentiment label prediction results are:
[0026] ,
[0027] in, The source domain sentiment category prediction results generated by the second sentiment classifier, is the weight matrix of the second sentiment classifier, is the bias vector of the second sentiment classifier;
[0028] Step 2-3: Through joint supervision, the source domain speech samples True emotional label The prediction results of the two classifiers and Conduct joint optimization;
[0029] Calculate sentiment label prediction results and and real emotion labels The difference between them, the sentiment classification loss function is defined as:
[0030] ,
[0031] in, is the sentiment classification loss function, is the number of source domain speech samples, is the true emotion label of the i-th source domain speech sample, The emotion category prediction result of the i-th source domain speech sample generated by the first emotion classifier, The emotion category prediction result of the i-th source domain speech sample generated by the second emotion classifier;
[0032] Step 2-4: The target domain deep speech features extracted in step 2-1 are Input the result to the decision boundary alignment network constructed in step 2-2, and use two emotion classifiers to predict the emotion category of the target domain speech sample.
[0033] The first emotion classifier is used to classify the target domain speech samples The sentiment label prediction results are:
[0034] ,
[0035] in, The target domain sentiment category prediction result generated by the first sentiment classifier;
[0036] The second emotion classifier is used to classify the target domain speech samples The sentiment label prediction results are:
[0037] ,
[0038] in, The target domain sentiment category prediction results generated by the second sentiment classifier;
[0039] Step 2-5: Calculate the prediction probability of the two sentiment classifiers on the target domain samples and Difference, construct difference detection loss function;
[0040] Step 2-6: Build a feature distribution alignment network to extract the source domain deep speech features from step 2-1. and target domain deep speech features Input into the feature distribution alignment network to reduce the difference in feature distribution between the source domain and the target domain;
[0041] Step 2-7, the domain category prediction result output by the feature distribution alignment network in step 2-6 On this basis, a domain label supervision mechanism is introduced to convert speech samples The real domain label and domain category prediction results Perform joint optimization and construct domain classification loss function.
[0042] Steps 2-5 include: the target domain emotion category prediction result generated by the first emotion classifier Expands to:
[0043] ,
[0044] in, The probability of the first emotion classifier predicting the first emotion of the target domain sample, The first sentiment classifier predicts the target domain sample The probability of class emotion, is the number of common sentiment categories between the source domain and the target domain; in the cross-database task The value of is 5;
[0045] The target domain sentiment category prediction results generated by the second sentiment classifier Expands to:
[0046] ,
[0047] in, The probability of the first category of emotion of the target domain sample is predicted by the second emotion classifier, Predict the target domain sample for the second sentiment classifier Probability of class emotion;
[0048] Calculate the difference in the prediction results of the two sentiment classifiers and define the difference detection loss function as:
[0049] ,
[0050] in, is the difference detection loss function, is the number of target domain speech samples, To take the absolute value; is the predicted probability distribution of the first emotion classifier for the nth category emotion of the mth target domain speech sample, is the predicted probability distribution of the nth category emotion of the mth target domain speech sample by the second emotion classifier.
[0051] Steps 2-6 include: the feature distribution alignment network includes two or more fully connected layers and nonlinear activation functions. By introducing a gradient reversal layer (GRL), the shared feature extraction network is guided to learn domain-invariant features to achieve the unification of the feature space between the source domain and the target domain. The domain category prediction result is:
[0052] ,
[0053] in, Represents the domain category prediction result of the generated input sample features, is the Sigmoid function used to convert the domain category score into a probability distribution, is the weight matrix of the feature distribution alignment network, is the ReLU activation function used to enhance domain features and suppress irrelevant features. are the input source domain and target domain speech samples, is the bias vector of the feature distribution alignment network.
[0054] Steps 2-7 include: Field category prediction results Expands to:
[0055] ,
[0056] in, is the predicted probability that the domain label is the target domain, is the predicted probability that the domain label is the source domain;
[0057] The real domain label d is expanded to:
[0058] ,
[0059] in, is the true probability that the domain label is the target domain, is the true probability that the domain label is the source domain, is the kth speech sample input, when From the source domain corpus ,but ;when From the target domain corpus ,but ;
[0060] Calculate the difference between the domain category prediction result and the true domain label, and define the domain classification loss function as:
[0061] ,
[0062] in, is the domain classification loss function, is the true domain label of the k-th speech sample, The domain category prediction result of the k-th speech sample.
[0063] Step 3 includes:
[0064] In step 3-1, during training, a stochastic gradient descent (SGD) optimizer is used with a momentum factor of 0.9 to accelerate model convergence and reduce gradient fluctuations, ensuring cross-database training stability. Initialize the parameters of the shared feature extraction network, decision boundary alignment network, and feature distribution alignment network. Use an exponentially decaying acoustic perception learning rate strategy to match the high-frequency attenuation of the human ear's Mel-scale. This strategy focuses on fundamental frequency learning in the early stages of training and gradually optimizes high-frequency resonance peaks in the later stages, thereby improving model recognition performance. The learning rate formula is:
[0065] ,
[0066] in, is the adjusted learning rate, is the initial learning rate, is the current number of training times, is the total number of training sessions;
[0067] Step 3-2: During the first phase of training, freeze the parameters of the first sentiment classifier of the decision boundary alignment network (weight matrix and the bias vector ) and the parameters of the second sentiment classifier (weight matrix and the bias vector ), retaining the difference in the initial decision perspective; dynamically adjusting the feature optimization direction through the gradient reversal mechanism, and updating the parameters of the feature alignment network (weight matrix and the bias vector ), gradient reversal coefficient and shared feature extraction network parameters ;
[0068] First, according to the domain classification loss function , update the parameters of the feature alignment network to eliminate the difference in feature distribution between corpora:
[0069] ,
[0070] ,
[0071] in, The weight matrix of the updated feature distribution alignment network, The bias vector of the network is aligned to the updated feature distribution; the symbol Indicates that the update is performed using the formula on the right; represents partial differential;
[0072] Then, dynamically adjust the backpropagation layer Gradient reversal coefficient , gradually strengthen the learning of target domain features and balance the migration capabilities of cross-domain features:
[0073] ,
[0074] in, is the ratio of the current training times to the total training times; e is a natural constant;
[0075] Finally, according to the sentiment classification loss function , domain classification loss function and the gradient reversal coefficient , update the parameters of the shared feature extraction network :
[0076] ,
[0077] in, is the weight that balances the sentiment classification loss function and the domain classification loss function, Extract network parameters for the updated deep features;
[0078] Step 3-3, during the second phase of training, freeze the parameters of the shared feature extraction network ;
[0079] Maximize the difference in prediction results between the two sentiment classifiers and update the parameters of the first sentiment classifier of the decision boundary alignment network (weight matrix and the bias vector ) and the parameters of the second sentiment classification (weight matrix and the bias vector );
[0080] Based on the constructed sentiment classification loss function and difference detection loss function , update the decision boundary alignment network parameters and filter out the target domain speech samples with incorrect mapping of emotion labels due to the misalignment of the decision boundary of cross-database speech emotion recognition:
[0081] , ,
[0082] , ,
[0083] in, is the weight that balances the sentiment classification loss function and the difference detection loss function, is the updated weight matrix of the first sentiment classifier, is the updated bias vector of the first sentiment classifier, is the updated weight matrix of the second sentiment classifier, is the updated bias vector of the second sentiment classifier;
[0084] Step 3-4, during the third stage of training, freeze the parameters of the first sentiment classifier of the decision boundary alignment network (weight matrix and the bias vector ) and the parameters of the second sentiment classifier (weight matrix and the bias vector ), to maintain the stability of the decision boundary, avoid excessive updates to the classifier parameters, and allow the model to focus on the feature extraction process; minimize the difference in prediction results between the two sentiment classifiers and update the shared feature extraction network parameters ;
[0085] Difference detection loss function based on construction , update the shared feature extraction network parameters , correct the incorrect sentiment label mapping caused by the misalignment of the decision boundary in the target domain:
[0086] ;
[0087] Step 3-5, step 3-2 ~ step 3-4 are three stages of training, which are performed alternately. The model parameters are gradually optimized and adjusted until the model reaches a convergence state or the number of training times reaches the preset maximum value.
[0088] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the method.
[0089] The present invention also provides a storage medium storing a computer program or instruction, which executes the steps of the method when the computer program or instruction is run on a computer.
[0090] To address the problem of decision boundary misalignment caused by differences in age, language, expression, etc. in cross-database speech emotion recognition, this paper designs a decision boundary alignment network to match the decision boundaries of the source domain and the target domain, thereby reducing the incorrect mapping of target sample emotion categories near the boundary.
[0091] To address the problem that the decision boundary alignment network is heavily dependent on the source domain sentiment classifier, this paper introduces a feature distribution alignment network to obtain corpus invariance features, thereby reducing the performance degradation of the decision boundary alignment network caused by excessive differences in domain feature distribution.
[0092] The present invention has the following beneficial effects: 1. The present invention adopts a decision boundary alignment network and a feature distribution alignment network, which reduces the feature distribution differences between domains and innovatively solves the decision boundary misalignment problem in cross-database speech emotion recognition.
[0093] 2. The present invention considers the collaborative confrontation between the decision boundary alignment network and the feature distribution alignment network, quantifies the contribution of different adversarial migration methods, so as to cope with the distribution of new speech samples and enhance the generalization ability of the model.
[0094] 3. The present invention constructs a cross-database speech emotion recognition method based on collaborative boundary-aware adversarial learning to obtain speech emotion features that are more suitable for the target domain, thereby improving the recognition accuracy of cross-database speech emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1 It is a flow chart of the method of the present invention.
[0096] Figure 2 It is a schematic diagram of the method of the present invention.
[0097] Figure 3 It is the statistical data of the subtasks in the cross-database speech emotion recognition experiment provided by the present invention. DETAILED DESCRIPTION
[0098] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more apparent.
[0099] like Figure 1 、 Figure 2 As shown, an embodiment of the present invention provides a cross-database speech emotion recognition method based on collaborative boundary-aware adversarial learning, comprising the following steps:
[0100] Step 1: Obtain source and target domain speech samples, perform preprocessing such as amplitude normalization, pre-emphasis filtering, frame segmentation and windowing, and convert the preprocessed speech samples into Mel spectrogram features. This includes the following steps:
[0101] Step 1-1: Randomly select two databases from multiple public speech emotion databases with significant heterogeneity to form the source domain corpus and the target domain corpus , to simulate actual multi-source heterogeneous data source scenarios. A label consistency screening method is used to retain only speech samples with the same emotion category, thereby avoiding decision boundary misalignment caused by category shift. This ensures that the source and target domains share a unified emotion label space, resolving the issue of label system incompatibility in cross-database speech emotion recognition.
[0102] Step 1-2: source domain speech samples and target domain speech samples Amplitude normalization, pre-emphasis filtering, and frame windowing are performed. Specific methods include: using maximum normalization to scale the sample amplitude to [-1, 1] to reduce the impact of acquisition environment variations on sample amplitude; applying a Gaussian filter to enhance high-frequency information in speech and reduce background noise; and using a Hanning window with a 25ms frame length and a 10ms frame shift to reduce spectral leakage. Subsequently, the data is converted to a Mel frequency scale to preserve emotional information and remove redundant features, achieving consistency between the source and target domain speech sample features. Furthermore, the data is converted to a Mel frequency scale that conforms to human hearing, ensuring consistency between heterogeneous source and target domain data and stability of emotional features.
[0103] Steps 1-3: Considering the large number of speech samples without emotion labels in reality, a supervised source domain corpus is selected. and unsupervised target domain corpus Conduct collaborative training to improve the model's cross-database generalization ability. Source domain corpus With emotional labels, is the source domain speech sample, To correspond to the emotional label, guide the model to obtain accurate emotional knowledge; target domain corpus No emotional labels, For the target domain speech samples, simulate the actual speech samples without emotion labels.
[0104] Step 2: Build a cross-database speech emotion recognition model based on collaborative boundary-aware adversarial learning to reduce the differences between source and target domain speech samples from the dual dimensions of feature distribution and decision boundary, and achieve accurate transfer of emotional knowledge.
[0105] Step 2-1: The source domain speech sample and target domain speech samples The Mel spectrogram is input into the shared feature extraction network (VGG11 network is the backbone network) to extract the source domain deep speech features with strong migration ability and discriminativeness. and target domain deep speech features To address the distribution inconsistency issue in cross-corpus speech emotion recognition, the VGG11 network was structurally pruned and parameterized, removing the last three fully connected layers and retaining only the convolutional and pooling layers, thereby reducing the risk of overfitting in heterogeneous corpora.
[0106] Step 2-2: Build a decision boundary alignment network and use two sentiment classifiers with the same structure but different initialization parameters to classify the source domain deep speech features extracted in step 2-1. Predict emotion categories from different perspectives. Both emotion classifiers use a multi-layer perceptron architecture, with their weight matrices and bias vectors initialized from a Gaussian distribution. However, different random seeds are used to generate parameters, thereby forming a diverse emotion decision boundary in the feature space. This mechanism allows the two classifiers to produce different predictions even for the same input, effectively identifying uncertain samples near the emotion boundary and alleviating the boundary underfitting problem caused by single-perspective discrimination.
[0107] The first emotion classifier is used to classify the source domain speech samples The sentiment label prediction results are:
[0108] ,
[0109] The second emotion classifier is used to classify the source domain speech samples The sentiment label prediction results are:
[0110] ;
[0111] Step 2-3: Through joint supervision, the source domain speech samples True emotional label The prediction results of the two classifiers and Joint optimization is performed to achieve accurate emotional knowledge learning. The two differentiated sentiment classifiers can converge to consistent but complementary decision boundaries during training, thereby improving the stability and generalization ability of sentiment discrimination.
[0112] Calculate sentiment label prediction results and and real emotion labels The difference between them, the sentiment classification loss function is defined as:
[0113] ,
[0114] in, is the sentiment classification loss function, is the number of source domain speech samples, is the true emotion label of the i-th source domain speech sample, The emotion category prediction result of the i-th source domain speech sample generated by the first emotion classifier, The emotion category prediction result of the i-th source domain speech sample generated by the second emotion classifier;
[0115] Step 2-4: The target domain deep speech features extracted in step 2-1 are This is fed into the decision boundary alignment network constructed in step 2-2, where two differentiated sentiment classifiers are used to predict the emotion category of the target domain speech sample. This dual-branch structure captures speech samples with ambiguous emotion labels in the target domain from different discriminative perspectives, providing a discriminative basis for subsequent inconsistency detection and adversarial optimization.
[0116] The first emotion classifier is used to classify the target domain speech samples The sentiment label prediction results are:
[0117] ,
[0118] in, The target domain sentiment category prediction result generated by the first sentiment classifier;
[0119] The second emotion classifier is used to classify the target domain speech samples The sentiment label prediction results are:
[0120] ,
[0121] in, The target domain sentiment category prediction results generated by the second sentiment classifier;
[0122] Step 2-5: To further enhance the recognition ability of samples with ambiguous sentiment labels in the target domain, calculate the prediction probabilities of the two sentiment classifiers on the target domain samples. and The difference is detected and the difference detection loss function is constructed based on it. This mechanism can identify speech samples with ambiguous emotion labels near the decision boundary and provide a sample selection basis for subsequent adversarial optimization.
[0123] The target domain sentiment category prediction results generated by the first sentiment classifier Expands to:
[0124] ,
[0125] in, The probability of the first emotion classifier predicting the first emotion of the target domain sample, The first sentiment classifier predicts the target domain sample The probability of class emotion, is the number of common sentiment categories between the source domain and the target domain; in the cross-database task The value of is 5;
[0126] The target domain sentiment category prediction results generated by the second sentiment classifier Expands to:
[0127] ,
[0128] in, The probability of the first category of emotion of the target domain sample is predicted by the second emotion classifier, Predict the target domain sample for the second sentiment classifier Probability of class sentiment;
[0129] Calculate the difference in the prediction results of the two sentiment classifiers and define the difference detection loss function as:
[0130] ,
[0131] in, is the difference detection loss function, is the number of target domain speech samples, To take the absolute value; is the predicted probability distribution of the first emotion classifier for the nth category emotion of the mth target domain speech sample, is the predicted probability distribution of the nth category emotion of the mth target domain speech sample by the second emotion classifier.
[0132] Step 2-6: Build a feature distribution alignment network to extract the source domain deep speech features from step 2-1. and target domain deep speech features The input is fed into the feature distribution alignment network to reduce the difference in feature distribution between the source and target domains and enhance transfer capabilities under heterogeneous corpus conditions. The feature distribution alignment network is composed of multiple fully connected layers and nonlinear activation functions. By introducing a gradient reversal layer (GRL), the shared feature extraction network is guided to learn domain-invariant features, thereby unifying the feature space between the source and target domains. The domain category prediction results are:
[0133] ,
[0134] in, Represents the domain category prediction result of the generated input sample features, is the Sigmoid function used to convert the domain category score into a probability distribution, is the weight matrix of the feature distribution alignment network, is the ReLU activation function used to enhance domain features and suppress irrelevant features. It is a shared feature extraction network composed of multiple convolutional layers. are the input source domain and target domain speech samples, Bias vector for aligning the network to feature distribution;
[0135] Step 2-7, the domain prediction result output by the feature distribution alignment network in step 2-6 On this basis, a domain label supervision mechanism is introduced to convert speech samples The true domain label d and the domain category prediction results A joint optimization is performed to construct a domain classification loss function. This consistency constraint further reduces the difference in the distribution of discriminant features between the source and target domains, enhancing the model's domain-invariant feature extraction capabilities.
[0136] Field category prediction results Expands to:
[0137] ,
[0138] in, is the predicted probability that the domain label is the target domain, is the predicted probability that the domain label is the source domain;
[0139] The real domain label d is expanded to:
[0140] ,
[0141] in, is the true probability that the domain label is the target domain, is the true probability that the domain label is the source domain, is the kth speech sample input, when From the source domain corpus ,but ;when From the target domain corpus ,but ;
[0142] Calculate the difference between the domain category prediction result and the true domain label, and define the domain classification loss function as:
[0143] ,
[0144] in, is the domain classification loss function, is the true domain label of the k-th speech sample, The domain category prediction result of the k-th speech sample.
[0145] In step 3, to minimize the problem of incorrect mapping of emotion labels to target domain speech samples caused by differences in decision boundaries between corpora, a maximum-minimization difference detection loss function is used. Unlike the traditional method of directly minimizing the difference loss function, this method not only focuses on the currently misclassified target domain samples, but also deeply explores potential emotion label mismapped samples, thereby fine-tuning the decision boundaries between categories. Therefore, a multi-stage training strategy of collaborative adversarial learning is adopted. In the first stage, feature alignment is achieved through emotion classification loss, domain classification loss, and gradient reversal coefficient; in the second stage, the difference between the two emotion classifiers is maximized to filter target domain samples with ambiguous emotion labels; in the third stage, the difference between the classifiers is minimized, the decision boundary is refined, and the target domain sample labels are corrected.
[0146] In step 3-1, a stochastic gradient descent (SGD) optimizer is used during training, with a momentum factor set to 0.9 to accelerate model convergence and reduce gradient fluctuations, ensuring cross-database training stability. Initialize the parameters of the shared feature extraction network, decision boundary alignment network, and feature distribution alignment network, and adopt an exponentially decaying acoustic perception learning rate strategy to match the high-frequency attenuation of the human ear's Mel scale. This strategy focuses on fundamental frequency learning in the early stages of training and gradually optimizes high-frequency resonance peaks in the later stages, thereby improving model recognition performance. The learning rate formula is:
[0147] ,
[0148] in, is the adjusted learning rate, is the initial learning rate, is the current number of training times, is the total number of training sessions;
[0149] Step 3-2, the first stage training freezes the decision boundary alignment network parameters of the first sentiment classifier (weight matrix and the bias vector ) and the parameters of the second sentiment classifier (weight matrix and the bias vector ), retaining the difference in the initial decision perspective. Dynamically adjust the feature optimization direction through the gradient reversal mechanism and update the parameters of the feature alignment network (weight matrix and the bias vector ), gradient reversal coefficient and shared feature extraction network parameters .
[0150] First, according to the domain classification loss function , update the parameters of the feature alignment network to eliminate the difference in feature distribution between corpora:
[0151] ,
[0152] ,
[0153] in, The weight matrix of the updated feature distribution alignment network, The bias vector of the network is aligned to the updated feature distribution; the symbol Indicates that the update is performed using the formula on the right; represents partial differential;
[0154] Then, dynamically adjust the backpropagation layer Gradient reversal coefficient , gradually strengthen the learning of target domain features and balance the migration capabilities of cross-domain features:
[0155] ,
[0156] in, is the ratio of the current training times to the total training times; e is a natural constant;
[0157] Finally, according to the sentiment classification loss function , domain classification loss function and the gradient reversal coefficient , update the parameters of the shared feature extraction network :
[0158] ,
[0159] in, is the weight that balances the sentiment classification loss function and the domain classification loss function, Extract network parameters for the updated deep features; ensure that the weight ratio of the two loss functions, sentiment classification loss and domain classification loss, remains stable in the first stage of training through a fixed region weight parameter search strategy;
[0160] Step 3-3, the second stage of training freezes the shared feature extraction network parameters , ensure the consistency and stability of the feature extraction process, reduce possible interference in the decision boundary adjustment process, and make the model focus on adjusting the decision boundary. Maximize the difference in prediction results between the two sentiment classifiers, update the parameters of the first sentiment classifier of the decision boundary alignment network (weight matrix and the bias vector ) and the parameters of the second sentiment classification (weight matrix and the bias vector ).
[0161] Based on the constructed sentiment classification loss function and difference detection loss function , update the decision boundary alignment network parameters and filter out the target domain speech samples with incorrect mapping of emotion labels due to the misalignment of the decision boundary of cross-database speech emotion recognition:
[0162] , ,
[0163] , ,
[0164] in, is the weight that balances the sentiment classification loss function and the difference detection loss function, is the updated weight matrix of the first sentiment classifier, is the updated bias vector of the first sentiment classifier, is the updated weight matrix of the second sentiment classifier, is the bias vector of the updated second sentiment classifier; by using a fixed region weight parameter search strategy, the weight ratio of the two loss functions of sentiment classification loss and difference detection loss in the second stage training is ensured to remain stable;
[0165] Step 3-4, the third stage trains the parameters of the first sentiment classifier of the frozen decision boundary alignment network (weight matrix and the bias vector ) and the parameters of the second sentiment classifier (weight matrix and the bias vector ), in order to maintain the stability of the decision boundary, avoid excessive updates to the classifier parameters, and allow the model to focus on the feature extraction process. Minimize the difference in prediction results between the two sentiment classifiers and update the shared feature extraction network parameters .
[0166] Difference detection loss function based on construction , update the shared feature extraction network parameters , correct the incorrect sentiment label mapping caused by the misalignment of the decision boundary in the target domain:
[0167] ;
[0168] Steps 3-2 to 3-4 are performed alternately, by gradually optimizing and adjusting the model parameters until the model reaches a convergence state or the number of training times reaches the preset maximum value.
[0169] This embodiment uses three speech emotion databases, EMO-DB, CASIA, and eNTERFACE, to conduct a large number of experiments to verify the effectiveness of the model. Among them, EMO-DB is a German emotion corpus simulated by 10 professional actors, covering seven different emotional states: anger, boredom, disgust, fear, happiness, neutrality, and sadness. CASIA is a Chinese emotion corpus recorded by the Institute of Automation of the Chinese Academy of Sciences, which contains six different emotional states: anger, fear, happiness, neutrality, sadness, and surprise. eNTERFACE is a multimodal English audition emotion corpus, covering six different emotional states: anger, disgust, fear, happiness, sadness, and surprise. The present invention uses B to represent the EMO-DB speech emotion database, C to represent the CASIA speech emotion database, and E to represent the eNTERFACE speech emotion database.
[0170] This embodiment selects two of the databases as the training database and the test database respectively, and selects speech samples of the same emotion category in the two databases to form the source domain speech samples and the target domain speech samples. The present invention designs 6 groups of cross-database speech emotion recognition sub-experiments, which are represented as , , , , , The left side of the arrow is the training database, and the right side of the arrow is the test database. The number of speech samples and the types of speech samples in the cross-database speech emotion recognition sub-experiment are as follows: Figure 3 shown.
[0171] The network model in this paper was built using the PyTorch 1.11.0 deep learning framework, and the hardware operating system used an NVIDIA RTX 3090 graphics card. The network parameters were set to a maximum of 200 epochs, and the SGD optimizer was used to update network parameters, with a momentum of 0.9 and a learning rate of 0.01. All experiments used the unweighted average recall (UAR) as the evaluation metric.
[0172] To further validate the effectiveness of the proposed model, this paper compares five subspace-based cross-database speech emotion recognition methods with five deep learning-based cross-database speech emotion recognition methods. The subspace-based cross-database speech emotion recognition methods specifically include Transfer Component Analysis (TCA), Geodesic Flow Kernel (GFK), Subspace Alignment (SA), Domain Adaptive Subspace Learning (DoSL), and Joint Distribution Adaptive Regression (JDAR). Furthermore, the INTERSPEECH 2009 Emotion Challenge (IS09) feature set and the extended Geneva Minimization Acoustic Parameter Set (eGeMAPS) were used to characterize speech samples. To ensure a fair comparison, this example uses support vector machines (SVMs) as emotion classifiers for methods lacking classifiers (such as TCA, GFK, and SA). The deep learning-based cross-database speech emotion recognition methods specifically include Deep Adaptive Network (DAN), Joint Adaptive Network (JAN), Domain Adversarial Neural Network (DANN), Conditional Domain Adversarial Network (CDAN), and Deep Subdomain Adaptive Network (DSAN). Furthermore, this paper compares the model with a pre-trained model (VGG-11) on source domain data.
[0173] TCA, GFK and SA introduce a d-dimensional common subspace, where the dimension d ranges from a predefined interval [5: 5: ] in the menu. is the total number of elements in the acoustic parameter set. In addition, d and Parameters used to balance the original classification loss function and the feature distribution elimination term. d and The parameter range of is [5: 5: 200]. The experimental results are shown in Table 1.
[0174] DAN, JAN, DSAN, CDAN, and DANN explore different trade-off parameter values in a fixed interval [0.001: 0.001: 0.01, 0.02: 0.01: 0.1, 0.2: 0.1: 1] To adjust the classification loss function and domain adaptation term, the experimental results are shown in Table 2.
[0175] The present invention uses and To balance the classification loss function with the dual classifier-based adversarial learning module and the domain discriminator-based adversarial learning module. and The search range is [0.001:0.001:0.01, 0.02:0.01:0.1,0.2:0.1:1].
[0176] In the table, the best results of each cross-database speech emotion recognition task are marked in bold.
[0177] Table 1
[0178]
[0179] Table 2
[0180]
[0181] The present invention provides a cross-database speech emotion recognition method based on collaborative boundary-aware adversarial learning. There are many methods and approaches to implement this technical solution. The above is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements and modifications without departing from the principles of the present invention. These improvements and modifications should also be considered within the scope of protection of the present invention. All components not specified in this embodiment can be implemented using existing technologies.
Claims
1. A cross-database speech emotion recognition method based on collaborative boundary-aware adversarial learning, characterized by: The following steps are involved: Step 1: Obtain source and target domain speech samples, perform preprocessing, including amplitude normalization, pre-emphasis filtering, frame segmentation and windowing, and convert the preprocessed speech samples into Mel spectrogram features; Step 2: Build a cross-database speech emotion recognition model based on collaborative boundary-aware adversarial learning to reduce the differences between source and target domain speech samples from the dual dimensions of feature distribution and decision boundary, thereby achieving emotional knowledge transfer. Step 3: adopting a multi-stage training strategy, which includes first-stage training, second-stage training, and third-stage training, using a maximum-minimization difference detection loss function, obtaining all target domain speech samples with incorrect mapping of emotion labels due to decision boundary differences, and adjusting the decision boundary between categories; Step 2 includes: Step 2-1: The source domain speech sample and target domain speech samples The Mel spectrogram is input into the shared feature extraction network to extract the source domain deep speech features with strong migration ability and discriminativeness. and target domain deep speech features ; The shared feature extraction network uses the VGG11 network as the backbone network; The VGG11 network is structured and parameterized, removing the last three fully connected layers and retaining only the convolutional and pooling layers. Step 2-2, build a decision boundary alignment network: use two sentiment classifiers with the same structure but different initialization parameters to classify the source domain deep speech features extracted in step 2-1 Predict emotion categories from different perspectives; Both sentiment classifiers are multi-layer perceptron structures, denoted as the first sentiment classifier and the second sentiment classifier respectively. The weight matrices and bias vectors of the two sentiment classifiers are initialized through Gaussian distribution, but different random seeds are used to generate parameters, thereby forming a sentiment decision boundary with discriminative diversity in the feature space. The first emotion classifier is used to classify the source domain speech samples The sentiment label prediction results are: , in, The source domain sentiment category prediction result generated by the first sentiment classifier, is the Softmax function used to convert the sentiment category score into a probability distribution, is the weight matrix of the first sentiment classifier, is the ReLU activation function used to enhance emotional features and suppress irrelevant features. It is a shared feature extraction network composed of multiple convolutional layers. is the bias vector of the first sentiment classifier; The second emotion classifier is used to classify the source domain speech samples The sentiment label prediction results are: , in, The source domain sentiment category prediction results generated by the second sentiment classifier, is the weight matrix of the second sentiment classifier, is the bias vector of the second sentiment classifier; Step 2-3: Through joint supervision, the source domain speech samples True emotional label The prediction results of the two classifiers and Conduct joint optimization; Calculate sentiment label prediction results and and real emotion labels The difference between them, the sentiment classification loss function is defined as: , in, is the sentiment classification loss function, is the number of source domain speech samples, is the true emotion label of the i-th source domain speech sample, The emotion category prediction result of the i-th source domain speech sample generated by the first emotion classifier, The emotion category prediction result of the i-th source domain speech sample generated by the second emotion classifier; Step 2-4: The target domain deep speech features extracted in step 2-1 are Input the result to the decision boundary alignment network constructed in step 2-2, and use two emotion classifiers to predict the emotion category of the target domain speech sample. The first emotion classifier is used to classify the target domain speech samples The sentiment label prediction results are: , in, The target domain sentiment category prediction result generated by the first sentiment classifier; The second emotion classifier is used to classify the target domain speech samples The sentiment label prediction results are: , in, The target domain sentiment category prediction results generated by the second sentiment classifier; Step 2-5: Calculate the prediction probability of the two sentiment classifiers on the target domain samples and Difference, construct difference detection loss function; Step 2-6: Build a feature distribution alignment network to extract the source domain deep speech features from step 2-1. and target domain deep speech features Input into the feature distribution alignment network to reduce the difference in feature distribution between the source domain and the target domain; Step 2-7, the domain category prediction result output by the feature distribution alignment network in step 2-6 On this basis, a domain label supervision mechanism is introduced to convert speech samples The real domain label and domain category prediction results Perform joint optimization and construct domain classification loss function; Step 3 includes: In step 3-1, a stochastic gradient descent optimizer is used during training to initialize the parameters of the shared feature extraction network, the decision boundary alignment network, and the feature distribution alignment network. An exponentially decaying acoustic perception learning rate strategy is used to match the high-frequency attenuation of the human ear Mel scale. The learning rate formula is: , in, is the adjusted learning rate, is the initial learning rate, is the current number of training times, is the total number of training sessions; Step 3-2: During the first stage of training, freeze the decision boundary and align the parameters of the first sentiment classifier and the second sentiment classifier in the network to preserve the difference in the initial decision perspective. Dynamically adjust the feature optimization direction through the gradient reversal mechanism, update the parameters of the feature alignment network and the gradient reversal coefficient and shared feature extraction network parameters ; First, according to the domain classification loss function , update the parameters of the feature alignment network to eliminate the difference in feature distribution between corpora: , , in, The weight matrix of the updated feature distribution alignment network, The bias vector of the network is aligned to the updated feature distribution; the symbol Indicates that the update is performed using the formula on the right; represents partial differential; Then, dynamically adjust the backpropagation layer Gradient reversal coefficient , gradually strengthen the learning of target domain features and balance the migration capabilities of cross-domain features: , in, is the ratio of the current training times to the total training times; e is a natural constant; Finally, according to the sentiment classification loss function , domain classification loss function and the gradient reversal coefficient , update the parameters of the shared feature extraction network : , in, is the weight that balances the sentiment classification loss function and the domain classification loss function, Extract network parameters for the updated deep features; Step 3-3, during the second phase of training, freeze the parameters of the shared feature extraction network Maximize the difference in prediction results between the two sentiment classifiers and update the decision boundary to align the parameters of the first sentiment classifier with the parameters of the second sentiment classifier. Based on the constructed sentiment classification loss function and difference detection loss function , update the decision boundary alignment network parameters and filter out the target domain speech samples with incorrect mapping of emotion labels due to the misalignment of the decision boundary of cross-database speech emotion recognition: , , , , in, is the weight that balances the sentiment classification loss function and the difference detection loss function, is the updated weight matrix of the first sentiment classifier, is the updated bias vector of the first sentiment classifier, is the updated weight matrix of the second sentiment classifier, is the updated bias vector of the second sentiment classifier; Step 3-4: During the third stage of training, freeze the decision boundary and align the parameters of the first sentiment classifier with the parameters of the second sentiment classifier; minimize the difference in prediction results between the two sentiment classifiers and update the shared feature extraction network parameters. ; Difference detection loss function based on construction , update the shared feature extraction network parameters , correct the incorrect sentiment label mapping caused by the misalignment of the decision boundary in the target domain: ; Step 3-5, step 3-2 ~ step 3-4 are three stages of training, which are performed alternately. The model parameters are gradually optimized and adjusted until the model reaches a convergence state or the number of training times reaches the preset maximum value.
2. The method according to claim 1, characterized in that Step 1 includes: Step 1-1: Randomly select two databases from the public speech emotion database with significant heterogeneity to form the source domain corpus and the target domain corpus , to simulate the actual multi-source heterogeneous data source scenario; A label consistency screening method is used to retain only speech samples with the same emotion category, ensuring that the source domain and the target domain share a unified emotion label space. Step 1-2: source domain speech samples and target domain speech samples Perform amplitude normalization, pre-emphasis filtering, frame division and windowing processing; Steps 1-3: Select a supervised source domain corpus and unsupervised target domain corpus Conduct collaborative training; Source domain corpus With emotional labels, for Corresponding sentiment labels; target domain corpus No emotional labels, Simulate actual speech samples without emotion labels.
3. The method according to claim 2, characterized in that Steps 1-2 include: using maximum normalization to scale the sample amplitude to [-1, 1]; using a Gaussian filter to enhance the high-frequency information of the speech; using a Hanning window for framing and windowing; then, converting to a Mel frequency scale to retain emotional information and remove redundant features, achieving consistency in the source domain and target domain speech sample features, and converting to a Mel frequency scale that conforms to the auditory characteristics of the human ear.
4. The method according to claim 3, characterized in that Steps 2-5 include: the target domain emotion category prediction result generated by the first emotion classifier Expands to: , in, The probability of the first emotion classifier predicting the first emotion of the target domain sample, The first sentiment classifier predicts the target domain sample The probability of class emotion, is the number of common sentiment categories between the source and target domains; The target domain sentiment category prediction results generated by the second sentiment classifier Expands to: , in, The probability of the first category of emotion of the target domain sample is predicted by the second emotion classifier, Predict the target domain sample for the second sentiment classifier Probability of class sentiment; Calculate the difference in the prediction results of the two sentiment classifiers and define the difference detection loss function as: , in, is the difference detection loss function, is the number of target domain speech samples, To take the absolute value; is the predicted probability distribution of the first emotion classifier for the nth category emotion of the mth target domain speech sample, is the predicted probability distribution of the nth category emotion of the mth target domain speech sample by the second emotion classifier.
5. The method according to claim 4, characterized in that Steps 2-6 include: the feature distribution alignment network includes two or more fully connected layers and nonlinear activation functions, and introduces a gradient reversal mechanism to guide the shared feature extraction network to learn domain-invariant features to achieve the unification of the feature space between the source domain and the target domain. The domain category prediction result is: , in, Represents the domain category prediction result of the generated input sample features, is the Sigmoid function used to convert the domain category score into a probability distribution, is the weight matrix of the feature distribution alignment network, is the ReLU activation function used to enhance domain features and suppress irrelevant features. are the input source domain and target domain speech samples, is the bias vector of the feature distribution alignment network.
6. The method according to claim 5, characterized in that Steps 2-7 include: Field category prediction results Expands to: , in, is the predicted probability that the domain label is the target domain, is the predicted probability that the domain label is the source domain; The real domain label d is expanded to: , in, is the true probability that the domain label is the target domain, is the true probability that the domain label is the source domain, is the kth speech sample input, when From the source domain corpus ,but ;when From the target domain corpus ,but ; Calculate the difference between the domain category prediction result and the true domain label, and define the domain classification loss function as: , in, is the domain classification loss function, is the true domain label of the k-th speech sample, The domain category prediction result of the k-th speech sample.
7. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 6.
8. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 6 are executed.
Citation Information
Patent Citations
PDAN-based cross-library speech emotion recognition method and device
CN115512721A
Bionic signal processing method based on wavelet transform and generative adversarial network
CN116403590A