A zero-shot speech emotion recognition method based on reconstructed prototype and generative learning
By adopting prototype reconstruction and generative learning methods in zero-sample speech emotion recognition, the domain gap problem between speech and auxiliary mode is solved, and the changing characteristics of speech data are taken into account, effectively identifying speech signals of unknown emotions categories is achieved.
Patent Information
- Application Number
- CN202310418118.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-04-19
AI Technical Summary
In the existing zero-sample speech emotion recognition methods, direct use of the original prototype leads to a domain gap between speech and auxiliary modes, and most strategies fail to fully consider the changing characteristics of speech data.
Using a method based on prototype reconstruction and generative learning, the semantic embedding prototype is reconstructed and a prototype matching the speech signal domain is generated. Combined with the generative learning method, segment sample sublingual characteristics of unknown emotional categories are generated.
Effective recognition of speech signals of unknown emotional categories is achieved, the domain gap between speech and auxiliary mode is reduced, and the changing characteristics of speech data are fully taken into account.
Smart Images

Figure CN116434785B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of speech signal emotion recognition, and in particular relates to a zero-sample speech emotion recognition method based on prototype reconstruction and generative learning. Background Art
[0002] Speech Emotion Recognition (SER) has a wide range of applications in human-computer interaction and other fields. It can determine the subjective intention of the speaker in the speech segment and the speaker's deeper emotional expression by studying the emotional information in the speech signal. In addition, it can also synthesize the emotional expression of the speech signal by analyzing the emotional information in the speech. In the diagnosis of mental illness, relevant technologies can be used to achieve preliminary screening of patients with depression and provide a basis for further diagnosis and treatment; in virtual reality, it can enable robots to have more powerful emotional analysis and expression capabilities.
[0003] Different from traditional speech emotion recognition research, zero-shot speech emotion recognition focuses on autonomously identifying invisible emotions in speech through machine learning, enabling machines to perceive unknown emotional speech without knowing any emotional state samples. Zero-shot speech emotion recognition makes it possible for machines to autonomously perceive emotional information through information transfer of external knowledge, which helps autonomous emotion computing based on audio.
[0004] In the public zero-shot speech emotion recognition schemes, such as the public literature: Xu X, Deng J, Cummins N, et al. Exploring zero-shot emotion recognition in speech using semantic-embedding prototypes [J]. IEEE Transactions on Multimedia, 2022, 24: 2752-2765., there are mainly two shortcomings: First, the existing work uses the original prototypes obtained directly from the auxiliary modality, which may lead to the domain gap between speech and auxiliary modalities; second, most strategies study the information transfer from known domains to unknown domain classifiers, without fully considering the changing characteristics of speech data. Summary of the invention
[0005] Aiming at the problem that directly using the original prototype in the prior art leads to differences between the paralinguistic features and auxiliary modalities of the speech signal, and most of the existing strategies do not fully consider the changing characteristics of speech data, the present invention proposes a zero-sample speech emotion recognition method based on reconstructed prototype and generative learning.
[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0007] A zero-shot speech emotion recognition method based on reconstructed prototype and generative learning, including data acquisition steps, model optimization and application steps;
[0008] The data acquisition step is to first establish a speech emotion database, which includes several speech samples, each sample has its corresponding emotion category label; divide the speech emotion database into a training set consisting of samples of known emotion categories, and a test set consisting of samples of unknown emotion categories; each sample has a known and unique emotion category label.
[0009] The model optimization and application steps include the following steps performed sequentially:
[0010] Step 1: Extract and generate n F N-dimensional original features: Each paragraph sample in the training sample set is processed separately to extract the corresponding paralinguistic features as the original features, and the original features are regularized to obtain N (S) The regularized features corresponding to the training samples
[0011] Step 2: Perform semantic embedding mapping on the known emotion category names to generate semantic embedding prototypes for each known emotion category where c (S) is the number of known emotion categories, n A is the semantic embedding dimension of the emotion category name;
[0012] Step 3: Regularized paralinguistic features X extracted from samples of known emotion categories (S) And the emotion category label of the corresponding sample And the known emotion category prototype A (S) , train the prototype reconstruction model, and obtain the reconstruction prototype of the known emotion category segment sample and the optimal prototype reconstruction model, and then use the optimal prototype reconstruction model to combine the unknown emotion category semantic embedding prototype to obtain the unknown emotion category segment sample reconstruction prototype
[0013] Step 4: Use the regularized paralinguistic features X extracted from the known emotion category segment samples (S) and the emotion category label Y of the corresponding sample (S) , and the prototype reconstructed by the known emotion category segment sample Generate learning by training the generator f G (·) and the discriminator f D (·,·), solve the optimal generative learning model corresponding to the generator Then, the optimal generative learning model is used to reconstruct the prototype in combination with the unknown emotion category. Generate paralinguistic features of samples generated from unknown emotion categories;
[0014] Step 5: After the paralinguistic features of the samples generated by the unknown emotion category are regularized as described in step 1, they are sent to the classifier for training to obtain the optimal classifier, and then the paralinguistic features are extracted for the test segment samples of the unknown emotion category, and the regularization method described in step 1 is used to obtain the regularized paralinguistic features x of the test segment samples of the unknown emotion category. (U) , and use the optimal classifier to make classification decisions.
[0015] Furthermore, the method of regularization processing in step 1 is as follows:
[0016] The feature column vector of any sample in all paragraph samples before normalization is x (0) , where N (S) The training sample set consisting of feature column vectors of known emotion category training samples is set up for The j-th characteristic element of
[0017] For any sample feature column vector x (0) , feature j corresponds to the element The calculation formula for the regularization process is:
[0018]
[0019] in Represents X (0) The largest element in row j, Represents X (0) The smallest element in the jth row; x ·j for The result after normalization processing;
[0020] All elements in any sample are calculated according to formula (1) to obtain the normalized feature column vector x = [x ·1 ,x ·2 ,...,x ·n ] T , where the normalized feature vectors of the speech signal samples belonging to the known emotion category training sample set constitute the normalized feature vector set of the training sample
[0021] Furthermore, the semantic embedding mapping described in step 2 can be achieved by using a word vector pre-training model for the emotion category name:
[0022] For each emotion category, the five nearest neighboring words given in the SenticNet5 model are used to represent the emotion. Secondly, the semantic embeddings of these neighboring words are obtained by inputting the pre-trained fastText model, and the semantic embeddings of the five neighboring words corresponding to each emotional state are averaged as the semantic embedding prototype of the emotion category.
[0023] By using the above method, we input the emotion category name into SenticNet5 and fastText models and get the n corresponding to the category. A dimensional semantic embedding prototype of emotion categories. For the known emotion categories corresponding to the training set, the semantic embedding prototypes of each category are expressed as For the c to be predicted in the test sample set (U) unknown emotion categories, and the semantic embedding prototypes of each category are expressed as
[0024] Furthermore, the optimal prototype reconstruction model in step 3 is
[0025]
[0026] Where ψ(·) represents the prototype reconstruction mapping on the original prototype, J(·, ·) represents the alignment loss function between the corresponding columns of the two parameters, which is implemented by linear ν-Support Vector Regression (ν-SVR). is a known emotion category example; the linear mapping matrix Linear dimensionality reduction is performed on visible emotion samples through principal component analysis (PCA). For the c-th known emotion category, the column vector and N c (S) The normalized paralinguistic features of samples of the known emotion category Get the center of the cth known emotion category As examples of each known emotion category.
[0027] Semantic embedding prototypes A for known and unknown emotion categories (S) and A (U) , reconstruct the model using the optimal prototype Get the reconstructed prototypes of known and unknown emotion categories respectively and According to the known emotion category segment sample label Y (S) ,Will Assign to the corresponding known emotion category paragraph samples, and obtain the known emotion category reconstruction prototype corresponding to each known emotion category paragraph sample
[0028] Furthermore, the generator f in step 4 G (·) is used to learn the optimal generative learning model and reconstruct the prototype according to any emotion category to obtain the paralinguistic features of the generated samples of the corresponding emotion category. Suppose any known emotion category reconstructs the prototype Its network structure is:
[0029]
[0030] Among them, f G1 (·,·) is the output of the first layer, and its formula is:
[0031]
[0032] Among them, n q The dimensional noise q follows a normal distribution, ReLU(·) and LeakyReLU(·) are the rectified linear unit (ReLU) and leaky ReLU activation functions, respectively. and represents the linear weights of the generator network, and is the corresponding bias, and the hidden layer includes n G nodes.
[0033] Furthermore, the discriminator f in step 4 D (·,·) is used to distinguish X (S) Any known emotion category of the paragraph sample or the generated sample paralinguistic features x (S) , and the corresponding reconstructed prototype The formula is:
[0034]
[0035] Among them, f D1 (·,·) is the output of the hidden layer of the discriminator, and its formula is
[0036]
[0037] in and Represented as the linear weight of the discriminator network, b D1 and b D2 is the corresponding bias, and n is used in the hidden layer D nodes.
[0038] Furthermore, in the solution of the optimal generative model described in step 4, including the training of the generator and the discriminator, the loss function used is:
[0039]
[0040] Among them, weight λ3>0; Paralinguistic features representing samples generated from known emotion categories In the corresponding label The classification error rate under Represents the known emotion category c (S) N G The known emotion category labels of the generated samples (each category N G Generate samples). Cls(·,·) is the sample X in the known emotion category segment (S) The Softmax classifier trained on , uses the Negative Log Likelihood (NLL) loss.
[0041] Wasserstein Generative Adversarial Network (WGAN) loss L WGAN Use Wasserstein Generative Adversarial Network-Gradient Penalty (WGAN-GP) loss:
[0042]
[0043] The real sample features of known emotion categories Generate sample features based on known emotion categories The weight λ1>0 is used to balance the relationship between real data and generated data. The features of the synthesized samples of known emotion categories are The weight λ2 follows a 0-1 uniform distribution.
[0044] Furthermore, the classifier in step five may be any supervised learning classifier, and the supervised learning classifier is preferably a Softmax classifier or a support vector machine (Support Vector Machine, SVM for short) classifier.
[0045] Beneficial effects: Figure 1As shown, the present invention provides a zero-shot speech emotion recognition method based on prototype reconstruction and generative learning. By reconstructing the semantic embedding prototype and generating paralinguistic features of unknown emotion category paragraph samples, the classification judgment of unknown emotion category is realized for paragraph signal samples with unknown emotion state information. Specifically, in the prototype reconstruction stage, paralinguistic features are extracted from known emotion category paragraph samples, and the optimal prototype reconstruction model is trained by combining the known emotion category semantic embedding prototype and the known emotion category paragraph sample label. Then, the known and unknown emotion category semantic embedding prototypes are combined to obtain the known and unknown emotion category reconstruction prototypes respectively; in the generative learning stage, paralinguistic features are extracted from known emotion category paragraph samples, and then the prototype is reconstructed according to the known emotion category paragraph sample label and combined with the known emotion category to obtain the optimal generative learning model, and then the prototype is reconstructed by combining the unknown emotion category to obtain the paralinguistic features of the unknown emotion category generated sample; in the supervised learning stage, the optimal classifier is trained by using the paralinguistic features of the unknown emotion category generated sample, and the unknown emotion category test paragraph sample is judged for the unknown emotion category.
[0046] There are two problems in the existing zero-shot speech emotion recognition methods, such as the public document: Xu X, Deng J, Cummins N, et al. Exploring zero-shot emotion recognition in speech using semantic-embedding prototypes [J]. IEEE Transactions on Multimedia, 2022, 24: 2752-2765.: directly using the original prototype leads to the domain gap between speech and auxiliary modalities, and most existing strategies do not fully consider the inherent changes in speech data. Therefore, the present invention provides a zero-shot speech emotion recognition method, which adopts a method based on reconstruction of prototypes and generative learning, and realizes zero-shot speech emotion recognition by reconstructing prototypes that match the speech signal domain, and a generative learning method that can provide a diversity of generated samples, that is, using known emotion category segment samples to realize classification judgment of unknown emotion category segment samples.
[0047] Experiments have shown that in the aspect of speech signal emotion recognition, the present invention proposes a zero-sample speech emotion recognition method based on prototype reconstruction and generative learning, which can effectively identify the unknown emotion category to which speech signal segment samples of unknown emotion category belong. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 This is a flow chart of a zero-sample speech emotion recognition method based on prototype reconstruction and generative learning of the present invention. DETAILED DESCRIPTION
[0049] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.
[0050] like Figure 1 As shown, in the prototype reconstruction stage, the method of the present invention first extracts paralinguistic features from known emotion category paragraph samples and regularizes them, and then trains the prototype reconstruction model in combination with the known emotion category semantic embedding prototype obtained by processing the known emotion category name and the known emotion category paragraph sample label, and solves to obtain the optimal prototype reconstruction model, and then uses the optimal prototype reconstruction model to process the known and unknown emotion category semantic embedding prototypes to obtain the known and unknown emotion category reconstruction prototypes respectively; in the generative learning stage, the regularized paralinguistic features of the known emotion category paragraph samples, the known emotion category paragraph sample labels, and the known emotion category reconstruction prototype are used to train the generative learning model to obtain the optimal generative learning model, and the optimal generative learning model is used to generate the paralinguistic features of the unknown emotion category generation sample in combination with the unknown emotion category reconstruction prototype; in the supervised learning stage, the paralinguistic features of the unknown emotion category generation sample and its corresponding unknown emotion category label are used to train the classifier to obtain the optimal classifier, and then the optimal classifier is used to extract and regularize the paralinguistic features of the unknown emotion category test paragraph sample to obtain the unknown emotion category judgment result of the paragraph sample.
[0051] The following experimental method compares the unweighted accuracy (UA) recognition rate and macro F1-score index between the method of the present invention and the existing zero-sample speech emotion recognition method.
[0052] The experiment uses part of the speech signal in the DEMoS (Database of Elicited Mood in Speech) database to verify the effectiveness of the method of the embodiment of the present invention.
[0053] The database is recorded in Italian, with a total of 9697 samples belonging to 68 speakers, including 23 women. This experiment uses 8 emotion categories, namely neutral (neu), guilt (col), disgust (dis), happiness (gio), fear (pau), anger (rab), surprise (sor), and sadness (tri); the dataset uses all samples of every two categories of emotions as unknown emotion test segment sample sets, and other emotion category samples as known emotion training segment sample sets. There are 28 different sample type combinations, so this experiment conducted 28 training tests in total.
[0054] The original paralinguistic features of the experiment use the extended Geneva Minimalistic Acoustic Parameter Set (eGeMAPS) as the feature set, and the original feature dimension n F = 88, the features come from 25 low-level descriptors (LLDs) combined with high-level statistical functions (HSFs), as well as temporal features and equivalent sound levels, and the feature set is obtained through expert selection. In this specific experiment, the openSMILE toolbox is used to extract these features.
[0055] Semantic embedding prototypes of known and unknown emotion categories using n A = 300-dimensional English word semantic vectors, these prototypes are based on the SenticNet 5 model and the fastText pre-trained model. The semantic embedding model used in the experiment is fastText, which uses 2 million word vectors trained on Common Craw.
[0056] In the experiment, in order to reflect the effect of the method of the present invention, the methods used for comparison are: SSE (SemanticSimilarity Embedding), LatEm (Latent Embedding), ESZSL (Embarrassingly SimpleZero-Shot Learning), SynC (Synthesised Classifiers), EXEM (EXEMplarsynthesis), FGN (Feature Generating Network), LisGAN (Lever-aging invariant side GAN) recognition models, and the SynC method under the condition of randomly selecting 1000 virtual prototypes (denoted as "SynC-rand") is also considered, and the best result of the EXEM method in all reduction dimensions is further considered (denoted as "EXEM-best"). A zero-sample speech emotion recognition method based on prototype reconstruction and generative learning of the present invention is expressed as: the method of the present invention (Example 1).
[0057] In the experiment, for the known emotion category paragraph samples as the training set, the optimal parameter selection was performed using the emotion category independent three-fold cross validation. As Example 1, the specific range of parameters selected in Example 1 is: the regularization coefficient is {2 -3 ,2 -2 ,...,2 3}, ν value is {2 -8 ,2 -7 ,...,2 0}, the kernel scale parameter is {2 -4 ,2 -3 ,...,2 4 In the embodiment of the present invention, the parameters are selected as follows: In the generation learning process, the adaptive moment estimation (Adam) optimization operator is used, and the initial learning rate is 10 -6 , the maximum number of training rounds is 30, the batch size is 64; in the generator and discriminator, n G and n D Set to 4096, n q =312; for the Softmax classifier in Cls(·,·), the initial learning rate of the Adam optimization operator is set to 10 -3 , the maximum number of training rounds is 50, and the batch size is 100; the weights λ1 and λ3 are set to 10 and 0.01 respectively, and λ2 is randomly selected according to the uniform distribution; by using the trained generator, each unseen emotion class generates N G = 600 samples; for the classifier that performs supervised learning on the generated unseen emotion samples, the initial learning rate of the Adam optimizer in 25 rounds of training is set to 10 -4 .
[0058] The optimal UA average results of these state-of-the-art zero-shot speech emotion recognition methods on the DEMoS database are shown in Table 1. As shown in Table 1, compared with other related comparative examples, the method of the present invention in Example 1 can achieve better UA performance and F1-score for the recognition of unknown emotions in speech signals.
[0059] Table 1
[0060] Adoption Method UA F1-score Comparative Example 1 SSE 51.8±3.7 - Comparative Example 2 LatEm 52.0±2.9 - Comparative Example 3 ESZSL 52.0±3.8 - Comparative Example 4 SynC 54.6±5.9 51.9±6.6 Comparative Example 5 SynC-rand 57.8±4.6 54.5±6.4 Comparative Example 6 EXEM 54.3±6.0 52.4±6.8 Comparative Example 7 EXEM-best 61.8±4.8 59.2±4.5 Comparative Example 8 FGN 58.7±3.1 57.5±3.0 Comparative Example 9 LisGAN 59.6±2.8 58.2±2.7 Example 1 Method of the present invention 63.9±2.7 62.5±2.4
[0061] Furthermore, in order to investigate the performance of EXEM-best, FGN, LisGAN and the method of the present invention in the recognition of each emotion category, the average UA results (%) in the experiment when each emotion category is used as an unknown emotion category are compared and made into Table 2. It can be seen from Table 2 that compared with the existing prototype reconstruction and generative learning methods, the method of the present invention has better performance for the recognition of most emotion categories.
[0062] Table 2
[0063]
[0064] In summary, the method of the present invention adopted in this embodiment 1 first reconstructs the semantic embedding prototype, and then generates the paralinguistic features of the unknown emotion category samples through the optimal generative learning model obtained through training, and finally trains the classifier for the unknown emotion category through supervised learning, thereby achieving better performance in the zero-sample speech emotion recognition problem.
[0065] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A zero-shot speech emotion recognition method based on reconstructed prototype and generative learning, characterized in that: The zero-sample speech emotion recognition method comprises: The data acquisition step is to first establish a speech emotion database, which includes several paragraph samples, each of which has its corresponding emotion category label; divide the speech emotion database into a training set consisting of samples of known emotion categories and a test set consisting of samples of unknown emotion categories; each paragraph sample has a known and unique emotion category label; Model optimization and application steps, the specific model optimization and application steps include the following specific steps: Step 1: Extract and generate n F N-dimensional original features: For each paragraph sample in the training sample set, the corresponding paralinguistic features are extracted as original features, and the original features are regularized to obtain N (S) The regularized features corresponding to the training samples Step 2: Perform semantic embedding mapping on the known emotion category names to generate semantic embedding prototypes for each known emotion category where c (S) is the number of known emotion categories, n A is the semantic embedding dimension of the emotion category name; Step 3: extract the regularized paralinguistic features X from the known emotion category segment samples (S) And the emotion category label of the corresponding sample And the known emotion category prototype A (S) , train the prototype reconstruction model, and obtain the reconstruction prototype of the known emotion category segment sample and the optimal prototype reconstruction model, and then use the optimal prototype reconstruction model to combine the unknown emotion category semantic embedding prototype to obtain the unknown emotion category segment sample reconstruction prototype Step 4: Use the regularized paralinguistic features X extracted from the known emotion category segment samples (S) and the emotion category label Y of the corresponding sample (S) , and the prototype reconstructed by the known emotion category segment sample Generate learning by training the generator f G (·) and the discriminator f D (·,·), solve the optimal generative learning model corresponding to the generator Then, the optimal generative learning model is used to reconstruct the prototype in combination with the unknown emotion category. Generate paralinguistic features of samples generated from unknown emotion categories; Step 5: After the paralinguistic features of the samples generated by the unknown emotion category are regularized as described in step 1, they are sent to the classifier for training to obtain the optimal classifier, and then the paralinguistic features are extracted for the test segment samples of the unknown emotion category, and the regularization method described in step 1 is used to obtain the regularized paralinguistic features x of the test segment samples of the unknown emotion category. (U) , and use the optimal classifier to make classification decisions.
2. A zero-sample speech emotion recognition method based on reconstructed prototype and generative learning according to claim 1, characterized in that: The method of regularization processing described in step 1 is as follows: The feature column vector of any sample in all paragraph samples before normalization is x (0) , where N (S) The training sample set consisting of feature column vectors of known emotion category training samples is set up for The j-th characteristic element of For any sample feature column vector x (0) , feature j corresponds to the element The calculation formula for the regularization process is: in Represents X (0) The largest element in row j, Represents X (0) The smallest element in the jth row; x ·j for The result after normalization processing; All elements in any sample are calculated according to formula (1) to obtain the normalized feature column vector x = [x ·1 ,x ·2 ,...,x ·n ] T , where the normalized feature vectors of the speech signal samples belonging to the known emotion category training sample set constitute the normalized feature vector set of the training sample 3. A zero-sample speech emotion recognition method based on reconstructed prototype and generative learning according to claim 1, characterized in that: The semantic embedding mapping described in step 2 can be achieved by using a word vector pre-training model for the emotion category names: For each emotion category, the five nearest neighboring words given in the SenticNet 5 model are used to represent the emotion. Secondly, the semantic embeddings of these neighboring words are obtained by inputting the pre-trained fastText model, and the semantic embeddings of the five neighboring words corresponding to each emotional state are averaged as the semantic embedding prototype of the emotion category. By using the above method, we can input the emotion category name into SenticNet 5 and fastText models and get the n corresponding to the category. A dimensional semantic embedding prototype of emotion categories; for the known emotion categories corresponding to the training set, the semantic embedding prototypes of each category are expressed as For the c to be predicted in the test sample set (U) unknown emotion categories, and the semantic embedding prototypes of each category are expressed as 4. The zero-sample speech emotion recognition method based on reconstructed prototype and generative learning according to claim 1, characterized in that: The optimal prototype reconstruction model described in step 3 is Where ψ(·) represents the prototype reconstruction mapping on the original prototype, J(·, ·) represents the alignment loss function between the corresponding columns of the two parameters, which is implemented by linear ν-support vector regression; is a known emotion category example; the linear mapping matrix Linear dimensionality reduction is performed on visible emotion samples through principal component analysis PCA; for the c-th known emotion category, the column vector and The normalized paralinguistic features of samples of the known emotion category Get the center of the cth known emotion category As examples of each known emotion category; Semantic embedding prototypes A for known and unknown emotion categories (S) and A (U) , reconstruct the model using the optimal prototype Get the reconstructed prototypes of known and unknown emotion categories respectively and According to the known emotion category segment sample label Y (S) ,Will Assign to the corresponding known emotion category paragraph samples, and obtain the known emotion category reconstruction prototype corresponding to each known emotion category paragraph sample 5. The zero-sample speech emotion recognition method based on reconstructed prototype and generative learning according to claim 1, characterized in that: The generator f described in step 4 G (·) is used to learn the optimal generative learning model and reconstruct the prototype according to any emotion category to obtain the paralinguistic features of the generated samples of the corresponding emotion category; suppose any known emotion category reconstructs the prototype Its network structure is: Among them, f G1 (·,·) is the output of the first layer, and its formula is: Among them, n q The dimensional noise q follows a normal distribution, ReLU(·) and LeakyReLU(·) are the rectified linear unit ReLU and Leaky ReLU activation functions, respectively. and represents the linear weights of the generator network, and is the corresponding bias, and the hidden layer includes n G nodes.
6. A zero-sample speech emotion recognition method based on reconstructed prototype and generative learning according to claim 1, characterized in that: The discriminator f described in step 4 D (·,·) is used to distinguish X (S) Any known emotion category of the paragraph sample or the generated sample paralinguistic features x (S) , and the corresponding reconstructed prototype The formula is: Among them, f D1 (·,·) is the output of the hidden layer of the discriminator, and its formula is in and Represented as the linear weight of the discriminator network, b D1 and b D2 is the corresponding bias, and n is used in the hidden layer D nodes.
7. The zero-sample speech emotion recognition method based on reconstructed prototype and generative learning according to claim 1, characterized in that: The process of solving the optimal generative learning model corresponding to the generator described in step 4 includes training the generator and the discriminator, and the loss function used is: Among them, weight λ3>0; Paralinguistic features representing samples generated from known emotion categories In the corresponding label The classification error rate under Represents the known emotion category c (S) N G The known emotion category labels of the generated samples, each category N G Generate samples; Cls(·,·) is a sample X in a known emotion category segment (S) Softmax classifier trained on , using negative log-likelihood NLL loss; Wasserstein Generative Adversarial Network (WGAN) Loss L WGAN Generate adversarial network gradient penalty WGAN-GP loss using Wasserstein: The real sample features of known emotion categories Generate sample features based on known emotion categories The weight λ1>0 is used to balance the relationship between real data and generated data. The features of the synthesized samples of known emotion categories are The weight λ2 follows a 0-1 uniform distribution.
8. The zero-sample speech emotion recognition method based on reconstructed prototype and generative learning according to claim 1, characterized in that: The classifier described in step 5 is a supervised learning classifier.
9. A zero-sample speech emotion recognition method based on reconstructed prototype and generative learning according to claim 8, characterized in that: The supervised learning classifier is a Softmax classifier or a support vector machine classifier.