Unsupervised Emotional Speech Synthesis Device and Method

Through the unsupervised emotional speech synthesis method, the characteristics of the emotional hidden variables extracted by the audio are modeled posteriorly and randomly sampled to obtain multi-dimensional emotional labels, solving the problem of dependence on manual annotation and poor expression in the existing technology, and achieving high expressive emotional speech synthesis.

CN114842881BActive Publication Date: 2025-06-17四川启睿克科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210634438.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-07
Publication Date
2025-06-17
Estimated Expiration
2042-06-07

AI Technical Summary

Technical Problem

Existing emotional speech synthesis methods rely on artificially labeled emotional labels, and single-dimensional emotional labels cannot meet the requirements of high expressive emotional speech.

Method used

Unsupervised method is used to model the posterior distribution of the emotional hidden variable features extracted from audio, and obtain multi-dimensional emotional labels through random sampling to achieve high expressive emotional speech synthesis.

Benefits of technology

Without artificial emotional annotation, high expressive emotional speech synthesis is achieved, solving the problem of dependence on manual annotation and poor expression in emotional speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842881B_ABST
    Figure CN114842881B_ABST
Patent Text Reader

Abstract

The present invention discloses an unsupervised emotional speech synthesis device, comprising: an emotion extraction module for extracting emotion latent variable features from the input real audio features; a random sampling module for obtaining emotion classification labels and the posterior distribution of each category by using the emotion latent variable features; a parameter fitting module for fitting the parameters of the speech synthesis module according to the prior distribution, the posterior distribution, the real audio features and the predicted audio features; and a speech synthesis module for generating emotional speech for the text to be synthesized and the target emotion label. The present invention also discloses an unsupervised emotional speech synthesis method. The present invention does not require manual emotional annotation of data, adopts an unsupervised method to model the posterior distribution of the emotion latent variable features extracted from the audio, and performs random sampling to obtain multi-dimensional emotion labels, realizing high-expression emotional speech synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech synthesis, and particularly to an unsupervised emotional speech synthesis device and method. Background Art

[0002] Speech synthesis is a technology that converts text information into speech information, that is, converts text information into any audible speech. It involves multiple disciplines such as acoustics, linguistics, and computer science. With the development of technology, emotional speech synthesis has begun to become a research hotspot. However, previous emotional speech synthesis methods usually require manually labeled emotional tags, and single-dimensional emotional tags cannot meet the requirements of highly expressive emotional speech. Summary of the Invention

[0003] To solve the problems existing in the prior art, the object of the present invention is to provide an unsupervised emotional speech synthesis device and method. The present invention does not require manual emotional annotation of data, uses an unsupervised method to model the posterior distribution of emotional latent variable features extracted from audio, and performs random sampling to obtain multi-dimensional emotional tags, realizing highly expressive emotional speech synthesis, and solving the problems of dependence on manual emotional annotation and poor expressiveness in emotional speech synthesis.

[0004] To achieve the above object, the technical solution adopted by the present invention is: an unsupervised emotional speech synthesis device, comprising:

[0005] An emotion extraction module, configured to extract emotional latent variable features from the input real audio features;

[0006] A random sampling module, configured to obtain an emotion classification label and the posterior distribution of each category by using the emotional latent variable features;

[0007] A parameter fitting module, configured to fit the parameters of the speech synthesis module according to the prior distribution, the posterior distribution, the real audio features, and the predicted audio features;

[0008] A speech synthesis module, configured to generate emotional speech for the text to be synthesized and the target emotional label.

[0009] The present invention also provides an unsupervised emotional speech synthesis method, which is implemented by using the above-mentioned unsupervised emotional speech synthesis device. The method includes the following steps:

[0010] S1. Obtain the input text and the corresponding real audio features, and extract the emotional latent variable features from the real audio features through the emotion extraction module;

[0011] S2. Randomly sample the emotional latent variable features obtained in step S1 by using the random sampling module to obtain an emotion classification label, and calculate the posterior distribution of each category at the same time;

[0012] S3. Obtain the speech synthesis module, and obtain the predicted audio features according to the speech synthesis module, the input text in step S1, and the sentiment classification label in step S2;

[0013] S4. Obtain the prior distribution of the sentiment label, and use the parameter fitting module to train the parameters of the speech synthesis module according to the prior distribution, the posterior distribution obtained in step S2, the predicted audio features obtained in step S3, and the real audio features in step S1;

[0014] S5. Obtain the emotional speech according to the text to be synthesized, the target sentiment label, and the speech synthesis module in step S4.

[0015] As a further improvement of the present invention, the emotion extraction module adopts a convolutional neural network or a recurrent neural network, and its network parameters are jointly optimized with the speech synthesis module through gradient backpropagation, and are used to extract the sentence-level emotion latent variable features from the real audio features.

[0016] As a further improvement of the present invention, in step S2, the process of random sampling is to specify the number of sentiment classifications K and the number of distributions N under each classification. By inputting the emotion latent variable features into a linear neural network, K×N posterior means are obtained, and then the Gambel-Softmax method is used to obtain the classified emotion labels and the posterior distribution of each category.

[0017] As a further improvement of the present invention, in step S3, the speech synthesis module adopts an end-to-end structure of VITS, where the input text is the first input item, and the text features are obtained through the text encoding part of the speech synthesis module. The real audio is the second input item. The emotion latent variable features extracted from the real audio are sampled by Gambel-Softmax to obtain the classified emotion labels, and the classified emotion labels are added to the text features, and the predicted audio features are obtained through the speech synthesis module.

[0018] As a further improvement of the present invention, in step S4, the prior distribution of the sentiment label is a uniform distribution; the mean square loss function of the predicted audio features and the real audio features is used as the first distance metric A, and the KL divergence of the prior distribution and the posterior distribution of the sentiment label is used as the second distance metric B. Minimizing αA + βB is used as the optimization objective, where α and β are given weights, and the parameters of the speech synthesis module are trained through gradient backpropagation.

[0019] As a further improvement of the present invention, in step S5, the generation method of the target sentiment label is:

[0020] The emotional latent variable features are extracted from the input target audio, and then the emotional latent variable features are randomly sampled to obtain emotional labels or manually specified emotional labels.

[0021] As a further improvement of the present invention, the step S5 is specifically as follows:

[0022] The text to be synthesized is the first input item, and text features are obtained through the text encoding part of the speech synthesis module. The emotional label is the second input item. After being added to the text vector, emotional speech is obtained through the decoding part of the speech synthesis module.

[0023] The beneficial effects of the present invention are:

[0024] The present invention does not require manual emotional annotation of data. It uses an unsupervised method to model the posterior distribution of the emotional latent variable features extracted from the audio, and performs random sampling to obtain multi-dimensional emotional labels, realizing high-expression emotional speech synthesis. Description of the Drawings

[0025] Figure 1 It is a schematic flow chart of unsupervised emotional speech synthesis in an embodiment of the present invention;

[0026] Figure 2 It is a training flow chart of the speech synthesis module in an embodiment of the present invention;

[0027] Figure 3 It is an inference flow chart of emotional speech in an embodiment of the present invention. Detailed Embodiments

[0028] The embodiments of the present invention will be described in detail below with reference to the drawings.

[0029] Embodiment

[0030] An unsupervised emotional speech synthesis device includes:

[0031] An emotion extraction module, configured to extract emotional latent variable features from input audio features;

[0032] Optionally, the emotion extraction module includes, but is not limited to, using a convolutional neural network and a recurrent neural network; it can be understood that its network parameters are jointly optimized with the speech synthesis module through gradient backpropagation, and are used to extract sentence-level emotional latent variable features from audio features;

[0033] A random sampling module, configured to obtain an emotional classification label and the posterior distribution of each category for the emotional latent variable features;

[0034] Optionally, the random sampling process is as follows: specify the number of emotion classifications K and the number of distributions under each classification N. By inputting the emotion latent variable features into a linear neural network, K×N posterior means are obtained, and then the Gambel-Softmax method is used to obtain the classification emotion labels and the posterior distribution of each category;

[0035] A parameter fitting module for fitting the parameters of the speech synthesis module according to the prior distribution, posterior distribution, true audio features, and predicted audio features;

[0036] Optionally, the prior distribution of the emotion labels includes, but is not limited to, a uniform distribution; the mean square loss function of the predicted audio features and the true audio features is used as the first distance metric A, and the KL divergence between the prior distribution and the posterior distribution of the emotion labels is used as the second distance metric B. Minimizing αA + βB is used as the optimization objective (α and β are given weights), and the parameters of the speech synthesis module are trained through gradient backpropagation

[0037] A speech synthesis module for generating emotional speech for the text to be synthesized and the target emotion label.

[0038] Optionally, there are two ways to generate the target emotion label. The first way is to extract the emotion latent variable features by inputting the target audio, and then randomly sample the emotion latent variable features to obtain the emotion label. The second way is to manually specify the emotion label; the text to be synthesized is the first input item, and the text features are obtained through the text encoding part of the speech synthesis module. The emotion label is the second input item, which is added to the text features, and the emotional speech is obtained through the speech synthesis module.

[0039] As Figure 1 shown, this embodiment also discloses an unsupervised emotional speech synthesis method, including the following steps:

[0040] S1. Obtain the input text and the corresponding true audio features, and extract the emotion latent variable features from the true audio features through the emotion extraction module;

[0041] Optionally, the emotion extraction module includes, but is not limited to, using a convolutional neural network or a recurrent neural network; it is understandable that its network parameters are jointly optimized with the speech synthesis module through gradient backpropagation, and are used to extract the sentence-level emotion latent variable features from the audio features;

[0042] S2. Randomly sample the emotion latent variable features obtained in S1 to obtain the classification emotion labels, and at the same time calculate the posterior distribution of each category;

[0043] Optionally, the random sampling process is as follows: specify the number of emotion classifications K and the number of distributions under each classification N. By inputting the emotion latent variable features into a linear neural network, K×N posterior means are obtained, and then the Gambel-Softmax method is used to obtain the classification emotion labels and the posterior distribution of each category;

[0044] For example, specify the number of emotion classifications as 5 and the number of distributions under each classification as 10. Input the emotion latent variable features into a linear neural network to obtain 50 posterior means μ. Then use the Gambel method to sample 50 random numbers as variances σ, and use the reparameterization trick μ+σ to generate 50 sampling values. Then calculate the posterior probability through Softmax;

[0045] S3. Obtain a speech synthesis module, and obtain predicted audio features according to the speech synthesis module, the input text of S1, and the classification emotion labels of S2;

[0046] Optionally, the speech synthesis module adopts an end-to-end structure of VITS. Among them, the text is the first input item, and text features are obtained through the text encoding part of the speech synthesis module. The real audio is the second input item. The emotion labels are obtained by sampling the emotion latent variable features extracted from the real audio through Gambel-Softmax, and the emotion labels are added to the text features, and the predicted audio features are obtained through the decoding part of the speech synthesis module;

[0047] S4. Obtain the prior distribution of the emotion labels. As Figure 2 shown, train the parameters of the speech synthesis module according to the prior distribution, the posterior distribution obtained in S2, the predicted audio features obtained in S3, and the real audio features in S1;

[0048] Optionally, the prior distribution of the emotion labels includes but is not limited to a uniform distribution; use the mean square loss function of the predicted audio features and the real audio features as the first distance metric A, and use the KL divergence of the prior distribution and the posterior distribution of the emotion labels as the second distance metric B. Minimize αA+βB as the optimization objective (α, β are given weights), and train the parameters of the speech synthesis module through gradient backpropagation;

[0049] S5. As Figure 3 shown, obtain emotion speech according to the text to be synthesized, the target emotion labels, and the speech synthesis module in S4;

[0050] Optionally, there are two ways to generate the target emotion label. The first way is to extract the emotion latent variable feature from the input target audio, and then randomly sample the emotion latent variable feature to obtain the emotion label. The second way is to manually specify the emotion label. The text to be synthesized is the first input item, and the text feature is obtained through the text encoding part of the speech synthesis module. The emotion label is the second input item, which is added to the text feature, and the emotion speech is obtained through the speech synthesis module.

[0051] For example, the emotion classification label categories are 5 (respectively: "happy", "sad", "surprised", "neutral", "angry"), and the distribution quantity of each category is 10. Adopting the first way, the emotion latent variable feature is extracted from the target audio with the emotion of "happy", and then the emotion latent variable feature is randomly sampled to obtain an emotion classification label of 5×10, and the posterior probability of the emotion of "happy" is obtained. Adopting the second way, the classification emotion label is manually specified in the way with the largest posterior probability of the emotion of "happy". According to the text to be synthesized, the classification emotion label, and the speech synthesis module, the emotion speech is generated.

[0052] Through this embodiment, there is no need for manual emotion annotation of data. The posterior distribution of the emotion latent variable feature extracted from the audio is modeled in an unsupervised way, and random sampling is performed to obtain multi-dimensional emotion labels, realizing high-expression emotion speech synthesis.

[0053] The above embodiments only represent the specific implementation manners of the present invention, and the description is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. An unsupervised emotional speech synthesis method, characterized in that, It is implemented by an unsupervised emotion speech synthesis device, and the device includes: An emotion extraction module for extracting emotion latent variable features from the input real audio features; A random sampling module for obtaining emotion classification labels and the posterior distribution of each category by using the emotion latent variable features; A parameter fitting module for fitting the parameters of the speech synthesis module according to the prior distribution, the posterior distribution, the real audio features, and the predicted audio features; A speech synthesis module for generating emotion speech for the text to be synthesized and the target emotion label; The method includes the following steps: S1. Obtain the input text and the corresponding real audio features, and extract the emotion latent variable features from the real audio features through the emotion extraction module; S2. Randomly sample the emotion latent variable features obtained in step S1 by using the random sampling module to obtain emotion classification labels, and at the same time calculate the posterior distribution of each category; S3. Obtain the speech synthesis module, and obtain the predicted audio features according to the speech synthesis module, the input text in step S1, and the emotion classification label in step S2; S4. Obtain the prior distribution of the emotion label, and use the parameter fitting module to train the parameters of the speech synthesis module according to the prior distribution, the posterior distribution obtained in step S2, the predicted audio features obtained in step S3, and the real audio features obtained in step S1; In step S4, the prior distribution of the emotion label is a uniform distribution; the mean square loss function of the predicted audio features and the real audio features is used as the first distance metric A, and the KL divergence of the prior distribution and the posterior distribution of the emotion label is used as the second distance metric B. Minimizing αA + βB is used as the optimization objective, where α and β are given weights, and the parameters of the speech synthesis module are trained through gradient backpropagation; S5. Obtain the emotion speech according to the text to be synthesized, the target emotion label, and the speech synthesis module in step S4.

2. The unsupervised emotional speech synthesis method according to claim 1, characterized in that, The emotion extraction module uses a convolutional neural network or a recurrent neural network, and its network parameters are jointly optimized with the speech synthesis module through gradient backpropagation, and are used to extract sentence-level emotion latent variable features from the real audio features.

3. The unsupervised emotional speech synthesis method according to claim 1, characterized in that, In step S2, the process of random sampling is to specify the number of emotion classifications K and the number of distributions N under each classification. By inputting the emotion latent variable features into a linear neural network, K×N posterior means are obtained, and then the Gambel-Softmax method is used to obtain the classification emotion labels and the posterior distribution of each category.

4. The unsupervised emotional speech synthesis method according to claim 3, characterized in that, In step S3, the speech synthesis module adopts the end-to-end structure of VITS, where the input text is the first input item, and the text features are obtained through the text encoding part of the speech synthesis module. The real audio is the second input item. The emotion latent variable features extracted from the real audio are sampled through Gambel-Softmax to obtain the classification emotion labels, and the classification emotion labels are added to the text features to obtain the predicted audio features through the speech synthesis module.

5. The unsupervised emotional speech synthesis method according to claim 1, characterized in that, In step S5, the generation method of the target emotion label is: Extract the emotion latent variable features by inputting the target audio, and then randomly sample the emotion latent variable features to obtain the emotion label or manually specify the emotion label.

6. The unsupervised emotional speech synthesis method according to claim 5, characterized in that, The specific steps of step S5 are as follows: The text to be synthesized is the first input item. The text features are obtained through the text encoding part of the speech synthesis module. The emotion label is the second input item. After being added to the text vector, the emotion speech is obtained through the decoding part of the speech synthesis module.

Citation Information

Patent Citations

  • Voice synthesis model training method and device, storage medium and electronic equipment

    CN112289299A

  • Speech synthesis model training method and speech synthesis method

    CN114387946A