Model training and speech synthesis methods, devices, equipment and media

By determining the emotion vector and probability vector through an emotion classification model, the problems of complex training and reliance on manual annotation in existing technologies are solved, and efficient training of emotion speech synthesis models and high-quality speech synthesis are achieved.

CN116206591BActive Publication Date: 2025-10-31BEIJING ORION STAR TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111451540.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-01
Publication Date
2025-10-31
Estimated Expiration
2041-12-01

AI Technical Summary

Technical Problem

Existing technologies require extensive manual annotation and rely heavily on the quality of speech samples when training speech synthesis models, making the training of target emotion speech synthesis models complex and difficult, and unable to guarantee the quality of speech synthesis.

Method used

By pre-configuring emotional probability vectors and acoustic features, the emotional classification model is used to determine the emotional vectors and probability vectors, train the emotional classification model and obtain the emotional extraction model, and combine non-speech factors to improve the accuracy of emotional extraction and reduce the workload of manual annotation.

Benefits of technology

This reduces the training difficulty of the target emotional speech synthesis model, improves the accuracy and stability of emotional speech synthesis, and ensures the quality of synthesized speech data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116206591B_ABST
    Figure CN116206591B_ABST
Patent Text Reader

Abstract

This invention discloses a model training and speech synthesis method, apparatus, device, and medium. Since an emotion extraction model can be obtained from the trained first emotion classification model, it facilitates subsequent training of the original speech synthesis model based on this emotion extraction model, reducing the difficulty of obtaining the speech synthesis model with the target emotion. The first emotion vector corresponding to the TTS data, determined by the first network layer included in the first emotion classification model, is determined based on the emotion weight vector corresponding to the emotion possessed by the TTS data and pre-configured emotion association parameters. This achieves the determination of the emotion vector of the TTS data by combining non-speech information, improving the accuracy of the emotion vector. This, in turn, facilitates the subsequent training of the original speech synthesis model based on this emotion vector, thereby improving the accuracy of the speech synthesis model used to synthesize synthesized speech data with emotion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech synthesis technology, and in particular to a model training and speech synthesis method, apparatus, device and medium. Background Technology

[0002] In existing technologies, to synthesize emotionally natural speech data, training a speech synthesis model typically requires pre-acquiring a sample set containing text-to-speech (TTS) data with emotion tags, and a corresponding reference sample set. This reference sample set contains natural speech data with different emotion tags. The emotion tag identifies the probability value of each pre-configured emotion in the speech data (including the TTS data in the sample set and the natural speech data in the reference sample set). The TTS data corresponds to text feature samples and a first acoustic feature. For each TTS data in the sample set, natural speech data with the corresponding emotion tag is determined from the reference sample set based on that TTS data's emotion tag. Based on each TTS data, its corresponding text feature samples, its corresponding first acoustic feature, and its corresponding natural speech data, the original speech synthesis model, the original emotion extraction model, and the original emotion classification model are jointly trained to obtain the trained speech synthesis model and emotion extraction model.

[0003] This method requires collecting a reference sample set corresponding to the sample set and labeling the natural speech data contained in the reference sample set, which increases the workload of staff and relies too much on the quality of the collected speech samples, thus failing to guarantee the conversion quality of the speech synthesis model of the target emotion. Summary of the Invention

[0004] This invention provides a model training and speech synthesis method, apparatus, device, and medium to solve the problem that the training process of existing speech synthesis models for target emotions is complex and difficult.

[0005] This invention provides a model training method, the method comprising:

[0006] For each text-to-speech (TTS) data in the first sample set, the TTS data corresponds to a first emotion probability vector and a first acoustic feature, wherein the first emotion probability vector includes the probability value of each emotion pre-configured for each TTS data.

[0007] For any TTS data, a first emotion vector corresponding to the emotion possessed by the TTS data is determined through a first network layer included in the first emotion classification model. This first emotion vector is determined based on the emotion weight vector corresponding to the emotion possessed by the TTS data and pre-configured emotion association parameters. The emotion weight vector includes the weight values ​​corresponding to each emotion association parameter, and each emotion association parameter is a non-speech-related emotion auxiliary vector used in the first emotion classification model to determine the emotion possessed by the TTS data. Furthermore, a second emotion probability vector corresponding to the TTS data is determined based on the first emotion vector through a second network layer included in the first emotion classification model. This second emotion probability vector includes the probability values ​​of each pre-configured emotion possessed by the TTS data, predicted by the first emotion classification model.

[0008] Based on the second emotion probability vector and the corresponding first emotion probability vector, the first emotion classification model is trained to obtain the trained first emotion classification model, and the emotion extraction model is determined according to the first network layer contained in the trained first emotion classification model.

[0009] This invention provides a speech synthesis method, the method comprising:

[0010] Based on the target emotion extraction model corresponding to the target emotion, the emotion vector of the target emotion is obtained;

[0011] Using a target speech synthesis model, the acoustic feature vector corresponding to the text to be processed is obtained based on the text features of the text to be processed and the emotion vector.

[0012] Using a vocoder, based on the acoustic feature vector, synthesized speech data with the target emotion corresponding to the text to be processed is obtained.

[0013] This invention provides a model training apparatus, the apparatus comprising:

[0014] The acquisition unit is used for each text-to-speech (TTS) data in the first sample set, which corresponds to a first emotion probability vector and a first acoustic feature, wherein the first emotion probability vector includes the probability value of each emotion pre-configured in the TTS data.

[0015] The training unit is configured to, for any TTS data, determine a first emotion vector corresponding to the emotion possessed by the TTS data through a first network layer included in a first emotion classification model. The first emotion vector is determined based on an emotion weight vector corresponding to the emotion possessed by the TTS data and pre-configured emotion association parameters. The emotion weight vector includes weight values ​​corresponding to each emotion association parameter, and each emotion association parameter is a non-speech-related emotion auxiliary vector used in the first emotion classification model to determine the emotion possessed by the TTS data. The training unit is also configured to, through a second network layer included in the first emotion classification model, determine a second emotion probability vector corresponding to the TTS data based on the first emotion vector. The second emotion probability vector includes the probability values ​​of each pre-configured emotion possessed by the TTS data predicted by the first emotion classification model. Based on the second emotion probability vector and the corresponding first emotion probability vector, the first emotion classification model is trained to obtain a trained first emotion classification model. Finally, an emotion extraction model is determined based on the first network layer included in the trained first emotion classification model.

[0016] This invention provides a speech synthesis device, the device comprising:

[0017] The acquisition module is used to acquire the sentiment vector of the target sentiment based on the target sentiment extraction model corresponding to the target sentiment;

[0018] The first processing module is used to obtain the acoustic feature vector corresponding to the text to be processed based on the text features of the text to be processed and the emotion vector through the target speech synthesis model.

[0019] The second processing module is used to obtain, through a vocoder, synthesized speech data corresponding to the text to be processed and carrying the target emotion, based on the acoustic feature vector.

[0020] This invention provides an electronic device, which includes at least a processor and a memory. The processor is used to execute a computer program stored in the memory to implement the steps of the model training method described above, or to implement the steps of the speech synthesis method described above.

[0021] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the model training method described above, or the steps of the speech synthesis method described above.

[0022] Since the first emotion classification model can be pre-trained using each TTS data point in the first sample set, an emotion extraction model can be obtained based on the trained first emotion classification model. This facilitates subsequent training of the original speech synthesis model based on the emotion extraction model, reducing the difficulty of obtaining the speech synthesis model with the target emotion. Furthermore, when processing TTS data, the first emotion vector corresponding to the TTS data, determined by the first network layer included in the first emotion classification model, is determined based on the emotion weight vector corresponding to the emotion in the TTS data and pre-configured emotion association parameters. This allows for the determination of the emotion vector of the TTS data by combining non-speech information, improving the accuracy of the emotion vector. This, in turn, facilitates the subsequent acquisition of the emotion vector corresponding to the emotion in the TTS data based on the emotion extraction model. This emotion vector is then incorporated into the training process of the original speech synthesis model, improving the accuracy of the speech synthesis model used to synthesize emotionally-infused speech data and enhancing the quality and stability of the acquired emotionally-infused speech data. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A schematic diagram of a training process provided in an embodiment of the present invention;

[0025] Figure 2 A schematic diagram illustrating a specific model training process provided in this embodiment of the invention;

[0026] Figure 3 A schematic diagram of a speech synthesis process provided in an embodiment of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of a model training device provided in an embodiment of the present invention;

[0028] Figure 5 This is a schematic diagram of the structure of a speech synthesis device provided in an embodiment of the present invention;

[0029] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention;

[0030] Figure 7 This is a schematic diagram of another electronic device provided in an embodiment of the present invention. Detailed Implementation

[0031] The present invention will now be described in further detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0032] For ease of understanding, some concepts involved in the embodiments of the present invention are explained below:

[0033] BERT (Bidirectional Encoder Representation from Transformers) is a text pre-trained model proposed by Google AI in 2018. It is currently the most widely applicable model in the field of NLP, achieving remarkable results across various tasks. BERT's network architecture uses a multi-layer Transformer structure. Its most significant feature is that it abandons traditional Recurrent Neural Networks (RNNs) and Convolutional Neural Networks (CNNs), using an attention mechanism to convert the distance between two words at any position into 1, effectively solving the intractable long-term dependency problem in NLP. BERT is a multi-task model consisting of two self-supervised tasks: MLM (Masked Language Model) and NSP (Next Sentence Prediction). After BERT is pre-trained on a large corpus, the pre-trained model can be applied to various NLP tasks.

[0034] To improve the quality of emotion synthesis and emotion conversion, embodiments of the present invention provide a model training and speech synthesis method, apparatus, device, and medium.

[0035] Example 1: Figure 1 A schematic diagram of a model training process provided in an embodiment of the present invention includes:

[0036] S101: For each text-to-speech (TTS) data in the first sample set, the TTS data corresponds to a first emotion probability vector and a first acoustic feature, wherein the first emotion probability vector includes the probability value of each emotion pre-configured for each TTS data.

[0037] The model training method provided in this embodiment of the invention is applied to electronic devices, which may be intelligent devices such as robots or servers.

[0038] In this application, an emotion extraction model can be trained to assist in training a speech synthesis model to synthesize speech data with specific emotions. To train the emotion extraction model, a sample set (denoted as the first sample set) can be pre-acquired for training the model. This first sample set should contain text-to-speech (TTS) data, and the emotion extraction model can be trained using the TTS data in the first sample set.

[0039] For each TTS data point in the first sample set, there is a corresponding emotion probability vector (denoted as the first emotion probability vector) and an acoustic feature (denoted as the first acoustic feature). The first emotion probability vector includes the probability values ​​of each pre-configured emotion for each TTS data point. Based on each probability value contained in the first emotion probability vector, the true emotion possessed by the TTS data can be reflected. The first acoustic feature can be the Mel-frequency cepstral coefficients (MFCCs) of the speech data, or the Mel spectrum, etc., and can be flexibly set according to actual needs during implementation.

[0040] For example, the pre-configured emotions are happiness, sadness, and anger. The first emotion probability vector corresponding to a certain TTS data is [0, 1, 0]. The first 0 from left to right in the first emotion probability vector indicates that the probability value of the TTS data having the emotion of happiness is 0. The 1 in the first emotion probability vector indicates that the probability value of the TTS data having the emotion of sadness is 1. The second 0 from left to right in the first emotion probability vector indicates that the probability value of the TTS data having the emotion of anger is 0. Based on each probability value contained in the first emotion probability vector, it can be determined that the actual emotion of the TTS data is sadness.

[0041] In one possible implementation, the first acoustic feature can be obtained through a model, such as a text-to-speech (TTS) model, or through a physical algorithm. The specific method for obtaining the first acoustic feature can be flexibly set according to actual needs and is not specifically limited here.

[0042] In one example, the first sentiment probability vector corresponding to any TTS data can be obtained through manual annotation or through a model. The specific method for obtaining the first sentiment probability vector can be flexibly set according to actual needs, and will not be elaborated here.

[0043] For example, if the first sentiment probability vector corresponding to TTS data is obtained through a model, a pre-trained sentiment classification model with high accuracy in sentiment classification tasks (denoted as the second sentiment classification model) can be used to obtain the first sentiment probability vector corresponding to each TTS data in the first sample set. Specifically, for each TTS data in the first sample set, the second sentiment classification model determines the first sentiment probability vector corresponding to the TTS data based on the first acoustic feature corresponding to the TTS data. This facilitates subsequent training of the first sentiment classification model, reduces the workload required for manually labeling the first sentiment probability vector corresponding to the TTS data, achieves unsupervised acquisition of the first sentiment probability vector corresponding to the TTS data, and improves the efficiency of model training.

[0044] S102: For any TTS data, a first emotion vector corresponding to the emotion possessed by the TTS data is determined through a first network layer included in the first emotion classification model. The first emotion vector is determined based on the emotion weight vector corresponding to the emotion possessed by the TTS data and pre-configured emotion association parameters. The emotion weight vector includes the weight values ​​corresponding to each emotion association parameter. Each emotion association parameter is a non-speech-related emotion auxiliary vector used in the first emotion classification model to determine the emotion possessed by the TTS data. A second emotion probability vector corresponding to the TTS data is determined based on the first emotion vector through a second network layer included in the first emotion classification model. The second emotion probability vector includes the probability values ​​of each pre-configured emotion possessed by the TTS data predicted by the first emotion classification model.

[0045] In this application, an emotion classification model (denoted as the first emotion classification model) is obtained in advance. The trained first emotion classification model is obtained by training the first emotion classification model. Then, an emotion extraction model is obtained based on the trained first emotion classification model.

[0046] In the specific implementation process, any TTS data point from the first sample set is acquired and input into the first sentiment classification model. The first sentiment classification model then processes the TTS data accordingly, obtaining its output, which is used to train the first sentiment classification model.

[0047] In one example, the first sentiment classification model includes a first network layer and a second network layer to accurately process the input TTS data. The first network layer is connected to the second network layer. The first network layer processes the input TTS data to obtain a first sentiment vector corresponding to the sentiment of the TTS data. The second network layer processes the first sentiment vector obtained by the first network layer to obtain a sentiment probability vector (denoted as the second sentiment probability vector) corresponding to the TTS data.

[0048] In one possible implementation, emotions can generally be expressed through voice, facial expressions, and body posture. Therefore, the emotions expressed by a person can be determined not only by the voice data they produce but also through non-vocal factors. That is, emotions can be expressed through means other than voice data, such as facial expressions and body posture, which can be obtained through image acquisition and processing. Furthermore, when expressing different emotions, not only will the voice data they produce differ, but the non-vocal factors they use while speaking will also vary. For example, compared to the expressions and movements of a person speaking sadly, a person speaking excitedly will have a smiling expression and may also use gestures while speaking. Therefore, in this embodiment of the invention, the first network layer can be configured with a non-speech-related emotion auxiliary vector (denoted as emotion association parameter) for determining the emotion possessed by the input TTS data. That is, the first emotion classification model includes pre-configured emotion association parameters. Based on these pre-configured emotion association parameters and the emotion weight vector corresponding to the emotion possessed by the input TTS data, the first emotion vector is obtained. This achieves accurate extraction of the emotion possessed by the input TTS data by combining non-speech factors, thereby improving the accuracy of emotion extraction. Here, each emotion association parameter is the weight value corresponding to each emotion association parameter contained in the emotion weight vector of the first emotion classification model.

[0049] In one example, the first network layer may include a feature extraction layer, a first encoding sub-network, and a second encoding sub-network. The first encoding sub-network is connected to both the feature extraction layer and the second encoding sub-network. The feature extraction layer contains multiple sub-networks connected sequentially. This feature extraction layer extracts features from the first acoustic features corresponding to the input TTS data, and the output data of the last sub-network in the feature extraction layer is the emotional feature extracted based on the first acoustic features corresponding to the TTS data. The first encoding sub-network processes the output data of the feature extraction layer to obtain a basic emotional vector corresponding to the TTS data. This basic emotional vector includes a speech-related emotional auxiliary vector used to determine the emotion possessed by the TTS data. The second encoding sub-network contains pre-configured emotional association parameters. The second encoding sub-network processes the basic emotional vector obtained by the first encoding sub-network to obtain an emotional weight vector corresponding to the emotion possessed by the TTS data, and determines the first emotional vector corresponding to the emotion possessed by the TTS data based on the emotional weight vector and the pre-configured emotional association parameters.

[0050] In the specific implementation process, any TTS data point from the first sample set is acquired, and the first acoustic feature corresponding to the TTS data is input into the first sentiment classification model. Through the feature extraction layer of the first network layer in the first sentiment classification model, based on the first acoustic feature corresponding to the TTS data, the output data of the last sub-network in the feature extraction layer is obtained. Through the first encoding sub-network (such as a Gate Recurrent Unit (GRU)) in the first network layer, based on the output data, the basic sentiment vector corresponding to the TTS data is obtained. Then, through the second encoding sub-network in the first network layer, based on the basic sentiment vector, the sentiment weight vector (Attention) corresponding to the sentiment of the TTS data is obtained, and based on the sentiment weight vector and the various sentiment association parameters (Emotional Tokens) pre-configured in the second encoding sub-network, the first sentiment vector (EmotionEmbedding) corresponding to the sentiment of the TTS data is determined.

[0051] In one example, when obtaining the sentiment weight vector corresponding to the sentiment of the TTS data based on the base sentiment vector through the second encoding subnetwork contained in the first network layer, the sentiment weight vector corresponding to the sentiment of the TTS data can be obtained through the multihead self-attention mechanism (MLP attention) network in the second encoding subnetwork based on the base sentiment vector.

[0052] It should be noted that each sentiment association parameter has a corresponding weight value in the sentiment weight vector, so that the intensity of each sentiment association parameter in the sentiment to be expressed can be represented by the weight value corresponding to each sentiment association parameter.

[0053] In one example, when determining the first emotion vector corresponding to the emotion of the TTS data using the second encoding subnetwork included in the first network layer, based on the emotion weight vector and the various emotion association parameters pre-configured in the second encoding subnetwork, each emotion association parameter in the second encoding subnetwork corresponds to a weight value in the emotion weight vector. For example, if the second encoding subnetwork contains M emotion association parameters, each emotion association parameter corresponds to a weight value in the emotion weight vector, such as the first emotion association parameter corresponding to the first weight value in the emotion weight vector, the second emotion association parameter corresponding to the second weight value in the emotion weight vector, and so on. For each emotion association parameter in the second encoding subnetwork, the emotion association parameter is multiplied by its corresponding weight value in the emotion weight vector to obtain a product vector. Based on each obtained product vector, the first emotion vector corresponding to the emotion of the TTS data is determined.

[0054] After obtaining the first emotion vector based on the above embodiments, the second emotion probability vector corresponding to the TTS data is determined by the second network layer contained in the first emotion classification model, based on the first emotion vector.

[0055] In one example, the second network layer in the first sentiment classification model can be composed of a deep neural network (DNN). Using this DNN, a second sentiment probability vector corresponding to the TTS data is obtained based on the first sentiment vector determined by the first network layer. This second sentiment probability vector includes the probability values ​​of each pre-configured sentiment in the TTS data predicted by the first sentiment classification model.

[0056] S103: Based on the second emotion probability vector and the corresponding first emotion probability vector, train the first emotion classification model to obtain the trained first emotion classification model, and determine the emotion extraction model according to the first network layer contained in the trained first emotion classification model.

[0057] After obtaining the second sentiment probability vector based on the above embodiments, a loss value can be determined based on the second sentiment probability vector and the corresponding first sentiment probability vector. Then, based on the determined loss value, the parameter values ​​of the parameters in the first sentiment classification model are adjusted to obtain the trained first sentiment classification model.

[0058] Since the accuracy of a model is directly proportional to the number of parameters it contains, high-accuracy models typically contain a very large number of parameters, such as hundreds of thousands or millions. If the accuracy of the first sentiment classification model is improved by including a large number of parameters, it will consume a significant amount of computational resources during task execution (such as obtaining the second sentiment probability vector), making training and application of the first sentiment classification model in tasks inconvenient. Therefore, in this embodiment of the invention, knowledge distillation can be used to compress the second sentiment classification model, which contains a large number of parameters and has high accuracy in sentiment classification tasks. This reduces the number of parameters in the second sentiment classification model without affecting its accuracy. Then, based on the compressed model, the first sentiment classification model is obtained, achieving the same accuracy as the second model while reducing the training difficulty and computational resource consumption during task execution, and also simplifying deployment of the first sentiment classification model in smart devices.

[0059] To compress the second sentiment classification model, each first target network can be identified from the networks included in the second sentiment classification model. The first target network represents the incompressible key networks in the second sentiment classification model. Furthermore, to ensure that the first sentiment classification network has similar accuracy to each first target network in the second sentiment classification model, it needs to include subnetworks that can implement the functions of each first target network. That is, each subnetwork in the first sentiment classification model corresponds to a first target network in the second sentiment classification model, and the output data of each subnetwork must be identical to the output data of its corresponding first target network. Then, for each TTS data point in the first sample set, the TTS data is input into both the first and second sentiment classification models to obtain the output data of each subnetwork in the first sentiment classification model and the output data of each first target network in the second sentiment classification model. Based on the output data of each subnetwork and each first target network, the second sentiment classification model is compressed.

[0060] In one example, considering that the second sentiment classification model typically contains a large number of parameters used for feature extraction from the input TTS data, the target networks (denoted as the first target networks) in the second sentiment classification model can be determined based on the network layers to which the parameters used for feature extraction belong.

[0061] In determining the first target networks in the second sentiment classification model based on the network layers to which the parameters used for feature extraction belong, the number of network layers separating each pair of adjacent first target networks in the second sentiment classification model should be as uniform as possible to ensure the accuracy and stability of the compressed model. For example, if the second sentiment classification model contains 12 network layers for feature extraction, and these 12 layers are connected sequentially, layers 3, 6, 9, and 12 can be identified as the first target networks, namely First Target Network 1, First Target Network 2, First Target Network 3, and First Target Network 4, with each pair of adjacent first target networks separated by two network layers. Correspondingly, the feature extraction layer in the first sentiment classification model contains 4 sub-networks, which are connected sequentially. In this model, according to the direction of input data transmission, the first sub-network corresponds to the first target network 1 in the second sentiment classification model, the second sub-network corresponds to the second target network 2 in the second sentiment classification model, the third sub-network corresponds to the second target network 3 in the second sentiment classification model, and the fourth sub-network corresponds to the second target network 4 in the second sentiment classification model.

[0062] It should be noted that the number of the first target network can be set differently depending on the scenario. If you want to compress the number of parameters in the model as much as possible, you can set the number to be smaller; if you want to ensure the accuracy of the compressed model as much as possible, you can set the number to be larger. The specific number of the first target network can be flexibly configured according to actual needs, and no specific limitation is made here.

[0063] In the specific implementation process, for each TTS data point contained in the first sample set, the TTS data is input into both the first sentiment classification model and the second sentiment classification model. Using the pre-trained second sentiment classification model, based on the first acoustic features corresponding to the TTS data, the output data of each first target network contained in the second sentiment classification model is obtained; and using the first sentiment classification model, based on the first acoustic features corresponding to the TTS data, the output data of each sub-network contained in the feature extraction layer is obtained. Based on the second sentiment probability vector and the first sentiment probability vector corresponding to the TTS data, a first loss value is determined. And based on the output data of each sub-network and the output data of the first target network corresponding to each sub-network, a second loss value is determined. Based on the first and second loss values, the parameter values ​​in the first sentiment classification model are adjusted to obtain the trained first sentiment classification model. This allows the first sentiment classification model to achieve similar accuracy to the second sentiment classification model with fewer parameters, reducing the training difficulty of the first sentiment classification model.

[0064] In practice, when adjusting the parameter values ​​of the first sentiment classification model based on the first loss value and the second loss value, a gradient descent algorithm can be used to backpropagate the gradients of the parameters in the first sentiment classification model, thereby updating the parameter values. Specifically, the parameter values ​​of the first sentiment classification model can be adjusted based on the weighted sum of the first and second loss values.

[0065] In this embodiment of the invention, the first sample set contains a large amount of TTS data. For each TTS data, the above steps are performed. When a preset convergence condition (denoted as the first convergence condition) is met, the first sentiment classification model is trained. The preset first convergence condition can be that the weights corresponding to each TTS data in the current iteration are less than a set first loss threshold, or that the number of iterations for training the first sentiment classification model reaches a set first maximum number of iterations. These conditions can be flexibly set in specific implementations and are not specifically limited here.

[0066] In the specific implementation process, when training the first sentiment classification model, the TTS data in the first sample set is divided into training samples and test samples. The first sentiment classification model is first trained based on the training samples, and then the reliability of the trained first sentiment classification model is verified based on the test samples.

[0067] Once the first emotion classification model has been trained based on the above embodiments, an emotion extraction model can be obtained based on the first network layer contained in the first emotion classification model. The emotion vector output by the emotion extraction model can then be used to assist in the training of the speech synthesis model, thereby improving the ability and accuracy of the speech synthesis model to synthesize emotional speech data.

[0068] Example 2: To improve the ability of speech synthesis models to synthesize emotional speech data, based on the above examples, this embodiment of the invention further includes:

[0069] For each TTS data, based on the sentiment extraction model, a second sentiment vector is obtained that contains the sentiment features of the TTS data.

[0070] Based on the text feature samples corresponding to TTS data, the first acoustic feature of TTS data, and the second emotion vector, the original speech synthesis model and the emotion extraction model are jointly trained to obtain the target speech synthesis model and the target emotion extraction model.

[0071] In specific implementation, after obtaining the emotion extraction model based on the above embodiments, for each TTS data, the emotion vector (denoted as the second emotion vector) of the emotion possessed by the TTS data can be obtained based on the emotion extraction model, that is, the emotion features of the emotion possessed by the TTS data can be obtained. Subsequently, the original speech synthesis model is trained based on the second emotion vector and the TTS data.

[0072] In one example, for each TTS data point, the second sentiment vector of the sentiment possessed by the TTS data is obtained based on the sentiment extraction model in at least one of the following ways:

[0073] Method 1: For each TTS data point, there is a corresponding sentiment tag (for ease of description, denoted as the target sentiment tag), which identifies the sentiment possessed by the TTS data. Based on the sentiment tag of the TTS data, from the first sample set, determine each TTS data point with the corresponding sentiment tag (for ease of description, denoted as reference TTS data). Obtain the reference sentiment vector corresponding to each reference TTS data point, and then determine the second sentiment vector of the sentiment possessed by the TTS data point based on each reference sentiment vector. This effectively avoids the problem that the sentiment extraction model cannot accurately identify TTS data outside the first sample set, thus reducing the accuracy of the obtained second sentiment vector, improving the accuracy of the obtained second sentiment vector, and eliminating the need to collect other speech data, reducing the workload of staff. The reference sentiment vector corresponding to any reference TTS data point is obtained by processing the reference TTS data point using the sentiment extraction model. The sentiment tag corresponding to the TTS data point can be obtained through manual annotation or through model or physical algorithm identification, such as a sentiment recognition model.

[0074] In this embodiment of the invention, the emotion tag can be represented by numbers, strings, or other forms. Any representation that can uniquely identify the emotion can be applied to this embodiment of the invention.

[0075] It should be noted that, to save time in obtaining reference sentiment vectors, each TTS data point in the first sample set can be pre-processed using a sentiment extraction model to obtain and save the corresponding reference sentiment vector for each TTS data point. Subsequently, the reference sentiment vector for each reference TTS data point can be obtained based on the saved reference sentiment vectors for each TTS data point. Alternatively, to save storage space on electronic devices, after determining the reference TTS data points, the sentiment extraction model can be used to process each reference TTS data point separately to obtain the corresponding reference sentiment vector for each reference TTS data point. The method for obtaining the reference sentiment vector for each reference TTS data point can be flexibly set according to actual needs and is not specifically limited here.

[0076] Method 2: During speech data collection, random speech data with different emotions is also collected. This random speech data is not part of the TTS data in the first sample set. Emotional tags are obtained for each random speech data. For each TTS data, based on its emotional tag, any random speech data with the corresponding emotional tag is obtained. This random speech data is then processed using an emotion extraction model to obtain a reference emotion vector. This reference emotion vector is then designated as the second emotion vector. This method, using random speech data, increases the diversity of the obtained second emotion vectors, thereby improving the robustness of the target speech synthesis model and the target emotion extraction model.

[0077] The method for obtaining sentiment tags has been described in the above embodiments and will not be repeated here.

[0078] Method 3: Since the sentiment vector is determined based on the sentiment weight vector and various sentiment-related parameters, for each TTS data point, the sentiment tag corresponding to that TTS data can be obtained to determine the sentiment it possesses, and the corresponding sentiment weight vector can be obtained. Then, at least one weight value in this sentiment weight vector is adjusted. Based on the adjusted sentiment weight vector and the various sentiment-related parameters in the sentiment extraction model, a second sentiment vector representing the sentiment of that TTS data is determined. This allows for flexible adjustment of the weight values ​​in the sentiment weight vector according to the needs of the staff, thereby increasing the diversity of the obtained second sentiment vector and thus improving the robustness of the target speech synthesis model and the target sentiment extraction model.

[0079] To save time in obtaining sentiment weight vectors, a sentiment extraction model can be used to pre-process one TTS data point for each sentiment in the first sample set, obtaining and saving the sentiment weight vector corresponding to each sentiment. Subsequently, based on the saved sentiment weight vectors for each sentiment, the sentiment weight vector corresponding to the sentiment of the TTS data can be obtained. Alternatively, to save storage space on electronic devices, after determining the sentiment of the TTS data, any TTS data point with that sentiment (denoted as auxiliary TTS data) can be obtained from the first sample set. The sentiment extraction model can then be used to process this auxiliary TTS data to obtain the sentiment weight vector corresponding to the sentiment of the auxiliary TTS data. The method for obtaining sentiment weight vectors can be flexibly set according to actual needs and is not specifically limited here.

[0080] It should be noted that the process of obtaining the sentiment weight vector through the sentiment extraction model is similar to the process of obtaining the sentiment weight vector through the first sentiment classification model, and will not be elaborated here.

[0081] In one possible implementation, at least one weight value contained in the emotion weight vector can be adjusted based on human experience.

[0082] Method 4: Based on human experience or needs, pre-set emotion weight vectors (denoted as preset emotion weight vectors) corresponding to each emotion. For each TTS data point, determine the emotion possessed by the TTS data based on its corresponding emotion tag, and obtain the preset emotion weight vector corresponding to that emotion. Based on the preset emotion weight vector and the various emotion association parameters included in the emotion extraction model, determine the second emotion vector of the emotion possessed by the TTS data. This allows for flexible configuration of the obtained second emotion vector according to the needs of the staff, enabling the generation of second emotion vectors for any emotion based on the staff's requirements. This increases the diversity of the obtained second emotion vectors, thereby improving the robustness of the obtained target speech synthesis model and target emotion extraction model.

[0083] After obtaining the second emotion vector based on the above embodiments, this second emotion vector can be applied to the training process of the original speech synthesis model, thereby improving the speech synthesis model's ability to synthesize emotional speech data. Since the input of the speech synthesis model is text features and the output is acoustic features, each TTS data also corresponds to a text feature sample. For each TTS data, the text feature sample corresponding to the TTS data and the second emotion vector corresponding to the TTS data are input into the original speech synthesis model. The acoustic features corresponding to the input text feature sample are obtained through the original speech synthesis model. Based on the acoustic features and the corresponding first acoustic features, a third loss value is determined. Based on the third loss value, the original speech synthesis model and the emotion extraction model are jointly trained, and the parameter values ​​in the original speech synthesis model and the emotion extraction model are adjusted to obtain the target speech synthesis model and the target emotion extraction model.

[0084] The order in which the parameter values ​​in the original speech synthesis model and the emotion extraction model are adjusted can be determined based on the input position of the second emotion vector in the original speech synthesis model, that is, the embedding position of the second emotion vector in the original speech synthesis model. For example, the second emotion vector can be embedded after the output of the encoder in the original speech synthesis model and before the input of the projection layer in the original speech synthesis model.

[0085] In this embodiment of the invention, the first sample set contains a large amount of TTS data. For each TTS data, the above steps are performed. When a preset convergence condition (denoted as the second convergence condition) is met, the training of the original speech synthesis model and the emotion extraction model is completed. The second convergence condition is based on the following: the sum of the weight values ​​corresponding to each TTS data in the second sample set in the current iteration is less than a set second loss threshold; the number of iterations for training the original speech synthesis model or the emotion extraction model reaches a set second maximum number of iterations, etc. These conditions can be flexibly set in specific implementations and are not specifically limited here.

[0086] As one possible implementation, when training the original speech synthesis model, the TTS data in the first sample set can be divided into training samples and test samples. First, the original speech synthesis model and the emotion extraction model are jointly trained based on the training samples, and then the reliability of the trained speech synthesis model and emotion extraction model is verified based on the test samples.

[0087] Example 3: In order to accurately obtain the second sentiment classification model, based on the above embodiments, in this embodiment of the invention, the second sentiment classification model is obtained in the following manner:

[0088] Obtain any speech sample from the second sample set, wherein each speech sample corresponds to a first emotion label and a speech type label. The first emotion label is used to identify the emotion of the speech sample, and the speech type label is used to identify the speech type to which the speech sample belongs. The speech type includes at least one of TTS type and speech emotion recognition SER type.

[0089] Using the original emotion classification model, based on the second acoustic features corresponding to the speech samples, a third emotion probability vector and a type probability vector are determined for each speech sample. The third emotion probability vector includes the probability value of each pre-configured emotion that the speech sample has, as determined by the original emotion classification model. The type probability vector includes the probability value of each pre-configured speech type that the speech sample belongs to, as determined by the original emotion classification model. Based on the third emotion probability vector and the first emotion label, and the type probability vector and the speech type label, the original emotion classification model is trained to obtain a second emotion classification model.

[0090] To facilitate the acquisition of the first emotion probability vector or the output data of each first target network, in this embodiment of the invention, a sample set (denoted as the second sample set) for training the second emotion classification model is pre-collected. The original emotion classification model is trained using the speech samples included in the second sample set, thereby obtaining the second emotion classification model based on the trained model. Each speech sample corresponds to an emotion label (denoted as the first emotion label). The first emotion label is used to identify the emotion possessed by the speech sample.

[0091] In one example, the speech samples in the second sample set may contain a large amount of TTS data to improve the ability of the second emotion classification model to recognize emotions from TTS data. The TTS data in the second sample set can be completely identical to the TTS data in the first sample set, meaning the second sample set contains all the TTS data from the first sample set and contains no other TTS data besides those in the first sample set; or it can be completely different, meaning the second sample set does not contain any TTS data from the first sample set. Of course, the TTS data in the second sample set can also be partially identical to the TTS data in the first sample set. For example, the second sample set may contain all the TTS data from the first sample set and include other TTS data besides those in the first sample set; or the second sample set may contain some of the TTS data from the first sample set and include other TTS data besides those in the first sample set; or the first sample set may contain all the TTS data from the second sample set and include other TTS data besides those in the second sample set. Specifically, when collecting the first and second sample sets, the settings can be flexibly configured according to actual needs, and no specific limitations are made here.

[0092] In another example, since acquiring a large amount of TTS data is costly and difficult, in this embodiment of the invention, the speech samples in the second sample set may also include Speech Emotion Recognition (SER) data to reduce the difficulty and cost of acquiring a large number of speech samples. To distinguish the speech types of the speech samples in the second sample set, each speech sample also has a corresponding speech type label, which identifies the speech type to which the speech sample belongs. For example, the speech type label for TTS data is TTS, and the speech type label for SER data is SER. The number of SER data in the second sample set can be greater than or significantly greater than the number of TTS data. The speech type label can be represented by numbers, strings, or other forms; any representation that uniquely identifies the speech type can be applied to this embodiment of the invention.

[0093] In the specific implementation process, any speech sample from the second sample set is obtained, and the acoustic features corresponding to this speech sample (denoted as the second acoustic feature) are input into the original emotion classification model. The original emotion classification model processes the second acoustic feature corresponding to the input speech sample to obtain the emotion probability vector corresponding to the speech sample (denoted as the third emotion probability vector). This third emotion probability vector includes the probability values ​​of each pre-configured emotion determined by the original emotion classification model for the speech sample. Based on the third emotion probability vector and the first emotion label corresponding to the speech sample, an emotion loss value is determined. Based on this emotion loss value, the parameter values ​​in the original emotion classification model are adjusted to obtain a trained emotion classification model. Then, based on this trained emotion classification model, a second emotion classification model is determined.

[0094] In one example, considering that TTS data requires text data and high-precision TTS technology, while SER data is obtained through natural speech data, collecting a large amount of SER data is relatively easy, while collecting a large amount of TTS data is difficult and costly. This results in a larger quantity of SER data than TTS data in the second sample set. Consequently, the sentiment classification model trained based on the speech samples in this second sample set affects the accuracy of the subsequent training of the sentiment classification model in processing SER data. Since the second sentiment classification model determined by this trained model is primarily used for processing TTS data, when processing the second acoustic features corresponding to the input speech sample using the original sentiment classification model, a type probability vector corresponding to the speech sample can be obtained. This type probability vector, along with the corresponding speech type label, can be used to train the original sentiment classification model, further improving its accuracy. The type probability vector includes the probability values ​​determined by the original sentiment classification model for each pre-configured speech type. Based on the probability vector of the type and the speech type label corresponding to the speech sample, a type loss value is determined. Based on this emotion loss value and the type loss value, the parameter values ​​in the original emotion classification model are adjusted to obtain a trained emotion classification model. Then, based on this trained emotion classification model, a second emotion classification model is determined.

[0095] For example, the original emotion classification model may include a feature extraction layer, a category recognition layer, and an emotion classification layer. The feature extraction layer is connected to both the category recognition layer and the emotion classification layer. The feature extraction layer extracts features from the input speech sample to obtain the corresponding emotion feature vector. The category recognition layer processes the emotion feature vector obtained by the feature extraction layer to obtain the corresponding type probability vector for the speech sample. The emotion classification layer processes the emotion feature vector obtained by the feature extraction layer to obtain the corresponding third emotion probability vector for the speech sample. In specific implementation, the feature extraction layer in the original emotion classification model extracts features from the second acoustic features corresponding to the input speech sample to obtain the corresponding emotion feature vector. Based on the emotion feature vector, the emotion classification layer in the original emotion classification model obtains the corresponding third emotion probability vector for the speech sample, and based on the emotion feature vector, the category recognition layer in the original emotion classification model obtains the corresponding type probability vector for the speech sample.

[0096] To improve the accuracy of the trained sentiment classification model in processing TTS data, it is necessary for the model to treat input speech samples of different speech types indiscriminately. In this embodiment of the invention, when determining the type loss value based on the type probability vector and the corresponding speech type label of the speech sample, the gradient inversion method can be used to maximize the classification error function based on the type probability vector and the corresponding speech type label to determine the type loss value. This ensures that the trained sentiment classification model cannot distinguish between speech samples of different speech types, thereby transferring the ability of the trained sentiment classification model to process SER data to process TTS data.

[0097] For example, when training the original emotion classification model based on the third emotion probability vector and its corresponding first emotion label, and the type probability vector and its corresponding speech type label, the emotion loss value can be determined using the cross-entropy loss function based on the third emotion probability vector and the first emotion label; and the type loss value can be determined by maximizing the classification error function using gradient inversion based on the type probability vector and the speech type label corresponding to the speech sample. Then, based on the emotion loss value and the type loss value, the parameter values ​​in the original emotion classification model are adjusted to obtain the trained emotion classification model. Based on the trained emotion classification model, a second emotion classification model is determined. For example, the parameter values ​​in the original emotion classification model are adjusted based on the weighted sum of the emotion loss value and the type loss value to obtain the trained emotion classification model.

[0098] Since the second sample set contains a large number of speech samples, the above steps can be performed for each speech sample. When the preset convergence condition (denoted as the third convergence condition) is met, the emotion classification model is trained. The second convergence condition is based on the following: the weights corresponding to each speech sample in the current iteration are less than the set third loss threshold; the number of iterations for training the emotion classification model reaches the set third maximum number of iterations. These conditions can be flexibly set in practice and are not specifically limited here.

[0099] As one possible implementation, when training the emotion classification model, the speech samples in the second sample set can be divided into training samples and test samples. The emotion classification model is first trained based on the training samples, and then the reliability of the trained emotion classification model is verified based on the test samples.

[0100] In one possible implementation, when determining the second emotion classification model based on the trained emotion classification model, the second emotion classification model can be determined based on the feature extraction layer and the emotion classification layer contained in the trained emotion classification model.

[0101] Once the second sentiment classification model is determined, each first target network can be determined based on the sub-networks contained in the feature extraction layer of the second sentiment classification model. That is, the feature extraction layer of the second sentiment classification model includes each first target network.

[0102] Example 4: In order to accurately obtain the second sentiment classification model, based on the above examples, in this embodiment of the invention, the original sentiment classification model is obtained in the following way:

[0103] Obtain any two SER data points from the third sample set, where each SER data point corresponds to a second sentiment label; using the original sentiment discrimination model, based on the third acoustic features corresponding to the two SER data points respectively, determine whether the sentiments of the two SER data points are consistent; and based on the second sentiment labels corresponding to the two SER data points respectively, determine whether the sentiments of the two SER data points are consistent; and based on the first and second results, train the original sentiment discrimination model to obtain the original sentiment classification model.

[0104] In this embodiment of the invention, the original emotion classification model can be set based on work experience or obtained through training. In specific implementation, it can be flexibly set according to needs.

[0105] In one example, if the original sentiment classification model is obtained through training, a sample set (denoted as the third sample set) is pre-collected for training this model. Based on the large amount of SER data contained in this third sample set, the original sentiment discrimination model is trained. Then, based on the trained sentiment discrimination model, the original sentiment classification model is obtained. This pre-trains the parameter values ​​in the feature extraction layer of the original sentiment classification model, improving its ability to extract sentiment features from input speech data (including TTS data and SER data). Each SER data point in the third sample set corresponds to a sentiment label (denoted as the second sentiment label). This second sentiment label identifies the speech type to which the SER data belongs.

[0106] It should be noted that the SER data in the third sample set can be completely identical to the SER data in the second sample set, meaning the third sample set contains all the SER data from the second sample set and contains no other SER data besides those from the second sample set; or they can be completely different, meaning the third sample set does not contain any SER data from the second sample set. Of course, the SER data in the third sample set can also be partially identical to the SER data in the second sample set. For example, the third sample set may contain all the SER data from the second sample set and contain other SER data besides those from the second sample set; or, the third sample set may contain some of the SER data from the second sample set and contain other SER data besides those from the second sample set; or, the second sample set may contain all the SER data from the third sample set and contain other SER data besides those from the third sample set. The specific settings for collecting the second and third sample sets can be flexibly configured according to actual needs, and no specific limitations are imposed here.

[0107] In the specific implementation process, any two SER data points are acquired from the third sample set and input into the original sentiment discrimination model. Using this original sentiment discrimination model, based on the acoustic features corresponding to the two SER data points (denoted as the third acoustic feature), the consistency of the sentiments possessed by the two SER data points is determined (denoted as the first result). Furthermore, based on the second sentiment tags corresponding to the two SER data points, the consistency of the sentiments possessed by the two SER data points is determined (denoted as the second result).

[0108] For example, the original sentiment discrimination model may include a feature extraction layer and a sentiment discrimination layer. The feature extraction layer is connected to the sentiment discrimination layer. The feature extraction layer is used to extract features from the two input SER data respectively, obtaining the sentiment feature vectors corresponding to the two SER data respectively. The sentiment discrimination layer is used to process the sentiment feature vectors corresponding to the two SER data respectively obtained by the feature extraction layer, and obtain a first result of whether the two SER data have the same sentiment. In specific implementation, the feature extraction layer in the original sentiment discrimination model can determine the sentiment feature vectors corresponding to the two SER data respectively based on the third acoustic features corresponding to the two SER data respectively. The sentiment discrimination layer in the original sentiment discrimination model determines the first result of whether the two SER data have the same sentiment based on the two sentiment feature vectors.

[0109] In one possible implementation, a second result can be determined by whether the emotions of the two SER data are consistent, based on whether the second sentiment tags corresponding to the two SER data are the same.

[0110] After obtaining the first and second results based on the above embodiments, a loss value (denoted as the fourth loss value) can be determined based on the first and second results. Then, according to the third loss value, the parameter values ​​in the original sentiment discrimination model are adjusted to obtain the trained sentiment discrimination model.

[0111] Since the third sample set contains a large amount of SER data, the above steps are performed for any two SER data points in the third sample set. The sentiment discrimination model is trained when a preset convergence condition (denoted as the fourth convergence condition) is met. The preset second convergence condition can be that the sum of the fourth loss values ​​corresponding to any two SER data points is less than a set fourth loss value threshold, or that the number of iterations for training the sentiment discrimination model reaches a set fourth maximum number of iterations. These conditions can be flexibly set in practice and are not specifically limited here.

[0112] As one possible implementation, when training the sentiment discrimination model, the SER data in the third sample set is divided into training samples and test samples. The sentiment discrimination model is first trained based on the training samples, and then the reliability of the trained sentiment discrimination model is verified based on the test samples.

[0113] Once the trained sentiment discrimination model is obtained based on the above embodiments, the original sentiment classification model can be determined based on the feature extraction layer contained in the trained sentiment discrimination model.

[0114] For example, the original emotion classification model can be determined by connecting the emotion classification layer and the category recognition layer to the feature extraction layer contained in the trained emotion discrimination model.

[0115] In one possible implementation, the first sentiment classification model can be pre-set or obtained through pre-training. To reduce the difficulty and workload of subsequent training of the first sentiment classification model, it can be obtained through pre-training.

[0116] For example, after obtaining the trained sentiment discrimination model based on the above embodiments, the original sentiment model can be trained based on the SER data contained in the third sample set and the pre-trained sentiment discrimination model to obtain the trained sentiment model. Then, based on the trained sentiment model, a first sentiment classification model is obtained. This reduces the workload of the first sentiment classification model in learning to extract the sentiment features contained in the input speech data, reduces the difficulty for the first sentiment classification model to accurately extract the sentiment features contained in the input speech data, and ensures that the trained sentiment model's ability to extract sentiment features from the input speech data is the same as the trained sentiment discrimination model's ability to extract sentiment features from the input speech data.

[0117] In one possible implementation, the original sentiment model includes a feature extraction layer for extracting sentiment features from the input SER data. The number of sub-networks included in this feature extraction layer can be set based on practical experience. Alternatively, an original sentiment model containing a different number of sub-networks can be set. Then, each original sentiment model is trained separately, and based on the training results, it is determined which trained sentiment model is used to determine the first sentiment classification model. The specific method for setting the number of sub-networks included in the original sentiment model can be flexibly set according to actual needs and will not be elaborated here.

[0118] It should be noted that the original sentiment model contains fewer subnetworks than the pre-trained sentiment discrimination model.

[0119] To ensure that the original sentiment model achieves the same accuracy as the trained sentiment discrimination model—that is, the same ability to extract sentiment features—a target network (denoted as the second target network) can be determined from the sub-networks of the trained sentiment discrimination model, based on the number of sub-networks (denoted as the third target network) in the original sentiment model. Each third target network corresponds to one of the second target networks in the pre-trained sentiment discrimination model.

[0120] In one example, considering that a pre-trained sentiment discrimination model typically contains a large number of parameters used for feature extraction from the input SER data, a second target network can be determined from each sub-network (e.g., a dense layer) within the feature extraction layer of the original sentiment model, based on the number of third target networks included in the original sentiment model. Each third target network corresponds to one second target network in the pre-trained sentiment discrimination model.

[0121] After determining the second target network in the pre-trained sentiment discrimination model, any SER data from the third sample set is obtained and input into the pre-trained sentiment discrimination model. Using this pre-trained sentiment discrimination model, based on the third acoustic features corresponding to the input SER data, the output data of each second target network included in the feature extraction layer of the sentiment discrimination model is determined. Then, using the original sentiment model, based on the third acoustic features corresponding to the SER data, the output data of each third target network included in the feature extraction layer of the original sentiment model is determined. Next, based on the output data of each third target network and the output data of the corresponding second target network, a loss value (denoted as the fifth loss value) is determined. Based on this fifth loss value, the parameter values ​​in the original sentiment model are adjusted to obtain the trained sentiment model. Finally, based on the trained sentiment model, the first sentiment classification model is obtained.

[0122] Since the third sample set contains a large amount of SER data, the above steps for training the original sentiment model are performed for each SER data. The sentiment model training is complete when a preset convergence condition (denoted as the fifth convergence condition) is met. The preset fifth convergence condition can be that the sum of the fifth loss values ​​corresponding to each SER data in the current iteration is less than a set fifth loss value threshold, or that the number of iterations for training the sentiment model reaches a set fifth maximum number of iterations. These conditions can be flexibly set in practice and are not specifically limited here.

[0123] As one possible implementation, when training the sentiment model, the SER data in the third sample set can be divided into training samples and test samples. The sentiment model is first trained based on the training samples, and then the reliability of the trained sentiment model is verified based on the test samples.

[0124] Once the trained sentiment model is obtained based on the above embodiments, the sentiment classification layer and the feature extraction layer in the trained sentiment model can be spliced ​​together, and the spliced ​​model can be determined as the first sentiment classification model.

[0125] Example 5: The model training method provided by the present invention will be described below through specific examples. Figure 2 This is a schematic diagram of a specific model training process provided in an embodiment of the present invention. The process includes:

[0126] S201: Obtain the trained sentiment discrimination model.

[0127] The sentiment discrimination model is based on SER data with second sentiment labels in the third sample set and the original sentiment discrimination model. The specific training process has been described in Example 4 above, and will not be repeated here.

[0128] The original sentiment discrimination model can be obtained based on a Big BERT model containing a large number of parameters. For example, the sentiment discrimination layer can be connected to this Big BERT model, and the Big BERT model with the sentiment discrimination layer connected is determined as the original sentiment discrimination model.

[0129] S202: Obtain the sentiment model based on the sentiment discrimination model obtained in S201.

[0130] The sentiment model can be obtained based on SER data with second sentiment labels from the third sample set and the trained sentiment discrimination model. The specific training process has been described in Example 4 above, and will not be repeated here.

[0131] S203: Based on the sentiment discrimination model obtained in S201, obtain the original sentiment classification model, and based on the original sentiment classification model, obtain the second sentiment classification model.

[0132] In one example, the sentiment classification layer, category recognition layer, and feature extraction layer from the sentiment discrimination model can be concatenated to determine the original sentiment classification model. For instance, if the feature extraction layer in the sentiment discrimination model is determined by a network included in a Big Bert model, the sentiment classification layer and category recognition layer can be directly connected to that Big Bert model to determine the original sentiment classification model. The process of obtaining a second sentiment classification model based on the original sentiment classification model has already been described in Example 3 above, and will not be repeated here.

[0133] S204: Based on the sentiment model obtained in S202, obtain the first sentiment classification model. Based on the first sentiment classification model and the second sentiment classification model obtained in S203, obtain the trained second sentiment classification model.

[0134] In one example, a first encoding sub-network (such as a BERT encoder), a second encoding sub-network (such as an emotion encoder), the networks contained in the emotion model, and a second network layer (such as an emotion classifier) ​​can be sequentially connected to form the first emotion classification model. The first network layer in this first emotion classification model can be determined based on the first encoding sub-network, the second encoding sub-network, and the networks contained in the emotion model. The process of obtaining the trained second emotion classification model based on the first emotion classification model and the second emotion classification model obtained in S203 has been described in Examples 1-2 above, and will not be repeated here.

[0135] S205: Based on the first sentiment classification model obtained in S204, obtain the sentiment extraction model, and based on the sentiment extraction model and the TTS data contained in the first sample set, jointly train the original speech synthesis model and the sentiment extraction model to obtain the target speech synthesis model and the target sentiment extraction model.

[0136] The specific process of jointly training the original speech synthesis model and the emotion extraction model based on the emotion extraction model and the TTS data contained in the first sample set has been described in the above embodiments 1-2, and will not be repeated here.

[0137] Example 6: This embodiment of the invention also provides a speech synthesis method based on the model obtained in the above embodiments. Figure 3 A schematic diagram of a speech synthesis process provided in an embodiment of the present invention includes:

[0138] S301: Obtain the emotion vector of the target emotion based on the target emotion extraction model. This target emotion is the emotion specified by the user when synthesizing speech based on the speech synthesis model.

[0139] The speech synthesis method provided in this invention is applied to electronic devices, which can be intelligent devices such as robots, or servers. The electronic device performing speech synthesis in this invention can be the same as or different from the electronic device used for model training described above.

[0140] In one possible implementation, the model training (including the target emotion extraction model and the target speech synthesis model) is generally performed offline. Once the trained model is obtained, it is stored on the electronic device used for speech synthesis.

[0141] It should be noted that the specific training processes for the target emotion extraction model and the target speech synthesis model have been described in Examples 1-5 above, and the repetitions will not be repeated.

[0142] When acquiring speech data with a target emotion that needs to be synthesized based on text information (i.e., when TTS processing is required), a pre-trained target emotion extraction model can be obtained. This model then yields the emotion vector of the target emotion.

[0143] The methods for obtaining the sentiment vector of the target sentiment through this target sentiment extraction model include at least one of the following:

[0144] Method 1: Based on the sentiment tag corresponding to the target sentiment, determine each reference TTS data with the sentiment tag corresponding to the target sentiment from each TTS data contained in the first sample set. Obtain the reference sentiment vector corresponding to each reference TTS data, and then determine the sentiment vector of the target sentiment based on each reference sentiment vector. The reference sentiment vector corresponding to any reference TTS data is obtained by processing the reference TTS data using a target sentiment extraction model.

[0145] It should be noted that, to save time in obtaining reference sentiment vectors, each TTS data point in the first sample set can be pre-processed using a target sentiment extraction model to obtain and save the corresponding reference sentiment vector for each TTS data point. Subsequently, the reference sentiment vector for each reference TTS data point can be obtained based on the saved reference sentiment vectors for each TTS data point. Alternatively, to save storage space on electronic devices, after determining the reference TTS data points, the target sentiment extraction model can be used to process each reference TTS data point separately to obtain the corresponding reference sentiment vector for each reference TTS data point. The method for obtaining the reference sentiment vector for each reference TTS data point can be flexibly set according to actual needs and is not specifically limited here.

[0146] Method 2: Pre-collect random speech data with different emotions. This random speech data is not part of the TTS data in the first sample set, and each random speech data point corresponds to an emotion tag. From all the random speech data, determine any random speech data point with an emotion tag corresponding to the target emotion. Process this random speech data point using the target emotion extraction model to obtain a reference emotion vector corresponding to this random speech data point, and determine this reference emotion vector as the emotion vector of the target emotion.

[0147] Method 3: Since the sentiment vector is determined based on the sentiment weight vector and various sentiment association parameters, the sentiment weight vector corresponding to the target sentiment can be obtained. Then, at least one weight value contained in the sentiment weight vector is adjusted. Based on the adjusted sentiment weight vector and the various sentiment association parameters contained in the target sentiment extraction model, the sentiment vector of the target sentiment is determined.

[0148] To save time in obtaining sentiment weight vectors, a target sentiment extraction model can be used to pre-process one TTS data point for each sentiment in the first sample set, obtaining and saving the sentiment weight vector corresponding to each sentiment. Subsequently, based on the saved sentiment weight vectors for each sentiment, the sentiment weight vector corresponding to the sentiment in the TTS data can be obtained. Alternatively, to save storage space on electronic devices, after determining the target sentiment, any auxiliary TTS data point containing the target sentiment can be obtained from each TTS data point in the first sample set. This auxiliary TTS data can then be processed using the target sentiment extraction model to obtain the corresponding sentiment weight vector for the target sentiment. The method for obtaining sentiment weight vectors can be flexibly configured according to actual needs and is not specifically limited here.

[0149] It should be noted that the process of obtaining the sentiment weight vector through the target sentiment extraction model is similar to the process of obtaining the sentiment weight vector through the first sentiment classification model, and will not be elaborated here.

[0150] In one possible implementation, at least one weight value contained in the emotion weight vector can be adjusted based on human experience.

[0151] Method 4: Based on human experience or needs, preset emotion weight vectors are pre-set for each emotion. Obtain the preset emotion weight vector corresponding to the target emotion. Based on the preset emotion weight vector and the various emotion-related parameters included in the target emotion extraction model, determine the second emotion vector for the target emotion.

[0152] S302: Using the target speech synthesis model, based on the text features and sentiment vector of the text to be processed, obtain at least one acoustic feature vector corresponding to the text to be processed.

[0153] Since a target speech synthesis model capable of synthesizing speech data with the target emotion has been pre-trained through the above embodiments, after obtaining the emotion vector of the target emotion, the text features of the text to be processed and the emotion vector are input into the target speech synthesis model. Based on the text features and the emotion vector, the target speech synthesis model performs corresponding processing to obtain the acoustic feature vector corresponding to the text features.

[0154] In one example, considering that the emotions inherent in natural speech data have varying intensities—for example, sadness of varying intensities includes heartbreak and grief—an emotion intensity vector is pre-configured to further differentiate the emotions conveyed in the speech data. Therefore, an emotion intensity vector is pre-configured. After obtaining the emotion vector (Emotion Embedding) of the target emotion based on the above embodiment, the Emotion Embedding can be multiplied by the pre-configured emotion intensity vector (Emotion Scalar) to obtain a vector product (Style Embedding). Then, the Emotion Embedding of the target emotion is updated based on the Style Embedding. The updated Emotion Embedding is then used as input data for the target speech synthesis model, thereby controlling the intensity of the emotions in the synthesized speech data, making the obtained synthesized speech data closer to natural language. The Emotion Embedding and the Emotion Scalar have the same dimension. The Emotion Scalar can be flexibly set based on work experience and actual needs.

[0155] In one possible implementation, since the acoustic feature vector predicted by the target speech synthesis model during training is a normalized acoustic feature vector, meaning that the value of each element in the acoustic feature vector is between [0,1], while in practical applications, the acoustic feature vector of normal speech information has a certain bit depth. Therefore, after obtaining the acoustic feature vector corresponding to the text features based on the above embodiment, it is necessary to perform inverse normalization processing on the acoustic feature vector using a preset inverse normalization function, such as the inverse minmax algorithm or the inverse mean normalization algorithm, so that the value of each element in the acoustic feature vector is within a preset value range, that is, within a certain bit depth range, so that the subsequently obtained speech information is more natural. The inverse normalized acoustic feature vector can also be a regularized minmax file.

[0156] S303: Using a vocoder, based on acoustic feature vectors, obtain synthesized speech data with target emotion corresponding to the text to be processed.

[0157] Based on the acquisition of acoustic feature vectors and vocoders, such as the WORLD vocoder and the Linear Prediction LPC vocoder, the audio data of the target speaker uttering text information in the target language is determined. The use of acoustic feature vectors and vocoders is existing technology and will not be elaborated upon here.

[0158] Example 7: This embodiment of the invention also provides a model training device. Figure 4This is a schematic diagram of a model training device provided in an embodiment of the present invention. The device includes:

[0159] The acquisition unit 41 is used for each text-to-speech (TTS) data in the first sample set, which corresponds to a first emotion probability vector and a first acoustic feature, wherein the first emotion probability vector includes the probability value of each emotion pre-configured in the TTS data.

[0160] Training unit 42 is configured to, for any TTS data, determine a first emotion vector corresponding to the emotion possessed by the TTS data through a first network layer included in a first emotion classification model. The first emotion vector is determined based on the emotion weight vector corresponding to the emotion possessed by the TTS data and pre-configured emotion association parameters. The emotion weight vector includes the weight values ​​corresponding to each emotion association parameter, and each emotion association parameter is a non-speech-related emotion auxiliary vector used in the first emotion classification model to determine the emotion possessed by the TTS data. Furthermore, it is configured to, through a second network layer included in the first emotion classification model, determine a second emotion probability vector corresponding to the TTS data based on the first emotion vector. The second emotion probability vector includes the probability values ​​of each pre-configured emotion possessed by the TTS data predicted by the first emotion classification model. Based on the second emotion probability vector and the corresponding first emotion probability vector, the first emotion classification model is trained to obtain a trained first emotion classification model, and an emotion extraction model is determined according to the first network layer included in the trained first emotion classification model.

[0161] In one possible implementation, the training unit 42 is specifically configured to, for any TTS data, obtain the output data of the last sub-network in the feature extraction layer included in the first network layer based on the first acoustic features corresponding to the TTS data; wherein, the output data is the emotion feature extracted based on the first acoustic features corresponding to the TTS data; obtain the basic emotion vector corresponding to the TTS data based on the output data through the first encoding sub-network included in the first network layer; the basic emotion vector includes a speech-related emotion auxiliary vector used to determine the emotion possessed by the TTS data; obtain the emotion weight vector corresponding to the emotion possessed by the TTS data based on the basic emotion vector through the second encoding sub-network included in the first network layer, and determine the first emotion vector corresponding to the emotion possessed by the TTS data according to the emotion weight vector and various emotion association parameters pre-configured in the second encoding sub-network.

[0162] In one possible implementation, the acquisition unit 41 is specifically used to acquire the first sentiment probability vector corresponding to each TTS data point in the following manner:

[0163] For each TTS data, a first emotion probability vector is determined based on the first acoustic feature corresponding to the TTS data using a pre-trained second emotion classification model.

[0164] In one possible implementation, the acquisition unit 41 is further configured to, for each TTS data, acquire the output data of each first target network included in the second emotion classification model based on the first acoustic features corresponding to the TTS data using a pre-trained second emotion classification model; and acquire the output data of each sub-network included in the feature extraction layer based on the first acoustic features corresponding to the TTS data using the first emotion classification model; the first target network is a part of the network in the second emotion classification model, and each sub-network in the first emotion classification model corresponds to a first target network in the second emotion classification model;

[0165] Training unit 42 is specifically used to train the first emotion classification model based on the second emotion probability vector and its corresponding first emotion probability vector, the output data of each sub-network and the output data of the first target network corresponding to each sub-network, so as to obtain the trained first emotion classification model.

[0166] In one possible implementation, the acquisition unit 41 is further configured to, for each TTS data, acquire a second emotion vector of the emotion possessed by the TTS data based on an emotion extraction model, wherein the second emotion vector contains the emotion features of the emotion possessed by the TTS data.

[0167] Training unit 42 is also used to jointly train the original speech synthesis model and the emotion extraction model based on the text feature samples corresponding to the TTS data, the first acoustic features of the TTS data, and the second emotion vector, so as to obtain the target speech synthesis model and the target emotion extraction model.

[0168] In one possible implementation, the acquisition unit 41 is specifically configured to acquire a second sentiment vector of the sentiment possessed by the TTS data based on a sentiment extraction model in at least one of the following ways:

[0169] The emotion extraction model processes any random speech data with an emotion tag corresponding to the TTS data to obtain a reference emotion vector corresponding to the random speech data; and the reference emotion vector is determined as the second emotion vector. Here, the random speech data is not TTS data in the first sample set, and the emotion tag is used to identify the emotion of the TTS data.

[0170] Based on the reference sentiment vectors corresponding to each TTS data with the sentiment tag corresponding to the TTS data in the first sample set, the second sentiment vector of the sentiment is determined. The reference sentiment vector is obtained by processing the TTS data with the sentiment through the sentiment extraction model.

[0171] Based on the sentiment tag corresponding to the TTS data, determine the sentiment of the TTS data; obtain the sentiment weight vector corresponding to the sentiment; wherein, the sentiment weight vector is obtained by processing any speech data with sentiment through the sentiment extraction model; adjust at least one weight value contained in the sentiment weight vector; based on the adjusted sentiment weight vector and the various sentiment association parameters contained in the sentiment extraction model, determine the second sentiment vector of the sentiment, wherein, the speech data includes random speech data and any one of the TTS data in the first sample set;

[0172] Based on the sentiment tags corresponding to the TTS data, the sentiment of the TTS data is determined; based on the preset sentiment weight vector corresponding to the sentiment and the various sentiment association parameters contained in the sentiment extraction model, the second sentiment vector of the sentiment is determined.

[0173] In one possible implementation, the acquisition unit 41 is further configured to acquire any speech sample in the second sample set, wherein the speech sample corresponds to a first emotion tag and a speech type tag, the first emotion tag is used to identify the emotion of the speech sample, and the speech type tag is used to identify the speech type to which the speech sample belongs, the speech type including at least one of TTS type and speech emotion recognition SER type.

[0174] Training unit 42 is also used to determine the third emotion probability vector and type probability vector corresponding to the speech sample based on the second acoustic features corresponding to the speech sample using the original emotion classification model. The third emotion probability vector includes the probability value of each pre-configured emotion that the speech sample has, as determined by the original emotion classification model. The type probability vector includes the probability value of each pre-configured speech type that the speech sample belongs to, as determined by the original emotion classification model. Based on the third emotion probability vector and the first emotion label, and the type probability vector and the speech type label, the original emotion classification model is trained to obtain the second emotion classification model.

[0175] In one possible implementation, the training unit 42 is specifically used to determine the emotion loss value based on the third emotion probability vector and the first emotion label using the cross-entropy loss function; and to determine the type loss value based on the type probability vector and the speech type label using the gradient inversion to maximize the classification error function; and to train the original emotion classification model based on the emotion loss value and the type loss value.

[0176] In one possible implementation, the acquisition unit 41 is further configured to acquire any two SER data in the third sample set; each SER data corresponds to a second sentiment tag.

[0177] Training unit 42 is also used to determine, based on the third acoustic features corresponding to the two SER data respectively, whether the emotions of the two SER data are consistent, using the original emotion discrimination model; and to determine, based on the second emotion labels corresponding to the two SER data respectively, whether the emotions of the two SER data are consistent, using the second emotion labels; and to train the original emotion discrimination model based on the first and second results to obtain the original emotion classification model.

[0178] In one possible implementation, the acquisition unit 41 is further configured to acquire any SER data in the third sample set;

[0179] Training unit 42 is also used to determine the output data of each second target network included in the feature extraction layer of the emotion discrimination model based on the third acoustic features corresponding to the SER data, using a pre-trained emotion discrimination model. Each second target network is a partial network included in the feature extraction layer. It also determines the output data of each third target network included in the feature extraction layer of the original emotion model based on the third acoustic features. Each third target network corresponds to a second target network in the emotion discrimination model. Based on the output data of each third target network and the output data of the second target networks corresponding to each third target network, the original emotion model is trained, and a first emotion classification model is obtained based on the trained emotion model.

[0180] Example 8: This embodiment of the invention also provides a speech synthesis device. Figure 5 This is a schematic diagram of a speech synthesis device provided in an embodiment of the present invention. The device includes:

[0181] The acquisition module 51 is used to acquire the sentiment vector of the target sentiment based on the target sentiment extraction model corresponding to the target sentiment;

[0182] The first processing module 52 is used to obtain the acoustic feature vector corresponding to the text to be processed based on the text features and emotion vector of the text to be processed through the target speech synthesis model.

[0183] The second processing module 53 is used to obtain synthesized speech data with target emotion corresponding to the text to be processed based on acoustic feature vectors through a vocoder.

[0184] In one possible implementation, the acquisition module 51 is specifically used to acquire the sentiment vector of the target sentiment based on the target sentiment extraction model corresponding to the target sentiment through at least one of the following methods:

[0185] The target emotion extraction model processes any random speech data with an emotion tag corresponding to the target emotion to obtain a reference emotion vector. This reference emotion vector is then determined as the emotion vector of the target emotion. The random speech data is not TTS data with the target emotion from the first sample set used to train the target emotion extraction model. The emotion vector of the target emotion is determined based on the reference emotion vectors corresponding to each TTS data with the emotion tag corresponding to the target emotion from the first sample set. The reference emotion vector is obtained by processing the TTS data with the target emotion using the target emotion extraction model. Finally, the emotion weight vector corresponding to the target emotion is obtained. The target emotion extraction model is obtained by processing any speech data containing the target emotion; adjusting at least one weight value contained in the emotion weight vector; and determining the emotion vector of the target emotion based on the adjusted emotion weight vector and the various emotion association parameters contained in the target emotion extraction model. The emotion weight vector contains the weight values ​​corresponding to each emotion association parameter; each emotion association parameter is a non-speech-related emotion auxiliary vector used in the target emotion extraction model to determine the target emotion; the speech data includes any one of random speech data and TTS data from the first sample set; and the emotion vector of the target emotion is determined based on the preset emotion weight vector corresponding to the target emotion and the various emotion association parameters contained in the target emotion extraction model.

[0186] In one possible implementation, the first processing module 52 is further configured to, after obtaining the emotion vector of the target emotion based on the target emotion extraction model corresponding to the target emotion, and before obtaining the acoustic feature vector corresponding to the text to be processed based on the text features of the text to be processed and the emotion vector through the target speech synthesis model, update the emotion vector according to the vector product of the emotion vector and the pre-configured emotion intensity vector, and use the updated emotion vector as the input data of the target speech synthesis model; wherein the emotion vector and the emotion intensity vector have the same dimension.

[0187] Example 9: Based on the above examples, this embodiment of the invention also provides an electronic device. Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 6 As shown, it includes: processor 61, communication interface 62, memory 63 and communication bus 64, wherein processor 61, communication interface 62 and memory 63 communicate with each other through communication bus 64;

[0188] The memory 63 stores a computer program, which, when executed by the processor 61, causes the processor 61 to perform the following steps:

[0189] For each text-to-speech (TTS) data point in the first sample set, each TTS data point corresponds to a first emotion probability vector and a first acoustic feature. The first emotion probability vector includes the probability values ​​of each pre-configured emotion for each TTS data point. For any given TTS data point, the first emotion vector corresponding to the emotion possessed by the TTS data point is determined through the first network layer included in the first emotion classification model. The first emotion vector is determined based on the emotion weight vector corresponding to the emotion possessed by the TTS data point and pre-configured emotion association parameters. The emotion weight vector includes the weight values ​​corresponding to each emotion association parameter, and each emotion association parameter is a first emotion probability vector. The model includes a non-speech-related emotion auxiliary vector used to determine the emotions possessed by the TTS data; and a second emotion probability vector corresponding to the TTS data determined based on the first emotion vector using a second network layer included in the first emotion classification model, wherein the second emotion probability vector includes the probability values ​​of each pre-configured emotion in the TTS data predicted by the first emotion classification model; and the first emotion classification model is trained based on the second emotion probability vector and the corresponding first emotion probability vector to obtain a trained first emotion classification model, and an emotion extraction model is determined based on the first network layer included in the trained first emotion classification model.

[0190] Since the principle of the above-mentioned electronic device in solving the problem is similar to that of the model training method, the implementation of the above-mentioned electronic device can be found in Examples 1-5 of the method, and the repeated parts will not be described again.

[0191] Example 10: Based on the above examples, this embodiment of the invention also provides an electronic device. Figure 7 This is a schematic diagram of the structure of another electronic device provided in an embodiment of the present invention, such as... Figure 7 As shown, it includes: processor 71, communication interface 72, memory 73 and communication bus 74, wherein processor 71, communication interface 72 and memory 73 communicate with each other through communication bus 74;

[0192] The memory 73 stores a computer program. When the program is executed by the processor 71, the processor 71 performs the following steps: obtaining the emotion vector of the target emotion based on the target emotion extraction model; obtaining the acoustic feature vector of the text to be processed based on the text features and emotion vector of the text to be processed through the target speech synthesis model; and obtaining the synthesized speech data with the target emotion corresponding to the text to be processed based on the acoustic feature vector through a vocoder.

[0193] Since the principle of the above-mentioned electronic device in solving the problem is similar to that of the speech synthesis method, the implementation of the above-mentioned electronic device can be found in Embodiment 6 of the method, and the repeated parts will not be described again.

[0194] The communication bus mentioned in the above-mentioned electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not indicate that there is only one bus or one type of bus. Communication interface 72 is used for communication between the above-mentioned electronic device and other devices. The memory can include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.

[0195] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0196] Example 11: Based on the above embodiments, this embodiment of the invention also provides a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to perform the following steps:

[0197] For each text-to-speech (TTS) data in the first sample set, the TTS data corresponds to a first emotion probability vector and a first acoustic feature, wherein the first emotion probability vector includes the probability value of each emotion pre-configured for each TTS data.

[0198] For any TTS data, a first emotion vector corresponding to the emotion possessed by the TTS data is determined through a first network layer included in a first emotion classification model. The first emotion vector is determined based on the emotion weight vector corresponding to the emotion possessed by the TTS data and pre-configured emotion association parameters. The emotion weight vector includes the weight values ​​corresponding to each emotion association parameter, and each emotion association parameter is a non-speech-related emotion auxiliary vector used in the first emotion classification model to determine the emotion possessed by the TTS data. A second emotion probability vector corresponding to the TTS data is determined based on the first emotion vector through a second network layer included in the first emotion classification model. The second emotion probability vector includes the probability values ​​of each pre-configured emotion possessed by the TTS data, predicted by the first emotion classification model.

[0199] Based on the second emotion probability vector and the corresponding first emotion probability vector, the first emotion classification model is trained to obtain the trained first emotion classification model, and the emotion extraction model is determined according to the first network layer contained in the trained first emotion classification model.

[0200] Since the principle of solving the problem using the aforementioned computer-readable storage medium is similar to that of the model training method, the implementation of the aforementioned computer-readable storage medium can be found in Implementation 1-5 of the method, and the repeated parts will not be described again.

[0201] Example 12: Based on the above embodiments, this embodiment of the invention also provides a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to perform the following steps:

[0202] Based on the target emotion extraction model corresponding to the target emotion, the emotion vector of the target emotion is obtained; through the target speech synthesis model, based on the text features and emotion vector of the text to be processed, the acoustic feature vector of the text to be processed is obtained; and through a vocoder, based on the acoustic feature vector, the synthesized speech data with the target emotion corresponding to the text to be processed is obtained.

[0203] Since the principle of the computer-readable storage medium in solving the problem is similar to that of the speech synthesis method, the implementation of the computer-readable storage medium can be found in Implementation 6 of the method, and the repeated parts will not be described again.

[0204] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0205] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0206] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0207] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0208] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A model training method, characterized in that, The method includes: For each text-to-speech (TTS) data in the first sample set, the TTS data corresponds to a first emotion probability vector and a first acoustic feature, wherein the first emotion probability vector includes the probability values ​​of each emotion pre-configured for each TTS data. For any TTS data, a first emotion vector corresponding to the emotion possessed by the TTS data is determined through a first network layer included in the first emotion classification model. This first emotion vector is determined based on the emotion weight vector corresponding to the emotion possessed by the TTS data and pre-configured emotion association parameters. The emotion weight vector includes the weight values ​​corresponding to each emotion association parameter, and each emotion association parameter is a non-speech-related emotion auxiliary vector used in the first emotion classification model to determine the emotion possessed by the TTS data. Furthermore, a second emotion probability vector corresponding to the TTS data is determined based on the first emotion vector through a second network layer included in the first emotion classification model. This second emotion probability vector includes the probability values ​​of each pre-configured emotion possessed by the TTS data, predicted by the first emotion classification model. Based on the second emotion probability vector and the corresponding first emotion probability vector, the first emotion classification model is trained to obtain the trained first emotion classification model, and the emotion extraction model is determined according to the first network layer contained in the trained first emotion classification model. Wherein, after determining the first sentiment vector corresponding to the sentiment of any TTS data through the first network layer included in the first sentiment classification model, the method further includes: For each TTS data, a pre-trained second sentiment classification model is used to obtain the output data of each first target network contained in the second sentiment classification model based on the first acoustic features corresponding to the TTS data; and the first sentiment classification model is used to obtain the output data of each sub-network contained in the feature extraction layer contained in the first network layer based on the first acoustic features corresponding to the TTS data; wherein, the first target network is a part of the network in the second sentiment classification model, and each sub-network in the first sentiment classification model corresponds to one of the first target networks in the second sentiment classification model; The step of training the first emotion classification model based on the second emotion probability vector and the corresponding first emotion probability vector includes: training the first emotion classification model based on the second emotion probability vector and its corresponding first emotion probability vector, the output data of each sub-network and the output data of the first target network corresponding to each sub-network, so as to obtain the trained first emotion classification model.

2. The method according to claim 1, characterized in that, For any TTS data, the first sentiment vector corresponding to the sentiment of the TTS data is determined through the first network layer contained in the first sentiment classification model, including: Through the feature extraction layer included in the first network layer, based on the first acoustic feature corresponding to the TTS data, the output data of the last sub-network in the feature extraction layer is obtained; wherein, the output data is the emotional feature extracted based on the first acoustic feature corresponding to the TTS data; The basic sentiment vector corresponding to the TTS data is obtained based on the output data through the first coding subnetwork included in the first network layer; wherein, the basic sentiment vector includes a speech-related sentiment auxiliary vector used to determine the sentiment of the TTS data. By using the second encoding subnetwork included in the first network layer, based on the basic sentiment vector, the sentiment weight vector corresponding to the sentiment of the TTS data is obtained, and according to the sentiment weight vector and the various sentiment association parameters pre-configured in the second encoding subnetwork, the first sentiment vector corresponding to the sentiment of the TTS data is determined.

3. The method according to claim 1, characterized in that, The first sentiment probability vector corresponding to each TTS data is obtained in the following way: For each TTS data, a first emotion probability vector is determined based on the first acoustic feature corresponding to the TTS data using a pre-trained second emotion classification model.

4. The method according to claim 1, characterized in that, The method further includes: For each TTS data, based on the emotion extraction model, a second emotion vector of the emotion possessed by the TTS data is obtained, wherein the second emotion vector contains the emotion features of the emotion possessed by the TTS data; Based on the text feature samples corresponding to the TTS data, the first acoustic feature of the TTS data, and the second emotion vector, the original speech synthesis model and the emotion extraction model are jointly trained to obtain the target speech synthesis model and the target emotion extraction model.

5. The method according to claim 4, characterized in that, Based on the sentiment extraction model, obtain the second sentiment vector of the TTS data using at least one of the following methods: The emotion extraction model is used to process any random speech data with an emotion tag corresponding to the TTS data to obtain a reference emotion vector corresponding to the random speech data; and the reference emotion vector is determined as the second emotion vector, wherein the random speech data is not TTS data in the first sample set, and the emotion tag is used to identify the emotion of the TTS data; Based on the reference sentiment vectors corresponding to each TTS data with the sentiment tag corresponding to the TTS data in the first sample set, a second sentiment vector of the sentiment is determined, wherein the reference sentiment vector is obtained by processing the TTS data with the sentiment through the sentiment extraction model; Based on the sentiment tag corresponding to the TTS data, the sentiment of the TTS data is determined; the sentiment weight vector corresponding to the sentiment is obtained; wherein, the sentiment weight vector is obtained by processing any speech data with the sentiment through the sentiment extraction model; at least one weight value contained in the sentiment weight vector is adjusted; based on the adjusted sentiment weight vector and the various sentiment association parameters contained in the sentiment extraction model, a second sentiment vector of the sentiment is determined, wherein, the speech data includes any one of the random speech data and the TTS data in the first sample set; Based on the sentiment tag corresponding to the TTS data, the sentiment of the TTS data is determined; based on the preset sentiment weight vector corresponding to the sentiment and the various sentiment association parameters included in the sentiment extraction model, the second sentiment vector of the sentiment is determined.

6. The method according to claim 1, characterized in that, The second sentiment classification model was obtained in the following way: Obtain any speech sample from the second sample set, wherein the speech sample corresponds to a first emotion label and a speech type label, the first emotion label is used to identify the emotion of the speech sample, and the speech type label is used to identify the speech type to which the speech sample belongs, the speech type including at least one of TTS type and speech emotion recognition SER type; Using the original emotion classification model, based on the second acoustic features corresponding to the speech sample, the third emotion probability vector and the type probability vector corresponding to the speech sample are determined. The third emotion probability vector includes the probability value of each pre-configured emotion that the speech sample has, as determined by the original emotion classification model. The type probability vector includes the probability value of each pre-configured speech type that the speech sample belongs to, as determined by the original emotion classification model. The original emotion classification model is trained based on the third emotion probability vector and the first emotion label, the type probability vector and the speech type label to obtain the second emotion classification model.

7. The method according to claim 6, characterized in that, The second sentiment classification model includes a feature extraction layer and a sentiment classification layer contained in the trained sentiment classification model, wherein the feature extraction layer includes each of the first target networks.

8. The method according to claim 6, characterized in that, The step of training the original emotion classification model based on the third emotion probability vector and the first emotion label, and the type probability vector and the speech type label includes: Based on the third emotion probability vector and the first emotion label, the emotion loss value is determined using the cross-entropy loss function; and based on the type probability vector and the speech type label, the type loss value is determined using the gradient inversion maximization classification error function. The original sentiment classification model is trained based on the sentiment loss value and the type loss value.

9. The method according to claim 6, characterized in that, The original sentiment classification model was obtained in the following way: Obtain any two SER data from the third sample set, wherein each of the SER data corresponds to a second sentiment tag; Using the original sentiment discrimination model, based on the third acoustic features corresponding to the two SER data sets respectively, a first result is determined as to whether the sentiments of the two SER data sets are consistent; and Based on the second sentiment tags corresponding to the two SER data respectively, a second result is determined as to whether the sentiments of the two SER data are consistent. Based on the first result and the second result, the original sentiment discrimination model is trained to obtain the original sentiment classification model.

10. The method according to claim 9, characterized in that, The original sentiment classification model includes a sentiment classification layer, a category recognition layer, and a feature extraction layer contained in the trained sentiment discrimination model.

11. The method according to claim 10, characterized in that, The first sentiment classification model was obtained in the following way: Obtain any SER data from the third sample set; Using a pre-trained sentiment discrimination model, based on the third acoustic features corresponding to the SER data, the output data of each second target network included in the feature extraction layer of the sentiment discrimination model is determined, wherein each second target network is a partial network included in the feature extraction layer; and using the original sentiment model, based on the third acoustic features, the output data of each third target network included in the feature extraction layer of the original sentiment model is determined, wherein each third target network corresponds to a second target network in the sentiment discrimination model. Based on the output data of each third target network and the output data of the second target network corresponding to each third target network, the original sentiment model is trained, and the first sentiment classification model is obtained based on the trained sentiment model.

12. A speech synthesis method based on a model obtained by the model training method as described in any one of claims 1-11, characterized in that, The method includes: Based on the target emotion extraction model corresponding to the target emotion, the emotion vector of the target emotion is obtained; Using a target speech synthesis model, the acoustic feature vector corresponding to the text to be processed is obtained based on the text features of the text to be processed and the emotion vector. Using a vocoder, based on the acoustic feature vector, synthesized speech data with the target emotion corresponding to the text to be processed is obtained.

13. The method according to claim 12, characterized in that, After obtaining the emotion vector of the target emotion based on the target emotion extraction model, and before obtaining the acoustic feature vector corresponding to the text to be processed based on the text features of the text to be processed and the emotion vector through the target speech synthesis model, the method further includes: The emotion vector is updated based on the vector product of the emotion vector and the pre-configured emotion intensity vector, and the updated emotion vector is used as the input data of the target speech synthesis model; wherein the emotion vector and the emotion intensity vector have the same dimension.

14. A model training device, characterized in that, The device includes: The acquisition unit is used for each text-to-speech (TTS) data in the first sample set, which corresponds to a first emotion probability vector and a first acoustic feature, wherein the first emotion probability vector includes the probability value of each emotion pre-configured in the TTS data. The training unit is configured to, for any TTS data, determine a first emotion vector corresponding to the emotion possessed by the TTS data through a first network layer included in a first emotion classification model. The first emotion vector is determined based on an emotion weight vector corresponding to the emotion possessed by the TTS data and pre-configured emotion association parameters. The emotion weight vector includes weight values ​​corresponding to each emotion association parameter, and each emotion association parameter is a non-speech-related emotion auxiliary vector used in the first emotion classification model to determine the emotion possessed by the TTS data. The unit also determines a second emotion probability vector corresponding to the TTS data based on the first emotion vector through a second network layer included in the first emotion classification model. The second emotion probability vector includes the probability values ​​of each pre-configured emotion possessed by the TTS data predicted by the first emotion classification model. Based on the second emotion probability vector and the corresponding first emotion probability vector, the first emotion classification model is trained to obtain a trained first emotion classification model. Finally, an emotion extraction model is determined based on the first network layer included in the trained first emotion classification model. The training unit is further configured to, for each TTS data, obtain the output data of each first target network contained in the second emotion classification model based on the first acoustic feature corresponding to the TTS data using a pre-trained second emotion classification model; and obtain the output data of each sub-network contained in the feature extraction layer contained in the first network layer based on the first acoustic feature corresponding to the TTS data using the first emotion classification model; wherein, the first target network is a part of the network in the second emotion classification model, and each sub-network in the first emotion classification model corresponds to one of the first target networks in the second emotion classification model; The training unit is specifically used to train the first emotion classification model based on the second emotion probability vector and its corresponding first emotion probability vector, the output data of each sub-network and the output data of the first target network corresponding to each sub-network, so as to obtain the trained first emotion classification model.

15. A speech synthesis device based on a model obtained from the model training apparatus as described in claim 14, characterized in that, The device includes: The acquisition module is used to acquire the sentiment vector of the target sentiment based on the target sentiment extraction model corresponding to the target sentiment; The first processing module is used to obtain the acoustic feature vector corresponding to the text to be processed based on the text features of the text to be processed and the emotion vector through the target speech synthesis model. The second processing module is used to obtain, through a vocoder, synthesized speech data corresponding to the text to be processed and carrying the target emotion, based on the acoustic feature vector.

16. An electronic device, characterized in that, The electronic device includes at least a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the steps of the model training method as described in any one of claims 1-11, or to implement the steps of the speech synthesis method as described in any one of claims 12-13.

17. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the model training method as described in any one of claims 1-11, or the steps of the speech synthesis method as described in any one of claims 12-13.

Citation Information

Patent Citations

  • Sensitivity adjustability-based speech emotion recognition method and system

    CN108564942A

  • Emotional audio generation method and device

    CN112837700A