A training method, device, equipment and medium for a speech synthesis model
By introducing the GST model to the TTS model to recognize the text emotionally and embed the emotion vector, the problem of the existing TTS model's demand for high-quality speech data sets is solved, and the effect of converting text into emotional speech without emotional annotation is achieved.
Patent Information
- Application Number
- CN202111138448.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-09-27
AI Technical Summary
The demand for high-quality and high-number speech data sets during training makes it difficult to manually annotate the data sets, and it is difficult to convert text into emotional speech without emotional annotation.
By introducing a global style label (GST) model into the speech synthesis model, the training text information is emotionally recognized, and the emotion vector is obtained, and it is embedded in the TTS model to perform speech synthesis processing to realize emotional speech synthesis.
Without emotional annotation of training samples, the TTS model can be used to convert text into emotional speech, which improves the training efficiency of the speech synthesis model.
Smart Images

Figure CN113889072B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training method, device, equipment and medium for a speech synthesis model. Background Art
[0002] With the development of deep learning technology, deep learning can be applied to the task of converting text to speech, which mainly includes Text To Speech (TTS) technology. Although the TTS model can convert text into a more natural speech signal, in terms of converting text into emotional speech, due to the need for high-quality and large-volume speech data sets during the training process of the TTS model, it becomes difficult to manually perform emotion annotation on the data sets used by the TTS model. Therefore, how to use the TTS model to convert text into emotional speech without emotionally annotating the data sets is a problem to be solved. Summary of the invention
[0003] The embodiments of the present application provide a method, apparatus, device and medium for training a speech synthesis model, which can realize the use of a TTS model to convert text into emotional speech without emotionally annotating training samples, thereby improving the training efficiency of the speech synthesis model.
[0004] On the one hand, an embodiment of the present application provides a method for training a speech synthesis model, the method comprising:
[0005] Acquire a training sample, where the training sample includes first training text information and training voice information corresponding to the first training text information;
[0006] Performing emotion recognition processing on the first training text information through a global style token (GST) model in the speech synthesis model to obtain an emotion vector of the first training text information, and embedding the emotion vector of the first training text information into a TTS model in the speech synthesis model;
[0007] Performing speech synthesis processing on the first training text information and the emotion vector of the first training text information through the TTS model to obtain predicted speech information corresponding to the first training text information;
[0008] Compare the predicted speech information corresponding to the first training text information with the training speech information to obtain a speech synthesis loss value;
[0009] Based on the speech synthesis loss value, the parameters of the TTS model and the parameters of the GST model are adjusted to train the speech synthesis model to obtain a trained speech synthesis model, which includes a trained GST model and a trained TTS model.
[0010] In one embodiment, before performing emotion recognition processing on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information, the following process may be further implemented:
[0011] Performing emotion classification processing on the first training text information through a cross-domain speech emotion recognition (SER) model to obtain a first emotion label of the first training text information, where the first emotion label is used to indicate an emotion category of the first training text information;
[0012] Performing sentiment classification processing on the first training text information through the GST model to obtain a second sentiment label of the first training text information;
[0013] Compare the first emotion label with the second emotion label to obtain an emotion loss value;
[0014] Adjust the parameters of the GST model based on the sentiment loss value to obtain an adjusted GST model;
[0015] The specific implementation process of performing sentiment classification processing on the first training text information through the GST model in the speech synthesis model to obtain the sentiment vector of the first training text information is as follows:
[0016] The first training text information is subjected to sentiment recognition processing through the adjusted GST model to obtain a sentiment vector of the first training text information.
[0017] In one embodiment, the specific implementation process of performing sentiment classification processing on the first training text information through the cross-domain SER model to obtain the first sentiment label of the first training text information is:
[0018] Selecting second training text information whose similarity to the first training text information is greater than a preset ratio threshold based on a Maximum Mean Discrepancy (MMD) algorithm;
[0019] The second training text information is subjected to sentiment classification processing through the cross-domain SER model to obtain a first sentiment label.
[0020] In one embodiment, the specific implementation process of performing emotion recognition processing on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information is:
[0021] Encode the first training text information through a reference encoder in the GST model to obtain a reference vector;
[0022] Use the attention mechanism to calculate the similarity between the reference vector and the initialization vector to obtain a set of weight values;
[0023] A set of weight values is weightedly added to the initialization vector to obtain the sentiment vector of the first training text information.
[0024] In one embodiment, the specific implementation process of performing speech synthesis processing on the first training text information and the emotion vector of the first training text information through the TTS model in the speech synthesis model to obtain the predicted speech information corresponding to the first training text information is:
[0025] Performing language learning on the first training text information through the TTS model to obtain underlying structural features of the first training text information;
[0026] The underlying structural features of the first training text information and the sentiment vector of the first training text information are decoded through the TTS model to obtain predicted speech information corresponding to the first training text information.
[0027] In one embodiment, the parameters of the TTS model and the GST model in the speech synthesis model are adjusted based on the speech synthesis loss value to train the speech synthesis model. After the trained speech synthesis model is obtained, the following process may be further implemented:
[0028] Get target text information;
[0029] The trained GST model is used to predict the sentiment of the target text information and obtain the sentiment vector of the target text information;
[0030] The trained TTS model is used to perform speech synthesis processing on the target text information and the emotion vector of the target text information to obtain the predicted speech information corresponding to the target text information.
[0031] In one embodiment, the specific implementation process of obtaining the target text information is:
[0032] Display target text information;
[0033] Upon receiving a speech synthesis instruction for target text information, obtaining the target text information;
[0034] After performing speech synthesis processing on the target text information and the emotion vector of the target text information through the trained TTS model to obtain predicted speech information corresponding to the target text information, the method further includes:
[0035] Output the predicted speech information corresponding to the target text information.
[0036] On the other hand, an embodiment of the present application provides a training device for a speech synthesis model, the training device for the speech synthesis model comprising:
[0037] An acquisition unit, configured to acquire a training sample, wherein the training sample includes first training text information and training voice information corresponding to the first training text information;
[0038] A processing unit, configured to perform emotion recognition processing on the first training text information through a GST model in a speech synthesis model, obtain an emotion vector of the first training text information, and embed the emotion vector of the first training text information into a TTS model;
[0039] The processing unit is further used to perform speech synthesis processing on the first training text information and the emotion vector of the first training text information through a TTS model in the speech synthesis model to obtain predicted speech information corresponding to the first training text information;
[0040] The processing unit is further used to compare the predicted speech information with the training speech information to obtain a speech synthesis loss value;
[0041] The processing unit is also used to adjust the parameters of the TTS model and the GST model in the speech synthesis model based on the speech synthesis loss value to train the speech synthesis model to obtain a trained speech synthesis model, wherein the trained speech synthesis model includes a trained GST model and a trained TTS model.
[0042] On the other hand, an embodiment of the present application provides an electronic device, including a processor, a memory and a communication interface, wherein the processor, the memory and the communication interface are interconnected, wherein the memory is used to store a computer program that supports a terminal to execute the above method, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the following steps: obtaining a training sample, the training sample including first training text information and training voice information corresponding to the first training text information; performing emotion recognition processing on the first training text information through a GST model in a speech synthesis model to obtain an emotion vector of the first training text information, and embedding the emotion vector of the first training text information into a TTS model in the speech synthesis model; performing speech synthesis processing on the first training text information and the emotion vector of the first training text information through the TTS model to obtain predicted voice information corresponding to the first training text information; comparing the predicted voice information corresponding to the first training text information with the training voice information to obtain a speech synthesis loss value; adjusting the parameters of the TTS model and the parameters of the GST model based on the speech synthesis loss value to train the speech synthesis model to obtain a trained speech synthesis model, wherein the trained speech synthesis model includes a trained GST model and a trained TTS model.
[0043] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the above-mentioned training method of the speech synthesis model.
[0044] In an embodiment of the present application, emotion recognition processing is performed on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information, and the emotion vector of the first training text information is embedded in the TTS model in the speech synthesis model, and then speech synthesis is performed based on the first training text information and the emotion vector through the TTS model. It can be achieved that without emotion annotation of the training samples, the TTS model can still be used to convert text into emotional speech, thereby improving the training efficiency of the speech synthesis model. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0046] Figure 1 It is a flowchart of a method for training a speech synthesis model provided in an embodiment of the present application;
[0047] Figure 2 It is a schematic diagram of the architecture of a training system for a speech synthesis model provided in an embodiment of the present application;
[0048] Figure 3 It is a flowchart of a speech synthesis method provided in an embodiment of the present application;
[0049] Figure 4 It is a structural schematic diagram of a training device for a speech synthesis model provided in an embodiment of the present application;
[0050] Figure 5 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] The embodiment of the present application provides a method for training a speech synthesis model, wherein the GST model in the speech synthesis model performs emotion recognition processing on the first training text information to obtain the emotion vector of the first training text information, and the emotion vector of the first training text information is embedded in the TTS model in the speech synthesis model, and then speech synthesis is performed based on the first training text information and the emotion vector through the TTS model. Training samples without manual emotion annotations can still be used for emotional speech synthesis, saving a lot of human resources and time resources, and the required emotional TTS can be established more quickly. Based on the trained speech synthesis model, emotion-rich speech can be generated without affecting the quality of the predicted speech information obtained by speech synthesis.
[0052] The training method of the speech synthesis model in the embodiment of the present application can be applied to a first electronic device, wherein the first electronic device can be any one or more of a smart phone, a tablet computer, a laptop computer, a desktop computer, an intelligent vehicle-mounted device, and an intelligent wearable device. Optionally, the first electronic device can also be a server, which can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. In other words, the server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0053] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0054] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0055] See also Figure 1 , Figure 1 is a flow chart of a method for training a speech synthesis model provided in an embodiment of the present application; Figure 1The training method of the speech synthesis model shown can be performed by a first electronic device, and the scheme includes but is not limited to steps S101 to S105, wherein:
[0056] S101, obtaining a training sample, where the training sample includes first training text information and training voice information corresponding to the first training text information.
[0057] The first electronic device can obtain training samples. For example, if a user inputs audio data about "Happy Birthday to you", the first electronic device can use the audio data as training voice information. The first training text information corresponding to the training voice information can be "Happy Birthday to you".
[0058] It is understandable that the training sample can be input by the user to the first electronic device, for example, the first electronic device collects training voice information through a microphone, and collects training text information corresponding to the training voice information through the input device of the first electronic device (such as a touch panel or keyboard, etc.). Optionally, the training sample can also be obtained by the first electronic device from a local memory, or obtained by the first electronic device from other devices, or downloaded by the first electronic device through the Internet, which is not limited by the embodiments of the present application.
[0059] S102, performing emotion recognition processing on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information, and embedding the emotion vector of the first training text information into the TTS model in the speech synthesis model.
[0060] The GST model in the embodiment of the present application adds an auxiliary supervised emotion prediction task based on style token weights relative to the traditional GST model, so that the GTS model can better model the style features related to emotions. The first electronic device inputs the acquired first training text information into the GST model, and the GST model can add corresponding emotion conditions to the character strings of different emotions in the first training text information to obtain the emotion vector of the first training text information, and the emotion vector includes the emotion unit vectors corresponding to the character strings of different emotions in the first training text information.
[0061] As a feasible implementation method, during the training process, the embodiment of the present application can provide emotion labels for the TTS model through a cross-domain SER model. Specifically, the first electronic device can perform emotion classification processing on the first training text information through the cross-domain SER model to obtain a first emotion label of the first training text information, and the first emotion label is used to indicate the emotion category of the first training text information; perform emotion classification processing on the first training text information through the GST model to obtain a second emotion label of the first training text information; compare the first emotion label and the second emotion label to obtain an emotion loss value; adjust the parameters of the GST model based on the emotion loss value to obtain an adjusted GST model. Then, the first electronic device can perform emotion recognition processing on the first training text information through the adjusted GST model to obtain an emotion vector of the first training text information.
[0062] As a feasible implementation mode, the first electronic device may perform sentiment classification processing on the first training text information through a cross-domain SER model, and a method for obtaining a first sentiment label of the first training text information may be: the first electronic device selects second training text information whose similarity with the first training text information is greater than a preset ratio threshold based on the MMD algorithm, and performs sentiment classification processing on the second training text information through a cross-domain SER model to obtain a first sentiment label.
[0063] In an embodiment of the present application, since the data sets used by the TTS model and the cross-domain SER model differ in speakers, recording devices, and recording environments, in order to ensure the similarity in distribution of the two data sets, the MMD algorithm can be used to reduce the distribution difference between the two data sets by minimizing the MMD loss.
[0064] Specifically, the MMD algorithm is a method for testing whether two distributions are similar, and is verified to be applicable to the cross-domain SER model. The embodiment of the present application uses the MMD algorithm to reduce the difference between the two data sets, which is actually to use the MMD algorithm to select a SER data set (i.e., the second training sample) that is similar to the TTS data set (i.e., the first training sample).
[0065] After obtaining the features of the SER and TTS datasets, the MMD algorithm is used to take the SER dataset as D s , taking the TTS dataset as D t , the distribution difference between the two datasets is reduced by reducing the MMD loss value (MMD loss value) calculated by the following formula.
[0066]
[0067] Where S i , S j D s, D t The features obtained in, m, n are D s , D t The number of samples in the data set, k(.,.) is the kernel function, which is a linear combination of multiple RBF kernels.
[0068] Based on this, the first electronic device performs sentiment classification processing on the first training text information through the cross-domain SER model to obtain the first sentiment label of the first training text information. Specifically, the cross-domain SER model uses the MMD algorithm to select the second training text information that is relatively similar to the first training text information, and uses the Convolutional Neural Network (CNN)-Recurrent Neural Network (RNN) in the cross-domain SER model to obtain the feature vector of the second training text information, inputs the feature vector into the sentiment classifier in the cross-domain SER model, and performs sentiment classification processing on the feature vector through the sentiment classifier to obtain the first sentiment label.
[0069] The cross-domain SER model is pre-trained, and the cross-domain SER model may not be trained in the embodiment of the present application. The training process of the cross-domain SER module is as follows: obtaining training samples, the training samples include SER data, and the emotion labels of the SER data manually annotated; then obtaining the feature vector of the SER data through the cross-domain SER model, inputting the feature vector into the emotion classifier, obtaining the emotion classification of the feature vector, comparing the obtained emotion classification with the emotion label in the third training sample, obtaining the loss value of the emotion classifier, and training the emotion classifier according to the loss value to realize the training of the cross-domain SER module.
[0070] As a feasible implementation mode, the first electronic device performs emotion recognition processing on the first training text information through the GST model in the speech synthesis model, and the implementation process of obtaining the emotion vector of the first training text information can be: encoding the first training text information through the reference encoder in the GST model to obtain a reference vector, using the attention mechanism to calculate the similarity between the reference vector and the initialization vector to obtain a set of weight values, and performing a weighted operation on the set of weight values and the initialization vector to obtain the emotion vector of the first training text information.
[0071] In the specific implementation, the first electronic device inputs the first training text information into the GST model. The GST model can use the reference encoder to encode the first training text information to obtain a reference vector, and then pass the reference vector into the styletoken layer. The attention mechanism is used to calculate the similarity between the reference vector and each token (where the token is a set of randomly initialized vectors, that is, the rhythm is decomposed into tokens), and finally outputs a set of weight values (indicating the contribution of each style token to the reference vector). Then the GST model embeds the weighted sum of the weight and the token, that is, the style vector, into the TTS model, that is, adding a second emotion label to the original text feature.
[0072] Compared with the traditional GST model, the embodiment of the present application adds an auxiliary sentiment predictor to the style token layer, inputs the token weight obtained from the style token layer into the auxiliary sentiment predictor, uses DNN to classify the first training text information into different sentiment categories, and then compares the sentiment category obtained at this time with the first sentiment label originally obtained by the cross-domain SER model to obtain the sentiment loss value, and adjusts the token weight of the GST model according to the sentiment loss value. Then, the weighted sum of the weight and the token (i.e., the sentiment vector) is embedded in the TTS model.
[0073] S103, performing speech synthesis processing on the first training text information and the emotion vector of the first training text information through a TTS model to obtain predicted speech information corresponding to the first training text information.
[0074] As a feasible implementation, the TTS model may be a Tacotron2 model. The Tacotron2 model may include a text encoder, an attention mechanism module, and a decoder. Since the encoding-decoding method is a sequence-to-sequence conversion, it may be that the sequences cannot be aligned. Therefore, the attention mechanism module can ensure the alignment of sequence elements.
[0075] Exemplarily, the attention mechanism module may include SENet (Squeeze-and-Excitation Networks) or CBAM (Convolutional Block Attention Module). The principle of SENet is to automatically obtain the importance of each feature channel through learning, and then enhance the useful features and suppress the features that are not very useful for the current task according to this importance. CBAM contains two independent sub-modules, the Channel Attention Module (CAM) and the Spatial Attention Module (SAM), which perform channel and spatial attention respectively. This not only saves parameters and computing power, but also ensures that it can be integrated into the existing network architecture as a plug-and-play module.
[0076] It can be understood that the TTS model in the embodiments of the present application includes but is not limited to the Tacotron2 model. For example, the TTS model can be an end-to-end adversarial TTS model (EATS) or a ClariNet model, etc. The ClariNet model refers to a parallel audio waveform (raw audio waveform) generation model based on WaveNet, and the Wavenet model is a sequence generation model.
[0077] S104: Compare the predicted speech information corresponding to the first training text information with the training speech information to obtain a speech synthesis loss value.
[0078] In a specific implementation, the first electronic device may obtain audio features of the predicted voice information and audio features of the training voice information, and then compare the audio features of the predicted voice information with the audio features of the training voice information to obtain a voice synthesis loss value.
[0079] S105, adjusting the parameters of the TTS model and the parameters of the GST model based on the speech synthesis loss value to train the speech synthesis model to obtain a trained speech synthesis model, wherein the trained speech synthesis model includes a trained GST model and a trained TTS model.
[0080] In a specific implementation, the first electronic device can adjust the parameters of the TTS model based on the speech synthesis loss value to obtain a trained TTS model, and adjust the style token weight of the GST model based on the speech synthesis loss value to obtain a trained GST model, thereby realizing GST-TTS joint training, thereby realizing the training of the speech synthesis model and obtaining a trained speech synthesis model.
[0081] In the embodiment of the present application, in terms of converting text into emotional speech, due to the TTS model's requirement for high-quality and large quantities of speech data sets, it becomes difficult to manually perform emotion annotation on the data sets used by the TTS model. However, the embodiment of the present application can achieve the use of the TTS model to convert text into emotional speech without performing emotion annotation on the training samples through the above-mentioned training process.
[0082] In an embodiment of the present application, emotion recognition processing is performed on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information, and the emotion vector of the first training text information is embedded in the TTS model in the speech synthesis model, and then speech synthesis is performed based on the first training text information and the emotion vector through the TTS model. It can be achieved that without emotion annotation of the training samples, the TTS model can still be used to convert text into emotional speech, thereby improving the training efficiency of the speech synthesis model.
[0083] As a feasible implementation method, based on Figure 1 The training method of the speech synthesis model described in the present application provides a training system for the speech synthesis model, such as Figure 2 As shown, Figure 2The present invention is a schematic diagram of the architecture of a training system for a speech synthesis model. The training system for the speech synthesis model may include a GST model, a TTS model, and a cross-domain SER model. The training system for the speech synthesis model may be run in a first electronic device, and the first electronic device may obtain a training sample, and the training sample includes a first training text information and its corresponding training speech information. Then the first electronic device may perform sentiment classification processing on the first training text information through a cross-domain SER model to obtain a first sentiment label of the first training text information, and the first sentiment label is used to indicate the sentiment category of the first training text information. The first electronic device may perform sentiment classification processing on the first training text information through a GST model to obtain a second sentiment label of the first training text information. The first electronic device may compare the first sentiment label with the second sentiment label to obtain a sentiment loss value, and adjust the style token weight of the GST model based on the sentiment loss value. The GST model obtains a sentiment vector based on the adjusted style token weight, and embeds the sentiment vector into a TTS model. The TTS model may include a text encoder, an attention mechanism module, and a decoder, and the text encoder may perform language learning on the first training text information to determine the underlying structural features of the first training text information. The attention mechanism model can align the underlying structural features and sentiment vectors of the first training text information, and then obtain the predicted speech information corresponding to the target text information through the decoder. The TTS model compares the processed predicted speech information with the training speech information to obtain the speech synthesis loss value, and adjusts the parameters of the TTS model and the style token weights of the GST model based on the speech synthesis loss value to obtain the trained TTS model and the trained GST model, thus realizing GST-TTS joint training.
[0084] See also Figure 3 , Figure 3 is a flow chart of a speech synthesis method provided in an embodiment of the present application; Figure 3 The speech synthesis method shown can be executed by the second electronic device, and the scheme includes but is not limited to steps S301 to S303, wherein:
[0085] S301, obtaining target text information.
[0086] In one example, the second electronic device runs a reading client that provides an audiobook function. If a user submits an audiobook instruction for a certain text information (such as a novel or poem, etc.), the second electronic device can obtain the target text information after detecting the audiobook instruction.
[0087] In another example, the second electronic device runs an instant messaging client. When the user is driving or in a bumpy environment where it is inconvenient to browse the device, a conversation interface in the instant messaging client includes at least one text message. If the user needs to convert a certain text message into voice, the user can submit a voice conversion instruction for the text message. After detecting the voice conversion instruction, the second electronic device can obtain the target text information.
[0088] In another example, when a user interacts with an intelligent customer service client in a second electronic device, if the user submits interaction information (the type of interaction information may be text or voice) to the intelligent customer service client through the second electronic device, the intelligent customer service client may determine the target text information to be output to the user based on the interaction information.
[0089] In another example, during the process of intelligent diagnosis and treatment or remote consultation, if the patient is unable to browse the device due to physical reasons (for example, the patient cannot move his body and there is a certain distance between the second electronic device and the patient), then the text information input by the peer user, where the text information is the target text information, illustratively, taking intelligent diagnosis and treatment as an example, the peer user may refer to an intelligent diagnosis and treatment assistant; taking remote consultation as an example, the peer user may refer to a doctor, which is not limited by the embodiments of the present application.
[0090] The second electronic device may be any one or more of a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart vehicle-mounted device, and a smart wearable device. The second electronic device may be the same device as the first electronic device, or the second electronic device may be a different device from the first electronic device, which is not limited by the embodiments of the present application.
[0091] S302, performing sentiment prediction on the target text information through the trained GST model to obtain a sentiment vector of the target text information.
[0092] In a specific implementation, the trained speech synthesis model includes a trained GST model and a trained TTS model. The trained GST model adds an auxiliary supervised emotion prediction task based on style token weights relative to the traditional GST model, so that GTS can better model and emotion-related style features. The second electronic device inputs the acquired target text information into the trained GST model. The trained GST model can add corresponding emotion conditions to the character strings of different emotions in the target text information to obtain the emotion vector of the target text information. The emotion vector includes the emotion unit vector corresponding to the character strings of different emotions in the target text information.
[0093] S303, performing speech synthesis processing on the target text information and the emotion vector of the target text information through the trained TTS model to obtain predicted speech information corresponding to the target text information.
[0094] In a specific implementation, the second electronic device can input the target text information into the trained TTS model, and the trained TTS model can perform language learning on the target text information to determine the underlying structural features of the target text information. At the same time, after the trained GST model obtains the emotional vector of the text information, the emotional vector can be input into the trained TTS model. After obtaining the underlying structural features and emotional vectors of the target text information, the trained TTS model can decode the underlying structural features and emotional vectors of the target text information to obtain the predicted speech information corresponding to the target text information.
[0095] In a feasible embodiment, after the second electronic device obtains the predicted voice information corresponding to the target text information, the predicted voice information can be output. In a specific implementation, after the second electronic device obtains the predicted voice information, the predicted voice information can be displayed. After the user performs a play operation on the predicted voice information (for example, single-clicking or long pressing the predicted voice information, etc.), the second electronic device can generate a play instruction in response to the play operation and play the predicted voice information. Alternatively, after the second electronic device obtains the predicted voice information, the predicted voice information can be played directly. By directly playing the predicted voice information, the embodiment of the present application can facilitate the user to know the specific content of the target text information without browsing the second electronic device.
[0096] The audio of the speech synthesized by the traditional speech synthesis model has a serious mechanical feeling, or even if the intonation conforms to the human speaking pattern, the lack of emotion makes it impossible for the audience to fully immerse themselves in the scene (such as listening to a book), resulting in user loss. Based on this, the embodiment of the present application uses the trained speech synthesis model to predict the emotion of the target text information, obtains the emotion vector of the target text information, and processes the target text information and the emotion vector through the trained speech synthesis model to obtain the predicted speech information corresponding to the target text information, which can ensure that the predicted speech information obtained by speech synthesis is rich in emotion, thereby improving user stickiness.
[0097] An embodiment of the present application further provides a computer storage medium, in which program instructions are stored. When the program instructions are executed, they are used to implement the corresponding methods described in the above embodiments.
[0098] See also Figure 4 , Figure 4 It is a structural schematic diagram of a training device for a speech synthesis model provided in an embodiment of the present application.
[0099] In one implementation of the device of the embodiment of the present application, the device includes the following structure.
[0100] An acquisition unit 401 is used to acquire a training sample, where the training sample includes first training text information and training voice information corresponding to the first training text information;
[0101] The processing unit 402 is used to perform emotion recognition processing on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information, and embed the emotion vector of the first training text information into the TTS model;
[0102] The processing unit 402 is further configured to perform speech synthesis processing on the first training text information and the emotion vector of the first training text information through a TTS model in the speech synthesis model to obtain predicted speech information corresponding to the first training text information;
[0103] The processing unit 402 is further used to compare the predicted speech information with the training speech information to obtain a speech synthesis loss value;
[0104] The processing unit 402 is also used to adjust the parameters of the TTS model and the GST model in the speech synthesis model based on the speech synthesis loss value to train the speech synthesis model to obtain a trained speech synthesis model, and the trained speech synthesis model includes a trained GST model and a trained TTS model.
[0105] In one embodiment, the processing unit 402 is further used to perform emotion classification processing on the first training text information through a cross-domain SER model to obtain a first emotion label of the first training text information before performing emotion recognition processing on the first training text information through a GST model in a speech synthesis model to obtain an emotion vector of the first training text information, where the first emotion label is used to indicate an emotion category of the first training text information;
[0106] The processing unit 402 is further used to perform sentiment classification processing on the first training text information through the GST model to obtain a second sentiment label of the first training text information;
[0107] The processing unit 402 is further configured to compare the first emotion label with the second emotion label to obtain an emotion loss value;
[0108] The processing unit 402 is further configured to adjust the parameters of the GST model based on the sentiment loss value to obtain an adjusted GST model;
[0109] The processing unit 402 performs sentiment classification processing on the first training text information through the GST model in the speech synthesis model to obtain a sentiment vector of the first training text information, including:
[0110] The first training text information is subjected to sentiment recognition processing through the adjusted GST model to obtain a sentiment vector of the first training text information.
[0111] In one implementation, the processing unit 402 performs sentiment classification processing on the first training text information through the cross-domain SER model to obtain a first sentiment label of the first training text information, including:
[0112] Selecting, based on the MMD algorithm, second training text information whose similarity to the first training text information is greater than a preset ratio threshold;
[0113] The second training text information is subjected to sentiment classification processing through the cross-domain SER model to obtain a first sentiment label.
[0114] In one embodiment, the processing unit 402 performs emotion recognition processing on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information, including:
[0115] Encode the first training text information through a reference encoder in the GST model to obtain a reference vector;
[0116] Use the attention mechanism to calculate the similarity between the reference vector and the initialization vector to obtain a set of weight values;
[0117] A set of weight values is weightedly added to the initialization vector to obtain the sentiment vector of the first training text information.
[0118] In one embodiment, the processing unit 402 performs speech synthesis processing on the first training text information and the emotion vector of the first training text information through a TTS model in the speech synthesis model to obtain predicted speech information corresponding to the first training text information, including:
[0119] Performing language learning on the first training text information through the TTS model to obtain underlying structural features of the first training text information;
[0120] The underlying structural features of the first training text information and the sentiment vector of the first training text information are decoded through the TTS model to obtain predicted speech information corresponding to the first training text information.
[0121] In one embodiment, the acquisition unit 401 is further used to adjust the parameters of the TTS model and the GST model in the speech synthesis model based on the speech synthesis loss value in the processing unit 402 to train the speech synthesis model, and after obtaining the trained speech synthesis model, obtain the target text information;
[0122] The processing unit 402 is further used to perform sentiment prediction on the target text information through the trained GST model to obtain a sentiment vector of the target text information;
[0123] The processing unit 402 is further configured to perform speech synthesis processing on the target text information and the emotion vector of the target text information through the trained TTS model to obtain predicted speech information corresponding to the target text information.
[0124] In one embodiment, the acquiring unit 401 acquires the target text information, including:
[0125] Display the target text information;
[0126] Upon receiving a speech synthesis instruction for the target text information, acquiring the target text information;
[0127] The device may also include:
[0128] The output unit 403 is used to perform speech synthesis processing on the target text information and the emotion vector of the target text information through the trained TTS model in the processing unit 402, and after obtaining the predicted speech information corresponding to the target text information, output the predicted speech information corresponding to the target text information.
[0129] In an embodiment of the present application, emotion recognition processing is performed on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information, and the emotion vector of the first training text information is embedded in the TTS model in the speech synthesis model, and then speech synthesis is performed based on the first training text information and the emotion vector through the TTS model. It can be achieved that without emotion annotation of the training samples, the TTS model can still be used to convert text into emotional speech, thereby improving the training efficiency of the speech synthesis model.
[0130] See also Figure 5 , Figure 5 1 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device in the embodiment of the present application includes a power supply module and other structures, and includes a processor 501, a memory 502, and a communication interface 503. The processor 501, the memory 502, and the communication interface 503 can exchange data, and the processor 501 implements the corresponding data processing solution.
[0131] The memory 502 may include a volatile memory, such as a random-access memory (RAM); the memory 502 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the memory 502 may also include a combination of the above-mentioned types of memory.
[0132] The processor 501 may be a central processing unit (CPU) 501. The processor 501 may also be a combination of a CPU and a GPU. In an electronic device, multiple CPUs and GPUs may be included as needed to perform corresponding data processing. In one embodiment, the memory 502 is used to store program instructions. The processor 501 may call program instructions to implement various methods involved in the embodiments of the present application.
[0133] In a first possible implementation, the processor 501 of the electronic device calls the program instructions stored in the memory 502 to perform the following operations:
[0134] Acquire a training sample, where the training sample includes first training text information and training voice information corresponding to the first training text information;
[0135] Performing emotion recognition processing on the first training text information through the GST model in the speech synthesis model to obtain an emotion vector of the first training text information, and embedding the emotion vector of the first training text information into the TTS model;
[0136] Performing speech synthesis processing on the first training text information and the emotion vector of the first training text information by a TTS model in the speech synthesis model to obtain predicted speech information corresponding to the first training text information;
[0137] Compare the predicted speech information with the training speech information to obtain the speech synthesis loss value;
[0138] Based on the speech synthesis loss value, the parameters of the TTS model and the parameters of the GST model in the speech synthesis model are adjusted to train the speech synthesis model to obtain a trained speech synthesis model, which includes a trained GST model and a trained TTS model.
[0139] In one embodiment, before the processor 501 performs emotion recognition processing on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information, it can also implement the following process:
[0140] Performing sentiment classification processing on the first training text information through a cross-domain SER model to obtain a first sentiment label of the first training text information, where the first sentiment label is used to indicate a sentiment category of the first training text information;
[0141] Performing sentiment classification processing on the first training text information through the GST model to obtain a second sentiment label of the first training text information;
[0142] Compare the first emotion label with the second emotion label to obtain an emotion loss value;
[0143] Adjust the parameters of the GST model based on the sentiment loss value to obtain an adjusted GST model;
[0144] The processor 501 performs sentiment classification processing on the first training text information through the GST model in the speech synthesis model to obtain the sentiment vector of the first training text information. The specific implementation process may be:
[0145] The first training text information is subjected to sentiment recognition processing through the adjusted GST model to obtain a sentiment vector of the first training text information.
[0146] In one implementation, the processor 501 performs sentiment classification processing on the first training text information through the cross-domain SER model to obtain the first sentiment label of the first training text information. The specific implementation process may be:
[0147] Selecting, based on the MMD algorithm, second training text information whose similarity to the first training text information is greater than a preset ratio threshold;
[0148] The second training text information is subjected to sentiment classification processing through the cross-domain SER model to obtain a first sentiment label.
[0149] In one embodiment, the processor 501 performs emotion recognition processing on the first training text information through the GST model in the speech synthesis model, and the specific implementation process of obtaining the emotion vector of the first training text information may be:
[0150] Encode the first training text information through a reference encoder in the GST model to obtain a reference vector;
[0151] Use the attention mechanism to calculate the similarity between the reference vector and the initialization vector to obtain a set of weight values;
[0152] A set of weight values is weightedly added to the initialization vector to obtain the sentiment vector of the first training text information.
[0153] In one embodiment, the processor 501 performs speech synthesis processing on the first training text information and the emotion vector of the first training text information through the TTS model in the speech synthesis model to obtain the predicted speech information corresponding to the first training text information. The specific implementation process may be:
[0154] Performing language learning on the first training text information through the TTS model to obtain underlying structural features of the first training text information;
[0155] The underlying structural features of the first training text information and the sentiment vector of the first training text information are decoded through the TTS model to obtain predicted speech information corresponding to the first training text information.
[0156] In one embodiment, the processor 501 adjusts the parameters of the TTS model and the GST model in the speech synthesis model based on the speech synthesis loss value to train the speech synthesis model. After obtaining the trained speech synthesis model, the following process may be further implemented:
[0157] Acquire target text information through the communication interface 503;
[0158] The trained GST model is used to predict the sentiment of the target text information and obtain the sentiment vector of the target text information;
[0159] The trained TTS model is used to perform speech synthesis processing on the target text information and the emotion vector of the target text information to obtain the predicted speech information corresponding to the target text information.
[0160] In one embodiment, the specific implementation process of the processor 501 acquiring the target text information through the communication interface 503 may be:
[0161] Display target text information;
[0162] Upon receiving a speech synthesis instruction for target text information, obtaining the target text information;
[0163] After the processor 501 performs speech synthesis processing on the target text information and the emotion vector of the target text information through the trained TTS model to obtain the predicted speech information corresponding to the target text information, the processor 501 may further implement the following process:
[0164] The predicted speech information corresponding to the target text information is outputted through the communication interface 503 .
[0165] In an embodiment of the present application, emotion recognition processing is performed on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information, and the emotion vector of the first training text information is embedded in the TTS model in the speech synthesis model, and then speech synthesis is performed based on the first training text information and the emotion vector through the TTS model. It can be achieved that without emotion annotation of the training samples, the TTS model can still be used to convert text into emotional speech, thereby improving the training efficiency of the speech synthesis model.
[0166] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM). The computer-readable storage medium can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of blockchain nodes, etc.
[0167] Among them, the blockchain referred to in this application is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.
[0168] The above disclosure is only part of the embodiments of the present application, which certainly cannot be used to limit the scope of rights of the present application. Ordinary technicians in this field can understand that all or part of the processes of implementing the above embodiments and making equivalent changes according to the claims of this application are still within the scope of the invention.
Claims
1. A method for training a speech synthesis model, characterized in that: include: Acquire a training sample, where the training sample includes first training text information and training voice information corresponding to the first training text information; Performing emotion recognition processing on the first training text information through a global style tag (GST) model in a speech synthesis model to obtain an emotion vector of the first training text information, and embedding the emotion vector of the first training text information into a text-to-speech conversion (TTS) model in the speech synthesis model; The emotion recognition process includes: encoding the first training text information by a reference encoder in the GST model to obtain a reference vector; using an attention mechanism to calculate the similarity between the reference vector and the initialization vector to obtain a set of weight values; performing a weighted operation on the set of weight values and the initialization vector to obtain an emotion vector of the first training text information; Performing language learning on the first training text information by using the TTS model to obtain underlying structural features of the first training text information; The underlying structural features of the first training text information and the sentiment vector of the first training text information are aligned by the TTS model to obtain predicted speech information corresponding to the first training text information; the TTS model includes an attention mechanism module, the attention mechanism module includes a compression and excitation network SENet or a convolution block-based attention mechanism CBAM, and the attention mechanism module is determined based on the current task requirements; the SENet enhances features that are useful for the current task and suppresses features that are not very useful for the current task based on the importance of each feature channel, and the CBAM includes a channel attention module CAM and a spatial attention module SAM, which are used to perform channel and spatial attention mechanisms respectively; Comparing the predicted speech information corresponding to the first training text information with the training speech information to obtain a speech synthesis loss value; Adjusting the parameters of the TTS model and the GST model based on the speech synthesis loss value to train the speech synthesis model to obtain a trained speech synthesis model, wherein the trained speech synthesis model includes a trained GST model and a trained TTS model; An instant messaging client is running on the second electronic device, and when the user is driving or in a bumpy environment, if the conversation interface in the instant messaging client includes at least one text message, in response to a speech conversion instruction for the text message, target text message corresponding to the speech conversion instruction is acquired, predicted speech information corresponding to the target text message is determined using the trained speech synthesis model, and the predicted speech information is played; During intelligent diagnosis and treatment or remote consultation, if it is detected that the patient is unable to move his body and the second electronic device reaches a preset distance from the patient, the text information input by the other user is used as the target text information, and the trained speech synthesis model is used to determine the predicted voice information corresponding to the target text information, and the predicted voice information is played.
2. The method according to claim 1, characterized in that Before performing emotion recognition processing on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information, the method further includes: Performing emotion classification processing on the first training text information through a cross-domain speech emotion recognition (SER) model to obtain a first emotion label of the first training text information, where the first emotion label is used to indicate an emotion category of the first training text information; Performing sentiment classification processing on the first training text information through the GST model to obtain a second sentiment label of the first training text information; Compare the first emotion label with the second emotion label to obtain an emotion loss value; Adjusting the parameters of the GST model based on the emotional loss value to obtain an adjusted GST model; The performing sentiment classification processing on the first training text information by using the GST model in the speech synthesis model to obtain the sentiment vector of the first training text information includes: The first training text information is subjected to emotion recognition processing by using the adjusted GST model to obtain an emotion vector of the first training text information.
3. The method according to claim 2, characterized in that The performing sentiment classification processing on the first training text information by using the cross-domain SER model to obtain a first sentiment label of the first training text information includes: Selecting second training text information whose similarity to the first training text information is greater than a preset ratio threshold based on a maximum mean difference (MMD) algorithm; The second training text information is subjected to sentiment classification processing by using the cross-domain SER model to obtain the first sentiment label.
4. The method according to claim 1, characterized in that The method further comprises adjusting the parameters of the TTS model and the GST model in the speech synthesis model based on the speech synthesis loss value to train the speech synthesis model and obtain the trained speech synthesis model. Get target text information; Performing sentiment prediction on the target text information by using the trained GST model to obtain a sentiment vector of the target text information; The trained TTS model is used to perform speech synthesis processing on the target text information and the emotion vector of the target text information to obtain predicted speech information corresponding to the target text information.
5. The method according to claim 4, characterized in that The step of obtaining target text information includes: Display the target text information; Upon receiving a speech synthesis instruction for the target text information, acquiring the target text information; After performing speech synthesis processing on the target text information and the emotion vector of the target text information through the trained TTS model to obtain predicted speech information corresponding to the target text information, the method further includes: Output the predicted speech information corresponding to the target text information.
6. A training device for a speech synthesis model, characterized in that: The device comprises: An acquiring unit, configured to acquire a training sample, wherein the training sample includes first training text information and training voice information corresponding to the first training text information; A processing unit is used to perform emotion recognition processing on the first training text information through the GST model in the speech synthesis model to obtain the emotion vector of the first training text information, and embed the emotion vector of the first training text information into the TTS model; the emotion recognition processing includes: encoding the first training text information through the reference encoder in the GST model to obtain a reference vector; using an attention mechanism to calculate the similarity between the reference vector and the initialization vector to obtain a set of weight values; performing a weighted operation on the set of weight values and the initialization vector to obtain the emotion vector of the first training text information; The processing unit is further used to perform language learning on the first training text information through the TTS model in the speech synthesis model to obtain the underlying structural features of the first training text information; align the underlying structural features of the first training text information and the sentiment vector of the first training text information through the TTS model to obtain the predicted speech information corresponding to the first training text information; the TTS model includes an attention mechanism module, the attention mechanism module includes a compression and excitation network SENet or a convolution block-based attention mechanism CBAM, and the attention mechanism module is determined based on the current task requirements; the SENet enhances the features that are useful for the current task and suppresses the features that are not very useful for the current task based on the importance of each feature channel, and the CBAM includes a channel attention module CAM and a spatial attention module SAM, which are used to perform channel and spatial attention mechanisms respectively; The processing unit is further used to compare the predicted speech information with the training speech information to obtain a speech synthesis loss value; The processing unit is further used to adjust the parameters of the TTS model and the GST model in the speech synthesis model based on the speech synthesis loss value to train the speech synthesis model to obtain a trained speech synthesis model, wherein the trained speech synthesis model includes a trained GST model and a trained TTS model; The processing unit is further configured to, when an instant messaging client is running on the second electronic device and the user is driving or in a bumpy environment, if the conversation interface in the instant messaging client includes at least one text message, in response to a speech conversion instruction for the text message, obtain target text message corresponding to the speech conversion instruction, determine predicted speech message corresponding to the target text message by using the trained speech synthesis model, and play the predicted speech message; The processing unit is also used to use the text information input by the other user as the target text information during intelligent diagnosis and treatment or remote consultation, if it is detected that the patient cannot move his body and the second electronic device reaches a preset distance from the patient, to use the trained speech synthesis model to determine the predicted voice information corresponding to the target text information, and play the predicted voice information.
7. An electronic device, characterized in that: It includes a processor, a memory and a communication interface, wherein the processor, the memory and the communication interface are interconnected, wherein the memory is used to store computer program instructions, and the processor is configured to execute the program instructions to implement the training method of the speech synthesis model as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, which, when executed by a processor, are used to execute the training method for a speech synthesis model as described in any one of claims 1-5.
Citation Information
Patent Citations
Emotional speech generating method and apparatus for controlling emotional intensity
US20210090551A1