Training methods, style generation methods, and devices for text-to-speech style generation models
By training a text reading style generation model across modalities, speaker style is decoupled, solving the problem that text reading style information is limited to a specific speaker, and achieving efficient improvement in speech expressiveness and consistent emotional expression.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-20
- Publication Date
- 2026-04-03
AI Technical Summary
In existing technologies, text reading style information is limited by a specific speaker, which affects the speech performance of the speech synthesis system, and the cost of acquiring training data is high.
By acquiring multiple audio sentence samples and sentence text samples of text reading aloud, audio features and average speaker reading features are extracted, and text encoders and audio encoders are trained. A cross-modal model is used to decouple speaker style and generate a text reading style generation model.
It significantly reduces the cost of acquiring training data, improves the speech performance of speech synthesis systems, achieves consistent emotional expression across speakers, and simplifies the model application process.
Smart Images

Figure CN116741142B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and audio technology, and in particular to a training method for a text reading style generation model, a text reading style generation method, a computer device, and a storage medium. Background Technology
[0002] With the development of artificial intelligence and audio technology, technologies for acquiring text reading style information have emerged. Text reading style can include emotional categories such as happiness, anger, sadness, surprise, fear, and disgust, as well as the emotional intensity corresponding to each emotional category. Text reading style information can be used to improve the speech performance of speech synthesis systems.
[0003] Current methods for obtaining text reading style information require model training and text reading style prediction based on audio data recorded by a specific speaker in a recording studio. This results in the text reading style being limited by the specific speaker, which affects the speech performance of the speech synthesis system. Summary of the Invention
[0004] Therefore, it is necessary to provide a training method, a text reading style generation method, a computer device, and a storage medium for a text reading style generation model to address the aforementioned technical problems.
[0005] Firstly, this application provides a method for training a text-to-speech style generation model. The method includes:
[0006] Obtain multiple audio sentence samples and multiple sentence text samples for text reading aloud, wherein one of the audio sentence samples for text reading aloud and one of the sentence text samples have a corresponding relationship;
[0007] Obtain multiple audio features from the multiple text-to-speech audio sentence samples, and obtain the average speaker reading features from the multiple text-to-speech audio sentence samples;
[0008] The multiple sentence text samples are input into the text encoder to be trained, and the first text reading style prediction information corresponding to each sentence text sample is obtained from the output of the text encoder to be trained.
[0009] The multiple audio features of the multiple text-reading audio sentence samples and the average speaker reading features are input into the audio encoder to be trained to obtain the second text reading style prediction information output by the audio encoder to be trained, which corresponds to each of the text-reading audio sentence samples.
[0010] Based on the similarity between each first text reading style prediction information and each second text reading style prediction information, the text encoder and the audio encoder to be trained are trained; when the similarity between the first text reading style prediction information and the second text reading style prediction information that have a corresponding relationship is greater than or equal to a first similarity threshold, and the similarity between the first text reading style prediction information and the second text reading style prediction information that do not have a corresponding relationship is less than a second similarity threshold, the trained text encoder is obtained as the text reading style generation model.
[0011] In one embodiment, obtaining multiple audio sentence samples and multiple sentence text samples includes:
[0012] Acquire text-to-speech audio data and corresponding text data; the text-to-speech audio data and corresponding text data come from a text-to-speech audio publishing platform; based on the text-to-speech audio data, acquire multiple text-to-speech audio sentence samples that meet preset audio sentence duration conditions; based on the multiple text-to-speech audio sentence samples and the corresponding text data, acquire the sentence text sample corresponding to each text-to-speech audio sentence sample.
[0013] In one embodiment, obtaining multiple text-reading audio sentence samples that meet preset audio sentence duration conditions based on the text-reading audio data includes: performing volume equalization processing on the text-reading audio data to obtain volume-equalized text-reading audio data; and obtaining multiple text-reading audio sentence samples that meet preset audio sentence duration conditions based on the volume-equalized text-reading audio data.
[0014] In one embodiment, obtaining multiple text-to-speech audio sentence samples and multiple sentence text samples includes: obtaining text-to-speech audio data and corresponding text data; the text-to-speech audio data and corresponding text data are from a text-to-speech audio publishing platform; obtaining multiple sentence text samples based on the corresponding text data; and obtaining the multiple text-to-speech audio sentence samples based on the multiple sentence text samples and the text-to-speech audio data.
[0015] In one embodiment, obtaining the plurality of text-to-speech audio sentence samples based on the plurality of sentence text samples and the text-to-speech audio data includes: performing volume equalization processing on the text-to-speech audio data to obtain volume-equalized text-to-speech audio data; and obtaining the plurality of text-to-speech audio sentence samples based on the plurality of sentence text samples and the volume-equalized text-to-speech audio data.
[0016] In one embodiment, acquiring text-to-speech audio data includes: acquiring raw text-to-speech audio data from the text-to-speech audio publishing platform; determining the language distribution information, speaker characteristic information, and accompaniment information of the raw text-to-speech audio data; if the raw text-to-speech audio data meets preset language distribution conditions based on the language distribution information, meets preset speaker conditions based on the speaker characteristic information, and meets preset accompaniment conditions based on the accompaniment information, then the raw text-to-speech audio data is identified as the text-to-speech audio data.
[0017] In one embodiment, obtaining the average speaker reading features of the plurality of text-to-speech audio sentence samples includes: obtaining the average speaker reading features of the plurality of text-to-speech audio sentence samples based on the average fundamental frequency and / or average speech rate of the plurality of text-to-speech audio sentence samples; wherein the average fundamental frequency is obtained by averaging multiple fundamental frequency sequences of the plurality of text-to-speech audio sentence samples, and the average speech rate is obtained by the total reading duration corresponding to the plurality of text-to-speech audio sentence samples and the total number of characters corresponding to the plurality of sentence text samples.
[0018] In one embodiment, inputting the plurality of sentence text samples into the text encoder to be trained includes: for each of the plurality of sentence text samples, performing masking processing on the text content of the sentence text sample according to a first preset ratio to obtain a plurality of masked sentence text samples; inputting the plurality of masked sentence text samples into the text encoder to be trained; and / or, inputting a plurality of audio features of the plurality of text-reading audio sentence samples and the average speaker reading features into the audio encoder to be trained includes: for each of the plurality of audio features, performing masking processing on the feature content of the audio feature according to a second preset ratio to obtain a plurality of masked audio features; inputting the plurality of masked audio features of the plurality of text-reading audio sentence samples and the average speaker reading features into the audio encoder to be trained.
[0019] Secondly, this application provides a method for generating a text-to-speech style. The method includes: acquiring a text to be read aloud; inputting the text to be read aloud into a trained text-to-speech style generation model; the trained text-to-speech style generation model being trained according to the method described in any of the above embodiments; and acquiring the text-to-speech style information corresponding to the text to be read aloud, output by the trained text-to-speech style generation model.
[0020] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0021] A plurality of audio sentence samples and multiple sentence text samples are obtained, wherein one audio sentence sample and one sentence text sample have a corresponding relationship; multiple audio features of the plurality of audio sentence samples are obtained, and the average speaker reading features of the plurality of audio sentence samples are obtained; the plurality of sentence text samples are input into a text encoder to be trained, and the first text reading style prediction information corresponding to each sentence text sample output by the text encoder to be trained is obtained; the multiple audio features of the plurality of audio sentence samples and the average speaker reading features are input into the audio encoder to be trained, and the first text reading style prediction information corresponding to each sentence text sample is obtained. The audio encoder outputs second text reading style prediction information corresponding to each of the text reading audio sentence samples; based on the similarity between each first text reading style prediction information and each second text reading style prediction information, the text encoder to be trained and the audio encoder to be trained are trained; when the similarity between the first text reading style prediction information and the second text reading style prediction information with a corresponding relationship is greater than or equal to a first similarity threshold, and the similarity between the first text reading style prediction information and the second text reading style prediction information without a corresponding relationship is less than a second similarity threshold, the trained text encoder is obtained as the text reading style generation model.
[0022] Fourthly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0023] Obtain the text to be read aloud; input the text to be read aloud into a trained text reading style generation model; the trained text reading style generation model is trained according to the method described in any of the above embodiments; obtain the text reading style information corresponding to the text to be read aloud output by the trained text reading style generation model.
[0024] Fifthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0025] A plurality of audio sentence samples and multiple sentence text samples are obtained, wherein one audio sentence sample and one sentence text sample have a corresponding relationship; multiple audio features of the plurality of audio sentence samples are obtained, and the average speaker reading features of the plurality of audio sentence samples are obtained; the plurality of sentence text samples are input into a text encoder to be trained, and the first text reading style prediction information corresponding to each sentence text sample output by the text encoder to be trained is obtained; the multiple audio features of the plurality of audio sentence samples and the average speaker reading features are input into the audio encoder to be trained, and the first text reading style prediction information corresponding to each sentence text sample is obtained. The audio encoder outputs second text reading style prediction information corresponding to each of the text reading audio sentence samples; based on the similarity between each first text reading style prediction information and each second text reading style prediction information, the text encoder to be trained and the audio encoder to be trained are trained; when the similarity between the first text reading style prediction information and the second text reading style prediction information with a corresponding relationship is greater than or equal to a first similarity threshold, and the similarity between the first text reading style prediction information and the second text reading style prediction information without a corresponding relationship is less than a second similarity threshold, the trained text encoder is obtained as the text reading style generation model.
[0026] Sixthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0027] Obtain the text to be read aloud; input the text to be read aloud into a trained text reading style generation model; the trained text reading style generation model is trained according to the method described in any of the above embodiments; obtain the text reading style information corresponding to the text to be read aloud output by the trained text reading style generation model.
[0028] The above-mentioned text-to-speech style generation model training method, text-to-speech style generation method, computer equipment, and storage medium acquire multiple text-to-speech audio sentence samples and multiple sentence text samples, wherein a text-to-speech audio sentence sample and a sentence text sample have a corresponding relationship. Multiple audio features of the multiple text-to-speech audio sentence samples are acquired, as well as the average speaker reading features of the multiple text-to-speech audio sentence samples. The multiple sentence text samples are input into a text encoder to be trained, and the output of the encoder, corresponding to the first text-to-speech style prediction information for each sentence text sample, is acquired. The multiple audio features and the average speaker reading features of the multiple text-to-speech audio sentence samples are input into an audio encoder to be trained, and the output of the audio encoder, corresponding to the second text-to-speech style prediction information for each text-to-speech audio sentence sample, is acquired. The text encoder and audio encoder are trained based on the similarity of the first and second text-to-speech style prediction information. When the similarity between the corresponding first and second text-to-speech style prediction information is greater than or equal to a first similarity threshold, and the similarity between the non-corresponding first and second text-to-speech style prediction information is less than a second similarity threshold, the trained text encoder is obtained as the text-to-speech style generation model. The training data for this scheme can come from text-to-speech audio publishing platforms, without relying on audio data recorded by specific speakers in recording studios. This significantly saves on the cost of acquiring training data and allows for the extraction of average speaker reading features for cross-modal model training. On the one hand, this enables the text reading style information predicted by the trained model to be better decoupled from the speaker's style. Integrating this text reading style information into speech synthesis systems of different speakers can also yield more consistent emotional expression, improving the speech performance of the speech synthesis system. On the other hand, when applying the model, only the text to be read needs to be input to obtain the text reading style information, which is convenient for integrating into the speech synthesis system to guide it in synthesizing high-expressive speech. Attached Figure Description
[0029] Figure 1 This is a diagram illustrating the application environment of the training method for the text reading style generation model in the embodiments of this application;
[0030] Figure 2 This is a flowchart illustrating the training method of the text reading style generation model in the embodiments of this application;
[0031] Figure 3 This is a schematic diagram of the model training process in an embodiment of this application;
[0032] Figure 4 This is a schematic diagram of the audio encoder processing data in an embodiment of this application;
[0033] Figure 5 This is a flowchart illustrating the steps for obtaining audio data for text reading in an embodiment of this application;
[0034] Figure 6 This is a flowchart illustrating the training method of a text-to-speech style generation model in another embodiment of this application;
[0035] Figure 7(a) is an internal structure diagram of the computer device in an embodiment of this application;
[0036] Figure 7(b) is an internal structural diagram of a computer device in another embodiment of this application. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0038] The training method and text-to-speech style generation method of the text-to-speech style generation model provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, the environment may include a terminal 110 and a server 120. The terminal 110 can communicate with the server 120 via a network. A data storage system can store the data that the server 120 needs to process. The data storage system can be integrated onto the server 120 or located in the cloud or on other network servers. The terminal 110 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. The server 120 can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0039] Specifically, the training method for the text-to-speech style generation model provided in this application embodiment can be executed by server 120. Server 120 can train a trained text encoder as a text-to-speech style generation model according to the training method for the text-to-speech style generation model provided in this application embodiment. Then, server 120 can deploy the text-to-speech style generation model locally or send it to terminal 110. Therefore, the text-to-speech style generation method provided in this application embodiment can be executed by server 120 or terminal 110. Server 120 can obtain the text to be read from terminal 110, and then input the text to be read into the aforementioned trained text-to-speech style generation model to obtain the text-to-speech style information corresponding to the text to be read output by the trained text-to-speech style generation model. The text-to-speech style information can be further integrated into a speech synthesis system to improve speech performance. Terminal 110 can also obtain the text to be read, and then input the text to be read into the aforementioned trained text-to-speech style generation model received from server 120 to obtain the text-to-speech style information corresponding to the text to be read output by the trained text-to-speech style generation model. The text-to-speech style information can be further integrated into a speech synthesis system to improve speech performance.
[0040] Speech synthesis technology, as a crucial component of human-computer interaction, aims to achieve synthesized effects comparable to those of real people. High-expression speech synthesis is gradually becoming a future trend, with significant application value in the production of audiobooks and the creation of virtual digital humans. High-expression speech can possess remarkable characteristics such as natural rhythm, rich emotional style, and clear sound quality. However, current technologies still have a significant, easily discernible gap between themselves and real people in terms of natural rhythm representation and rich emotional style. The method in this application, through the training and application of a text reading style generation model, can predict text reading style information from the text. Applying this information to a speech synthesis system can significantly improve the expressiveness of synthesized audiobooks and the naturalness of speech interaction technology.
[0041] Current methods for acquiring text reading style information inevitably require the collection of large amounts of high-quality audio data to construct the speech synthesis system in order to improve generalization ability. This data cost is enormous, and the text reading style information obtained through training is limited to a specific individual speaker. That is, the same reading style yields significantly different emotional expressions on different speakers' speech synthesis systems, affecting the speech performance of the speech synthesis system. In contrast, the method of this application does not rely on audio data recorded by a specific speaker in a recording studio, which can significantly save the cost of acquiring training data. Furthermore, the method of this application extracts average speaker reading features for cross-modal model training. On the one hand, the text reading style information predicted by the trained model can be better decoupled from the speaker style. Integrating this text reading style information into speech synthesis systems of different speakers can also yield more consistent emotional expressions, improving the speech performance of the speech synthesis system. On the other hand, when applying the model, only the text to be read needs to be input to obtain the text reading style information, which is convenient for integrating into the speech synthesis system to guide the synthesis of high-expressive speech.
[0042] The following section is based on, for example Figure 1 The application environment shown, along with various embodiments and corresponding figures, will be used to further explain the training method and text reading style generation method of the text reading style generation model of this application.
[0043] In one embodiment, such as Figure 2 As shown, a training method for a text-to-speech style generation model is provided. This method can be achieved by, for example, training a text-to-speech style generation model. Figure 1 The method, executed by server 120, includes the following steps:
[0044] Step S201: Obtain multiple audio sentence samples and multiple sentence text samples for text reading.
[0045] In this system, there is a correspondence between a text-to-speech audio sentence sample and a corresponding text sample. Specifically, server 120 can segment the text-to-speech audio data and its corresponding text data into sentences. The duration of a sentence can be limited to a certain range, such as 0 to 10 seconds. This allows server 120 to obtain multiple text-to-speech audio sentences and multiple corresponding texts for each audio sentence. Because these texts correspond to the audio sentences, they are called sentence texts. Server 120 uses these multiple audio sentences and their corresponding sentence texts as training samples for model training. Therefore, the audio sentences are further denoted as text-to-speech audio sentence samples, and the sentence texts are denoted as sentence text samples. Thus, server 120 obtains multiple audio sentence samples and their corresponding sentence text samples. For example, “Ah! This is truly a happy thing!” can be considered a text sample, and the audio of “Ah! This is truly a happy thing!” is the corresponding text-to-speech audio sentence sample. The audio data for text reading can include the audio data of the text to be read by the user, and the corresponding text data can include the text content corresponding to the audio data of the text to be read by the user. In practical applications, the longer the total duration of the audio data for text reading, the more speakers it covers, and the more types of text to be read, the more accurate the prediction effect of the trained text reading style generation model will be on the text reading style information.
[0046] Step S202: Obtain multiple audio features of multiple text-to-speech audio sentence samples, and obtain the average speaker reading features of multiple text-to-speech audio sentence samples.
[0047] This step requires acquiring two types of features: firstly, the audio features of each text-reading audio sentence sample itself; and secondly, the average speaker reading features, which reflect the average reading features of all speakers covered by the multiple text-reading audio sentence samples as a whole. The audio features can be obtained using Mel spectra. Specifically, for each text-reading audio sentence sample, the Mel spectrum corresponding to that sample is obtained, thus the server 120 obtains multiple audio features corresponding to multiple text-reading audio sentence samples. The reason for extracting Mel spectra as audio features is that the frequency range audible to the human ear is 20 to 20,000 Hz, but the human ear does not perceive this scale unit of Hertz linearly. Mel spectrum extraction first involves framing and windowing the text-reading audio sentence samples, then calculating the linear spectrum using Fourier transform, and finally transforming the linear spectrum into a Mel spectrum using a Mel-scale filter bank. In a practical implementation, 80 sets of Mel triangular filters can be used to extract features from each text-reading audio sentence sample.
[0048] The average speaker reading characteristics may include at least one of the average fundamental frequency and average speech rate of each speaker. In some embodiments, the method of this application may further include the following steps:
[0049] Based on multiple audio sentence samples of text reading, multiple fundamental frequency sequences corresponding to the multiple audio sentence samples of text reading are obtained; the multiple fundamental frequency sequences corresponding to the multiple audio sentence samples of text reading are averaged to obtain the average fundamental frequency.
[0050] In other words, the average fundamental frequency can be obtained by averaging multiple fundamental frequency sequences from multiple text-to-speech audio sentence samples. The solution in this embodiment is mainly used to obtain the average fundamental frequency of each speaker. Specifically, for the average fundamental frequency, the pYin fundamental frequency extraction method can be used. For each text-to-speech audio sentence sample in the multiple text-to-speech audio sentence samples, the frame-level fundamental frequency sequence corresponding to the text-to-speech audio sentence sample can be obtained, thus obtaining multiple fundamental frequency sequences corresponding to multiple text-to-speech audio sentence samples. This covers the fundamental frequency sequences of all text-to-speech audio sentence samples for each speaker. Averaging these multiple fundamental frequency sequences corresponding to multiple text-to-speech audio sentence samples, i.e., averaging the fundamental frequency sequences of all text-to-speech audio sentence samples for each speaker, yields the average fundamental frequency.
[0051] In other embodiments, the method of this application may further include the following steps:
[0052] The average speaking speed is obtained by considering the total reading time of multiple audio sentence samples and the total number of words in multiple sentence text samples.
[0053] That is, the average speech rate can be obtained from the total reading time of multiple text-to-speech audio sentence samples and the total number of words in the text samples. The solution in this embodiment is mainly used to obtain the average speech rate of each speaker. Specifically, for the average speech rate, the server 120 can obtain the total reading time of multiple text-to-speech audio sentence samples and the total number of words in the text samples. Then, it calculates the average speech rate based on the total reading time and the total number of words. In a specific implementation, the server 120 can calculate the average speech rate of multiple text-to-speech audio sentence samples based on the total reading time of all text-to-speech audio samples from all speakers and the total number of words in all text samples covered in the text-to-speech audio data.
[0054] Based on this, in some embodiments, obtaining the average speaker reading features of multiple text-to-speech audio sentence samples in step S202 may include: obtaining the average speaker reading features of multiple text-to-speech audio sentence samples based on the average fundamental frequency and / or average speech rate of the multiple text-to-speech audio sentence samples. Specifically, when only the average fundamental frequency or average speech rate is obtained, the server 120 can use the average fundamental frequency or average speech rate as the average speaker reading features of the multiple text-to-speech audio sentence samples; when both the average fundamental frequency and average speech rate are obtained, the server 120 can use both the average fundamental frequency and average speech rate as the average speaker reading features of the multiple text-to-speech audio sentence samples. The use of average speaker reading features, including average fundamental frequency and average speech rate, in this application is significant in avoiding the influence of the overall reading style of different speakers on the subsequent model's extraction of text reading style.
[0055] Step S203: Input multiple sentence text samples into the text encoder to be trained, and obtain the first text reading style prediction information corresponding to each sentence text sample output by the text encoder to be trained.
[0056] Step S204: Input multiple audio features of multiple text-reading audio sentence samples and average speaker reading features into the audio encoder to be trained, and obtain the second text reading style prediction information output by the audio encoder to be trained, which corresponds to each text-reading audio sentence sample.
[0057] Step S205: Based on the similarity between each first text reading style prediction information and each second text reading style prediction information, train the text encoder and the audio encoder to be trained; when the similarity between the first text reading style prediction information and the second text reading style prediction information that have a corresponding relationship is greater than or equal to the first similarity threshold, and the similarity between the first text reading style prediction information and the second text reading style prediction information that do not have a corresponding relationship is less than the second similarity threshold, the trained text encoder is obtained as the text reading style generation model.
[0058] Steps S203 to S205 described above constitute the main process by which this application trains the model after obtaining multiple sentence text samples, multiple audio features corresponding to multiple text reading audio sentence samples, and average speaker reading features. In this regard, combined with... Figure 3To explain, in step S203, multiple sentence text samples are input into the text encoder to be trained, and the first text reading style prediction information (T1, T2, T3, ..., Tn) corresponding to each sentence text sample is obtained from the output of the text encoder to be trained, where n represents the number of multiple sentence text samples. In step S204, multiple audio features (Mel spectrograms) and average speaker reading features (average fundamental frequency, average speech rate) of multiple text reading audio sentence samples are input into the audio encoder to be trained, and the second text reading style prediction information (A1, A2, A3, ..., An) corresponding to each text reading audio sentence sample is obtained from the output of the audio encoder to be trained.
[0059] For the text encoder, the open-source BERT (Bidirectional Encoder Representation from Transformers) model can be used. This model is a pre-trained language representation model. The final output of the text encoder is the hidden state of the last layer, with dimensions [batch_size, embed_dim]. Here, batch_size is the number of samples n in a single batch of training, and embed_dim is the vector dimension corresponding to the predicted first text reading style information, which can be set to 256.
[0060] For the audio encoder, the input can specifically include multiple Mel spectra, as well as the average fundamental frequency and average speech rate. As mentioned earlier, the significance of using the average fundamental frequency and average speech rate is to avoid the overall reading style of different speakers affecting the extraction of the reading style prediction information of the second text. Let the dimension of the Mel spectra be [T, 80], where T represents time, and the average fundamental frequency and average speech rate are quantized single values. After performing word embedding operations on each, word embedding vectors of dimension [1, 80] can be obtained. Then, the word embedding vectors are copied T times and concatenated with the Mel spectra to obtain a feature representation of dimension [T, 240]. Further combining... Figure 4The feature representation [T, 240] is input into the residual network of the audio encoder. Then, different convolutions (kernel sizes of 1, 3, 5, 7, and 9, with a maximum of 128 channels) are used in the convolutional layers to capture local features. The convolutional results are then concatenated to obtain a vector of [T, 640]. A subsequent linear layer and Rectified Linear Unit (ReLU) reduce the feature dimension to [T, 256]. Finally, a Bi-Gated Recurrent Unit (Bi-GRU) network further enhances the temporal modeling capability. After extracting the last state layer of the Bi-GRU network, a linear layer is used for transformation to obtain the second text reading style prediction information, which has the same vector dimension of 256 as the first text reading style prediction information. Note that the input audio Mel spectrum also undergoes a 20% random masking operation to improve the modeling capability of the audio encoder.
[0061] Based on this, combined Figure 3 In step S205, the similarity (X11, X12, ..., Xnn) between each first text reading style prediction information (T1, T2, T3, ..., Tn) and each second text reading style prediction information (A1, A2, A3, ..., An) can be obtained first, and the text encoder and audio encoder to be trained can be trained based on the similarity (X11, X12, ..., Xnn).
[0062] Specifically, the principle of training the model in this application mainly involves enabling the model to predict whether a given audio feature and a given sentence text sample are paired. A cross-modal model can be pre-trained using contrastive learning as the loss function. This cross-modal model can predict the first text reading style prediction information corresponding to the sentence text sample through the aforementioned text encoder, and predict the second text reading style prediction information through the aforementioned audio encoder based on audio features and average speaker reading features. When a given audio feature and a given sentence text sample are a pair, their similarity is high; when they are not a pair, their similarity is low. Therefore, for a batch containing n audio feature-sentence text sample pairs, the positive samples are audio features and sentence text samples with corresponding relationships, totaling n, while all other combinations of audio features and sentence text samples are unpaired, meaning there are n×nn negative samples. Therefore, after calculating the similarity between each first text reading style prediction information and each second text reading style prediction information, the objective function of contrastive learning is to make the similarity of positive sample pairs higher and the similarity of negative sample pairs lower. Specifically, in this step, when the similarity between the first text reading style prediction information and the second text reading style prediction information that have a corresponding relationship is greater than or equal to the first similarity threshold, and the similarity between the first text reading style prediction information and the second text reading style prediction information that do not have a corresponding relationship is less than the second similarity threshold, the server 120 can obtain the trained text encoder and use the trained text encoder as the text reading style generation model.
[0063] The training method for the above-mentioned text-to-speech style generation model involves acquiring multiple audio sentence samples and multiple sentence text samples, wherein one audio sentence sample and one sentence text sample have a corresponding relationship. Multiple audio features of the multiple audio sentence samples are acquired, as well as the average speaker reading features of the multiple audio sentence samples. These multiple sentence text samples are input into a text encoder to be trained, and the output of the encoder yields first text-to-speech style prediction information corresponding to each sentence text sample. The multiple audio features and the average speaker reading features of the multiple audio sentence samples are input into an audio encoder to be trained, and the output of the audio encoder yields second text-to-speech style prediction information corresponding to each audio sentence sample. The text encoder and audio encoder are trained based on the similarity between the first and second text-to-speech style prediction information. When the similarity between the corresponding first and second text-to-speech style prediction information is greater than or equal to a first similarity threshold, and the similarity between the non-corresponding first and second text-to-speech style prediction information is less than a second similarity threshold, the trained text encoder is obtained as the text-to-speech style generation model. The training data for this scheme can come from text-to-speech audio publishing platforms, without relying on audio data recorded by specific speakers in recording studios. This significantly saves on the cost of acquiring training data and allows for the extraction of average speaker reading features for cross-modal model training. On the one hand, this enables the text reading style information predicted by the trained model to be better decoupled from the speaker's style. Integrating this text reading style information into speech synthesis systems of different speakers can also yield more consistent emotional expression, improving the speech performance of the speech synthesis system. On the other hand, when applying the model, only the text to be read needs to be input to obtain the text reading style information, which is convenient for integrating into the speech synthesis system to guide it in synthesizing high-expressive speech.
[0064] In one embodiment, step S201, obtaining multiple audio sentence samples and multiple sentence text samples for text reading, includes:
[0065] Acquire the text-to-speech audio data and the corresponding text data; based on the text-to-speech audio data, acquire multiple text-to-speech audio sentence samples that meet the preset audio sentence duration conditions; based on the multiple text-to-speech audio sentence samples and the corresponding text data, acquire the sentence text sample corresponding to each text-to-speech audio sentence sample.
[0066] In this embodiment, server 120 can acquire text-to-speech audio data and corresponding text data from a text-to-speech audio publishing platform. Specifically, the acquired text-to-speech audio data and corresponding text data can both originate from the text-to-speech audio publishing platform. This platform refers to a platform for various speakers (users) to publish text-to-speech audio content; specifically, it can be a platform for publishing various audiobook reading data. With relevant authorization, server 120 can acquire the text-to-speech audio data and corresponding text data provided by the platform. In practical applications, the longer the total duration of the text-to-speech audio data, the more speakers it covers, and the more types of the read-aloud text, the more accurate the prediction effect of the trained text-to-speech style generation model on text-to-speech style information will be. Then, server 120 can first acquire multiple text-to-speech audio sentence samples that meet the preset audio sentence duration conditions based on the text-to-speech audio data. Then, based on the multiple text-to-speech audio sentence samples and their corresponding text data, and according to the temporal correspondence, it can extract multiple sentence texts corresponding to the multiple text-to-speech audio sentence samples from the text data, thereby obtaining the corresponding multiple sentence text samples. Among them, the preset audio sentence duration condition can be a preset duration, which can be set to a duration within the range of 0 to 10 seconds. This allows relevant personnel to flexibly set the required duration of the text reading audio sentence sample and thereby obtain multiple corresponding sentence text samples.
[0067] Furthermore, in some embodiments, obtaining multiple text-to-speech audio sentence samples that meet preset audio sentence duration conditions based on the text-to-speech audio data in the above embodiments may include:
[0068] The audio data of the text reading is subjected to volume equalization processing to obtain the volume equalization processed audio data of the text reading; based on the volume equalization processed audio data of the text reading, multiple audio sentence samples of the text reading that meet the preset audio sentence duration conditions are obtained.
[0069] In this embodiment, before acquiring multiple text-to-speech audio sentence samples, the text-to-speech audio data can be processed by volume equalization. Specifically, the sampling rate can be unified to 16000Hz, and then the loudness of the text-to-speech audio data can be set to a uniform value to obtain the volume-equalized text-to-speech audio data. Then, based on the volume-equalized text-to-speech audio data, multiple text-to-speech audio sentence samples that meet the preset audio sentence duration conditions can be acquired, thereby avoiding the introduction of volume differences between different audio data and improving the prediction accuracy of the trained model.
[0070] In another embodiment, obtaining multiple text-to-speech audio sentence samples and multiple sentence text samples in step S201 may include:
[0071] Obtain the audio data of the text-to-speech function and the corresponding text data; obtain multiple sentence text samples based on the corresponding text data; and obtain multiple audio sentence samples of the text-to-speech function based on the multiple sentence text samples and the audio data of the text-to-speech function.
[0072] In this embodiment, server 120 can acquire text-to-speech audio data and corresponding text data from a text-to-speech audio publishing platform. Specifically, the acquired text-to-speech audio data and corresponding text data can both originate from the text-to-speech audio publishing platform. This platform refers to a platform for various speakers (users) to publish text-to-speech audio content; specifically, it can be a platform for publishing various audiobook reading data. Server 120 can acquire the text-to-speech audio data and corresponding text data provided by the text-to-speech audio publishing platform with the relevant authorization. In practical applications, the longer the total duration of the text-to-speech audio data, the more speakers it covers, and the more types of text it reads, the more accurate the prediction effect of the trained text-to-speech style generation model on text-to-speech style information will be. Then, server 120 can first obtain multiple sentence text samples based on the corresponding text data. Specifically, it can segment multiple sentence text samples from the corresponding text data on a sentence-by-sentence basis. Then, based on the multiple sentence text samples and the text reading audio data, and based on the temporal correspondence, it can obtain multiple text reading audio sentence samples corresponding to the multiple sentence text samples from the text reading audio data. In this way, high-quality corresponding sentence text samples and text reading audio sentence samples can be obtained on a sentence-by-sentence basis.
[0073] Furthermore, in some embodiments, obtaining multiple text-to-speech audio sentence samples based on multiple sentence text samples and text-to-speech audio data in the above embodiments may include:
[0074] The audio data of the text reading is subjected to volume equalization processing to obtain the audio data of the text reading after volume equalization; based on multiple sentence text samples and the audio data of the text reading after volume equalization, multiple audio sentence samples of the text reading are obtained.
[0075] In this embodiment, similarly, before acquiring multiple text-to-speech audio sentence samples, the text-to-speech audio data is first subjected to volume equalization processing. Specifically, the sampling rate can be unified to 16000Hz, and then the loudness of the text-to-speech audio data is set to a uniform value to obtain volume-equalized text-to-speech audio data. Then, based on the time correspondence between the multiple sentence text samples and the volume-equalized text-to-speech audio data, multiple text-to-speech audio sentence samples corresponding to the multiple sentence text samples are obtained from the volume-equalized text-to-speech audio data, thereby avoiding the introduction of volume differences between different audio data and improving the prediction accuracy of the trained model.
[0076] In some embodiments, such as Figure 5 As shown, the acquisition of text-to-speech audio data in the above embodiments specifically includes:
[0077] Step S501: Obtain the original text-to-speech audio data from the text-to-speech audio publishing platform.
[0078] In this step, as mentioned above, server 120 can obtain text-to-speech audio data provided by the text-to-speech audio publishing platform with the relevant authorization. Server 120 first sets the text-to-speech audio data as the original text-to-speech audio data.
[0079] Step S502: Determine the language distribution information, speaker characteristic information, and accompaniment information of the original text reading audio data.
[0080] Specifically, server 120 can analyze the original text-to-speech audio data using relevant recognition models to determine the language distribution information, speaker characteristic information, and accompaniment information of the original text-to-speech audio data. The language distribution information can indicate how many languages the original text-to-speech audio data covers, the speaker characteristic information can indicate which types and how many speakers the original text-to-speech audio data covers, and the accompaniment information can indicate whether the original text-to-speech audio data contains accompaniment.
[0081] Step S503: If the original text reading audio data meets the preset language distribution conditions based on the language distribution information, meets the preset speaker conditions based on the speaker characteristic information, and meets the preset accompaniment conditions based on the accompaniment information, then the original text reading audio data is determined as text reading audio data.
[0082] In this step, server 120 can determine whether the original text-to-speech audio data meets preset language distribution conditions based on language distribution information. These preset language distribution conditions can be conditions used to determine whether the original text-to-speech audio data covers several specified languages and whether the audio data corresponding to these languages reaches a predetermined duration, such as 12 hours for English audio data. It can also determine whether the original text-to-speech audio data meets preset speaker conditions based on speaker characteristic information. These preset speaker conditions can be conditions used to determine whether the original text-to-speech audio data covers specified speaker types and their corresponding numbers, such as including a male speaker and a female speaker. Furthermore, it can determine whether the original text-to-speech audio data meets preset accompaniment conditions based on accompaniment information. These preset accompaniment conditions can be conditions used to determine whether the original text-to-speech audio data does not contain accompaniment. Therefore, if server 120 determines that the original text-to-speech audio data meets the preset language distribution conditions, preset speaker conditions, and preset accompaniment conditions, then server 120 can identify the original text-to-speech audio data as text-to-speech audio data. The solution in this embodiment can improve the accuracy of model training with the obtained text reading audio data, and compared with the audio data required for traditional speech synthesis, the sound quality requirements and corresponding text quality requirements are lower, significantly reducing the data acquisition cost.
[0083] In one embodiment, step S203, inputting multiple sentence text samples into the text encoder to be trained, may include:
[0084] For each sentence text sample in the multiple sentence text samples, the text content in the sentence text sample is masked according to a first preset ratio to obtain multiple masked sentence text samples; the multiple masked sentence text samples are then input into the text encoder to be trained.
[0085] Specifically, as mentioned earlier, the text encoder can use the open-source BERT model. Unlike traditional methods that use one-way language models or shallowly concatenate two one-way language models, in this embodiment, for each sentence text sample among multiple sentence text samples, a text content masking method is used to generate a deep language representation. Specifically, for each sentence text sample, the text content (such as words) in the sentence text sample is randomly masked according to a first preset ratio (such as 30%), thereby obtaining multiple masked sentence text samples. Then, these multiple masked sentence text samples are input into the text encoder to be trained (which can be composed of multiple Transformer modules), thereby enhancing the text encoder's ability to model text reading style information.
[0086] In one embodiment, step S204, inputting multiple audio features of multiple text-reading audio sentence samples and average speaker reading features into the audio encoder to be trained, may include:
[0087] For each of the multiple audio features, the feature content in the audio feature is masked according to a second preset ratio to obtain multiple masked audio features; the multiple masked audio features of multiple text reading audio sentence samples and the average speaker reading features are input into the audio encoder to be trained.
[0088] Similarly, in this embodiment, for each of the multiple audio features, before inputting it to the audio encoder, the server 120 can randomly mask the feature content of the audio feature according to a second preset ratio (such as 20%), thereby obtaining multiple masked audio features. Then, the multiple masked audio features and the average speaker reading features are input into the audio encoder to be trained, thereby improving the audio encoder's ability to model text reading style information.
[0089] In one embodiment, such as Figure 6 As shown, a method for generating text-to-speech style is provided, which can be achieved by, for example... Figure 1 The method, executed by terminal 110, may include the following steps:
[0090] Step S601: Obtain the text to be read aloud.
[0091] In this step, terminal 110 can obtain the text to be read aloud, either provided by the user or by server 120.
[0092] Step S602: Input the text to be read aloud into the trained text reading style generation model.
[0093] Specifically, the trained text-to-speech style generation model can be trained by server 120 according to the training method of the text-to-speech style generation model as described in any of the above embodiments, and then sent by server 120 to terminal 110 for use. In this step, terminal 110 inputs the text to be read aloud into the trained text-to-speech style generation model. It should be noted that there is no need to perform random masking processing on the text to be read aloud. In the model application stage, it is only necessary to input the text to be read aloud into the trained text-to-speech style generation model, i.e., the text encoder.
[0094] Step S603: Obtain the text reading style information corresponding to the text to be read out, output by the trained text reading style generation model.
[0095] In this step, server 120 obtains the text reading style generation model, which has been trained, and outputs the text reading style information corresponding to the text to be read based on the input text to be read. This text reading style information can be connected to the speech synthesis system to guide it in synthesizing highly expressive speech.
[0096] The solution in this embodiment can construct a cross-modal model for unsupervised comparative learning by using a large number of audiobooks with emotional readings and texts. In the text reading style prediction stage, the text to be read can be directly input into the trained text encoder to predict the corresponding text reading style information. This text reading style information can be used by the speech synthesis system to control aspects such as speech rate, pitch, and emotion, thereby achieving highly expressive speech synthesis.
[0097] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0098] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram is shown in Figure 7(a). The computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores data such as text-to-speech audio data and text data. The input / output interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a training method and a text-to-speech style generation method for a text-to-speech style generation model.
[0099] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as shown in Figure 7(b). The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for wired or wireless communication with external terminals; wireless communication can be achieved through WIFI, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a text-to-speech style generation method. The display unit of the computer device is used to form a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0100] Those skilled in the art will understand that the structures shown in Figures 7(a) and 7(b) are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements.
[0101] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0102] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0103] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0104] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0105] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0106] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A training method for a text-to-speech style generation model, characterized in that, The method includes: Obtain multiple audio sentence samples and multiple sentence text samples for text reading aloud, wherein one of the audio sentence samples for text reading aloud and one of the sentence text samples have a corresponding relationship; Obtain multiple audio features from the multiple text-to-speech audio sentence samples, and obtain the average speaker reading features from the multiple text-to-speech audio sentence samples; The multiple sentence text samples are input into the text encoder to be trained, and the first text reading style prediction information corresponding to each sentence text sample is obtained from the output of the text encoder to be trained. The multiple audio features of the multiple text-reading audio sentence samples and the average speaker reading features are input into the audio encoder to be trained to obtain the second text reading style prediction information output by the audio encoder to be trained, which corresponds to each of the text-reading audio sentence samples. Based on the similarity between each first text reading style prediction information and each second text reading style prediction information, the text encoder and the audio encoder to be trained are trained; when the similarity between the first text reading style prediction information and the second text reading style prediction information that have a corresponding relationship is greater than or equal to a first similarity threshold, and the similarity between the first text reading style prediction information and the second text reading style prediction information that do not have a corresponding relationship is less than a second similarity threshold, the trained text encoder is obtained as the text reading style generation model.
2. The method according to claim 1, characterized in that, The acquisition of multiple audio sentence samples and multiple sentence text samples includes: Acquire text-to-speech audio data and corresponding text data; the text-to-speech audio data and corresponding text data come from a text-to-speech audio publishing platform; Based on the text-to-speech audio data, obtain multiple text-to-speech audio sentence samples that meet the preset audio sentence duration conditions; Based on the multiple audio sentence samples for text reading and the corresponding text data, obtain the sentence text sample corresponding to each audio sentence sample for text reading.
3. The method according to claim 2, characterized in that, The step of obtaining multiple text-to-speech audio sentence samples that meet preset audio sentence duration conditions based on the text-to-speech audio data includes: The text-to-speech audio data is subjected to volume equalization processing to obtain volume equalization-processed text-to-speech audio data. Based on the text-to-speech audio data after volume equalization processing, obtain multiple text-to-speech audio sentence samples that meet the preset audio sentence duration conditions.
4. The method according to claim 1, characterized in that, The acquisition of multiple audio sentence samples and multiple sentence text samples includes: Acquire text-to-speech audio data and corresponding text data; the text-to-speech audio data and corresponding text data come from a text-to-speech audio publishing platform; Based on the corresponding text data, obtain multiple sentence text samples; Based on the multiple sentence text samples and the text reading audio data, obtain the multiple text reading audio sentence samples.
5. The method according to claim 4, characterized in that, The step of obtaining the multiple text-to-speech audio sentence samples based on the multiple sentence text samples and the text-to-speech audio data includes: The text-to-speech audio data is subjected to volume equalization processing to obtain volume equalization-processed text-to-speech audio data. Based on the multiple sentence text samples and the text reading audio data after volume equalization processing, the multiple text reading audio sentence samples are obtained.
6. The method according to any one of claims 2 to 5, characterized in that, The acquisition of text-to-speech audio data includes: Obtain raw text-to-speech audio data from the aforementioned text-to-speech audio publishing platform; Determine the language distribution information, speaker characteristic information, and accompaniment information of the original text reading audio data; If the original text reading audio data satisfies the preset language distribution conditions based on the language distribution information, satisfies the preset speaker conditions based on the speaker characteristic information, and satisfies the preset accompaniment conditions based on the accompaniment information, then the original text reading audio data is determined to be the text reading audio data.
7. The method according to claim 1, characterized in that, The step of obtaining the average speaker reading features of the multiple text-to-speech audio sentence samples includes: The average speaker reading features of the multiple text-reading audio sentence samples are obtained based on the average fundamental frequency and / or average speech rate of the multiple text-reading audio sentence samples; wherein the average fundamental frequency is obtained by averaging multiple fundamental frequency sequences of the multiple text-reading audio sentence samples, and the average speech rate is obtained by the total reading time corresponding to the multiple text-reading audio sentence samples and the total number of characters of the text corresponding to the multiple sentence text samples.
8. The method according to claim 1, characterized in that, The step of inputting the multiple sentence text samples into the text encoder to be trained includes: For each of the multiple sentence text samples, the text content in the sentence text sample is masked according to a first preset ratio to obtain multiple masked sentence text samples. Input the sentence text samples processed by the multiple masks into the text encoder to be trained; And / or, The step of inputting multiple audio features of the multiple text-reading audio sentence samples and the average speaker reading features into the audio encoder to be trained includes: For each of the multiple audio features, the feature content of the audio feature is masked according to a second preset ratio to obtain multiple masked audio features. The audio features processed by multiple masks from the multiple text-reading audio sentence samples, along with the average speaker reading features, are input into the audio encoder to be trained.
9. A method for generating text-to-speech style, characterized in that, The method includes: Get the text to be read aloud; The text to be read aloud is input into a trained text reading style generation model; the trained text reading style generation model is trained according to the method described in any one of claims 1 to 8; Obtain the text reading style information corresponding to the text to be read, output by the trained text reading style generation model.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 8 or claim 9.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 8 or claim 9.
Citation Information
Patent Citations
Speech synthesis method and device of text, electronic equipment and storage medium
CN112908292A
Controlling expressivity in end-to-end speech synthesis system
CN114175143A