An end-to-end Chinese speech cloning method with timbre and emotion transfer

Through end-to-end speech synthesis technology based on phrase speech training, timbre emotion encoder and vocoder are constructed, and the transfer of tone and emotion in Chinese voice is realized, solving the problem of unnatural and lack of realistic voice in the prior art. The generated speech is natural and realistic, and the characteristics of the speaker are embedded.

CN115359775BActive Publication Date: 2025-05-16SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210846358.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2025-05-16
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

The existing Chinese pronunciation synthesis technology is difficult to effectively deal with the complex problems of tone change, polyphonic characters and rhythm in Chinese pronunciation, and the pronunciation cloning technology has shortcomings in embedding the characteristics of the speaker, resulting in unnatural and lack of realistic feeling of synthesized pronunciation.

Method used

The end-to-end speech synthesis technology based on phrase speech training is adopted to realize the transfer of tone and emotion by constructing timbre and emotion. The method includes collecting short speech data of multi-speakers, pre-processing of data, constructing a Chinese speech cloning synthesis model, and generating speech with specific tones and emotions through WaveRNN and Griffin-Lim vocoders.

Benefits of technology

The requirement for the duration of a single speaker training voice in speech synthesis technology is reduced, and end-to-end tone and emotional transfer is realized. The generated voice is natural and realistic, and can effectively embed the speaker's characteristics, solving the problem of large demand for speech synthesis training corpus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359775B_ABST
    Figure CN115359775B_ABST
Patent Text Reader

Abstract

The present invention discloses an end-to-end Chinese speech cloning method for timbre and emotion transfer, and the steps are as follows: collect Chinese speech recorded by users as training data, and extract required speech features; train a speech cloning synthesis model, including three parts: a timbre and emotion encoder, a synthesizer, and a vocoder; use the trained speech cloning synthesis model to generate the speech of a designated speaker already in the speech cloning synthesis model according to the speech or text content input by the user; or quickly clone the timbre and emotion in the user's speech according to the short-term speech input by the user. The present invention realizes end-to-end speech synthesis and cloning, and through a multi-speaker model, synthesizes speech with different emotions and timbre with the same model and different speaker vector embedding. The present invention uses the speaker embedding vector generated by short speech, combined with a generative model trained with more corpus to perform speech cloning, and realizes speech cloning that can reflect the timbre and emotion of a specific speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer speech synthesis, and in particular to a speech synthesis method based on end-to-end timbre and emotion transfer of short speech training. Background Art

[0002] Speech synthesis is a key technology necessary for realizing human-computer voice communication and establishing a spoken language system with listening and speaking capabilities. Enabling computers to have human-like speaking abilities is an important competitive market in today's information industry. Speech synthesis can convert any text information into standard and fluent speech in real time. The text can be commonly used language characters or SSML markup languages, while the sound is a continuous analog signal. The synthesis process that the computer needs to perform is to simulate it through digital signals. In the process of speech synthesis, pronunciation, alignment, rhythm, intonation and other issues are the key to synthesizing speech, and the difficult problems in Chinese speech such as tone changes, polyphonic characters, and complex rhythms are the focus and difficulty of current Chinese speech synthesis technology.

[0003] As an extension of speech synthesis, speech cloning aims to achieve real-time cloning of the speaker's timbre and emotion, so that the speech synthesized by the system is no longer the timbre and emotion of the speaker data during model training, but is different from the person, with the timbre and emotion of the target speaker. At present, it is urgent to propose a Chinese speech cloning method, the goal of which is to embed speaker characteristics, namely speaker timbre and emotion characteristics, into the synthesized speech, so that the synthesized speech is natural and realistic, corresponding to the specified speaker. Summary of the invention

[0004] The purpose of the present invention is to reduce the requirement for the length of speech training for a single speaker in speech synthesis technology, and to propose an end-to-end speech synthesis technology based on short speech training and a speech synthesis method for timbre and emotion transfer, so as to realize speech synthesis (speech cloning) with timbre and emotion transfer through short speech.

[0005] The purpose of the present invention can be achieved by adopting the following technical solutions:

[0006] A Chinese voice cloning method for end-to-end timbre and emotion migration, the Chinese voice cloning method comprising the following steps:

[0007] S1. Collecting voice data: Collect multiple Chinese short sentence voice files from multiple speakers. Each speaker records multiple short sentence voices according to the given text, and creates corresponding text tags for each voice file. Each voice is no longer than 15 seconds, and the total voice duration is no less than 30 hours. The voice is recorded in a quiet environment.

[0008] S2, data preprocessing: processing the voice files collected in step S1, unifying the sampling rate, format, bit depth and number of channels of the voice files, obtaining the required audio files, and generating a JSON file containing a recording file tag, a corresponding voice text tag and a speaker tag;

[0009] S3. Constructing a Chinese speech cloning synthesis model: The Chinese speech cloning synthesis model includes a timbre emotion encoder, a synthesizer, and a vocoder;

[0010] S4. Construct a timbre emotion encoder: The timbre emotion encoder includes three layers of LSTM networks connected in sequence, calculates the frequency domain feature Mel spectrum of the audio file as the input of the timbre emotion encoder, and obtains a fixed-dimensional speaker embedding vector as the output of the timbre emotion encoder;

[0011] S5. Training synthesizer: The synthesizer is composed of an encoder and a decoder connected in sequence, wherein the encoder includes a preprocessing network composed of a fully connected layer, a word embedding module, three one-dimensional convolutional layers connected in sequence, and a bidirectional LSTM network, the JSON file is used as the input of the encoder, and the encoder hidden state is used as the output of the encoder; the decoder includes a preprocessing network, two layers of LSTM networks connected in sequence, a projection layer composed of a linear mapping layer, and a post-processing network, the encoder hidden state is concatenated with the speaker embedding vector output by the timbre emotion coding as the input of the decoder, and the Mel spectrum of the synthesized speech is obtained as the output of the decoder;

[0012] S6. Training vocoder: The vocoder is composed of a parallel WaveRNN vocoder and a Griffin-Lim vocoder. The Mel spectrum of the synthesized speech output by the decoder is used as the input of the vocoder, and the waveform prediction of the synthesized speech is used as the output of the vocoder.

[0013] S7, generating cloned speech: using the text input by the user or the speech input by the user through speech recognition to obtain the text, using different speaker embedding vectors according to the speaker specified by the user, and passing through the synthesizer and vocoder to obtain the output speech;

[0014] Or fast voice cloning: pre-process the user audio, input it into the timbre emotion encoder, obtain the speaker embedding vector, and save the speaker embedding vector for generating cloned voice.

[0015] Furthermore, the pre-processing process of step S2 is as follows:

[0016] S2.1. Perform speech processing on multiple short sentence voice files, and convert multiple short sentence recording files into audio files with an audio sampling rate of 16000Hz, an audio format of wav format, a bit depth of 16 bits, and a monophonic channel. Unifying the relevant parameters and formats of audio files can speed up the extraction and processing of audio file data by the Chinese speech cloning synthesis model, improve training efficiency, and achieve better results;

[0017] S2.2. Generate a JSON file containing tags, and concatenate the text tags, speakers, speaker IDs, and audio file tags obtained by speech processing to obtain one or more JSON format files, where the text tags refer to the Chinese text corresponding to the audio content, the speaker ID refers to the number tag for the speaker, and the audio file tag refers to the audio file name corresponding to the speaker and the speech content. The generated JSON file provides the data set required for training the Chinese speech cloning synthesis model, namely, the speech and the text content of the speech, and at the same time, it also corresponds the speech information to the speaker information one by one.

[0018] Furthermore, the working process of the timbre emotion encoder in step S4 is as follows:

[0019] S4.1. For a given short recording of multiple sentences, calculate the Mel spectrum according to the following formula:

[0020]

[0021] Among them, f is the frequency of multiple short speech sentences, and m is the Mel spectrum of multiple short speech sentences. Mel spectrum can largely retain the information required for human ears to understand the original speech. Therefore, using Mel spectrum as the input of the timbre emotion encoder can improve the accuracy of the timbre emotion encoder;

[0022] S4.2. Input the Mel spectrum of multiple short speech sentences into the timbre emotion encoder and output a fixed-dimensional speaker embedding vector. The process is as follows:

[0023] S4.2.1. Input the Mel spectrograms of multiple short speech sentences into three layers of LSTM networks connected in sequence, and map the output of each frame of the last layer of LSTM network to a 256-dimensional fixed-length vector, where each frame refers to a fixed time unit. The 3-layer LSTM network can effectively extract the speaker's timbre and emotional characteristics in the speech, while also avoiding the additional memory overhead caused by too many layers;

[0024] S4.2.2. Perform mean and normalization processing on the output of all time units obtained in S4.2.1 to obtain the final fixed-dimensional speaker embedding vector, where the speaker embedding vector is used to distinguish the speaker corresponding to the speech from other speakers, and the speaker embedding vector can preserve the speaker's timbre and emotion. The mean and normalization processing scales the data to the same range, reduces the amount of calculation, and can speed up the training efficiency of the speaker encoder.

[0025] Furthermore, the training process of the timbre emotion encoder in step S4 is as follows:

[0026] The speaker embedding vector output by the speaker's timbre emotion encoder in a certain training iteration is compared with the corresponding speaker sample. When the comparison result obtained based on the speaker embedding vector can distinguish the speaker, it means that the timbre emotion encoder has extracted features that can distinguish the speaker, and the timbre emotion encoder parameters are retained. Otherwise, iterative training continues. The iterative training method can ensure that the speaker embedding vector obtained is a feature that can represent the timbre and emotion of a certain speaker and can be distinguished from the timbre and emotion of other speakers.

[0027] Furthermore, the working process of the encoder of the synthesizer in step S5 is as follows:

[0028] S5.1.1. Obtain the input of the encoder: Translate the text tags in the JOSN file generated in step S2 into phoneme sequences, where Chinese is converted into corresponding pinyin, concatenate the obtained phoneme sequence with the JSON file into the JSON file, and use the obtained JSON file as the input of the encoder; converting Chinese text data into a phoneme sequence helps the encoder better retain the key information in the speech into the word vector, thereby improving the encoder's ability to extract information from the speech content.

[0029] S5.1.2, Generate word embedding vector: First, input the JSON obtained in step S5.1.1 into the preprocessing network for analysis and transformation, and output the preprocessed sequence. The preprocessing operation can further convert the input sequence into a format that allows the encoder to better extract voice information, which is conducive to the preprocessing network to adjust parameters by itself and extract information more effectively; then perform word embedding operation on the preprocessed sequence, and calculate the weight of each phoneme in the phoneme sequence corresponding to the remaining phonemes. This operation incorporates the correlation between each phoneme into the training to improve the accuracy of the results, and finally outputs a 512-dimensional word vector. The word vector contains the location information and content information of the speech;

[0030] S5.1.3, obtain the intermediate state: input the word vector obtained in step S5.1.2 into three one-dimensional convolutional layers connected in sequence for convolution operation, and perform BatchNorm and Dropout operations on the output after each convolutional layer, and obtain the intermediate state as output in the last convolutional layer; the BatchNorm operation is to keep the input of each layer of the neural network in the same distribution during the deep neural network training process, which can speed up the training speed, and the Dropout operation is to temporarily block the neural network unit with a certain probability to avoid overfitting problems. These two operations can better retain the speech information in the word vector;

[0031] S5.1.4, Obtaining the encoder hidden state: Input the intermediate state obtained in step S5.1.3 into a bidirectional LSTM network, and the output of the bidirectional LSTM network is used as the encoder hidden state. The bidirectional LSTM network can obtain the temporal correlation of the input data and further obtain complete speech information;

[0032] S5.1.5, timbre and emotion feature embedding: The speaker embedding vector output by the timbre emotion encoder is concatenated with the encoder hidden state obtained in S5.1.4 to obtain the final encoder hidden state. The concatenated encoder hidden state contains the information of the speech and the timbre and emotion information of the speaker.

[0033] Furthermore, the working process of the decoder of the synthesizer in step S5 is as follows:

[0034] S5.2.1. The decoder runs in a loop, and each loop is called a time step. In each time step, the attention mechanism is operated on the encoder hidden state obtained in S5.1.5 to measure the unit similarity in the encoder hidden state. The attention mechanism is a matrix composed of context weight vectors. The input data is scored in each dimension, and then the features are weighted according to the scores to highlight the impact of important features on downstream models or modules. Then, the encoder hidden state after the attention mechanism operation is normalized to obtain the context vector, which contains the information in the encoder hidden state.

[0035] S5.2.2, the preprocessing network in the decoder takes the context vector outputted in the previous time step as input, analyzes and processes it, and inputs the processed context vector into the two layers of LSTM networks connected in sequence in the decoder for processing. The output of the last layer of LSTM is used as a new context vector, and finally the new context vector is input into the projection layer to obtain the spectrogram frame and the end probability as output, where the spectrogram frame contains the predicted Mel spectrum. The context vector transmission process ensures the accuracy of the speech information extracted by the decoder;

[0036] S5.2.3. The spectrogram frame obtained in S5.2.2 is used as the input of the post-processing network in the decoder, where the post-processing network consists of three convolutional layers connected in sequence, and the output of the last convolutional layer is used as the Mel spectrum of the synthesized speech. The post-processing network can improve the quality of the Mel spectrum, which improves the guarantee for the subsequent use of the Mel spectrum to synthesize high-fidelity speech.

[0037] Furthermore, the working process of the WaveRNN vocoder in step S6 is as follows:

[0038] The WaveRNN vocoder consists of a single-layer RNN network and a double softmax layer connected in sequence. The Mel spectrum of the synthesized speech obtained in step S5 is used as input. After being processed by a single-layer RNN, the output is divided into two parts and input into the corresponding softmax layer respectively. The output of the two softmax layers is concatenated to obtain a predicted 16-bit audio sequence. The synthesis process of WaveRNN ensures the quality of the synthesized speech while taking into account the speed of the synthesized speech.

[0039] Furthermore, the working process of the Griffin-Lim vocoder in step S6 is as follows:

[0040] The Griffin-Lim vocoder uses iterative processing. First, a time domain graph is randomly initialized. The Mel spectrum of the synthesized speech obtained in step S5 and the initialized time domain graph are inversely short-time Fourier transformed to obtain a new time domain graph. Then, the new time domain graph is inversely short-time Fourier transformed to obtain a new time domain graph and spectrum graph. The new spectrum graph is discarded, and then the new time domain graph and the known spectrum graph are used to synthesize speech. The above operations are continuously iterated to obtain the optimal synthesized speech as output. The iterative method is used to reconstruct the phase information from the given time-frequency spectrum, reduce the mean square error between the time-frequency spectrum of the reconstructed signal and the given time-frequency spectrum, and improve the quality of the synthesized speech.

[0041] Furthermore, the process of generating cloned voice in step S7 is as follows:

[0042] S7.1.1. Obtaining an input sequence: The input sequence is Chinese text or Chinese speech input by a user, wherein the Chinese speech is converted into Chinese text using an existing speech recognition method;

[0043] S7.1.2, obtain speaker embedding vector: after the data collected by S1 is preprocessed by S2, the speaker embedding vectors corresponding to different speakers are extracted through the timbre emotion encoder trained in S3, and the corresponding speaker embedding vector is selected according to the speaker that the user needs to clone; the corresponding speaker embedding vector saves the timbre emotion information of the speaker;

[0044] S7.1.3, generate the Mel spectrum of the synthesized speech: the input sequence obtained in S7.1.1 and the speaker embedding vector obtained in S7.1.2 are used as the input of the synthesizer, and the output of the synthesizer is the Mel spectrum of the synthesized speech; the input speaker embedding vector can make the output synthesized speech have the timbre and emotional characteristics of the speaker;

[0045] S7.1.4. Input the Mel spectrum obtained in S7.1.3 into the vocoder, and output a synthesized cloned speech, wherein the synthesized cloned speech has the timbre and emotion of the speaker specified in S7.1.2.

[0046] Furthermore, the process of rapid voice cloning in step S7 is as follows:

[0047] S7.2.1. Obtain input voice: collect the user's Chinese voice;

[0048] S7.2.2, preprocessing: Use the preprocessing method in S2 to process the Chinese speech obtained in step S7.2.1 to obtain the corresponding JSON file and the preprocessed audio file, that is, convert them into an input format that can be received by the Chinese speech cloning synthesis model;

[0049] S7.2.3, input the audio file obtained in S7.2.2 into the timbre emotion encoder, the output of the timbre emotion encoder is the speaker embedding vector, save the speaker vector and the corresponding speaker identifier for subsequent synthesis of cloned speech. Subsequent cloning of a speaker's speech can directly call the corresponding speaker embedding vector saved.

[0050] Compared with the prior art, the present invention has the following advantages and effects:

[0051] (1) The present invention realizes end-to-end speech synthesis and cloning, and simplifies the synthesis process.

[0052] (2) The present invention uses a multi-speaker model to synthesize speech with different emotions and timbres using the same model and different speaker vector embedding.

[0053] (3) The present invention uses speaker embedding vectors generated by short speech and combines them with a generative model trained with a large amount of corpus to perform speech cloning, so that speech cloning that reflects emotions and the timbre of a specific speaker can be achieved using short speech, thereby solving the problem of a large amount of speech synthesis training corpus required. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0055] Figure 1 It is a processing flow chart of the Chinese speech cloning synthesis model disclosed in the present invention;

[0056] Figure 2 is a network structure diagram of an encoder in a synthesizer of the present invention;

[0057] Figure 3 is a network structure diagram of a decoder in a synthesizer of the present invention;

[0058] Figure 4 It is a processing flow chart of the Chinese speech cloning synthesis model without the timbre emotion encoder of the present invention. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the method in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] Example 1

[0061] This embodiment discloses a Chinese speech cloning method based on end-to-end timbre and emotion migration. Figure 1 is a processing flow chart of the Chinese speech cloning synthesis model disclosed in the present invention, Figure 2 is the encoder structure in the synthesizer in this embodiment, Figure 3 It is the decoder structure in the synthesizer in this embodiment. The user can input voice or text to determine the content of the synthesized voice, and select the corresponding speaker embedding vector according to the timbre and emotion of the speaker to be simulated to synthesize the target voice. In this embodiment, taking the computer as an example, the processing flow of generating cloned voice is specifically introduced, including the following steps:

[0062] Step 101, obtaining user input. The user can input Chinese text or Chinese voice; if the input is Chinese voice, it is first converted into Chinese text through the existing speech recognition method.

[0063] Step 102: Obtain a speaker embedding vector of the speaker to be imitated.

[0064] The steps of obtaining the speaker embedding vector in step 102 are as follows:

[0065] 1) Collect voice data: Collect the user’s Chinese voice and record it in a quiet environment.

[0066] 2) Preprocessing: The collected voice files are processed and converted into audio files with an audio sampling rate of 16000 Hz, an audio format of WAV format, a bit depth of 16 bits and a mono channel.

[0067] 3) Input the above audio file into the timbre emotion encoder, and the output of the timbre emotion encoder is the speaker embedding vector. Save the speaker vector and the corresponding speaker identifier for subsequent synthesis of cloned speech.

[0068] Step 103, generating a Mel spectrum of the synthesized speech: the user input obtained in step 101 and the speaker embedding vector obtained in step 102 are used as inputs of the synthesizer, and the output of the synthesizer is the Mel spectrum of the synthesized speech.

[0069] Step 103 of obtaining the synthesized Mel spectrum is as follows:

[0070] 1) Obtaining the input of the encoder in the synthesizer: translating the input sequence obtained in step 101 into a phoneme sequence, wherein Chinese characters are converted into corresponding pinyin, and using the obtained phoneme sequence as the input of the encoder;

[0071] 2) Generate word embedding vector: First, input the phoneme sequence obtained in step 1) into the preprocessing network for analysis and transformation, and output the preprocessed sequence; then perform word embedding operation on the preprocessed sequence, calculate the weight of each phoneme in the phoneme sequence corresponding to the other phonemes, and output a 512-dimensional word vector;

[0072] 3) Obtaining the intermediate state: Input the word vector obtained in step 2) into three sequentially connected one-dimensional convolutional layers for convolution operation, and perform BatchNorm and Dropout operations on the output after each convolutional layer, and obtain the intermediate state as the output in the last convolutional layer; the BatchNorm operation is to keep the input of each layer of the neural network in the same distribution during the deep neural network training process, and the Dropout operation is to temporarily block the neural network unit with a certain probability;

[0073] 4) Obtaining the encoder hidden state: Input the intermediate state obtained in step 3) into a bidirectional LSTM network, and the output of the bidirectional LSTM network is used as the encoder hidden state;

[0074] 5) Timbre and emotion feature embedding: The speaker embedding vector obtained in step 102 is concatenated with the encoder hidden state obtained in step 4) to obtain the final encoder hidden state.

[0075] 6) Input the encoder hidden state obtained in step 5) into the decoder, and the decoder runs in a loop, each loop is called a time step, and in each time step, the encoder hidden state obtained in step 5) is operated by the attention mechanism to measure the unit similarity in the encoder hidden state, where the attention mechanism is a matrix composed of context weight vectors, and then the encoder hidden state after the attention mechanism operation is normalized to obtain the context vector;

[0076] 7) The preprocessing network in the decoder takes the context vector outputted in the previous time step as input, analyzes and processes it, and inputs the processed context vector into the two layers of LSTM networks connected in sequence in the decoder for processing. The output of the last layer of LSTM is used as a new context vector, and finally the new context vector is input into the projection layer to obtain a spectrogram frame and an end probability as output, wherein the spectrogram frame contains the predicted Mel spectrum;

[0077] 8) The spectrogram frame obtained in step 7) is used as the input of the post-processing network in the decoder, where the post-processing network consists of three convolutional layers connected in sequence, and the output of the last convolutional layer is used as the Mel spectrum of the synthesized speech.

[0078] Step 104: input the Mel spectrum generated in step 103 into a vocoder to obtain the required speech.

[0079] Step 104 of obtaining the required voice is as follows:

[0080] In the network architecture of speech synthesis, WaveRNN vocoder and Griffin-Lim vocoder are used, both of which receive Mel spectrum as input to predict the waveform of speech synthesis. We input the Mel spectrum generated in step 103 into the two vocoders to output speech with a faster synthesis speed, and the synthesized speech has the timbre and emotional characteristics of the specified speaker.

[0081] In summary, this embodiment corresponds to the speaker embedding vector according to the text or speech input by the user and the speaker to be simulated. The content text is converted into a pinyin phoneme sequence, input into the encoder to obtain the encoder hidden state, and spliced ​​with the speaker embedding vector to obtain the encoder hidden state with timbre and emotion, and then input into the decoder to generate the Mel spectrum, and the Mel spectrum is input into the vocoder to obtain the target speech. The synthesized target speech has the timbre and emotion of the specified speaker. The present invention solves the problem of large demand for speech synthesis training corpus by using the speaker embedding vector generated by short speech and combining it with the generative model trained with more corpus for speech cloning.

[0082] Example 2

[0083] This embodiment removes the timbre emotion encoder in the Chinese speech cloning synthesis model, such as Figure 4 As shown, the effect of the timbre emotion encoder part in the present invention in retaining the timbre and emotion information of the speaker is demonstrated. In this embodiment, taking the computer end as an example, the processing flow of the speech synthesis method is specifically introduced, which includes the following steps:

[0084] Step 101 is consistent with that in Example 1, and reference is made to step 101 in Example 1.

[0085] Step 102, generating a Mel spectrum of the synthesized speech: the user input obtained in step 101 is used as an input of a synthesizer, and the output of the synthesizer is a Mel spectrum of the synthesized speech.

[0086] The steps of step 102 to obtain the synthesized Mel spectrum are as follows:

[0087] The step of obtaining the encoder hidden state refers to 1) to 4) of step 103 in Example 1. Since the timbre emotion encoder is removed, the generated encoder hidden state is not combined with the speaker embedding vector and is directly input into the decoder. The step of generating the synthesized Mel spectrum refers to 6) to 8) of step 103 in Example 1.

[0088] Step 103 synthesizes the required speech with reference to step 104 in Example 1. Since the timbre and emotion encoder is removed, the Chinese speech cloning synthesis model lacks the ability to extract the speaker's timbre and emotion characteristics in the speech. The synthesizer and vocoder structure can only complete the task of speech synthesis. The speech synthesized in this embodiment obviously does not match the speaker's timbre, lacks emotional characteristics, and is relatively stiff.

[0089] For further verification, two groups of speech with the same content were synthesized using the original Chinese speech cloning synthesis model and the Chinese speech cloning synthesis model without the timbre and emotion encoder, and the two groups were put on an online webpage together with the real speech for a questionnaire survey. The survey was conducted in the form of a blind test, that is, the evaluators did not know the source of the speech, and the most authoritative international standard MOS score for judging speech quality was used for scoring, with a full score of 5. The final score shows that the speech synthesized by the Chinese speech cloning synthesis model without the timbre and emotion encoder has a lower score, as shown in Table 1 below.

[0090] Table 1. Comparison of MOS scores of different speech sources

[0091] Voice Source MOS score Ground Truth 4.78±0.05 Original Chinese speech cloning synthesis model 4.62±0.05 Chinese speech cloning synthesis model without timbre emotion encoder 4.36±0.04

[0092] It can be concluded from this embodiment that the Chinese speech cloning synthesis of the present invention has the ability to transfer the timbre and emotional characteristics of the speaker, and can better improve the quality of the synthesized speech.

[0093] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A Chinese speech cloning method with end-to-end timbre and emotion migration, characterized in that: The Chinese voice cloning method comprises the following steps: S1. Collecting voice data: Collect multiple Chinese short sentence voice files from multiple speakers. Each speaker records multiple short sentence voices according to the given text, and creates corresponding text tags for each voice file. Each voice is no longer than 15 seconds, and the total voice duration is no less than 30 hours. The voice is recorded in a quiet environment. S2, data preprocessing: Process the voice files collected in step S1, unify the sampling rate, format, bit depth and number of channels of the voice files, obtain the required audio files, and generate a JSON file containing the recording file tag, the corresponding voice text tag and the speaker tag; the process is as follows: S2.

1. Perform voice processing on the multiple short sentence voice files, and convert the multiple short sentence recording files into audio files with an audio sampling rate of 16000 Hz, an audio format of WAV format, a bit depth of 16 bits, and a mono channel; S2.

2. Generate a JSON file containing tags. Concatenate the text tags, speakers, speaker IDs, and audio file tags obtained by speech processing to obtain one or more JSON format files. The text tags refer to the Chinese text corresponding to the audio content, the speaker ID refers to the number tag for the speaker, and the audio file tag refers to the audio file name corresponding to the speaker and the speech content. S3. Constructing a Chinese speech cloning synthesis model: The Chinese speech cloning synthesis model includes a timbre emotion encoder, a synthesizer, and a vocoder; S4. Construct a timbre emotion encoder: The timbre emotion encoder includes three layers of LSTM networks connected in sequence, calculates the frequency domain feature Mel spectrum of the audio file as the input of the timbre emotion encoder, and obtains a fixed-dimensional speaker embedding vector as the output of the timbre emotion encoder; S5. Training synthesizer: The synthesizer is composed of an encoder and a decoder connected in sequence, wherein the encoder includes a preprocessing network composed of a fully connected layer, a word embedding module, three one-dimensional convolutional layers connected in sequence, and a bidirectional LSTM network, the JSON file is used as the input of the encoder, and the encoder hidden state is used as the output of the encoder; the decoder includes a preprocessing network, two layers of LSTM networks connected in sequence, a projection layer composed of a linear mapping layer, and a post-processing network, the encoder hidden state is concatenated with the speaker embedding vector output by the timbre emotion coding as the input of the decoder, and the Mel spectrum of the synthesized speech is obtained as the output of the decoder; S6. Training vocoder: The vocoder is composed of a parallel WaveRNN vocoder and a Griffin-Lim vocoder. The Mel spectrum of the synthesized speech output by the decoder is used as the input of the vocoder, and the waveform prediction of the synthesized speech is used as the output of the vocoder. S7, generating cloned speech: using the text input by the user or the speech input by the user through speech recognition to obtain the text, using different speaker embedding vectors according to the speaker specified by the user, and passing through the synthesizer and vocoder to obtain the output speech; Or fast voice cloning: pre-process the user audio, input it into the timbre emotion encoder, obtain the speaker embedding vector, and save the speaker embedding vector for generating cloned voice.

2. The Chinese speech cloning method with end-to-end timbre and emotion migration according to claim 1, characterized in that: The working process of the timbre emotion encoder in step S4 is as follows: S4.

1. For a given short recording of multiple sentences, calculate the Mel spectrum according to the following formula: Among them, f is the frequency of multiple short speech sentences, and m is the Mel spectrum of multiple short speech sentences; S4.

2. Input the Mel spectrum of multiple short speech sentences into the timbre emotion encoder and output a fixed-dimensional speaker embedding vector. The process is as follows: S4.2.

1. Input the Mel-spectrograms of multiple short speech sentences into three layers of sequentially connected LSTM networks, and map the output of each frame of the last layer of LSTM networks to a 256-dimensional fixed-length vector, where each frame refers to a fixed time unit. S4.2.

2. Average and normalize the outputs of all time units obtained in S4.2.1 to obtain the final fixed-dimensional speaker embedding vector, where the speaker embedding vector is used to distinguish the speaker corresponding to the speech from other speakers, and the speaker embedding vector can preserve the speaker's timbre and emotion.

3. The Chinese speech cloning method with end-to-end timbre and emotion migration according to claim 1 is characterized in that: The training process of the timbre emotion encoder in step S4 is as follows: The speaker embedding vector output by the speaker's timbre emotion encoder in a certain training iteration is compared with the corresponding speaker sample. When the comparison result obtained based on the speaker embedding vector can distinguish the speaker, it means that the timbre emotion encoder has extracted features that can distinguish the speaker. The timbre emotion encoder parameters are retained, otherwise the iterative training continues.

4. The Chinese speech cloning method with end-to-end timbre and emotion migration according to claim 1, characterized in that: The working process of the encoder of the synthesizer in step S5 is as follows: S5.1.

1. Obtaining the input of the encoder: translating the text tags in the JSON file generated in step S2 into phoneme sequences, wherein Chinese characters are converted into corresponding pinyin, concatenating the obtained phoneme sequence with the JSON file into the JSON file, and using the obtained JSON file as the input of the encoder; S5.1.

2. Generate word embedding vector: First, input the JSON obtained in step S5.1.1 into the preprocessing network for analysis and transformation, and output the preprocessed sequence; then perform word embedding operation on the preprocessed sequence, calculate the weight of each phoneme in the phoneme sequence corresponding to the remaining phonemes, and output a 512-dimensional word vector; S5.1.3, obtain the intermediate state: input the word vector obtained in step S5.1.2 into three one-dimensional convolutional layers connected in sequence for convolution operation, and perform BatchNorm operation and Dropout operation on the output after each convolutional layer, and obtain the intermediate state as output in the last convolutional layer; the BatchNorm operation is to keep the input of each layer of the neural network in the same distribution during the deep neural network training process, and the Dropout operation is to temporarily block the neural network unit with a certain probability; S5.1.4, obtain the encoder hidden state: input the intermediate state obtained in step S5.1.3 into a bidirectional LSTM network, and the output of the bidirectional LSTM network is used as the encoder hidden state; S5.1.5, timbre and emotion feature embedding: The speaker embedding vector output by the timbre emotion encoder is concatenated with the encoder hidden state obtained in S5.1.4 to obtain the final encoder hidden state.

5. The method for Chinese speech cloning with end-to-end timbre and emotion migration according to claim 4, characterized in that: The working process of the decoder of the synthesizer in step S5 is as follows: S5.2.1, the decoder runs in a loop, each loop is called a time step, in each time step, the attention mechanism operation is performed on the encoder hidden state obtained in S5.1.5, and the unit similarity in the encoder hidden state is measured, where the attention mechanism is a matrix composed of context weight vectors, and then the encoder hidden state after the attention mechanism operation is normalized to obtain the context vector; S5.2.2, the preprocessing network in the decoder takes the context vector outputted in the previous time step as input, analyzes and processes it, and inputs the processed context vector into the two layers of LSTM networks connected in sequence in the decoder for processing, and the output of the last layer of LSTM is used as a new context vector, and finally the new context vector is inputted into the projection layer to obtain the spectrogram frame and the end probability as output, wherein the spectrogram frame contains the predicted Mel spectrum; S5.2.

3. Use the spectrogram frame obtained in S5.2.2 as the input of the post-processing network in the decoder, where the post-processing network consists of three convolutional layers connected in sequence, and the output of the last convolutional layer is used as the Mel spectrum of the synthesized speech.

6. The Chinese speech cloning method with end-to-end timbre and emotion migration according to claim 1, characterized in that: The working process of the WaveRNN vocoder in step S6 is as follows: The WaveRNN vocoder consists of a single-layer RNN network and a double softmax layer connected in sequence. The Mel spectrum of the synthesized speech obtained in step S5 is used as input. After being processed by a single-layer RNN, the output is divided into two parts and input into the corresponding softmax layer respectively. The outputs of the two softmax layers are concatenated to obtain a predicted 16-bit audio sequence.

7. The method for Chinese speech cloning with end-to-end timbre and emotion migration according to claim 1, characterized in that: The operation process of the Griffin-Lim vocoder in step S6 is as follows: The Griffin-Lim vocoder adopts iterative processing. First, a time domain graph is randomly initialized. The Mel spectrum of the synthesized speech obtained in step S5 and the initialized time domain graph are inversely short-time Fourier transformed to obtain a new time domain graph. Then, the new time domain graph is inversely short-time Fourier transformed to obtain a new time domain graph and spectrum graph. The new spectrum graph is discarded, and then the new time domain graph and the known spectrum graph are used to synthesize speech. The above operations are continuously iterated to obtain the optimal synthesized speech as output.

8. The Chinese speech cloning method with end-to-end timbre and emotion migration according to claim 1, characterized in that: The process of generating cloned voice in step S7 is as follows: S7.1.

1. Obtaining an input sequence: The input sequence is Chinese text or Chinese speech input by a user, wherein the Chinese speech is converted into Chinese text using an existing speech recognition method; S7.1.2, obtain speaker embedding vector: after the data collected by S1 is preprocessed by S2, the speaker embedding vectors corresponding to different speakers are extracted through the timbre emotion encoder trained in S3, and the corresponding speaker embedding vector is selected according to the speaker that the user needs to clone; S7.1.

3. Generate Mel spectrum of synthesized speech: The input sequence obtained in S7.1.1 and the speaker embedding vector obtained in S7.1.2 are used as input of the synthesizer, and the output of the synthesizer is the Mel spectrum of the synthesized speech; S7.1.

4. Input the Mel spectrum obtained in S7.1.3 into the vocoder, and output a synthesized cloned speech, wherein the synthesized cloned speech has the timbre and emotion of the speaker specified in S7.1.

2.

9. The method for Chinese speech cloning with end-to-end timbre and emotion migration according to claim 1, characterized in that: The process of rapid voice cloning in step S7 is as follows: S7.2.

1. Obtain input voice: collect the user's Chinese voice; S7.2.2, preprocessing: use the preprocessing method in S2 to process the Chinese speech obtained in S7.2.1 to obtain the corresponding JSON file and the preprocessed audio file; S7.2.

3. Input the audio file obtained in S7.2.2 into the timbre emotion encoder. The output of the timbre emotion encoder is the speaker embedding vector. The speaker vector and the corresponding speaker identifier are saved for subsequent synthesis of cloned speech.

Citation Information

Patent Citations

  • Speech synthesis method and system for new tone generation

    CN112802448A

  • End-to-end emotion speech synthesis method in business hall environment

    CN112951201A