Two-Level Text-to-Speech System Using Synthetic Training Data
The method uses a trained voice cloning system to generate synthetic speech representations for a TTS system, addressing the challenge of transferring accents and dialects, enabling accurate speech synthesis with diverse speaking styles.
Patent Information
- Application Number
- JP2024501888
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-07-14
- Filing Date
- 2022-07-01
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2042-07-01
AI Technical Summary
Existing speech synthesis systems face challenges in transferring features between speech models due to significant development costs, architectural constraints, and design limitations, making it difficult to generate synthetic speech with different accents or dialects without sufficient training data.
A computer-implemented method using a trained voice cloning system to generate synthetic speech representations of a target speaker in a different accent or dialect, which is then used to train a text-to-speech system to produce expressive synthetic speech that clones the target speaker's voice in the desired accent or dialect.
Enables the generation of synthetic speech with accurate accents and dialects by leveraging a trained voice cloning system to create training data for a TTS system, allowing it to convert input text into speech that mimics the target speaker's voice in a different accent or dialect.
Smart Images

Figure 0007713087000001 
Figure 0007713087000002 
Figure 0007713087000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a two-level text reading system using synthetic training data.
Background Art
[0002] Speech synthesis systems use speech models to generate synthetic audio from text and / or audio inputs and are becoming increasingly popular on mobile devices. There are various different speech models, each with its own efficiencies and features such as speaking style, rhythm, language, accent, etc. Depending on the scenario, it may be useful to implement one of these developed features in another speech model. However, the specific training data required to train a speech model may not be available. In other cases, it may be useful to transfer one or more of these features between speech models. However, here, it can be particularly difficult to transfer features between speech models because a particular speech model has significant development costs, architectural constraints, and / or design limitations.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Means for Solving the Problems
[0005] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations. The operations include obtaining training data including a plurality of training audio signals and corresponding transcripts. Each training audio signal corresponds to a reference utterance spoken by a target speaker in a first accent / dialect. Each transcript includes a textual representation of the corresponding reference utterance. The operations include, for each training audio signal of the training data, generating a training synthetic speech representation of the corresponding reference utterance spoken by the target speaker by a trained voice cloning system configured to receive, as input, the training audio signal corresponding to the reference utterance spoken by the target speaker in the first accent / dialect. The training synthetic speech representation includes the voice of the target speaker in a second accent / dialect different from the first accent / dialect. Here, for each training audio signal of the training data, the operations also include training a text-to-speech (TTS) system based on the corresponding transcript of the training audio signal and the training synthetic speech representation of the corresponding reference utterance generated by the trained voice cloning system. The operations also include receiving an input text utterance to be synthesized into speech in the second accent / dialect. The operations also include obtaining a conditional input including a speaker embedding representing the voice characteristics of the target speaker and an accent / dialect identifier identifying the second accent / dialect. The operations also include generating, by using the trained TTS system conditioned on the obtained conditional input and processing the input text utterance, an output audio waveform corresponding to a synthetic speech representation of the input text utterance that creates a clone of the voice of the target speaker in the second accent / dialect.
[0006] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the step of training the TTS system includes training the encoder portion of the TTS model of the TTS system to encode the corresponding reference utterance training synthesis speech representation generated by the trained voice cloning system into an utterance embedding representing the prosody captured by the training synthesis speech representation. In these implementations, the step of training the TTS system also includes training the decoder portion of the TTS system by decoding the utterance embedding using the corresponding transcript of the training audio signal to generate the predicted output audio signal of the expressive speech. In some examples, the step of training the TTS system further includes training the synthesizer of the TTS system using the predicted output audio signal to create a clone of the target speaker's voice in a second / accent dialect and generate the predicted synthesis speech representation of the input text utterance having the prosody represented by the utterance embedding, generating a gradient / loss between the predicted synthesis speech representation and the training synthesis speech representation, and backpropagating the gradient / loss through the TTS model and the synthesizer.
[0007] This operation may further include sampling a sequence of fixed-length reference frames that provide reference prosodic features representing the prosody captured by the training synthetic speech representation from the training synthetic speech representation. Here, the step of training the encoder portion of the TTS model includes training the encoder portion to encode a sequence of fixed-length reference frames sampled from the training synthetic speech representation for utterance embedding. In some implementations, the step of training the decoder portion of the TTS model includes decoding the utterance embedding into a sequence of fixed-length predicted frames that provide predicted prosodic features of the transcript representing the prosody represented by the utterance embedding, using the corresponding transcript of the training audio signal. Optionally, the TTS model may be trained such that the number of fixed-length predicted frames decoded by the decoder portion is equal to the number of fixed-length reference frames sampled from the training synthetic speech representation.
[0008] In some implementations, the training synthetic speech representation of the reference utterance includes a sequence of audio waveforms or mel-frequency spectrograms. The trained voice cloning system may be further configured to receive the corresponding transcript of the training audio signal as input when generating the training synthetic speech representation. In some examples, the training audio signal corresponding to the reference utterance spoken by the target speaker includes the input audio waveform of human speech, the training synthetic speech representation includes the output audio waveform of synthetic speech creating a clone of the target speaker's voice in a second accent / dialect, and the trained voice cloning system includes an end-to-end neural network configured to directly convert the input audio waveform into the corresponding output audio waveform.
[0009] In some implementations, the TTS system includes a TTS model configured to generate an output audio signal of expressive speech by decoding an utterance embedding into a sequence of fixed-length predicted frames that are conditioned by the conditional input and provide prosodic features using the input text utterance. The utterance embedding is selected to specify the prosody intended for the input text utterance, and the prosodic features represent the intended prosody specified by the utterance embedding. In these implementations, the TTS system also includes a waveform synthesizer configured to receive as input a sequence of fixed-length predicted frames and generate as output an output audio waveform corresponding to a synthetic speech representation of the input text utterance that clones the voice of the target speaker in a second accent / dialect. The prosodic features representing the intended prosody can include duration, pitch contour, energy contour, and / or mel-frequency spectrogram contour.
[0010] Another aspect of the present disclosure provides a system that includes data processing hardware and memory hardware that stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include obtaining training data that includes a plurality of training audio signals and corresponding transcripts. Each training audio signal corresponds to a reference utterance spoken by a target speaker in a first accent / dialect. Each transcript includes a text representation of the corresponding reference utterance. The operations include, for each training audio signal of the training data, generating, by a trained voice cloning system configured to receive, as an input, the training audio signal corresponding to the reference utterance spoken by the target speaker in the first accent / dialect, a training synthetic speech representation of the corresponding reference utterance spoken by the target speaker. The training synthetic speech representation includes the voice of the target speaker in a second accent / dialect that is different from the first accent / dialect. Here, for each training audio signal of the training data, the operations also include training a text-to-speech (TTS) system based on the corresponding transcript of the training audio signal and the training synthetic speech representation of the corresponding reference utterance generated by the trained voice cloning system. The operations also include receiving an input text utterance to be synthesized into speech in the second accent / dialect. The operations also include obtaining a conditioned input that includes a speaker embedding representing the voice characteristics of the target speaker and an accent / dialect identifier that identifies the second accent / dialect. The operations also include generating, using the trained TTS system conditioned on the obtained conditioned input and by processing the input text utterance, an output audio waveform corresponding to a synthetic speech representation of the input text utterance that creates a clone of the voice of the target speaker in the second accent / dialect.
[0011] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, training the TTS system involves training the encoder portion of the TTS model of the TTS system to encode the corresponding reference utterance training synthesis speech representation generated by the trained voice cloning system into an utterance embedding that represents the prosody captured by the training synthesis speech representation. In these implementations, training the TTS system also involves training the decoder portion of the TTS system by decoding the utterance embedding using the corresponding transcript of the training audio signal to generate a predicted output audio signal of expressive speech. In some examples, training the TTS system further includes creating a clone of the target speaker's voice in a second / accent dialect and generating a predicted synthesis speech representation of the input text utterance having the prosody represented by the utterance embedding, training the synthesizer of the TTS system using the predicted output audio signal, generating a gradient / loss between the predicted synthesis speech representation and the training synthesis speech representation, and backpropagating the gradient / loss through the TTS model and the synthesizer.
[0012] This operation may further include sampling a sequence of fixed-length reference frames that provide reference prosodic features representing the prosody captured by the training synthetic speech representation from the training synthetic speech representation. Here, training the encoder portion of the TTS model includes training the encoder portion to encode a sequence of fixed-length reference frames sampled from the training synthetic speech representation into the utterance embedding. In some implementations, training the decoder portion of the TTS model includes decoding the utterance embedding into a sequence of fixed-length predicted frames that provide predicted prosodic features of the transcript representing the prosody represented by the utterance embedding, using the corresponding transcript of the training audio signal. Optionally, the TTS model may be trained such that the number of fixed-length predicted frames decoded by the decoder portion is equal to the number of fixed-length reference frames sampled from the training synthetic speech representation.
[0013] In some implementations, the training synthetic speech representation of the reference utterance includes a sequence of audio waveforms or mel-frequency spectrograms. The trained voice cloning system may be further configured to receive the corresponding transcript of the training audio signal as an input when generating the training synthetic speech representation. In some examples, the training audio signal corresponding to the reference utterance spoken by the target speaker includes the input audio waveform of human speech, the training synthetic speech representation includes the output audio waveform of synthetic speech that creates a clone of the target speaker's voice in a second accent / dialect, and the trained voice cloning system includes an end-to-end neural network configured to directly convert the input audio waveform into the corresponding output audio waveform.
[0014] In some implementations, the TTS system is conditioned on conditional input and configured to generate an output audio signal of expressive speech by decoding an utterance embedding into a sequence of fixed-length predicted frames that provide prosodic features using the input text utterance. The utterance embedding is selected to specify the prosody intended for the input text utterance, and the prosodic features represent the intended prosody specified by the utterance embedding. In these implementations, the TTS system also includes a waveform synthesizer configured to receive as input a sequence of fixed-length predicted frames and generate as output an output audio waveform corresponding to a synthetic speech representation of the input text utterance that clones the voice of the target speaker in a second accent / dialect. The prosodic features representing the intended prosody may include duration, pitch contour, energy contour, and / or mel-frequency spectrogram contour.
[0015] Another aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations including obtaining training data including a plurality of text utterances. For each training text utterance of the training data, the operations also include generating a training synthetic speech representation of the corresponding training text utterance by a trained voice cloning system configured to receive the training text utterance as an input, and training a text-to-speech (TTS) system to learn a method of generating synthetic speech having target speech characteristics based on the corresponding training text utterance and the training synthetic speech representation generated by the trained voice cloning system. The training synthetic speech representation is within the voice of the target speaker and has the target speech characteristics. The operations also include receiving an input text utterance to be synthesized into speech having the target speech characteristics, and generating a synthetic speech representation of the input text utterance using the trained TTS system, wherein the synthetic speech representation has the target speech characteristics.
[0016] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the operations further include obtaining a conditional input including a speaker identifier indicative of the voice characteristics of the target speaker. Here, when generating the synthetic speech representation of the input text utterance, the trained TTS system is conditioned by the obtained conditional input, and the synthetic speech representation having the target speech characteristics creates a clone of the voice of the target speaker. The target speech characteristics may include a target accent / dialect or a target prosody / style. In some examples, when generating the training synthetic speech representation of the corresponding training text utterance, the trained voice cloning system is further configured to receive a speaker identifier indicative of the voice characteristics of the target speaker.
[0017] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims.
Brief Description of the Drawings
[0018]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0019] Like reference numerals in the various drawings indicate like elements.
[0020] A text-to-speech (TTS) system, which is often used by a speech synthesis system, usually only receives text input without a reference acoustic representation during execution, and in order to generate synthetic speech that sounds real, it is necessary to substitute many language elements not provided by the text input. A subset of these language elements is collectively called prosody and can include intonation (changes in pitch), stress (stressed and unstressed syllables), length of sounds, volume, tone, rhythm, and speaking style. Prosody can indicate the emotional state of speech, the form of speech (e.g., statement, question, command, etc.), the presence of irony or sarcasm in speech, the uncertainty of speech knowledge, or other language elements that cannot be encoded by the grammar or vocabulary choice of the input text. Language elements can also convey an accent / dialect associated with the way a speaker in a particular geographical region pronounces words / terms in a given language. For example, an English speaker in Boston, Massachusetts, has a "Boston accent" and pronounces words / terms differently from the way an English speaker in Fargo, North Dakota, pronounces the same terms. Thus, a given text input can generate synthetic speech in a given language that spans various different accents / dialects and / or different speaking styles, and can also generate synthetic speech that spans different languages.
[0021] In some cases, the TTS system is trained using human speech spoken by one or more target speakers. For example, each target speaker may be a professional voice actor who speaks in a particular style and a particular accent / dialect that is native to the target speaker (e.g., an American English accent). Using a corpus of training utterances spoken by the target speaker (e.g., a professional advertisement reader), the TTS system can learn how to generate synthetic speech that matches the voice, speaking style, and accent / dialect associated with the target speaker. In some situations, it may be useful for the TTS system to create a clone of the target speaker's voice but generate synthetic speech with a speaking style and / or accent / dialect that is different from the speaking style and / or accent / dialect that is native to the target speaker. Returning to the example where the target speaker is a professional voice actor who speaks with an American English accent, there may be cases where it is desirable for the TTS system to generate synthetic speech that includes the voice of the voice actor (e.g., the target speaker) but has a British English accent. Here, unless the TTS system is trained based on reference utterances spoken by the target speaker with a British English accent, it is not possible to generate synthetic speech that creates a clone of the target speaker's voice with a British English accent. Further, a professional voice actor who natively speaks with an American English accent may not be able to generate speech that accurately pronounces / enunciates terms associated with a British English accent, and it may not even be possible to train the TTS system based on reference utterances spoken by the voice actor with a British English accent. The situation of not being able to obtain sufficient training data is further exacerbated in situations where it is desirable for the TTS system to generate synthetic speech that creates a clone of the target speaker's voice across multiple different accents / dialects that are not native to the voice actor.
[0022] The implementations of this specification are directed to using a trained voice cloning system to generate a training synthetic speech representation for creating a clone of a target speaker's voice in a target accent / dialect that the target speaker does not speak natively, and using the training synthetic speech representation to train a TTS system to learn a method of generating expressive synthetic speech that creates a clone of the target speaker's voice in the target accent / dialect. More specifically, the trained voice cloning system obtains training data including a plurality of training audio signals and corresponding transcripts, where each training audio signal corresponds to a reference utterance spoken by the target speaker in a first accent / dialect that is native to the target speaker. For each training audio signal, the trained voice cloning system generates a training synthetic speech representation of the corresponding reference utterance spoken by the target speaker. Here, the training synthetic speech representation includes the voice of the target speaker in a second accent / dialect that is different from the first accent / dialect. That is, the training synthetic speech representation is associated with an accent / dialect that is different from the first accent / dialect associated with the reference utterance spoken by the target speaker.
[0023] An untrained TTS system is trained based on transcripts of training audio signals and training synthetic speech representations to learn a method for generating synthetic speech that creates a clone of the target speaker's voice in a second accent / dialect. That is, in an untrained state, the TTS system cannot transfer the voice of the target speaker that spans different accents / dialects in the synthetic speech generated from the input text. However, after using a voice cloning system to generate training synthetic speech representations that create a clone of the target speaker's voice in different accents / dialects and using the training synthetic speech representations to train the TTS system, the trained TTS system can be used during inference to convert an input text utterance into a corresponding expressive synthetic speech that creates a clone of the target speaker's voice in a second accent / dialect. Here, during inference, the trained TTS system receives a conditional input that includes a speaker embedding representing the voice characteristics of the target speaker and an accent / dialect identifier that identifies the second accent / dialect, such that the TTS system can convert the input text utterance into an output audio waveform that creates a clone of the target speaker's voice in the second accent / dialect.
[0024] FIG. 1 shows an exemplary system 100 for training an untrained text-to-speech (TTS) system 300 and running the trained TTS system 300 to synthesize an input text utterance 320 into expressive speech 152 that includes the voice of a target speaker with a target accent / dialect. Examples herein are directed to generating synthetic speech 152 with a particular voice for different accents / dialects, but implementations herein can be equally applied to generate synthetic speech 152 with a particular voice for different ways of speaking, in addition to, or instead of, different accents / dialects. System 100 includes a computing system (interchangeably referred to as a “computing device”) 120 having data processing hardware 122 and memory hardware 124 that communicates with the data processing hardware 122 and stores instructions executable by the data processing hardware 122 to cause operations to be performed by the data processing hardware 122.
[0025] In some implementations, the computing system 120 (e.g., data processing hardware 122) provides a trained voice cloning system 200 configured to generate a training synthetic speech representation 202 for use in training an untrained TTS system 300. The trained voice cloning system 200 obtains training data 10 that includes a plurality of training audio signals 102 and corresponding transcripts 106. Each training audio signal 102 includes an utterance of human speech spoken by a target speaker in a first accent / dialect. For example, the training audio signal 102 may be spoken with an American English accent by the target speaker. Thus, the first accent / dialect associated with the utterance of human speech spoken by the target speaker may correspond to the native accent / dialect of the target speaker. Each transcript 106 includes a text representation for the corresponding reference utterance. The training data 10 may also include a plurality of speaker embeddings (also referred to as "speaker identifiers") 108, each representing the speaker characteristics (e.g., native accent, speaker identifier, male / female, etc.) of the corresponding target speaker. That is, the speaker embedding / identifier 108 may represent the speaker characteristics of the target speaker. The speaker embedding / identifier 108 may include a numerical vector representing the speaker characteristics of the target speaker, or may include an identifier associated with the target speaker that simply instructs the voice cloning system 200 trained to generate the training synthetic speech representation 202 with the voice of the target speaker. In the latter case, the speaker identifier may be converted into a corresponding speaker embedding used by the system 200. In some examples, the trained voice cloning system 200 includes a voice conversion system that directly converts each training audio signal 102 (e.g., a reference utterance of human speech) into a corresponding training synthetic speech representation 202.In another example, the trained voice cloning system 200 includes a text-to-speech voice cloning system that converts a corresponding transcript 106 into a corresponding training synthetic speech representation 106 that creates a clone of the reference utterance voice in a second accent / dialect different from the first accent / dialect associated with the training audio signal 102.
[0026] For simplicity, the examples in this specification are directed to a trained voice cloning system 200 that generates a training synthetic speech representation 202 that creates a clone of the target speaker's voice in a target accent / dialect (e.g., the second accent / dialect). However, the implementations herein are equally applicable to a trained voice cloning system 200 that creates a clone of the target speaker's voice and generates a training synthetic speech representation 202 having any target speech characteristics. Thus, the target speech characteristics can include at least one of a target accent / dialect, a target prosody / style, or some other speech characteristic. As will become apparent, the training synthetic speech representation 202 having the target speech characteristics generated by the trained voice cloning system is used to train an untrained TTS system 300 to learn a method for generating synthetic speech 202 having the target speech characteristics.
[0027] For each training audio signal 102 of the training data 10, the trained voice clone creation system 200 generates a corresponding training synthetic speech representation 202 of a reference utterance spoken by the target speaker. Here, the training synthetic speech representation 202 includes the voice of the target speaker with a second accent / dialect different from the first accent / dialect of the training audio signal 102. That is, the trained voice clone creation system 200 receives, as input, the training audio signal 102 corresponding to the reference utterance spoken by the target speaker in the first accent / dialect, and generates, as output, the training synthetic speech representation 202 of the training audio signal 102 in the second accent / dialect. Thus, the trained voice clone creation system 200 generates a corresponding training synthetic speech representation 202 for each of the plurality of training audio signals 102 of the training data 10 to create a plurality of training synthetic speech representations 202 for use when training the untrained TTS system 300. In some examples, the trained voice clone creation system 200 determines the speaker characteristics of the training synthetic speech representation 202 from the speaker embedding / identifier 108.
[0028] In some implementations, when the trained voice clone creation system 200 includes a TTS voice clone creation system 200, the training data 10 includes a plurality of training text utterances 106, and the TTS voice clone creation system 200 converts each training text utterance 106 into a training synthetic speech representation 202 in the target speech characteristics. The target speech characteristics may include a second accent / dialect. Alternatively, the target speech characteristics may include a target prosody / style. That is, the TTS voice clone creation system 200 may generate a training synthetic speech representation 202 from text only. Thus, the training text utterance 106 may correspond to an unspoken text utterance that is not paired with the corresponding audio signal of human speech. Therefore, the unspoken text utterance can be derived manually or from a language model. The TTS voice clone creation system 200 may also receive a speaker embedding / identifier 108 that conditions the TTS voice clone creation system 200 to create a clone of the target speaker's voice to generate a training synthetic speech representation 202 having the target speech characteristics. The TTS voice clone creation system 200 may also receive a target speech characteristic identifier that identifies the target speech characteristics. For example, the target speech characteristic identifier may include an accent / dialect identifier 109 that identifies the target accent / dialect (e.g., a second accent / dialect) of the resulting training synthetic speech representation 202, and / or a prosody / style identifier (i.e., an utterance embedding 204) that indicates the target prosody / style of the resulting training synthetic speech representation 202.
[0029] For each training audio signal 102 of the training data 10, the untrained TTS system 300 is trained based on the corresponding transcript 106 of the training audio signal 102 and the corresponding training synthetic speech representation 202 output from the trained voice cloning system 200 that includes the voice of the target speaker in a second dialect / language. More specifically, training the untrained TTS system 300 may include training both the TTS model 400 and the synthesizer 150 of the untrained TTS system 300 to learn a method of generating synthetic speech from the input text such that the synthetic speech creates a clone of the voice of the target speaker in a second dialect / accent. That is, the TTS system 300 including the TTS model 400 and the synthesizer 150 is trained to generate synthetic speech 152 that matches each training synthetic speech representation 202. During training, the TTS system 300 may learn to predict an utterance embedding 204 for the training synthetic speech representation 202. Here, each utterance embedding 204 may represent prosody information and / or accent / dialect information associated with the training synthetic speech representation 202 that the TTS system 300 aims to replicate. Further, a plurality of TTS systems 300, 300A - N may be trained based on the training synthetic speech representations 202 output from the trained voice cloning system 200. Here, each TTS system 300 is trained based on a corresponding set of training synthetic speech representations 202 that may include the voices of different target speakers, different ways of speaking / prosody, and / or different accents / dialects. Thereafter, each of the plurality of trained TTS systems 300 is configured to generate expressive speech 152 for each respective target voice of the corresponding accent / dialect.The computing device 120 may store each trained TTS system 300 in the data storage 180 (e.g., the memory hardware 124) for later use during inference.
[0030] During inference, the computing device 120 may use the trained TTS system 300 to synthesize the input text utterance 320 into expressive speech 152 that creates a clone of the target speaker's voice in the target accent / dialect (or conveys some other target speech characteristic in addition to or instead of the target accent / dialect). In particular, the TTS model 400 of the trained TTS system 300 may receive a conditional input that includes a speaker embedding / identifier 108 that represents the voice characteristics of the target speaker and an accent / dialect identifier 109 that identifies the intended accent / dialect (e.g., British English or American English). The conditional input may further include a speech prosody / style identifier that represents the particular speech vertical that the resulting synthetic speech 152 should include. The TTS model 400 conditioned on the speaker embedding / identifier 108 and the accent / dialect identifier 109 processes the input text utterance 320 to generate an output audio waveform 402. Here, the speaker embedding / identifier 108 includes the speaker characteristics of the target speaker, and the accent / dialect identifier 109 includes the target accent / dialect (e.g., American English, British English, etc.). The output audio waveform 402 conveys the target accent / dialect and the voice characteristics of the target speaker, enabling the speech synthesizer 150 to generate synthetic speech 152 from the output audio waveform 402. The TTS model 400 may also generate some predicted frames 280 corresponding to the output audio waveform 402.
[0031] FIG. 2A shows an example of a trained voice clone creation system 200, 200a of the system 100. The trained voice clone creation system 200a receives a training audio signal 102 corresponding to a reference utterance spoken by a target speaker in a first accent / dialect and a corresponding transcription 106 of the reference utterance, and generates a training synthetic speech representation 202 that creates a clone of the target speaker's voice in a second accent / dialect different from the first accent / dialect. The trained voice clone creation system 200a includes an inference network 210, a synthesizer 220, and an adversarial loss module 230. The inference network 210 includes a residual encoder 212 configured to consume an input training audio signal 102 corresponding to a reference utterance spoken by a target speaker in a first accent / dialect, and outputs a residual encoding 214 of the training audio signal 102. The training audio signal 102 may include a mel spectrogram representation. In some examples, a feature representation (i.e., a mel spectrogram sequence) is extracted from the training audio signal 102 and provided as input to the residual encoder 212 to generate a corresponding residual encoding 214 therefrom.
[0032] The synthesizer 220 includes a text encoder 222, a speaker embedding / identifier 108, a language embedding 224, a decoder neural network 500, and a waveform synthesizer 228. The text encoder 222 may include an encoder neural network having a convolutional subnetwork and a bidirectional long short-term memory (LSTM) layer. The decoder neural network 500 is configured to receive, as input, the outputs 225 from the text encoder 222, the speaker embedding / identifier 108, and the language embedding 224 in order to generate an output mel spectrogram 502. The speaker embedding / identifier 108 may represent the voice characteristics of the target speaker, and the language embedding 224 may specify language information associated with at least one of the language of the training audio signal, the language of the generated training synthetic speech utterance 204, and the accent / dialect identifier 109 that identifies the accent / dialect associated with the training audio signal 102 and the training synthetic speech representation. Finally, the waveform synthesizer 228 may convert the mel spectrogram 502 output from the decoder neural network 500 into a time-domain waveform (e.g., the training synthetic speech representation 202). The training synthetic speech representation 202 includes the voice of the target speaker in a second accent / dialect different from the first accent / dialect spoken in the reference utterance of the training data by the same target speaker. Thus, the voice cloning system 200a retains the voice of the target speaker who spoke the reference utterance in the first accent / dialect and outputs a training synthetic speech representation 202 that converts the first accent / dialect spoken in the reference utterance to the second / accent dialect. Each training synthetic speech representation 202 generated by the voice cloning system 200a may also be associated with the language embedding 224, the accent / dialect identifier 109, and / or the speaker embedding / identifier 108 for use as a conditioning input when training the TTS system 300 based on the training synthetic speech representation 202. In some implementations, the waveform synthesizer 228 is a Griffin-Lim synthesizer. In some other implementations, the waveform synthesizer 228 is a Vocoder.For example, the waveform synthesizer 228 may include a WaveRNN vocoder. Here, the WaveRNN vocoder may generate a 16-bit signal sampled at 24 kHz, conditioned on the spectrogram predicted by the trained voice cloning system 200. In some other implementations, the waveform synthesizer 228 is an inverter from a trainable spectrogram to a waveform. After the waveform synthesizer 125 generates a waveform, the audio output system can use the waveform to generate a training synthetic speech representation 202. In some examples, a WaveNet neural vocoder replaces the waveform synthesizer 228. The WaveNet neural vocoder may provide a different audio fidelity of the training synthetic speech representation 202 compared to the training synthetic speech representation 202 generated by the waveform synthesizer 228.
[0033] The text encoder 222 is configured to encode the corresponding transcription 106 of the training audio signal 102 into a sequence of text encodings 225, 225a - n. In some implementations, the text encoder includes an attention network configured to receive sequential feature representations of the transcription 106 in order to generate, for each output step of the decoder neural network 500, the corresponding text encoding as a fixed-length context vector. That is, the attention network in the text encoder 222 may generate fixed-length context vectors 225, 225a - n for each frame of the mel spectrogram 502 that the decoder neural network 500 will generate later. A frame is a unit of the mel spectrogram 502 based on a small portion of the input signal, e.g., a 10-millisecond sample of the input signal. The attention network may determine weights for each element of the text encoder 222 output and generate the fixed-length vector 225 by determining the weighted sum of each element. The attention weights may vary for each time step of the decoder neural network 500.
[0034] Accordingly, decoder neural network 500 is configured to receive a fixed-length vector (e.g., text encoding) 225 as input and generate, as output, corresponding frames of mel-frequency spectrogram 502. Mel-frequency spectrogram 502 is a frequency-domain representation of sound. The mel-frequency spectrogram emphasizes low frequencies that are important for speech intelligibility while emphasizing high frequencies that are dominated by fricatives and other noise bursts and generally need not be modeled with high fidelity.
[0035] In some implementations, decoder neural network 500 includes an attention-based sequence-to-sequence model configured to generate a sequence of output log mel spectrogram frames, e.g., output mel spectrogram 502, based on transcription 106. For example, decoder neural network 500 may be based on the Tacotron 2 model (see, e.g., “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions” by J. Shen et al. in https: / / arxiv.org / abs / 1712.05884, which is incorporated herein by reference). The trained voice cloning system 200a provides an enhanced multi-lingual trained voice cloning system that enhances decoder neural network 500 using additional speaker input (e.g., speaker embedding / identifier 108), and optionally, language embedding 224, an adversarially trained speaker classifier (e.g., speaker classifier 234), and a variational autoencoder style residual encoder (e.g., residual encoder 212).
[0036] An enhanced, trained voice cloning system 200a that enhances an attention-based sequence-to-sequence decoder neural network 500 using one or more of a speaker classifier 234, a residual encoder 212, a speaker embedding / identifier 108, and / or a language embedding 224 yields particularly many positive results. That is, the trained voice cloning system 200a enables the use of a phoneme input representation of a transcription 106 to facilitate sharing of model capabilities across different natural languages and different accents / dialects, and incorporates an adversarial loss term 233 to facilitate the trained voice cloning system 200a in unraveling from the speech content how the identity of a speaker that fully correlates with the language used in the training data 10 is represented.
[0037] Figure 2B shows an exemplary trained voice cloning system 200, 200b configured to convert an input training audio signal 102 corresponding to a reference utterance spoken by a target speaker in a first accent / dialect into an output mel spectrogram 502 representing the voice of the target speaker in a second accent / dialect. That is, the trained voice cloning system 200b includes a speech-to-speech (S2S) conversion model. The training voice cloning system 200b is contrasted with the training voice cloning system 200a (Figure 2A) that uses a corresponding transcription 106 as input to generate the output mel spectrogram 502. The S2S conversion model 200b is configured to directly convert the training audio signal 102 into the output mel spectrogram 502 without performing speech recognition or requiring the generation of an intermediate discrete representation (e.g., text or phonemes) from the training audio signal 102. The S2S conversion model 200b includes a spectrogram encoder 240 configured to encode the training audio signal 102 into a hidden feature representation (e.g., a series of vectors) and a spectrogram decoder 500 configured to decode the hidden representation into the output mel spectrogram 502. For example, when the spectrogram decoder 500 receives the input training audio signal 102 corresponding to the reference utterance, the spectrogram decoder 500 may process the frames provided by the audio and convert five frames of that audio into ten vectors. The vectors are a mathematical representation of the frames of the training audio signal 102 rather than a transcription of the frames of the training audio signal 102. Next, the spectrogram decoder 500 may generate the output mel spectrogram 502 corresponding to the training synthetic speech representation based on the vectors received from the spectrogram encoder 240. For example, the spectrogram decoder 500 may receive ten vectors representing five frames of audio from the spectrogram encoder 240.Here, the spectrogram decoder 500 can generate five frames of an output mel spectrogram 502 corresponding to the speech representation of a reference utterance containing the intended word or part of a word, as five frames of the training audio signal 102 in the second / accent dialect.
[0038] In some examples, the S2S conversion model 200b also includes a text decoder (not shown) that decodes the hidden representation into a text representation, such as phonemes or graphemes. In these examples, the spectrogram decoder 500 and the text decoder can correspond to parallel decoding branches of a trained voice cloning system 200 that each receive the hidden representation encoded by the spectrogram encoder 240 and output the output mel spectrogram 502 or the text representation in parallel. Similar to the TTS-based voice cloning system 200a of FIG. 2A, the S2S conversion system 200b can further include a waveform synthesizer 228 or a vocoder for synthesizing the output mel spectrogram 502 into a time-domain waveform for audible output. The time-domain audio waveform includes an audio waveform that defines the amplitude of the audio signal over time. The waveform synthesizer 228 can include a unit selection module or a WaveNet module for synthesizing the output mel spectrogram 502 into the time-domain waveform of the training synthetic speech representation 202. In some implementations, the vocoder 228, i.e., the neural vocoder, is separately trained and conditioned based on the mel-frequency spectrogram for conversion to the time-domain audio waveform (e.g., of the training synthetic speech representation 202).
[0039] In the illustrated example, the target speaker associated with the training data 10 speaks with a first accent / dialect (e.g., an American English accent). The trained voice cloning system (e.g., an S2S voice conversion model) 200b is trained to directly convert the training audio signal 102 of the training data 10 spoken with the first accent / dialect into a training synthetic speech representation 202 that includes the voice of a target speaker with a second accent / dialect (e.g., a British English accent). Without departing from the scope of the present disclosure, the trained voice cloning system 200b can be trained to convert the training audio signal 102 corresponding to a reference utterance spoken by the target speaker in a first language or way of speaking into a training synthetic speech representation 202 that retains the voice of the target speaker but is a different second language or way of speaking.
[0040] FIG. 3 shows an exemplary training process 301 for training a TTS system 300 based on a trained speech clone creation system 200-generated training synthetic speech representation 202. The trained speech clone creation system 200 obtains training data 10 that includes a training audio signal 102 and a corresponding transcript 106. Each training signal 102 can be associated with conditioning inputs including a speaker embedding / identifier 108 and an accent / dialect identifier 109. Here, the training audio signal 102 of the training data 10 represents human speech in a first accent / dialect (e.g., American English). Based on the training audio signal 102 (and optionally the corresponding transcript), the trained speech clone creation system 200 is configured to generate a training synthetic speech representation 202 that includes the speech of a target speaker in a second accent / dialect different from the first accent / dialect. The training synthetic speech representation 202 can include a sequence of audio waveforms or mel-frequency spectrograms. The trained speech clone creation system 200 provides a training synthetic speech representation 202 for training an untrained TTS model 300.
[0041] The untrained TTS system 300 includes a TTS model 400 and a synthesizer 150. The TTS model 400 includes an encoder part 400a and a decoder part 400b. The TTS model 400 may further include a variational layer. The encoder part 400a is trained to learn a method of encoding a training synthetic speech representation 202 into a corresponding utterance embedding 204 that represents the prosody and / or a second accent / dialect captured by the training synthetic speech representation 202. During training, the decoder part 400b is conditioned on a transcript 106 and conditional inputs (e.g., a speaker embedding / identifier 108 and an accent / dialect identifier), and is configured to decode the utterance embedding 204 encoded from the training synthetic speech representation 202 by the encoder part 400a into a predicted output audio signal 402. During training, the decoder part 400b receives the transcript 106 of the training data and the utterance embedding 204 to generate the predicted output audio signal. The goal of training is to minimize the loss between the predicted output audio signal 402 and the training synthetic speech representation 202. The decoder part 400b may also generate some predicted frames 280 corresponding to the predicted output audio signal 402. That is, the decoder part 400b decodes the utterance embedding 204 into a sequence of fixed-length predicted frames 280 (alternatively referred to as "predicted frames") that provide prosody features and / or accent / dialect information. The prosody features represent the prosody of the training synthetic speech representation 202 and include duration, pitch contour, energy contour, and / or mel-frequency spectrogram contour.
[0042] In some implementations, the synthesizer 150 is trained to learn a method for generating a predicted synthetic speech representation 152 from the number of predicted frames 280 corresponding to the predicted output audio signal 402 from the TTS model 400. Here, the predicted synthetic speech representation may create a clone of the target speaker's voice in a second accent / dialect and further include the prosody captured by the training synthetic speech representation 202. More specifically, the synthesizer 150 receives, as a ground truth label, the training synthetic speech representation 202 output from the voice cloning system 200, similar to the TTS model 400, in order to teach the synthesizer 150 to generate a predicted synthetic speech representation 152 that matches the training synthetic speech representation 202. During training, the synthesizer 150 generates a gradient / loss 154 between the predicted synthetic speech representation 152 and the training synthetic speech representation 202. In some examples, the synthesizer 150 backpropagates the gradient / loss 154 through the TTS model 400 and the synthesizer 150.
[0043] Once the synthesizers of the TTS model 400 and the TTS system 300 are trained, the trained TTS system 300 applies only the decoder portion 400b to generate synthetic speech 152 in a second accent / dialect from the input text utterance 320. That is, the decoder portion 400b may decode the input text utterance 320 and the selected utterance embedding 204 conditioned on the conditioning inputs 108, 109 into the output audio waveform 402 and the corresponding number of predicted frames 280. Thereafter, the synthesizer 150 uses the number of predicted frames 280 to generate synthetic speech 152 that creates a clone of the target speaker's voice in a second accent / dialect.
[0044] Figures 4A and 4B show the TTS model 400 of FIG. 3, which represents the input text utterance 320 by a hierarchical language structure for synthesizing into expressive speech that creates a clone of the target speaker's voice in the target accent / dialect. As will become apparent, the TTS model 400 has the target accent / dialect for each syllable of a given input text utterance 320 and can be trained to predict together the duration of the syllable and the pitch (F0) and energy (C0) contours of that syllable to generate the synthesized speech 152 in the voice of the target speaker without relying on a unique mapping from the given input text utterance or other language specifications.
[0045] During training, the hierarchical language structure of the TTS model 400 includes an encoder portion 400a (FIG. 4A) that encodes a plurality of fixed-length reference frames 211 sampled from the training synthesized speech representation 202 into a fixed-length utterance embedding 204, and a decoder portion 400b (FIG. 4B) that learns a way to decode the fixed-length utterance embedding 204. The decoder portion 400b can decode the fixed-length utterance embedding 204 into an output audio waveform 402 that includes some predicted frames 280 of expressive speech. As will become apparent, the TTS model 400 is trained such that the number of predicted frames 280 output from the decoder portion 400b is equal to the number of reference frames 211 input to the encoder portion 400a. Further, the TTS model 400 is trained such that the accent / dialect and prosody information associated with the reference frames 211 and the predicted frames 280 substantially match each other.
[0046] Referring to FIGS. 3 and 4A, the encoder portion 400a receives a sequence of fixed-length reference frames 211 sampled from the synthetic speech representation 202 output from the trained voice cloning system 200. The training synthetic speech representation 202 includes the voice of a target speaker with a target accent / dialect. The reference frame 211 includes a duration of 5 milliseconds (ms) and may represent either the pitch contour (F0) or the energy contour (C0) (and / or the spectral characteristic contour (M0)) of the synthetic speech representation 202. In parallel, the encoder portion 400a may also receive a second sequence of reference frames 211, each including a duration of 5 milliseconds, representing the other of the pitch contour (F0) or the energy contour (C0) (and / or the spectral characteristic contour (M0)) of the synthetic speech representation 202. Thus, the sequence of reference frames 211 sampled from the synthetic speech representation 202 provides a duration, pitch contour, energy contour, and / or spectral characteristic contour for representing the target accent / dialect and / or prosody of the synthetic speech representation 202. The length or duration of the synthetic speech representation 202 correlates to the sum total of the total number of reference frames 211.
[0047] The encoder portion 400a includes hierarchical levels of the reference frame 211 of the synthetic speech representation 202, phonemes 421, 421a, syllables 430, 430a, words 440, 440a, and sentences 450, 450a that are clock-controlled with respect to each other. For example, the level associated with the sequence of the reference frame 211 clock-controls faster than the next level associated with the sequence of the phonemes 421. Similarly, the level associated with the sequence of the syllables 430 clock-controls slower than the level associated with the sequence of the phonemes 421 and faster than the level associated with the sequence of the words 440. Thus, in order to essentially provide a sequence-to-sequence encoder, the slower clocking layer receives the output from the faster clocking layer as input so that the output after the faster subsequent final clock (i.e., state) is taken as input to the corresponding slower layer. In the illustrated example, the hierarchical levels include long short-term memory (LSTM) levels.
[0048] In the illustrated example, the synthetic speech representation 202 includes one sentence 450, 450A that includes three words 440, 440A - C. The first word 440, 440A includes two syllables 430, 430Aa - Ab. The second word 440, 440B includes one syllable 430, 430Ba. The third word 440, 440a includes two syllables 430, 430Ca - Cb. The first syllable 430, 430Aa of the first word 440, 440A includes two phonemes 421, 421Aa1 - Aa2. The first syllable 430, 430Ba of the second word 440, 440B includes three phonemes 421, 421Ba1 - Ba3. The first syllable 430, 430Ca of the third word 440, 440C includes one phoneme 421, 421Ca1. The second syllable 430, 430Cb of the third word 440, 440C includes two phonemes 421, 421Cb1 - Cb2.
[0049] In some implementations, the encoder portion 400a first encodes the sequence of the reference frame 211 into frame-based syllable embeddings 432, 432Aa - Cb. Each frame-based syllable embedding 432 may represent a reference prosodic feature represented as a numerical vector indicating the duration, pitch (F0), and / or energy (C0) associated with the corresponding syllable 430. In some implementations, the reference frame 211 defines a sequence of phonemes 421Aa1 - 421Cb2. Here, instead of encoding a subset of the reference frame 211 into one or more phonemes, the encoder portion 400a takes into account the phoneme 421 by encoding the single-phoneme level linguistic features 422, 422Aa1 - Cb2 into single-phoneme feature-based syllable embeddings 434, 434Aa - Cb. Each phoneme-level linguistic feature 422 may indicate the position of the phoneme, while each phoneme feature-based syllable embedding 434 includes a vector indicating the position of each phoneme within the corresponding syllable 430 and the number of phonemes 421 within the corresponding syllable 430. For each syllable 430, each respective syllable embedding 432, 434 may be concatenated with the respective syllable-level linguistic features 436, 436Aa - Cb of the corresponding syllable 430 and encoded. Further, each syllable embedding 432, 434 indicates the corresponding state at the syllable 430 level.
[0050] Continuing to refer to FIG. 4A, blocks within the hierarchical layer that include a diagonal hatching pattern correspond to language features (excluding word level 440) at a particular level of the hierarchy. The hatching pattern at word level 440 includes word embeddings 442 extracted as language features from input text utterance 320 (during inference), or WP embeddings 442 output from a bidirectional encoder representation from Transformer (BERT) model 470 based on word units 472 obtained from transcript 106. Since there is no concept of word pieces in the recurrent neural network (RNN) portion of encoder 400a, the WP embedding 442 corresponding to the first word piece of each word can be selected to represent a word that may include one or more syllables 430. Using frame-based syllable embeddings 432 and single-phoneme feature-based syllable embeddings 434, encoder portion 400a encodes these syllable embeddings 432, 434 along with other language features 436, 453, 442 (or WP embedding 442). For example, encoder portion 400a encodes syllable embeddings 432, 434 concatenated with syllable-level language features 436, 436Aa~Cb, word-level language features (or WP embeddings 432, 432A~C output from BERT model 470), and / or sentence-level language features 452, 452A. By encoding syllable embeddings 432, 434 along with language features 436, 452, 442 (or WP embedding 442), encoder portion 400a generates an utterance embedding 204 for the synthetic speech representation 202. The utterance embedding 204 can be stored in data storage 180 (FIG. 1) along with the transcript 106 (e.g., text representation) of the synthetic speech representation 202. From training data 10, language features 432, 442, 452 can be extracted and stored for use in adjusting the training of the hierarchical language structure. Language features (e.g., language features 422, 436, 442, 452) can include, but are not limited to, the individual sounds for each phoneme and / or the position of each phoneme within a syllable, whether each syllable is emphasized, syntactic information for each word, whether the utterance is a question or a phrase, and / or the gender of the speaker of the utterance.As used herein, any reference to word-level linguistic features 442 regarding the encoder portion 400a and decoder portion 400b of the TTS model 400 can be replaced with WP embeddings from the BERT model 470.
[0051] In the example of FIG. 4A, encoding blocks 422, 422Aa - Cb are shown to illustrate the encoding between the linguistic features 436, 442, 452 and the syllable embeddings 432, 434. Here, block 422 is sequence - encoded at the syllable rate to generate the utterance embedding 204. By way of example, the first block 422Aa is supplied as an input to the second block 422Ab. The second block 422Ab is supplied as an input to the third block 422Ba. The third block 422Ba is supplied as an input to the fourth block 422Ca. The fourth block 422Ca is supplied to the fifth block 422Cb. In some configurations, the utterance embedding 204 includes a mean μ and a standard deviation σ that relate to the training data of a plurality of training synthetic speech representations 202.
[0052] In some implementations, each syllable 430 receives, as input, the corresponding encoding of a subset of the reference frames 211 and includes a duration equal to the number of reference frames 211 within the encoded subset. In the illustrated example, the first seven fixed-length reference frames 211 are encoded into syllable 430Aa, the next four fixed-length reference frames 211 are encoded into syllable 430Ab, the next eleven fixed-length reference frames 211 are encoded into syllable 430Ba, the next three fixed-length reference frames 211 are encoded into syllable 430Ca, and the last six fixed-length reference frames 211 are encoded into syllable 430Cb. Thus, each syllable 430 within the syllable sequence 430 can include a corresponding duration based on the number of reference frames 211 encoded into the syllable 430, as well as a corresponding pitch and / or energy contour. For example, syllable 430Aa includes a duration equal to 35 milliseconds (i.e., seven reference frames 211 each having a fixed length of 5 milliseconds), and syllable 430Ab includes a duration equal to 20 milliseconds (i.e., four reference frames 211 each having a fixed length of 5 milliseconds). Thus, the level of the reference frames 211 clock-controls a total of 10 times for a single clocking between syllable 430Aa and the next syllable 430Ab at the level of the syllable 430. The duration of the syllable 430 can indicate the timing of the syllable 430 and the pauses between adjacent syllables 430.
[0053] In some examples, the utterance embedding 204 generated by the encoder portion 400a is a fixed-length utterance embedding 204 that includes a numerical vector representing the accent / dialect and / or prosody of the synthetic speech representation 202. In some examples, the fixed-length utterance embedding 204 includes a numerical vector having a value equal to "128" or "256".
[0054] Referring now to FIGS. 3 and 4B, during training, the decoder portion 400b of the TTS model 400 is configured to generate a plurality of fixed-length syllable embeddings 435 by first decoding a fixed-length utterance embedding 204 that specifies the target accent / dialect and prosody of the transcript 106. More specifically, the utterance embedding 204 represents the target accent / dialect and prosody held by the synthetic speech representation 202 output from the trained voice cloning system 200. Further, the decoder portion 400b decodes the fixed-length utterance embedding 204 associated with the transcript 106 using the received speaker embedding / identifier 108 indicating the voice characteristics of the target speaker and / or the accent / dialect identifier 109 indicating the target accent / dialect of the resulting synthetic speech 152. Thus, the decoder portion 400b is configured to backpropagate the utterance embedding 204 to generate a plurality of fixed-length predicted frames 280 that exactly match the plurality of fixed-length reference frames 211 encoded by the encoder portion 400a of FIG. 4A. For example, the fixed-length predicted frames 280 for both pitch (F0) and energy (C0) can be generated in parallel to represent a target accent / dialect (e.g., a predicted accent) that substantially matches the target accent / dialect prosody held by the training synthetic speech representation 202. In some examples, the speech synthesizer 150 uses the fixed-length predicted frames 280 to generate synthetic speech 152 that creates a clone of the target speaker's voice in the intended accent / dialect based on the fixed-length utterance embedding 204. For example, the unit selection module or WaveNet module of the speech synthesizer 150 can use some of the predicted frames 280 to generate synthetic speech 152 having the intended accent and / or the intended prosody.In particular, as described above, the intended accent / dialect generated in the synthetic speech 152 includes an accent / dialect that is not native to the target speaker and has not been spoken by the target speaker in any of the reference utterances of the training data 10.
[0055] In the illustrated example, the decoder portion 400b decodes the utterance embedding 204 received from the encoder portion 400a into the hierarchical levels of words 440, 440b, syllables 430, 430b, phonemes 421, 421b, and the fixed-length predicted frames 280. Specifically, the fixed-length utterance embedding 204 corresponds to the variational layer of the hierarchical input data of the decoder portion 400b, and each of the stacked hierarchical levels includes a long short-term memory (LSTM) processing cell that is variably clock-controlled according to the length of the hierarchical input data. For example, the syllable level 430 is clock-controlled faster than the word level 440 and slower than the phoneme level 421. The rectangular blocks at each level correspond to the LSTM processing cells for words, syllables, phonemes, or frames, respectively. Advantageously, the trained voice cloning system 200 provides the LSTM processing cell at the word level 440 with memory over the last 1000 words, the LSTM cell at the syllable level 430 with memory over the last 100 syllables, the LSTM cell at the phoneme level 421 with memory over the last 100 phonemes, and the LSTM cell for the fixed-length pitch and / or energy frames 280 with memory over the last 100 fixed-length frames 280. If the fixed-length frames 280 each include a duration of 5 milliseconds (e.g., frame rate), the corresponding LSTM processing cell provides memory over the last 500 milliseconds (e.g., 0.5 seconds).
[0056] In the illustrated example, the decoder portion 400b of the hierarchical language structure simply backpropagates the fixed-length utterance embedding 204 encoded by the encoder portion 400a into sequences of three words 440A - 440C, sequences of five syllables 430Aa - 430Cb, and sequences of nine phonemes 421Aa1 - 421Cb2 in order to generate a sequence of predicted fixed-length frames 280. The decoder portion 400b is conditioned on the language features of the training data 10 during training and on the input text utterance 320 during inference. In contrast to the encoder portion 400a of FIG. 4A where the output from a faster clocking layer is received as input by a slower clocking layer, the decoder portion 400b includes the output from a slower clocking layer that is fed to a faster clocking layer such that the output of the slower clocking layer has a timing signal added thereto at each clock cycle and is distributed to the input of the faster clocking layer. Further details of the TTS model 400 are described with reference to U.S. Patent Application No. 16 / 867,427, filed May 5, 2020, the entire contents of which are incorporated by reference.
[0057] Referring to FIG. 4B, in some implementations, the hierarchical language structure of the TTS model 400 is adapted to provide a controllable model for predicting the mel-spectral information of the input text utterance 320 during inference, while effectively controlling the accents / dialects and prosodies implicitly represented in the mel-spectral information. Specifically, the TTS model 400 can predict the mel-frequency spectrogram 502 of the input text utterance and provide the mel-frequency spectrogram 502 as an input to the vocoder network 155 of the speech synthesizer 150 for conversion to a time-domain audio waveform. The time-domain audio waveform includes an audio waveform that defines the amplitude of the audio signal over time. As will become apparent, the speech synthesizer 150 can generate the synthesized speech 152 from the input text utterance 320 using the TTS system 300 trained based on the sample transcript 106 and the training synthesized speech representation 202 output from the trained voice cloning system 200. That is, the TTS system 300 does not receive complex linguistic and acoustic features that require expertise in areas critical for generation. Instead, it can use an end-to-end deep neural network to convert the input text utterance 320 into the mel-frequency spectrogram 502. The vocoder network 155, i.e., the neural vocoder, can be separately trained and conditioned based on the mel-frequency spectrogram for conversion to a time-domain audio waveform.
[0058] The mel-frequency spectrogram includes a frequency-domain representation of the sound. The mel-frequency spectrogram emphasizes the low frequencies that are important for speech intelligibility while de-emphasizing the high frequencies that are dominated by fricatives and other noise bursts and generally do not need to be modeled with high fidelity. The vocoder network 155 can be any network configured to receive the mel-frequency spectrogram and generate audio output samples based on the mel-frequency spectrogram. For example, the vocoder network 155 can be based on the parallel feed-forward neural network described in "Parallel WaveNet: Fast High-Fidelity Speech Synthesis" by van den Oord, which is available at https: / / arxiv.org / pdf / 1711.10433.pdf and is hereby incorporated by reference. Alternatively, the vocoder network 155 can be an autoregressive neural network.
[0059] Referring now to FIG. 5, the spectrogram decoder 500 (also referred to as the decoder portion 500) of the trained voice cloning system 200 can include an architecture having a prenet 510, a long short-term memory (LSTM) subnetwork 520, a linear projection 530, and a convolutional postnet 540. The prenet 510 through which the mel-frequency predictions from previous time steps pass can include two fully-connected layers of rectified linear units (ReLUs). The prenet 510 functions as an information bottleneck for learning attention to increase the convergence rate of the speech synthesis system during training and improve its generalization ability. Dropout with a probability of 0.5 can be applied after the prenet 510 to introduce output variability.
[0060] The LSTM sub-network 520 may include two or more LSTM layers. At each time step, the LSTM sub-network 520 receives the concatenation of the outputs of the pre-network 510, and the fixed-length context vector 225 (e.g., the text encoding output from the encoders of FIGS. 2A and 2B) is projected to a scalar to predict that the output sequence of the mel spectrogram 502 is complete and passes through a sigmoid activation. The LSTM layer can be normalized using zoneout with a probability of, for example, 0.1. The linear projection receives the output of the LSTM sub-network 520 as input and generates a prediction of the mel frequency spectrograms 502, 502P.
[0061] The convolutional post-network 540 having one or more convolutional layers processes the predicted mel frequency spectrogram 502P at each time step to predict a residual 542 that is added to the predicted mel frequency spectrogram 502P predicted in the adder 550. This improves the overall reconstruction. After each convolutional layer except the last convolutional layer, batch normalization and hyperbolic tangent (TanH) activation may follow. The convolutional layer is normalized using dropout with a probability of, for example, 0.5. The residual 542 is added to the predicted mel frequency spectrogram 502P generated by the linear projection 520, and the sum (i.e., the mel frequency spectrogram 502) can be provided to the speech synthesizer 150. In some implementations, in parallel with the decoder portion 500 predicting the mel frequency spectrogram 502 for each time step, the concatenation of the output of the LSTM sub-network 520, [utterance embedding], and a portion of the training data 10 (e.g., the character embedding generated by a text encoder (not shown)) is projected to a scalar to predict the probability that the output sequence of the mel frequency spectrogram 502 is complete and passes through a sigmoid activation. The output sequence mel frequency spectrogram 502 corresponds to the training synthetic speech representation 202 of the training data 10 and includes the intended rhythm and the intended accent of the target speaker.
[0062] This "stop token" prediction is used during inference so that the trained voice cloning system 200 can dynamically determine when to end generation rather than always generating over a fixed period. When the stop token indicates that generation has ended, i.e., when the probability of the stop token exceeds a threshold, the decoder portion 500 stops predicting the mel-frequency spectrogram 502P and returns the mel-frequency spectrogram predicted up to that point as the training synthetic speech representation 202. Alternatively, the decoder portion 500 can always generate a mel-frequency spectrogram 502 of the same length (e.g., 10 seconds).
[0063] FIG. 6 is a flowchart of an exemplary configuration of the operations of a method 600 for synthesizing input text utterances into expressive speech having an intended accent / dialect and creating a clone of the voice of a target speaker 432. Data processing hardware 122 (FIG. 1) may execute the operations of method 600 by executing instructions stored in memory hardware 124. In operation 602, method 600 includes the step of obtaining training data 10 including a plurality of training audio signals 102 and corresponding transcripts 106. Each training audio signal 102 corresponds to a reference utterance spoken by the target speaker in a first accent / dialect. Each transcript 106 includes a text representation of the corresponding reference utterance. For each training audio signal 102 of the training audio signals 102, method 600 performs operations 604 and 606. In operation 604, method 600 includes the step of generating a training synthetic speech representation 202 of the corresponding reference utterance spoken by the target speaker by a trained voice cloning system 200 configured to receive as input a training audio signal 102 corresponding to a reference utterance spoken by the target speaker in a first accent / dialect. Here, the training synthetic speech representation 202 includes the voice of the target speaker in a second accent / dialect different from the first accent / dialect. In operation 606, method 600 includes the step of training a text-to-speech (TTS) system 300 based on the corresponding transcript 106 of the training audio signal 102 and the training synthetic speech representation 202 of the corresponding reference utterance generated by the trained voice cloning system 200.
[0064] In operation 608, method 600 includes receiving an input text utterance 320 that is synthesized into richly expressive speech 152 in a second accent / dialect. In operation 610, method 600 includes obtaining a conditional input that includes a speaker embedding / identifier 108 representing the voice characteristics of a target speaker and an accent / dialect identifier 109 that identifies the second accent / dialect. In operation 612, method 600 includes generating an output audio waveform 402 corresponding to a synthesized speech representation 202 of the input text utterance 320 that creates a clone of the voice of the target speaker in the second accent / dialect, using the trained TTS system 300 conditioned on the obtained conditional input and by processing the input text utterance 320. In some implementations, the step of obtaining the conditional input (108, 109) of operation 610 is optional, and the step of performing operation 612 may include generating the synthesized speech representation 202 of the input text utterance 320 using the trained TTS system 300 without conditioning the trained TTS system 300 on any conditional input (108, 109).
[0065] A software application (i.e., a software resource) can refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application", an "app", or a "program". Examples of applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and game applications.
[0066] A non-transitory memory may be a physical device used to temporarily or persistently store programs (e.g., instruction sequences) or data (e.g., program state information) used by a computing device. The non-transitory memory may be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.
[0067] FIG. 7 is a schematic diagram of an exemplary computing device 700 that may be used to implement the systems and methods described herein. Computing device 700 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended only as examples and are not intended to limit implementations of the inventions described and / or claimed herein.
[0068] Computing device 700 includes processor 710, memory 720, storage device 730, high-speed interface / controller 740 connected to memory 720 and high-speed expansion port 750, and low-speed interface / controller 760 connected to low-speed bus 770 and storage device 730. Each of components 710, 720, 730, 740, 750, and 760 is interconnected using various buses and may be mounted on a common motherboard or in other manners as required. Processor 710 can process instructions to be executed within computing device 700, including instructions stored in memory 720 or storage device 730 for displaying graphical information on an external input / output device such as display 780 coupled to high-speed interface 740 in a graphical user interface (GUI). In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and memory types, as required. Also, multiple computing devices 700 may be connected, with each device providing a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).
[0069] Memory 720 stores information non-temporarily within computing device 700. Memory 720 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-temporary memory 720 may be a physical device used to store temporarily or persistently a program (e.g., a sequence of instructions) or data (e.g., program state information) used by computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.
[0070] Storage device 730 can provide mass storage to computing device 700. In some implementations, storage device 730 is a computer-readable medium. In various different implementations, storage device 730 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices including a device in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer or machine-readable medium such as memory 720, storage device 730, memory on processor 710, etc.
[0071] The high-speed controller 740 manages operations that consume a large amount of the bandwidth of the computing device 700, while the low-speed controller 760 manages operations that consume little bandwidth. Such an assignment of duties is merely an example. In some implementations, the high-speed controller 740 is coupled to a high-speed expansion port 750 that can accept a memory 720, a display 780 (e.g., through a graphics processor or accelerator), and various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to a storage device 730 and a low-speed expansion port 790. The low-speed expansion port 790, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or a router, e.g., through a network adapter.
[0072] As shown in the figure, the computing device 700 can be implemented in many different forms. For example, it can be implemented as a standard server 700a, or multiple times within a group of such servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.
[0073] Various implementations of the systems and techniques described herein can be realized in digital electronic circuits and / or optical circuits, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations of one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device, which may be either special purpose or general purpose.
[0074] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD) media) used to provide machine instructions and / or data to a programmable processor for use in receiving the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0075] The processes and logical flows described herein can be performed by one or more programmable processors, also referred to as data processing hardware, which execute one or more computer programs to perform functions by operating on input data to generate output. The processes and logical flows can also be performed by dedicated logic circuitry, such as, for example, an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory, a random access memory, or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer also includes, or is operatively coupled to receive data from, or transfer data to, one or more mass storage devices for storing data, such as, magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, dedicated logic circuitry.
[0076] To provide interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, and optionally a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, tactile feedback, etc., and the input from the user can be received in any form, such as acoustic, voice, tactile input, etc. Further, the computer can interact with the user by transmitting and receiving documents to and from the devices used by the user, such as by transmitting a web page to a web browser on the user's client device in response to a request received from the web browser.
[0077] Numerous implementations have been described. Nevertheless, it will be understood that various changes may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are also included within the scope of the claims.
Description of the Reference Numerals
[0078] 10 Training Data 100 System 102 Training Audio Signal 106 Transcription 106 Transcript 106 Training Text Utterance 108 Speaker Embedding / Identifier 109 Accent / Dialect Identifier 120 Computing System 122 Data Processing Hardware 124 Memory Hardware 150 Synthesizer 150 Speech Synthesizer 152 Speech 152 Synthetic Speech Representation 152 Synthetic Speech 153 Gradient / Loss 154 Gradient / Loss 155 Vocoder Network 180 Data Storage 200 Voice Cloning System 200a Voice Cloning System 200b Voice Cloning System 200b S2S Conversion Model 200b S2S Conversion System 202 Training Synthetic Speech Representation 204 Utterance Embedding 204 Training Synthetic Speech Utterance 210 Inference Network 211 Reference Frame 212 Residual Encoder 214 Residual Encoding 220 Synthesizer 222 Text Encoder 224 Language Embedding 225 Output 225 Text Encoding 225 Fixed-Length Context Vector 225 Fixed-Length Vector 225a~n Text Encoding 225a~n Fixed-Length Context Vector 228 Waveform Synthesizer 228 Vocoder 230 Adversarial Loss Module 233 Adversarial Loss Term 234 Speaker Classifier 240 Spectrogram Encoder 280 Predicted Frame 280 Fixed-Length Frame 300 Text-to-Speech (TTS) System 300A~N TTS System 301 Training Process 320 Input Text Speech 400 TTS Model 400a Encoder Part 400a Encoder 400b Decoder Part 402 Output Audio Waveform, Output Audio Signal 421 Phoneme 421a Phoneme 421Aa1~Aa2 Phonemes 421Aa1~421Cb2 Phonemes 421b Phoneme 421Ba1~Ba3 Phonemes 421Ca1 Phoneme 421Cb1~Cb2 Phonemes 422 Phoneme-Level Linguistic Features 422 Linguistic Features 422 Encoding Block 422Aa~Cb Encoding Block 422Aa1~Cb2 Phoneme-Level Linguistic Features 422Aa First Block 422Ab Second Block 422Ba Third Block 422Ca Fourth Block 422Cb Fifth Block 430 Syllable 430 Syllable Sequence 430a Syllable 430Aa First Syllable 430Ab Syllable 430Ba First Syllable 430b Syllable 430Ca First Syllable 430Ca Syllable 430Cb Second Syllable 430Cb Syllable 432 Syllable Embedding 432 Linguistic Features 432Aa~Cb Syllable Embedding 434 Syllable Embedding 434 Aa~Cb Phoneme Feature-Based Syllable Embedding 435 Syllable Embedding 436 Syllable-Level Linguistic Features 436 Linguistic Features 436 Aa~Cb Syllable-Level Linguistic Features 440 Word 440 Word Level 440a Word 440A First Word 440B Second Word 440b Word 440C Third Word 442 WP Embedding 442 Word-Level Linguistic Features 450 Sentence 450a Sentence 452 Sentence-Level Linguistic Features 452 Linguistic Features 452A Sentence-Level Linguistic Features 453 Linguistic Features 470 Bidirectional Encoder Representations from Transformers (BERT) Model 472 Word Unit 500 Spectrogram Decoder 500 Decoder Part 500 Decoder Neural Network 502 Output Mel-Spectrogram 502 Mel-Spectrogram 502 Mel-Frequency Spectrogram 502P Mel-Frequency Spectrogram 510 Planet 520 Long Short-Term Memory (LSTM) Subnetwork 530 Linear Projection 540 Convolutional Postnet 542 Residual 550 Adder 600 Method 700 Computing Device 700a Standard Server 700b Laptop Computer 700c Rack Server System 710 Processor 720 Memory 730 Storage Device 740 High-Speed Interface / Controller 750 High-Speed Expansion Port 760 Low-Speed Interface / Controller 770 Low-Speed Bus 780 Display 790 Low-Speed Expansion Port
Claims
1. When executed on data processing hardware (122), the data processing hardware (122) is caused to obtain training data (10) including a plurality of training audio signals (102) and corresponding transcripts (106), each training audio signal (102) corresponding to a reference utterance spoken by a target speaker in a first accent / dialect, and each transcript (106) including a text representation of the corresponding reference utterance, the step of for each training audio signal (102) of the training data (10), generate a training synthetic speech representation (202) of the corresponding reference utterance spoken by the target speaker by means of a trained voice cloning system (200) configured to receive, as input, the training audio signal (102) corresponding to the reference utterance spoken by the target speaker in the first accent / dialect, the training synthetic speech representation (202) including the voice of the target speaker in a second accent / dialect different from the first accent / dialect, the step of training a text-to-speech (TTS) system (300) based on the corresponding transcript (106) of the training audio signal (102) and the training synthetic speech representation (202) of the corresponding reference utterance generated by the trained voice cloning system (200); receive an input text utterance (320) to be synthesized into speech (152) in the second accent / dialect; obtain a conditioning input (108, 109) including a speaker embedding (108) representing the voice characteristics of the target speaker and an accent / dialect identifier (109) for identifying the second accent / dialect; generate an output audio waveform (152) corresponding to the synthetic speech representation (202) of the input text utterance (320) by using the trained TTS system (300) conditioned by the obtained conditioning input (108, 109) and processing the input text utterance (320) to create a clone of the voice of the target speaker in the second accent / dialect A computer-implemented method (600) for causing an operation to be performed that includes **Claim 2** The step of training the TTS system (300) includes Training the encoder portion (400a) of the TTS model (400) of the TTS system (300) to encode the training synthesis speech representation (202) of the corresponding reference utterance generated by the trained voice cloning system (200) into an utterance embedding (204) representing the prosody captured by the training synthesis speech representation (202); Training the decoder portion (400b) of the TTS system (300) by decoding the utterance embedding (204) using the corresponding transcript (106) of the training audio signal (102) to generate a predicted output audio signal (402) of expressive speech The computer-implemented method (600) according to claim 1, including **Claim 3** The step of training the TTS system (300) includes Training the synthesizer (150) of the TTS system (300) using the predicted output audio signal (402) to generate a predicted synthesis speech representation (152) of the input text utterance (320), where the predicted synthesis speech representation (152) creates a clone of the voice of the target speaker in the second accent / dialect and has the prosody represented by the utterance embedding (204); Generating a gradient / loss (154) between the predicted synthesis speech representation (152) and the training synthesis speech representation (202); Backpropagating the gradient / loss (153) through the TTS model (400) and the synthesizer (150) The computer-implemented method (600) according to claim 2, further including **Claim 4** The operation further includes Sampling a sequence of fixed-length reference frames that provide reference prosody features representing the prosody captured by the training synthesis speech representation (202) from the training synthesis speech representation (202) The step of training the encoder part (400a) of the TTS model (400) includes the step of training the encoder part (400a) to encode a sequence of the fixed-length reference frames sampled from the training synthetic speech representation (202) into the utterance embedding (204). The computer-implemented method (600) according to claim 2.
5. The step of training the decoder part (400b) of the TTS model (400) includes the step of decoding the utterance embedding (204) into a sequence of fixed-length predicted frames (280) that provide predicted prosodic features of the transcript (106) representing the prosody represented by the utterance embedding (204) using the corresponding transcript (106) of the training audio signal (102). The computer-implemented method (600) according to claim 4.
6. The TTS model (400) is trained such that the number of fixed-length predicted frames decoded by the decoder part (400b) is equal to the number of fixed-length reference frames sampled from the training synthetic speech representation (202). The computer-implemented method (600) according to claim 5.
7. The training synthetic speech representation (202) of the reference utterance includes a sequence of audio waveforms or mel-frequency spectrograms. The computer-implemented method (600) according to claim 1.
8. The trained voice cloning system (200) is further configured to receive the corresponding transcript (106) of the training audio signal (102) as an input when generating the training synthetic speech representation (202). The computer-implemented method (600) according to claim 1.
9. The training audio signal (102) corresponding to the reference utterance spoken by the target speaker includes an input audio waveform of human speech. The training synthetic speech representation (202) includes an output audio waveform of synthetic speech that creates a clone of the voice of the target speaker in the second accent / dialect. The computer-implemented method (600) according to claim 1, wherein the trained voice clone creation system (200) comprises an end-to-end neural network configured to directly convert an input audio waveform into a corresponding output audio waveform.
10. The TTS system (300) A TTS model (400) configured to generate an output audio signal (402) of expressive speech by decoding an utterance embedding (204) into a sequence of fixed-length predicted frames (502) that are conditioned by the conditional input and provide prosodic features using the input text utterance (320), wherein the utterance embedding (204) is selected to specify the intended prosody of the input text utterance (320), and the prosodic features represent the intended prosody specified by the utterance embedding (204), A waveform synthesizer (228) configured to receive as input a sequence of the fixed-length predicted frames (502) and generate as output the output audio waveform corresponding to the synthetic speech representation (202) of the input text utterance (320) that creates a clone of the voice of the target speaker in the second accent / dialect The computer-implemented method (600) according to claim 1, comprising:
11. The computer-implemented method (600) according to claim 10, wherein the prosodic features representing the intended prosody include duration, pitch contour, energy contour, and / or mel-frequency spectrogram contour.
12. Data processing hardware (122); A memory hardware (124) that communicates with the data processing hardware (122) and stores instructions that, when executed by the data processing hardware (122), cause the data processing hardware (122) to perform the method according to any one of claims 1 to 11. A system (100) comprising:
13. A computer program that, when executed by data processing hardware (122), causes the data processing hardware (122) to perform the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Speech synthesis prosody using a BERT model
US11881210B2
Voice-transformation based data augmentation for prosodic classification
US20190272818A1
Speech style transfer
US20200410976A1
Text-to-speech processing
US20210097976A1
Variational embedding capacity in expressive end-to-end speech synthesis
WO2020236990A1