Massive multilingual speech-text joint semi-supervised learning for text-to-speech

Through a large-scale multilingual speech-text joint semi-supervised learning method, the TTS model is trained to generate synthetic speech in multiple different languages, solving the problem of low-resource language pronunciation synthesis in the existing technology, and achieving efficient adaptation of the model in a multilingual environment.

CN120153418APending Publication Date: 2025-06-13GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380076314.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-26
Filing Date
2023-10-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing multilingual text-to-speech (TTS) models are difficult to effectively learn and generate synthetic speech in low-resource languages, resulting in increased complexity and cost in the application of the model in multiple different languages.

Method used

A large-scale multilingual speech-text joint semi-supervised learning method is adopted, by receiving multiple sets of training data, each set of data is associated with different languages, a shared encoder output is generated using text encoder and speech encoder, and a TTS model is trained based on reconstruction loss to teach the model to learn speech synthesis in multiple languages.

Benefits of technology

It realizes the flexible generation of synthetic speech between multiple different languages, reduces the need for high-quality paired training data, and improves the adaptability and efficiency of the TTS model in low-resource languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153418A_ABST
    Figure CN120153418A_ABST
Patent Text Reader

Abstract

A method (600) includes receiving training data (301) including a plurality of sets of text-to-speech (TTS) spoken utterances (510), each set of TTS spoken utterances being associated with a respective language and including TTS utterances of synthetic speech including a corresponding reference speech representation (504) paired with a corresponding input text sequence (502). For each TTS utterance, the method includes generating a corresponding TTS encoded text representation for the corresponding input text sequence (512); generating a corresponding speech code for the TTS utterance of the corresponding synthetic speech (514); generating a shared encoder output (532, 534); generating a predicted speech representation for the TTS utterance of the corresponding synthetic speech (522); and determining a reconstruction loss (545). The method further includes training a TTS model based on the reconstruction loss determined for the TTS utterance in each set of TTS spoken training utterances (501).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to large-scale multilingual speech-text joint semi-supervised learning for text-to-speech. Background Art

[0002] Text-to-speech (TTS) systems read digital text aloud to users and are becoming increasingly popular on mobile devices. Certain TTS models are designed to synthesize various aspects of speech, such as speaking style and language, to produce natural-sounding speech similar to that of a human. Some TTS models are multilingual, enabling the TTS model to output synthesized speech in multiple different languages. However, even these multilingual TTS models are only compatible with a relatively small fraction of all languages in the world. Specifically, the lack of sufficient training data for other languages, especially low-resource languages, hinders the TTS model from learning to generate synthesized speech in these other languages. Thus, training a multilingual TTS model to generate synthesized speech in multiple different languages, even for low-resource languages, further increases the use of the TTS model. Summary of the Invention

[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations for large-scale multilingual speech-text joint semi-supervised learning for text-to-speech. The operations include receiving training data that includes multiple sets of text-to-speech (TTS) utterances. Each set of TTS utterances is associated with a respective language in multiple different languages, the respective language being different from the respective language associated with each other set of TTS utterances, and each set of TTS utterances includes TTS utterances of synthesized speech spoken in the respective language. Each TTS utterance of synthesized speech includes a corresponding reference speech representation paired with a corresponding input text sequence. For each TTS utterance in each set of TTS utterances of the received training data, the operations include: generating a corresponding TTS encoded text representation for the corresponding input text sequence using a text encoder; generating a corresponding speech encoding for the TTS utterance of the corresponding synthesized speech using a speech encoder; generating a shared encoder output using a shared encoder configured to receive the corresponding TTS encoded text representation or the corresponding speech encoding; and determining a reconstruction loss based on a predicted speech representation and a corresponding reference speech representation of the corresponding TTS utterance. The operations further include training a TTS model based on the reconstruction losses determined for the TTS utterances in each set of TTS training utterances to teach the TTS model to learn how to synthesize speech in each of multiple different languages.

[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, for each TTS utterance in each set of TTS spoken training utterances of the received training data, the operations further include: obtaining a corresponding speaker embedding that represents the speaker characteristics of the corresponding speaker who speaks the TTS utterance of the synthetic speech in the corresponding language; and obtaining a corresponding variational embedding that specifies the expected prosody / style of the predicted speech representation generated for the TTS utterance of the corresponding synthetic speech. In these implementations, the text encoder is configured to receive a concatenation of the corresponding input text sequence and the corresponding speaker embedding when generating the corresponding TTS encoded text representation for the corresponding input text sequence, and the speech decoder is conditioned on the corresponding variational embedding and the corresponding speaker embedding when generating the predicted speech representation for the TTS utterance of the corresponding synthetic speech. In some examples, for each TTS utterance in each set of TTS spoken training utterances of the received training data, the operations further include: using an automatic speech recognition (ASR) decoder configured to receive a shared encoder output to generate a sequence of speech recognition hypotheses representing candidate transcriptions of the TTS utterance of the corresponding synthetic speech; and determining an ASR loss based on the sequence of speech recognition hypotheses and the corresponding input text sequence. Here, training the TTS model is further based on the ASR loss determined for the TTS utterances in each set of TTS spoken training utterances.

[0005] The training data may further include multiple sets of automatic speech recognition (ASR) transcribed utterances, each set of ASR transcribed utterances being associated with a corresponding language that is different from the corresponding language associated with each other set of ASR transcribed utterances, and each set of ASR transcribed utterances including ASR utterances of non-synthetic speech spoken in the corresponding language, where each ASR utterance of non-synthetic speech is paired with a corresponding transcription, and training the TTS model includes training the TTS model on the multiple sets of ASR transcribed utterances. The speech decoder may include a recurrent neural network transducer (RNN-T) architecture. In some implementations, the operations further include determining a consistency loss between the speech encoding generated for the TTS utterance of the synthetic speech using the speech encoder and the TTS encoded text representation generated for the input text sequence. In these implementations, training the TTS model is further based on the consistency loss.

[0006] In some examples, the operation further includes determining a modality matching loss between a speech encoding generated for a TTS utterance of synthesized speech using a speech encoder and a TTS encoded text representation generated for an input text sequence. In these examples, the TTS model is trained further based on the modality matching loss. In some implementations, for each TTS utterance in each set of TTS verbal training utterances of the received training data, the operation further includes: obtaining a sequence representation of the corresponding input text sequence concatenated with a variational embedding; using a duration model network to predict the duration of the input text sequence based on the sequence representation and upsample the sequence representation to an upsampled output of a specified number of frames; and determining a duration loss based on the predicted duration and the ground truth duration of the input text sequence. In these implementations, generating a predicted speech representation for a TTS utterance of corresponding synthesized speech using a speech decoder configured to receive a shared encoder output is based on the upsampled output, and training the TTS model further includes training the TTS model on the duration loss determined for the TTS utterances in each set of TTS verbal training utterances.

[0007] The operation may further include: obtaining a masked language modeling (MLM) loss of a speech encoding generated for a TTS utterance of synthesized speech using a speech encoder; and obtaining an aligned MLM loss of a TTS encoded text representation generated for an input text sequence using a text encoder. Here, training the TTS model further includes training the TTS model on the MLM loss and the aligned MLM loss. In some examples, the training data further includes non-verbal text utterances in corresponding multiple different languages, where each non-verbal text utterance is not paired with any verbal utterance of corresponding synthesized speech, and for each non-verbal text utterance, the operation further includes: generating a corresponding non-verbal encoded text representation for the corresponding non-verbal text utterance using a text encoder; and obtaining an aligned masked language modeling (MLM) loss of the corresponding non-verbal encoded text representation generated for the corresponding non-verbal text utterance. In these examples, training the TTS model further includes training the TTS model based on the aligned MLM loss obtained for the non-verbal encoded text representation.

[0008] In some implementations, the training data further includes untranscribed non-synthetic speech utterances in a respective plurality of different languages, where each untranscribed non-synthetic speech utterance is not paired with a corresponding transcription, and for each untranscribed non-synthetic speech utterance, the operations further include: generating, using a speech encoder, a corresponding speech encoding for the corresponding untranscribed non-synthetic speech utterance; and obtaining a masked language modeling (MLM) loss for the corresponding speech encoding generated for the corresponding untranscribed non-synthetic speech utterance. In these implementations, training the TTS model further includes training the TTS model based on the MLM loss obtained for the corresponding speech encoding. The TTS model can include a text encoder and a speech decoder. In some examples, each corresponding input text sequence includes a sequence of graphemes, word-piece model units, phonemes, or bytes.

[0009] Another aspect of the present disclosure provides a system that includes data processing hardware and memory hardware that stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving training data that includes multiple sets of text-to-speech (TTS) spoken utterances. Each set of TTS spoken utterances is associated with a respective language in a plurality of different languages, the respective language being different from the respective language associated with each other set of TTS spoken utterances, and each set of TTS spoken utterances includes a TTS utterance of synthesized speech spoken in the respective language. Each TTS utterance of synthesized speech includes a corresponding reference speech representation paired with a corresponding input text sequence. For each TTS utterance in each set of TTS spoken utterances of the received training data, the operations include: generating, using a text encoder, a corresponding TTS encoded text representation for the corresponding input text sequence; generating, using a speech encoder, a corresponding speech encoding for the corresponding TTS utterance of synthesized speech; generating, using a shared encoder configured to receive the corresponding TTS encoded text representation or the corresponding speech encoding, a shared encoder output; and determining a reconstruction loss based on a predicted speech representation and a corresponding reference speech representation of the corresponding TTS utterance. The operations also include training a TTS model based on the reconstruction losses determined for the TTS utterances in each set of TTS training spoken utterances to teach the TTS model to learn how to synthesize speech in each of the plurality of different languages.

[0010] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, for each TTS utterance in each set of TTS oral training utterances of the received training data, the operations further include: obtaining a corresponding speaker embedding that represents the speaker characteristics of the corresponding speaker who speaks the TTS utterance of the synthesized speech in the corresponding language; and obtaining a corresponding variational embedding that specifies the expected prosody / style of the predicted speech representation generated for the TTS utterance of the corresponding synthesized speech. In these implementations, the text encoder is configured to receive the concatenation of the corresponding input text sequence and the corresponding speaker embedding when generating the corresponding TTS encoded text representation for the corresponding input text sequence, and the speech decoder conditions on the corresponding variational embedding and the corresponding speaker embedding when generating the predicted speech representation for the TTS utterance of the corresponding synthesized speech. In some examples, for each TTS utterance in each set of TTS oral training utterances of the received training data, the operations further include: using an automatic speech recognition (ASR) decoder configured to receive a shared encoder output to generate a sequence of speech recognition hypotheses representing a candidate transcription of the TTS utterance of the corresponding synthesized speech; and determining an ASR loss based on the sequence of speech recognition hypotheses and the corresponding input text sequence. Here, training the TTS model is further based on the ASR loss determined for the TTS utterances in each set of TTS oral training utterances.

[0011] The training data may further include multiple sets of automatic speech recognition (ASR) transcribed utterances, each set of ASR transcribed utterances being associated with a corresponding language that is different from the corresponding language associated with each other set of ASR transcribed utterances, and each set of ASR transcribed utterances including ASR utterances of non-synthesized speech spoken in the corresponding language, where each ASR utterance of non-synthesized speech is paired with a corresponding transcription, and training the TTS model includes training the TTS model on the multiple sets of ASR transcribed utterances. The speech decoder may include a recurrent neural network transducer (RNN-T) architecture. In some implementations, the operations further include determining a consistency loss between the speech encoding generated for the TTS utterance of the synthesized speech using the speech encoder and the TTS encoded text representation generated for the input text sequence. In these implementations, training the TTS model is further based on the consistency loss.

[0012] In some examples, the operation further includes determining a modality matching loss between a speech encoding generated for a TTS utterance of synthesized speech using a speech encoder and a TTS encoded text representation generated for an input text sequence. In these examples, training the TTS model is further based on the modality matching loss. In some implementations, for each TTS utterance in each set of TTS verbal training utterances of the received training data, the operation further includes: obtaining a sequence representation of the corresponding input text sequence concatenated with a variational embedding; using a duration model network to predict the duration of the input text sequence based on the sequence representation and upsample the sequence representation to an upsampled output of a specified number of frames; and determining a duration loss based on the predicted duration and the ground truth duration of the input text sequence. In these implementations, generating a predicted speech representation for a TTS utterance of corresponding synthesized speech using a speech decoder configured to receive a shared encoder output is based on the upsampled output, and training the TTS model further includes training the TTS model on the duration loss determined for the TTS utterances in each set of TTS verbal training utterances.

[0013] The operation may further include: obtaining a masked language modeling (MLM) loss of a speech encoding generated for a TTS utterance of synthesized speech using a speech encoder; and obtaining an aligned MLM loss of a TTS encoded text representation generated for an input text sequence using a text encoder. Here, training the TTS model further includes training the TTS model on the MLM loss and the aligned MLM loss. In some examples, the training data further includes non-verbal text utterances in respective multiple different languages, where each non-verbal text utterance is not paired with any verbal utterance of corresponding synthesized speech, and for each non-verbal text utterance, the operation further includes: generating a corresponding non-verbal encoded text representation for the corresponding non-verbal text utterance using a text encoder; and obtaining an aligned masked language modeling (MLM) loss of the corresponding non-verbal encoded text representation generated for the corresponding non-verbal text utterance. In these examples, training the TTS model further includes training the TTS model based on the aligned MLM loss obtained for the non-verbal encoded text representation.

[0014] In some implementations, the training data further includes untranscribed non-synthetic speech utterances in respective multiple different languages, where each untranscribed non-synthetic speech utterance is not paired with a corresponding transcription, and for each untranscribed non-synthetic speech utterance, the operations further include: generating, using a speech encoder, a corresponding speech encoding for the corresponding untranscribed non-synthetic speech utterance; and obtaining a masked language modeling (MLM) loss for the corresponding speech encoding generated for the corresponding untranscribed non-synthetic speech utterance. In these implementations, training the TTS model further includes training the TTS model based on the MLM loss obtained for the corresponding speech encoding. The TTS model may include a text encoder and a speech decoder. In some examples, each corresponding input text sequence includes a sequence of graphemes, word-piece model units, phonemes, or bytes.

[0015] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic diagram of an example speech recognition system.

[0017] Figure 2 is a schematic diagram of an example automatic speech recognition model.

[0018] Figures 3A to 3C is a schematic diagram of an example training process for training a text-to-speech (TTS) model using ASR-transcribed utterance groups.

[0019] Figure 4 is a schematic diagram of an example alignment model used during an example training process.

[0020] Figures 5A to 5C is a schematic diagram of an example training process for training a TTS model using TTS-transcribed utterance groups.

[0021] Figure 6 is a flowchart of an example operational arrangement of a method for training a large-scale multilingual TTS model.

[0022] Figure 7 is a schematic diagram of an example computing device that may be used to implement the systems and methods described herein.

[0023] In the various figures, like reference numerals indicate like elements. DETAILED DESCRIPTION

[0024] Text-to-speech is the process of generating synthetic speech based on input text data. In some cases, the TTS model is multilingual, whereby the TTS model can receive text inputs and generate synthetic speech corresponding to text inputs in a variety of different languages. Recently, TTS models have made significant progress in synthesizing high-quality human-like speech in multiple languages. However, even multilingual TTS models are only able to generate synthetic speech in a few different languages. A major obstacle preventing TTS models from scaling to hundreds or even thousands of different languages is the difficulty in collecting large amounts of high-quality paired training data for each different language required to train the TTS model. Specifically, low-resource languages have a very sparse (or even zero) amount of paired training data, further increasing the difficulty of scaling the TTS model to these low-resource languages.

[0025] Accordingly, the implementations herein relate to methods and systems for training a large-scale multilingual TTS model using speech-text joint semi-supervised learning. That is, the training process can receive training data that includes multiple sets of TTS utterances. Each set of TTS utterances is associated with a corresponding language, which is different from the corresponding language associated with every other set of TTS utterances. Additionally, each set of TTS utterances includes TTS utterances of synthetic speech in the corresponding language. Here, each TTS utterance of synthetic speech includes a corresponding reference speech representation paired with a corresponding input text sequence. For each TTS utterance in each set of TTS training utterances, the training process uses a text encoder to generate a corresponding TTS encoded text representation, uses a speech encoder to generate a corresponding speech encoding, uses a shared encoder to generate a shared encoder output based on the corresponding TTS encoded text representation or the corresponding speech encoding, uses a speech decoder to generate a predicted speech representation based on the shared encoder output, and determines a reconstruction loss based on the predicted speech representation and the corresponding reference speech representation.

[0026] Notably, the training process can employ one or more components of an automatic speech recognition (ASR) model (e.g., a speech encoder and / or a text encoder) to train the multilingual TTS model. In some examples, the ASR model and the TTS model share the same text encoder. In other examples, the ASR model and the TTS model each include a corresponding text encoder. Finally, the training process trains the multilingual TTS model based on the reconstruction losses determined for the TTS utterances in each set of TTS training utterances to teach the TTS model to learn how to synthesize speech in each of a variety of different languages. More specifically, the training process can update the parameters of the text encoder of the TTS model based on the reconstruction losses.

[0027] Figure 1An example system 100 implementing an automatic speech recognition (ASR) model 200 and a text-to-speech (TTS) model 501 is shown, which resides on a user device 102 of a user 104 and / or on a remote computing device 201 (e.g., one or more servers of a distributed system executed in a cloud computing environment) communicating with the user device 102. Although the user device 102 is depicted as a mobile computing device (e.g., a smartphone), the user device 102 may correspond to any type of computing device, such as but not limited to a tablet device, a laptop / desktop computer, a wearable device, a digital assistant device, a smart speaker / display, a smart home appliance, an automotive infotainment system, or an Internet of Things (IoT) device, and is equipped with data processing hardware 111 and memory hardware 113.

[0028] The user device 102 includes an audio subsystem 108, which is configured to receive utterances 106 spoken by the user 104 (e.g., the user device 102 may include one or more microphones for recording the spoken utterances 106) and convert the utterances 106 into a corresponding digital format associated with input acoustic frames 110 that can be processed by the ASR system 100. In the example shown, the user speaks the corresponding utterance 106 of the phrase “What is the weather in New York City?” in natural language English, and the audio subsystem 108 converts the utterance 106 into a corresponding acoustic frame 110 for input into the ASR system 100. Thereafter, the ASR model 200 receives the acoustic frame 110 corresponding to the utterance 106 as input and generates / predicts a corresponding transcription 120 (e.g., a recognition result / hypothesis) of the utterance 106 as output. In the example shown, the user device 102 and / or the remote computing device 201 also execute a user interface generator 107, which is configured to present a representation of the transcription 120 of the utterance 106 to the user 104 of the user device 102. In some configurations, the transcription 120 output from the ASR system 100 is processed by a natural language understanding (NLU) module, e.g., executed on the user device 102 or the remote computing device 201, to execute user commands. Additionally or alternatively, the TTS model 501 (e.g., executed on any combination of the user device 102 or the remote computing device 201) may convert the transcription into synthetic speech for audible output by the audio subsystem 108 or another device. For example, the original utterance 106 may correspond to a message that the user 104 is sending to a friend, where the transcription 120 is converted into synthetic speech for audible output to the friend to hear the message conveyed in the original utterance 106.

[0029] The TTS model 501 receives a text input 112 corresponding to a word or sequence of words as input and generates a corresponding speech representation 520 of the text input as output. Specifically, the TTS model 501 can generate a text encoding based on the text input 112 and decode the text encoding 520 to produce the speech representation 520. The user 104 can provide the text input 112 to the user device 102 via a user input. In some examples, the user 104 directly provides the text input 112 by typing on the screen of the user device 102. In other examples, the user 104 can speak utterances 106 such that the ASR model 200 generates a transcription 120 based on the utterances 106 that serve as the text input 112. Without departing from the scope of the present disclosure, the text input 112 can correspond to a response, notification, or other communication that the digital assistant conveys to the user 104. The user 104 can also select a target embedding for the TTS model 501 to use to generate a synthetic voice with the speaker characteristics of the target speaker. Additionally or alternatively, the user 104 can further specify the expected prosody / style of the resulting synthetic voice. The audio subsystem 108 including a vocoder can receive the speech representation 520 and generate an audible output of the text input 112 (e.g., via one or more speakers of the user device 102).

[0030] Reference Figure 2 , in some examples, the ASR model 200 includes a recurrent neural network transducer (RNN-T) model architecture that adheres to the latency constraints associated with interactive applications. The use of the RNN-T model architecture is exemplary, and the ASR model 200 can include other architectures such as transformer-transducer and conformer-transducer model architectures. The RNN-T model architecture provides a small computational footprint and uses fewer memory requirements than conventional ASR architectures, making the RNN-T model architecture suitable for performing speech entirely on the user device 102 (e.g., without the need to communicate with a remote server). The RNN-T model architecture of the ASR model 200 includes an encoder network 210, a prediction network 220, and a joint network 230. The encoder network 210, which is roughly analogous to the acoustic model (AM) in a traditional ASR system, includes a stack of self-attention layers (e.g., Conformer layers or Transformer layers) or a recurrent network of stacked long short-term memory (LSTM) layers. For example, the encoder network 210 reads a sequence of d-dimensional feature vectors (e.g., acoustic frames 110( Figure 1 ))x=(x 1 ,x 2 ,···,x T ), where and generate higher-order feature representations at each output step. This higher-order feature representation is denoted as

[0031] Similarly, the prediction network 220 is also an LSTM network, which processes the non-empty symbol sequence y 0 ,...,y ui-1 output by the final Softmax layer 240 currently into a dense representation like a language model (LM). Finally, using the RNN-T model architecture, the representations generated by the encoder network 210 and the prediction / decoder network 220 are combined by the joint network 230. The prediction network 220 can be replaced by an embedded lookup table to improve latency by outputting the looked-up sparse embeddings instead of processing the dense representations. Then the joint network predicts which is a distribution over the next output symbol. In other words, the joint network 230 generates a probability distribution over possible speech recognition hypotheses at each output step (e.g., time step). Here, "possible speech recognition hypotheses" corresponds to a set of output labels that represent symbols / characters using a specified natural language. For example, when the natural language is English, the set of output labels can include twenty-seven (27) symbols, e.g., one label for each of the 24 letters in the English alphabet, and one label specifying a space. Thus, the joint network 230 can output a set of values that indicate the likelihood of each output label in a set of predetermined output labels occurring. The set of values can be a vector and can indicate a probability distribution over the set of output labels. In some cases, the output labels are graphemes (e.g., individual characters, and possibly punctuation marks and other symbols), but the set of output labels is not limited to this. For example, in addition to or instead of graphemes, the set of output labels can also include word pieces, phonemes, and / or whole words. The output distribution of the joint network 230 can include posterior probability values for each of the different output labels. Thus, if there are 100 different output labels representing different graphemes or other symbols, the output y i of the joint network 230 can include 100 different probability values, one probability value for each output label. Then, the probability distribution can be used (e.g., by the Softmax layer 240) to select candidate orthographic elements (e.g., graphemes, word pieces, and / or words) during a beam search process and assign scores to them for determining the transcription 120.

[0032] The Softmax layer 240 can use any technique to select the output label / symbol with the highest probability in the distribution as the next output symbol predicted by the ASR model 200 at the corresponding output step. In this way, the RNN-T model architecture of the ASR model 200 does not make a conditional independence assumption. Instead, the prediction of each symbol is conditioned not only on the acoustics but also on the sequence of labels output so far. The ASR model 200 does assume that the output symbols are independent of future acoustic frames 110, which allows the RNN-T model architecture of the ASR model 200 to be employed in a streaming fashion.

[0033] In some examples, the encoder network (i.e., audio encoder) 210 of the ASR model 200 includes a stack of self-attention layers / blocks, such as conformer blocks. Here, each conformer block includes a series of multi-head self-attention layers, depthwise convolutional layers, and feed-forward layers. The prediction network 220 can have two 2,048-dimensional LSTM layers, where each layer is also followed by a 440-dimensional projection layer. Alternatively, the prediction network 220 can include a stack of transformer blocks or conformer blocks, or an embedding lookup table instead of the LSTM layers. Finally, the joint network 230 can also have 440 hidden units. The Softmax layer 240 can be composed of a unified set of word pieces or graphemes, which is generated using all the unique word pieces or graphemes in multiple training datasets.

[0034] Figures 3A to 3C An example training process 300 for training a TTS model 501 using ASR transcripts of a set of utterances 310 is shown. Specifically, the training process 300 can train the text encoder 202 of the TTS model 501. The TTS model 501 and the ASR model 200 can share the text encoder 204. It is evident that the training process 300 can use training data 301 including multiple sets of ASR utterances 310 to train the TTS model 400. More specifically, each set of ASR utterances 310 in the multiple sets of ASR utterances 310 includes a set of non-verbal text utterances (X 文本 ) 308, a set of transcribed non-synthesized speech utterances (X sup ) 304, and / or untranscribed non-synthesized speech utterances (X unsup) 306. Each non-verbal text utterance 308 includes plain text data (i.e., unpaired data), such that each non-verbal text utterance 308 is not paired with any corresponding verbal audio representation (i.e., speech) of the utterance. The non-verbal text utterance 308 can include any sequence of text blocks, including words, word pieces, phonemes, bytes, and / or graphemes. Each untranscribed non-synthetic speech utterance 306 (also simply referred to as "untranscribed speech utterance 306") includes plain audio data (i.e., unpaired data), such that the untranscribed speech utterance 306 is not paired with any corresponding transcription. On the other hand, each transcribed non-synthetic speech utterance 304 (also simply referred to as "transcribed speech utterance 304") includes a corresponding transcription 302 paired with the corresponding non-synthetic speech representation of the transcribed speech utterance 304.

[0035] In addition, each set of ASR utterances 310 is associated with a corresponding language, which is different from the corresponding language associated with each other set of ASR utterances 310, and each set of ASR utterances includes ASR utterances of non-synthetic speech spoken in the corresponding language. For example, in the illustrated example, the training data 301 includes a first set of ASR utterances 310, 310a, including transcriptions 302, transcribed speech utterances 304, untranscribed speech utterances 306, and non-verbal text utterances 308, each associated with a first corresponding language (e.g., English). Continuing with the illustrated example, the training data 301 also includes a second set of ASR utterances 310, 310b, including transcriptions 302, transcribed speech utterances 304, untranscribed speech utterances 306, and non-verbal text utterances 308, each associated with a second corresponding language (e.g., Chinese). For clarity only, the illustrated example includes two sets of ASR utterances 310 associated with two corresponding languages, as it can be understood that the training data 301 can include multiple sets of ASR utterances 310 associated with any number of languages.

[0036] For simplicity, the training process 300 includes a contrastive self-supervised loss portion 300a ( Figure 3A ), an ASR supervised loss portion 300b ( Figure 3B ), and a consistency regularization portion 300c ( Figure 3C ). The training process 300 trains the TTS model 501 for the total loss based on: a contrastive loss (L 文本 ) 316 derived using the contrastive self-supervised loss portion 300a from a corpus of non-verbal training text utterances (X sup ) 308, transcribed non-synthetic speech utterances (X unsup ) 304, and untranscribed non-synthetic speech utterances (X w2v ) 306; an ASR supervised loss (L 文本)306 and the transcribed non-synthetic speech utterance (X sup ) the supervision loss (L) derived from 304 aux ) 342, 344; and the consistency loss derived using the consistency regularization portion 300c 352.

[0037] In some examples, the training process 300 employs an alignment model 400 that is configured to generate an alignment output (i.e., a text representation) 402 for a corresponding one of the plurality of non-verbal training text utterances 308, transcriptions 302, and / or input text sequences 502 at each of a plurality of output steps. Thus, the alignment model 400 can generate a corresponding alignment output 402 for each of the non-verbal text utterances 308, transcriptions 302, and / or input text sequences 502. Thereafter, the training process 300 uses the generated alignment output 402 to train the TTS model 501.

[0038] Now refer to Figure 4 , in some examples, the alignment model 400 includes an embedding extractor 410, a duration predictor 420, and an upsampler 430. The embedding extractor 410 receives a corresponding one of the non-verbal text utterances 308, transcriptions 302, and / or input text sequences 502. Here, the non-verbal text utterances 308, transcriptions 302, and input text sequences 502 can each include a sequence of text chunks, including words, word pieces, phonemes, bytes, and / or graphemes. Thus, the embedding extractor 410 extracts a corresponding initial text representation (e t)412. For example, the embedding extractor 410 may receive a corresponding input text sequence 502 and extract an initial text representation 412 from the corresponding input text sequence 502. The initial text representation 412 includes embedded vocabulary information from the sequence of text chunks. The duration predictor 420 receives the initial text representation 412 from the embedding extractor 410 and predicts the corresponding text chunk durations (i.e., word, word-piece, phoneme, and / or grapheme durations) 422. The text chunk durations 422 indicate the duration for which the corresponding text chunk will be spoken when a human (or a text-to-speech system) speaks the non-verbal text utterance 308. For example, the input text sequence 502 may include a sequence of phonemes, and the duration predictor 420 predicts the phoneme durations 422 for each phoneme in the sequence of phonemes. In this example, the duration predictor 420 predicts the phoneme durations 422 by predicting the probability of a non-zero duration for each phoneme and the probability of the successive phoneme durations for each phoneme. Since the sequence of phonemes includes regular phonemes, silences between word boundaries, and punctuation marks, only the regular phonemes are associated with non-zero durations, while the silences and punctuation marks are typically associated with successive phoneme durations. Thus, the duration predictor 420 may use a sigmoid activation after the first of two independent activations to predict the probability of a non-zero duration and a softplus activation after the second of two independent projections to predict the successive text chunk durations 422 for each text chunk. The duration predictor 420 determines whether the probability of a non-zero duration for each text chunk is less than a threshold, and when the probability of a non-zero duration is less than the threshold, a multiplier may zero out the successive text chunk durations 422 predicted by the softplus activation for the corresponding text chunk. Otherwise, when the probability of a non-zero duration is not less than the threshold, the predicted text chunk durations 422 may be set equal to the successive phoneme durations predicted by the softplus activation.

[0039] The upsampler 430 receives each corresponding initial text representation 412 output by the embedding extractor 410 and the corresponding predicted text chunk durations 422, and generates an aligned output with multiple frames by upsampling the initial text representation 412 using the corresponding predicted text chunk durations 422. 402. In some examples, the alignment model 400 sends the alignment output 402 to the text encoder 202. In other examples (not shown), the alignment model 400 sends the alignment output 402 to the shared encoder 250 of the frequency encoder 210 (e.g., bypassing the text encoder 202). In these other examples, the alignment output 402 is used as the encoded text representation 312 such that the shared encoder 250 can directly receive the alignment output 402 from the alignment model. In some additional examples, paired training data is available and the upsampler 430 generates the alignment output 402 as follows.

[0040]

[0041] Here, the upsampler includes a resampler and a refiner layer that aligns the initial text embedding 412 to directly align with the corresponding encoded audio representation 314. In other examples, paired training data is not available and the upsampler 430 generates the alignment output 402 as follows.

[0042]

[0043] Specifically, the number of frames of the alignment output 402 indicates the predicted speech duration of the corresponding one of the non-verbal text utterance 308, the transcription 302, or the input text sequence 502. In other words, the number of frames of the alignment output 402 maps (i.e., aligns) the sequence of text chunks of the text input to the speech frames. Here, the upsampler 430 includes a resampler and a refiner layer that replicates the initial text embedding 412 to match the predicted text chunk duration 422 (i.e., the speech duration). Thus, the alignment output 402 includes a text representation of the text input (e.g., the non-verbal text utterance 308, the transcription 302, and / or the input text sequence 502) that has a timing component aligned with how a human would speak the text input.

[0044] Notably, in most instances, the TTS system (i.e., the auxiliary TTS system) generates an audible output to give the text input the timing component of human speech such that the training process can use the audible output (i.e., the synthesized speech) to train the encoder 210. Thus, since the alignment model 400 generates the alignment output 402 that directly maps the sequence of text chunks to the speech frames, the training process 300 does not require speech synthesis of the speech to generate the alignment output 402. That is, the alignment model 400 does not convert the input text into synthesized speech.

[0045] Now specifically referring to Figure 3A In some implementations, the encoder 210 includes reference Figure 3B and Figure 3CA more detailed description of the speech encoder 204 and the text encoder 202. In the example shown, the speech encoder 204 processes audio inputs (e.g., transcribed speech utterances 304 and untranscribed speech utterances 306), and the text encoder 206 processes text inputs (e.g., non-verbal text 308). Each of the speech encoder 204 and the text encoder 202 includes a Conformer encoder, which includes a stack of conformer blocks, and each conformer block includes a series of multi-head self-attention layers, depthwise convolutional layers, and feed-forward layers. Alternatively, the audio encoder 210 may include another type of encoder having a stack of self-attention layers / blocks, such as a Transformer encoder. Each of the speech encoder 204 and the text encoder 202 can be naturally divided into: a feature encoder, including a convolutional subsampling block 212; and a context network, including a stack of linear layers 214 and Conformer blocks 216. In some implementations, the convolutional subsampling block 212 has two two-dimensional convolutional layers, both of which have a stride of (2,2), resulting in a reduction of the feature sequence length to 1 / 4. The convolutional subsampling block 212 receives an input feature / vector sequence (e.g., a Mel spectrogram, such as Figure 1 the acoustic frames 110) associated with each transcribed non-synthetic speech utterance 304 and each untranscribed non-synthetic speech utterance 306 as input, and for each of a plurality of output steps, generates an encoded audio feature 211 corresponding to the corresponding one in the transcribed non-synthetic speech utterance 304 or the untranscribed non-synthetic speech utterance 306 as output. The convolutional subsampling block 212 can receive each alignment output 402 as input, and for each of a plurality of output steps, generate an encoded text feature 213 corresponding to the corresponding one in the alignment output 402 as output.

[0046] The encoded audio features 211 and the encoded text features 213 (i.e., interchangeably referred to as "encoded features 211, 213") output from the convolutional subsampling block 212 can be fed to the masking module 218, in which some of the encoded features 211, 213 are randomly selected and replaced with a trained feature vector shared among all the masked time steps to provide corresponding masked encoded audio features 211, 211m and masked encoded text features 213, 213m. In some examples, the masking module 218 masks the randomly selected encoded features 211, 213 to be masked by randomly sampling a proportion p of the time steps in all the time steps as starting indices without replacement and then masking the subsequent M consecutive time steps from each sample index, whereby some spans may overlap. After applying the masking, the linear layer 214 and the Conformer block 216 of the context network receive the masked encoded features 211m (or the encoded features 211, 213 not selected by the masking module 218), and output corresponding contrast context vectors (i.e., encoded representations) 215 from the masked encoded features 211m, 213m. In addition, the quantizer 217 receives the encoded features 211, 213 as inputs and generates a quantization vector (i.e., a target context vector) 219 as an output. Thereafter, the contrast loss module 315 derives a contrast loss (L w2v ) 316 between the contrast context vector 215 at the masked position and the target context vector 219 as follows.

[0047]

[0048] where c t is the contrast context vector 215 centered on the masked time step t, and q t represents the target context vector 219 in a set of K + 1 candidate target context vectors 219 including q t and K distractors at the time step t. The distractors can be uniformly sampled from other masked time steps of the same utterance.

[0049] The contrast loss 316 is optimized between the contrast context vector 215 at the masked position and the target context vector 219. After the encoder 210 converges on the untranscribed non-synthetic speech utterance 306, the training process is repeated for both the aligned output 402 corresponding to the non-verbal text utterance 308 and the transcribed non-synthetic speech utterance 304. Thus, the contrast loss (L w2v ) is optimized for both the real / human (non-synthetic) and the non-verbal text utterance 308 represented by the aligned output 402, while having an additional auxiliary loss on the transcribed non-synthetic speech utterance 304 and the aligned output 402, as described below with reference to Figure 3Bwill be described in more detail. Thus, the contrastive portion 300a of the training process 300 trains the speech encoder 204 and the text encoder 202 against the derived contrastive loss 316, which is applied to the corresponding encoded features 211, 213 associated with each alignment output 402 provided as input to the encoder 210, each transcribed non-synthetic speech utterance 304, and each untranscribed non-synthetic speech utterance 306. Training the encoder 210 can include updating the parameters of the encoder 210 based on the contrastive loss 316. In some implementations, the contrastive loss module 315 determines the masked language modeling (MLM) loss 318 of the speech input (e.g., the transcribed speech utterance 304 and the untranscribed speech utterance 306) by comparing the contrastive context vectors 215 generated from the masked encoded features with the contrastive context vectors 215 generated from the corresponding unmasked encoded features. Thus, the MLM loss 318 compares the encodings generated for the masked and unmasked encoded features.

[0050] Now referring to Figure 3B , the ASR supervision loss portion 300b of the training process 300 is configured to inject lexical information into the text encoder 204 of the TTS model 501 during pre-training based on the supervision loss terms 342, 344 derived from the transcribed non-synthetic speech utterances 304 and the alignment outputs 402 corresponding to the non-verbal text utterances 308 output by the alignment model 400. It is noted that the ASR supervision loss portion 300b utilizes one or more ASR decoders 390 to generate the supervision loss terms (i.e., the ASR loss) 342, 344. The ASR decoders 390 can include a connectionist temporal classification (CTC) decoder, a listen-attention-spell (LAS) decoder, or an RNN-T decoder. These ASR decoders 390 can include at least one of a phoneme decoder configured to decode a sequence of phonemes or a word-piece decoder configured to decode a sequence of word pieces. The ASR decoders 390 can also include a grapheme decoder configured to decode a sequence of graphemes.

[0051] During the ASR supervision loss portion 300b, the text encoder 202 is configured to receive an alignment output 402 (i.e., text embedding) from the alignment model 400, and the speech encoder 204 is configured to receive the transcribed non-synthetic speech utterance 304. That is, the text encoder 202 generates an encoded text representation 312 for the alignment output 402 (e.g., corresponding to the non-verbal text utterance 308), and the speech encoder 204 of the encoder 210 generates an encoded audio representation 314 for the speech input (i.e., the transcribed non-synthetic speech utterance 304). Here, neither the encoded text representation 312 nor the encoded audio representation 314 may be compatible with the ASR decoder 390. In some examples, the text encoder 202 obtains a corresponding speaker embedding 326, which characterizes the speaker characteristics of the corresponding speaker who speaks the ASR utterance in the corresponding language, and the text encoder generates the corresponding encoded text representation 312 based on the concatenation of the corresponding alignment output 402 and the corresponding speaker embedding 326.

[0052] Accordingly, the ASR supervision loss portion 300b can employ a shared encoder 250 that receives the encoded text representation 312 as input and generates a first encoded shared representation 322 (e 文本 ) as output. Similar to the text encoder 202, the TTS model 501 and the ASR model 200 can share the shared encoder 250. Additionally, the shared encoder 250 receives the encoded audio representation 314 as input and generates a second encoded shared representation (e sup ) 324 as output. Thus, the shared encoder 250 generates the first encoded shared representation 322 and the second encoded shared representation 324 into a shared latent representation space that is compatible with the ASR decoder 390.

[0053] Specifically, the shared encoder 250 receives each encoded text representation 312 corresponding to the alignment output 402 generated from the non-verbal text utterance 308 as input, and for each of a plurality of time steps, generates a first encoded shared representation (e 文本)322 as the output. An ASR decoder 390 including a phoneme decoder or a word-piece decoder receives each first encoded shared representation 332 output from the shared encoder 250 as input and generates a first probability distribution 392 over possible speech recognition hypotheses for the corresponding aligned output 402 at the corresponding output step as output. In some examples, the first probability distribution 392 over possible speech recognition hypotheses includes one of possible phoneme labels, possible word-piece labels, or possible grapheme labels. Thereafter, the ASR supervision loss module 340 can determine an alignment output loss term 342 based on the first probability distribution 392 over possible speech recognition hypotheses for the aligned output 402 corresponding to the non-verbal text utterance 308. Here, the corresponding non-verbal text utterance 308 from which the aligned output 402 is generated also serves as the ground truth transcription 302. Since the aligned output 402 may be masked, the alignment output loss term 342 also serves as an alignment MLM loss. The ASR supervision loss portion 300b can train the text encoder 202 and / or the speech encoder 204 for the alignment output loss term 342 by updating the parameters of the text encoder 202 and / or the speech encoder 204 based on the alignment output loss term 342.

[0054] Similarly, during the ASR supervision loss portion 300b, the shared encoder 250 receives each transcribed encoded audio representation 314 corresponding to the non-synthesized speech utterance 304 as input and, for each of a plurality of time steps, generates a second encoded shared representation (e sup )334 as the output. An ASR decoder 390 including a phoneme decoder or a word-piece decoder receives each second encoded shared representation 334 output from the shared encoder 250 as input and generates a second probability distribution 394 over possible non-synthesized speech recognition hypotheses for the corresponding transcribed non-synthesized speech utterance 304 at the corresponding time step as output. In some examples, the second probability distribution 394 over possible non-synthesized speech recognition hypotheses includes one of possible phoneme labels, possible word-piece labels, or possible grapheme labels. Thereafter, the ASR supervision loss module 340 can determine a non-synthesized speech loss term 344 based on the second probability distribution 394 over possible non-synthesized speech recognition hypotheses and the corresponding transcription 302 paired with the transcribed non-synthesized speech utterance 304. Here, the corresponding transcription 302 serves as the ground truth transcription and can include a sequence of target phonemes, target word-pieces, and / or target graphemes. The ASR supervision loss portion 300b can train the text encoder 202 and / or the speech encoder 204 for the non-synthesized speech loss term 344 by updating the parameters of the text encoder 202 and / or the speech encoder 204 based on the non-synthesized speech loss term 344.

[0055] The untranscribed non-synthetic speech utterance 306 and the non-verbal text utterance 308 each correspond to "unpaired" training data, whereby the contrastive loss (L 文本 ) 316 derived from the non-verbal text utterance (X w2v ) 308 can be combined with the supervised loss associated with the alignment output loss term 342 to obtain the non-verbal text loss function as follows.

[0056]

[0057] Similarly, the contrastive loss (L unsup ) 316 derived from the untranscribed non-synthetic speech utterance (X w2v ) 306 can be used to express the unsupervised speech loss function as follows.

[0058]

[0059] During the training of the text encoder 202 and the speech encoder 204, the alignment output 402 and the untranscribed non-synthetic utterance 306 can be separated or mixed within each batch. To force the text encoder 202 to learn a representation that is effective for both the alignment output 402 corresponding to the non-verbal text utterance 308 and non-synthetic (human / real) speech, a loss mask σ is applied when combining the loss functions of Equation 5 and Equation 6 to obtain the unpaired data loss function as follows.

[0060]

[0061] The transcribed non-synthetic speech utterance 304 corresponds to "paired" and "supervised" training data, whereby the derived contrastive loss L w2v and the derived supervised loss associated with the non-synthetic speech loss term 344 can be combined to obtain the paired data loss function as follows.

[0062]

[0063] Reference Figure 3C , the consistency regularization part (i.e., the modality matching part) 300c of the training process 300 is configured to generate a consistency loss term 352 promotes the text encoder 202 and the speech encoder 204 to learn consistent predictions between non-synthetic speech (e.g., real / human speech) and the aligned output 402 corresponding to the non-verbal text utterance 308. Each training utterance pair includes the corresponding one in the transcribed non-synthetic speech utterance (X sup ) and the paired aligned output 404 of the same utterance as the corresponding transcribed non-synthetic speech utterance 304. Thus, the non-synthetic speech utterance 304 and the paired aligned output 404 in each training utterance pair 303 are associated with the same ground truth transcription. In short, the consistency loss term 352 between the transcribed non-synthetic speech utterance 304 and the paired aligned output 404 of the same training utterance provides an unsupervised training aspect by encouraging the encoder 210 to behave consistently regardless of whether the training utterance belongs to non-synthetic speech (i.e., speech training data) or the aligned output (i.e., text training data), and this unsupervised training aspect is independent of the supervised loss terms between the ground truth transcription 302 and each of the following: the non-synthetic speech recognition hypothesis output by the auxiliary decoder 390; and the speech recognition hypothesis output by the auxiliary decoder 390.

[0064] Similar to Figure 3B the aligned output 402 generated from the non-verbal text utterance 308 in, the alignment model 400 can use the corresponding transcription 302 paired with the transcribed non-synthetic speech utterance 304 to generate each paired aligned output 404. Here, the non-synthetic speech representation 304 is associated with the paired aligned output 404 generated by mapping the non-verbal text utterance 308 into speech frames by the alignment model 400.

[0065] During the consistency regularization section 300c, the text encoder 202 receives each paired aligned output 404 as input and, for each of a plurality of time steps, generates a coded text representation 313 corresponding to the paired aligned output 404 at the corresponding output step as output. In some examples, the text encoder 202 obtains the corresponding speaker embedding 326, which characterizes the speaker features of the corresponding speaker who speaks the ASR utterance in the corresponding language, and the text encoder generates the corresponding coded text representation 312 based on the concatenation of the corresponding aligned output 402 and the corresponding speaker embedding 326. The shared encoder 250 receives the coded text representation 313 as input and generates a first coded shared representation (e * sup)323 as an output. An auxiliary decoder 390 including a phoneme decoder or a word-piece decoder receives each first encoded shared representation 323 output from the shared encoder 250 as an input, and generates a first probability distribution 311 of possible speech recognition hypotheses regarding the aligned output 404 for the corresponding pair at the corresponding output step as an output. In some examples, the first probability distribution 311 of possible speech recognition hypotheses includes one of possible phoneme labels or possible word-piece labels.

[0066] Similarly, the speech encoder 204 receives each transcribed non-synthetic speech utterance 304 as a sequence of features / vectors (e.g., Mel spectrogram, such as Figure 1 the acoustic frames 110) as an input, and for each of a plurality of time steps, generates an encoded audio representation 314 corresponding to the transcribed non-synthetic speech utterance 304 at the corresponding output step as an output. The shared encoder 250 receives the encoded audio representation 314 as an input, and generates a second encoded shared representation (e sup )324 as an output. An auxiliary decoder 390 including a phoneme decoder or a word-piece decoder receives each second encoded shared representation 324 output from the shared encoder 250 as an input, and generates a second probability distribution 394 of possible non-synthetic speech recognition hypotheses regarding the corresponding transcribed non-synthetic speech utterance 304 at the corresponding time step as an output. In some examples, the second probability distribution 394 of possible non-synthetic speech recognition hypotheses includes one of possible phoneme labels or possible word-piece labels.

[0067] Continuing to refer to Figure 3C , the consistency regularization part 300c of the training process 300 further determines a consistency loss term 352 for each output step among the plurality of output steps of each training utterance pair 301, based on the first probability distribution 311 of possible speech recognition hypotheses and the second probability distribution 394 of possible non-synthetic speech recognition hypotheses. For example, the training process 300 may employ a consistency loss term module 350 configured to receive the corresponding non-synthetic speech and speech recognition results 311, 394 output by the auxiliary decoder 390 at each time step, and determine the consistency loss term 352 for the corresponding training utterance pair 301 at the time step.

[0068] In some examples, the consistency regularization part 300c of the training process 300 determines the consistency loss term 352 based on the Kullback-Leibler divergence (D KL ) between the first probability distribution 311 of possible speech recognition hypotheses and the second probability distribution 394 of possible non-synthetic speech recognition hypotheses. Based on D KLThe consistency loss term 352 can be expressed by the following equation.

[0069]

[0070] Here, the consistency loss term 352 determined for the training utterance pairs 301 at each time step provides an "unsupervised" loss term independent of the accuracy of the auxiliary decoder 390 (e.g., independent of Figure 3B the supervised loss terms 342, 344), and thus, can be used to update the parameters of the encoder 210 to promote the consistency between the non-synthetic speech representation and the aligned output of the same utterance. In batch training, the consistency loss term 352 can correspond to the average loss term obtained for the batch. In other words, the consistency loss term 352 allows the text encoder 202 and the speech encoder 204 to learn to behave the same, e.g., make consistent encoding representation predictions for both the non-synthetic speech (e.g., real / human speech) and the aligned output of the same training utterance, regardless of whether the training utterance belongs to the non-synthetic speech or the aligned output.

[0071] Finally, the training process 300 can combine the unpaired data loss function the paired data loss function and the consistency loss term to obtain the overall loss term that can be expressed as follows

[0072]

[0073] where λ 1 can be equal to 1.0, and λ 2 is equal to 0.1. The training process 300 can use the overall loss term to pre-train the audio encoder speech encoder 204 and the text encoder 202 by updating the parameters of the speech encoder 204 and the text encoder 202 to effectively teach the speech encoder 204 and the text encoder 202 to learn the shared representation between speech and text. After pre-training the speech encoder 204 and the text encoder 202, the training process 300 can fine-tune the pre-trained speech encoder 204 and the text encoder 202 for the transcribed speech utterances, which can include supervised training samples of both the aligned output corresponding to the non-verbal text utterances 308 and the non-synthetic (e.g., human speech).

[0074] In some implementations, the training process 300 for pre-training the speech encoder 204 and the text encoder 202 applies encoder consistency regularization. Different from decoder consistency regularization that is applied to the auxiliary decoder during the consistency regularization section 300c and requires assuming labels (e.g., transcriptions 302 and non-verbal text utterances 308), encoder consistency regularization does not require assuming labels and thus has the advantage of allowing it to be applied to all training data 304, 306, 308. Encoder consistency regularization can be applied via the hierarchical contrastive consistency regularization (HCCR) technique, where the encoder activations e, e* from the original / non-augmented and augmented speech are projected via an auxiliary network to generate z and z*. Thereafter, the positive and negative pairs are constructive, and the contrastive loss is calculated as follows.

[0075]

[0076] Specifically for HCCR, a convolutional neural network (CNN) projection network can calculate projections on the duration-increased segments (30, 50, 120 ms) of the encoder activation e to produce 3 views (V), and draw negative examples from the same utterance of the short segment and from other utterances in the batch with 120-ms segments. Thus, the HCCR loss can be calculated for the transcribed non-synthetic speech utterance 304 (paired speech), the untranscribed non-synthetic speech utterance 306 (unpaired speech), and the aligned output 402 generated from the non-verbal text utterance 308 as follows:

[0077]

[0078] The HCCR loss calculated by Equation 11 can be added to Equation 9, where the coefficient 1e-3 is used as part of the overall loss term for pre-training the speech encoder 204 and the text encoder 202.

[0079] In short, the training process 300 trains the TTS model 501 using the ASR utterance 310 set by training the speech decoder 204, the text encoder 202, and / or the shared encoder 250 based on any loss derived from the training process 300. Although the TTS model 501 may not adopt the speech decoder 204 and the shared encoder 240 during inference, the training process 300 also trains these components to learn a better shared representation between speech and text, thereby further training the TTS model 501 (e.g., the text encoder 202 of the TTS model 501) to generate an encoding that accurately represents human speech.

[0080] Figures 5A to 5Cillustrates an example training process 500 for training a TTS model 501 using a set of TTS utterances 510. Similar to the training process 300 ( Figures 3A to 3C ), the training process 500 trains the text encoder 202 of the TTS model 501. However, the training process 500 also trains the speech decoder 520 of the TTS model 501. It is apparent that the training process 500 can use the training data 301, which also includes a set of TTS utterances 510, to train the TTS model 501. Contrary to the ASR utterances 310, the TTS utterances 510 can include synthesized speech, while the ASR utterances 510 include non-synthesized or human speech.

[0081] Each set of TTS utterances 510 in the set of TTS utterances 510 includes a TTS utterance of synthesized speech spoken in the corresponding language. Specifically, each TTS utterance of non-synthesized speech includes a corresponding reference speech representation 504 paired with a corresponding input text sequence 502. Here, the reference speech representation 504 includes audio data paired with the corresponding input text sequence 502, thereby forming labeled training data for training the TTS model 501. The reference speech representation 504 and the TTS utterance 504 can be used interchangeably. In some examples, the reference speech representation 504 and the input text sequence 502 are the same as the transcribed speech utterance 304 and the transcription 302 ( Figures 3A to 3C ). In other examples, the reference speech representation 504 and the input text sequence 502 are different from the transcribed speech utterance 304 and the transcription 302. Each TTS utterance of non-synthesized speech can include a speaker embedding 326 that characterizes the speaker characteristics of the corresponding speaker who speaks the TTS utterance in the corresponding language.

[0082] In addition, each set of TTS utterances 510 is associated with a corresponding language in a plurality of different languages, and the corresponding language is different from the corresponding language associated with each other set of TTS utterances 510. For example, in the illustrated example, the training data 301 includes a first set of TTS utterances 510, 510a, including an input text sequence 502 and a reference speech representation 504, each associated with a first corresponding language (e.g., English), and a second set of TTS utterances 510, 510b, an input text sequence 502, and a reference speech representation 504, each associated with a second corresponding language (e.g., Chinese). For clarity only, the illustrated example includes two sets of TTS utterances 510 associated with two corresponding languages, as it can be understood that the training data 301 can include a set of TTS utterances 510 associated with any number of languages. Each set of TTS utterances 510 can include a corresponding speaker embedding 326.

[0083] For simplicity, the training process 501 includes a contrastive self-supervised loss portion 500a (Figure 5A ) TTS supervision loss part 500b Figure 5B ) and consistency regularization part 500c Figure 5C ) The training process 500 trains the TTS model 501 for the total loss based on the following: the contrastive loss (L w2v ) 516 derived from the reference speech representation 504 and the input text sequence 502 using the contrastive self-supervised loss part 500a; the supervision loss (L aux ) 542, 544 and reconstruction loss 545 derived from the TTS supervision loss part 500b derived from the reference speech representation 504 and the input text sequence 502; and the consistency loss 552 derived using the consistency regularization part 500c. As described above, the training process 500 can employ the alignment model 400 to generate the aligned output 402 of the input text sequence 502 at each of the multiple output steps.

[0084] Now referring specifically to Figure 5A , in some implementations, the speech encoder 204 processes the audio input (e.g., the reference speech representation 504), and the text encoder 202 processes the text input (e.g., the input text sequence 502). Each of the speech encoder 204 and the text encoder 202 includes a Conformer encoder, which includes a stack of conformer blocks, and each conformer block includes a series of multi-head self-attention layers, depthwise convolutional layers, and feed-forward layers. Each of the speech encoder 204 and the text encoder 202 can be naturally divided into: a feature encoder, including a convolutional subsampling block 212; and a context network, including a stack of linear layers 214 and conformer blocks 216. In some implementations, the convolutional subsampling block 212 has two two-dimensional convolutional layers, both of which have a stride of (2, 2), resulting in the feature sequence length being reduced to 1 / 4. The convolutional subsampling block 212 receives the input feature / vector sequence (e.g., mel spectrogram, such as Figure 1 the acoustic frames 110) associated with each transcribed non-synthetic speech utterance 304, each reference speech representation 504, and the input text sequence 502 as input, and for each of the multiple output steps, generates the encoded audio feature 211 corresponding to the corresponding one in the reference speech representation 504 as output. The convolutional subsampling block 212 can receive each aligned output 402 as input, and for each of the multiple output steps, generate the encoded text feature 213 corresponding to the corresponding one in the aligned output 402 as output.

[0085] The encoded audio features 211 and encoded text features 213 output from the convolutional subsampling block 212 (i.e., interchangeably referred to as "encoded features 211, 213") can be fed into the masking module 218, where some of the encoded features 211, 213 are randomly selected and replaced with trained feature vectors shared among all masking time steps to provide corresponding masked encoded audio features 211, 211m and masked encoded text features 213, 213m. In some examples, the masking module 218 masks the randomly selected encoded features 211, 213 to be masked by randomly sampling a proportion p of the time steps in all time steps as starting indices without replacement and then masking the subsequent M consecutive time steps from each sample index, whereby some spans may overlap. After applying the masking, the linear layer 214 and the Conformer block 216 of the context network receive the masked encoded features 211m (or the encoded features 211, 213 not selected by the masking module 218) and output corresponding contrast context vectors (i.e., encoded representations) 215 from the masked encoded features 211m, 213m. Additionally, the quantizer 217 receives the encoded features 211, 213 as inputs and generates quantized vectors (i.e., target context vectors) 219 as outputs. Thereafter, the contrast loss module 515 derives a contrast loss (L w2v ) 516 between the contrast context vector 215 and the target context vector 219 at the masked positions as follows according to Equation 3.

[0086] The contrast loss 516 is optimized between the contrast context vector 215 and the target context vector 219 at the masked positions. The contrast loss (L w2v ) is optimized for both the synthetic speech and the input text sequence 502 represented by the alignment output 402. Thus, the contrast part 500a of the training process 500 trains the speech encoder 204 and the text encoder 202 with respect to the derived contrast loss 516, which is applied to the corresponding encoded features 211, 213 associated with each alignment output 402 and each reference speech representation 504 provided as inputs to the speech encoder 204 or the text encoder 202. Training the speech encoder 204 and / or the text encoder 202 may include updating the parameters of the speech encoder 204 and / or the text encoder 202 based on the contrast loss 516. In some implementations, the contrast loss module 515 determines a masked language modeling (MLM) loss 518 for the speech input (e.g., the reference speech representation 504) by comparing the contrast context vector 215 generated from the masked encoded features with the contrast context vector 215 generated from the corresponding unmasked encoded features. Thus, the MLM loss 518 compares the encodings generated for the masked and unmasked encoded features.

[0087] Now refer to Figure 5B, the TTS supervision loss part 500b of the training process 500 is configured to inject lexical information into the text encoder 202 of the TTS model 501 during training based on supervision loss terms 542, 544 derived from a reference speech representation 504 and an alignment output 402 corresponding to the input text sequence 502. Compared with the ASR supervision loss part 300b ( Figure 3B ), the TTS supervision loss part 500 also employs a speech decoder 520 and determines a reconstruction loss 545. The TTS supervision loss part 500b utilizes one or more ASR decoders 390 to generate supervision loss terms (i.e., ASR losses) 542, 544. The ASR decoders 390 may include a connectionist temporal classification (CTC) decoder, a listen attend spell (LAS) decoder, or an RNN-T decoder. These ASR decoders 390 may include at least one of a phoneme decoder configured to decode a phoneme sequence or a word-piece decoder configured to decode a word-piece sequence. The ASR decoders 390 may also include a grapheme decoder configured to decode a grapheme sequence.

[0088] During the TTS supervision loss part 500b, the text encoder 202 is configured to receive the alignment output 402 (i.e., text embedding) from the alignment model 400, and the speech encoder 204 is configured to receive the reference speech representation 504. That is, the text encoder 202 generates an encoded text representation 512 for the alignment output 402 (e.g., corresponding to the input text sequence 502), and the speech encoder 204 generates an encoded audio representation 514 for the speech input (i.e., the reference speech representation 504 of the TTS utterance). Here, neither the encoded text representation 512 nor the encoded audio representation 514 may be compatible with the ASR decoders 390. In some examples, the text encoder 202 obtains a corresponding speaker embedding 326 that characterizes the speaker features of the corresponding speaker who speaks the ASR utterance in the corresponding language, and the text encoder generates the corresponding encoded text representation 512 based on the concatenation of the corresponding alignment output 402 (or the corresponding input text sequence 502) and the corresponding speaker embedding 326.

[0089] Thus, the TTS supervision loss part 500b may employ a shared encoder 250 that receives the encoded text representation 512 as input and generates a first encoded shared representation 532 (e 文本 ) as output. Similar to the text encoder 202, the TTS model 501 and the ASR model 200 may share the shared encoder 250. Additionally, the shared encoder 250 receives the encoded audio representation 514 as input and generates a second encoded shared representation (e sup)534 as output. Thus, the shared encoder 250 generates the first encoded shared representation 532 and the second encoded shared representation 534 into a shared latent representation space that is compatible with the ASR decoder 390.

[0090] Specifically, the shared encoder 250 receives each encoded text representation 512 corresponding to the aligned output 402 generated from the input text sequence 502 as input, and for each of the multiple time steps, generates a first encoded shared representation (e 文本 )532 as output. The ASR decoder 390, including a phoneme decoder or a word-piece decoder, receives each first encoded shared representation 532 output from the shared encoder 250 as input, and generates a first probability distribution 592 of possible speech recognition hypotheses regarding the corresponding aligned output 402 at the corresponding output step as output. The first probability distribution 592 may represent a candidate transcription of the corresponding TTS utterance. In some examples, the first probability distribution 592 of possible speech recognition hypotheses includes one of possible phoneme labels, possible word-piece labels, or possible grapheme labels. Thereafter, the TTS supervision loss module 540 may determine an alignment output loss term 542 based on the first probability distribution 592 of possible speech recognition hypotheses regarding the aligned output 402 corresponding to the input text sequence 502. Here, the corresponding input text sequence 502 from which the aligned output 402 is generated also serves as the ground truth transcription. Since the aligned output 402 may be masked ( Figure 4 ), the alignment output loss term 542 also serves as an alignment MLM loss. The TTS supervision loss section 500b may train the text encoder 202 and / or the speech encoder 204 for the alignment output loss term (i.e., the ASR loss) 542 by updating the parameters of the text encoder 202 and / or the speech encoder 204 based on the alignment output loss term 542.

[0091] Similarly, during the TTS supervision loss section 500b, the shared encoder 250 receives each transcribed encoded audio representation 514 corresponding to the reference speech representation 504 as input, and for each of the multiple time steps, generates a second encoded shared representation corresponding to the reference speech representation 504 at the corresponding time step (e sup)534 is output. An ASR decoder 390, including a phoneme decoder or a word-piece decoder, receives each second encoded shared representation 534 output from the shared encoder 250 as input and generates a second probability distribution 594 over possible synthetic speech recognition hypotheses for the corresponding reference speech representation 504 at the corresponding time step as output. The first probability distribution 592 may represent a candidate transcription of the corresponding TTS utterance. In some examples, the second probability distribution 594 over possible synthetic speech recognition hypotheses includes one of possible phoneme labels, possible word-piece labels, or possible grapheme labels. Thereafter, the TTS supervision loss module 540 may determine a synthetic speech loss term 544 based on the second probability distribution 594 over possible synthetic speech recognition hypotheses and the corresponding input text sequence 502 paired with the transcribed reference speech representation 504. Here, the corresponding input text sequence 502 serves as the ground truth transcription and may include a sequence of target phonemes, target word-pieces, and / or target graphemes. The TTS supervision loss portion 500b may train the text encoder 202 and / or the speech encoder 204 for the synthetic speech loss term (i.e., the ASR loss) 544 by updating the parameters of the text encoder 202 and / or the speech encoder 204 based on the synthetic speech loss term 544.

[0092] In some examples, the TTS supervision loss portion 500b determines a modality matching loss 505 between the speech encoding 514 generated using the speech encoder 204 for the TTS utterance and the TTS encoded text representation 512 generated for the input text sequence 502. That is, the TTS supervision loss portion 500b compares the speech encoding 514 and the TTS encoded text representation 512, each corresponding to the same utterance to determine the modality matching loss 505. Thereafter, the supervision loss portion 500b trains the speech encoder 204 and / or the text encoder 202 based on the modality matching loss 505.

[0093] The TTS supervision loss portion also employs a speech decoder 520 that can include an RNN-T architecture. The speech decoder 520 can be part of the TTS model 501, whereby the speech decoder 520 is configured to receive a first encoded shared representation 532 or a second encoded shared representation 534 (collectively, the shared encoder outputs 532, 534), and generate a predicted speech representation 522 for a TTS utterance of the corresponding synthetic speech represented by the reference speech representation 504 or the aligned output 402 generated from the input text sequence 502. In some examples, the speech decoder 520 obtains a corresponding variational embedding 528 that specifies the expected prosody / style of the predicted speech representation 522, whereby the speech encoder 520 is conditioned on the corresponding variational embedding 528 and the corresponding speaker embedding 326. The predicted speech representation 522 represents the features of the synthetic speech that the TTS model 501 will generate for the TTS utterance 510. Accordingly, the reconstruction loss 545 is based on the predicted speech representation 522 and the corresponding reference speech representation 504, which serves as the ground truth label from which the predicted speech representation 522 is generated. The training process 500 trains the speech encoder 202, text encoder 204, shared encoder 250, and / or speech decoder 520 based on the reconstruction loss 545 generated for each TTS utterance 510.

[0094] Reference Figure 5C , the consistency regularization portion (i.e., the modality matching portion) 500c of the training process 500 is configured to promote consistent predictions between the text encoder 202 and the speech encoder 204 for non-synthetic speech and the aligned output 402 corresponding to the input text sequence 502 by generating a consistency loss term 552. Each training utterance pair includes a corresponding one in the reference speech representation 504 and the paired aligned output 404 of the same utterance as the corresponding reference speech representation 504. Accordingly, the reference speech representation and the paired aligned output 404 in each training utterance pair 503 are associated with the same ground truth transcription. In short, the consistency loss term 552 between the reference speech representation 504 and the paired aligned output 404 of the same training utterance provides an unsupervised training aspect by encouraging the speech encoder 204 and the text encoder 202 to behave consistently regardless of whether the training utterance belongs to synthetic speech or the aligned output (i.e., text training data), and this unsupervised training aspect is independent of the supervised loss term between the ground truth transcription (i.e., the input text sequence) 502 and each of the speech recognition hypotheses output by the auxiliary decoder 390.

[0095] Similar to Figure 5BFor the aligned output 402 generated from the input text sequence 502 in [description], the alignment model 400 can generate each paired aligned output 404 using the corresponding input text sequence 502 paired with the reference speech representation 503. Here, the reference speech representation 504 is associated with the paired aligned output 404 generated by mapping the input text sequence 502 into speech frames through the alignment model 400.

[0096] During the consistency regularization section 500c, the text encoder 202 receives each paired aligned output 404 as input and, for each of the multiple time steps, generates an encoded text representation 513 corresponding to the paired aligned output 404 at the corresponding output step as output. In some examples, the text encoder 202 obtains a corresponding speaker embedding 326 that characterizes the speaker characteristics of the corresponding speaker who speaks the ASR utterance in the corresponding language, and the text encoder generates the corresponding encoded text representation 512 based on the concatenation of the corresponding aligned output 402 (or the corresponding input text sequence 502) and the corresponding speaker embedding 326. The shared encoder 250 receives the encoded text representation 513 as input and generates a first encoded shared representation (e * sup ) 523 as output. The auxiliary decoder 390 including a phoneme decoder or a word-piece decoder receives each first encoded shared representation 523 output from the shared encoder 250 as input and generates a first probability distribution 511 of possible speech recognition hypotheses regarding the paired aligned output 404 at the corresponding output step as output. In some examples, the first probability distribution 511 of possible speech recognition hypotheses includes one of possible phoneme labels or possible word-piece labels.

[0097] Similarly, the speech encoder 204 receives each reference speech representation 504 as a sequence of features / vectors (e.g., Mel spectrogram, such as Figure 1 the acoustic frame 110) as input and, for each of the multiple time steps, generates an encoded audio representation 514 corresponding to the reference speech representation 504 at the corresponding output step as output. The shared encoder 250 receives the encoded audio representation 514 as input and generates a second encoded shared representation (e sup ) 534 as output. The auxiliary decoder 390 including a phoneme decoder or a word-piece decoder receives each second encoded shared representation 534 output from the shared encoder 250 as input and generates a second probability distribution 594 of possible non-synthetic speech recognition hypotheses regarding the corresponding reference speech representation 504 at the corresponding time step as output. In some examples, the second probability distribution 594 of possible non-synthetic speech recognition hypotheses includes one of possible phoneme labels or possible word-piece labels.

[0098] Continue to refer to Figure 5C In the consistency regularization part 500c of the training process 500, at each of multiple output steps of each training utterance pair 503, a consistency loss term for the corresponding training utterance pair 503 is determined based on a first probability distribution 511 over possible speech recognition hypotheses and a second probability distribution 594 over possible non-synthesized speech recognition hypotheses 352. For example, the training process 500 can employ a consistency loss term module 550 that is configured to receive the corresponding non-synthesized speech and speech recognition results 511, 594 output by the auxiliary decoder 390 at each time step and determine the consistency loss term 552 for the corresponding training utterance pair 503 at the time step

[0099] In some examples, the consistency regularization part 500c of the training process 500 determines the consistency loss term 552 based on the Kullback-Leibler divergence (D KL ) between the first probability distribution 511 over possible speech recognition hypotheses and the second probability distribution 594 over possible non-synthesized speech recognition hypotheses. The consistency loss term 552 based on D KL can be represented by Equation 8. Here, the consistency loss term 552 determined for the training utterance pair 503 at each time step provides an "unsupervised" loss term independent of the accuracy of the auxiliary decoder 390, and thus, can be used to update the parameters of the speech encoder 204 and / or the text encoder 202 to promote the consistency between the synthetic speech representation and the aligned output of the same utterance. In batch training, the consistency loss term 552 can correspond to the average loss term obtained for the batch. In other words, the consistency loss term 552 allows the text encoder 202 and the speech encoder 204 to learn to behave the same, e.g., make consistent encoding representation predictions for both the synthetic speech and the aligned output of the same training utterance, regardless of whether the training utterance belongs to non-synthesized speech or aligned output

[0100] In short, training processes 300 and 500 train the TTS model 500, which includes a text encoder 202 and a speech decoder 520, during inference. Training process 300 uses ASR utterances of non-synthetic speech to train the TTS model 500, and the ASR utterances include transcribed speech utterances, untranscribed speech utterances, and non-verbal text. Training process 500 uses TTS utterances of synthetic speech to train the TTS model 500, and the TTS utterances include speech representations paired with input text sequences. In addition, training processes 300 and 500 use training data from multiple different languages to train the TTS model 500, so that training processes 300 and 500 train the TTS model 500 to be multilingual. By training the TTS model 500 on each loss (or any combination of losses) derived from training processes 300 and 500, the TTS model 500 can be scaled to a large-scale multilingual TTS model, even for languages with little or no training data. Specifically, training processes 300 and 500 utilize text input training data to train the TTS model 500 by generating alignment outputs 402. That is, the alignment outputs 402 enable the TTS model 500 to be trained on text inputs without synthesizing the text inputs.

[0101] Figure 6 is a flowchart of an exemplary operational arrangement of a computer-implemented method 600 for large-scale multilingual speech-text joint semi-supervised learning for text-to-speech. The method 600 can be executed on data processing hardware 710 ( Figure 7 ) using instructions stored in memory hardware 720 ( Figure 7 ). The data processing hardware 710 and the memory hardware 720 can reside on Figure 1 each corresponding to a computing device 700 ( Figure 7 ) of the user device 102 and / or the remote computing device 201.

[0102] At operation 602, method 600 includes receiving training data 301, which includes multiple sets of TTS spoken utterances 510. Each set of TTS spoken utterances 510 is associated with a corresponding language in a plurality of different languages, and the corresponding language is different from the corresponding languages associated with each other set of TTS spoken utterances 510. Additionally, each set of TTS spoken utterances includes TTS utterances 510 of synthesized speech spoken in the corresponding language. Each TTS utterance 510 of synthesized speech includes a corresponding reference speech representation 504 paired with a corresponding input text sequence 502. For each TTS utterance 510 in each set of TTS spoken training utterances 510 of the received training data 301, method 600 performs operations 604-612. At operation 604, method 600 includes generating a corresponding TTS encoded text representation 512 for the corresponding input text sequence 502 using text encoder 202. At operation 604, method 600 includes generating a corresponding speech encoding 514 for the TTS utterance 510 of the corresponding synthesized speech using speech encoder 204, and at operation 608, method 600 includes generating a shared encoder output 532, 534 using a shared encoder 250 configured to receive the corresponding TTS encoded text representation 512 or the corresponding speech encoding 514. At operation 610, method 600 includes generating a predicted speech representation 522 of the TTS utterance 510 of the corresponding synthesized speech using a speech decoder 520 configured to receive the shared encoder output 532, 534. At operation 612, method 600 includes determining a reconstruction loss 545 based on the predicted speech representation 522 of the corresponding TTS utterance 510 and the corresponding reference speech representation 504. At operation 614, method 600 includes training a TTS model 501 based on the reconstruction loss 545 determined for the TTS utterances 510 in each set of TTS spoken training utterances 510 to teach the TTS model 501 to learn how to synthesize speech in each of a plurality of different languages.

[0103] Figure 7 is a schematic diagram of an example computing device 700 that can be used to implement the systems and methods described in this document. Computing device 700 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functions are only intended to be exemplary and are not intended to limit the implementation of the invention described and / or claimed in this document.

[0104] The computing device 700 includes a processor 710, a memory 720, a storage device 730, a high-speed interface / controller 740 connected to the memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 740 connected to a low-speed bus 770 and the storage device 730. Each of the components 710, 720, 730, 740, 750, and 740 is interconnected using various buses and can be mounted on a common motherboard or otherwise as appropriate. The processor 710 can process instructions for execution within the computing device 700, including instructions stored in the memory 720 or on the storage device 730, to display graphical information of a graphical user interface (GUI) on an external input / output device (such as a display 780 coupled to the high-speed interface 740). In other implementations, multiple processors and / or multiple buses and multiple memories and multiple types of memories may be used as appropriate. Additionally, multiple computing devices 700 can be connected (e.g., as a server bank, blade server group, or multi-processor system), where each device provides a portion of the necessary operations.

[0105] The memory 720 stores information non-transitorily within the computing device 700. The memory 720 can be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transitory memory 720 can be a physical device for storing programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 700 on a temporary or permanent basis. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and magnetic disks or tapes.

[0106] The storage device 730 is capable of providing large-capacity storage for the computing device 700. In some implementations, the storage device 730 is a computer-readable medium. In various different implementations, the storage device 730 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices (including devices in a storage area network or other configuration). In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or a machine-readable medium, such as the memory 720, the storage device 730, or the memory on the processor 710.

[0107] The high-speed controller 740 manages the bandwidth-intensive operations of the computing device 700, while the low-speed controller 740 manages the less bandwidth-intensive operations. Such a division of responsibilities is merely exemplary. In some implementations, the high-speed controller 740 is coupled to the memory 720, the display 780 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 750 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 740 is coupled to the storage device 730 and a low-speed expansion port 790. The low-speed expansion port 790, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner; or to a networking device, such as a switch or router, e.g., via a network adapter.

[0108] The computing device 700 can be implemented in many different forms, as shown. For example, the computing device can be implemented as a standard server 700a or multiple times in a group of such servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.

[0109] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuitry, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor which may be special purpose or general purpose and coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0110] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, device, and / or apparatus (e.g., a disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal for providing machine instructions and / or data to a programmable processor.

[0111] The processes and logical flows described in this specification can be executed by one or more programmable processors, also referred to as data processing hardware, which execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be executed by special-purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). By way of example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from one or more mass storage devices or to transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.

[0112] In order to provide for interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), an LCD (liquid crystal display) monitor, or a touch screen) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, a computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on a client device of the user in response to a request received from the web browser.

[0113] A variety of implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.

Claims

1. A computer-implemented method (600) that, when executed on data processing hardware (710), causes the data processing hardware (710) to perform operations, the operations comprising: receiving training data (301), the training data including multiple sets of text-to-speech (TTS) spoken utterances (510), each set of the TTS spoken utterances (510) being associated with a respective language in a plurality of different languages, the respective language being different from the respective languages associated with each other set of the TTS spoken utterances (510), and each set of the TTS spoken utterances (510) including a TTS utterance (510) of synthesized speech spoken in the respective language, each TTS utterance (510) of synthesized speech including a corresponding reference speech representation (504) paired with a corresponding input text sequence (502); for each TTS utterance (510) in each set of the TTS spoken utterances (510) of the received training data (301): generating, using a text encoder (202), a corresponding TTS encoded text representation (512) for the corresponding input text sequence (502); generating, using a speech encoder (504), a corresponding speech encoding (514) for the TTS utterance (510) of the corresponding synthesized speech; generating, using a shared encoder (250) configured to receive the corresponding TTS encoded text representation (512) or the corresponding speech encoding (514), a shared encoder output (532, 534); generating, using a speech decoder (520) configured to receive the shared encoder output (532, 534), a predicted speech representation (522) of the TTS utterance (510) of the corresponding synthesized speech; and determining a reconstruction loss (545) based on the predicted speech representation (522) and the corresponding reference speech representation (504) of the corresponding TTS utterance (510); and training a TTS model (501) based on the reconstruction losses determined for the TTS utterances (510) in each set of the TTS spoken training utterances (510) to teach the TTS model (500) to learn how to synthesize speech in each of the plurality of different languages.

2. The computer-implemented method (600) according to claim 1, wherein the operations further comprise, for each TTS utterance (510) in each set of the TTS spoken training utterances (510) of the received training data (301): obtaining a corresponding speaker embedding (326) that characterizes speaker characteristics of a corresponding speaker who speaks the TTS utterance (510) of the synthesized speech in the respective language; and obtaining a corresponding variational embedding (528) that specifies an expected prosody / style of the predicted speech representation (522) generated for the TTS utterance (510) of the corresponding synthesized speech, where, when generating the corresponding TTS encoded text representation (512) for the corresponding input text sequence (502), the text encoder (202) is configured to receive a concatenation of the corresponding input text sequence (502) and the corresponding speaker embedding (326), and where, when generating the predicted speech representation (522) for the TTS utterance (510) of the corresponding synthesized speech, the speech decoder (520) is conditioned on the corresponding variational embedding (528) and the corresponding speaker embedding (326).

3. The computer-implemented method (600) according to claim 1 or 2, wherein the operation further comprises, for each TTS utterance (510) in each set of the TTS oral training utterances (510) of the received training data (310): using an automatic speech recognition ASR decoder (390) configured to receive the shared encoder outputs (532, 534) to generate a sequence of speech recognition hypotheses (592, 594) representing candidate transcriptions of the TTS utterance (510) of the corresponding synthesized speech; and determining an ASR loss (542, 544) based on the sequence of speech recognition hypotheses (592, 594) and the corresponding input text sequence (502), where training the TTS model (500) is further based on the ASR losses (542, 544) determined for the TTS utterances (510) in each set of the TTS oral training utterances (510).

4. The computer-implemented method (600) according to any one of claims 1 to 3, wherein: the training data (301) further comprises multiple sets of automatic speech recognition ASR transcribed utterances (310), each set of the ASR transcribed utterances (310) being associated with a respective language different from the respective language associated with every other set of the ASR transcribed utterances (310), and each set of the ASR transcribed utterances (310) comprises ASR utterances (304) of non-synthesized speech spoken in the respective language, each ASR utterance (304) of non-synthesized speech being paired with a corresponding transcription (302); and training the TTS model (500) further comprises training the TTS model (500) on the multiple sets of ASR transcribed utterances (310).

5. The computer-implemented method (600) according to any one of claims 1 to 4, wherein the speech decoder (520) comprises a recurrent neural network transducer RNN-T architecture.

6. The computer-implemented method (600) according to any one of claims 1 to 5, wherein the operation further comprises: determining a consistency loss (552) between the speech encoding (514) generated for the TTS utterance (510) of the synthesized speech using the speech encoder (204) and the TTS encoded text representation (512) generated for the input text sequence (522), Wherein, training the TTS model (500) is further based on the consistency loss (500).

7. The computer-implemented method (600) according to any one of claims 1 to 6, wherein the operation further comprises: determining a modality matching loss (505) between the speech encoding (5114) generated for the TTS utterance (510) of the synthetic speech using the speech encoder (204) and the TTS encoded text representation (512) generated for the input text sequence (502), wherein training the TTS model (500) is further based on the modality matching loss (505).

8. The computer-implemented method (600) according to any one of claims 1 to 7, wherein the operation further comprises: obtaining a masked language modeling MLM loss (518) of the speech encoding (215) generated for the TTS utterance (510) of the synthetic speech using the speech encoder (204); and obtaining an aligned MLM loss (542) of the TTS encoded text representation (512) generated for the input text sequence (502) using the text encoder (202), wherein training the TTS model (501) further comprises training the TTS model (501) on the MLM loss (518) and the aligned MLM loss (542).

9. The computer-implemented method (600) according to any one of claims 1 to 8, wherein: the training data (301) further comprises non-verbal text utterances (308) in respective multiple different languages, and each non-verbal text utterance (308) is not paired with any corresponding verbal utterance of the synthetic speech; and the operation further comprises, for each non-verbal text utterance (308): generating a corresponding non-verbal encoded text representation (312) for the corresponding non-verbal text utterance (308) using the text encoder (202); and obtaining an aligned masked language modeling MLM loss (342) of the corresponding non-verbal encoded text representation (312) generated for the corresponding non-verbal text utterance (308), wherein training the TTS model (501) further comprises training the TTS model (501) based on the aligned MLM loss (342) obtained for the non-verbal encoded text representation (308).

10. The computer-implemented method (600) according to any one of claims 1 to 9, wherein: the training data (301) further comprises untranscribed non-synthetic speech utterances (306) in respective multiple different languages, and each untranscribed non-synthetic speech utterance (306) is not paired with a corresponding transcription; and the operation further comprises, for each untranscribed non-synthetic speech utterance (306): generating a corresponding speech encoding (314) for the corresponding untranscribed non-synthetic speech utterance (306) using the speech encoder (204); and Obtain a masked language modeling (MLM) loss (318) for the corresponding speech encoding (314) generated for the corresponding untranscribed non-synthetic speech utterance (306). Wherein training the TTS model (501) further includes training the TTS model (501) based on the MLM loss (318) obtained for the corresponding speech encoding (314).

11. The computer-implemented method (600) according to any one of claims 1 to 10, wherein the TTS model (501) includes the text encoder (202) and the speech decoder (520).

12. The computer-implemented method (600) according to any one of claims 1 to 11, wherein each corresponding input text sequence (502) includes a sequence of graphemes, word-piece model units, phonemes, or bytes.

13. A system (100), comprising: data processing hardware (710); and memory hardware (720) communicatively coupled to the data processing hardware (710), the memory hardware (720) storing instructions that, when executed on the data processing hardware (720), cause the data processing hardware (710) to perform operations including: Receiving training data (301), the training data including multiple sets of text-to-speech (TTS) spoken utterances (510), each set of the TTS spoken utterances (510) being associated with a respective language in a plurality of different languages, the respective language being different from the respective language associated with each other set of the TTS spoken utterances (510), and each set of the TTS spoken utterances (510) including a TTS utterance (510) of synthesized speech spoken in the respective language, each TTS utterance (510) of synthesized speech including a corresponding reference speech representation (504) paired with a corresponding input text sequence (502); For each TTS utterance (510) in each set of the TTS spoken utterances (510) of the received training data (301): Using a text encoder (202) to generate a corresponding TTS encoded text representation (512) for the corresponding input text sequence (502); Using a speech encoder (504) to generate a corresponding speech encoding (514) for the TTS utterance (510) of the corresponding synthesized speech; Using a shared encoder (250) configured to receive the corresponding TTS encoded text representation (512) or the corresponding speech encoding (514) to generate a shared encoder output (532, 534); Using a speech decoder (520) configured to receive the shared encoder output (532, 534) to generate a predicted speech representation (522) of the TTS utterance (510) of the corresponding synthesized speech; and Determining a reconstruction loss (545) based on the predicted speech representation (522) of the corresponding TTS utterance (510) and the corresponding reference speech representation (504); and Train a TTS model (501) based on the reconstruction loss determined for each TTS utterance (510) in each set of the TTS oral training utterances (510) to teach the TTS model (500) how to synthesize speech in each of the multiple different languages.

14. The system (100) of claim 13, wherein the operation further comprises, for each TTS utterance (510) in each set of the received training data (301) of the TTS oral training utterances (510): Obtain a corresponding speaker embedding (326), the speaker embedding characterizing the speaker features of the corresponding speaker who uttered the TTS utterance (510) of the synthesized speech in the corresponding language; and Obtain a corresponding variational embedding (528), the variational embedding specifying the expected prosody / style of the predicted speech representation (522) generated for the TTS utterance (510) of the corresponding synthesized speech, wherein when generating the corresponding TTS encoded text representation (512) for the corresponding input text sequence (502), the text encoder (202) is configured to receive the concatenation of the corresponding input text sequence (502) and the corresponding speaker embedding (326), and wherein when generating the predicted speech representation (522) for the TTS utterance (510) of the corresponding synthesized speech, the speech decoder (520) is conditioned on the corresponding variational embedding (528) and the corresponding speaker embedding (326).

15. The system (100) of claim 13 or 14, wherein the operation further comprises, for each TTS utterance (510) in each set of the received training data (310) of the TTS oral training utterances (510): Use an automatic speech recognition ASR decoder (390) configured to receive the shared encoder outputs (532, 534) to generate a sequence of speech recognition hypotheses (592, 594) representing candidate transcriptions of the TTS utterance (510) of the corresponding synthesized speech; and Determine an ASR loss (542, 544) based on the sequence of speech recognition hypotheses (592, 594) and the corresponding input text sequence (502), wherein training the TTS model (500) is further based on the ASR loss (542, 544) determined for each TTS utterance (510) in each set of the TTS oral training utterances (510).

16. The system (100) of any one of claims 13 to 15, wherein: The training data (301) further includes multiple sets of automatic speech recognition (ASR) transcribed utterances (310), each set of the ASR transcribed utterances (310) being associated with a respective language that is different from the respective language associated with each other set of the ASR transcribed utterances (310), and each set of the ASR transcribed utterances (310) including ASR utterances (304) of non-synthetic speech spoken in the respective language, each ASR utterance (304) of non-synthetic speech being paired with a corresponding transcription (302); and Training the TTS model (500) further includes training the TTS model (500) on the multiple sets of ASR transcribed utterances (310).

17. The system (100) according to any one of claims 13 to 16, wherein the speech decoder (520) includes a recurrent neural network transducer (RNN-T) architecture.

18. The system (100) according to any one of claims 13 to 17, wherein the operation further includes: determining a consistency loss (552) between the speech encoding (514) generated using the speech encoder (204) for the TTS utterance (510) of the synthetic speech and the TTS encoded text representation (512) generated for the input text sequence (522), wherein training the TTS model (500) is further based on the consistency loss (500).

19. The system (100) according to any one of claims 13 to 18, wherein the operation further includes: determining a modality matching loss (505) between the speech encoding (5114) generated using the speech encoder (204) for the TTS utterance (510) of the synthetic speech and the TTS encoded text representation (512) generated for the input text sequence (502), wherein training the TTS model (500) is further based on the modality matching loss (505).

20. The system (100) according to any one of claims 13 to 19, wherein the operation further includes: obtaining a masked language modeling (MLM) loss (518) of the speech encoding (215) generated using the speech encoder (204) for the TTS utterance (510) of the synthetic speech; and obtaining an aligned MLM loss (542) of the TTS encoded text representation (512) generated using the text encoder (202) for the input text sequence (502), wherein training the TTS model (501) further includes training the TTS model (501) on the MLM loss (518) and the aligned MLM loss (542).

21. The system (100) according to any one of claims 13 to 20, wherein: the training data (301) further includes non-verbal text utterances (308) in respective multiple different languages, each non-verbal text utterance (308) not being paired with any corresponding synthetic speech verbal utterance; and The operation further includes, for each non-verbal text utterance (308): using the text encoder (202) to generate a corresponding non-verbal encoded text representation (312) for the corresponding non-verbal text utterance (308); and obtaining an aligned masked language modeling MLM loss (342) of the corresponding non-verbal encoded text representation (312) generated for the corresponding non-verbal text utterance (308), where training the TTS model (501) further includes training the TTS model (501) based on the aligned MLM loss (342) obtained for the non-verbal encoded text representation (308).

22. The system (100) according to any one of claims 13 to 21, wherein: the training data (301) further includes untranscribed non-synthetic speech utterances (306) in respective multiple different languages, each untranscribed non-synthetic speech utterance (306) not paired with a corresponding transcription; and the operation further includes, for each untranscribed non-synthetic speech utterance (306): using the speech encoder (204) to generate a corresponding speech encoding (314) for the corresponding untranscribed non-synthetic speech utterance (306); and obtaining a masked language modeling MLM loss (318) of the corresponding speech encoding (314) generated for the corresponding untranscribed non-synthetic speech utterance (306), where training the TTS model (501) further includes training the TTS model (501) based on the MLM loss (318) obtained for the corresponding speech encoding (314).

23. The system (100) according to any one of claims 13 to 22, wherein the TTS model (501) includes the text encoder (202) and the speech decoder (520).

24. The system (100) according to any one of claims 13 to 23, wherein each corresponding input text sequence (502) includes a sequence of graphemes, word-piece model units, phonemes, or bytes.