Teacherless Parallel Tacotron Non-Autoregressive Controllable Text-to-Speech

The non-autoregressive TTS model with a VAE-based residual encoder efficiently predicts mel-spectrogram sequences, addressing inefficiencies and errors in autoregressive models by disentangling style and rhythm, resulting in high-quality synthetic speech with controlled prosody and speaker-specific characteristics.

JP7709545B2Active Publication Date: 2025-07-16GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023558226
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-22
Filing Date
2021-05-20
Publication Date
2025-07-16
Estimated Expiration
2041-05-20

AI Technical Summary

Technical Problem

Autoregressive text-to-speech (TTS) models are inefficient during inference and prone to mismatches between training and inference, leading to degraded synthetic speech quality due to errors like mumbling, early cutoffs, and word repetitions, especially for longer texts.

Method used

A non-autoregressive TTS model enhanced with a variational autoencoder (VAE)-based residual encoder disentangles style and rhythm information, allowing for efficient prediction of mel-spectrogram sequences that convey intended prosody and style without relying on external aligners, using a duration model network to predict phoneme durations and an attention context representation.

Benefits of technology

The model generates high-quality synthetic speech with controlled rhythm and style, reducing inefficiencies and errors, enabling efficient conversion of text to speech with intended prosody and speaker-specific characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007709545000008
    Figure 0007709545000008
  • Figure 0007709545000009
    Figure 0007709545000009
  • Figure 0007709545000010
    Figure 0007709545000010
Patent Text Reader

Abstract

A method (600) for training a non-autoregressive TTS model (300) includes obtaining a sequence representation (224) of an encoded text sequence (219) concatenated with a variational embedding (220). The method also includes predicting a phoneme duration (240) for each phoneme represented by the encoded text sequence. The method also includes learning interval representations and auxiliary attention context representations based on the predicted phoneme durations, and upsampling the sequence representation to an upsampled output (258) using the interval representation and auxiliary attention context representation. The method also includes generating one or more predicted mel-frequency spectrogram sequences (302) of the encoded text sequence based on the upsampled output. The method also includes determining a final spectrogram loss (280) based on the predicted mel-frequency spectrogram sequence and a reference mel-frequency spectrogram sequence (202), and training a TTS model based on the final spectrogram loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to non-autoregressive and controllable text reading of a teacherless parallel tacotron.

Background Art

[0002] Text-to-speech (TTS) systems read digital text aloud to users and are becoming increasingly popular on mobile devices. Certain TTS models aim to synthesize various aspects of speech, such as speaking style, to produce a voice that sounds natural like a human. Synthesis in a TTS model is a one-to-many mapping problem because there can be multiple speech outputs for different rhythms of text input. Many TTS systems utilize autoregressive models that predict the current value based on previous values. Although autoregressive TTS models can synthesize text and generate a very natural speech output, they require hundreds of calculations, resulting in a decrease in efficiency during inference.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Means for Solving the Problems

[0004] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations for training a non-autoregressive text-to-speech (TTS) model. The operations include obtaining a sequence representation of encoded text sequences concatenated with variational embeddings. The operations also include predicting, using a duration model network, a phoneme duration for each phoneme represented by the encoded text sequences based on the sequence representation. The operations include learning an interval representation using a first function conditioned on the sequence representation and learning an auxiliary attention context representation using a second function conditioned on the sequence representation, based on the predicted phoneme durations. The operations also include upsampling the sequence representation to an upsampled output that specifies the number of frames, using the interval representation matrix and the auxiliary attention context representation. The operations also include generating one or more predicted mel-frequency spectrogram sequences of the encoded text sequences as an output from a stack of one or more self-attention blocks and based on the upsampled output. The operations also include determining a final spectrogram loss based on one or more predicted mel-frequency spectrogram sequences and a reference mel-frequency spectrogram sequence, and training the TTS model based on the final spectrogram loss.

[0005] Implementations of the present disclosure may include one or more of any of the following features. In some implementations, each of the first function and the second function includes a multi-layer perception-based learnable function. This operation may further include determining a global phoneme duration loss based on the predicted phoneme duration and the average phoneme duration. Here, the step of training the TTS model is further based on the global phoneme duration loss. In some examples, the step of training the TTS model based on the final spectrogram loss and the global phoneme duration loss includes training a duration model network to predict the phoneme duration for each phoneme without using the labeled phoneme duration with a teacher extracted from an external aligner.

[0006] In some implementations, this operation further includes generating respective start boundaries and end boundaries based on the predicted phoneme duration for each phoneme represented by the encoded text sequence using a duration model network, and mapping each start boundary and end boundary generated for each phoneme to respective grid matrices based on the number of phonemes represented by the encoded text sequence and the number of reference frames in the reference mel-frequency spectrogram sequence. Here, the step of learning the interval representation is based on each grid matrix mapped from the start boundary and the end boundary, and the step of learning the auxiliary attention context is based on each grid matrix mapped from the start boundary and the end boundary. The step of upsampling the sequence representation to the upsampled output may include determining the product of the interval representation matrix and the sequence representation, determining the Einstein sum (einsum) of the interval representation matrix and the auxiliary attention context representation, and summing the product of the interval representation matrix and the sequence representation and the projection of the einsum to generate the upsampled output.

[0007] In some implementations, the operation includes receiving training data including a reference audio signal and a corresponding input text sequence, where the reference audio signal includes a spoken utterance and the input text sequence corresponds to a transcript of the reference audio signal; using a residual encoder to encode the reference audio signal into a variational embedding, where the variational embedding disentangles style / rhythm information from the reference audio signal; and using a text encoder to encode the input text sequence into an encoded text sequence. In some examples, the residual encoder includes a global variational autoencoder (VAE). In these examples, the step of encoding the reference audio signal into a variational embedding includes sampling a reference mel-frequency spectrogram sequence from the reference audio signal and using the global VAE to encode the reference mel-frequency spectrogram sequence into a variational embedding. Optionally, the residual encoder may include a phoneme-level fine-grained variational autoencoder (VAE). Here, the step of encoding the reference audio signal into a variational embedding includes sampling a reference mel-frequency spectrogram sequence from the reference audio signal, aligning the reference mel-frequency spectrogram sequence with each phoneme in a sequence of phonemes extracted from the input text sequence, and using the phoneme-level fine-grained VAE and based on aligning the reference mel-frequency spectrogram sequence with each phoneme in the sequence of phonemes, encoding a sequence of phoneme-level variational embeddings.

[0008] The residual encoder includes a stack of lightweight convolutional (LConv) blocks. Each LConv block in the stack of LConv blocks includes a gated linear unit (GLU) layer, an LConv layer configured to receive the output of the GLU layer, a residual connection configured to concatenate the output of the LConv layer with the input to the GLU layer, and a final feed-forward layer configured to receive as input the residual connection that concatenates the output of the LConv layer with the input to the GLU layer. In some implementations, this operation further includes the step of concatenating an encoded text sequence, a variational embedding, and a reference speaker embedding representing the identification of the reference speaker who uttered the reference audio signal, and the step of generating a sequence representation based on a duration modeling network that receives as input the concatenation of the encoded text sequence, the variational embedding, and the reference speaker embedding. In some examples, the input text sequence includes a sequence of phonemes. In these examples, the step of encoding the input text sequence into an encoded text sequence includes the steps of receiving, from a phoneme lookup table, the respective embeddings of each phoneme in the sequence of phonemes, processing each embedding using the encoder pre-net neural network of the text encoder to generate a respective transformed embedding for each phoneme, processing each transformed embedding using a bank of convolutional blocks to generate a convolutional output, and processing the convolutional output using a stack of self-attention blocks to generate an encoded text sequence. Optionally, each self-attention block in the stack of self-attention blocks includes the same lightweight convolutional (LConv) block. Each self-attention block in the stack of self-attention blocks includes the same transformer block.

[0009] Another aspect of the present disclosure provides a system for training a non-autoregressive text-to-speech (TTS) model that includes data processing hardware and memory hardware that stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include obtaining a sequence representation of an encoded text sequence concatenated with variational embeddings. The operations also include using a duration model network to predict a phoneme duration for each phoneme represented by the encoded text sequence based on the sequence representation. The operations include learning an interval representation using a first function conditioned on the sequence representation and learning an auxiliary attention context representation using a second function conditioned on the sequence representation, based on the predicted phoneme durations. The operations also include upsampling the sequence representation to an upsampled output that specifies a number of frames, using the interval representation matrix and the auxiliary attention context representation. The operations also include generating one or more predicted mel-frequency spectrogram sequences of the encoded text sequence as an output from a stack of one or more self-attention blocks and based on the upsampled output. The operations also include determining a final spectrogram loss based on the one or more predicted mel-frequency spectrogram sequences and a reference mel-frequency spectrogram sequence, and training the TTS model based on the final spectrogram loss.

[0010] Implementations of the present disclosure may include one or more of any of the following features. In some implementations, each of the first function and the second function includes a multi-layer perception-based learnable function. This operation may further include determining a global phoneme duration loss based on the predicted phoneme duration and the average phoneme duration. Here, the step of training the TTS model is further based on the global phoneme duration loss. In some examples, the step of training the TTS model based on the final spectrogram loss and the global phoneme duration loss includes training a duration model network to predict the phoneme duration for each phoneme without using the labeled phoneme duration with a teacher extracted from an external aligner.

[0011] In some implementations, this operation further includes generating respective start and end boundaries based on the predicted phoneme duration for each phoneme represented by the encoded text sequence using a duration model network, and mapping each of the generated start and end boundaries for each phoneme to respective grid matrices based on the number of phonemes represented by the encoded text sequence and the number of reference frames in the reference mel-frequency spectrogram sequence. Here, the step of learning the interval representation is based on each grid matrix mapped from the start and end boundaries, and the step of learning the auxiliary attention context is based on each grid matrix mapped from the start and end boundaries. The step of upsampling the sequence representation to the upsampled output may include determining the product of the interval representation matrix and the sequence representation, determining the Einstein sum (einsum) of the interval representation matrix and the auxiliary attention context representation, and summing the product of the interval representation matrix and the sequence representation and the projection of einsum to generate the upsampled output.

[0012] In some implementations, the operation includes receiving training data that includes a reference audio signal and a corresponding input text sequence, where the reference audio signal includes spoken utterances and the input text sequence corresponds to a transcript of the reference audio signal; encoding the reference audio signal into a variational embedding using a residual encoder, where the variational embedding disentangles style / rhythm information from the reference audio signal; and encoding the input text sequence into an encoded text sequence using a text encoder. In some examples, the residual encoder includes a global variational autoencoder (VAE). In these examples, encoding the reference audio signal into a variational embedding includes sampling a reference mel-frequency spectrogram sequence from the reference audio signal and encoding the reference mel-frequency spectrogram sequence into a variational embedding using the global VAE. Optionally, the residual encoder may include a phoneme-level fine-grained variational autoencoder (VAE). Here, encoding the reference audio signal into a variational embedding includes sampling a reference mel-frequency spectrogram sequence from the reference audio signal, aligning the reference mel-frequency spectrogram sequence with each phoneme in a sequence of phonemes extracted from the input text sequence, and encoding a sequence of phoneme-level variational embeddings using the phoneme-level fine-grained VAE and based on aligning the reference mel-frequency spectrogram sequence with each phoneme in the sequence of phonemes.

[0013] The residual encoder includes a stack of lightweight convolutional (LConv) blocks, where each LConv block in the stack of LConv blocks includes a gated linear unit (GLU) layer, an LConv layer configured to receive the output of the GLU layer, a residual connection configured to concatenate the output of the LConv layer with the input to the GLU layer, and a final feedforward layer configured to receive as input the residual connection that concatenates the output of the LConv layer with the input to the GLU layer. In some implementations, this operation further includes concatenating an encoded text sequence, a variational embedding, and a reference speaker embedding representing the identification of the reference speaker who uttered the reference audio signal, and generating a sequence representation based on a duration modeling network that receives as input the concatenation of the encoded text sequence, the variational embedding, and the reference speaker embedding. In some examples, the input text sequence includes a sequence of phonemes. In these examples, encoding the input text sequence into an encoded text sequence includes receiving, from a phoneme lookup table, the respective embeddings of each phoneme in the sequence of phonemes, processing each embedding using the encoder prenet neural network of the text encoder to generate the respective transformed embeddings for each phoneme in the sequence of phonemes, processing the respective transformed embeddings using a bank of convolutional blocks to generate a convolutional output, and processing the convolutional output using a stack of self-attention blocks to generate the encoded text sequence. Optionally, each self-attention block in the stack of self-attention blocks includes the same lightweight convolutional (LConv) block. Each self-attention block in the stack of self-attention blocks includes the same transformer block.

[0014] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the following description. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims.

Brief Description of the Drawings

[0015]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Best Mode for Carrying Out the Invention

[0016] Like reference numerals in the various drawings indicate like elements.

[0017] Realistic human speech synthesis is an open problem in that there are an infinite number of valid speech realizations for the same text input. End-to-end neural network-based approaches have advanced to the point of matching human performance for short utterances such as those of an assistant, but neural network models can be considered more difficult to interpret and control compared to conventional models that include multiple processing steps acting on refined linguistic or phonetic representations. Sources of variation in speech include intonation, stress, rhythm, prosodic characteristics of style, as well as speaker and channel characteristics. The prosodic features of spoken utterances convey linguistic, semantic, and affective meaning beyond what is present in the lexical representation (e.g., the transcript of the spoken utterance).

[0018] A neural network is a machine learning model that uses one or more layers of non-linear units to predict an output for a received input. For example, an end-to-end neural network-based text-to-speech (TTS) model can convert input text into output speech. A neural network TTS model offers the possibility of robustly synthesizing speech by predicting linguistic elements corresponding to prosody not provided by the text input. As a result, it is possible to generate synthetic speech that sounds real rather than monotonous in many applications such as audiobook narration, newsreaders, voice design software, and conversational assistants.

[0019] Many neural end-to-end TTS models utilize autoregressive models that predict the current value based on previous values. For example, many autoregressive models are based on recurrent neural networks that use some or all of the internal state of the network from the previous time step when calculating the output at the current time step. Examples of recurrent neural networks are long short-term (LSTM) neural networks that include one or more LSTM memory blocks. Each LSTM memory block can include one or more cells each containing an input gate, a forget gate, and an output gate that allows the cell to remember the previous state of the cell for use, for example, in generating the current activation or for providing to other components of the LSTM neural network.

[0020] Autoregressive TTS models can synthesize text and generate very natural-sounding speech outputs, but architectures through a series of one-directional LSTM-based decoder blocks with soft attention are inherently less efficient for both training and inference when implemented on state-of-the-art parallel hardware compared to a fully feed-forward architecture. Further, since autoregressive models are trained via teacher forcing by applying ground truth labels at each time step, autoregressive models further tend to have a mismatch between training and when the trained model is applied during inference. Combined with the soft attention mechanism, these mismatches can lead to degraded synthetic speech outputs such as synthetic speech that exhibits robustness errors such as mumbling, early cutoffs, word repetitions, and word skips. The degradation in the quality of the synthetic speech in autoregressive TTS models can further worsen as the size of the synthetic text increases.

[0021] To mitigate the aforementioned drawbacks of autoregressive-based TTS models, embodiments of the present specification are directed to non-autoregressive neural TTS models enhanced with a variational autoencoder (VAE)-based residual encoder. As will become apparent, the VAE-based residual encoder can disentangle the latent representation / state from a reference audio signal that conveys residual information such as style / rhythm information that cannot be represented by the input text to be synthesized (e.g., a phoneme sequence), or the speaker identifier (ID) of the speaker who spoke the reference audio signal. That is, the latent representation enables the output synthetic speech generated by the TTS model to sound like the reference audio signal input to the residual encoder.

[0022] The non-autoregressive neural TTS model enhanced with a VAE-based residual encoder provides a controllable model for predicting the mel-spectrogram information of an input text utterance (e.g., a predicted sequence of mel-frequency spectrograms), and at the same time effectively controls the rhythm / style represented in the mel-spectrogram information. For example, to represent the intended rhythm / style for synthesizing an input text utterance into expressive speech, a selected variational embedding learned by the VAE-based residual encoder is used, and the spectrogram decoder of the TTS model predicts the mel-frequency spectrogram of the input text utterance and provides it as an input to a synthesizer (e.g., a waveform synthesizer or a vocoder network) for converting the mel-frequency spectrogram into a time-domain audio waveform representing the synthetic speech with the intended rhythm / style. As will become apparent, since the non-autoregressive TTS model is trained on a sample input text sequence and the corresponding reference mel-frequency spectrogram sequence of human speech only, the trained TTS model can convert the input text utterance into a sequence of mel-frequency spectrograms with the intended rhythm / style conveyed by the learned prior variational embedding.

[0023] FIG. 1 shows an exemplary system 100 for training a deep neural network 200 enhanced with a VAE-based residual encoder 180 to provide a non-autoregressive neural TTS model (or simply “TTS model”) 300, and for predicting a spectrogram (i.e., a mel-frequency spectrogram sequence) 302 of a text utterance 320 using the TTS model 300. The system 100 includes a computing system 120 having data processing hardware 122 and memory hardware 124 that communicates with the data processing hardware 122 and stores instructions for causing the data processing hardware 122 to perform operations. In some implementations, the computing system 120 (e.g., the data processing hardware 122) or the user computing device 10 that executes the trained TTS model 300 provides the predicted mel-frequency spectrogram 302 predicted by the TTS model 300 from the input text utterance 320 to a synthesizer 155 to convert it into a time-domain audio waveform representing a synthetic speech 152 that can be aurally output as an oral representation of the input text utterance 320. The time-domain audio waveform includes an audio waveform that defines the amplitude of the audio signal over time. The synthesizer 155 can be separately trained and conditioned based on the mel-frequency spectrogram to convert it into a time-domain audio waveform.

[0024] The mel-frequency spectrogram includes a frequency-domain representation of the sound. The mel-frequency spectrogram emphasizes the low frequencies that are important for speech intelligibility while de-emphasizing the high frequencies that are dominated by fricatives and other noise bursts and generally do not need to be modeled with high fidelity. The synthesizer 155 can include a vocoder neural network that is configured to receive the mel-frequency spectrogram and generate an audio output sample (e.g., a time-domain audio waveform) based on the mel-frequency spectrogram. For example, the vocoder network 155 can be based on the parallel feed-forward neural network described in van den Oord, Parallel WaveNet: Fast High-Fidelity Speech Synthesis, available at https: / / arxiv.org / pdf / 1711.10433.pdf, which is incorporated herein by reference. Alternatively, the vocoder network 155 can be an autoregressive neural network. The synthesizer 155 can include a waveform synthesizer such as a Griffin-Lim synthesizer or a trainable spectrogram / waveform inverter. The choice of the synthesizer 155 does not affect the rhythm / style resulting in the synthesized speech 152 and, in fact, only affects the audio fidelity of the synthesized speech 152.

[0025] For input text utterance 320, since there is no way to convey the context, semantics, and pragmatics that guide the desired prosody / style of the synthesized speech 152, the TTS model 300 can apply the variational embedding 220 as a latent variable that specifies the intended prosody / style to predict the mel-frequency spectrogram 302 of the text utterance 320 that conveys the intended prosody / style specified by the variational embedding 220. In some examples, the computing system 120 implements the TTS model 300. Here, the user can access the TTS model 300 through the user computing device 10 and provide the input text utterance 320 to the TTS model 300 to synthesize into an expressive speech 152 having the intended prosody / style specified by the variational embedding 220. The variational embedding 220 can be selected by the user and correspond to a previous variational embedding 220 sampled from the residual encoder 180. The variational embedding 220 can be a speaker-specific variational embedding 220 that the user can select by providing a speaker identifier (ID) that identifies the speaker who speaks using the intended prosody / style (i.e., through an interface executed on the user computing device 10). Here, each speaker ID can be mapped to a respective speaker-specific variational embedding 220 pre-learned by the residual encoder 180. Additionally or alternatively, the user can provide an input that specifies a particular vertical associated with each prosody / style. Here, different verticals (e.g., news caster, sports caster, etc.) can each be mapped to a respective variational embedding 220 pre-learned by the residual encoder 180 that conveys the respective prosody / style associated with that vertical. In these examples, the synthesizer 155 can reside on the computing system 120 or the user computing device 10. When the synthesizer 155 resides on the computing system 120, the computing system 120 can send the time-domain audio waveform representing the synthesized speech 152 to the user computing device 10 for audible playback.In other examples, the user computing device 10 implements the TTS model 300. The computing system may include a distributed system (e.g., a cloud computing environment).

[0026] In some implementations, the deep neural network 200 is trained on a large set of reference audio signals 201. Each reference audio signal 201 may include a spoken utterance of human speech recorded by a microphone and having a prosody / style representation. During training, the deep neural network 200 may receive multiple reference audio signals 201 with different prosodies / styles for the same spoken utterance (i.e., the same utterance can be spoken in multiple different ways). Here, the reference audio signal 201 is a variable-length signal with different lengths of the spoken utterance even though the content is the same. The deep neural network 200 may also receive multiple sets of reference audio signals 201, each set including reference audio signals 201 for utterances that have the same prosody / style spoken by the same respective speaker but convey different language contents. The deep neural network 200 enhanced using the VAE-based residual encoder 180 is configured to encode / compress the prosody / style representation associated with each reference audio signal 201 into a corresponding variational embedding 220. The variational embedding 220 may include a fixed-length variational embedding 220. The deep neural network 200 stores each variational embedding 220 together with (e.g., on the memory hardware 124 of the computing system 120) a corresponding speaker embedding y that represents the speaker identity 205 (FIG. 2) of the reference speaker who uttered the reference audio signal 201 associated with the variational embedding 220 in storage 185. The variational embedding 220 may be a speaker-wise variational embedding that includes an aggregate (e.g., an average) of multiple variational embeddings 220 encoded by the residual encoder 180 from reference audio signals 201 spoken by the same speaker. s and can be stored in storage 185. The variational embedding 220 may be a speaker-wise variational embedding that includes an aggregate (e.g., an average) of multiple variational embeddings 220 encoded by the residual encoder 180 from reference audio signals 201 spoken by the same speaker.

[0027] During inference, computing system 120 or user computing device 10 may use trained TTS model 300 to predict the mel-frequency spectrogram sequence 302 of text utterance 320. TTS model 300 may select variational embedding 220 representing the intended rhythm / style of text utterance 320 from storage 185. Here, variational embedding 220 may correspond to a previous variational embedding 220 sampled from VAE-based residual encoder 180. TTS model 300 may use the selected variational embedding 220 to predict the mel-frequency spectrogram sequence 302 of text utterance 320. In the illustrated example, synthesizer 155 uses the predicted mel-frequency spectrogram sequence 302 to generate synthetic speech 152 having the intended rhythm / style specified by variational embedding 220.

[0028] In a non-limiting example, an individual may train the deep neural network 200 to learn a speaker-specific variational embedding that conveys prosody / style expressions associated with a particular speaker. For example, Brian McCullough, the host of the Techmeme Ride Home podcast, may train the deep neural network 200 with respect to the reference audio signal 201 that includes previous episodes of the podcast, along with an input text sequence 206 corresponding to the transcript of the reference audio signal 201. During training, the VAE-based residual encoder 180 may learn a speaker-specific variational embedding 220 that represents Brian's prosody / style of narrating the Ride Home podcast. Brian may then apply this speaker-specific variational embedding 220 for use by the trained TTS model 300 (executed on the computing system 120 or the user computing device 10) to predict a sequence of mel-frequency spectrograms 302 for a text utterance 320 corresponding to the transcript of a new episode of the Ride Home podcast. The predicted sequence of mel-frequency spectrograms 302 may be provided as input to a synthesizer 155 to generate a synthetic speech 152 having Brian's unique prosody / style specified by the speaker-specific variational embedding 220. That is, the resulting synthetic speech 152 may sound exactly like Brian's voice and may have Brian's prosody / style for narrating episodes of the Ride Home podcast. Thus, to broadcast a new episode, Brian may simply provide the transcript of the episode and use the trained TTS model 300 to generate the synthetic speech 152 that may be streamed to loyal listeners of the Ride Home podcast (also called the Mutant Podcast Army).

[0029] Figure 2 shows a non-autoregressive neural network (e.g., the deep neural network of FIG. 1) 200 for training the non-autoregressive TTS model 300. The deep neural network 200 includes a VAE-based residual encoder 180, a text encoder 210, a duration model network 230, and a spectrogram decoder 260. The deep neural network 200 can be trained based on training data including a plurality of reference audio signals 201 and corresponding input text sequences 206. Each reference audio signal 201 includes a spoken utterance of human speech, and the corresponding input text sequence 206 corresponds to a transcription of the reference audio signal 201. In the illustrated example, the VAE-based residual encoder 180 is configured to encode the reference audio signal 201 into a variational embedding 220 (z). Specifically, the residual encoder 180 receives a reference mel-frequency spectrogram sequence 202 sampled from the reference audio signal 201, and encodes the reference mel-frequency spectrogram sequence 202 into the variational embedding 220, whereby the variational embedding 220 unlocks style / rhythm information from the reference audio signal 201 corresponding to the spoken utterance of human speech. Thus, the variational embedding 220 corresponds to the latent state of the reference speaker, such as the rhythm, emotion, and / or emotions and intentions contributing to the speaking style of the reference speaker. As used herein, the variational embedding 220 includes both style information and rhythm information. In some examples, the variational embedding 220 includes a numerical vector having a capacity represented by the number of bits within the variational embedding 220. The reference mel-frequency spectrogram sequence 202 sampled from the reference audio signal 201 may have a length L R and dimension D R The reference mel-frequency spectrogram sequence 202, as used herein, includes a plurality of fixed-length reference mel-frequency spectrogram frames sampled / extracted from the reference audio signal 201. Each reference mel-frequency spectrogram frame may include a duration of 5 milliseconds.

[0030] The VAE-based residual encoder 180 corresponds to a posterior network that enables unsupervised learning of the latent representation of the speech style (i.e., variational embedding (z)). When learning variational embeddings using a VAE network, favorable properties of entanglement resolution, scaling, and combination for simplifying style control are obtained compared to heuristic-based systems. Here, the residual encoder 180 includes a projection layer with a rectified linear unit (ReLU) 410 and a stack of lightweight convolutional (LConv) blocks 420 with multi-head attention, including a phoneme-level fine-grained VAE network. The phoneme-level fine-grained VAE network 180 is configured to encode spectrogram frames from the reference mel-frequency spectrogram sequence 202 associated with each phoneme in the input text sequence 206 into respective phoneme-level variational embeddings 220. More specifically, the phoneme-level fine-grained VAE network can align the reference mel-frequency spectrogram sequence 202 with each phoneme in the sequence of phonemes extracted from the input text sequence 206 and can encode a sequence of phoneme-level variational embeddings 220. Thus, each phoneme-level variational embedding 220 in the sequence of phoneme-level variational embeddings 220 encoded by the phoneme-level fine-grained VAE network 180 encodes a respective subset of one or more spectrogram frames from the reference mel-frequency spectrogram sequence 202 that includes the respective phoneme in the sequence of phonemes extracted from the input text sequence 206. The phoneme-level fine-grained VAE network 180 first aligns the reference mel-frequency spectrogram sequence 202 with the speaker embedding y representing the speaker who spoke the utterance associated with the reference mel-frequency spectrogram sequence 202 sand can be concatenated with a sine wave position embedding 214 indicating the phoneme position information for each phoneme in the sequence of phonemes extracted from the input text sequence 206. In some examples, the residual encoder 180 includes a position embedding (not shown) instead of the sine wave position embedding 214. Each sine wave position embedding 214 can include a fixed-length vector containing information about the specific position of each phoneme in the sequence of phonemes extracted from the input text sequence 206. Subsequently, this concatenation is applied to a stack of five 8-head 17×1 LConv blocks 420 to calculate attention using the layer-normalized encoded text sequence 219 output from the text encoder 210.

[0031] FIG. 5 shows a schematic diagram 500 of an exemplary LConv block (e.g., a stack of LConv blocks 420) having a gated linear unit (GLU) layer 502, an LConv layer 504 configured to receive the output of the GLU layer 502, and a feed-forward (FF) layer 506. The exemplary LConv block 500 also includes a first residual connection 508 (e.g., a first connector 508) configured to concatenate the output of the LConv layer 504 with the input to the GLU layer 502. The FF layer 506 is configured to receive as input the first residual connection 508 that concatenates the output of the LConv layer 504 with the input to the GLU layer 502. The exemplary LConv block 500 also includes a second residual connection 510 (e.g., a second connector 510) configured to concatenate the output of the FF layer 506 with the first residual connection 508. The exemplary LConv block 500 can perform FF mixing in the FF layer 506 using the structure ReLU(W1X + b1)W2 + b2, where W1 increases the dimension by a factor of four.

[0032] In another implementation form, the VAE-based residual encoder 180 may include a global VAE network that can employ a non-autoregressive deep neural network 200 instead of the phoneme-level fine-grained VAE network shown in FIG. 2. The global VAE network encodes the reference mel-frequency spectrogram sequence 202 into a global variational embedding 220 at the utterance level. Here, the global VAE network includes two stacks of lightweight convolutional (LConv) blocks each having multi-head attention. Each LConv block in the first stack and the second stack of the global VAE network 180 may include eight heads. In some examples, the first stack of LConv blocks includes three 17×1 LConv blocks, and the second stack of LConv blocks following the first stack includes five 17×1 LConv blocks interleaved using a 3×1 convolution. This configuration of the two stacks of LConv blocks enables the global VAE network 180 to continuously downsample the latent representation before applying global average pooling to obtain the final global variational embedding 220. The projection layer may project the dimension of the global variational embedding 220 output from the second stack of LConv blocks. For example, the global variational embedding 220 output from the second stack of LConv blocks may have eight dimensions, and the projection layer may project that dimension to 32.

[0033] Continuing to refer to FIG. 2, the text encoder 210 encodes the input text sequence 206 into a text encoding sequence 219. The text encoding sequence 219 includes an encoded representation of a sequence of speech units (e.g., phonemes) extracted from the input text sequence 206. The input text sequence 206 can include words each having one or more phonemes, silence at all word boundaries, and punctuation. Thus, the input text sequence 206 includes a sequence of phonemes, and the text encoder 210 can receive, from the token embedding lookup table 207, a respective token embedding for each phoneme in the sequence of phonemes. Here, each token embedding includes a phoneme embedding. However, in other examples, the token embedding lookup table 207 can obtain token embeddings for other types of speech inputs associated with the input text sequence 206, such as, but not limited to, sub - phonemes (e.g., senomes), graphemes, word pieces, or words in the utterance, instead of phonemes. After receiving the respective token embedding for each phoneme in the sequence of phonemes, the text encoder 210 processes each token embedding and uses the encoder pre - net neural network 208 to generate a respective transformed embedding 209 for each phoneme. Thereafter, a bank of convolutional (Conv) blocks 212 can process the respective transformed embeddings 209 to generate a convolutional output 213. In some examples, the bank of Conv blocks 212 includes three identical 5×1 Conv blocks. FIG. 4 shows a schematic diagram 400 of an exemplary Conv block having a Conv layer 402, a batch normalization layer 404, and a dropout layer 406. During training, the batch normalization layer 404 can apply batch normalization to reduce internal covariate shift. The dropout layer 406 can reduce overfitting. Finally, a stack of self - attention blocks 218 processes the convolutional output 213 to generate the encoded text sequence 219. In the illustrated example, the stack of self - attention blocks 218 includes six transformer blocks.In other examples, the self-attention block 218 may include an LConv block instead of a transformer block.

[0034] In particular, each convolutional output 213 flows through the stack of self-attention blocks 218 simultaneously, and the stack of self-attention blocks 218 has no knowledge of the position / order of each phoneme in the input text utterance 206. Thus, in some examples, the sine wave position embedding 214 is combined with the convolutional output 213 to insert the necessary position information indicating the order of each phoneme in the input text sequence 206. In other examples, an encoded position embedding is used instead of the sine wave position embedding 214. In contrast, an autoregressive encoder incorporating a recurrent neural network (RNN) inherently considers the order of each phoneme because each phoneme is parsed sequentially from the input text sequence. However, the text encoder 210 integrating a stack of self-attention blocks 218 using multi-head self-attention avoids the recurrence of the autoregressive encoder and significantly shortens the training time, and theoretically captures longer dependencies in the input text sequence 206.

[0035] Continuing to refer to FIG. 2, the concatenator 222 concatenates the variational embedding 220 from the residual encoder 180, the encoded text sequence 219 output from the text encoder 210, and the reference speaker embedding y representing the speaker identity 205 of the reference speaker who uttered the reference audio signal. sConnect it to concatenation 224. The duration model network 230 is configured to receive the concatenation 224 and generate an upsampled output 258 that specifies the number of frames of the encoded text sequence 219 from the concatenation 224. In some implementations, the duration model network 230 includes a stack of self-attention blocks 232, followed by two independent small convolutional blocks 234, 238, and a projection with a softplus activation 236. In the illustrated example, the stack of self-attention blocks includes four 3×1 LConv blocks 232. As described above with reference to FIG. 5, each LConv block in the stack of self-attention blocks 232 may include a GLU unit 502, an LConv layer 504, and an FF layer 506 with the remaining connections. The stack of self-attention blocks 232 is based on the concatenation 224 of the encoded text sequence 219, the variational embedding 220, and the reference speaker embedding y s to generate a sequence representation V. Here, the sequence representation V represents a sequence of M×1 column vectors (e.g., V = {v l , …, v k}).

[0036] In some implementations, the first convolutional block 234 generates an output 235 and the second convolutional block 238 generates an output 239 from the sequence representation V. The convolutional blocks 234, 238 can each include a 3×1 Conv block with a kernel width of 3 and an output dimension of 3. The projection by the softplus activation 236 can predict the phoneme durations 240 for each phoneme represented by the encoded text sequence 219 (e.g., {d l , …, d k}). Here, the softplus activation 236 receives the sequence representation V as input to predict the phoneme durations 240. The duration model network 230 is the predicted phoneme durations 240 and,

[0037]

Number

[0038] Calculate the global phoneme duration loss 241 (e.g., L1 loss term 241) between the target average duration 245 represented by and the above formula, where L dur represents the global phoneme duration loss 241 (e.g., L1 loss term 241), K represents the total number of phonemes (e.g., tokens) of the input text sequence 206, and d k represents the phoneme duration 240 of a specific phoneme k from the total number of phonemes K, and T represents the total target frame duration from the reference mel-frequency spectrogram sequence 202.

[0039] In some examples, the training of the TTS model 300 is based on the global phoneme duration loss 241 from Equation 1. The individual target durations for each phoneme are unknown, and thus the duration model network 230 determines the target average duration 245 based on the ratio of the total frame duration of Ts from the entire reference mel-frequency spectrogram sequence 202 to the total number of phonemes (e.g., tokens) of Ks in the input text sequence 206. That is, the target average duration 245 is the average duration of all phonemes that use the reference mel-frequency spectrogram sequence 202 and the input text sequence 206. Next, the L1 loss term 241 is determined between the predicted phoneme duration 240 and the target average duration 245 determined using the reference mel-frequency spectrogram sequence 202 and the input text sequence 206. Thus, the duration model network 230 learns to predict the phoneme duration 240 in a teacherless manner without using the teacher-provided phoneme duration labels from an external aligner. The external aligner can provide a reasonable alignment between phonemes and mel-spectrum frames, but rounding of the phoneme durations by a length adjuster is required to upsample the phonemes in the input text sequence 206 according to their durations, which leads to possible rounding errors remaining. In some cases, using the teacher-provided duration labels from the external aligner during training and the predicted durations during inference results in a mismatch in phoneme durations between the training of the TTS model 300 (FIG. 2) and the inference of the TTS model 300 (FIG. 3). Further, such rounding operations are not differentiable, and thus the error gradients cannot propagate through the duration model network 230.

[0040] The duration model network 230 determines token boundaries from the predicted phoneme duration 240

[0041] [Number]

[0042] It includes a matrix generator 242 for defining. The matrix generator 242 determines token boundaries (e.g., phoneme boundaries) from the predicted phoneme durations 240 as follows.

[0043]

Number

[0044] e k = s k + d k (3)

[0045] In Equation 2, s k represents the start of the token boundary of a specific phoneme k (also referred to as the start boundary s k in this specification). In Equation 3, e k represents the end of the token boundary of a specific phoneme k (also referred to as the end boundary e k in this specification). The matrix generator 242 maps the start boundary s k and end boundary e k from Equations 2 and 3 to two token boundary grid matrices S and E as follows. S tk = t - s k (4) E tk = e k - t (5)

[0046] Equation 4 maps each start boundary s k to the S tk grid matrix and gives the distance to the start boundary s k of token k at time t. Equation 5 maps each end boundary e k to the E tk grid matrix and gives the distance to the end boundary e k of token k at time t. The matrix generator 242 generates the start boundary s k and end boundary e k (collectively also referred to as token boundaries s k , e k ), and the token boundaries s k , e kAre respectively mapped to the start token boundary grid matrix S and the end token boundary grid matrix E (collectively referred to as the grid matrix 243). Here, the size of the grid matrix 243 is T×K, where K represents the number of phonemes in the input text sequence 206, and T represents the number of frames and the total frame duration of the reference mel-frequency spectrogram sequence 202. The matrix generator 242, based on the number of phonemes represented by the encoded text sequence 219 and the number of reference frames in the reference mel-frequency spectrogram sequence 202, for each phoneme in the phoneme sequence, each start boundary s k and end boundary e k Can be respectively mapped to each grid matrix 243 (for example, the start token boundary grid matrix S and the end token boundary grid matrix E).

[0047] The duration model network 230 includes a first function 244 for learning the interval representation matrix W. In some implementations, the first function 244 includes two projection layers with bias and Swish activation. In the illustrated example, both projections of the first function 244 are projected and output using dimension 16 (i.e., P = 16). The first function 244 receives the concatenation 237 as input from the concatenator 222. The concatenator 222 concatenates the output 235 from the convolutional block 234 and the grid matrix 243 to generate the concatenation 237. The first function 244 projects the output 247 using the concatenation 237. Then, the first function 244 generates the interval representation matrix W (for example, a T×K attention matrix) from the projection with the softplus activation 246 of the output 247. The interval representation matrix W is based on each grid matrix 243 mapped from the start boundary s k and end boundary e k Can be learned as follows based on each mapped grid matrix 243. W = Softmax(MLP(S, E, Conv 1D(V))) (6)

[0048] Equation 6 uses the softmax function, a multi-layer perceptron-based (MLP) learnable function of the grid matrices 243 (e.g., the start token boundary grid matrix S and the end token boundary grid matrix E), and the sequence representation V to compute the interval representation W (e.g., the T×K attention matrix). The MLP learnable function includes a third projection layer of output dimension 1 that is fed into the softmax activation function, and Conv1D(V) includes a kernel width of 3, an output dimension of 8, batch normalization, and the Swish activation. Here, the (k, t)-th element of the grid matrix 243 gives the attention probability between the k-th token (e.g., phoneme) and the t-th frame, and W() is a learnable function that maps the grid matrix 243 and the sequence representation V using a small 1D convolutional layer.

[0049] The duration model network 230 can learn an auxiliary attention context tensor C (e.g., C = [C l ,..., C p ) using a second function 248 conditioned on the sequence representation V. Here, C p includes the T×K matrix from the auxiliary attention context tensor C. The auxiliary attention context tensor C can include auxiliary multi-head attention-like information for the spectrogram decoder 260. The concatenator 222 concatenates the grid matrix 243 and the output 239 to produce the concatenation 249. The second function 248 can include two projection layers with bias and the Swish activation. In the illustrated example, the projection of the second function 248 projects the output of dimension 2 (i.e., P = 2). The second function 248 receives the concatenation 249 as input to generate the auxiliary attention context tensor C based on each of the grid matrix 243 and the sequence representation V mapped from the start boundary s k , the end boundary e k as follows. C = MLP(S, E, Conv1D(V)) (7)

[0050] Equation 7 uses a multi-layer perceptron-based (MLP) learnable function of the grid matrix 243 (e.g., the start token boundary grid matrix S and the end token boundary grid matrix E) and the sequence representation V to calculate the auxiliary attention context tensor C. The auxiliary attention context tensor C can help smooth the optimization and converge the stochastic gradient descent (SGD). The duration model network 230 can upsample the sequence representation V using several frames to the output 258 (e.g., 0 = {o l , …, o T}). Here, the number of frames of the upsampled output 258 corresponds to the predicted length of the predicted mel-frequency spectrogram 302 determined by the predicted phoneme durations 240 of the corresponding input text sequence 206. Upsampling the sequence representation V to the upsampled output 258 includes determining the product 254 of the interval representation matrix W and the sequence representation V using the multiplier 253.

[0051] The duration model network 230 uses the einsum operator 255 to determine the Einstein sum (einsum) 256 of the interval representation matrix W and the auxiliary attention context tensor C. The projection 250 projects the einsum 256 to the projection output 252, and the projection output 252 is added to the product 254 in the adder 257 to generate the upsampled output 258. Here, the upsampled output 258 can be represented as follows. O = WV + [(W Θ C1)1 k ...(W Θ C p )1 k A (8) In the above equation, O represents the upsampled output 258, Θ represents element-wise multiplication, 1 k represents a K×1 column vector containing all elements equal to 1, and A is a P×M projection matrix.

[0052] Continuing to refer to FIG. 2, the spectrogram decoder 260 is configured to receive, as input, the upsampled output 258 of the duration model network 230 and to generate, as output, one or more predicted mel-frequency spectrogram sequences 302 for the input text sequence 206. The spectrogram decoder 260 may include a stack of multiple self-attention blocks 262, 262a - n with multi-head attention. In some examples, the spectrogram decoder 260 includes six 8-head 17×1 LConv blocks with a 0.1 dropout. The spectrogram decoder 260 may include more or fewer than six LConv blocks. As described above with reference to FIG. 5, each LConv block in the stack of self-attention blocks 232 may include a GLU unit 502, an LConv layer 504, and an FF layer 506 with the remaining connections. In other examples, each self-attention block 262 in the stack includes the same transformer block.

[0053] In some implementations, the spectrogram decoder 260 generates respective predicted mel-frequency spectrogram sequences 302, 302a-n as outputs from each self-attention block 262 in the stack of self-attention blocks 262, 262a-n. The network 200 can be trained such that the number of frames of each predicted mel-frequency spectrogram sequence 302 is equal to the number of frames of the reference mel-frequency spectrogram sequence 202 input to the residual encoder 180. In the illustrated example, each self-attention block 262a-n is paired with a corresponding projection layer 264a-n that projects the output 263 from the self-attention block 262 to generate respective predicted mel-frequency spectrogram sequences 302a-n having dimensions that match the dimensions of the reference mel-frequency spectrogram sequence 202. In some examples, the projection layer 264 projects the predicted mel-frequency spectrogram sequence 302 of 128 bins. By predicting a plurality of mel-frequency spectrogram sequences 302a-n, the non-autoregressive neural network 200 can be trained using a soft dynamic time warping (soft DTW) loss. That is, since the predicted mel-frequency spectrogram sequences 302 can have a different length (e.g., number of frames) than the reference mel-frequency spectrogram sequence 202, the spectrogram decoder 260 cannot determine a normal Laplace loss. Instead, the spectrogram decoder 260 determines a soft DTW loss between the reference mel-frequency spectrogram sequence 202 and the predicted mel-frequency spectrogram sequences 302, which can include different lengths. In particular, for each respective predicted mel-frequency spectrogram sequence 302a-n, the spectrogram decoder 260 determines respective spectrogram losses 270, 270a-n based on the corresponding predicted mel-frequency spectrogram sequence 302 and the reference mel-frequency spectrogram sequence 202. Each spectrogram loss 270 can include a soft DTW loss term determined recursively as follows.

[0054]

Number

[0055] In Equation 9, r i,j represents the distance between the reference mel-frequency spectrogram sequence frames from 1 to i and the predicted mel-frequency spectrogram sequence frames from 1 to j with the best alignment. Here, min γ includes a generalized minimum operation using the smoothing parameter γ, warp includes a warp penalty, and x i and x j are the reference mel-frequency spectrogram frame and the predicted mel-frequency spectrogram sequence frame at times i and j, respectively.

[0056] The soft DTW loss term recursion is computationally expensive, O(T 2can include the complexity of (0), a diagonal band width fixed at 60, a warp penalty of 128, and a smoothing parameter γ of 0.05. For example, the first spectrogram loss 270a can be determined based on the first predicted mel-frequency spectrogram sequence 302a and the reference mel-frequency spectrogram sequence 202. Here, the first predicted mel-frequency spectrogram sequence 302a and the reference mel-frequency spectrogram sequence 202 may have the same length or different lengths. The second spectrogram loss 270b can be determined based on the first predicted mel-frequency spectrogram sequence 302a and the reference mel-frequency spectrogram sequence 202, and the same applies until all of the respective spectrogram losses 270a to n are repeatedly determined for the predicted mel-frequency spectrogram sequences 302a to n. The spectrogram loss 270 including the soft DTW loss term can be aggregated to generate the final soft DTW loss 280. The final soft DTW loss 280 can correspond to the iterative soft DTW loss term 280. The final soft DTW loss 280 can be determined from any combination of predicted mel-frequency spectrogram sequences 302 and reference mel-frequency spectrogram sequences 202 of the same length and / or different lengths as follows.

[0057] [Number]

[0058] In Equation 10, L includes the final soft DTW loss 280,

[0059] [Number]

[0060] includes the soft DTW L1 spectrogram reconstruction loss for the l-th iteration in the spectrogram decoder, L dur includes the L1 loss of the average duration, D KLincludes the KL divergence between before and after the residual encoder. The training of the deep neural network 200 aims to minimize the final soft DTW loss 280 in order to reduce the difference in phoneme durations between the predicted mel-frequency spectrogram sequence 302 and the reference mel-frequency spectrogram sequence 202. By minimizing the final soft DTW loss 280, the trained TTS model 300 can generate a predicted mel-frequency spectrogram sequence 302 that includes the intended rhythm / style based on the reference mel-frequency spectrogram sequence 202. Aggregating the spectrogram losses 270a~n may include summing the spectrogram losses 270a~n to obtain the final soft DTW loss 280. Optionally, aggregating the spectrogram losses 270 may include averaging the spectrogram losses 270.

[0061] The deep neural network 200 can be trained such that the number of frames in each predicted mel-frequency spectrogram sequence 302 is equal to the number of frames in the reference mel-frequency spectrogram sequence 202 input to the residual encoder 180. Further, the deep neural network 200 is trained such that the data associated with the reference mel-frequency spectrogram sequence 202 and the predicted mel-frequency spectrogram sequence 302 substantially match each other. The predicted mel-frequency spectrogram sequence 302 may implicitly provide the rhythm / style representation of the reference audio signal 201.

[0062] FIG. 3 shows an example of a non-autoregressive TTS model 300 trained by the non-autoregressive deep neural network 200 of FIG. 2. Specifically, FIG. 3 shows a TTS model 300 that uses a variational embedding 220 selected to predict a mel-frequency spectrogram sequence 302 of an input text utterance 320, whereby the selected variational embedding 220 represents the intended prosody / style of the text utterance 320. During inference, the trained TTS model 300 is executed on the computing system 120 or the user computing device 10 and may use the selected variational embedding 220 to predict a mel-frequency spectrogram sequence 302 corresponding to the input text utterance 320. Here, the TTS model 300 selects a variational embedding 220 representing the intended prosody / style of the text utterance 320 from the storage 185. In some examples, the user provides a user input instruction indicating a selection of an intended prosody / style that the user wishes to convey for the resulting synthetic speech 152 for the text utterance 320, and the TTS model 300 selects an appropriate variational embedding 220 representing the intended prosody / style from the data storage 185. In these examples, the intended prosody / style may be selected by the user by indicating a speaker identity 205 associated with a particular speaker who speaks in the intended prosody / style and / or by specifying a particular prosodic vertical (e.g., news caster, sports caster, etc.) corresponding to the intended prosody / style. The selected variational embedding 220 may correspond to a previous variational embedding 220 sampled from the VAE-based residual encoder 180. The trained TTS model 300 uses the synthesizer 155 to generate a synthetic speech 152 having the intended prosody / style (e.g., the selected variational embedding 220) for each input text utterance 320. That is, the selected variational embedding 220 may include the intended prosody / style (e.g., news caster, sports caster, etc.) stored in the residual encoder 180.The selected variational embedding 220 conveys the intended prosody / style via the synthetic speech 152 of the input text utterance 320.

[0063] In an additional implementation form, the trained TTS model 300 uses the residual encoder 180 during inference to extract / predict on-the-fly the variational embedding 220 used when predicting the mel-frequency spectrogram sequence 302 of the input text utterance 320. For example, the residual encoder 180 may receive a reference audio signal 201 (FIG. 2) uttered by a human user conveying the intended prosody / style (e.g., "Please say it like this") and extract / predict the corresponding variational embedding 220 representing the intended prosody / style. Thereafter, the trained TTS model 300 may use the variational embedding 220 to effectively transfer the intended prosody / style conveyed by the reference audio signal 201 to the mel-frequency spectrogram sequence 302 predicted for the input text utterance 320. Thus, the input text utterance 320 synthesized into the expressive speech 152 and the reference audio signal 201 conveying the intended prosody / style transferred to the expressive speech 152 may contain different language contents.

[0064] Specifically, the text encoder 210 encodes the sequence of phonemes extracted from the text utterance 320 into the encoded text sequence 219. The text encoder 210 may receive respective token embeddings for each phoneme in the sequence of phonemes extracted from the text utterance 320 from the token embedding lookup table 207. After receiving the respective token embeddings for each phoneme in the sequence of phonemes extracted from the text utterance 320, the text encoder 210 uses the encoder pre-net neural network 208 to process the respective token embeddings to generate the respective transformed embeddings 209 for each phoneme. Thereafter, a bank of Conv blocks 212 (e.g., three identical 5×1 Conv blocks) processes the respective transformed embeddings 209 to generate the convolutional output 213. Finally, a stack of self-attention blocks 218 processes the convolutional output 213 to generate the encoded text sequence 219. In the illustrated example, the stack of self-attention blocks 218 includes six transformation blocks. In other examples, the self-attention block 218 may include LConv blocks instead of transformer blocks. Specifically, since each convolutional output 213 flows through the stack of self-attention blocks 218 simultaneously, the stack of self-attention blocks 218 does not have knowledge about the position / order of each phoneme in the input text utterance. Thus, in some examples, the sine wave position embedding 214 is combined with the convolutional output 213 to insert the necessary position information indicating the order of each phoneme in the input text sequence 206. In other examples, an encoded position embedding is used instead of the sine wave position embedding 214.

[0065] Continuing to refer to FIG. 3, the concatenator 222 concatenates the selected variational embedding 220, the encoded text sequence 219, and optionally the reference speaker embedding y s to generate the concatenation 224. Here, the reference speaker embedding y smay represent the speaker identity 205 of the reference speaker who uttered one or more reference audio signals 201 associated with the selected variational embedding 220, or the speaker identity 205 of another reference speaker having the voice characteristics conveyed in the resulting synthetic speech 152. The duration model network 230 is configured to decode the concatenation 224 of the encoded text sequence 219, the selected variational embedding 220, and the reference speaker embedding y in order to predict the phoneme durations 240 for each phoneme in the sequence of phonemes in the input text utterance 320 and to generate the upsampled output 258 of the duration model network 230. s is configured to.

[0066] The duration model network 230 is configured to generate an upsampled output 258 that specifies the number of frames of the encoded text sequence 219 from the concatenation 224. In some implementations, the duration model network 230 includes a stack of self-attention blocks 232, followed by two independent small convolutional blocks 234, 238, and a projection with a softplus activation 236. In the illustrated example, the stack of self-attention blocks includes four 3×1 LConv blocks 232. As described above with reference to FIG. 5, each LConv block within the stack of self-attention blocks 232 may include a GLU unit 502, an LConv layer 504, and an FF layer 506 with the remaining connections. The stack of self-attention blocks 232 generates a sequence representation V based on the concatenation 224 of the encoded text sequence 219, the variational embedding 220, and the reference speaker embedding y. s is generated.

[0067] In some implementations, the first convolutional block 234 generates an output 235, and the second convolutional block 238 generates an output 239 from the sequence representation V. The convolutional blocks 234, 238 can include 3×1 Conv blocks with a kernel width of 3 and an output dimension of 3. The duration model network 230 includes a projection with a softplus activation 236 to predict the per-phoneme phoneme durations 240 (e.g., {d l , …, d k}) represented by the encoded text sequence 219. Here, the softplus activation 236 receives the sequence representation V as input to predict the phoneme durations 240.

[0068] As described above with reference to FIG. 2, the duration model network 230 includes a matrix generator 242 for defining token boundaries

[0069]

Number

[0070] from the predicted phoneme durations 240. The duration model network 230 includes a first function 244 for learning the interval representation matrix W. In some implementations, the first function 244 includes two projection layers with bias and Swish activation. In the illustrated example, both projections of the first function 244 are projected and output using dimension 16 (i.e., P = 16). The first function 244 receives the concatenation 237 as input from the concatenator 222. The concatenator 222 concatenates the output 235 from the convolutional block 234 and the grid matrix 243 to generate the concatenation 237. Then, the first function 244 uses Equation 6 to generate the interval representation matrix W (e.g., a T×K attention matrix) from a projection with a softplus activation 246 of the output 247.

[0071] The duration model network 230 may learn an auxiliary attention context tensor C (e.g., C = [C l ,..., C p ) using a second function 248 conditioned on the sequence representation V. Here, C p is a T×K matrix from the auxiliary attention content tensor C. The auxiliary attention context tensor C may contain auxiliary multi-head attention-like information for the spectrogram decoder 260. The concatenator 222 concatenates the grid matrix 243 and the output 239 to generate a concatenation 249. The second function 248 may include two projection layers with biases and Swish activations. The second function 248 receives the concatenation 249 as input to generate the auxiliary attention context tensor C based on each grid matrix 243 and the sequence representation V mapped from the start boundary s k , end boundary e k using Equation 7.

[0072] The duration model network 230 may upsample the sequence representation V to an output 258 (e.g., 0 = {o l , …, o T}) using several frames. Here, the number of frames of the upsampled output is corresponding to the predicted length of the predicted mel-frequency spectrogram 302 determined by the predicted phoneme durations 240 of the corresponding input text sequence 206. Upsampling the sequence representation V to the upsampled output 258 is based on determining the product 254 of the interval representation W and the sequence representation V using the multiplier 253 and determining the einsum 256 of the interval representation matrix W and the auxiliary attention context sensor C using the einsum operator 255. The projection 250 projects the einsum 356 to the projected output 252, and the adder 257 sums the product 254 to generate the upsampled output 258 represented by Equation 8.

[0073] The spectrogram decoder 260 is configured to generate a predicted mel-frequency spectrogram sequence 302 of the text utterance 320 based on the upsampled output 258 of the duration model network 230 and the predicted phoneme durations 240. Here, the predicted mel-frequency spectrogram sequence 302 has the intended rhythm / style specified by the selected variational embedding 220. The predicted mel-frequency spectrogram sequence 302 of the text utterance 320 is based on the auxiliary attention context tensor C, the sequence representation V, the interval representation w, and the upsampled output 258 to the number of frames of the duration model network 230.

[0074] The spectrogram decoder 260 generates respective predicted mel-frequency spectrogram sequences 302 as the output from the last self-attention block 262 in the stack of self-attention blocks 262a - n. Here, each self-attention block 262 in the stack of self-attention blocks 262 of the spectrogram decoder 260 includes one of the same LConv block or the same transformer block. In some examples, the spectrogram decoder 260 includes six 8-head 17×1 LConv blocks with a 0.1 dropout. The output of the last feed-forward (FF) layer 506 (FIG. 5) for each self-attention block 262 is provided as the input to the subsequent self-attention block 262. That is, the GLU unit 502 (FIG. 5) and the LConv layer 504 (FIG. 5) of the first self-attention block 262a in the stack of self-attention blocks 262a - n process the output 238 from the duration model network 230 and the predicted phoneme durations 240, and the output from the last FF layer 506 of the first self-attention block 262a is provided as the input to the subsequent second self-attention block 262b in the stack of self-attention blocks 262. The output of the last FF layer of each self-attention block 262 is provided as the input to the subsequent self-attention block 262 until it reaches the last self-attention block 262n in the stack of self-attention blocks 262. The last self-attention block 262n (e.g., the sixth self-attention block 262) in the stack of self-attention blocks 262 is paired with a corresponding projection layer 264 that projects the output 263 from the last self-attention block 262 to generate respective predicted mel-frequency spectrogram sequences 302.

[0075] The predicted mel-frequency spectrogram sequence 302 generated by the spectrogram decoder 260 corresponds to the input text utterance 320 and conveys the intended prosody / style indicated by the selected variational embedding 220. The trained TTS model 300 provides the predicted mel-frequency spectrogram sequence 302 of the input text utterance 320 to the synthesizer 155 to convert it into a time-domain audio waveform representing the synthesized speech 152. The synthesized speech 152 can be aurally output as an oral representation of the input text utterance 320, including the intended prosody / style indicated by the selected variational embedding 220.

[0076] Figure 6 is a flowchart of an exemplary configuration of the operations of a computer-implemented method 600 for training a non-autoregressive text-to-speech (TTS) model. In operation 602, the method 600 includes the step of obtaining a sequence representation V of a sequence of encoded text sequences 219 concatenated with a variational embedding 220. In operation 604, the method 600 includes the step of predicting, using a duration model network, a phoneme duration 240 for each phoneme represented by the encoded text sequence 219 based on the sequence representation V. Based on the predicted phoneme duration 240, the method 600 includes, in operation 606, the step of learning an interval representation matrix W using a first function 244 conditioned on the sequence representation V. In operation 608, the method 600 includes the step of learning an auxiliary attention context representation C using a second function 248 conditioned on the sequence representation V.

[0077] In operation 610, method 600 includes the step of upsampling sequence representation V to upsampled output 258 that specifies the number of frames, using interval representation matrix W and auxiliary attention context representation C. In operation 612, method 600 includes the step of generating one or more predicted mel-frequency spectrogram sequences 302 for encoded text sequence 219 as output from spectrogram decoder 260 that includes a stack of one or more self-attention blocks 262, 262a - n, based on the upsampled output 258. In operation 614, method 600 includes the step of determining final spectrogram loss 280 based on one or more predicted mel-frequency spectrogram sequences 302 and reference mel-frequency spectrogram sequence 202. In operation 616, method 600 includes the step of training TTS model 300 based on final spectrogram loss 280.

[0078] FIG. 7 is a schematic diagram of an exemplary computing device 700 that may be used to implement the systems and methods described herein. Computing device 700 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown herein, their connections and relationships, and their functions, are intended only as examples and are not intended to limit implementations of the inventions described and / or claimed herein.

[0079] Computing device 700 includes a processor 710, a memory 720, a storage device 730, a high-speed interface / controller 740 connected to the memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connected to a low-speed bus 770 and the storage device 730. Each of the components 710, 720, 730, 740, 750, and 760 may be interconnected using various buses and may be mounted on a common motherboard or in other manners as desired. The processor 710 can process instructions for execution within the computing device 700, including instructions stored in the memory 720 or the storage device 730 for displaying graphical information for a graphical user interface (GUI) of an external input / output device such as a display 780 connected to the high-speed interface 740. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and memory types, as desired. Also, multiple computing devices 700 may be connected, with each device providing a portion of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).

[0080] Memory 720 stores information non - transiently within computing device 700. Memory 720 may be a computer - readable medium, a volatile memory unit, or a non - volatile memory unit. The non - transient memory 720 may be a physical device used to store temporarily or persistently a program (e.g., a sequence of instructions) or data (e.g., program state information) used by computing device 700. Examples of non - volatile memory include, but are not limited to, flash memory and read - only memory (ROM) / programmable read - only memory (PROM) / erasable programmable read - only memory (EPROM) / electrically erasable programmable read - only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase - change memory (PCM), and disks or tapes.

[0081] Storage device 730 can provide large - capacity storage to computing device 700. In some implementations, storage device 730 is a computer - readable medium. In various different implementations, storage device 730 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid - state memory device, or an array of devices including devices within a storage area network or other configuration. In additional implementations, a computer program product is specifically incorporated into an information carrier. The computer program product includes instructions that, when executed, implement one or more of the methods as described above. The information carrier is a computer - readable or machine - readable medium such as memory 720, storage device 730, or memory on processor 710.

[0082] The high-speed controller 740 manages operations that consume a large amount of the bandwidth of the computing device 700, while the low-speed controller 760 manages operations that consume less bandwidth. Such an assignment of duties is merely an example. In some implementations, the high-speed controller 740 is coupled to a high-speed expansion port 750 that can accept the memory 720, the display 780 (e.g., through a graphics processor or accelerator), and various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to the storage device 730 and the low-speed expansion port 790. The low-speed expansion port 790, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, etc., or a networking device such as a switch or a router, e.g., through a network adapter.

[0083] As shown in the drawings, the computing device 700 can be implemented in many different forms. For example, it can be implemented as a standard server 700a, or multiple times as a laptop computer 700b in a group of such servers 700a, or as part of a rack server system 700c.

[0084] Various implementations of the systems and techniques described herein can be realized in digital electronic circuits and / or optical circuits, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device, which can be either dedicated or general purpose.

[0085] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application", an "app", or a "program". Examples of applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and game applications.

[0086] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural languages and / or object-oriented programming languages and / or assembly language / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable media logic device (PLD)) used to provide machine instructions and / or data to a programmable processor that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0087] The processes and logic flows described herein can be implemented by one or more programmable processors, also called data processing hardware, which execute one or more computer programs to perform functions by operating on input data to generate output. The processes and logic flows can also be implemented by dedicated logic circuits, such as, for example, an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory, a random access memory, or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer also includes, or is operatively coupled to for transferring data to and / or receiving data from, one or more mass storage devices for storing data, such as, magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as, EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, dedicated logic circuits.

[0088] To provide interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, and optionally a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form, including acoustic, voice, or tactile input. Further, the computer can interact with the user by sending and receiving documents to and from the devices used by the user. For example, in response to a request received from a web browser, the computer can send a web page to a web browser on the user's client device.

[0089] Numerous implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are also within the scope of the appended claims.

Description of Reference Numerals

[0090] 10 User computing device 100 System 120 Computing system 122 Data processing hardware 124 Memory hardware 152 Synthetic voice 155 Synthesizer 155 Vocoder network 180 Residual encoder 180 Global VAE network 180 Phoneme-level fine-grained VAE network 185 Data storage 200 Deep neural network 200 Non-autoregressive neural network 201 Reference audio signal 202 Reference mel-frequency spectrogram sequence 205 Speaker identity 206 Input text sequence 207 Token embedding lookup table 208 Encoder pre-net neural network 209 Transformed embedding 210 Text encoder 212 Convolution (Conv) block 213 Convolution output 214 Sine wave position embedding 216 Position embedding 218 Self-attention block 219 Text encoding sequence 219 Encoded text sequence 220 Variational embedding 222 Connector 224 Concatenation 230 Duration model network 232 Self-attention block 234 Convolution block 234 First convolution block 235 Output 236 Softplus activation 236 Projection 237 Concatenation 238 Small convolution block 238 Second convolution block 239 Output 240 Phoneme duration 241 Global phoneme duration loss 241 L1 loss term 242 Matrix generator 243 Grid matrix 244 First function 245 Target average duration 246 Softplus activation 247 Output 248 Second function 249 Concatenation 250 Projection 252 Projected output 253 Multiplier 254 Product 255 Einsum operator 256 Einstein summation (einsum) 257 Adder 258 Upsampled output 260 Spectrogram decoder 262 Self-attention block 262a~n Self-attention block 263 Output 264 Projection layer 264a~n Projection layer 270 Spectrogram loss 270a~n Spectrogram loss 270b Second spectrogram loss 280 Final spectrogram loss 280 Soft DTW loss 300 TTS model 300 Non-autoregressive neural TTS model 302 Mel-frequency spectrogram sequence 302 Predicted mel-frequency spectrogram 302a First predicted mel-frequency spectrogram sequence 302a~n Mel-frequency spectrogram sequence 320 Text-to-speech 320 Input text-to-speech 400 Schematic diagram of Conv block 402 Conv layer 404 Batch normalization layer 406 Dropout layer 420 LConv block 430 Projection Layer 500 LConv Block 500 Schematic Diagram 502 Gate Linear Unit (GLU) Layer 504 LConv Layer 506 Feed Forward (FF) Layer 508 First Residual Connection 508 First Connector 510 Second Residual Connection 510 Second Connector 600 Computer-Implemented Method 600 Method 700 Computing Device 700a Standard Server 700b Laptop Computer 700c Rack Server System 710 Processor 720 Memory 730 Storage Device 740 High-Speed Interface / Controller 750 High-Speed Expansion Port 760 Low-Speed Interface / Controller 770 Low-Speed Bus 780 Display 790 Low-Speed Expansion Port

Claims

1. A computer-implemented method (600) that, when executed on data processing hardware (122), causes the data processing hardware (122) to perform operations for training a non-autoregressive text-to-speech (TTS) model (300), the operations comprising: obtaining a sequence representation (224) of an encoded text sequence (219) concatenated with variational embeddings (220); using a duration model network (230) to predict a phoneme duration (240) for each phoneme represented by the encoded text sequence based on the sequence representation (224); based on the predicted phoneme durations (240), learning an interval representation matrix using a first function (244) conditioned on the sequence representation (224); learning an auxiliary attention context representation using a second function (248) conditioned on the sequence representation (224); upsampling the sequence representation (224) to an upsampled output (258) specifying a number of frames using the interval representation and the auxiliary attention context representation; generating one or more predicted mel-frequency spectrogram sequences (302) of the encoded text sequence (219) based on the upsampled output (258) as an output from a spectrogram decoder (260) comprising a stack of one or more self-attention blocks; determining a final spectrogram loss (280) based on the one or more predicted mel-frequency spectrogram sequences (302) and a reference mel-frequency spectrogram sequence (202); training the TTS model (300) based on the final spectrogram loss (280); and wherein the operations further comprise: receiving training data comprising a reference audio signal (201) and a corresponding input text sequence (206), the reference audio signal (201) comprising an utterance that was spoken and the input text sequence (206) corresponding to a transcript of the reference audio signal (201). ​ Encoding the reference audio signal (201) into a variational embedding (220) using a residual encoder (180), wherein the variational embedding (220) unravels style / rhythm information from the reference audio signal (201); Encoding the input text sequence (206) into the encoded text sequence (219) using a text encoder (210); A computer-implemented method (600) further comprising: **Claim 2** The computer-implemented method (600) according to claim 1, wherein the first function (244) and the second function (248) each comprise a multi-layer perceptron-based learnable function. **Claim 3** The operation further comprises: Determining a global phoneme duration loss (241) based on the predicted phoneme durations (240) and the average phoneme durations (240); The step of training the TTS model (300) is further based on the global phoneme duration loss (241), the computer-implemented method (600) according to claim 1 or 2. **Claim 4** The step of training the TTS model (300) based on the final spectrogram loss (280) and the global phoneme duration loss (241) comprises training the duration model network (230) to predict the phoneme durations (240) for each phoneme without using teacher-aligned phoneme duration labels extracted from an external aligner, the computer-implemented method (600) according to claim 3. **Claim 5** The operation further comprises using the duration model network (230) to: Generate respective start and end boundaries for each phoneme represented by the encoded text sequence (219) based on the predicted phoneme durations (240); and Map the respective start and end boundaries generated for each phoneme to respective grid matrices (243) based on the number of phonemes represented by the encoded text sequence (219) and the number of reference frames in the reference mel-frequency spectrogram sequence (202). The computer-implemented method (600) further comprises: Learning the interval representation is based on the respective grid matrices (243) mapped from the start boundary and the end boundary. Learning the auxiliary attention context representation is based on the respective grid matrices (243) mapped from the start boundary and the end boundary, the computer-implemented method (600) according to any one of claims 1 to 4. **Claim 6** The step of upsampling the sequence representation (224) to the upsampled output (258) is determining the product (254) of the interval representation matrix and the sequence representation (224); determining the Einstein sum (einsum) (256) of the interval representation matrix and the auxiliary attention context representation; summing the projection (252) of the product (254) of the interval representation matrix and the sequence representation (224) and the einsum (256) to generate the upsampled output (258). The computer-implemented method (600) according to any one of claims 1 to 5, comprising: **Claim 7** The residual encoder (180) comprises a global variational autoencoder (VAE). The step of encoding the reference audio signal (201) into the variational embedding (220) is sampling the reference mel-frequency spectrogram sequence (202) from the reference audio signal (201); encoding the reference mel-frequency spectrogram sequence (202) into the variational embedding (220) using the global VAE. The computer-implemented method (600) according to claim 1, comprising: **Claim 8** The residual encoder (180) comprises a phoneme-level fine-grained variational autoencoder (VAE). The step of encoding the reference audio signal (201) into the variational embedding (220) is sampling the reference mel-frequency spectrogram sequence (202) from the reference audio signal (201); aligning the reference mel-frequency spectrogram sequence (202) with each phoneme in the sequence of phonemes extracted from the input text sequence (206). Encoding a sequence of variational embeddings (220) at the phoneme level, based on the step of aligning the reference mel-frequency spectrogram sequence (202) with each phoneme in the sequence of phonemes, using the finely-grained VAE at the phoneme level The computer-implemented method (600) according to claim 1, comprising

9. The residual encoder (180) comprises a stack of lightweight convolutional (LConv) blocks (420), and each LConv block (420) in the stack of LConv blocks (420) comprises a gated linear unit (GLU) layer (502), and an LConv layer (504) configured to receive the output of the GLU layer (502), and a residual connection configured to concatenate the output of the LConv layer (504) with the input to the GLU layer (502), and a final feed-forward layer configured to receive, as input, the residual connection that concatenates the output of the LConv layer (504) with the input to the GLU layer (502) The computer-implemented method (600) according to any one of claims 1, 7 or 8, comprising

10. The operation further comprises concatenating the encoded text sequence (219), the variational embedding (220), and a reference speaker embedding representing the identification of the reference speaker who uttered the reference audio signal (201), and generating the sequence representation (224) based on the duration model network (230) that receives, as input, the concatenation of the encoded text sequence (219), the variational embedding (220), and the reference speaker embedding The computer-implemented method (600) according to any one of claims 1, 7, 8, or 9, further comprising

11. The input text sequence (206) includes a sequence of phonemes, and the step of encoding the input text sequence (206) into the encoded text sequence (219) comprises receiving, from a phoneme lookup table, respective embeddings (207) of each phoneme in the sequence of phonemes For each phoneme in the sequence of phonemes, using the encoder prenet neural network of the text encoder (210) to process each respective embedding (207) to generate a respective transformed embedding (209) of the phoneme; Using a bank of convolutional blocks to process each respective transformed embedding (209) to generate a convolutional output (213); Using a stack of self-attention blocks to process the convolutional output (213) to generate the encoded text sequence (219); The computer-implemented method (600) according to any one of claims 1, 7, 8, 9 or 10, comprising: **Claim 12** The computer-implemented method (600) according to any one of claims 1 to 10, wherein each self-attention block in the stack of self-attention blocks comprises the same lightweight convolutional (LConv) block. **Claim 13** The computer-implemented method (600) according to any one of claims 1 to 10, wherein each self-attention block in the stack of self-attention blocks comprises the same transformer block. **Claim 14** A system (100) for training a non-autoregressive text-to-speech (TTS) model (300), comprising: Data processing hardware (122); Memory hardware (124) communicating with the data processing hardware (122), which, when executed by the data processing hardware (122), causes the data processing hardware (122) to: Obtain a sequence representation (224) of an encoded text sequence (219) concatenated with a variational embedding (220); Using a duration model network (230) Predict a phoneme duration (240) for each phoneme represented by the encoded text sequence (219) based on the sequence representation (224); Based on the predicted phoneme durations (240), Learn an interval representation matrix using a first function (244) conditioned on the sequence representation (224); Learn an auxiliary attention context representation using a second function (248) conditioned on the sequence representation (224); Using the interval representation and the auxiliary attention context representation, upsample the sequence representation (224) to an upsampled output that specifies the number of frames. Based on the upsampled output, as an output from a spectrogram decoder (260) comprising a stack of one or more self-attention blocks, generate one or more predicted mel-frequency spectrogram sequences of the encoded text sequence (219). Based on the one or more predicted mel-frequency spectrogram sequences and a reference mel-frequency spectrogram sequence, determine a final spectrogram loss (280). Train the TTS model (300) based on the final spectrogram loss (280). A memory hardware (124) storing instructions for performing an operation comprising: Comprising: The operation is: Receiving training data including a reference audio signal (201) and a corresponding input text sequence (206), wherein the reference audio signal (201) comprises an utterance that is spoken, and the input text sequence (206) corresponds to a transcript of the reference audio signal (201). Encoding the reference audio signal (201) into a variational embedding (220) using a residual encoder (180), wherein the variational embedding (220) unpacks style / rhythm information from the reference audio signal (201). Encoding the input text sequence (206) into the encoded text sequence (219) using a text encoder (210). A system (100) further comprising: Claims 15 The system (100) according to claim 14, wherein the first function (244) and the second function (248) each comprise a multi-layer perceptron-based learnable function. Claims 16 The operation further comprises: Determining a global phoneme duration loss (241) based on the predicted phoneme duration (240) and the average phoneme duration (240). The system (100) according to claim 14 or 15, wherein training the TTS model (300) is further based on the global phoneme duration loss (241).

17. Training the TTS model (300) based on the final spectrogram loss (280) and the global phoneme duration loss (241) comprises training the duration model network (230) to predict the phoneme duration (240) for each phoneme without using teacher - labeled phoneme duration labels extracted from an external aligner. The system (100) according to claim 16.

18. The operation uses the duration model network (230) to generate respective start boundaries and end boundaries for each phoneme represented by the encoded text sequence (219) based on the predicted phoneme duration (240); and map the respective start boundaries and end boundaries generated for each phoneme to respective grid matrices (243) based on the number of phonemes represented by the encoded text sequence (219) and the number of reference frames in the reference mel - frequency spectrogram sequence. further comprising learning the interval representation based on the respective grid matrices (243) mapped from the start boundaries and end boundaries; learning the auxiliary attention context representation based on the respective grid matrices (243) mapped from the start boundaries and end boundaries. The system (100) according to any one of claims 14 to 17.

19. Upsampling the sequence representation (224) to the upsampled output comprises determining the product of the interval representation matrix and the sequence representation (224); determining the Einstein sum of the interval representation matrix and the auxiliary attention context representation; and summing the projection of the product of the interval representation matrix and the sequence representation (224) and the Einstein sum to generate the upsampled output. The system (100) according to any one of claims 14 to 18.

20. The residual encoder (180) comprises a global variational autoencoder (VAE), encoding the reference audio signal (201) into the variational embedding (220) comprises sampling the reference mel-frequency spectrogram sequence from the reference audio signal (201), and encoding the reference mel-frequency spectrogram sequence into the variational embedding (220) using the global VAE The system (100) according to claim 14, comprising. **Claim 21** The residual encoder (180) comprises a phoneme-level fine-grained variational autoencoder (VAE), encoding the reference audio signal (201) into the variational embedding (220) comprises sampling the reference mel-frequency spectrogram sequence from the reference audio signal (201), aligning the reference mel-frequency spectrogram sequence with each phoneme in a sequence of phonemes extracted from the input text sequence (206), and encoding a sequence of phoneme-level variational embeddings (220) based on aligning the reference mel-frequency spectrogram sequence with each phoneme in the sequence of phonemes using the phoneme-level fine-grained VAE The system (100) according to claim 14, comprising. **Claim 22** The residual encoder (180) comprises a stack of lightweight convolutional (LConv) blocks, and each LConv block in the stack of LConv blocks (420) comprises a gated linear unit (GLU) layer, an LConv layer (504) configured to receive an output of the GLU layer (502), a residual connection configured to concatenate an output of the LConv layer (504) with an input to the GLU layer (502), and a final feed-forward layer configured to receive as input the residual connection that concatenates the output of the LConv layer (504) with the input to the GLU layer (502) The system (100) according to any one of claims 14, 20 or 21, comprising. **Claim 23** The operation comprises concatenating the encoded text sequence (219), the variational embedding (220), and a reference speaker embedding representing an identification of a reference speaker who uttered the reference audio signal (201) Based on a duration modeling network that receives, as input, the concatenation of the encoded text sequence (219), the variational embedding (220), and the reference speaker embedding, generate a sequence representation (224). The system (100) according to any one of claims 14, 20, 21, or 22, further comprising... **Claim 24** The input text sequence (206) includes a sequence of phonemes, Encoding the input text sequence (206) into the encoded text sequence (219) includes: Receiving, from a phoneme lookup table, respective embeddings for each phoneme in the sequence of phonemes; For each phoneme in the sequence of phonemes, processing the respective embeddings to generate respective transformed embeddings (209) of the phonemes using an encoder pre - neural network of the text encoder (210); Processing the respective transformed embeddings (209) using a bank of convolutional blocks to generate a convolutional output (213); Processing the convolutional output (213) using a stack of self - attention blocks to generate the encoded text sequence (219). The system (100) according to any one of claims 14, 20, 21, 22, or 23, comprising... **Claim 25** The system (100) according to any one of claims 21 to 23, wherein each self - attention block in the stack of self - attention blocks comprises an identical lightweight convolutional (LConv) block. **Claim 26** The system (100) according to any one of claims 21 to 23, wherein each self - attention block in the stack of self - attention blocks comprises an identical transformer block.

Citation Information

Patent Citations

  • Voice synthesis processing device, voice synthesis processing method, and, program

    JP2021012351A

  • Clockwork hierarchical variational encoder

    WO2019217035A1

  • Variational embedding capacity in expressive end-to-end speech synthesis

    WO2020236990A1

  • Duration informed attention network (durian) for audio-visual synthesis

    WO2021040989A1