Predicting parametric vocoder parameters from prosodic features

CN115943460BActive Publication Date: 2026-09-29GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180052153.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-26
Filing Date
2021-06-21
Publication Date
2026-09-29
Estimated Expiration
2041-06-21

AI Technical Summary

Technical Problem

虽然这些预测韵律特征足以驱动对语言和韵律特征操作的基于大型神经网络的声学模型,诸如WaveNet或WaveRNN模型,但这些预测韵律特征不足以驱动需要许多附加声码器参数的参数化声码器

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115943460B_ABST
    Figure CN115943460B_ABST
Patent Text Reader

Abstract

The method (600) includes receiving a text utterance (320) having words (240), each word having one or more syllables (230), and each syllable having one or more phonemes (220). The method also includes receiving prosodic features (322) representing an intended prosody of the text utterance and language specifications (402) as inputs to a vocoder model (400). The prosodic features include a duration, a pitch contour, and an energy contour of the text utterance, while the language specifications include sentence-level language features (252), word-level language features (242), syllable-level language features (232), and phoneme-level language features (222). The method also includes predicting vocoder parameters (450) based on the prosodic features and the language specifications. The method also includes providing the predicted vocoder parameters and the prosodic features to a parametric vocoder (155) configured to generate a synthesized speech representation (152) of the text utterance having the intended prosody.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to predicting parameterized vocoder parameters from prosodic features. Background Technology

[0002] Speech synthesis systems use text-to-speech (TTS) models to generate speech from text input. The generated / synthesized speech should accurately convey the message (intelligibility) while sounding like human speech (naturalness) and possessing the expected prosody (expressiveness). While traditional cascaded and parametric synthesis models can provide intelligible speech, and recent advances in speech neural modeling have significantly improved the naturalness of synthesized speech, most existing TTS models are ineffective at modeling prosody, resulting in a lack of expressiveness in synthesized speech used in important applications. For example, applications such as conversational assistants and long-form readers expect to generate realistic speech by capturing prosodic features not conveyed in the input text, such as intonation, stress, rhythm, and style. For instance, a simple statement can be spoken in many different ways, depending on whether the statement is a question, an answer to a question, whether there is uncertainty in the statement, or whether it conveys any other meaning about the environment or context not specified in the input text.

[0003] Recently, variational autoencoders have been developed to predict prosodic features of duration, pitch contour, and energy contour to efficiently model the prosody of synthesized speech. While these predicted prosodic features are sufficient to drive acoustic models based on large neural networks, such as WaveNet or WaveRNN, which operate on language and prosodic features, they are insufficient to drive parameterized vocoders that require many additional vocoder parameters. Summary of the Invention

[0004] One aspect of this disclosure provides a method for predicting parameterized vocoder parameters from prosodic features. The method includes receiving, at data processing hardware, a text utterance having one or more words, each word having one or more syllables, and each syllable having one or more phonemes. The method further includes receiving, at data processing hardware, prosodic features output from a prosodic model and the linguistic specification of the text utterance as input to a vocoder model, the prosodic features representing the expected prosodicity of the text utterance. The prosodic features include the duration, pitch profile, and energy profile of the text utterance, while the linguistic specification of the text utterance includes sentence-level linguistic features of the text utterance, word-level linguistic features of each word of the text utterance, syllable-level linguistic features of each syllable of the text utterance, and phoneme-level linguistic features of each phoneme of the text utterance. The method further includes the data processing hardware predicting vocoder parameters as output from the vocoder model based on the prosodic features output from the prosodic model and the linguistic specification of the text utterance. The method also includes the data processing hardware providing the predicted vocoder parameters output from the vocoder model and the prosodic features output from the prosodic model to a parameterized vocoder. The parameterized vocoder is configured to generate a synthesized speech representation of the text utterance with the desired prosody.

[0005] Implementations of this disclosure may include one or more of the following optional features. In some implementations, the method further includes receiving, at data processing hardware, linguistic feature alignment activations of the language specification of the text utterance as input to a vocoder model. In these implementations, the predicted vocoder parameters are further based on the linguistic feature alignment activations of the language specification of the text utterance. In some examples, the linguistic feature alignment activations include word-level alignment activations and syllable-level alignment activations. Word-level alignment activations each align the activation of each word with the syllable-level linguistic features of each syllable of the word, and syllable-level alignment activations each align the activation of each syllable with the phoneme-level linguistic features of each phoneme of the syllable. Here, the activation of each word may be based on the word-level linguistic features of the corresponding word and the sentence-level linguistic features of the text utterance. In some examples, the word-level linguistic features include chunk embeddings obtained from a series of chunk embeddings generated from the text utterance by a bidirectional encoder representation (BERT) model of the transducer.

[0006] In some embodiments, the method further includes selecting a discourse embedding of the text utterance by data processing hardware, the discourse embedding representing a desired prosodic pattern. For each syllable, using the selected discourse embedding, the method includes: using a prosodic model to predict the duration of the syllable by encoding phoneme-level linguistic features of each phoneme in the syllable with a corresponding prosodic syllable embedding of the syllable; predicting the pitch of the syllable based on the predicted duration of the syllable by the data processing hardware; and generating a plurality of fixed-length predicted pitch frames based on the predicted duration of the syllable by the data processing hardware. Each fixed-length pitch frame represents the predicted pitch of the syllable, wherein the prosodic features received as input to the vocoder model include a plurality of fixed-length predicted pitch frames generated for each syllable of the text utterance. In some examples, the method further includes, for each syllable, using the selected discourse embedding: predicting the energy level of each phoneme in the syllable based on the predicted duration of the syllable by the data processing hardware; and generating a plurality of fixed-length predicted energy frames based on the predicted duration of the syllable by the data processing hardware for each phoneme in the syllable. Each fixed-length predicted energy frame represents the predicted energy level of the corresponding phoneme. Here, the prosodic features received as input to the vocoder model further include multiple fixed-length predicted energy frames generated for each phoneme in each syllable of the text utterance.

[0007] In some implementations, the prosodic model is incorporated into a hierarchical language structure to represent text utterances. The hierarchical language structure includes: a first level comprising a Long Short-Term Memory (LSTM) processing unit representing each word of the text utterance; a second level comprising an LSTM processing unit representing each syllable of the text utterance; a third level comprising an LSTM processing unit representing each phoneme of the text utterance; a fourth level comprising an LSTM processing unit representing each fixed-length predicted pitch frame; and a fifth level comprising an LSTM processing unit representing each fixed-length predicted energy frame. The LSTM processing units of the second level time faster than those of the first level, the third level faster than those of the second level, the fourth level faster than those of the third level, and the fifth level times at the same speed as the fourth level and faster than those of the third level. In some implementations, the first level of the hierarchical language structure generates activation for each word of the text utterance in a single first pass; the second level of the hierarchical language structure generates activation for each syllable of the text utterance in a single second pass after the first pass; the third level of the hierarchical language structure generates activation for each phoneme of the text utterance in a single third pass after the second pass; the fourth level of the hierarchical language structure generates activation for each fixed-length predicted pitch frame in a single fourth pass after the third pass; and the fifth level of the hierarchical language structure generates activation for each fixed-length predicted energy frame in a single fifth pass after the third pass.

[0008] In some examples, the vocoder model incorporates a hierarchical language structure to represent text utterances. The hierarchical language structure includes: a first level comprising a Long Short-Term Memory (LSTM) processing unit representing each word of the text utterance; a second level comprising an LSTM processing unit representing each syllable of the text utterance; a third level comprising an LSTM processing unit representing each phoneme of the text utterance; and a fourth level comprising an LSTM processing unit representing each of a plurality of fixed-length speech frames, wherein the LSTM processing units of the fourth level time faster than those of the third level. The LSTM processing units of the second level time faster than those of the first level, the third level faster than those of the second level, and the fourth level faster than those of the third level. In these examples, each of the plurality of fixed-length speech frames may represent a corresponding portion of the predicted vocoder parameters output from the vocoder model. Furthermore, in these examples, the first level of the hierarchical language structure can generate activation for each word of the text utterance in a single first pass, the second level of the hierarchical language structure can generate activation for each syllable of the text utterance in a single second pass after the first pass, the third level of the hierarchical language structure can generate activation for each phoneme of the text utterance in a single third pass after the second pass, and the fourth level of the hierarchical language structure can generate activation for each fixed-length speech frame in a single fourth pass after the third pass.

[0009] The method may further include receiving training data at data processing hardware, the training data comprising multiple reference audio signals and corresponding transcripts. Each reference audio signal comprises spoken utterances of speech and has a corresponding prosody, while each transcript comprises a text representation of the corresponding reference audio signal. For each pair of reference audio signals and corresponding transcripts, the method may include: obtaining, by the data processing hardware, a reference language specification for the corresponding transcript and reference prosodic features representing the corresponding prosody of the corresponding reference audio signal; and training a vocoder model by the data processing hardware using a deep neural network to generate a series of fixed-length predicted speech frames based on the reference language specification and the reference prosodic features, the series of fixed-length predicted speech frames providing Mel-frequency cepstral coefficients, aperiodic components, and vocalization components. In some examples, for each reference audio signal, training the vocoder model further includes: sampling a series of fixed-length reference speech frames from the corresponding reference audio signal, the series of fixed-length reference speech frames providing reference Mel-spectral coefficients, reference aperiodic components, and reference vocal components of the reference audio signal; generating a gradient / loss between a series of fixed-length predicted speech frames generated by the vocoder model and a series of fixed-length reference speech frames sampled from the corresponding reference audio signal; and backpropagating the gradient / loss through the vocoder model.

[0010] In some embodiments, the method further includes decomposing the predicted vocoder parameters output from the vocoder model into Mel-Cepstral coefficients, aperiodic components, and vocalization components by data processing hardware. In these embodiments, the method further includes denormalizing the Mel-Cepstral coefficients, aperiodic components, and vocalization components individually by data processing hardware. In these embodiments, the method further includes concatenating prosodic features, denormalized Mel-Cepstral coefficients, denormalized aperiodic components, and denormalized vocalization components output from the prosodic model into a vocoder vector by data processing hardware. In these embodiments, providing the predicted vocoder parameters output from the vocoder model and the prosodic features output from the prosodic model to the parameterized vocoder includes providing the vocoder vector to the parameterized vocoder as input for generating a synthesized speech representation of the text utterance.

[0011] Another aspect of this disclosure provides a system for predicting parameterized vocoder parameters from prosodic features. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving a text utterance having one or more words, each word having one or more syllables, and each syllable having one or more phonemes. The operations also include receiving prosodic features output from a prosodic model and the linguistic specification of the text utterance as input to the vocoder model, the prosodic features representing the expected prosodicity of the text utterance. The prosodic features include the duration, pitch profile, and energy profile of the text utterance, while the linguistic specification of the text utterance includes sentence-level linguistic features of the text utterance, word-level linguistic features of the text utterance, syllable-level linguistic features of the text utterance, and phoneme-level linguistic features of the text utterance. The operations further include predicting vocoder parameters as output from the vocoder model based on the prosodic features output from the prosodic model and the linguistic specification of the text utterance. The operation also includes providing predicted vocoder parameters output from the vocoder model and prosodic features output from the prosodic model to the parameterized vocoder. The parameterized vocoder is configured to generate a synthesized speech representation of the text utterance with the desired prosodicity.

[0012] Implementations of this disclosure may include one or more of the following optional features. In some implementations, the operation further includes receiving a language feature alignment activation of the language specification of the text utterance as input to the vocoder model. In these implementations, the predicted vocoder parameters are further based on the language feature alignment activation of the language specification of the text utterance. In some examples, the language feature alignment activation includes word-level alignment activation and syllable-level alignment activation. The word-level alignment activation aligns the activation of each word with the syllable-level language features of each syllable of the word, and the syllable-level alignment activation aligns the activation of each syllable with the phoneme-level language features of each phoneme of the syllable. Here, the activation of each word may be based on the word-level language features of the corresponding word and the sentence-level language features of the text utterance. In some examples, the word-level language features include chunk embeddings obtained from a series of chunk embeddings generated from the text utterance by a bidirectional encoder representation (BERT) model of the transducer.

[0013] In some embodiments, the operation further includes selecting a discourse embedding of the text utterance, the discourse embedding representing the expected prosody. In these embodiments, for each syllable, using the selected discourse embedding, the operation further includes: predicting the duration of the syllable using a prosodic model by encoding phoneme-level linguistic features of each phoneme in the syllable with the corresponding prosodic syllable embedding of the syllable; predicting the pitch of the syllable based on the predicted duration of the syllable; and generating a plurality of fixed-length predicted pitch frames based on the predicted duration of the syllable. Each fixed-length pitch frame represents the predicted pitch of the syllable, wherein the prosodic features received as input to the vocoder model include a plurality of fixed-length predicted pitch frames generated for each syllable of the text utterance.

[0014] In some examples, the operation further includes, for each syllable, using a selected utterance embedding: predicting the energy level of each phoneme in the syllable based on the predicted duration of the syllable; and for each phoneme in the syllable, generating multiple fixed-length predicted energy frames based on the predicted duration of the syllable. Each fixed-length predicted energy frame represents the predicted energy level of the corresponding phoneme. The prosodic features received as input to the vocoder model further include multiple fixed-length predicted energy frames generated for each phoneme in each syllable of the text utterance.

[0015] In some implementations, the prosodic model is incorporated into a hierarchical language structure to represent text utterances. The hierarchical language structure includes: a first level comprising a short-term memory (LSTM) processing unit representing each word of the text utterance; a second level comprising an LSTM processing unit representing each syllable of the text utterance; a third level comprising an LSTM processing unit representing each phoneme of the text utterance; a fourth level comprising an LSTM processing unit representing each fixed-length predicted pitch frame; and a fifth level comprising an LSTM processing unit representing each fixed-length predicted energy frame. The LSTM processing units of the second level time faster than those of the first level, the third level faster than those of the second level, the fourth level faster than those of the third level, and the fifth level times at the same speed as the fourth level and faster than those of the third level. In some implementations, the first level of the hierarchical language structure generates activation for each word of the text utterance in a single first pass; the second level of the hierarchical language structure generates activation for each syllable of the text utterance in a single second pass after the first pass; the third level of the hierarchical language structure generates activation for each phoneme of the text utterance in a single third pass after the second pass; the fourth level of the hierarchical language structure generates activation for each fixed-length predicted pitch frame in a single fourth pass after the third pass; and the fifth level of the hierarchical language structure generates activation for each fixed-length predicted energy frame in a single fifth pass after the third pass.

[0016] In some examples, the vocoder model incorporates a hierarchical language structure to represent text utterances. The hierarchical language structure includes: a first level comprising a Long Short-Term Memory (LSTM) processing unit representing each word of the text utterance; a second level comprising an LSTM processing unit representing each syllable of the text utterance; a third level comprising an LSTM processing unit representing each phoneme of the text utterance; and a fourth level comprising an LSTM processing unit representing each of a plurality of fixed-length speech frames, wherein the LSTM processing units of the fourth level time faster than those of the third level. The LSTM processing units of the second level time faster than those of the first level, the third level faster than those of the second level, and the fourth level faster than those of the third level. In these examples, each of the plurality of fixed-length speech frames may represent a corresponding portion of the predicted vocoder parameters output from the vocoder model. Furthermore, in these examples, the first level of the hierarchical language structure can generate activation for each word of the text utterance in a single first pass, the second level of the hierarchical language structure can generate activation for each syllable of the text utterance in a single second pass after the first pass, the third level of the hierarchical language structure can generate activation for each phoneme of the text utterance in a single third pass after the second pass, and the fourth level of the hierarchical language structure can generate activation for each fixed-length speech frame in a single fourth pass after the third pass.

[0017] The operation may further include: receiving training data, the training data comprising multiple reference audio signals and corresponding transcripts. Each reference audio signal comprises spoken utterances of speech and has a corresponding prosody, while each transcript comprises a text representation of the corresponding reference audio signal. For each pair of reference audio signals and corresponding transcripts, the operation may further include: obtaining a reference language specification for the corresponding transcript and reference prosodic features representing the corresponding prosody of the corresponding reference audio signal; and training a vocoder model using a deep neural network to generate a series of fixed-length predicted speech frames based on the reference language specification and reference prosodic features, the series of fixed-length predicted speech frames providing Mel-Cepstral coefficients, aperiodic components, and phonological components. In some examples, training the vocoder model further includes: sampling a series of fixed-length reference speech frames from the corresponding reference audio signals, the series of fixed-length reference speech frames providing reference Mel-Cepstral coefficients, reference aperiodic components, and reference phonological components of the reference audio signals; generating a gradient / loss between the series of fixed-length predicted speech frames generated by the vocoder model and the series of fixed-length reference speech frames sampled from the corresponding reference audio signals; and backpropagating the gradient / loss through the vocoder model.

[0018] In some embodiments, the operation further includes decomposing the predicted vocoder parameters output from the vocoder model into Mel-Cepstral coefficients, aperiodic components, and vocalization components. In these embodiments, the operation also includes individually denormalizing the Mel-Cepstral coefficients, aperiodic components, and vocalization components. In these embodiments, the operation also includes concatenating the prosodic features, denormalized Mel-Cepstral coefficients, denormalized aperiodic components, and denormalized vocalization components output from the prosodic model into a vocoder vector. In these embodiments, providing the predicted vocoder parameters output from the vocoder model and the prosodic features output from the prosodic model to the parameterized vocoder includes providing the vocoder vector to the parameterized vocoder as input for generating a synthesized speech representation of the text utterance. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of an example system for training a deep neural network to provide a vocoder model for predicting vocoder parameters based on prosodic features representing the expected prosody of a text utterance.

[0020] Figure 2 It is incorporated into a hierarchical language structure to represent textual discourse. Figure 1 A schematic diagram of a vocoder model.

[0021] Figure 3A This is a schematic diagram of an example autoencoder used to predict the duration and pitch profile of each syllable in a text utterance.

[0022] Figure 3B This is a schematic diagram of an example autoencoder used to predict the duration and energy profile of each phoneme in a text utterance.

[0023] Figure 4 This is a schematic diagram of an example deep neural network used to predict vocoder parameters based on prosodic features used to drive a parameterized vocoder.

[0024] Figure 5 This is a schematic diagram of updating the parameters of the vocoder model.

[0025] Figure 6 This is a flowchart of an example setup for predicting vocoder parameters of a text discourse based on prosodic features output from a prosodic model and the linguistic norms of the text discourse.

[0026] Figure 7 This is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.

[0027] Similar reference numerals in various figures indicate similar elements. Detailed Implementation

[0028] Text-to-speech (TTS) models, frequently used by speech synthesis systems, typically operate on text input without any reference acoustic representation. The ability to produce realistic-sounding synthesized speech is attributable to numerous linguistic factors not provided in the text input. A subset of these linguistic factors is collectively called prosody and can include intonation (pitch variation), stress (stressed and unstressed syllables), duration, loudness, pitch, rhythm, and style. Prosody can indicate the emotional state of speech, its form (e.g., statement, question, command), the presence of irony or satire, uncertainty in linguistic knowledge, or other linguistic elements that cannot be encoded through the grammatical or lexical selection of the input text. Therefore, a given text input with high prosodic variation can produce synthesized speech with local variations in pitch and duration to convey different semantics, and synthesized speech with global variations in the overall pitch trajectory to convey different moods and emotions.

[0029] Neural network models offer the potential for robust speech synthesis by predicting linguistic factors that correspond to prosody not provided in the text input. Consequently, many applications, such as audiobook narration, news readers, speech design software, and conversational assistants, are able to produce synthesized speech that sounds realistic, rather than monotonous utterances. Recent advances in variational autoencoders (VACs) have enabled the prediction of prosodic features—including duration, pitch contour, and energy contour—to effectively model the prosody of synthesized speech. While these predicted prosodic features are sufficient to drive acoustic models based on large, state-of-the-art neural networks, such as WaveNet or WaveRNN, which operate on linguistic and prosodic features, they are insufficient to drive parametric vocoders, which require many additional vocoder parameters. That is, in addition to prosodic features of pitch, energy, and phoneme duration, parametric vocoders require many additional vocoder parameters, including Mel-Cepstral Equations (MCEPs) for each speech unit, typically fixed-length frames (e.g., 5 milliseconds), aperiodic components, and speech components. Therefore, parametric vocoders cannot benefit from the improvements in prosodic modeling made by VACs. However, compared to acoustic models based on large, state-of-the-art neural networks that run on servers, parametric vocoders are associated with low processing and memory requirements, making parametric vocoder models the preferred choice for setups on devices where constraints on processing and memory are relaxed.

[0030] The implementation described in this paper relates to a two-stage speech synthesis system incorporating a prosodic model and a neural network vocoder model. During the first stage, the prosodic model is configured to predict prosodic features representing the intended prosodicity of the text utterance. These prosodic features represent the acoustic information of the text utterance in terms of pitch (F0), phoneme duration, and energy (C0). However, the prosodic features predicted by the prosodic model only include a portion of the large number of vocoder parameters required to drive the parametric vocoder. During the second stage, the neural network vocoder model is configured to receive the prosodic features predicted by the prosodic model as input and generate the remaining vocoder parameters as output to drive the parametric vocoder to produce a synthesized speech representation of the text utterance with the intended prosodicity. The prosodic model can be incorporated with a variational autoencoder (VAE) optimized for the predicted prosodic features to provide a higher quality prosodic representation of the text utterance than conventional statistical parametric models can produce. Besides the prosodic features of pitch, energy, and phoneme duration, the task of these conventional statistical parametric models is to generate all vocoder parameters, including MCEP, aperiodic components, and speech components.

[0031] Advantageously, the neural network vocoder model can leverage the prosodic model's ability to accurately predict prosodic features representing the intended prosodicity of text speech, and use these predicted prosodic features for a dual purpose: inputting them into the vocoder model to predict the remaining vocoder parameters needed to drive the parameterized vocoder; and iteratively passing them through the vocoder model to drive the parameterized vocoder in combination with the remaining vocoder parameters predicted by the vocoder module. In other words, the prosodic features predicted by the prosodic model and the remaining vocoder parameters predicted by the vocoder model can collectively provide all the necessary vocoder parameters to drive the parameterized vocoder, generating a synthesized speech representation of the text speech with the intended prosodicity. Therefore, by utilizing a prosodic model optimized for modeling prosodicity and incorporating it into the neural network vocoder model to predict the remaining vocoder parameters, an on-device parameterized vocoder can be used to produce a synthesized speech representation with improved prosodicity.

[0032] The VAE-based prosodic model disclosed herein includes a clock-based hierarchical variational autoencoder (CHiVE) with an encoder and a decoder. The encoder part of the CHiVE can be trained to represent a utterance embedding that represents prosodicity by encoding a number of reference audio signals conditioned on prosodic features and language specifications associated with each reference audio signal. As described above, prosodic features represent acoustic information about the reference audio signals in terms of pitch (F0), phoneme duration, and energy (C0). For example, prosodic features may include phoneme duration and fixed-length frames of pitch and energy sampled from the reference audio signals. Language specifications may include, but are not limited to: phoneme-level language features containing information about phoneme position in a syllable, phoneme identity, and multiple phonemes in a syllable; syllable-level language features containing information such as syllable identity and whether a syllable is stressed or unstressed; word-level language features containing information such as whether an indicator word is a noun / adjective / verb; and sentence-level language features containing information about the speaker, the speaker's gender, and / or whether the utterance is a question or a phrase. In some examples, the prosodic model includes a bidirectional encoder representation (BERT) model configured to output chunk embeddings as transformers. In these examples, chunk embeddings can replace word-level language features that would otherwise explicitly encode grammatical information about each word.

[0033] Each utterance embedding encoded by the encoder portion can be represented by a fixed-length numerical vector. In some implementations, the fixed-length numerical vector includes values ​​equal to 256. However, other implementations may use fixed-length numerical vectors with values ​​greater than or less than 256. The decoder portion can decode the fixed-length utterance embedding into a series of phoneme durations via a first decoder, and use the phoneme durations to decode the fixed-length utterance embedding into a series of fixed-length frames (e.g., five milliseconds) of pitch and energy. The fixed-length utterance embedding can represent the expected prosody of the input text to be synthesized into speech. The series of phoneme durations and the fixed-length frames of pitch and energy correspond to the prosodic features predicted by the decoder portion. During training, the prosodic features of the phoneme durations and the fixed-length frames of pitch and energy predicted by the decoder portion closely match the prosodic features of the phoneme durations and the fixed-length frames of pitch and energy sampled from a reference audio signal associated with the fixed-length utterance embedding.

[0034] The disclosed neural network vocoder model is trained to predict vocoder parameters conditioned on reference prosodic features and linguistic norms of the training text utterance. That is, the vocoder model receives prosodic features representing the expected prosodicity of the training text utterance and the linguistic norms of the training text utterance as input, and predicts vocoder parameters as output based on the reference prosodic features and linguistic norms of the text utterance. The vocoder parameters include each speech unit, such as the MCEP of a typically fixed-length frame (e.g., 5 milliseconds), an aperiodic component, and a speech component. The prosodic feature of the energy profile / level C0 is the 0th MCEP required to drive the parameterized vocoder. Therefore, the vocoder model can be configured to predict the remaining MCEPs C[1–n] used to drive the parameterized vocoder. th Each training text discourse may have one or more words, each word may have one or more syllables, and each syllable may have one or more phonemes. Therefore, the linguistic specification of each training text discourse includes sentence-level linguistic features, word-level linguistic features for each word, syllable-level linguistic features for each syllable, and phoneme-level linguistic features for each phoneme.

[0035] The prosodic model and neural network vocoder model based on CHiVE can each incorporate a hierarchical structure of stacked layers of Long Short-Term Memory (LSTM) units. Each layer of the LSTM unit incorporates the structure of the text utterance, such that one layer represents a fixed-length frame, the next layer represents a phoneme, the next layer represents a syllable, and another layer represents a word. Furthermore, the hierarchical structure of the stacked layers of LSTM units can be variably timed according to the length of the hierarchical input data. For example, if the input data (e.g., text utterance) contains three-syllable words followed by four-syllable words, the syllable layer of the hierarchical structure will be timed three times relative to a single clock cycle of the word layer of the first input word, and then the syllable layer will be timed four more times relative to a subsequent single clock cycle of the word layer of the second word.

[0036] During inference, the CHiVE-based prosodic model is configured to receive text utterances and select utterance embeddings for those utterances. These utterance embeddings can be categorized for different prosodic domains, including but not limited to news readers, sports broadcasters, speakers, or story readers. The utterance embeddings can also be more granular to include subdomains. For example, a story reader domain could include utterance embeddings for conveying suspense in a thriller, and utterance embeddings for conveying different emotions consistent with the context of a given chapter in an ebook. Users can choose utterance embeddings that convey the desired prosodicity, or they can choose utterance embeddings that are paired with text utterances that closely match the received text utterances to be synthesized into speech. The received text utterances have at least one word, each word has at least one syllable, and each syllable has at least one phoneme. Because text utterances lack contextual, semantic, and pragmatic information to guide the appropriate prosodicity for generating synthesized speech from the utterances, the CHiVE-based prosodic model uses the selected utterance embeddings as latent variables to represent the desired prosodicity. Subsequently, the CHiVE-based prosodic model concatenates the selected utterance embedding with sentence-level, word-level, and syllable-level linguistic features obtained from the text utterance to predict the duration of each syllable and, based on the predicted duration, predict the pitch of each syllable. Finally, the CHiVE-based prosodic model is configured to generate multiple fixed-length pitch frames based on the predicted duration of each syllable, such that each fixed-length pitch frame represents the predicted pitch of the syllable. These multiple fixed-length pitch frames can provide a logarithm f0 to represent the fundamental frequency of the text utterance on a logarithmic scale. Similarly, the CHiVE-based prosodic model can predict the energy (e.g., loudness) of each syllable based on its predicted duration and generate multiple fixed-length energy frames, each representing the predicted energy c0 of the syllable, where c0 is the 0th MCEP.

[0037] The linguistic features of the text discourse (e.g., sentence-level, word-level, syllable-level, and phoneme-level linguistic features) and fixed-length pitch and / or energy frames output from the prosodic model can be provided as input to a neural network vocoder model to generate predicted vocoder parameters as output, the predicted vocoder parameters including the MCEP (c[1–n) of each speech unit. thThe vocoder model is configured to predict multiple fixed-length speech frames (e.g., 5 ms frames), each representing a corresponding portion of the predicted vocoder parameters. Furthermore, the vocoder model can insert prosodic features—the pitch f0, energy c0, and phoneme duration predicted by the prosodic model—into appropriate fixed-length speech frames to drive a parameterized vocoder with all the required vocoder parameters. Here, the portion of the prosodic vocoder parameters driving the parameterized vocoder is obtained from the prosodic model optimized for prosodic modeling, and the remaining vocoder parameters are obtained from the vocoder model based on the prosodic features.

[0038] Figure 1An example system 100 is shown for training a deep neural network 200 to provide a vocoder model 400, and for using the trained vocoder module 400 to predict vocoder parameters 450 of text utterance 320 to drive a parameterized vocoder 155. During training, system 100 incorporates a computing system 120 having data processing hardware 122 in communication with data processing hardware 122 and memory hardware 124, and storing instructions that cause the data processing hardware 122 to perform operations. In some embodiments, computing system 120 (e.g., data processing hardware 122) provides a trained prosodic model 300 and a trained vocoder model 400 to a text-to-speech (TTS) system 150 based on the trained deep neural network 200 for controlling the prosody of synthesized speech 152 from input text utterance 320. In other words, the trained prosody model 300 and the trained vocoder model 400 work together to generate all the necessary vocoder parameters 322, 450 required to drive the parameterized vocoder 155 on the TTS system 150 to produce synthesized speech 152 with the desired prosody. In the illustrated example, the TTS system 150 resides on a user device 110, such as a smartphone, smartwatch, smart speaker / display, smart appliance, laptop computer, desktop computer, tablet computer, or other computing device associated with the user. In other examples, a computing system 120 implements the TTS system 150. The computing system 120 used to train the vocoder model 400 and optionally the prosody model 300 may include a distributed system (e.g., a cloud computing environment). User device 110 has data processing hardware 112 and memory hardware 114, the memory hardware communicating with the data processing hardware 112 and storing instructions that cause the data processing hardware 112 to perform operations, such as executing prosody and vocoder models 300, 400 to generate vocoder parameters 322, 450 from input text utterance 320, and driving a parameterized vocoder 155 according to the vocoder parameters 322, 450 to generate synthesized speech 152. The synthesized speech 152 can be audibly output by a speaker communicating with user device 110. For example, the speaker can reside on user device 110 or can be a separate component communicating with user device 110 via a wired or wireless connection.

[0039] Since the input text utterance 320 lacks a method to convey context, semantics, and pragmatics to guide the appropriate prosodicity of the synthesized speech 152, the prosodic model 300 can predict the prosodic representation 322 of the input text utterance 320 by conditioned the model 300 on the language norm 402 extracted from the text utterance 320 and using a fixed-length utterance embedding 204 as a latent variable representing the expected prosodicity of the text utterance 320. That is, during inference, the prosodic model 300 can use the selected utterance embedding 204 to predict the prosodic representation 322 of the text utterance 320. The prosodic representation 322 can include prosodic features of the predicted pitch, predicted timing, and predicted loudness (e.g., energy) of the text utterance 320. Therefore, the terms "prosodic representation" and "prosodic features" can be used interchangeably. Subsequently, the prosodic features 322 predicted by the prosodic model 300 are fed as input to the vocoder model 400 to predict the remaining vocoder parameters 450 required to drive the parameterized vocoder 155. In other words, the parameterized vocoder 155 cannot generate synthesized speech 152 of the input text utterance 322 from the prosodic features 322 predicted solely by the prosodic model 300, and requires a large number of additional vocoder parameters 450 to generate the synthesized speech 152. The additional vocoder parameters 450 predicted by the vocoder model 400 include the MCEP (c[1–n)) for each speech unit. th ( ), aperiodic components, and vocal components. See below for reference. Figure 2 and 4 In more detail, the neural network vocoder model 400 is configured to predict multiple fixed-length speech frames 280V0 (e.g., 5 ms frames), each representing a corresponding portion of the predicted vocoder parameters 450. Finally, the vocoder model 400 is configured to provide the predicted vocoder parameters 450 output from the vocoder model 400 and the prosodic features 322 output from the prosodic model 300 to the parameterized vocoder 155, thereby configuring the parameterized vocoder 155 to generate a synthesized speech representation 152 of the text utterance 320 with the expected prosodicity. The vocoder parameters 450 and the prosodic features 322 can be cascaded into a single output vector to drive the parameterized vocoder 155.

[0040] In some implementations, the deep neural network 200 incorporated into the vocoder model 400 is trained on a large set of training data 201 stored on a data storage device 180 (e.g., memory hardware 124). The training data 201 includes multiple reference audio signals 202 and corresponding transcriptions 206. Each reference audio signal 202 may include spoken utterances of speech (e.g., human speech recorded by a microphone) and has a prosodic representation. Each transcription 206 may include a textual representation of the corresponding reference audio signal 202. For each pair of reference audio signals 204 and corresponding transcriptions 206, the deep neural network 200 obtains a reference language specification 402R for the corresponding transcription 206 and a reference prosodic feature 322R representing the corresponding prosody of the corresponding reference audio signal 202. Subsequently, the deep neural network 200 trains the vocoder model 400 to generate additional vocoder parameters 450 as a series of fixed-length predicted speech frames based on the reference language specification 402R and the reference prosodic feature 322R, which provide the MCEP (c[1–n)) for each frame. th The vocoder model 400 comprises a periodic component and a vocal component. The vocal component of each frame (e.g., a speech unit) can indicate whether the corresponding frame is audible or silent. The true values ​​of the vocoder parameter 450 can be sampled from the reference audio signal 202 as a series of fixed-length predicted speech frames. In some examples, reference prosodic features 322R, including pitch, energy, and phoneme duration, are sampled from the corresponding reference audio signal 202. In other examples, the reference prosodic features 322R correspond to prosodic features 322 predicted by a fully trained prosodic model 300, which receives a reference language specification 402R and a corresponding transcription 206 as input and uses a utterance embedding 204 representing the expected prosodic. In some implementations (not shown), the prosodic model 300 and the vocoder model 400 are jointly trained on training data 201. Here, a prosodic model 300 can be trained to encode utterance embeddings 204, each utterance embedding representing the prosody of a corresponding reference audio signal 202, and each utterance embedding 204 conditioned on a reference language specification 402R is decoded to predict prosodic features 322. In these embodiments, the prosodic features 322 predicted by the prosodic model 300 during joint training are used as reference prosodic features 322R, which, together with the reference language specification 402R, are fed as training input to the vocoder model 400 for predicting additional vocoder parameters 450R.

[0041] In the illustrated example, the computing system 120 stores trained prosodic and vocoder models 300, 400 on the data storage device 180. The user device 110 can obtain the trained prosodic and vocoder models 300, 400 from the data storage device 180, or the computing system 120 can directly push models 300, 400 to the user device 110 after training and / or retraining any one or both of them. The TTS system 150 executing on the user device 110 can use a parameterized vocoder 155, configured to receive prosodic features 322 and residual vocoder parameters 450 as input, and generate a synthesized speech representation 152 of the text speech 320 with the desired prosodic pattern as output.

[0042] When predicting the prosodic features 322 of the expected prosodicity of the text discourse 320, the prosodic model 300 can select the discourse embedding 204 representing the expected prosodicity of the text discourse 320. (See below for reference.) Figure 3A and 3B In more detail, the prosodic model 300 can use the selected utterance embedding 204 to predict the prosodic representation 322 of the text utterance 320.

[0043] Figure 2 The hierarchical language structure used to represent the text discourse 320 to be synthesized is shown (e.g., Figure 1Each of the deep neural network 200, prosodic model 300 (i.e., clock-based hierarchical variational autoencoder (CHiVE) or simply "autoencoder"), and vocoder model 400 can be incorporated into the text utterance. Prosodic model 300 is incorporated into hierarchical language structure 200 to provide a controllable prosodic model for jointly predicting the duration of syllables (and / or the duration of each phoneme within a syllable) and the pitch (F0) and energy (C0) profiles of syllables for each syllable of a given input text, without relying on any unique mapping from the given input text or other language specifications to produce synthesized speech 152 with the expected / selected prosodic. Relative to prosodic model 300, hierarchical language structure 200 is configured to decode a fixed-length utterance embedding 204 representing the expected prosodic of a given input text into multiple fixed-length prediction frames 280 (e.g., to predict pitch (F0), energy (C0)). Relative to vocoder model 400, hierarchical language structure 200 is configured to predict multiple fixed-length speech frames 280, 280V based on the language specification 402 of text utterance 320 and multiple fixed-length prediction frames 280F0, 280C0 output from prosodic model 300, as prosodic features 322 representing the expected prosodicity of the input text. Each fixed-length speech frame 280V may include a corresponding portion of the MCEP ([C1–nth]), aperiodic components, and residual vocoder parameters 450 of the phonological components for each speech input (e.g., each frame) predicted by vocoder model 400.

[0044] The hierarchical language structure 200 represents text discourse 322 as a hierarchical hierarchy of sentences 250, words 240, syllables 230, phonemes 220, and fixed-length frames 280. More specifically, each stacked hierarchical level includes Long Short-Term Memory (LSTM) processing units that are variably timed according to the length of the hierarchical input data. For example, the syllable level 230 times faster than the word level 240 and slower than the phoneme level 220. The rectangular blocks in each level correspond to the LSTM processing units for the corresponding words, syllables, phonemes, or frames. Advantageously, the hierarchical language structure 200 provides the LSTM processing units for the word level 240 with the memory for the last 100 words, the LSTM units for the syllable level 230 with the memory for the last 100 syllables, the LSTM units for the phoneme level 220 with the memory for the last 100 phonemes, and the LSTM units for the fixed-length pitch and / or energy frames 280 with the memory for the last 100 fixed-length frames 280. When each of the fixed-length frames 280 comprises a duration of five milliseconds (e.g., frame rate), the corresponding LSTM processing unit provides memory for the last 500 milliseconds (e.g., half a second).

[0045] In the illustrated example, the hierarchical language structure 200 represents the text utterance 322 as a series of three words 240A–240C at the word level 240, a series of five syllables 230Aa–230Cb at the syllable level 230, and a series of nine phonemes 220Aa1–220Cb2 at the phoneme level 220, to generate a series of predicted fixed-length frames 280 at the frame level 280. In some embodiments, the prosodic model 300 and / or vocoder model 400 incorporated into the hierarchical language structure 200 receive language feature alignment activations 221, 231, 241 of the language specification 402 of the text utterance 320. For example, the unroller broadcaster 270 may provide the language feature alignment activations 221, 231, 241 as input to the prosodic model 300 and / or vocoder model 400 before processing occurs. For simplicity, each of the receiving language features in prosody model 300 and vocoder model 400 is aligned with activations 221, 231, and 241. However, only one of models 300 and 400 can receive aligned activations 221, 231, and 241.

[0046] In some examples, the expander 270 broadcasts word-level alignment activation 241 to the word level 240 of the hierarchical language structure 200, broadcasts syllable-level alignment activation 231 to the syllable level 230 of the hierarchical linguistic structure 200, and broadcasts phoneme-level alignment activation 221 to the phoneme level 220 of the hierarchical language structure 200. Each word-level alignment activation 241 aligns the activation 245 of each word 240 with the syllable-level linguistic feature 236 of each syllable 230 of the word 240. Each syllable-level alignment activation 231 aligns the activation 235 of each syllable 230 with the phoneme-level linguistic feature 222 of each phoneme 220 of the syllable 230. In these examples, the activation 245 of each word 240 is based on the word-level linguistic feature 242 and sentence-level linguistic feature 252 of the corresponding word 240.

[0047] In some implementations, models 300 and 400 align activations 221, 231, and 241 received by each variable-rate stacking level of the hierarchical language structure 200 in a single pass, i.e., as a single time-series batch. Here, the unroller 270 can broadcast the aligned activations 221, 231, and 241 by converting the increment between the loop start index and the end index into a series of clustered indices to correctly broadcast corresponding activations 225, 235, and 245 across multiple layers / levels of the hierarchical language structure 200. In these implementations, the word level 240 of the hierarchical language structure generates activations 245 for each word 240 of the text utterance 322 in a single first pass, the syllable level 230 of the hierarchical language structure generates activations for each syllable of the text utterance in a single second pass after the first pass, the phoneme level 220 of the hierarchical language structure generates activations for each phoneme of the text utterance in a single third pass after the second pass, and the frame level 280 of the hierarchical language structure generates activations for each fixed-length frame in a single fourth pass after the third pass. For the prosodic model, frame level 280 can generate activations of fixed-length pitch frames 280F0 and fixed-length energy frames 280C0 in a single pass in parallel. These implementations involve four-pass reasoning, which allows each layer 240, 230, 220, 280 of the hierarchical language structure 200 representing the entire text discourse 320 to be expanded in a single pass before moving to the next level in the hierarchical structure. On the other hand, by generating a variable number of syllables for each word, two-pass reasoning divides the word level, syllable level, and phoneme level 240, 230, 220 into a first pass, which in turn enables the generation of a variable number of phonemes. The second pass then runs on the phoneme output to form a frame level to produce the word output frame 280. These two passes will be repeated for each word in the text discourse 320 associated with the sentence. While this division in the two-pass processing improves the efficiency of servers where processing and memory resources are not constrained (e.g., Figure 1 The speed of executing prosody model 300 and vocoder model 400 on the computing system 120 is optimized, but four-pass inference is optimized on the device (e.g., Figure 1 The processing speed of the prosody model 300 and the vocoder model 400 executed on the user device 110. That is, when the prosody model 300 and the vocoder model 400 are implemented on the device, two-pass inference is shown to be 30 percent faster than four-pass inference.

[0048] refer to Figure 3A and 3BIn some implementations, the autoencoder (i.e., prosodic model) 300 uses a hierarchical language structure 200 to predict a prosodic representation 322 of a given text utterance 320 during inference by jointly predicting the duration of phonemes 220 and the pitch and / or energy profile of each syllable 230. Since the text utterance 320 does not provide any contextual, semantic, or pragmatic information to indicate the appropriate prosodicity of the text utterance, the autoencoder 300 selects utterance embeddings 206 as latent variables to represent the expected prosodicity of the text utterance 320.

[0049] Data can be embedded from speech into data storage device 180 ( Figure 1 Speech embedding 204 is selected from the corresponding variable-length reference audio signal 202 during training. Each speech embedding 204 in the storage device 180 can be selected from the corresponding variable-length reference audio signal 202 during training. Figure 1 The autoencoder 300 may include an encoder portion (not shown) that compresses the prosody of a variable-length reference audio signal 202 into fixed-length utterance embeddings 204 during training, and stores each utterance embedding 204 along with a transcription 206 of the corresponding reference audio signal 202 in a data storage device 180 for use during inference. In the example shown, the autoencoder 300 may first locate utterance embeddings 204 having transcriptions 206 that closely match the text utterance 320, and then select one of the utterance embeddings 204 to predict the prosodic representation 322 of a given text utterance 320. Figure 1 In some examples, a fixed-length discourse embedding 204 is selected by picking specific points in the latent space of embedding 204 that may represent the target prosodicity with particular semantic and pragmatic meaning. In other examples, the latent space is sampled to select a random discourse embedding 204 for representing the expected prosodicity of the text discourse 320. In yet another example, the autoencoder 300 models the latent space as a multidimensional unit Gaussian by selecting the mean of discourse embeddings 204 with closely matched transcriptions 206 to represent the most likely prosodicity of the language norm 402 associated with the text discourse 320. The last example of selecting the mean of discourse embeddings 204 is a reasonable choice if the prosodic variation of the training data is reasonably neutral. In some additional examples, the user of user device 110 selects the expected prosodicity, such as a specific prosodic domain (e.g., a news reader, a speaker, a sports broadcaster), using an interface that can be executed on user device 110. Based on the expected prosodicity selected by the user, the autoencoder 300 can select the most suitable discourse embedding 204.

[0050] Figure 3AA text discourse 320 is shown, comprising three words 240A, 240B, and 240C represented at word level 240 of hierarchical language structure 200. The first word 240A contains syllables 230Aa and 230Ab, the second word 240B contains one syllable 230Ba, and the third word 240C contains syllables 230Ca and 230Cb. Therefore, the syllable level 230 of hierarchical language structure 200 comprises a series of five syllables 230Aa–230Cb of text discourse 320. At the syllable level 230 of the LTSM processing unit, the autoencoder 300 is configured to generate / output corresponding syllable embeddings (e.g., syllable-level activations 235) 235Aa, 235Ab, 235Ba, 235Ca, 235Cb from the following inputs: fixed-length utterance embeddings 204; sentence-level linguistic features 252 associated with the text utterance 320; word-level linguistic features 242 associated with the word 240 containing the syllable 230 (which may correspond to the word chunk embeddings generated by the BERT model 270); and syllable-level linguistic features 236 for the syllable 230. Sentence-level linguistic features 252 may include, but are not limited to, whether the text utterance 320 is a question, an answer to a question, a phrase, a sentence, the speaker's gender, etc. Word-level linguistic features 242 may include, but are not limited to, whether the word is a noun, adjective, verb, or other part of speech. Syllable-level linguistic features 236 may include, but are not limited to, whether the syllable 240 is stressed or unstressed.

[0051] In the example shown, each syllable 230Aa, 230Ab, 230Ba, 230Ca, 230Cb in syllable level 230 can be associated with a corresponding LTSM processing unit, which embeds the corresponding syllable 235Aa, 235Ab, 235Ba, 235Ca, 235Cb and outputs it to a faster-timing phoneme level 220 for parallel decoding of individual fixed-length predicted pitch (F0) frames 280, 280F0 (Figure 3A) and for decoding individual fixed-length predicted energy (C0) frames 280, 280C0 (Figure 3A). Figure 3B ). Figure 3A Each syllable in syllable level 230 is shown, comprising multiple fixed-length predicted pitch (F0) frames 280F0 indicating the duration (timing and pauses) and pitch profile of syllable 230. Here, the duration and pitch profile correspond to the prosodic representation of syllable 230. Figure 3B Each phoneme in phoneme level 220 is shown, which includes multiple fixed-length predicted energy (C0) frames 280C0 that indicate the duration and energy profile of the phoneme.

[0052] The first syllable 230Aa in syllable level 230 (i.e., LSTM processing unit Aa) receives a fixed-length utterance embedding 204, a sentence-level linguistic feature 252 associated with the text utterance 320, a word-level linguistic feature 242A associated with the first word 240A, and a syllable-level linguistic feature 236Aa of syllable 236Aa as input for generating the corresponding syllable embedding 235Aa. The second syllable 230Ab in syllable level 230 receives a fixed-length utterance embedding 204, a sentence-level linguistic feature 252 associated with the text utterance 320, a word-level linguistic feature 242A associated with the first word 240A, and a corresponding syllable-level linguistic feature 236 (not shown) of syllable 230Ab as input for generating the corresponding syllable embedding 235Ab. Although the example only shows the syllable-level linguistic feature 232 associated with the first syllable 230Aa, for clarity, from... Figure 3A and 3B The view omits only the corresponding syllable-level linguistic features 232 associated with each of the other syllables 230Ab-230Cb in syllable level 230.

[0053] For simplicity, the corresponding syllable-level linguistic features 236 of the processing block input to syllable 230Ab are not shown. The LTSM processing unit (e.g., rectangular Ab) associated with the second syllable 230Ab also receives the state of the preceding first syllable 230Aa. The syllables 230Ba, 230Ca, and 230Cb of the remaining sequence in syllable level 230 each generate corresponding syllable embeddings 235Ba, 235Ca, and 235Cb in a similar manner. For simplicity, the corresponding syllable-level linguistic features 236 of the processing block input to each of syllables 230Ba, 230Ca, and 230Cb are not shown. Furthermore, each LTSM processing unit of syllable level 230 receives the state of the immediately preceding LTSM processing unit of syllable level 240.

[0054] refer to Figure 3A The phoneme level 220 of the hierarchical language structure 200 includes a series of nine phonemes 220Aa1-220Cb2, each phoneme associated with a corresponding predicted phoneme duration 234. Furthermore, the autoencoder 300 uses corresponding syllable embeddings 235 to encode the phoneme-level language features 222 associated with each phoneme 220Aa1–220Cb2 for predicting the corresponding predicted phoneme duration 234 and for predicting the corresponding pitch (f0) contour of the syllable containing the phoneme. The phoneme-level language features 222 may include, but are not limited to, the identity of the sound of the corresponding phoneme 230 and / or the position of the corresponding phoneme 230 in the syllable containing the phoneme. Although the example only shows the phoneme-level language features 222 associated with the first phoneme 220Aa1, for clarity, from... Figure 3A and 3BThe view only omits the phoneme-level linguistic features 222 associated with the other phonemes 220Aa2-220Cb2 in phoneme level 220.

[0055] The first syllable 230Aa contains phonemes 220Aa1 and 220Aa2 and includes a predicted syllable duration equal to the sum of the predicted phoneme durations 234 of phonemes 220Aa1 and 220Aa2. Here, the predicted syllable duration of the first syllable 230Aa determines the number of fixed-length predicted pitch (F0) frames 280F0 to be decoded for the first syllable 230Aa. In the illustrated example, the autoencoder 300 decodes a total of seven fixed-length predicted pitch (F0) frames 280F0 of the first syllable 230Aa based on the sum of the predicted phoneme durations 234 of phonemes 220Aa1 and 220Aa2. Therefore, the faster-timing syllable layer 230 assigns the first syllable embedding 235Aa as input to each phoneme 220Aa1 and 220Aa2 included in the first syllable 230Aa. A timing signal can also be attached to the first syllable embedding 235Aa. Syllable level 230 also passes the state of the first syllable 230Aa to the second syllable 230Ab.

[0056] The second syllable 230Ab contains a single phoneme 220Ab1, and therefore includes a predicted syllable duration equal to the predicted phoneme duration 234 of phoneme 220Ab1. Based on the predicted syllable duration of the second syllable 230Ab, the autoencoder 300 decodes a total of four fixed-length predicted pitch (F0) frames 280F0 for the second syllable 230Ab. Therefore, the faster-timing syllable layer 230 assigns the second syllable embedding 235Ab as input to phoneme 220Ab1. A timing signal can also be appended to the second syllable embedding 235Aa. The syllable level 230 also passes the state of the second syllable 230Ab to the third syllable 230Ba.

[0057] The third syllable 230Ba contains phonemes 220Ba1, 220Ba2, and 220Ba3 and includes a predicted syllable duration equal to the sum of the predicted phoneme durations 234 of phonemes 220Ba1, 220Ba2, and 220Ba3. In the example shown, the autoencoder 300 decodes a total of eleven fixed-length predicted pitch (F0) frames 280F0 of the third syllable 230Ba based on the sum of the predicted phoneme durations 234 of phonemes 220Ba1, 220Ba2, and 220Ba3. Therefore, the faster-timing syllable layer 230 assigns the third syllable embedding 235Ba as input to each phoneme 220Ba1, 220Ba2, and 220Ba3 included in the third syllable 230Ba. A timing signal can also be attached to the third syllable embedding 235Ba. The syllable level 230 also passes the state of the third syllable 230Ba to the fourth syllable 230Ca.

[0058] The fourth syllable 230Ca contains a single phoneme 220Ca1, and therefore includes a predicted syllable duration equal to the predicted phoneme duration 234 of phoneme 220Ca1. Based on the predicted syllable duration of the fourth syllable 230Ca, the autoencoder 300 decodes a total of three fixed-length predicted pitch (F0) frames 280F0 for the fourth syllable 230Ca. Therefore, the faster-timing syllable layer 240 assigns the fourth syllable embedding 235Ca as input to phoneme 220Ca1. A timing signal can also be appended to the fourth syllable embedding 235Ca. The syllable level 230 also passes the state of the fourth syllable 230Ba to the fifth syllable 230Cb.

[0059] Finally, the fifth syllable 230Cb contains phonemes 220Cb1 and 220Cb2 and includes a predicted syllable duration equal to the sum of the predicted phoneme durations 234 of phonemes 220Cb1 and 220Cb2. In the example shown, the autoencoder 300 decodes a total of six fixed-length predicted pitch (F0) frames 280F0 of the fifth syllable 230Cb based on the sum of the predicted phoneme durations 234 of phonemes 220Cb1 and 220Cb2. Therefore, the faster-timing syllable layer 230 assigns the fifth syllable embedding 235Cb as input to each phoneme 220Cb1 and 220Cb2 included in the fifth syllable 230Cb. A timing signal can also be appended to the fifth syllable embedding 235Cb.

[0060] Still refer to Figure 3ASimilarly, the autoencoder 300 decodes each remaining syllable embedding 235Ab, 235Ba, 235Ca, 235Cb output from the syllable level 230 into a fixed-length predicted pitch (F0) frame 280 for each corresponding syllable 230Ab, 230Ba, 230Ca, 230Cb. For example, the second syllable embedding 235Ab is further combined at the output of the phoneme level 220 with the encoding of the second syllable embedding 235Ab and the corresponding phoneme-level linguistic feature 222 associated with phoneme 220Ab1, while the third syllable embedding 235Ba is further combined at the output of the phoneme level 220 with the encoding of the third syllable embedding 235Ba and the corresponding phoneme-level linguistic feature 222 associated with each of phonemes 220Ba1, 220Ba2, 220Ba3. Furthermore, the fourth syllable embedding 235Ca is further combined at the output of phoneme level 220 with the encoding of the fourth syllable embedding 235Ca and the corresponding phoneme-level linguistic feature 222 associated with phoneme 220Ca1, while the fifth syllable embedding 235Cb is further combined at the output of phoneme level 220 with the encoding of the fifth syllable embedding 235Cb and the corresponding phoneme-level linguistic feature 222 associated with each of phonemes 220Cb1 and 220Cb2. Although the fixed-length predicted pitch (F0) frame 280F0 generated by the autoencoder 300 includes a frame-level LSTM, other configurations can replace the frame-level LSTM of the pitch (F0) frame 280F0 with a feedforward layer, such that the pitch (F0) of each frame in the corresponding syllable is predicted in one pass.

[0061] Now for reference Figure 3B The autoencoder 300 is further configured to encode phoneme-level linguistic features 222 associated with each phoneme 220Aa1–220Cb2 using corresponding syllable embeddings 235, for predicting the corresponding energy (C0) profile of each phoneme 220. For clarity, from Figure 3BThe view only omits the phoneme-level linguistic features 222 associated with phonemes 220Aa2–220Cb2 in phoneme level 220. The autoencoder 300 determines the number of fixed-length prediction energy (C0) frames 280, 280C0 to be decoded for each phoneme 220 based on the corresponding prediction phoneme duration 234. For example, the autoencoder 300 decodes / generates four (4) prediction energy (C0) frames 280C0 for the first phoneme 220Aa1, three (3) prediction energy (C0) frames 280C0 for the second phoneme 220Aa2, and four (4) prediction energy (C0) frames 280C0 for the third phoneme 220Ab1. C0, two (2) predictive energy (C0) frames 280C0 for the fourth phoneme 220Ba1, five (5) predictive energy (C0) frames 280C0 for the fifth phoneme 220Ba2, four (4) predictive energy (C0) frames 280C0 for the sixth phoneme 220Ba3, three (3) predictive energy (C0) frames 280C0 for the seventh phoneme 220Ca1, four (4) predictive energy (C0) frames 280C0 for the eighth phoneme 220Cb1, and two (2) predictive energy (C0) frames 280C0 for the ninth phoneme 220Cb2. Therefore, similar to the predictive phoneme duration 234, the predictive energy profile of each phoneme in the phoneme level 220 is based on the encoding between the syllable embedding 235 input from the corresponding syllable in the slower timing syllable level 230 containing the phoneme and the language features 222 associated with the phoneme.

[0062] Figure 4 It shows that it can be incorporated into Figure 1 Example neural network vocoder model 400 in the TTS system 150. Vocoder model 400 receives prosodic features 322 output from prosodic model 300 as input, the prosodic features representing the expected prosodic model of text utterance 320 and the language specification 402 of text utterance 320. The prosodic features 322 output from the prosodic model may include duration, pitch contour f0_log, energy contour c0, and the duration of each phoneme in the text utterance. The pitch contour f0_log can be derived from... Figure 3A A series of fixed-length predicted pitch frames 280F0 represent this, and the energy profile c0 can be derived from... Figure 3BA series of fixed-length prediction energy frames 280C0 are represented. The language specification 402 includes sentence-level language features 252 of the text utterance 320, word-level language features 242 of each word 240 of the text utterance 320, syllable-level language features 236 of each syllable 230 of the text utterance 320, and phoneme-level language features 222 of each phoneme 220 of the text utterance 320. In the example shown, the fully connected layer 410 receives the language specification 402 and generates a fully connected output that is input to the collector 412. The language specification 402 can be normalized before being input to the fully connected layer 410. Meanwhile, the unroller broadcaster 270 receives the prosodic features 322 output from the prosodic model 300 and generates language feature alignment activations 221, 231, and 241 of the language specification 402 and provides them to the collector 412. At the collector, the word-level alignment activations 241 each align each word 240 ( Figure 2 Activation 245 ( Figure 2 ) and each syllable of word 240 (230) Figure 2 Syllable-level linguistic features 236 Figure 2 Alignment, syllable-level alignment activation 231, each syllable's activation 235 ( Figure 2 ) and each phoneme of the syllable 220 ( Figure 2 Phoneme-level linguistic features 222 Figure 2 Alignment, and phoneme-level alignment activation 221 each activates activation 225 for each phoneme 220. Figure 2 ) and the corresponding fixed-length frame 280 of phoneme 220 Figure 2 Alignment.

[0063] The outputs of linguistic feature alignment activations 221, 231, and 241 from collector 412 are input to cascade 414 for concatenation with prosodic features 322 output from the prosodic model. Prosodic features 322 can be normalized before concatenation with the outputs from collector 412 at cascade 414. The concatenated output from cascade 414 is input to a first LSTM layer 420 of vocoder model 400, and the output of the first LSTM layer 420 is input to a second LSTM layer 430 of vocoder model 400. Subsequently, the output of the second LSTM layer 430 is input to a recurrent neural network (RNN) layer 440 of vocoder model 400, and a separator 445 splits the output of the RNN layer 440 into predicted additional vocoder parameters 450. (See above reference...) Figure 1 As described, the additional vocoder parameters 450 split by the splitter 445 include each speech unit, such as the MCEP (c[1–n) of a fixed-length speech frame 280V0. thThe prosodic model 400 is configured to predict multiple fixed-length speech frames 280V0 (e.g., 5 ms frames), each representing a corresponding portion of the predicted vocoder parameters 450. The prosodic features 322 predicted by the prosodic model 300, together with the additional vocoder parameters 450, provide all the vocoder parameters needed to drive the parameterized vocoder 155 to produce a synthesized speech representation 152 of the text speech 320 with the desired prosodicity. Therefore, after splitting the additional vocoder parameters 450, the vocoder model 400 inserts the prosodic features 322 of pitch f0 and energy c0 into appropriate speech units, allowing the cascade 455 to cascade the prosodic features 322 and the additional vocoder parameters 450 into a final vocoder vector 460 for each of the multiple fixed-length speech frames 280V0 to drive the parameterized vocoder 155. Before cascading, the additional vocoder parameter 450 can be denormalized.

[0064] Figure 5 This is a sample procedure 500 used to train the vocoder model 400. You can refer to [the example]. Figure 1 and 5 Describing process 500. As an example, vocoder model 400 can be trained to learn additional vocoder parameters 450 (e.g., MCEP(c[1–n)) to predict a segment of input text (e.g., a reference transcription 206 of a reference audio signal 202) using reference prosodic features 322R (e.g., pitch F0, energy C0, and phoneme duration) and reference language specification 402R as input. th The vocoder parameters 450 can be represented as a series of fixed-length predicted speech frames 280V0, each fixed-length predicted speech frame providing the MCEP, aperiodic component, and speech component of the corresponding portion of the transcription 206.

[0065] Process 500 executes loss module 510, which is configured to generate gradient / loss 520 between the predicted additional vocoder parameters 450 output by vocoder model 400 and reference speech frames 502 sampled from a reference audio signal 202 (e.g., utterance) associated with transcription 206. The reference speech frames 502 sampled from the reference audio signal 202 may include fixed-length reference speech frames (e.g., 5 ms), each providing reference (e.g., true) vocoder parameters sampled from a corresponding portion of the reference audio signal. Therefore, loss module 510 can generate gradient / loss between a series of fixed-length predicted speech frames 280V0 (representing the predicted additional vocoder parameters 450) generated by vocoder model 400 and a series of fixed-length reference speech frames 502 (representing the reference / true vocoder parameters 450) sampled from the corresponding reference audio signal 202. Here, the gradient / loss 520 can be backpropagated through the vocoder model 400 to update the parameters until the vocoder model 400 is fully trained.

[0066] Figure 6 This is a flowchart illustrating an example arrangement of the operation of a method 600 that uses prosodic features 322 to predict additional vocoder parameters 322 of text utterance 320. The additional vocoder parameters 322 and the prosodic features 322 constitute all the necessary vocoder parameters for driving the parameterized vocoder 155 to produce a synthesized speech representation 152 of text utterance 320 having the expected prosodicity conveyed by the prosodic features 322. See also... Figures 1 to 4 Method 600 is described. Memory hardware 114 residing on user device 110 can store instructions that cause data processing hardware 112 residing on user device 110 to perform the operations of method 600, as exemplified. At operation 602, method 600 includes receiving text utterances 320. Text utterances 320 have at least one word, each word has at least one syllable, and each syllable has at least one phoneme.

[0067] At operation 604, method 600 includes receiving prosodic features 322 output from prosodic model 300 and linguistic specifications 402 of text utterance 320 as input to vocoder model 400. Prosodic features 322 represent the expected prosodicity of text utterance 320 and include the duration, pitch profile, and energy profile of text utterance 320. Linguistic specifications 402 include sentence-level linguistic features 252 of text utterance 320, word-level linguistic features 242 for each word 240 of text utterance 320, syllable-level linguistic features 236 for each syllable 230 of text utterance 320, and phoneme-level linguistic features 222 for each phoneme 220 of text utterance 320.

[0068] At operation 606, method 600 includes predicting (attached) vocoder parameters 450 as output from vocoder model 400 based on prosodic features 322 and language specifications 402. At operation 608, method 600 includes providing the predicted vocoder parameters 450 output from vocoder model 400 and prosodic features 322 output from prosodic model 300 to parameterized vocoder 155. Parameterized vocoder 155 is configured to generate a synthesized speech representation 152 of text utterance 320 with the expected prosodicity. In other words, the attached vocoder parameters 450 and prosodic features 322 are configured to drive parameterized vocoder 155 to generate synthesized speech representation 152. (See reference...) Figure 4 As described in the vocoder model 400, the prosodic features 322 and additional vocoder parameters 450 output from the prosodic model 300 can be cascaded to the vocoder vector 460 of each speech unit (e.g., each of a plurality of fixed-length speech frames 280V0) to drive the parameterized vocoder 155.

[0069] Text-to-speech (TTS) system 150 may combine prosodic model 300, vocoder model 400, and parametric vocoder 155. TTS system 150 may reside on user device 110, i.e., execute on data processing hardware 112 of the user device. In some configurations, TTS system 150 resides on computing system (e.g., server) 120, i.e., executes on data processing hardware 122. In some examples, some portions of TTS system 150 execute on computing system 120 and the remainder of TTS system 150 execute on user device 110. For example, at least one of prosodic model 300 or vocoder model 400 may execute on computing system 120, while parametric vocoder 155 may be on user device.

[0070] Figure 7 This is a schematic diagram of an example computing device 700 that can be used to implement the systems and methods described in this document. The computing device 700 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components, connections, and relationships between components shown herein, as well as the functionality of the components, are intended to be exemplary only and are not intended to limit the implementation of the inventions described and / or claimed in this document.

[0071] Computing device 700 includes a processor 710, a memory 720, a storage device 730, a high-speed interface / controller 740 connected to the memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connected to a low-speed bus 770 and the storage device 730. Each of components 710, 720, 730, 740, 750, and 760 is interconnected using various buses, and the components may be mounted on a general-purpose motherboard or otherwise. The processor 710 is capable of processing instructions for execution within computing device 700, including instructions stored in memory 720 or on storage device 730 to display graphical information of a graphical user interface (GUI) on an external input / output device, such as a display 780 coupled to the high-speed interface 740. In other embodiments, multiple processors and / or multiple buses may be used with multiple memories and various types of memory, depending on the situation. Moreover, multiple computing devices 700 may be connected, with each device providing a portion of the necessary operation (e.g., as a server group, blade server cluster, or multiprocessor system).

[0072] Memory 720 stores information non-temporarily within computing device 700. Memory 720 may be a computer-readable medium, (a plurality of) volatile memory cells, or (a plurality of) non-volatile memory cells. Non-temporarily stored memory 720 may be a physical means for temporarily or permanently storing programs (e.g., instruction sequences) or data (e.g., program state information) for use by computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used in firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and magnetic disks or magnetic tapes.

[0073] Storage device 730 provides mass storage for computing device 700. In some embodiments, storage device 730 is a computer-readable medium. In various embodiments, storage device 730 may be a floppy disk device, hard disk device, optical disk device, magnetic tape device, flash memory or other similar solid-state storage device, or an array of devices, including devices in a storage area network or other configuration. In additional embodiments, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as memory 720, storage device 730, or memory on processor 710.

[0074] High-speed controller 740 manages ultra-wideband operation of computing device 700, while low-speed controller 760 manages lower ultra-wideband operation. This allocation of responsibilities is merely exemplary. In some embodiments, high-speed controller 740 is coupled to memory 720, display 780 (e.g., via a graphics processor or accelerometer), and high-speed expansion port 750 which can accept various expansion cards (not shown). In some embodiments, low-speed controller 760 is coupled to storage device 730 and low-speed expansion port 790. Low-speed expansion port 790, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, Wireless Ethernet), may be coupled to one or more input / output devices, such as keyboards, pointing devices, scanners, or networking devices such as switches or routers, for example, via a network adapter.

[0075] As shown in the figure, the computing device 700 can be implemented in a variety of different forms. For example, the computing device can be implemented as a standard server 700a, or multiple times in a group of such servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.

[0076] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can be included in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be coupled for dedicated or general purposes to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transfer data and instructions to the storage system, at least one input device, and at least one output device.

[0077] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0078] The processes and logic flows described in this specification can be executed by one or more programmable processors, also known as data processing hardware, which execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). For example, processors suitable for executing computer programs include both general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from or transfer data to said one or more mass storage devices, or both. However, a computer is not required to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory can be supplemented by or incorporated into dedicated logic circuitry.

[0079] To provide interaction with a user, one or more aspects of this disclosure can be implemented on a computer with a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or a touchscreen) to display information to the user, and optionally a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending web pages to a web browser on the user's client device in response to a request received from a web browser.

[0080] Several embodiments have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of this disclosure. Therefore, other embodiments are within the scope of the appended claims.

Claims

1. A method (600) for predicting parameterized vocoder parameters from prosodic features, comprising: The data processing hardware (122) receives a text utterance (320) having one or more words (240), each word (240) having one or more syllables (230), and each syllable (230) having one or more phonemes (220). The following items are received as inputs to the vocoder model (400) at the data processing hardware (122): The prosodic features (322) output from the prosodic model (300), the prosodic features representing the expected prosodic of the text utterance (320), the prosodic features (322) including the duration, pitch profile, and energy profile of the text utterance (320); and The language specification (402) of the text discourse (320) includes sentence-level language features (252) of the text discourse (320), word-level language features (242) of each word (240) of the text discourse (320), syllable-level language features (232) of each syllable (230) of the text discourse (320), and phoneme-level language features (222) of each phoneme (220) of the text discourse (320). The data processing hardware (122) predicts vocoder parameters (450) as the output from the vocoder model (400) based on the prosodic features (322) output from the prosodic model (300) and the language specification (402) of the text discourse (320); as well as The data processing hardware (122) provides the predicted vocoder parameters (450) output from the vocoder model (400) and the prosodic features (322) output from the prosodic model (300) to the parameterized vocoder (155), which is configured to generate a synthesized speech representation (152) of the text utterance (320) with the expected prosodic. The vocoder model (400) is incorporated into a hierarchical language structure (200) to represent the text utterance (320), wherein the hierarchical language structure (200) includes: The first level includes a Long Short-Term Memory (LSTM) processing unit representing each word (240) of the text discourse (320); The second level includes an LSTM processing unit representing each syllable (230) of the text utterance (320), and the LSTM processing unit of the second level times faster than the LSTM processing unit of the first level. The third level includes an LSTM processing unit representing each phoneme (220) of the text utterance (320), wherein the LSTM processing unit of the third level times faster than the LSTM processing unit of the second level; and The fourth level includes an LSTM processing unit representing each of a plurality of fixed-length speech frames (502), the LSTM processing unit of the fourth level timing faster than the LSTM processing unit of the third level.

2. The method (600) according to claim 1, further comprising: At the data processing hardware (122), the language feature alignment activation of the language specification (402) of the text utterance (320) is received as input to the vocoder model (400). The predicted vocoder parameters (450) are further activated based on the language features of the language specification (402) of the text discourse (320).

3. The method (600) according to claim 2, wherein, The language feature alignment activation includes: Word-level aligned activation, each of the word-level aligned activations aligns the activation of each word (240) with the syllable-level linguistic feature (232) of each syllable (230) of the word (240); and Syllable-level alignment activation, each of the syllable-level alignment activations aligns the activation of each syllable (230) with the phoneme-level linguistic feature (222) of each phoneme (220) of the syllable (230).

4. The method (600) according to claim 3, wherein, The activation of each word (240) is based on the word-level linguistic features (242) of the corresponding word (240) and the sentence-level linguistic features (252) of the text discourse (320).

5. The method (600) according to claim 1, wherein, The word-level language features (242) include word embeddings obtained from a word embedding sequence generated from the text discourse (320) by a bidirectional encoder representation model of the converter.

6. The method (600) according to claim 1, further comprising: The data processing hardware (122) selects the discourse embedding (204) of the text discourse (320), the discourse embedding (204) representing the expected prosody; For each syllable (230), the selected utterance embedding (204) is used: The duration of the syllable (230) is predicted by the prosodic model (300) by the data processing hardware (122) encoding the phoneme-level linguistic features (222) of each phoneme (220) in the syllable (230) using the corresponding prosodic syllable embedding of the syllable (230); The data processing hardware (122) predicts the fundamental tone of the syllable (230) based on the predicted duration of the syllable (230); and The data processing hardware (122) generates multiple fixed-length predicted pitch frames based on the predicted duration of the syllable (230), each fixed-length pitch frame representing the predicted pitch of the syllable (230). The prosodic features (322) received as input to the vocoder model (400) include the plurality of fixed-length predicted pitch frames generated for each syllable (230) of the text utterance (320).

7. The method (600) according to claim 6, further comprising, for each syllable (230), using a selected utterance embedding (204): The data processing hardware (122) predicts the energy level of each phoneme (220) in the syllable (230) based on the predicted duration of the syllable (230); and For each phoneme (220) in the syllable (230), the data processing hardware (122) generates a plurality of fixed-length predicted energy frames (280) based on the predicted duration of the syllable (230), each fixed-length predicted energy frame representing the predicted energy level of the corresponding phoneme (220). in, The prosodic features (322) received as input to the vocoder model (400) further include the plurality of fixed-length predicted energy frames (280) generated for each phoneme (220) in each syllable (230) of the text utterance (320).

8. The method (600) according to claim 7, wherein, The prosodic model (300) is incorporated into a hierarchical language structure (200) to represent the text discourse (320), the hierarchical language structure (200) comprising: The first level includes a Long Short-Term Memory (LSTM) processing unit representing each word (240) of the text discourse (320); The second level includes an LSTM processing unit representing each syllable (230) of the text utterance (320), and the LSTM processing unit of the second level times faster than the LSTM processing unit of the first level. The third level includes an LSTM processing unit representing each phoneme (220) of the text utterance (320), and the LSTM processing unit of the third level times faster than the LSTM processing unit of the second level. The fourth stage includes an LSTM processing unit representing each fixed-length predicted pitch frame, the LSTM processing unit of the fourth stage timing faster than the LSTM processing unit of the third stage; and The fifth stage includes an LSTM processing unit representing each fixed-length prediction energy frame, the LSTM processing unit of the fifth stage timing at the same speed as the LSTM processing unit of the fourth stage and timing faster than the LSTM processing unit of the third stage.

9. The method (600) according to claim 8, wherein: The first level of the hierarchical language structure (200) generates the activation of each word (240) of the text discourse (320) in a single first pass; The second level of the hierarchical language structure (200) generates activation of each syllable (230) of the text discourse (320) in a single second pass after the first pass; The third level of the hierarchical language structure (200) generates the activation of each phoneme (220) of the text discourse (320) in a single third pass after the second pass; The fourth level of the hierarchical language structure (200) generates activation of each fixed-length predicted pitch frame in a single fourth pass after the third pass; as well as The fifth level of the hierarchical language structure (200) generates activation of each fixed-length prediction energy frame in a single fifth pass after the third pass.

10. The method (600) according to claim 1, wherein, Each of the plurality of fixed-length speech frames (502) represents a corresponding portion of the predicted vocoder parameters (450) output from the vocoder model (400).

11. The method (600) according to claim 1, wherein: The first level of the hierarchical language structure (200) generates the activation of each word (240) of the text discourse (320) in a single first pass; The second level of the hierarchical language structure (200) generates activation of each syllable (230) of the text discourse (320) in a single second pass after the first pass; The third level of the hierarchical language structure (200) generates the activation of each phoneme (220) of the text discourse (320) in a single third pass after the second pass; as well as The fourth level of the hierarchical language structure (200) generates activation of each fixed-length speech frame in the plurality of fixed-length speech frames (502) in a single fourth pass after the third pass.

12. The method (600) according to claim 1, further comprising: Training data is received at the data processing hardware (122), the training data comprising a plurality of reference audio signals and corresponding transcripts (206), each reference audio signal comprising spoken utterances of speech and having corresponding prosody, and each transcript (206) comprising a text representation of the reference audio signals; and For each reference audio signal and its corresponding transcription (206) pair: The data processing hardware (122) obtains the reference language specification (402) of the corresponding transcription (206) and the reference prosodic features (322) representing the corresponding prosodic of the reference audio signal; and The data processing hardware (122) uses a deep neural network to train the vocoder model (400) to generate a fixed-length sequence of predicted speech frames (502) based on the reference language specification (402) and the reference prosodic features (322), the fixed-length sequence of predicted speech frames providing Mel-frequency cepstral coefficients, aperiodic components and vocal components.

13. The method (600) according to claim 12, wherein, Training the vocoder model (400) further includes, for each reference audio signal: A fixed-length sequence of reference speech frames (502) is sampled from the reference audio signal, the fixed-length sequence of reference speech frames providing reference Mel-spectral coefficients, reference aperiodic components and reference vocal components of the reference audio signal; The gradient or loss generated between the fixed-length predicted speech frame (502) sequence generated by the vocoder model (400) and the fixed-length reference speech frame (502) sequence sampled from the reference audio signal; and The gradient or loss is backpropagated through the vocoder model (400).

14. The method (600) according to any one of claims 1 to 13, further comprising: The data processing hardware (122) splits the predicted vocoder parameters (450) output from the vocoder model (400) into Mel-frequency cepstral coefficients, aperiodic components, and vocal components. The Mel-frequency cepstral coefficients, aperiodic components, and vocal components are individually normalized by the data processing hardware (122); as well as The data processing hardware (122) concatenates the prosodic features (322), denormalized Mel-Cepstral coefficients, denormalized aperiodic components, and denormalized vocal components output from the prosodic model (300) into a vocoder vector (460). Providing the predicted vocoder parameters (450) output from the vocoder model (400) and the prosodic features (322) output from the prosodic model (300) to the parameterized vocoder (155) includes providing the vocoder vector (460) to the parameterized vocoder (155) as input to the synthesized speech representation (152) for generating the text utterance (320).

15. A system (100) for predicting parameterized vocoder parameters from prosodic features, comprising: Data processing hardware (122); as well as A memory hardware (124) communicates with the data processing hardware (122), the memory hardware (124) stores instructions, and when the instructions are executed on the data processing hardware (122), the data processing hardware (122) performs an operation, the operation including: Receive a text utterance (320) having one or more words (240), each word (240) having one or more syllables (230), each syllable (230) having one or more phonemes (220); The following items are accepted as inputs to the vocoder model (400): The prosodic features (322) output from the prosodic model (300), the prosodic features representing the expected prosodic of the text utterance (320), the prosodic features (322) including the duration, pitch profile, and energy profile of the text utterance (320); and The language specification (402) of the text discourse (320) includes sentence-level language features (252) of the text discourse (320), word-level language features (242) of each word (240) of the text discourse (320), syllable-level language features (232) of each syllable (230) of the text discourse (320), and phoneme-level language features (222) of each phoneme (220) of the text discourse (320). Vocoder parameters (450) are predicted as the output of the vocoder model (400) based on the prosodic features (322) output from the prosodic model (300) and the language norms (402) of the text utterance (320); and The predicted vocoder parameters (450) output from the vocoder model (400) and the prosodic features (322) output from the prosodic model (300) are provided to a parameterized vocoder (155), which is configured to generate a synthesized speech representation (152) of the text utterance (320) with the expected prosodic. The vocoder model (400) is incorporated into a hierarchical language structure (200) to represent the text utterance (320), wherein the hierarchical language structure (200) includes: The first level includes a Long Short-Term Memory (LSTM) processing unit representing each word (240) of the text discourse (320); The second level includes an LSTM processing unit representing each syllable (230) of the text utterance (320), and the LSTM processing unit of the second level times faster than the LSTM processing unit of the first level. The third level includes an LSTM processing unit representing each phoneme (220) of the text utterance (320), wherein the LSTM processing unit of the third level times faster than the LSTM processing unit of the second level; and The fourth level includes an LSTM processing unit representing each of a plurality of fixed-length speech frames (502), the LSTM processing unit of the fourth level timing faster than the LSTM processing unit of the third level.

16. The system (100) according to claim 15, wherein, The operation further includes: The language feature alignment activation of the language specification (402) of the received text utterance (320) is used as the input of the vocoder model (400). The predicted vocoder parameters (450) are further activated based on the language features of the language specification (402) of the text discourse (320).

17. The system (100) according to claim 16, wherein, The language feature alignment activation includes: Word-level aligned activation, each of the word-level aligned activations aligns the activation of each word (240) with the syllable-level linguistic feature (232) of each syllable (230) of the word (240); and Syllable-level alignment activation, each of the syllable-level alignment activations aligns the activation of each syllable (230) with the phoneme-level linguistic feature (222) of each phoneme (220) of the syllable (230).

18. The system (100) according to claim 17, wherein, The activation of each word (240) is based on the word-level linguistic features (242) of the corresponding word (240) and the sentence-level linguistic features (252) of the text discourse (320).

19. The system (100) according to claim 15, wherein, The word-level language features (242) include word embeddings obtained from a word embedding sequence generated from the text discourse (320) by a bidirectional encoder representation model of the converter.

20. The system (100) according to claim 15, wherein, The operation further includes: Select the discourse embedding (204) of the text discourse (320), the discourse embedding (204) representing the expected prosody; For each syllable (230), the selected utterance embedding (204) is used: The duration of the syllable (230) is predicted using the prosodic model (300) by encoding the phoneme-level linguistic features (222) of each phoneme (220) in the syllable (230) with the corresponding prosodic syllable embedding of the syllable (230); The fundamental tone of the syllable (230) is predicted based on its predicted duration; and Multiple fixed-length predicted pitch frames are generated based on the predicted duration of the syllable (230), each fixed-length pitch frame representing the predicted pitch of the syllable (230). The prosodic features (322) received as input to the vocoder model (400) include the plurality of fixed-length predicted pitch frames generated for each syllable (230) of the text utterance (320).

21. The system (100) according to claim 20, wherein, The operation further includes, for each syllable (230), using the selected utterance embedding (204): The energy level of each phoneme (220) in the syllable (230) is predicted based on the predicted duration of the syllable (230); as well as For each phoneme (220) in the syllable (230), a plurality of fixed-length predicted energy frames (280) are generated based on the predicted duration of the syllable (230), each fixed-length predicted energy frame representing the predicted energy level of the corresponding phoneme (220). The prosodic features (322) received as input to the vocoder model (400) further include the plurality of fixed-length predicted energy frames (280) generated for each phoneme (220) in each syllable (230) of the text utterance (320).

22. The system (100) according to claim 21, wherein, The prosodic model (300) is incorporated into a hierarchical language structure (200) to represent the text discourse (320), the hierarchical language structure (200) comprising: The first level includes a Long Short-Term Memory (LSTM) processing unit representing each word (240) of the text discourse (320); The second level includes an LSTM processing unit representing each syllable (230) of the text utterance (320), and the LSTM processing unit of the second level times faster than the LSTM processing unit of the first level. The third level includes an LSTM processing unit representing each phoneme (220) of the text utterance (320), and the LSTM processing unit of the third level times faster than the LSTM processing unit of the second level. The fourth stage includes an LSTM processing unit representing each fixed-length predicted pitch frame, the LSTM processing unit of the fourth stage timing faster than the LSTM processing unit of the third stage; and The fifth stage includes an LSTM processing unit representing each fixed-length prediction energy frame, the LSTM processing unit of the fifth stage timing at the same speed as the LSTM processing unit of the fourth stage and timing faster than the LSTM processing unit of the third stage.

23. The system (100) according to claim 22, wherein: The first level of the hierarchical language structure (200) generates the activation of each word (240) of the text discourse (320) in a single first pass; The second level of the hierarchical language structure (200) generates activation of each syllable (230) of the text discourse (320) in a single second pass after the first pass; The third level of the hierarchical language structure (200) generates the activation of each phoneme (220) of the text discourse (320) in a single third pass after the second pass; The fourth level of the hierarchical language structure (200) generates activation of each fixed-length predicted pitch frame in a single fourth pass after the third pass; as well as The fifth level of the hierarchical language structure (200) generates activation of each fixed-length prediction energy frame in a single fifth pass after the third pass.

24. The system (100) according to claim 15, wherein, Each of the plurality of fixed-length speech frames (502) represents a corresponding portion of the predicted vocoder parameters (450) output from the vocoder model (400).

25. The system (100) according to claim 15, wherein: The first level of the hierarchical language structure (200) generates the activation of each word (240) of the text discourse (320) in a single first pass; The second level of the hierarchical language structure (200) generates activation of each syllable (230) of the text discourse (320) in a single second pass after the first pass; The third level of the hierarchical language structure (200) generates the activation of each phoneme (220) of the text discourse (320) in a single third pass after the second pass; as well as The fourth level of the hierarchical language structure (200) generates activation of each fixed-length speech frame in the plurality of fixed-length speech frames (502) in a single fourth pass after the third pass.

26. The system (100) according to claim 15, wherein, The operation further includes: Receive training data, the training data including multiple reference audio signals and corresponding transcripts (206), each reference audio signal including spoken utterances of speech and having corresponding prosody, each transcript (206) including a text representation of the reference audio signals; and For each reference audio signal and its corresponding transcription (206) pair: Obtain the reference linguistic specification (402) of the corresponding transcription (206) and the reference prosodic features (322) representing the corresponding prosodic of the reference audio signal; and The vocoder model (400) is trained using a deep neural network to generate a fixed-length sequence of predicted speech frames (502) based on the reference language specification (402) and the reference prosodic features (322), the fixed-length sequence of predicted speech frames providing Mel-frequency cepstral coefficients, aperiodic components, and vocalization components.

27. The system (100) according to claim 26, wherein, Training the vocoder model (400) further includes, for each reference audio signal: A fixed-length sequence of reference speech frames (502) is sampled from the reference audio signal, the fixed-length sequence of reference speech frames providing reference Mel-spectral coefficients, reference aperiodic components and reference vocal components of the reference audio signal; The gradient or loss generated between the fixed-length predicted speech frame (502) sequence generated by the vocoder model (400) and the fixed-length reference speech frame (502) sequence sampled from the reference audio signal; and The gradient or loss is backpropagated through the vocoder model (400).

28. The system (100) according to any one of claims 15 to 27, wherein, The operation further includes: The predicted vocoder parameters (450) output from the vocoder model (400) are decomposed into Mel-frequency cepstral coefficients, aperiodic components, and vocal components. Denormalize the Mel-frequency cepstral coefficients, aperiodic components, and vocal components separately; and The prosodic features (322), denormalized Mel-Cepstral coefficients, denormalized aperiodic components, and denormalized vocal components output from the prosodic model (300) are concatenated into a vocoder vector (460). Providing the predicted vocoder parameters (450) output from the vocoder model (400) and the prosodic features (322) output from the prosodic model (300) to the parameterized vocoder (155) includes providing the vocoder vector (460) to the parameterized vocoder (155) as input to the synthesized speech representation (152) for generating the text utterance (320).

Citation Information

Patent Citations

  • Methods and apparatus for predicting prosody in speech synthesis

    US20120191457A1

  • Speech synthesis apparatus, method, and computer-readable medium

    US20140052447A1