End-to-End Speech Waveform Generation by Estimation of Data Density Gradient
A non-autoregressive method directly generates waveforms from phoneme sequences, addressing inefficiencies in existing autoregressive models by reducing iterations and computational requirements, and enhancing generalization through end-to-end training.
Patent Information
- Application Number
- JP2022580985
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-02
- Filing Date
- 2021-09-02
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-09-02
AI Technical Summary
Existing machine learning models for generating waveforms, particularly in text-to-speech synthesis, are inefficient due to their autoregressive nature, requiring numerous iterations and significant computational resources, and often rely on intermediate representations like mel spectrograms, limiting their ability to generalize to new inputs.
A non-autoregressive method that directly generates waveforms from phoneme sequences using a noise estimation neural network, iteratively improving the waveform output through gradient-based sampling, eliminating the need for intermediate representations and reducing the number of required iterations.
This approach generates high-fidelity speech samples with reduced latency and computational resources, outperforming state-of-the-art autoregressive models while enabling end-to-end training and improved generalization to new inputs.
Smart Images

Figure 0007708793000018 
Figure 0007708793000019 
Figure 0007708793000020
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority to U.S. Patent Application No. 63 / 073,867, filed Sep. 2, 2020, and U.S. Patent Application No. 63 / 170,401, filed Apr. 2, 2021, the disclosures of which are incorporated herein by reference.
[0002] This specification relates to generating waveforms conditioned on text sequences using a machine learning model.
Background Art
[0003] A machine learning model receives an input and generates an output, for example, a predicted output, based on the received input. Some machine learning models are parametric models and generate an output based on the received input and the values of the parameters of the model.
[0004] Some machine learning models are deep models that use multiple layers of the model to generate an output for the received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non - linear transformation to the received input to generate an output.
Prior Art Documents
Non - Patent Documents
[0005]
Non - Patent Document 1
Summary of the Invention
Means for Solving the Problem
[0006] This specification describes a system implemented as a computer program on one or more computers at one or more locations that generates waveforms conditioned on text sequences.
[0007] According to a first aspect, a method performed by one or more computers includes: obtaining a phoneme sequence, the phoneme sequence including respective phoneme tokens at each of a plurality of input time steps; processing the phoneme sequence using an encoder neural network to generate a hidden representation of the phoneme sequence; generating a conditioning input from the hidden representation; initializing a current waveform output; and generating a final waveform output that defines an utterance of the phoneme sequence by a speaker by updating the current waveform output in each of a plurality of iterations, each iteration corresponding to a respective noise level, and the updating being performed using a noise estimation neural network configured to process a model input to generate a noise output in each iteration, the processing including processing a model input for the iteration that includes (i) the current waveform output and (ii) the conditioning input, the noise output including respective estimated values of noise for each value in the current waveform output, and updating the current waveform output using the estimated values of noise and the noise level for the iteration.
[0008]
[0009] In some implementations, updating the current waveform output using the estimated value of noise and the noise level for the iteration includes at least generating an update for the iteration from the estimated value of noise and the noise level corresponding to the iteration, and subtracting the update from the current waveform output to generate the first updated output waveform.
[0010] In some implementations, updating the current waveform output further includes modifying the first updated output waveform based on the noise level for the iteration to generate a modified first updated output waveform.
[0011] In some implementations, for the last iteration, the modified first updated output waveform is the updated output waveform after the last iteration, and for each iteration before the last iteration, the updated output waveform after the last iteration is generated by adding noise to the modified first updated output waveform.
[0012] In some implementations, the step of initializing the current waveform output includes sampling each of a plurality of initial values for the current waveform output from the corresponding noise distribution.
[0013] In some implementations, the model input for each iteration includes iteration-specific data that is different for each iteration.
[0014] In some implementations, the model input for each iteration includes the noise level corresponding to the iteration.
[0015] In some implementations, the model input for each iteration includes an aggregate noise level for the iteration generated from the noise levels corresponding to the iteration and any subsequent iterations among the plurality of iterations.
[0016] In some implementations, the noise estimation neural network includes a plurality of noise generation neural network layers and a noise generation neural network configured to process conditioning inputs to map the conditioning inputs to noise outputs, and an output waveform processing neural network including a plurality of output waveform processing neural network layers configured to process a current waveform output to generate an alternative representation of the current waveform output, wherein at least one of the noise generation neural network layers receives an input derived from (i) an output of another one of the noise generation neural network layers, (ii) an output of the corresponding output waveform processing neural network layer, and (iii) iteration-specific data regarding the iteration.
[0017] In some implementations, the final output waveform has a higher dimensionality than the conditioning input, and the alternative representation has the same dimensionality as the conditioning input.
[0018] In some implementations, the noise estimation neural network includes a respective feature-wise linear modulation (FiLM) module corresponding to each of at least one noise generation neural network layer, and the FiLM module corresponding to a given noise generation neural network layer is configured to process (i) an output of another one of the noise generation neural network layers, (ii) an output of the corresponding output waveform processing neural network layer, and (iii) iteration-specific data regarding the iteration to generate an input to the noise generation neural network layer.
[0019] In some implementations, a FiLM module corresponding to a given noise generation neural network layer generates a scale vector and a bias vector from (ii) the output of the corresponding output waveform processing neural network layer and (iii) iteration-specific data regarding the iteration, and is configured to generate an input to the given noise generation neural network layer by applying an affine transformation to another output of the noise generation neural network layer.
[0020] In some implementations, at least one of the noise generation neural network layers includes an activation function layer that applies a non-linear activation function to the input to the activation function layer.
[0021] In some implementations, another one of the noise generation neural network layers corresponding to the activation function layer is a residual connection layer or a convolutional layer.
[0022] In some implementations, the latent representation includes respective hidden vectors for each of a plurality of hidden time steps, and the step of generating a conditioning input from the latent representation includes, for each hidden time step, processing the latent representation using a duration predictor neural network to generate a predicted duration in the utterance characterized by the hidden vector at the hidden time step, and upsampling the latent representation according to the predicted duration so as to match the time scale of the final waveform output, to generate a conditioning input.
[0023] In some implementations, the conditioning input includes respective conditioning vectors for each of a plurality of quantized time segments within the final waveform output.
[0024] In some implementations, the method further includes, for each hidden time step, generating a predicted influence range in the utterance characterized by the hidden vector at the hidden time step, and upsampling the hidden representation includes applying Gaussian upsampling to the hidden representation using the predicted duration and the predicted influence range for the hidden time step.
[0025] Certain embodiments of the subject matter described herein can be implemented to realize one or more of the following advantages.
[0026] The techniques described generate output waveforms in a non-autoregressive manner directly from text sequences. Generally, autoregressive models have been shown to generate high-quality waveforms, but require many iterations, leading to large latency as well as consumption of resources such as memory and processing power. This is because autoregressive models generate each given output in the output waveform one by one, and each is conditioned on all of the outputs preceding the given output in the output waveform.
[0027] On the one hand, the described technique starts with an initial output waveform, e.g., a noisy output containing values sampled from a noise distribution, and iteratively improves the output waveform by a gradient-based sampler conditioned on a conditioning input generated from a text sequence, i.e., an iterative noise removal process may be used. As a result, the approach is not autoregressive and requires only a fixed number of generation steps during inference. For example, with respect to speech synthesis, the described technique can generate high-fidelity speech samples that rival or even exceed the speech samples generated by state-of-the-art autoregressive models with significantly reduced latency and using much fewer computational resources, with a very small number of iterations, e.g., 6 or fewer iterations. Further, the described technique can generate speech samples of higher quality (e.g., more faithful) than those generated by existing non-autoregressive models.
[0028] Furthermore, unlike previous approaches to non-autoregressive generation, the described technique generates speech directly from a sequence of phonemes, i.e., without requiring an intermediate structured representation such as the features of a mel spectrogram. This enables the system to generate speech without using a separate model to generate the features of the spectrogram. Also, this enables the system to be trained completely end-to-end, thus improving the system's ability to generalize to new inputs and perform better in various text-to-speech tasks after training.
[0029] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
Brief Description of the Drawings
[0030]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
[0031] Like reference numerals and reference designations in the various drawings indicate like elements.
[0032] FIG. 1 shows an exemplary end-to-end waveform system 100. The end-to-end waveform system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations where the systems, components, and techniques described below are implemented.
[0033] The end-to-end waveform system 100 generates a final waveform output 104 conditioned on a text sequence 102. The system 100 initializes a current waveform output 116 (e.g., by sampling from a noise distribution such as a Gaussian noise distribution) and updates the current waveform output in each of a plurality of iterations using a noise level 106 and a conditioning input 114 generated from the text sequence 102. The system 100 outputs the updated output waveform after the last iteration as the final output waveform 104.
[0034] In some implementations, the final waveform output 104 may represent an utterance in which the text sequence 102 is spoken by a speaker. The text sequence 102 may be represented by a series of phoneme tokens including silence tokens inserted at word boundaries and sequence end tokens inserted after each sentence. For example, the sequence of phonemes may be generated from the text or linguistic features of the text such that the waveform generated by the system represents the utterance of the text being spoken. The output waveform generated by the system can be a sequence of amplitude values at a specified output frequency, such as raw amplitude values, compressed amplitude values, or companded amplitude values. Once generated, the system can provide the output waveform to play the utterance, or play the utterance using an audio output device.
[0035] The end-to-end waveform system 100 uses an encoder neural network 108 and a duration predictor network 112 to generate a conditioning input 114 from the text sequence 102. The conditioning input 114 may include respective conditioning vectors for each of a plurality of quantized time segments within the final waveform output. The quantized time segments may represent the resolution of the final output waveform with respect to the conditioning input 114 (e.g., a 10 ms time segment). That is, for an exemplary final output waveform of 0.8 s, the final output waveform may be divided into a sequence of 80 10 ms quantized time segments. At that time, the conditioning input 114 can include 80 conditioning vectors, each conditioning vector corresponding to a respective 10 ms quantized time segment.
[0036] System 100 generates a hidden representation 110 of text sequence 102 by processing text sequence 102 using encoder neural network 108. The hidden representation may include respective hidden vectors for each of a plurality of hidden time steps. The final output waveform 104 can have a higher frequency than the hidden representation 110, and the hidden representation 110 can have the same frequency as the text sequence (e.g., can have the same frequency as a phoneme sequence representing the text sequence, where each hidden vector corresponds to a phoneme token). For example, each hidden vector can be represented by an ordered set of numerical values, such as a vector of numerical values.
[0037] Encoder neural network 108 can have any suitable neural network architecture that enables encoder neural network 108 to perform its described function, i.e., process a text sequence to generate a hidden representation of the text sequence. In particular, the encoder neural network can include any suitable number (e.g., 1, 5, or 25) of any suitable type of neural network layer (e.g., fully connected layer, attention layer, convolutional layer, etc.) connected in any suitable configuration (e.g., as a linear sequence of layers). For example, the encoder neural network can be a recurrent neural network, such as an LSTM or GRU neural network. In a particular example, the encoder neural network can include an embedding layer, followed by one or more convolutional layers, and then one or more bidirectional LSTM neural network layers.
[0038] System 100 generates the conditioning input 114 from the latent representation 110 by "upsampling" (i.e., increasing the frequency) the latent representation 110. To upsample the latent representation 110, System 100 generates, for each latent time step, the respective predicted duration of the utterance of the text sequence. The predicted duration of a latent time step corresponds to the predicted duration of the utterance characterized by the latent vector at the latent time step.
[0039] System 100 generates the predicted durations by processing the latent representation using the duration predictor network 112. The predicted durations can be represented, for example, by an ordered set of numerical values, such as a vector of numerical values, where each value corresponds to a different latent time step. That is, the duration predictor neural network 112 is configured to process the latent representation to generate, for each latent time step, the respective predicted duration, for example, the integer duration (e.g., number of frames) of the utterance characterized by the latent vector at the latent time step, or the integer duration measured in any unit of time (e.g., seconds) of the utterance characterized by the latent vector at the time step.
[0040] The duration predictor network 112 can have any suitable neural network architecture that enables the duration predictor network 112 to perform the described function, i.e., process the hidden representation of the text sequence to generate the respective predicted durations for each hidden time step within the hidden representation. In particular, the duration predictor neural network can include any suitable number (e.g., 1 layer, 5 layers, or 25 layers) of any suitable type of neural network layer (e.g., fully connected layer, attention layer, convolutional layer, etc.) connected in any suitable configuration (e.g., as a linear sequence of layers). For example, the duration predictor neural network can be a recurrent neural network, such as an LSTM or GRU neural network. In a particular example, the duration predictor neural network can include one or more bidirectional LSTM layers followed by a projection layer.
[0041] For example, system 100 can upsample the latent representation using Gaussian upsampling. To perform Gaussian upsampling, system 100 can further process the predicted duration and the latent representation to generate, for each latent time step, a respective predicted extent of influence of the utterance characterized by the latent vector at the latent time step (e.g., represented by a positive numerical value). The predicted extent of influence of the utterance can represent the extent of influence of the utterance. Then, the system can upsample the latent representation using the predicted duration and extent of influence (e.g., each predicted duration-extent of influence pair is, respectively, the mean and standard deviation of a Gaussian distribution with respect to the latent time step). An example of Gaussian upsampling is described in more detail with reference to Jonathan Shen et al., "Non-Attentive Tacotron: Robust And Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling", arXiv:2010.04301v4, May 11, 2021, which is incorporated herein by reference.
[0042] In some implementations, the predicted extent of influence can be generated by an extent of influence predictor neural network. The extent of influence predictor neural network can process the latent representation concatenated with the predicted duration to generate the predicted extent of influence. For example, the extent of influence predictor neural network can be a recurrent neural network, such as an LSTM or GRU. In a particular example, the extent of influence predictor neural network can include one or more bidirectional LSTM layers, followed by a projection layer, and then a softplus layer.
[0043] System 100 initializes the current waveform output 116 and updates the current waveform output in each of a plurality of iterations using the noise level 106 and the conditioning input 114 generated from the text sequence 102. System 100 outputs the updated waveform output after the last iteration as the final waveform output 104.
[0044] For example, System 100 can initialize the current output waveform 116 by sampling each value in the current output waveform from a corresponding noise distribution (e.g., a Gaussian distribution such as N(0, I) where I is the identity matrix) (i.e., generate the first instance of the current output waveform 116). That is, the first current output waveform 116 contains the same number of values as the final output waveform 104, but each value is sampled from the corresponding noise distribution.
[0045] And System 100 generates the final output waveform 104 by updating the current output waveform 116 in each of a plurality of iterations using the conditioning input 114 generated from the text sequence 102. In other words, the final output waveform 104 is the current output waveform 116 after the last iteration of the plurality of iterations.
[0046] In some cases, the number of iterations is determined.
[0047] In other cases, System 100 or another system can adjust the number of iterations based on the latency requirements for generating the final output waveform. That is, System 100 can select the number of iterations such that the final output waveform 104 is generated to meet the latency requirements.
[0048] In other cases, system 100 or another system can adjust the number of iterations based on the computational resource consumption requirements for generating the final output waveform 104, i.e., can select the number of iterations such that the final output waveform is generated to meet the requirements. For example, the requirement can be the maximum number of floating-point operations (FLOPS) executed as part of generating the final output waveform.
[0049] In each iteration, the system processes the model input for the iteration using the noise estimation neural network 300, which includes (i) the current output waveform 116, (ii) the conditioning input 114, and optionally (iii) iteration-specific data for the iteration. The iteration-specific data is generally derived from the noise level 106 (e.g., if each noise level corresponds to a particular iteration). The system can update the current output waveform using the noise level 106 as a scale for each iteration of the update. That is, each noise level of the noise level 106 can correspond to a particular iteration, and each noise level regarding the iteration can lead to the scale of the update for the current output waveform 116 in the iteration.
[0050] The noise estimation neural network 300 has parameters (“network parameters”) and is configured to process the model input according to the current values of the network parameters to generate a noise output 110 that includes an estimated value of the respective noise for each value within the current output waveform 116. Details of the noise estimation neural network are considered in more detail below in connection with FIG. 3.
[0051] Generally, the estimated value of the noise for a given value within the current output waveform is the estimated value of the noise added to the corresponding actual value within the actual output waveform for the text sequence to generate the given value. That is, the estimated value of the noise defines how the actual value needs to be modified to generate a given value within the current output waveform, given the noise level corresponding to the current iteration. In other words, a given value may be generated by applying the estimated value of the noise to the actual value according to the noise level for the current iteration.
[0052] This estimated value of the noise can be interpreted as an estimated value of the gradient of the data density, and thus the generation process can be considered as a process of repeatedly generating the output waveform by estimating the data density.
[0053] And the system 100 uses the update engine 120 to update the current output waveform 116 in the direction of the estimated value of the noise.
[0054] In particular, the update engine 120 uses the estimated value of the noise and the corresponding noise level for the iteration to update the current output waveform 116. That is, the update engine 120 updates each value of the current output waveform 116 using the corresponding estimated value of the noise of the noise output 110 in the iteration and the corresponding noise level, as will be discussed in more detail in relation to FIG. 2.
[0055] After the last iteration, the conditional output generation system 100 outputs the updated output waveform 116 as the final output waveform 104. For example, in an implementation where the final output waveform 104 represents an audio waveform, the system can play the audio using a speaker or transmit the audio for playback or the like. In some implementations, the system 100 can save the final output waveform 104 to a data store or transmit the final output waveform 104 to be stored remotely.
[0056] Before system 100 generates the final output waveform using the noise estimation neural network 300, the duration predictor neural network 112, and the encoder neural network 108, system 100 or another system trains the neural networks using training data. In implementations that include a range influence neural network, the system also trains the range influence neural network using training data. The training data can include a plurality of training text sequences and, for each training text sequence, a respective target duration and a respective ground truth waveform, as described with reference to FIG. 7.
[0057] The training data can include a respective target duration for each text sequence (e.g., for each phoneme of a phoneme sequence if each text sequence is represented by a respective phoneme sequence). The duration predictor neural network can be trained using a loss function that has at least a duration loss term that measures the error between the target duration and the predicted duration for each training text sequence. The predicted duration can be used for the loss function for training the predictor neural network, and the target duration can be used for upsampling the hidden representation during training.
[0058] FIG. 2 is a flow diagram of an exemplary process 200 for generating an output conditioned on a text sequence. For convenience, process 200 is described as being executed by a system of one or more computers located in one or more locations. For example, a conditional output generation system appropriately programmed according to this specification, such as conditional output generation system 100 of FIG. 1, can execute process 200.
[0059] The system obtains (202) a text sequence that conditions the final output waveform. For example, the text sequence can be represented by a sequence of phonemes (e.g., generated from text or linguistic features of the text), and the final output waveform can represent the utterance of the sequence of phonemes by a speaker.
[0060] The system generates (204) a hidden representation of the text sequence using an encoder neural network. The hidden representation can include hidden vectors for each of a plurality of hidden time steps. In a particular example, the encoder neural network can include an embedding layer, one or more convolutional layers, and one or more bidirectional LSTM layers.
[0061] The system generates a conditioning input from a latent representation (206). The system can use upsampling to generate a conditioning input from the latent representation. For example, the system can use Gaussian upsampling to generate a conditioning input. The system can generate a predicted duration for each latent time step within the latent representation by processing the latent representation using a duration predictor neural network. Further, the system can generate a predicted extent of influence for each latent time step within the latent representation by processing, for example, a concatenated latent representation combined with the predicted duration using an extent of influence predictor neural network. Then, the system can use Gaussian upsampling to generate a conditioning input from the latent representation, the predicted duration, and the predicted extent of influence. An example of Gaussian upsampling is described in more detail with reference to Jonathan Shen et al., "Non-Attentive Tacotron: Robust And Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling", arXiv:2010.04301v4, May 11, 2021, which is incorporated herein by reference.
[0062] The system initializes the current output waveform (208). For the final output waveform containing a plurality of values, the system can sample each value within the first current output waveform having the same number of values as the final output waveform from a noise distribution. For example, the system can initialize the current output waveform using a noise distribution represented by y N ~N(0, I) (e.g., a Gaussian noise distribution), where I is the identity matrix and N in y N represents the intended number of iterations. The system can update the first current output waveform in N iterations in descending order from iteration N to iteration 1.
[0063] At that time, the system updates the current output waveform in each of a plurality of iterations. Generally, the current output waveform in each iteration can be interpreted as the final output waveform with additional noise. That is, the current output waveform is a noisy version of the final output waveform. For example, for the first current output waveform y N with respect to, where N represents the number of iterations, the system can update the current output waveform in each of iterations from N to 1 by removing the estimated value of the noise corresponding to the iteration. That is, the system can improve the current output waveform in each iteration by determining the estimated value of the noise and updating the current output waveform according to the estimated value. The system may use a descending order of iterations until the final output waveform y0 is output.
[0064] In each of a plurality of iterations, the system uses a noise estimation neural network to generate a noise output for the iteration (210) by processing a model input that includes (1) the current output waveform, (2) a conditioning input, and optionally (3) iteration-specific data regarding the iteration. The iteration-specific data is generally derived from the noise level regarding the iteration, and each noise level corresponds to a specific iteration. The noise output may include an estimated value of the noise for each value within the current output waveform. For example, each estimated value of the noise regarding a specific value within the current output waveform may represent the estimated value of the noise added to the corresponding actual value within the actual output waveform regarding the phoneme sequence to generate the specific value. That is, the estimated value of the noise regarding a specific value represents how the actual value needs to be corrected if a corresponding noise level is given to generate the specific value when the actual value is known.
[0065] In each of multiple iterations, the system updates the current output waveform at the current iteration (212) using the noise output for the current iteration and the noise level corresponding to the current iteration. The system can update each value in the current output waveform using the estimated value of the corresponding noise in the noise output and the noise level for the current iteration. The system can generate an update for the iteration from the estimated value of the noise and the noise level for the iteration, and then subtract the update from the current output waveform to generate a first updated output waveform. And the system modifies the first updated output waveform based on the noise level for the iteration to obtain a modified first updated output waveform,
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0066] For the last iteration, the modified first updated output waveform is the updated output waveform after the last iteration, and for each iteration before the last iteration, the updated output waveform after the last iteration is generated by adding noise to the modified first updated output waveform. That is, when the iteration is not the last iteration (i.e., when n > 1), the system takes the modified first updated output waveform, y n-1 = y n-1 + σn z(3) is further updated as such, where n indexes the iteration, and σ n is the noise schedule
Number
[0067] The system determines whether an end criterion is met (214). For example, the end criterion may include having executed a specific number of iterations (determined to meet maximum computational resource requirements such as, e.g., a minimum performance metric, a maximum latency requirement, or a maximum number of FLOPS). If the specific number of iterations has not been executed, the system can start again from step (206) and perform another update on the current output waveform.
[0068] If the system determines that the end criterion is met, the system outputs the final output waveform, which is the updated output waveform after the last iteration (216).
[0069] Process 200 can be used to generate an output waveform in a non-autoregressive manner conditioned on a phoneme sequence. Generally, autoregressive models have been shown to generate high-quality output waveforms, but they require a large number of iterations, leading to significant latency as well as consumption of resources such as memory and processing power. This is because autoregressive models generate each given output in the output waveform one by one, and each is conditioned on all of the outputs preceding the given output in the output waveform. On the other hand, process 200 starts with an initial output waveform, e.g., a noisy output containing values sampled from a noise distribution, and iteratively improves the output waveform by a gradient-based sampler conditioned on the phoneme sequence. As a result, the approach is non-autoregressive and requires only a fixed number of generation steps during inference. For example, with respect to speech synthesis conditioned on a spectrogram, the described technique can generate high-fidelity speech samples that rival or even exceed the speech samples generated by state-of-the-art autoregressive models with significantly reduced latency and using far fewer computational resources, with very few iterations, e.g., six or fewer iterations.
[0070] Furthermore, unlike previous approaches to non-autoregressive generation, the described technique generates speech directly from a sequence of phonemes, i.e., without requiring an intermediate structured representation such as the features of a mel spectrogram. This enables the system to generate speech without using a separate model to generate the features of the spectrogram. Also, this enables the system to be trained completely end-to-end, thus improving the system's ability to generalize to new inputs and perform better in various text-to-speech tasks after training.
[0071] FIG. 3 shows an exemplary architecture of a noise estimation network 300.
[0072] The exemplary noise estimation network 300 includes multiple types of neural network layers and neural network blocks, such as a convolutional neural network layer, a noise generation neural network block, a Feature-wise Linear Modulation (FiLM) module neural network block, and an output waveform processing neural network block (e.g., each neural network block includes multiple neural network layers).
[0073] To generate the noise output 118, the noise estimation network 300 processes a model input that includes (1) the current output waveform 116, (2) the conditioning input 114, and (3) iteration-specific data including the aggregated noise level 306 corresponding to the current iteration. The current output waveform 116 has a higher dimension (e.g., frequency) than the conditioning input 114, and the noise output 118 has the same dimension (e.g., frequency) as the current output waveform 116. For example, the current output waveform can represent an audio waveform at 24 kHz, and the conditioning input can be represented at 80 Hz.
[0074] The noise estimation network 300 includes multiple output waveform processing blocks for processing the current output waveform 116 to generate alternative representations of the current output waveform 116.
[0075] The noise estimation network 300 also includes an output waveform processing block 400 for processing the current output waveform 116 to generate an alternative representation of the current output waveform, and the alternative representation has a smaller dimension (e.g., frequency) than the current output waveform.
[0076] The noise estimation network 300 further includes additional output waveform processing blocks (e.g., output waveform processing blocks 318, 316, 314, and 312) for processing the alternative representation generated by the previous output waveform processing block to generate another alternative representation having a smaller dimension (e.g., frequency) than the previous alternative representation (e.g., network 318 processes the alternative representation from block 400 to generate an alternative representation having a smaller dimension than the output of block 400, block 316 processes the alternative representation from block 318 to generate an alternative representation having a smaller dimension than the output of block 318, etc.). The alternative representation of the current output waveform generated from the last output waveform processing block (e.g., 312) has the same dimension (e.g., frequency) as the conditioning input 114.
[0077] For example, with respect to a current output waveform including a 24 kHz audio waveform and a conditioning input of 80 Hz, the blocks of the output waveform processing block can "downsample" (i.e., lower the frequency) the dimension by a factor of 2, 2, 1 / 3, 1 / 5, and 1 / 5 (e.g., by the output waveform processing blocks 400, 318, 316, 314, and 312, respectively) until the alternative representation generated by the last layer 312 is 80 Hz (i.e., reduced to 1 / 300 to match the conditioning input). The architecture of an exemplary output waveform processing block is considered in more detail in connection with FIG. 4.
[0078] The noise estimation block 300 includes a plurality of FiLM module neural network blocks for processing iteration-specific data corresponding to the current iteration (e.g., the aggregated noise level 306) and alternative representations from the output waveform processing neural network block to generate an input for the noise generation neural network block. Each FiLM module processes the aggregated noise level 306 and an alternative representation from each respective output waveform processing block to generate an input for each respective noise generation block (e.g., the FiLM module 500 processes an alternative representation from the output waveform processing block 400 to generate an input for the noise generation block 600, and the FILM module 328 processes an alternative representation from the output waveform processing block 318 to generate an input for the noise generation block 338, etc.). In particular, each FiLM module generates a scale vector and a bias vector as an input to each respective noise generation block (e.g., as an input to an affine transformation neural network layer within each respective noise generation block), as will be discussed in more detail with reference to FIG. 5.
[0079] The noise estimation network 300 includes a plurality of noise generation neural network blocks for processing the conditioning input 114 and the output from the FiLM module to generate the noise output 118. The noise estimation network 300 may include a convolutional layer 302 for processing the conditioning input 114 to generate an input to the first noise generation block 332, and a convolutional layer 304 for processing the output from the last noise generation block 600 to generate the noise output 118. Each noise generation block generates an output having a higher dimension (e.g., frequency) than the conditioning input 114. In particular, each subsequent noise generation block after the first generates an output having a higher dimension (e.g., frequency) than the output from the previous noise generation block. The last noise generation block generates an output having the same dimension (e.g., frequency) as the current output waveform 116.
[0080] The noise estimation network 300 includes a noise generation block 332 for processing the output from the convolutional layer 302 (i.e., the convolutional layer that processes the conditioning input 114) and the output from the FiLM module 332 to generate an input to the noise generation block 334. The noise estimation network 300 further includes noise generation blocks 336, 338, and 600. The noise generation blocks 334, 336, 338, and 600 process, respectively, the output from the respective previous noise generation block (e.g., block 334 processes the output from block 332, block 336 processes the output from block 334, etc.) and the output from the respective FiLM module (e.g., noise generation block 334 processes the output from FiLM module 324, noise generation block 336 processes the output from FiLM module 326, etc.) to generate an input for the next neural network block. The noise generation block 600 generates an input for the convolutional layer 304 that processes the input to generate the noise output 118. The architecture of an exemplary noise generation block (e.g., noise generation block 600) is discussed in further detail in relation to FIG. 6.
[0081] Each noise generation block prior to the last can generate an output having the same dimension (e.g., frequency) as the corresponding alternative representation of the current output waveform (e.g., noise generation block 332 generates an output having the same dimension as the alternative representation generated by the output waveform processing block 314, noise generation block 334 generates an output having the same dimension as the output from the output waveform processing block 316, etc.).
[0082] For example, with respect to a current output waveform including a 24 kHz audio waveform and an 80 Hz conditioning input, the noise generation block can “upsample” (i.e., increase the frequency) the dimensions 5, 5, 3, 2, and 2 times (e.g., by noise generation blocks 332, 334, 336, 338, and 600 respectively) until the output of the last noise generation block (e.g., noise generation block 600) becomes 24 kHz (i.e., is increased 300 times to match the current output waveform 116).
[0083] Figure 4 shows an exemplary architecture of the output waveform processing block 400.
[0084] The output waveform processing block 400 processes the current output waveform 116 to generate an alternative representation 402 of the current output waveform 116. The alternative representation has a smaller dimension than the current output waveform. The output waveform processing block 400 includes one or more neural network layers. The one or more neural network layers can include multiple types of neural network layers, such as a downsampling layer (e.g., for “downsampling” or reducing the dimension of the input), an activation layer with a non-linear activation function (e.g., a fully connected layer with a Leaky ReLU activation function), a convolutional layer, and a residual connection layer.
[0085] For example, the downsample layer can be a convolutional layer using a stride necessary to reduce the dimension of the input (i.e., “downsample” the input). In a particular example, stride X can be used to reduce the dimension of the input to 1 / X (e.g., stride 2 can be used to reduce the dimension of the input to 1 / 2, stride 5 can be used to reduce the dimension of the input to 1 / 5, etc.).
[0086] The left branch of the residual connection layer 420 includes a convolutional layer 402 and a downsampling layer 404. The convolutional layer 402 processes the current output waveform 116 to generate an input to the downsampling layer 404. The downsampling layer 404 processes the output from the convolutional layer 402 to generate an input to the residual connection layer 420. The output of the downsampling layer 404 has a reduced dimension compared to the current output waveform 116. For example, the convolutional layer 402 can include a filter of size 1x1 with a stride of 1 (i.e., to maintain the dimension), and the downsampling layer 404 can include a filter of size 2x1 with a stride of 2 to downsample the input dimension by a factor of 2.
[0087] The right branch of the residual connection layer 420 includes a downsampling layer 406 and three subsequent blocks where a convolutional layer follows an activation layer (e.g., activation layer 408, convolutional layer 410, activation layer 412, convolutional layer 414, activation layer 416, and convolutional layer 418). The downsampling layer 406 processes the current output waveform 116 to generate an input for the three subsequent blocks of the activation layer and the convolutional layer. The output of the downsampling layer 406 has a smaller dimension compared to the current output waveform 116. The three subsequent blocks process the output from the downsampling layer 406 to generate an input to the residual connection layer 420. For example, the downsampling layer 406 can include a filter of size 2x1 with a stride of 2 to reduce the input dimension by a factor of 2 (e.g., to exactly match the downsampling layer 404). The activation layers (e.g., 408, 412, and 416) can be fully connected layers with a Leaky ReLU activation function. The convolutional layers (e.g., 410, 414, and 418) can include filters of size 3x1 with a stride of 1 (i.e., to maintain the dimension).
[0088] The residual connection layer 420 combines the output from the left branch and the output from the right branch to generate an alternative representation 402. For example, the residual connection layer 420 can add (e.g., element-wise add) the output from the left branch and the output from the right branch to generate the alternative representation 402.
[0089] FIG. 5 shows an exemplary Feature-wise Linear Modulation (FiLM) module 500.
[0090] The FiLM module 500 processes the alternative representation 402 of the current output waveform and the aggregated noise level 306 corresponding to the current iteration to generate a scale vector 512 and a bias vector 516. The scale vector 512 and the bias vector 516 can be processed as inputs to a specific layer (e.g., an affine transformation layer) within their respective noise generation blocks (e.g., the noise generation block 600 of the noise estimation network 300 in FIG. 3). The FiLM module 500 includes a positional encoding function and one or more neural network layers. The one or more neural network layers can include multiple types of neural network layers including a residual connection layer, a convolutional layer, and an activation layer having a non-linear activation function (e.g., a fully-connected layer having a Leaky ReLU activation function).
[0091] The left branch of the residual connection layer 508 includes a positional encoding function 502. The positional encoding function 502 processes the aggregated noise level 306 to generate a positional encoding of the noise level. For example, the aggregated noise level 306 can be multiplied by a positional encoding function 502 that is a combination of a sine function for even-dimensional indices and a cosine function for odd-dimensional indices, similar to the preprocessing of a transformer model.
[0092] The right branch of the residual connection layer 508 includes a convolutional layer 504 and an activation layer 506. The convolutional layer 504 processes the alternative representation 402 to generate an input to the activation layer 506. The activation layer 506 processes the output from the convolutional layer 504 to generate an input to the residual connection layer 508. For example, the convolutional layer 504 can include a filter of size 3x1 with a stride of 1 (to maintain the dimension), and the activation layer 506 can be a fully connected layer with a Leaky ReLU activation function.
[0093] The residual connection layer 508 can combine the output from the left branch (e.g., the output from the positional encoding function 502) and the output from the right branch (e.g., the output from the activation layer 506) to generate an input to both the convolutional layer 510 and the convolutional layer 514. For example, the residual connection layer 508 can add (e.g., element-wise add) the output from the left branch and the output from the right branch to generate an input to two convolutional layers (e.g., 510 and 514).
[0094] The convolutional layer 510 processes the output from the residual connection layer 508 to generate a scale vector 512. For example, the convolutional layer 510 can include a filter of size 3x1 with a stride of 1 (to maintain the dimension).
[0095] The convolutional layer 514 processes the output from the residual connection layer 508 to generate a bias vector 516. For example, the convolutional layer 514 can include a filter of size 3x1 with a stride of 1 (to maintain the dimension).
[0096] Figure 6 shows an exemplary architecture of the noise generation block 600.
[0097] The noise generation block 600 processes the input 602 and the output from the FiLM module 500 to generate the output 310. The input 602 can be a phoneme sequence processed by one or more previous neural network layers (e.g., the noise generation blocks 338, 336, 334, 332 in FIG. 3 and the convolutional layer 302). The output 310 can be an input to a subsequent convolutional layer (e.g., the convolutional layer 304 in FIG. 3) that processes the output 310 to generate the noise output 118. The noise generation block 600 includes one or more neural network layers. The one or more neural network layers can include multiple types of neural network layers such as an activation layer with a non-linear activation function (e.g., a fully connected layer with a Leaky ReLU activation function), an upsampling layer (e.g., which "upsamples" or increases the dimension of the input), a convolutional layer, an affine transformation layer, and a residual connection layer.
[0098] For example, an upsampling layer can be a neural network layer that "upsamples" (i.e., increases) the dimension of the input. That is, the upsampling layer generates an output that has a higher dimension than the input to the layer. In a specific example, the upsampling layer can generate an output that has X copies of each value in the input in order to make the dimension of the output X times larger compared to the input (e.g., for the input (2, 7, -4), an output with two copies of each value like (2, 2, 7, 7, -4, -4), or an output with five copies of each value like (2, 2, 2, 2, 2, 7, 7, 7, 7, 7, -4, -4, -4, -4, -4), etc.). Generally, the upsampling layer can fill each additional spot in the output with the nearest value in the input.
[0099] The left branch of the residual connection layer 618 includes an upsampling layer 602 and a convolutional layer 604. The upsampling layer 602 processes the input 602 to generate an input to the convolutional layer 604. The input to the convolutional layer has a higher dimension than the input 602. The convolutional layer 604 processes the output from the upsampling layer 602 to generate an input to the residual connection layer 618. For example, the upsampling layer can double the dimension of the input by generating an output that has two copies of each value in the input 602. The convolutional layer 604 can include filters of dimension 3x1 and stride 1 (e.g., to maintain the dimension).
[0100] The right branch of the residual connection layer 618 includes an activation layer 606 (e.g., a fully-connected layer with a Leaky ReLU activation function), an upsampling layer 608, a convolutional layer 610 (e.g., with a filter size of 3x1 and a stride of 1), an affine transformation layer 612, an activation layer 614 (e.g., a fully-connected layer with a Leaky ReLU activation function), and a convolutional layer 616 (e.g., with a filter size of 3x1 and a stride of 1) in this order.
[0101] The activation layer 606 processes the input 602 to generate an input to the upsampling layer 608. The upsampling layer increases the dimension of the output from the activation layer 606 to generate an input to the convolutional layer 610 that has a higher dimension (e.g., twice as high as the input 602, to match the upsampling layer 602). The convolutional layer 610 processes the output from the upsampling layer 608 to generate an input to the affine transformation layer 612 (e.g., using filters of dimension 3x1 and stride 1 to maintain the dimension). The activation layer 614 and the convolutional layer 616 further process the output from the affine transformation layer 612 to generate an input to the residual connection layer 618 (e.g., using the Leaky ReLU function for the network 614 and filters of dimension 3x1 and stride 1 for the network 616).
[0102] For example, the affine transformation function can process the output from a preceding neural network layer (e.g., the convolutional layer 610 of the noise generation block 600) and the output from the FiLM module to generate an output. For example, the FiLM module can generate a scale vector and a bias vector. The affine transformation layer can add the bias vector to the result of scaling the output from the previous neural network layer (e.g., using the Hadamard product, or element-wise multiplication) using the scale vector from the FiLM module.
[0103] The affine transformation layer 612 can process the output from the convolutional layer 610 and the output from the FiLM module 500 to generate an input to the activation layer 614. For example, by adding the bias vector from the FiLM module 500 to the result of scaling the output from the convolutional layer 610 using the scale vector from the FiLM module 500.
[0104] The residual connection layer 618 combines the output from the left branch (e.g., the output from the convolutional layer 604) and the output from the right branch (e.g., the output from the convolutional layer 616) to generate an output. For example, the residual connection layer 618 can sum the output from the left branch and the output from the right branch to generate an output.
[0105] The left branch of the residual connection layer 632 includes the output from the residual connection layer 618. The left branch can be interpreted as the identity function of the output from the residual connection layer 618.
[0106] The right branch of the residual connection layer 632 processes the output from the residual connection layer 618 and includes two consecutive blocks in the order of an affine transformation layer, an activation layer, and a convolutional layer to generate an input to the residual connection layer 632. Specifically, the first block includes the affine transformation layer 620, the activation layer 622, and the convolutional layer 624. The second block includes the affine transformation layer 626, the activation layer 628, and the convolutional layer 630.
[0107] For example, for each block, each affine transformation layer can process the output from the FiLM module 500 and the output from each previous neural network layer to generate its respective output (e.g., the affine transformation layer 620 can process the output from the residual connection layer 618, and the affine transformation layer 626 can process the output from the convolutional layer 624). Each affine transformation layer can generate its respective output by scaling the output from the previous neural network layer using the scale vector from the FiLM module 500 and summing the result with the bias vector from the FiLM module 500. Each activation layer (e.g., 620 and 628) can be a respective fully-connected layer having a Leaky ReLU activation function. Each convolutional layer can include respective filters of dimension 3x1 and stride 1 (e.g., to maintain the dimension).
[0108] The residual connection layer 632 combines the output from the left branch (e.g., the identity function of the output from the residual connection layer 618) and the output from the right branch (e.g., the output from the convolutional layer 630) to generate the output 310. For example, the residual connection layer 632 can sum the output from the left branch and the output from the right branch to generate the output 310. The output 310 can be an input to a convolutional layer (e.g., the convolutional layer 304 in FIG. 3) that generates the noise output 118.
[0109] The noise generation block 600 can include a plurality of channels. Each noise generation block in FIG. 3 (e.g., 600, 338, 336, 334, and 332) can include a respective number of channels. For example, the noise generation blocks 600, 338, 336, 334, and 332 can include 128, 128, 256, 512, and 512 channels, respectively.
[0110] FIG. 7 is a flowchart of an exemplary process for training an end-to-end waveform system. For convenience, process 700 is described as being performed by a system of one or more computers located in one or more locations.
[0111] The system can execute process 700 in each of multiple training iterations to repeatedly update the values of the parameters of the neural network of the end-to-end waveform system.
[0112] The system obtains (702) a batch of triples of a training text sequence, each target duration, and the corresponding training output waveform. For example, the system can randomly sample triples of training from a data store. The training output waveform can represent an utterance (e.g., a recording of a speaker) of the training text sequence (e.g., represented by a sequence of phonemes generated from the text or linguistic features of the text). Each target duration can represent a set of ground truth durations to be predicted by a duration predictor neural network from the text sequence.
[0113] For each training triplet within a batch, the system generates a conditioning input from the training text sequence (704). The system can generate a conditioning input from the training text sequence in the same way as steps (204)-(206) from FIG. 2. For example, the system can use an encoder neural network to generate a hidden representation of the text sequence (e.g., including respective hidden vectors for each of a plurality of hidden time steps). Then, the system can use a duration predictor neural network to generate respective predicted durations for each hidden time step within the hidden representation. These predicted durations can be used for a duration loss term (e.g., an L2 loss term including the sum of squared errors between the predicted duration and the target duration for each pair of corresponding predicted and target durations of the text sequence) that measures the error between the predicted duration for a hidden time step within the hidden representation of the text sequence and the target duration. The predictor neural network can be trained using the gradient of a loss function (e.g., a linear combination of the loss function from step (714) and the duration prediction loss) that includes at least the duration prediction loss term, and an appropriate optimization method (e.g., ADAM).
[0114] The system can upsample (i.e., increase the dimensionality of) the hidden representation to generate a conditioning input using the target duration (e.g., using the influence range neural network and Gaussian upsampling as considered in FIG. 2).
[0115] For each training triplet within a batch, the system selects iteration-specific data from a set that includes iteration-specific data for all of the iterations (706). For example, the system can sample a particular iteration from a discrete uniform distribution that includes integers from 1 to the last iteration, and then select iteration-specific data based on the particular iteration sampled from the distribution. The iteration-specific data can include a noise level, an aggregated noise level (e.g., determined by Equation (2)), or the iteration number itself. Thus, the system can condition the noise estimation neural network on a discrete index, or can condition the noise estimation neural network on a continuous scalar indicating the noise level. Conditioning on a continuous scalar indicating the noise level can be advantageous because when the noise estimation neural network is trained, different numbers of improvement steps (i.e., iterations) can be used when generating the final network output during inference.
[0116] For each training triplet within a batch, the system samples a noisy output that includes respective noise values for each value within the training output waveform (708). For example, the system can sample a noisy output from a noise distribution. In a particular example, the noise distribution can be a Gaussian noise distribution (e.g., N(0, I) where I is the identity matrix of dimension n x n and n is the number of values within the training output waveform).
[0117] For each training triplet within a batch, the system generates a modified training output waveform from the noisy output and the corresponding training output waveform (710). The system can combine the noisy output and the corresponding training output waveform to generate the modified training output waveform. For example, the system can generate the modified training output waveform as
Number
Number
[0118] For each training triple in the batch, the system uses the noise estimation neural network according to the current values of the noise estimation neural network parameters to generate a training noise output (712) by processing a model input that includes (1) the modified training output waveform, (2) the training text sequence, and (3) iterative-specific data. The noise estimation neural network can process the model input to generate a training noise output as described in the process of FIG. 2. For example, the iterative-specific criterion can include an aggregated noise level
Number
[0119] The system determines updates to the network parameters of the noise estimation network, the duration predictor network, and the encoder network from the gradient of the objective function for the batch of training triples (714). The system determines the gradient of the objective function with respect to the neural network parameters for each training triple in the batch and then uses any of various suitable optimization methods such as momentum-based stochastic gradient descent or ADAM to backpropagate the gradient and update the current values of the neural network parameters using the gradient (e.g., a linear combination of gradients such as the average of the gradients).
[0120] The objective function can measure the error between the noisy output and the training noise output generated by the noise estimation network for each training triplet. For example, for a specific training triplet, the objective function can include a loss term that measures the L1 distance between the noisy output and the training noise output, such as
Number
Number
Number
[0121] The system can repeatedly execute steps (702) to (714) for a plurality of batches of training triplets (e.g., triplets of multiple training text sequences, their respective target durations, and training output waveforms).
[0122] This specification uses the term "configured" in relation to components of a system and computer programs. That one or more computer systems are configured to perform a particular operation or action means that the system has installed on it software, firmware, hardware, or a combination thereof that during operation causes the system to perform the operation or action. That one or more computer programs are configured to perform a particular operation or action means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operation or action.
[0123] Embodiments and functional operations of the subject matter described in this specification can be implemented in digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware, or one or more combinations of them, including the structures disclosed in this specification and their structural equivalents. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory recording medium for execution by, or to control the operation of, a data processing apparatus. The computer recording medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a receiver device suitable for execution by a data processing apparatus.
[0124] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of devices, apparatuses, and machines for processing data, including, by way of example, one programmable processor, one computer, or multiple processors or computers. The apparatus can be a dedicated logic circuit, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), or can further include such a dedicated logic circuit. Optionally, in addition to the hardware, the apparatus can include code for creating an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0125] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language, including compiler languages or interpreter languages, or declarative languages or procedural languages, and can be arranged in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use within a computing environment. The program may or may not correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data, such as one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple organized files, such as files that hold one or more modules, subprograms, or portions of code. A computer program can be arranged to execute on one computer or placed in one location, or distributed across multiple locations and executed on multiple computers interconnected by a data communication network.
[0126] As used herein, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine is implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers are dedicated to a particular engine, and in other cases, multiple engines can be installed and executed on the same single computer or multiple computers.
[0127] The processes and logical flows described in this specification can be executed by one or more programmable computers executing one or more computer programs to perform operations on input data and generate output. Also, the processes and logical flows can be executed by dedicated logic circuitry, such as an FPGA or ASIC, or by a combination of dedicated logic circuitry and one or more programmed computers.
[0128] Computers suitable for the execution of a computer program can be based on general purpose microprocessors or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit receives instructions and data from a read only memory, or a random access memory, or both. Essential elements of a computer are a central processing unit for executing or performing instructions, and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, dedicated logic circuitry. Also, generally, a computer is operatively coupled to one or more mass storage devices, such as magnetic disks, magneto-optical disks, or optical disks, for storing data, or receiving data from, or transferring data to, or both, such mass storage devices. However, a computer need not have such devices. Further, a computer can be incorporated in another device, such as a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive.
[0129] Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, as well as CD-ROM disks and DVD-ROM disks.
[0130] To provide interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, as well as a keyboard and a pointing device, e.g., a mouse or trackball, by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user, e.g., feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input received from the user can be in any form including acoustic, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the devices used by the user, e.g., by sending a web page to a web browser of the user's device in response to a request received from the web browser. Also, the computer can interact with the user by executing a messaging application to send a text message or other form of message to a personal device, e.g., a smartphone, and receiving a response message from the user in return.
[0131] A data processing apparatus for implementing a machine learning model can also include, for example, a dedicated hardware accelerator unit for processing computationally intensive parts of machine learning training or generation, i.e., inference workloads.
[0132] The machine learning model can be implemented and deployed using a machine learning framework, for example, the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework.
[0133] Embodiments of the subject matter described herein can be implemented in a computing system that includes back-end components, such as, for example, a data server, or middleware components, such as, for example, an application server, or front-end components, such as, for example, a graphical user interface, a web browser, or an app with which a user can interact with an implementation of the subject matter described herein, or any combination of one or more such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, for example, a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0134] A computing system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship between a client and a server is created by computer programs that are executed on respective computers and are in a client-server relationship with each other. In some embodiments, a server displays data to a user who interacts with a device acting as a client, for example, and sends data, such as an HTML page, to the user device for the purpose of receiving user input from such a user. Data generated in the user device, such as the result of a user interaction, can be received at the server from the device.
[0135] This specification includes many details of specific implementations, but these should not be considered as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. The specific features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable partial combination. Further, features may be described above as acting in certain combinations and even initially claimed as such, but one or more features of the claimed combination can in some cases be deleted from the combination, and the claimed combination can be directed to a partial combination or a variation of a partial combination.
[0136] Similarly, although the operations are shown in the drawings in a particular order and recited in the claims, it should not be understood that such operations are to be performed in the particular order shown or in sequential order, or that all of the operations shown are required to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Further, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and the described program components and systems may generally be integrated together into a single software product or packaged into multiple software products.
[0137] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired result. As one example, the processes shown in the accompanying figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.
Description of the Reference Numerals
[0138] 100 Conditional Output Generation System, End-to-End Waveform System 102 Text Sequence 104 Final Waveform Output, Final Output Waveform 106 Noise Level 108 Encoder Neural Network 110 Noise Output, Latent Representation 112 Duration Predictor Network 114 Conditioning Input 116 Current Waveform Output, Current Output Waveform 118 Noise Output 120 Update Engine 200 Process 300 Noise Estimation Neural Network, Noise Estimation Network 302 Convolutional Layer 304 Convolutional Layer 306 Aggregate Noise Level 310 Output 312 Output Waveform Processing Block 314 Output Waveform Processing Block 316 Output Waveform Processing Block 318 Output Waveform Processing Block 324 FILM Module 326 FILM Module 328 FILM Module 332 Noise Generation Block 334 Noise Generation Block 336 Noise Generation Block 338 Noise Generation Block 400 Output Waveform Processing Block 402 Alternative Representation, Convolutional Layer 404 Downsampling Layer 406 Downsampling Layer 408 Activation Layer 410 Convolutional Layer 412 Activation Layer 414 Convolutional Layer 416 Activation Layer 418 Convolutional Layer 420 Residual Connection Layer 500 FiLM Module 502 Position Encoding Function 504 Convolutional Layer 506 Activation Layer 508 Residual Connection Layer 510 Convolutional Layer 512 Scale Vector 514 Convolutional Layer 516 Bias Vector 600 Noise Generation Block 602 Input, Upsampling Layer 604 Convolutional Layer 606 Activation Layer 608 Upsampling Layer 610 Convolution layer 612 Affine transformation layer 614 Activation layer 616 Convolution layer 618 Residual connection layer 620 Affine transformation layer 622 Activation layer 624 Convolution layer 626 Affine transformation layer 628 Activation layer 630 Convolution layer 632 Residual connection layer 700 Process
Claims
1. A method performed by one or more computers to non-autoregressively generate a waveform output from a text sequence, comprising: obtaining a phoneme sequence representing the text sequence, the phoneme sequence including respective phoneme tokens at each of a plurality of input time steps; processing the phoneme sequence using an encoder neural network to generate a hidden representation of the phoneme sequence; generating a conditioning input from the hidden representation, the conditioning input including respective conditioning vectors for each of a plurality of quantized time segments in the final waveform output; initializing the current waveform output by sampling each value in the current waveform output from a corresponding noise distribution; generating the final waveform output that defines an utterance of the phoneme sequence by a speaker by gradually removing noise from the initialized current output waveform by updating the current waveform output in each of a plurality of iterations until an end criterion is met, wherein each iteration corresponds to a respective noise level, and the updating comprises, in each iteration: inputting, to a noise estimation neural network configured to process a model input to generate a noise output, the model input corresponding to the iteration and including (i) the current waveform output and (ii) the conditioning input, the noise output including respective estimated values of noise corresponding to each value in the current waveform output; and updating the current waveform output using the estimated values of the noise and the noise level corresponding to the iteration, wherein updating the current waveform output using the estimated values of the noise and the noise level corresponding to the iteration comprises: multiplying the estimated values of the noise by a coefficient based on the noise level corresponding to the iteration to generate an update for the iteration; and subtracting the update from the current waveform output. A method as claimed in claim 1.
2. The method according to claim 1, wherein the noise estimation neural network and the encoder neural network are end-to-end trained with training data including a plurality of training phoneme sequences and, for each training phoneme sequence, a respective ground truth waveform.
3. The updated output waveform is obtained by subtracting the update from the current waveform output, Updating the current waveform output, The method according to claim 1, further comprising modifying the updated output waveform based on the noise level corresponding to the iteration to generate a modified updated output waveform.
4. For the current iteration, the modified first updated output waveform is the updated output waveform after the current iteration, and for each iteration before the current iteration, the updated output waveform after the current iteration is generated by adding noise to the modified first updated output waveform. The method according to claim 3.
5. The method according to any one of claims 1 to 4, wherein the model input at each iteration includes iteration-specific data that is different for each iteration.
6. The method according to claim 5, wherein the model input for each iteration includes the noise level corresponding to the iteration.
7. The method according to claim 5, wherein the model input for each iteration includes an aggregated noise level for the iteration generated from the noise levels corresponding to the iteration and any iteration after the iteration among the plurality of iterations.
8. The noise estimation neural network comprises a noise generation neural network including a plurality of noise generation neural network layers configured to process the conditioning input to map the conditioning input to the noise output, and an output waveform processing neural network including a plurality of output waveform processing neural network layers configured to process the current waveform output to generate an alternative representation of the current waveform output. The method according to any one of claims 5 to 7, wherein at least one of the noise generation neural network layers receives an input derived from (i) an output of another one of the noise generation neural network layers, (ii) an output of a corresponding output waveform processing neural network layer, and (iii) data specific to the iteration regarding the iteration.
9. The method according to claim 8, wherein the final output waveform has a higher dimension than the conditioning input, and the alternative representation has the same dimension as the conditioning input.
10. The noise estimation neural network includes a respective Feature-wise Linear Transformation (FiLM) module corresponding to each of at least one noise generation neural network layer, and the FiLM module corresponding to a given noise generation neural network layer is configured to process (i) the output of the another one of the noise generation neural network layers, (ii) the output of the corresponding output waveform processing neural network layer, and (iii) the data specific to the iteration regarding the iteration, in order to generate the input to the noise generation neural network layer. The method according to claim 8 or 9.
11. The FiLM module corresponding to the given noise generation neural network layer (ii) generates a scale vector and a bias vector from the output of the corresponding output waveform processing neural network layer and the data specific to the iteration regarding the iteration, The method according to claim 10, wherein the FiLM module corresponding to the given noise generation neural network layer is configured to generate the input to the given noise generation neural network layer by applying an affine transformation to (i) the output of the another one of the noise generation neural network layers.
12. The method according to any one of claims 8 to 11, wherein at least one of the noise generation neural network layers includes an activation function layer, and the activation function layer applies a non-linear activation function to the input to the activation function layer.
13. The method according to claim 12, wherein the another one of the noise generation neural network layers corresponding to the activation function layer is a residual connection layer or a convolutional layer.
14. The hidden representation includes respective hidden vectors for each of a plurality of hidden time steps, and the step of generating a conditioning input from the hidden representation comprises: for each hidden time step, processing the hidden representation using a duration predictor neural network to generate a predicted duration in the utterance characterized by the hidden vector at the hidden time step; generating the conditioning input by upsampling the hidden representation according to the predicted duration so as to match the time scale of the final waveform output. The method according to any one of claims 1 to 13. **Claim 15** The method according to claim 14, wherein the conditioning input includes respective conditioning vectors for each of a plurality of quantized time segments within the final waveform output. **Claim 16** The method according to claim 14 or claim 15, further comprising, for each hidden time step, generating a predicted extent of influence in the utterance characterized by the hidden vector at the hidden time step, and wherein upsampling the hidden representation includes applying Gaussian upsampling to the hidden representation using the predicted duration and the predicted extent of influence for the hidden time step. **Claim 17** A system comprising one or more computers and one or more storage devices storing instructions operable to cause the one or more computers to perform the operations of the method according to any one of claims 1 to 16 when executed by the one or more computers. **Claim 18** A computer recording medium recording a computer program operable to cause one or more computers to perform the operations of the method according to any one of claims 1 to 16 when executed by the one or more computers.
Citation Information
Patent Citations
Learning device, acoustic generation device, method, and program
JP2019168608A
Generating Audio Using Neural Networks
JP2019532349A