High-Fidelity Speech Synthesis Using Adversarial Networks
By using a combined training method of feedforward generation neural network and conditional and unconditional discriminators, the problem of high computing resources and time consumption in the prior art is solved, and the efficient generation of realistic audio signals is achieved without explicitly modeling the audio data distribution.
Patent Information
- Application Number
- CN202080068264.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-25
- Filing Date
- 2020-09-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-09-25
AI Technical Summary
When generating audio data, especially autoregressive generation neural networks, the computing resources and time consumption are too high, and it is difficult to generate high-quality audio output. At the same time, the prior art is difficult to generate realistic audio signals without clearly modeling the audio data distribution.
The feedforward generation neural network is used for training in combination with conditional and unconditional discriminators, and the network parameters are updated by combining discriminator predictions, and the expansive convolutional neural network layer is used to broaden the sensory domain, reduce the computational complexity and improve the authenticity of audio samples.
Generating high-quality audio outputs in a single forward pass is achieved, reducing computing resources and time consumption, while the generated audio signals are realistic and in line with input text, avoiding explicit modeling of audio data distribution.
Smart Images

Figure CN114503191B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to using adversarial neural networks to generate audio data. Background Art
[0002] A neural network is a machine learning model that uses one or more layers of non-linear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to the output layer. The output of each hidden layer is used as an input to one or more other layers in the network (i.e., one or more other hidden layers, the output layer, or both). Each layer of the network generates an output from the received input according to the current values of a corresponding set of parameters. Summary of the Invention
[0003] This specification describes a system implemented as a computer program on one or more computers in one or more locations that uses a generative neural network to generate output audio examples. The generative neural network has been configured through training to receive a network input that includes conditioning input characterizing input text. The generative neural network processes the conditioning input to generate audio data corresponding to the input text, such as audio data characterizing a speaker who utters the input text.
[0004] This specification also describes a training system for training the generative neural network. Generally, the training system uses a set of one or more discriminative neural networks to train the generative neural network in an adversarial manner. That is, each discriminator neural network processes the audio examples generated by the generative neural network and predicts whether the audio example is a genuine example of audio data (e.g., a recording of a human speaker) or a synthetic example of audio data, i.e., whether the audio example was generated by the generative neural network. In this specification, the discriminator neural network is also simply referred to as a "discriminator".
[0005] In some implementations, the set of discriminators includes both a conditional discriminator and an unconditional discriminator. The conditional discriminator processes both the audio example and conditioning text input to generate a prediction, while the unconditional discriminator only processes the audio example without processing the conditioning text input to generate a prediction. In some implementations, each discriminator randomly samples different parts of the audio example and processes the random samples to generate a prediction.
[0006] The training system can combine the corresponding predictions of the set of discriminators to generate a combined prediction and update the parameters of both the feed-forward generative neural network and each discriminator based on the error of the combined prediction.
[0007] The subject matter described in this specification can be implemented in a particular embodiment so as to realize one or more of the following advantages.
[0008] In some implementations described in this specification, the generative neural network can be a feed-forward generative neural network. That is, the generative neural network can process the network input to generate an output audio example in a single forward pass. The feed-forward generative neural network as described in this specification can generate output examples faster than prior art that relies on an autoregressive generative neural network. The autoregressive neural network generates output examples across multiple time steps by performing a forward pass at each time step. At a given time step, the autoregressive neural network generates a new output sample that will be included in the output audio example, and the new output sample is conditioned on the output samples that have already been generated. This process can consume a large amount of computational resources and take a large amount of time. On the other hand, the feed-forward generative neural network can generate output examples in a single forward pass while maintaining the high quality of the generated output examples. This greatly reduces the time and amount of computational resources required to generate the output audio example.
[0009] Other prior art relies on a reversible feed-forward neural network trained by using a probability density extraction autoregressive model. Training in this way allows the reversible feed-forward neural network to generate speech signals that sound realistic and correspond to the input text without having to model every possible variation that occurs in the data. The feed-forward generative neural network as described in this specification can also generate realistic audio samples that faithfully follow the input text without having to explicitly model the data distribution of the audio data, but can do so without the extraction and reversibility requirements of the reversible feed-forward neural network.
[0010] Using both a conditional discriminator and an unconditional discriminator provides various advantages to the feed-forward generative neural network as described in this specification. The conditional discriminator can analyze to what extent the generated audio corresponds to the input text characterized by the conditioning text input, thereby allowing the feed-forward generative neural network to learn to generate audio examples that follow the input text. However, as described in more detail below, the endpoints of the random samples taken by the conditional discriminator need to be aligned with the input time steps of the conditioning text input in order for the conditional discriminator to evaluate the conditioning text input (at the frequency of the input time steps of the conditioning text input) and the generated audio (at the frequency of the output time steps of the audio example). As a specific example, if each input time step corresponds to 120 output time steps, the sampling that the conditional discriminator can perform is limited to a lower frequency by a factor of 120. On the other hand, the unconditional discriminator is not limited by such a sampling frequency constraint and is thus exposed to more diverse audio samples.
[0011] Using a discriminator that only processes samples of audio data can allow the system to discriminate between lower-dimensional distributions. Assigning a specific window size to each discriminator can allow the discriminator to operate on different frequencies of the audio samples, thereby increasing the authenticity of the audio samples generated by the feed-forward generative neural network. Using a discriminator that only processes samples of audio data can also reduce the computational complexity of the discriminator, which can allow the system to train the feed-forward generative neural network faster.
[0012] Using dilated convolutional neural network layers can also broaden the receptive fields of the feed-forward generative neural network and the discriminator, thereby allowing the corresponding networks to learn dependencies at various frequencies (e.g., long-term frequencies and short-term frequencies) in the audio examples.
[0013] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the following description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a diagram of an example training system for training a generative neural network.
[0015] Figure 2 is a diagram of an example generator block.
[0016] Figure 3 is a diagram of an example discriminator neural network system.
[0017] Figure 4 is a diagram of an example unconditional discriminator block and an example conditional discriminator block.
[0018] Figure 5 is a flowchart of an example process for training a generative neural network.
[0019] Like reference numerals and designations in the various drawings indicate like elements. DETAILED DESCRIPTION
[0020] This specification describes a system for training a generative neural network to generate output audio examples using conditioning text inputs. The system can use a discriminator neural network system including one or more discriminators to train the generative neural network in an adversarial manner.
[0021] Figure 1 is a diagram of an example training system 100 for training a generative neural network 110. Training system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented.
[0022] The training system 100 includes a generative neural network 110, a discriminator neural network system 120, and a parameter update system 130. The training system 100 is configured to train the generative neural network 110 to receive a conditioned text input 102 and process the conditioned text input 102 to generate an audio output 112. In some implementations, the generative neural network 110 is a feed-forward neural network, i.e., the generative neural network 110 generates the audio output 112 in a single forward pass.
[0023] The conditioned text input 102 represents the input text, and the audio output 112 depicts speech corresponding to the input text. In some implementations, the conditioned text input 102 includes the input text itself, such as a character-level embedding or a word-level embedding of the input text. Alternatively or additionally, the conditioned text input may include linguistic features characterizing the text input. For example, the conditioned text input may include a respective vector of linguistic features for each input time step in a sequence of input time steps. As a specific example, the linguistic features for each input time step may include i) phonemes and ii) the duration of the text at that input time step. The linguistic features may also include pitch information; for example, pitch may be represented by the log base frequency logF0 of the input time step.
[0024] The generative neural network 110 may have any suitable neural network architecture. As a specific example, the generative neural network 110 may include a sequence of groups of convolutional neural network layers referred to as "generator blocks". The first generator block in the sequence of generator blocks may receive the conditioned text input (or an embedding of the conditioned text input) as input and generate a block output. Each subsequent generator block in the sequence of generator blocks may receive the block output generated by the previous generator block in the sequence of generator blocks as input and generate a subsequent block output. The generator blocks are described in more detail below with reference to Figure 2 More detailed description of the generator blocks.
[0025] In some implementations, the generative neural network 110 may also receive a noise input 104 as input. For example, the noise input 104 may be randomly sampled from a pre-determined distribution (such as a normal distribution). The noise input 104 may ensure variability in the audio output 112.
[0026] In some implementations, the generative neural network 110 may also receive an identification of a class 106 to which the audio output 112 should belong as input. The class 106 may be a member of a set of possible classes. For example, the class 106 may correspond to a specific speaker that the audio output 112 should sound like. That is, the audio output 112 may depict a specific speaker saying the input text.
[0027] The audio output 112 may include audio samples of the audio wave at each output time step in the output time step sequence. For example, for each output time step, the audio output 112 may include the amplitude value of the audio wave. In some implementations, the amplitude value may be a compressed or companded amplitude value.
[0028] Generally, the input time step sequence and the output time step sequence characterize the same time period, such as 1, 2, 5, or 10 seconds. As a specific example, if the time period is 2 seconds, the conditioning input 102 may include 400 input time steps (resulting in a frequency of 200 Hz), while the audio output 112 may include 48,000 time steps (resulting in a frequency of 24 kHz). Thus, the generative neural network 110 may generate audio samples for multiple output time steps (in this case, 120) for each single input time step.
[0029] In some implementations in which the generative neural network 110 includes one or more sequences of generator blocks, due to the frequency difference between the input time steps and the output time steps, one or more of the generator blocks in the generative neural network 110 may include one or more corresponding upsampling layers. The dimension of the layer output of each upsampling layer is greater than the dimension of the layer input of the upsampling layer. The total degree of upsampling over all the generator blocks in the generative neural network 110 may be proportional to the ratio of the frequency of the output time steps to the frequency of the input time steps.
[0030] After the generative neural network 110 generates the audio output 112, the training system 100 may provide the audio output 112 to the discriminator neural network system 120. The training system 100 may train the discriminator neural network system 120 to process the audio samples and generate a prediction 122 as to whether the audio samples are real (i.e., audio samples captured in the real world) or synthetic (i.e., audio samples that have been generated by the generative neural network 110).
[0031] The discriminator neural network system 120 may have any suitable neural network architecture. As a specific example, the discriminator neural network system 120 may include one or more discriminators, each of which processes the audio output 112 and predicts whether the audio output 112 is real or synthetic. Each discriminator may include a sequence of groups of convolutional neural network layers, referred to as "discriminator blocks". Example discriminator blocks are described below with reference to Figure 3 Describe example discriminator blocks.
[0032] In some implementations, one or more discriminators of the discriminator neural network system 120 include one or more conditional discriminators and one or more unconditional discriminators. The conditional discriminator receives as inputs i) the audio output 112 generated by the generative neural network 110 and ii) the conditioning text input 102 used by the generative neural network 110 to generate the audio output 112. The unconditional discriminator receives the output audio 112 generated by the generative neural network 110 as an input, but does not receive the conditioning text input 102 as an input. Thus, the conditional discriminator can measure, in addition to the general authenticity of the audio output 112, to what extent the audio output 112 corresponds to the input text characterized by the conditioning text input 102, while the unconditional discriminator only measures the general authenticity of the audio output 112. The following describes an example discriminator neural network system in more detail with reference to Figure 4 Describe the example discriminator neural network system in more detail.
[0033] The discriminator neural network system 120 can combine the respective predictions of one or more discriminators to generate a prediction 122.
[0034] The parameter update system 130 can obtain the prediction 122 generated by the discriminator neural network system 120 and determine a parameter update 132 based on the error in the prediction 122. The training system can apply the parameter update 132 to the parameters of the generative neural network 110 and the discriminator neural network system 120. That is, the training system 100 can train the generative neural network 110 and the discriminators in the discriminator neural network system 120 simultaneously.
[0035] Generally, the parameter update system 130 determines a parameter update 132 for the parameters of the generative neural network 110 in order to increase the error in the prediction 122. For example, if the discriminator neural network system 120 correctly predicts that the audio output 112 is synthetic, the parameter update system 130 generates a parameter update 132 for the parameters of the generative neural network 110 in order to improve the authenticity of the audio output 112 such that the discriminator neural network system 120 may incorrectly predict the next audio output 112 as real.
[0036] Conversely, the parameter update system 130 determines a parameter update 132 for the parameters of the discriminator neural network system 120 in order to reduce the error in the prediction 122. For example, if the discriminator neural network system 120 incorrectly predicts that the audio output 112 is real, the parameter update system 130 generates a parameter update 132 for the parameters of the discriminator neural network system 120 in order to improve the prediction 122 of the discriminator neural network system 120.
[0037] In some implementations where the discriminator neural network system 120 includes multiple different discriminators, the parameter update system 130 uses the predictions 122 output by the discriminator neural network system 120 to determine parameter updates 132 for each discriminator. That is, the parameter update system 130 uses the same combined predictions 122 to determine parameter updates 132 for each particular discriminator regardless of the corresponding predictions generated by the particular discriminator, where the combined predictions 122 are generated by combining the corresponding predictions of the multiple discriminators.
[0038] In some other implementations where the discriminator neural network system 120 includes multiple different discriminators, the parameter update system 130 uses the corresponding predictions generated by each discriminator to determine parameter updates 132 for that discriminator. That is, the parameter update system 130 generates parameter updates 132 for a particular discriminator to improve the corresponding predictions generated by that particular discriminator (which will indirectly improve the combined predictions 122 output by the discriminator neural network system 120 since the combined predictions 122 are generated based on the corresponding predictions generated by the multiple discriminators).
[0039] During training, the training system 100 can also provide real audio samples 108 to the discriminator neural network system 120. Each discriminator in the discriminator neural network system 120 can process the real audio samples 108 to predict whether the real audio samples 108 are real examples of audio data or synthetic examples of audio data. Again, the discriminator neural network system 120 can combine the corresponding predictions of each discriminator to generate a second prediction 122. The parameter update system 130 can then determine second parameter updates 132 for the parameters of the discriminator neural network system 120 based on the error in the second prediction 122. Generally, the training system 100 does not use the second prediction corresponding to the real audio samples 108 to update the parameters of the generative neural network 110.
[0040] In some implementations where the generative neural network 110 has been trained to generate synthetic audio outputs 112 belonging to a category 106, the discriminator neural network system 120 does not receive the identity of the category 106 to which the received (real or synthetic) audio examples belong as an input. However, the real audio samples 108 received by the discriminator neural network system 120 can include audio samples belonging to each category 106 in the category set.
[0041] As a specific example, the parameter update system 130 can use the Wasserstein loss function to determine the parameter updates 132, which is:
[0042] D(x) - D(G(z)),
[0043] Where D(x) is the likelihood assigned by the discriminator neural network system 120 that the real audio sample 108 is real, G(z) is the synthetic audio output 112 generated by the generator neural network 110, and D(G(z)) is the likelihood assigned by the discriminator neural network system 120 that the synthetic audio output 112 is real. The purpose of the generator neural network 110 is to minimize the Wasserstein loss by maximizing D(G(z)), that is, by making the discriminator neural network system 120 predict that the synthetic audio output 112 is real. The purpose of the discriminator neural network system 120 is to maximize the Wasserstein loss, that is, to correctly predict real audio examples and synthetic audio examples.
[0044] As another specific example, the parameter update system 130 can use the following loss function:
[0045] log(D(x)) + log(1 - D(G(z)))
[0046] Where again, the goal of the generator neural network 110 is to minimize the loss, while the goal of the discriminator neural network system 120 is to maximize the loss.
[0047] The training system 100 can backpropagate the loss through both the generator neural network 110 and the discriminator neural network system 120, thereby training both networks simultaneously.
[0048] Figure 2 is a diagram of the example generator block 200. The generator block 200 can be a component of a generator neural network (such as Figure 1 the generator neural network 110 depicted in), which is configured to process the conditioning text input and generate an audio output. The generator neural network can include a sequence of generator blocks. In some implementations, each generator block in the sequence of generator blocks has the same architecture. In some other implementations, one or more of the generator blocks have an architecture different from the other generator blocks in the sequence of generator blocks.
[0049] The generator block is configured to receive i) a block input 202 and ii) a noise input 204 as inputs, and generate a block output 206. In implementations where the generator neural network also receives an identification of the class to which the audio output should belong as an input, one or more of the generator blocks of the generator neural network can also receive the identification of the class as an input; for simplicity, this is omitted from Figure 2 here.
[0050] In some implementations, the first generator block in the sequence of generator blocks receives the conditioning text input as the block input 202. In some other implementations, the first generator block in the sequence of generator blocks receives an embedding of the conditioning text input generated by one or more initial neural network layers of the generator neural network. Each subsequent generator block in the sequence of generator blocks receives the block output generated by the previous generator block in the sequence of generator blocks as the block input 202.
[0051] Layer Time dimension Frequency Number of channels Linguistic features 400 200 Hz 567 Input convolutional layer 400 200 Hz 768 G-block 400 200 Hz 768 G-block 400 200 Hz 768 G-block, upsampling x2 800 400 Hz 384 G-block, upsampling x2 1600 800 Hz 384 G-block, upsampling x2 3200 1600 Hz 384 G-block, upsampling x3 9600 4800 Hz 192 G-block, upsampling x5 48000 24 kHz 96 Output convolutional layer 48000 24 kHz 1
[0052] Table 1: Example Generator Neural Network Architecture
[0053] Table 1 describes an example architecture of the generator neural network. The examples shown in Table 1 are for illustrative purposes only, and many different configurations of the generator neural network are possible. The input to the generator neural network is a vector of linguistic features for each of the 400 input time steps corresponding to two seconds of audio (at a frequency of 200 Hz). Each vector of linguistic features corresponding to a respective input time step includes 567 channels.
[0054] The generator neural network includes an input convolutional neural network layer that generates an embedding of the linguistic features to provide to the first generator block in the sequence of generator blocks. The input convolutional neural network layer increases the number of channels per time step from 567 to 768. In some implementations, the generator neural network has multiple input neural network layers.
[0055] The generator neural network includes seven generator blocks ("G blocks"), although generally the generator neural network can include any number of generator blocks. The first two generator blocks do not include any upsampling layers, so the time dimension of the block output is the same as the time dimension of the block input. The next three generator blocks each upsample their respective block inputs by 2x, such that the block output of the fifth generator block has a time dimension of 3200 (at a frequency of 1600 Hz). The sixth generator block upsamples its block input by 3x, thereby generating a block output with a time dimension of 9600 (at a frequency of 4800 Hz). The last generator block upsamples its block input by 5x, thereby generating a block output with a time dimension of 48000 (at a frequency of 24 kHz). The sequence of generator blocks also reduces the number of channels per time step from 768 to 96.
[0056] The generator neural network includes an output convolutional neural network layer that processes the block output of the last generator block in the sequence of generator blocks and generates an audio output of the generator neural network. The audio output includes 48,000 output time steps, and each output time step has a single channel representing the amplitude of the audio wave at the output time step. In some implementations, the output convolutional neural network layer includes an activation function, e.g., a hyperbolic tangent activation function. In some implementations, the generator neural network has multiple output neural network layers.
[0057] Return reference Figure 2 , the generator block 200 uses a first stack 212 of neural network layers to process the block input 202. The first stack 212 of neural network layers includes a batch normalization layer 212a, an activation layer 212b, an upsampling layer 212c, and a convolutional neural network layer 212d.
[0058] In some implementations, the batch normalization layer 212a is a conditional batch normalization layer that is conditioned on a noise input 204. For example, the conditional batch normalization layer 212a can be conditioned on a linear embedding of the noise input 204 generated by a linear layer 214 of the generator block 200. In some implementations, the linear layer 214 combines (e.g., concatenates) i) a linear embedding of the noise input 204 and ii) an identity of the class to which the output audio is to belong (e.g., a particular speaker that the output audio is to sound like) to generate a combined representation and provides the combined representation to the conditional batch normalization layer 212a. For example, the class identity can be encoded as a one-hot vector, i.e., a vector whose elements have value 0 except for a single element corresponding to a particular class that has value 1. The conditional batch normalization layer 212a can then be conditioned on the combined representation.
[0059] In some implementations, the activation layer 212b can be a ReLU activation layer.
[0060] For a generator block that upsamples the block input 202, e.g., the last five generator blocks listed in Table 1, the upsampling layer 212c upsamples the output of the activation layer by a corresponding upsampling factor p. That is, the upsampling layer 212c generates a layer output that is p times the layer input in the time dimension. For example, the upsampling layer 212c can generate a layer output that includes p consecutive copies of the layer input for each element of the layer input. As another example, the upsampling layer 212c can linearly interpolate between each pair of consecutive elements in the layer input to generate (p - 1) corresponding additional elements in the layer output. As another example, the upsampling layer 212c can perform high-order interpolation on the elements of the layer input to generate the layer output. Although Figure 2The generator block 200 depicted in [reference] has a single upsampling layer 212c, but generally the generator block can have any number of upsampling layers throughout the generator block.
[0061] The convolutional neural network layer 212d processes a layer input that includes M channels for each element and generates a layer output that includes N channels for each output. M corresponds to the number of channels in the block input 202, while N corresponds to the number of channels in the block output 206. In some cases, M = N. Although Figure 2 the example depicted in [reference] shows that the first convolutional neural network layer 212d of the generator block 200 processes a layer input with M channels to generate a layer output with N channels while each subsequent convolutional neural network layer retains the number of channels, generally any one or more convolutional layers of the generator block can change the number of channels to produce a block output with N channels.
[0062] The generator block 200 includes a second stack 216 of neural network layers that processes the output of the first stack 212 of neural network layers. The second stack 216 of neural network layers includes a batch normalization layer 216a, an activation layer 216b, and a convolutional neural network layer 216c.
[0063] As described above, the batch normalization layer 216a can be a conditional batch normalization layer that is conditioned on the linear embedding of the noise input 204 generated by the linear layer 218. Generally, each of the linear layers 214, 218, 234, and 238 of the generator block 200 has different parameters, and thus generates different linear embeddings of the noise input 204.
[0064] The generator block 200 includes a first skip connection 226 that combines the input of the first stack 212 of neural network layers with the output of the second stack 216 of neural network layers. For example, the first skip connection 226 can add or concatenate the input of the first stack 212 of neural network layers with the output of the second stack 216 of neural network layers.
[0065] In the case where the first stack 212 or the second stack 216 of neural network layers includes an upsampling layer (e.g., the upsampling layer 212c), the generator block can include another upsampling layer 222 before the first skip connection 226 such that the two inputs to the first skip connection 226 have the same dimensions.
[0066] In the case where the first stack 212 or the second stack 216 of neural network layers includes a convolutional neural network layer (e.g., the convolutional neural network layer 212d) that changes the number of channels of the layer input, the generator block can include another convolutional neural network layer 224 before the first skip connection 226 such that the two inputs to the first skip connection 226 have the same number of channels.
[0067] The generator block 200 includes a third stack 232 of neural network layers that processes the output of the first skip connection 226. The third stack 232 of neural network layers includes a batch normalization layer 232a, an activation layer 232b, and a convolutional neural network layer 232c. As described above, the batch normalization layer 232a can be a conditional batch normalization layer that is conditioned on the linear embedding of the noise input 204 generated by the linear layer 234.
[0068] The generator block 200 includes a fourth stack 236 of neural network layers that processes the output of the third stack 232 of neural network layers. The fourth stack 236 of neural network layers includes a batch normalization layer 236a, an activation layer 236b, and a convolutional neural network layer 236c. As described above, the batch normalization layer 236a can be a conditional batch normalization layer that is conditioned on the linear embedding of the noise input 204 generated by the linear layer 238.
[0069] The generator block 200 includes a second skip connection 242 that combines, e.g., by addition or concatenation, the input of the third stack 232 of neural network layers and the output of the fourth stack of neural network layers. The block output 206 of the generator module 200 is the output of the second skip connection 242.
[0070] In some implementations, one or more of the convolutional neural network layers of the generator block 200 are dilated convolutional neural network layers. A dilated convolutional neural network layer is a convolutional layer in which a filter is applied over a region larger than the length of the filter by skipping input values with a certain stride defined by the dilation value of the dilated convolution. In some implementations, the generator block 200 can include multiple dilated convolutional neural network layers with increasing dilation. For example, for each dilated convolutional neural network layer, the dilation value can be doubled from an initial dilation and then returned to the initial dilation in the next generator block. In Figure 2 the example depicted, the first convolutional neural network layer 212d has a dilation value of 1, the second convolutional neural network layer 216c has a dilation value of 2, the third convolutional neural network layer 232c has a dilation value of 4, and the fourth convolutional neural network layer 236c has a dilation value of 8.
[0071] Figure 3 is a diagram of an example discriminator neural network system 300. The discriminator neural network system 300 can be a training system configured to train a generative neural network (e.g., Figure 1Components of the training system 100 depicted. The discriminator neural network system 300 has been trained to receive audio examples 302 and generate a prediction 306 as to whether the audio example is a real audio example or a synthetic audio example generated by a generator neural network. The audio example can include corresponding amplitude values at each output time step in an output time step sequence (referred to as an "output" time step because the audio example may already be the output of a generative neural network). The audio example corresponds to conditioning text input 304, which includes corresponding vectors of linguistic features at each input time step in an input time step sequence. Note that even if the audio example 302 is a real audio example, the conditioning text input 304 still corresponds to the audio example 302.
[0072] The discriminator neural network system 300 can include one or more unconditional discriminators, one or more conditional discriminators, or both. In Figure 3 the example depicted, the discriminator neural network system 300 includes five unconditional discriminators (including unconditional discriminator 320) and five conditional discriminators (including conditional discriminator 340).
[0073] In some implementations, instead of processing the entire audio example 302, each discriminator in the discriminator neural network system 300 processes a different subset of the audio example. For example, each discriminator can randomly sample a subset of the audio example 302. That is, each discriminator processes only the amplitudes of a subsequence of consecutive output time steps of the audio sample 302, where the subsequence of consecutive output time steps is randomly sampled from the entire sequence of output time steps. In some implementations, the size of the random sample (i.e., the number of output time steps sampled) is the same for each discriminator. In some other implementations, the size of the random sample is different for each discriminator and is referred to as the "window size" of the discriminator.
[0074] In some implementations, one or more discriminators in the discriminator neural network system 300 can have the same network architecture and the same parameter values. For example, the training system can update the one or more discriminators in the same way during the training of the discriminator neural network system.
[0075] In some implementations, different discriminators have different window sizes. For example, for each window size in a set of multiple window sizes, the system can include one conditional discriminator with that window size and one unconditional discriminator with that window size. As a specific example, the discriminator neural network 300 includes one conditional discriminator and one unconditional discriminator corresponding to each window size in a set of five window sizes (240, 480, 960, 1920, and 3600 output time steps).
[0076] Each conditional discriminator also obtains a sample of the conditioning text input 304 corresponding to a random sample of the audio sample 302 obtained by the conditional discriminator. Because the samples of the audio example 302 and the samples of the conditioning text input 304 must be aligned, the conditional discriminator can be constrained to sample a subsequence of the audio output 302 starting at the same point as the point at which the input time step of the conditioning text input 304 begins. That is, because the number of input time steps in the conditioning text input 304 is less than the number of output time steps in the audio example 302, each conditional discriminator can be constrained to sample the audio example 302 at points aligned with the input time steps of the conditioning text input 304. Because the unconditional discriminator does not process the conditioning text input 304, the unconditional discriminator does not have this constraint.
[0077] In some implementations, before processing the random samples of the audio example 302, each discriminator first uses a "shaping" layer to downsample the random samples by a factor proportional to the window size of the discriminator. Downsampling effectively allows discriminators with different window sizes to process the audio example 302 at different frequencies, where the frequency at which a particular discriminator operates is proportional to the window size of the particular discriminator. Downsampling by a factor proportional to the window size also allows each downsampled representation to have a common dimension, which allows each discriminator to have a similar architecture and similar computational complexity, despite having different window sizes. In Figure 3 the example depicted, the common dimension is 240 time steps; each discriminator is labeled with a factor k, and the discriminator downsamples its random samples by factor k (1, 2, 4, 8, and 15 respectively).
[0078] In some implementations, the discriminator uses a strided convolutional neural network layer to downsample the corresponding random samples of the audio example 302.
[0079] The unconditional discriminator 320 randomly samples a sample 312 from the audio example 302 that includes 480 output time steps. The unconditional discriminator 320 includes a shaping layer 322 that downsamples the random sample 312 by 2x to generate a network input that includes 240 time steps.
[0080] The unconditional discriminator 320 includes a sequence 324 of discriminator blocks ("DBlock"). Although the unconditional discriminator 320 is depicted as including five discriminator blocks, generally the unconditional discriminator can have any number of discriminator blocks. Each discriminator block in the unconditional discriminator 320 is an unconditional discriminator block. The example unconditional discriminator block is described below with reference to Figure 4 is described.
[0081] The first discriminator block in discriminator block sequence 324 is configured to receive the network input from the shaping layer 322 and process the network input to generate a block output. Each subsequent discriminator block in discriminator block sequence 324 is configured to receive the block output generated by the previous discriminator block in sequence 324 as input and generate a subsequent block output.
[0082] In some implementations, one or more discriminator blocks in the unconditional discriminator 320 perform further downsampling. For example, as Figure 3 shown, the second discriminator block in discriminator block sequence 324 performs 5x downsampling, while the third discriminator block in sequence 324 performs 3x downsampling. Generally, downsampling helps the unconditional discriminator 320 limit the dimension of the internal representation, thereby improving efficiency and allowing the unconditional discriminator 320 to learn the relationships between more distant elements of the samples 312 of the audio example 302. In some implementations, the number of downsampling layers in the corresponding unconditional discriminator and the degree of downsampling performed by the corresponding unconditional discriminator depend on the factor k of the corresponding unconditional discriminator.
[0083] The unconditional discriminator 320 includes an output layer 326 that receives the block output of the last discriminator block in discriminator block sequence 324 as input and generates a prediction 328. The prediction 328 predicts whether the audio example is a real audio example or a synthetic audio example. For example, the output layer 326 can be an average pooling layer that generates a scalar representing the likelihood that the audio example 302 is a real audio example, where a larger scalar value indicates a higher confidence that the audio example 302 is real.
[0084] The conditional discriminator 340 randomly samples samples 332 from the audio example 302 that include 3600 output time steps. The conditional discriminator 340 includes a shaping layer 342 that performs 15x downsampling on the random samples 332 to generate a network input that includes 240 time steps.
[0085] The conditional discriminator 340 includes a discriminator block sequence 344. Although the conditional discriminator 340 is depicted as including five unconditional discriminator blocks and one conditional discriminator block, generally the conditional discriminator can have any number of conditional discriminator blocks and unconditional discriminator blocks.
[0086] The first discriminator block in discriminator block sequence 344 is configured to receive the network input from the shaping layer 342 and process the network input to generate a block output. Each subsequent discriminator block in discriminator block sequence 344 is configured to receive the block output generated by the previous discriminator block in sequence 344 as input and generate a subsequent block output.
[0087] The conditional discriminator block is configured to receive as inputs i) the block output of the previous discriminator block in the sequence 344 and ii) a sample of the conditioning text input 304 corresponding to the sample 332 of the audio example. The following refers to Figure 4 Describes an example conditional discriminator block.
[0088] Due to the frequency difference between the output time steps and the input time steps, one or more unconditional discriminator blocks and / or the conditional discriminator block itself in the discriminator block sequence 344 may perform downsampling. For example, the total degree of downsampling across all discriminator blocks in the conditional discriminator 340 may be proportional to the ratio of the frequency of the input time steps to the frequency of the output time steps. In Figure 3 the example depicted, the conditional discriminator 340 performs 8x downsampling on the network input to reach a frequency of 200 Hz, which is the same frequency as the conditioning text input.
[0089] The conditional discriminator 340 includes an output layer 346 that receives the block output of the last discriminator block in the discriminator block sequence 344 as an input and generates a prediction 348 as to whether the audio example is a real audio example or a synthetic audio example. For example, the output layer 346 may be an average pooling layer that generates a scalar representing the likelihood that the audio example 302 is a real audio example.
[0090] After each discriminator in the discriminator neural network system 300 generates a prediction, the discriminator neural network system 300 may combine the corresponding predictions to generate a final prediction 306 as to whether the audio example is a real audio example or a synthetic audio example. For example, the discriminator neural network system 300 may determine the sum or average of the corresponding predictions of the discriminators. As another example, the discriminator neural network system 300 may generate the final prediction 306 according to a voting algorithm, e.g., predicting that the audio example 302 is a real audio example if and only if a majority of the discriminators predict the audio example 302 as a real audio example.
[0091] Note that the discriminator neural network system 300 does not receive the class associated with the audio input as an input. However, in some implementations, the discriminator neural network system 300 may do so, e.g., as a further conditioning input to the conditioning discriminator block of each conditional discriminator.
[0092] Figure 4 Is a diagram of an example unconditional discriminator block 400 and an example conditional discriminator block 450. The discriminator blocks 400 and 450 may be discriminator neural network systems (e.g., Figure 1Components of the discriminator neural network system 120 depicted. The discriminator neural network system can include one or more discriminators, and each discriminator can include one or more sequences of discriminator blocks. In some implementations, each discriminator block in the sequence of discriminator blocks in a discriminator has the same architecture. In some other implementations, one or more discriminator blocks have an architecture different from other discriminator blocks in the sequence of discriminator blocks in the discriminator.
[0093] The unconditional discriminator block 400 can be a component of one or more unconditional discriminators, one or more conditional discriminators, or both. The conditional discriminator block 450 can be a component of only one or more conditional discriminators. That is, a conditional discriminator can include an unconditional discriminator block, but an unconditional discriminator cannot include a conditional discriminator block.
[0094] The unconditional discriminator block 400 is configured to receive a block input 402 and generate a block output 404. In some implementations, if the unconditional discriminator block 400 is the first discriminator block in the sequence of discriminator blocks of a discriminator, the block input 402 is an audio example. In some other implementations, if the unconditional discriminator block 400 is the first discriminator block in the sequence of discriminator blocks of a discriminator, the block input 402 is an embedding of the audio example, e.g., an embedding generated by the input convolutional neural network layer of the discriminator. If the unconditional discriminator block 400 is not the first discriminator block in the sequence, the block input 402 can be the block output of the previous discriminator block in the sequence.
[0095] The unconditional discriminator block 400 includes a first stack 412 of neural network layers, and the first stack includes a downsampling layer 412a, an activation layer 412b, and a convolutional neural network layer 412c.
[0096] For an unconditional discriminator block that downsamples the block input 402 (e.g., Figure 3 the second and third discriminator blocks in the discriminator block sequence 324 depicted), the downsampling layer 412a downsamples the block input 402 by a corresponding downsampling factor. Although Figure 4 the unconditional discriminator block 400 depicted has a single downsampling layer 412a, generally an unconditional discriminator block can have any number of downsampling layers throughout the unconditional discriminator block.
[0097] In some implementations, the activation layer 412b can be a ReLU activation layer.
[0098] The convolutional neural network layer 412c processes a layer input of N channels for each element and generates a layer output of m·N channels for each element. N corresponds to the number of channels in the block input 202, and m corresponds to a multiplier. In some cases, m = 1. Generally, any one or more convolutional layers of the unconditional discriminator block can change the number of channels of the layer input.
[0099] The unconditional discriminator block 400 includes a second stack 414 of neural network layers that processes the output of the first stack 412 of neural network layers. The second stack 414 of neural network layers includes an activation layer 414a and a convolutional neural network layer 414b.
[0100] The unconditional discriminator block 400 includes a skip connection 426 that combines the input of the first stack 412 of neural network layers with the output of the second stack 414 of neural network layers. For example, the skip connection 426 can add or concatenate the input of the first stack 412 of neural network layers with the output of the second stack 414 of neural network layers.
[0101] In the case where the first stack 412 or the second stack 414 of neural network layers includes a convolutional neural network layer (e.g., convolutional neural network layer 412c) that changes the number of channels of the layer input, the unconditional discriminator block can include another convolutional neural network layer 422 before the skip connection 426 so that the two inputs to the skip connection 426 have the same number of channels.
[0102] In the case where the first stack 412 or the second stack 414 of neural network layers includes a downsampling layer (e.g., downsampling layer 412a), the unconditional discriminator block can include another downsampling layer 424 before the skip connection 426 so that the two inputs to the skip connection 426 have the same dimensions.
[0103] In some implementations, one or more convolutional neural network layers in the unconditional discriminator block 400 are dilated convolutional layers.
[0104] The conditional discriminator block 450 is configured to receive a block input 452 corresponding to an audio example and a conditioning text input 454 and generate a block output 456.
[0105] The conditional discriminator block 450 includes a first stack 462 of neural network layers that includes a downsampling layer 462a, an activation layer 462b, and a convolutional neural network layer 462c.
[0106] For the conditional discriminator block that downsamples the block input 452, the downsampling layer 462a downsamples the block input 452 by a corresponding downsampling factor. Although Figure 4The conditional discriminator block 450 shown in [Figure 0] has a single downsampling layer 462a, but generally a conditional discriminator block can have any number of downsampling layers throughout the conditional discriminator block. In particular, the conditional discriminator block 450 downsamples the layer input 452 such that it has the same dimension as the conditioning text input 454 in the temporal dimension.
[0107] In some implementations, the activation layer 462b can be a ReLU activation layer.
[0108] The convolutional neural network layer 462c processes a layer input that includes N channels for each element and generates a layer output that includes m·N channels for each element; in this example, m = 2. Generally, any one or more convolutional layers of the conditional discriminator block can change the number of channels of the layer input.
[0109] The conditional discriminator block 450 includes a convolutional neural network layer 472 that processes the conditioning text input 454 and generates a layer output that includes m·N channels for each element. That is, the convolutional neural network layer 472 changes the number of channels of the conditioning text input (in this case, 567) to match the output of the convolutional neural network layer 462c in the first stack 462 of neural network layers.
[0110] The conditional discriminator block 450 includes a combination layer 474 that combines the outputs of the convolutional neural network layers 462c and 472, for example, by addition or by concatenation.
[0111] The conditional discriminator block 450 includes a second stack 482 of neural network layers that processes the output of the combination layer 474. The second stack 482 of neural network layers includes an activation layer 482a and a convolutional neural network layer 482b.
[0112] The conditional discriminator block 450 includes a skip connection 496 that combines the input of the first stack 462 of neural network layers with the output of the second stack 482 of neural network layers. For example, the skip connection 496 can add or concatenate the input of the first stack 462 of neural network layers with the output of the second stack 482 of neural network layers.
[0113] In the case where the first stack 462 or the second stack 482 of neural network layers includes a convolutional neural network layer (e.g., the convolutional neural network layer 462c) that changes the number of channels of the layer input, the conditional discriminator block 450 can include another convolutional neural network layer 492 before the skip connection 496 such that the two inputs to the skip connection 496 have the same number of channels.
[0114] In the case where the first stack 462 or the second stack 482 of neural network layers includes a downsampling layer (e.g., downsampling layer 462a), the conditional discriminator block 450 may include another downsampling layer 494 before the skip connection 496, such that the two inputs to the skip connection 496 have the same dimensions.
[0115] In some implementations, one or more convolutional neural network layers in the conditional discriminator block 450 are dilated convolutional layers.
[0116] Figure 5 is a flowchart of an example process 500 for training a generative neural network. For convenience, process 500 will be described as being executed by a system of one or more computers located in one or more locations. For example, a training system (e.g., Figure 1 the training system 100 depicted) appropriately programmed according to this specification can execute process 500.
[0117] The generative neural network has a plurality of generative parameters. The generative neural network may include a sequence of groups of convolutional neural network layers, each group including one or more dilated convolutional neural network layers. One or more groups or blocks of the generative neural network may also include one or more upsampling layers to account for the ratio between the input time steps of the conditioning text input and the output time steps of the audio output.
[0118] The system obtains a training conditioning text input (step 502). The training conditioning text input may include corresponding linguistic feature representations at each of a plurality of input time steps. For example, the linguistic feature representation at each input time step may include phonemes, durations, and log fundamental frequencies at that time step.
[0119] The system uses the generative neural network to process a training generative input that includes the training conditioning text input to generate a training audio output (step 504). The system generates the training audio output according to the current values of the generative parameters of the generative neural network. The training audio output may include corresponding audio samples at each of a plurality of output time steps.
[0120] The training generative input may also include a noise input. The generative neural network may include one or more conditional batch normalization neural network layers conditioned on a linear embedding of the noise input. In some implementations, each conditional batch normalization neural network layer is conditioned on a different linear embedding of the noise input.
[0121] In some implementations, the generative input includes an identification of the class to which the output wave should belong. In some such implementations, the conditional batch normalization neural network layers are also conditioned on the identification of the class.
[0122] In some implementations, the system zero-pads the training conditioning text input (i.e., adds one or more zeros to the end of the training conditioning text input) before providing it to the generative neural network, such that each training conditioning text input has the same dimension. Instead of processing training examples of different sizes sequentially, it can be more efficient for a deep learning model to process batches of training examples of a fixed size.
[0123] However, because the generative neural network has multiple convolutional neural network layers, convolutions may propagate non-zero values into the zero-padded elements, which results in interference with non-zero-padded elements in later convolutions. To avoid this, the system can apply a convolutional mask to the input of each convolutional neural network layer. The convolutional mask can be a zero-one mask, where 0 is at the elements corresponding to the zero-padded elements of the input and 1 is at the elements corresponding to the non-zero-padded elements of the input. Applying the convolutional mask ensures that the zero-padded elements of the input do not interfere with the non-zero-padded elements of the input. The zero-one mask can be upsampled at the same rate as the conditional input, such that an appropriate number of zero-padded elements are processed before each corresponding convolutional neural network layer. Before outputting the final audio example, the system can remove the zero-padded elements. By zero-padding and then applying the convolutional mask, the system is able to generate audio examples of arbitrary length.
[0124] The system uses each of the multiple discriminators to process the training audio output to generate corresponding predictions as to whether the training audio output is real or synthetic (step 506).
[0125] In some implementations, the multiple discriminators include one or more conditional discriminators and one or more unconditional discriminators.
[0126] In some implementations, one or more of the multiple discriminators process different true subsets of the training audio output. In some such implementations, the size of the true subset for a particular discriminator is predetermined. In some such implementations, each discriminator downsamples the true subset by a predetermined downsampling factor corresponding to the size of the true subset. In some such implementations, each discriminator uses strided convolutions to downsample the true subset.
[0127] The system determines a combined prediction by combining the corresponding predictions of the multiple discriminators (step 508). For example, the system can determine the average of the predictions, or use a voting algorithm to process the predictions.
[0128] The system determines an update to the current value of the generation parameters to increase the error in the combined prediction (step 510). In some implementations, the system can also determine an update to the current value of the discriminator parameters of the multiple discriminators to reduce the error in the combined prediction.
[0129] This specification uses the term "configured" in connection with systems and computer program components. A system of one or more computers being configured to perform particular operations or actions means that the system has software, firmware, hardware, or a combination of them installed on it, which, when operating, cause the system to perform the operations or actions. One or more computer programs being configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operations or actions.
[0130] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly implemented computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory, or a combination of one or more of them. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiver device for execution by a data processing apparatus).
[0131] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of devices, equipment, and machines for processing data, including, by way of example, a programmable processor, a computer, or a multi-processor or computer. The apparatus can also be, or further include, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to hardware, the apparatus can optionally include code that creates an execution environment for the computer program, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0132] A computer program (which may also be referred to as or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or module, component, subroutine, or other unit suitable for use in a computing environment. A program can, but need not, correspond to a file in a file system. A program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts in a markup language document), in a single file dedicated to the program, or in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or multiple computers, which are located at one site or distributed among multiple sites and interconnected by a data communication network.
[0133] In this specification, the term "database" is used broadly to refer to any collection of data: the data need not be structured in any particular way, or structured at all, and can be stored on storage devices in one or more locations. Thus, for example, an indexed database can include multiple collections of data, each of which can be organized and accessed differently.
[0134] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and run on the same one or more computers.
[0135] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs, thereby implementing functions by operating on input data and generating output. The processes and logical flows can also be performed by dedicated logic circuitry (such as an FPGA or ASIC), or by a combination of dedicated logic circuitry and one or more programmed computers.
[0136] A computer suitable for executing a computer program can be based on a general or special purpose microprocessor or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a central processing unit for executing or running instructions and one or more memory devices for storing the instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will include one or more mass storage devices (such as, by way of example only, magnetic disks, magneto-optical disks or optical disks) for storing data or can be operatively coupled to one or more mass storage devices to receive data therefrom or to transfer data thereto or both. However, a computer need not have such devices. Additionally, a computer can be embedded in another device (such as, by way of example only, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver or a portable storage device (such as a universal serial bus (USB) flash drive)).
[0137] Computer-readable media for storing computer program instructions and data include all forms of non-volatile memory, media and storage devices, by way of example including semiconductor memory devices such as EPROM, EEPROM and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0138] To support interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user and a keyboard and a pointing device by which the user can provide input to the computer, the display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and the pointing device such as a mouse or a trackball. Other kinds of devices can also be used to support interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback such as visual feedback, auditory feedback or tactile feedback; and input from the user can be received in any form including acoustic, speech or tactile input. Additionally, a computer can interact with a user by sending files to and receiving documents from a device used by the user; for example, by sending a web page to a web browser in response to a request received from the web browser on the user's device. Additionally, a computer can interact with a user by sending a text message or other form of message to a personal device (such as a smart phone running a messaging application) and receiving a response message from the user in response.
[0139] The data processing apparatus for implementing a machine learning model may also include, for example, a dedicated hardware accelerator unit for processing the common and computationally intensive parts (i.e., inference, workload) of machine learning training or generation.
[0140] The machine learning model can be implemented and deployed using a machine learning framework (e.g., TensorFlow framework, Microsoft Cognitive Toolkit framework, Apache Singa framework, or Apache MXNet framework).
[0141] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface, a web browser, or an application through which a user can interact with an implementation of the subject matter described in this specification), or includes any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0142] The computing system can include a client and a server. Generally, the client and the server are located far apart from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on the respective computers and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to a user device acting as a client, for example, for the purpose of displaying data to a user interacting with the device and receiving user input from it. Data generated at the user device (e.g., the result of a user interaction) can be received at the server from the device.
[0143] Although this specification contains many specific implementation details, these details should not be construed as limiting the scope of any invention or the scope that may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features that are described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although features may be described as acting in certain combinations and initially claimed as such, in some cases, one or more features from a claimed combination can be excluded from the combination, and the claimed combination can relate to a sub-combination or a variant of a sub-combination.
[0144] Similarly, although the operations are depicted in the drawings in a particular order and recited in the claims, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed to achieve a desirable result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above-described embodiments should not be construed as requiring such separation in all embodiments, but rather it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.
[0145] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the acts recited in the claims may be performed in a different order and still achieve a desirable result. As one example, the processes depicted in the figures need not be in the particular order shown or in sequential order to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method for training a feed - forward generative neural network, the feed - forward generative neural network having a plurality of generative parameters and being configured to generate output audio examples using conditioning text inputs, Among them, where each conditioning text input includes a respective linguistic feature representation at each of a plurality of input time steps, wherein the feed - forward generative neural network is configured to receive a generative input including the conditioning text input and process the generative input to generate an audio output, the audio output including a respective audio sample at each of a plurality of output time steps, and wherein the training includes: obtaining training conditioning text inputs; processing a training generative input including the training conditioning text input using the feed - forward generative neural network according to a current value of the generative parameters to generate a training audio output; processing the training audio output using each of a plurality of discriminators, wherein: the plurality of discriminators includes one or more conditional discriminators, where each conditional discriminator processes a respective subset of the training audio output and the training conditioning text input to generate a prediction as to whether the training audio output is a real audio example or a synthetic audio example, and the plurality of discriminators includes one or more unconditional discriminators, where each unconditional discriminator processes a respective subset of the training audio output without processing the training conditioning text input to generate a prediction as to whether the training audio output is a real audio example or a synthetic audio example; determining a first combined prediction by combining the respective predictions of the plurality of discriminators; and determining an update to the current value of the generative parameters to increase a first error in the first combined prediction.
2. The method according to claim 1, wherein: each discriminator has a plurality of respective discriminator parameters, each conditional discriminator processes the training audio output and the training conditioning text input according to a current value of the respective discriminator parameters, each unconditional discriminator processes the training audio output according to a current value of the respective discriminator parameters without processing the training conditioning text input, and the method further includes determining an update to the current value of the discriminator parameters to reduce the first error in the first combined prediction.
3. The method according to claim 2, wherein, The training further includes: obtaining real audio examples and real conditioning text inputs including transcripts of the real audio examples; processing i) the real audio examples and the real conditioning text inputs using each of the conditional discriminators and processing ii) the real audio examples without processing the real conditioning text inputs using each of the unconditional discriminators, where each discriminator generates a prediction as to whether the real audio example is a real audio example or a synthetic audio example; determining a second combined prediction by combining the respective predictions of the plurality of discriminators; and determining an update to the current value of the discriminator parameters to reduce a second error in the second combined prediction.
4. The method according to claim 1, wherein The feed - forward generative neural network includes a sequence of groups of convolutional neural network layers, where each group includes one or more dilated convolutional layers.
5. The method according to claim 1, wherein, Each discriminator includes a discriminator neural network, and the discriminator neural network includes a sequence of groups of convolutional neural network layers, where each group includes one or more dilated convolutional layers.
6. The method according to claim 1, wherein The feed-forward generative neural network includes a sequence of groups of convolutional neural network layers, where one or more groups include one or more corresponding upsampling layers to account for a first ratio between the input time steps of the conditioning text input and the output time steps of the audio output.
7. The method according to claim 1, wherein Each discriminator includes a discriminator neural network, and the discriminator neural network includes a sequence of groups of convolutional neural network layers, where one or more groups include one or more corresponding downsampling layers to account for a second ratio between the output time steps of the audio output and the input time steps of the conditioning text input.
8. The method according to claim 1, wherein: The feed-forward generative neural network includes a sequence of groups of convolutional neural network layers, The method further includes zero-padding each training conditioning text input so that each training conditioning text input has a common dimension, and The method further includes using a zero-one masking to process the corresponding input to each convolutional neural network layer.
9. The method according to claim 1, wherein: Each corresponding subset of the training audio output is a proper subset of the training audio output, and At least two of the discriminators process different proper subsets of the training audio output.
10. The method according to claim 9, wherein, Processing the corresponding proper subset of the training audio output includes, for each discriminator: Obtaining a random sample of the training audio output, where the random sample includes a plurality of consecutive audio samples, and the size of the random sample for a given discriminator is predetermined; and Processing the random sample of the training audio output.
11. The method according to claim 10, wherein, For each conditional discriminator: Obtaining a random sample of the training audio output includes obtaining a random sample corresponding to a sequence of consecutive input time steps, and Processing the training conditioning text input includes processing the training conditioning text input with the sequence of consecutive input time steps.
12. The method according to claim 10, wherein Processing the random sample of the training audio output includes, for each discriminator, downsampling the random sample of the training audio output to generate a downsampled representation, where each discriminator downsamples the random sample by a predetermined downsampling factor.
13. The method according to claim 12, wherein: The corresponding predetermined downsampling factor for each discriminator corresponds to the size of the random sample for the discriminator; and Each downsampled representation has a common dimension for all discriminators.
14. The method according to claim 12, wherein, Downsampling the random sample of the training audio output includes using a strided convolutional neural network layer to process the random sample of the training audio output.
15. The method according to claim 1, wherein The corresponding linguistic feature representation of the conditioning text input at each of the input time steps includes one or more of the following: phonemes, durations, or log fundamental frequencies.
16. The method according to claim 1, wherein, The generative input further includes a noise input.
17. The method according to claim 16, wherein, The feed-forward generative neural network includes one or more conditional batch normalization neural network layers conditioned on a linear embedding of the noise input.
18. The method according to any one of claims 1-17, wherein The generative input further includes an identifier of the class to which the audio output should belong.
19. A method for generating an output audio example using a feed - forward generative neural network that has been trained using the method according to any one of claims 1 - 18.
20. A method for training a feed - forward generative neural network that has a plurality of generative parameters and is configured to generate an output audio example using a conditioning text input, Among them, where each conditioning text input includes a respective linguistic feature representation at each of a plurality of input time steps, wherein the feed - forward generative neural network is configured to receive a generative input that includes the conditioning text input and process the generative input to generate an audio output that includes a respective audio sample at each of a plurality of output time steps, and wherein the training includes: obtaining training conditioning text inputs; processing a training generative input that includes the training conditioning text input using the feed - forward generative neural network according to the current values of the generative parameters to generate a training audio output; processing the training audio output using each of a plurality of discriminators, wherein: each discriminator processes a respective first discriminative input that includes a respective true subset of the training audio output to generate a prediction as to whether the training audio output is a real audio example or a synthetic audio example, and at least two of the discriminators process different true subsets of the training audio output; determining a first combined prediction by combining the respective predictions of the plurality of discriminators; and determining an update to the current values of the generative parameters to increase a first error in the first combined prediction.
21. The method according to claim 20, wherein: each discriminator has a plurality of respective discriminative parameters, each discriminator processes the respective first discriminative input according to the current values of the respective discriminative parameters, and the method further includes determining an update to the current values of the discriminative parameters to reduce the first error in the first combined prediction.
22. The method according to claim 21, wherein The training further includes: obtaining real audio examples and real conditioning text inputs that include transcripts of the real audio examples; processing a second discriminative input that includes the real audio examples using each of the discriminators to generate a prediction as to whether the real audio examples are real audio examples or synthetic audio examples; determining a second combined prediction by combining the respective predictions of the plurality of discriminators; and determining an update to the current values of the discriminative parameters to reduce a second error in the second combined prediction.
23. The method according to claim 20, wherein, The feed - forward generative neural network includes a sequence of groups of convolutional neural network layers, where each group includes one or more dilated convolutional layers.
24. The method according to claim 20, wherein Each discriminator includes a discriminator neural network that includes a sequence of groups of convolutional neural network layers, where each group includes one or more dilated convolutional layers.
25. The method according to claim 20, wherein The feed - forward generative neural network includes a sequence of groups of convolutional neural network layers, where one or more groups include one or more respective upsampling layers to account for a first ratio between the input time steps of the conditioning text input and the output time steps of the audio output.
26. The method according to claim 20, wherein, Each discriminator includes a discriminator neural network, the discriminator neural network including a sequence of groups of convolutional neural network layers, where one or more of the groups include one or more respective downsampling layers to assume a second ratio between the output time steps of the audio output and the input time steps of the conditioning text input.
27. The method according to claim 20, wherein Processing the respective true subset of the training audio output includes, for each discriminator: Obtaining a random sample of the training audio output, where the random sample includes a plurality of consecutive audio samples, and where the size of the random sample for a given discriminator is predetermined; and Processing the random sample of the training audio output.
28. The method according to claim 27, wherein, Processing the random sample of the training audio output includes, for each discriminator, downsampling the random sample of the training audio output to generate a downsampled representation, where each discriminator downsamples the random sample by a predetermined downsampling factor.
29. The method according to claim 28, wherein: The respective predetermined downsampling factor for each discriminator corresponds to the size of the random sample for the discriminator; and Each downsampled representation has a common dimension for all discriminators.
30. The method according to claim 28, wherein, Downsampling the random sample of the training audio output includes using a strided convolutional neural network layer to process the random sample of the training audio output.
31. The method according to claim 20, wherein, The respective linguistic feature representations of the conditioning text input at each of the input time steps include one or more of the following: phonemes, durations, and log fundamental frequencies.
32. The method according to claim 20, wherein The generation input further includes a noise input.
33. The method according to claim 32, wherein The feedforward generation neural network includes one or more conditional batch normalization neural network layers conditioned on a linear embedding of the noise input.
34. The method according to any one of claims 20 - 33, wherein, The generation input further includes an identification of the class to which the audio output should belong.
35. A method for generating an output audio example using a feedforward generation neural network that has been trained using the method according to any one of claims 20 - 34.
36. A system including one or more computers and one or more storage devices, the one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the method according to any one of claims 1 - 35.
37. One or more non - transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the method according to any one of claims 1 - 35.
Citation Information
Patent Citations
Audio quality recovery system based on GAN
CN108877832A
Text-based insertion and replacement in audio narration
US20190130894A1