High fidelity speech synthesis using adversarial networks
By combining feedforward generative neural networks with adversarial training and using conditional and unconditional discriminators, the problems of high computational resource consumption and slow generation speed of autoregressive neural networks are solved, and the effect of quickly generating high-quality audio signals is achieved.
Patent Information
- Application Number
- CN202510888720.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-25
- Filing Date
- 2020-09-25
- Publication Date
- 2025-09-05
AI Technical Summary
In the existing technology, autoregressive neural networks consume large computing resources and have slow generation speed when generating audio signals, while reversible feedforward neural networks require explicit modeling of the audio data distribution, resulting in low efficiency in generating realistic audio signals.
A feedforward generative neural network is used in combination with conditional and unconditional discriminators for adversarial training. The receptive field is broadened by expanding the convolutional neural network layer. The conditional discriminator is used to evaluate the correspondence between audio and text, and the unconditional discriminator exposes more audio samples, reducing computational complexity and achieving rapid generation of high-quality audio signals.
It achieves the generation of high-quality audio signals in a single forward pass, reduces computing resources and time consumption, and does not require explicit modeling of the audio data distribution, thereby improving the efficiency and authenticity of the generated audio signals.
Smart Images

Figure CN120599997A_ABST
Abstract
Description
[0001] Description of the case
[0002] This application is a divisional application of Chinese invention patent application No. 202080068264.4, with the application date of September 25, 2020. Technical Field
[0003] This specification relates to using adversarial neural networks to generate audio data. Background Art
[0004] A neural network is a machine learning model that uses one or more layers of nonlinear units to predict outputs for received inputs. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to one or more other layers in the network (i.e., one or more other hidden layers, the output layer, or both). Each layer of the network generates an output from the received input based on the current values of the corresponding parameter set. Summary of the Invention
[0005] This specification describes a system implemented as a computer program on one or more computers in one or more locations that uses a generative neural network to generate output audio examples. The generative neural network has been configured through training to receive network inputs, including conditioning inputs representing input text. The generative neural network processes the conditioning inputs to generate audio data corresponding to the input text, such as audio data representing a speaker speaking the input text.
[0006] This specification also describes a training system for training a generative neural network. Generally speaking, the training system uses a set of one or more discriminator neural networks to train the generative neural network in an adversarial manner. That is, each discriminator neural network processes an audio example generated by the generative neural network and predicts whether the audio example is a real example of audio data (e.g., a recording of a human speaker) or a synthetic example of audio data, that is, whether the audio example was generated by the generative neural network. In this specification, the discriminator neural network is also referred to as a "discriminator."
[0007] In some implementations, the discriminator group includes both conditional and unconditional discriminators. The conditional discriminators process both the audio example and the conditioned text input to generate predictions, while the unconditional discriminators process only the audio example without the conditioned text input to generate predictions. In some implementations, each discriminator randomly samples different portions of the audio example and processes the random samples to generate predictions.
[0008] The training system may combine the respective predictions of the discriminator group to generate a combined prediction, and update parameters of both the feed-forward generative neural network and each discriminator based on the error of the combined prediction.
[0009] The subject matter described in this specification can be implemented in specific embodiments to realize one or more of the following advantages.
[0010] In some implementations described in this specification, the generative neural network can be a feedforward generative neural network. That is, the generative neural network can process the network input to generate an output audio example in a single forward pass. The feedforward generative neural network as described in this specification can generate output examples faster than the existing techniques that rely on autoregressive generative neural networks. The autoregressive neural network generates output examples across multiple time steps by performing a forward pass at each time step. At a given time step, the autoregressive neural network generates a new output sample to be included in the output audio example, which is adjusted based on the output samples that have already been generated. This process may consume a large amount of computing resources and take a lot of time. On the other hand, the feedforward generative neural network can generate an output example in a single forward pass while maintaining the high quality of the generated output example. This greatly reduces the time and amount of computing resources required to generate the output audio example.
[0011] Other existing techniques rely on reversible feedforward neural networks trained by using probability density extraction autoregressive models. Training in this manner allows the reversible feedforward neural network to generate speech signals that sound realistic and correspond to the input text without having to model every possible variation that occurs in the data. The feedforward generative neural network described in this specification can also generate realistic audio samples that faithfully follow the input text without having to explicitly model the data distribution of the audio data, but can do so without the extraction and reversibility requirements of the reversible feedforward neural network.
[0012] Using both conditional and unconditional discriminators provides various advantages to the feedforward generative neural network as described in this specification. The conditional discriminator can analyze the extent to which the generated audio corresponds to the input text represented by the conditioning text input, thereby allowing the feedforward generative neural network to learn to generate audio examples that follow the input text. However, as described in more detail below, the endpoints of the random samples obtained by the conditional discriminator need to be aligned with the input time steps of the conditioning text input so that the conditional discriminator evaluates the conditioning text input (at the frequency of the input time steps of the conditioning text input) and the generated audio (at the frequency of the output time steps of the audio examples). As a specific example, if each input time step corresponds to 120 output time steps, the sampling that the conditional discriminator can perform is limited to a lower frequency of 120 times. On the other hand, the unconditional discriminator is not limited to this sampling frequency constraint and is therefore exposed to a wider variety of audio samples.
[0013] Using a discriminator that processes only samples of audio data can allow the system to discriminate between lower-dimensional distributions. Assigning a specific window size to each discriminator can allow the discriminator to operate on different frequencies of the audio samples, thereby increasing the realism of the audio samples generated by the feedforward generative neural network. Using a discriminator that processes only samples of audio data can also reduce the computational complexity of the discriminator, which can allow the system to train the feedforward generative neural network more quickly.
[0014] Using dilated convolutional neural network layers can also broaden the receptive fields of the feedforward generator and discriminator, allowing the respective networks to learn dependencies in audio examples at various frequencies (e.g., long-term and short-term frequencies).
[0015] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the following description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a diagram of an example training system for training a generative neural network.
[0017] Figure 2 is a diagram of an example generator block.
[0018] Figure 3 is a diagram of an example discriminator neural network system.
[0019] Figure 4 is a diagram of an example unconditional discriminator block and an example conditional discriminator block.
[0020] Figure 5 is a flowchart of an example process for training a generative neural network.
[0021] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION
[0022] This specification describes a system for training a generative neural network to generate output audio examples using conditioned text input. The system can train the generative neural network in an adversarial manner using a discriminator neural network system comprising one or more discriminators.
[0023] Figure 1 is a diagram of an example training system 100 for training a generative neural network 110. Training system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations in which the systems, components, and techniques described below may be implemented.
[0024] The training system 100 includes a generative neural network 110, a discriminator neural network system 120, and a parameter update system 130. The training system 100 is configured to train the generative neural network 110 to receive a conditioned text input 102 and process the conditioned text input 102 to generate an audio output 112. In some implementations, the generative neural network 110 is a feed-forward neural network, i.e., the generative neural network 110 generates the audio output 112 in a single forward pass.
[0025] The conditioned text input 102 represents the input text, and the audio output 112 depicts speech corresponding to the input text. In some implementations, the conditioned text input 102 includes the input text itself, such as a character-level embedding or a word-level embedding of the input text. Alternatively or additionally, the conditioned text input may include linguistic features that characterize the text input. For example, the conditioned text input may include a corresponding vector of linguistic features for each input time step in the sequence of input time steps. As a specific example, the linguistic features for each input time step may include i) a phoneme and ii) the duration of the text at that input time step. The linguistic features may also include pitch information; for example, pitch may be determined by the logarithmic fundamental frequency of the input time step. express.
[0026] The generative neural network 110 can have any suitable neural network architecture. As a specific example, the generative neural network 110 can include a sequence of groups of convolutional neural network layers called "generator blocks." The first generator block in the sequence of generator blocks can receive as input a conditioned text input (or an embedding of a conditioned text input) and generate a block output. Each subsequent generator block in the sequence of generator blocks can receive as input a block output generated by a previous generator block in the sequence of generator blocks and generate a subsequent block output. Figure 2 Describe the generator block in more detail.
[0027] In some implementations, the generative neural network 110 may also receive a noise input 104 as input. For example, the noise input 104 may be randomly sampled from a predetermined distribution (e.g., a normal distribution). The noise input 104 may ensure variability in the audio output 112.
[0028] In some implementations, the generative neural network 110 may also receive as input an identification of a category 106 to which the audio output 112 should belong. The category 106 may be a member of a set of possible categories. For example, the category 106 may correspond to a particular speaker that the audio output 112 should sound like. That is, the audio output 112 may depict the particular speaker speaking the input text.
[0029] The audio output 112 may include an audio sample of the audio wave at each output time step in the sequence of output time steps. For example, for each output time step, the audio output 112 may include an amplitude value of the audio wave. In some implementations, the amplitude value may be a compressed or companded amplitude value.
[0030] In general, the sequence of input time steps and the sequence of output time steps represent the same time period, such as 1, 2, 5, or 10 seconds. As a specific example, if the time period is 2 seconds, the conditioning input 102 may include 400 input time steps (resulting in a frequency of 200 Hz), while the audio output 112 may include 48,000 time steps (resulting in a frequency of 24 kHz). Thus, the generative neural network 110 may generate audio samples for multiple output time steps (in this case, 120) for each single input time step.
[0031] In some implementations where the generative neural network 110 includes a sequence of one or more generator blocks, one or more of the generator blocks in the generative neural network 110 may include one or more corresponding upsampling layers due to a frequency difference between input time steps and output time steps. The dimension of the layer output of each upsampling layer is greater than the dimension of the layer input of the upsampling layer. The total degree of upsampling across all generator blocks in the generative neural network 110 may be proportional to the ratio of the frequency of the output time steps to the frequency of the input time steps.
[0032] After the generative neural network 110 generates the audio output 112, the training system 100 can provide the audio output 112 to the discriminator neural network system 120. The training system 100 can train the discriminator neural network system 120 to process the audio sample and generate a prediction 122 of whether the audio sample is real (i.e., an audio sample captured in the real world) or synthetic (i.e., an audio sample that has been generated by the generative neural network 110).
[0033] The discriminator neural network system 120 can have any suitable neural network architecture. As a specific example, the discriminator neural network system 120 can include one or more discriminators, each of which processes the audio output 112 and predicts whether the audio output 112 is real or synthesized. Each discriminator can include a sequence of groups of convolutional neural network layers, referred to as a "discriminator block." Figure 3 Describe an example discriminator block.
[0034] In some implementations, the one or more discriminators of the discriminator neural network system 120 include one or more conditional discriminators and one or more unconditional discriminators. The conditional discriminator receives as input i) the audio output 112 generated by the generative neural network 110 and ii) the conditioned text input 102 used by the generative neural network 110 to generate the audio output 112. The unconditional discriminator receives as input the output audio 112 generated by the generative neural network 110, but does not receive the conditioned text input 102 as input. Therefore, in addition to measuring the general authenticity of the audio output 112, the conditional discriminator can also measure the extent to which the audio output 112 corresponds to the input text represented by the conditioned text input 102, while the unconditional discriminator only measures the general authenticity of the audio output 112. Figure 4 An example discriminator neural network system is described in more detail.
[0035] Discriminator neural network system 120 may combine corresponding predictions of one or more discriminators to generate prediction 122 .
[0036] The parameter update system 130 can obtain the prediction 122 generated by the discriminator neural network system 120 and determine a parameter update 132 based on the error in the prediction 122. The training system can apply the parameter update 132 to the parameters of the generative neural network 110 and the discriminator neural network system 120. That is, the training system 100 can train the discriminator in the generative neural network 110 and the discriminator neural network system 120 simultaneously.
[0037] In general, the parameter update system 130 determines parameter updates 132 to the parameters of the generative neural network 110 to increase the error in the prediction 122. For example, if the discriminator neural network system 120 correctly predicts that the audio output 112 is synthetic, the parameter update system 130 generates a parameter update 132 to the parameters of the generative neural network 110 to improve the realism of the audio output 112 such that the discriminator neural network system 120 may incorrectly predict the next audio output 112 as real.
[0038] Conversely, the parameter update system 130 determines parameter updates 132 to the parameters of the discriminator neural network system 120 to reduce the error in the prediction 122. For example, if the discriminator neural network system 120 incorrectly predicts that the audio output 112 is real, the parameter update system 130 generates parameter updates 132 to the parameters of the discriminator neural network system 120 to improve the prediction 122 of the discriminator neural network system 120.
[0039] In some implementations where the discriminator neural network system 120 includes a plurality of different discriminators, the parameter update system 130 determines a parameter update 132 for each discriminator using the prediction 122 output by the discriminator neural network system 120. That is, the parameter update system 130 determines the parameter update 132 for each particular discriminator using the same combined prediction 122 generated by combining the respective predictions of the plurality of discriminators regardless of the respective predictions generated by the particular discriminator.
[0040] In some other implementations where the discriminator neural network system 120 includes multiple different discriminators, the parameter update system 130 uses the corresponding prediction generated by each discriminator to determine the parameter update 132 for the discriminator. That is, the parameter update system 130 generates the parameter update 132 for a particular discriminator in order to improve the corresponding prediction generated by that particular discriminator (which will indirectly improve the combined prediction 122 output by the discriminator neural network system 120 because the combined prediction 122 is generated based on the corresponding predictions generated by multiple discriminators).
[0041] During training, the training system 100 may also provide the real audio sample 108 to the discriminator neural network system 120. Each discriminator in the discriminator neural network system 120 may process the real audio sample 108 to predict whether the real audio sample 108 is a real example of audio data or a synthetic example of audio data. Again, the discriminator neural network system 120 may combine the respective predictions of each discriminator to generate a second prediction 122. The parameter update system 130 may then determine a second parameter update 132 for the parameters of the discriminator neural network system 120 based on the error in the second prediction 122. Generally speaking, the training system 100 does not use the second prediction corresponding to the real audio sample 108 to update the parameters of the generative neural network 110.
[0042] In some implementations where the generative neural network 110 has been trained to generate synthesized audio output 112 belonging to a class 106, the discriminator neural network system 120 does not receive as input an identification of the class 106 to which the received (real or synthesized) audio examples belong. However, the real audio samples 108 received by the discriminator neural network system 120 may include an audio sample belonging to each class 106 in the class set.
[0043] As a specific example, parameter update system 130 may use a Wasserstein loss function to determine parameter updates 132, which is:
[0044]
[0045] in is the likelihood that the real audio sample 108 is real as assigned by the discriminator neural network system 120, is the synthesized audio output 112 generated by the generative neural network 110, and is the likelihood that the synthesized audio output 112 is real, as assigned by the discriminator neural network system 120. The goal of the generative neural network 110 is to maximize , i.e., minimizing the Wasserstein loss by having the discriminator neural network system 120 predict that the synthesized audio output 112 is real. The goal of the discriminator neural network system 120 is to maximize the Wasserstein loss, i.e., correctly predict the real audio examples and the synthesized audio examples.
[0046] As another specific example, parameter updating system 130 may use the following loss function:
[0047]
[0048] Here again, the goal of the generator neural network 110 is to minimize the loss, while the goal of the discriminator neural network system 120 is to maximize the loss.
[0049] The training system 100 may backpropagate losses through both the generative neural network 110 and the discriminator neural network system 120, thereby training both networks simultaneously.
[0050] Figure 2 is a diagram of an example generator block 200. The generator block 200 may be a generator neural network (e.g. Figure 1 , which is configured to process conditioned text input and generate audio output. The generator neural network can include a sequence of generator blocks. In some implementations, each generator block in the sequence of generator blocks has the same architecture. In some other implementations, one or more generator blocks have a different architecture than other generator blocks in the sequence of generator blocks.
[0051] The generator block is configured to receive as input i) a block input 202 and ii) a noise input 204, and to generate a block output 206. In implementations where the generator neural network also receives as input an identification of the class to which the audio output should belong, one or more generator blocks of the generator neural network may also receive as input the identification of the class; for simplicity, this is omitted. Figure 2 Omitted in .
[0052] In some implementations, the first generator block in the generator block sequence receives the conditioned text input as the block input 202. In some other implementations, the first generator block in the generator block sequence receives an embedding of the conditioned text input generated by one or more initial neural network layers of the generator neural network. Each subsequent generator block in the generator block sequence receives the block output generated by the previous generator block in the generator block sequence as the block input 202.
[0053]
[0054] Table 1: Example generator neural network architecture
[0055] Table 1 describes an example architecture of a generator neural network. The example shown in Table 1 is for illustrative purposes only, and many different configurations of the generator neural network are possible. The input to the generator neural network is a vector of linguistic features corresponding to each of the 400 input time steps of two seconds of audio (at a frequency of 200 Hz). Each vector of linguistic features corresponding to the corresponding input time step includes 567 channels.
[0056] The generator neural network includes an input convolutional neural network layer that generates embeddings of linguistic features to provide to the first generator block in the sequence of generator blocks. The input convolutional neural network layer increases the number of channels per time step from 567 to 768. In some implementations, the generator neural network has multiple input neural network layers.
[0057] The generator neural network includes seven generator blocks ("G blocks"), although in general a generator neural network can include any number of generator blocks. The first two generator blocks do not include any upsampling layers, so the time dimension of the block output is the same as the time dimension of the block input. The next three generator blocks each upsample their respective block inputs by 2x, so that the block output of the fifth generator block has a time dimension of 3200 (at a frequency of 1600 Hz). The sixth generator block upsamples its block input by 3x, thereby generating a block output with a time dimension of 9600 (at a frequency of 4800 Hz). The last generator block upsamples its block input by 5x, thereby generating a block output with a time dimension of 48000 (at a frequency of 24 kHz). The sequence of generator blocks also reduces the number of channels per time step from 768 to 96.
[0058] The generator neural network includes an output convolutional neural network layer that processes the block output of the last generator block in the sequence of generator blocks and generates an audio output of the generator neural network. The audio output includes 48,000 output time steps, and each output time step has a single channel representing the amplitude of the audio wave at the output time step. In some implementations, the output convolutional neural network layer includes an activation function, such as a hyperbolic tangent activation function. In some implementations, the generator neural network has multiple output neural network layers.
[0059] Return Reference Figure 2 , the generator block 200 uses a first stack of neural network layers 212 to process the block input 202. The first stack of neural network layers 212 includes a batch normalization layer 212a, an activation layer 212b, an upsampling layer 212c, and a convolutional neural network layer 212d.
[0060] In some implementations, the batch normalization layer 212a is a conditional batch normalization layer that is conditioned on the noise input 204. For example, the conditional batch normalization layer 212a can be conditioned on the linear embedding of the noise input 204 generated by the linear layer 214 of the generator block 200. In some implementations, the linear layer 214 combines (e.g., concatenates) i) the linear embedding of the noise input 204 and ii) an identification of the class to which the output audio should belong (e.g., a specific speaker that the output audio should sound like) to generate a combined representation and provides the combined representation to the conditional batch normalization layer 212a. For example, the class identification can be encoded as a one-hot vector, i.e., a vector whose elements all have a value of 0 except for a single element corresponding to a specific class having a value of 1. The conditional batch normalization layer 212a can then be conditioned on the combined representation.
[0061] In some implementations, the activation layer 212b may be a ReLU activation layer.
[0062] For a generator block that upsamples the block input 202, such as the last five generator blocks listed in Table 1, the upsampling layer 212c is updated by the corresponding upsampling factor The output of the activation layer is upsampled. That is, the upsampling layer 212c generates the layer input in the time dimension For example, upsampling layer 212c may generate a layer output that includes, for each element of the layer input, As another example, the upsampling layer 212c may linearly interpolate each pair of consecutive elements in the layer input to produce As another example, the upsampling layer 212c may perform high-order interpolation on the elements of the layer input to generate the layer output. Figure 2 The generator block 200 depicted in FIG has a single upsampling layer 212 c , but in general a generator block may have any number of upsampling layers throughout the generator block.
[0063] The convolutional neural network layer 212d processes the channels of the layer input, and generates a The layer output has channels. corresponds to the number of channels in the block input 202, and corresponds to the number of channels in block output 206. In some cases, = .although Figure 2 The example depicted in FIG shows that the first convolutional neural network layer 212d of the generator block 200 processes a neural network having channels layer input to generate channels and each subsequent convolutional neural network layer preserves the number of channels, but in general any one or more convolutional layers of a generator block can change the number of channels to produce a network with Channel-by-channel block output.
[0064] The generator block 200 includes a second stack of neural network layers 216 that processes the output of the first stack of neural network layers 212. The second stack of neural network layers 216 includes a batch normalization layer 216a, an activation layer 216b, and a convolutional neural network layer 216c.
[0065] As described above, the batch normalization layer 216a can be a conditional batch normalization layer that is conditioned according to the linear embedding of the noise input 204 generated by the linear layer 218. In general, each linear layer 214, 218, 234, and 238 of the generator block 200 has different parameters and, therefore, generates a different linear embedding of the noise input 204.
[0066] Generator block 200 includes a first skip connection 226 that combines the input of first stack of neural network layers 212 with the output of second stack of neural network layers 216. For example, first skip connection 226 may add or concatenate the input of first stack of neural network layers 212 with the output of second stack of neural network layers 216.
[0067] In cases where the first stack 212 or the second stack 216 of neural network layers includes an upsampling layer (e.g., upsampling layer 212c), the generator block may include another upsampling layer 222 before the first skip connection 226 such that the two inputs to the first skip connection 226 have the same dimension.
[0068] In the case where the first stack 212 or the second stack 216 of neural network layers includes a convolutional neural network layer (e.g., convolutional neural network layer 212d) that changes the number of channels of the layer input, the generator block may include another convolutional neural network layer 224 before the first skip connection 226 such that the two inputs of the first skip connection 226 have the same number of channels.
[0069] The generator block 200 includes a third stack 232 of neural network layers that processes the output of the first skip connection 226. The third stack 232 of neural network layers includes a batch normalization layer 232a, an activation layer 232b, and a convolutional neural network layer 232c. As described above, the batch normalization layer 232a can be a conditional batch normalization layer that is conditioned according to the linear embedding of the noise input 204 generated by the linear layer 234.
[0070] The generator block 200 includes a fourth stack 236 of neural network layers that processes the output of the third stack 232 of neural network layers. The fourth stack 236 of neural network layers includes a batch normalization layer 236a, an activation layer 236b, and a convolutional neural network layer 236c. As described above, the batch normalization layer 236a can be a conditional batch normalization layer that is conditioned based on the linear embedding of the noise input 204 generated by the linear layer 238.
[0071] The generator block 200 includes a second skip connection 242 that combines the input of the third stack 232 of neural network layers and the output of the fourth stack of neural network layers, for example, by adding or concatenating. The block output 206 of the generator module 200 is the output of the second skip connection 242.
[0072] In some implementations, one or more of the convolutional neural network layers of the generator block 200 is a dilated convolutional neural network layer. A dilated convolutional neural network layer is a convolutional layer in which a filter is applied over an area larger than the length of the filter by skipping input values with a certain step defined by the dilation value of the dilated convolution. In some implementations, the generator block 200 may include multiple dilated convolutional neural network layers with increasing dilation. For example, for each dilated convolutional neural network layer, the dilation value may be doubled starting from the initial dilation and then fed back to the initial dilation in the next generator block. Figure 2In the example depicted in , the first convolutional neural network layer 212d has a dilation value of 1, the second convolutional neural network layer 216c has a dilation value of 2, the third convolutional neural network layer 232c has a dilation value of 4, and the fourth convolutional neural network layer 236c has a dilation value of 8.
[0073] Figure 3 is a diagram of an example discriminator neural network system 300. The discriminator neural network system 300 may be a training system (e.g., Figure 1 ) components of the training system 100 depicted in . The discriminator neural network system 300 has been trained to receive an audio example 302 and generate a prediction 306 as to whether the audio example is a real audio example or a synthetic audio example generated by a generator neural network. The audio example may include a corresponding amplitude value at each output time step in the sequence of output time steps (referred to as an "output" time step because the audio example may have been the output of the generator neural network). The audio example corresponds to a conditioning text input 304 that includes a corresponding vector of linguistic features at each input time step in the sequence of input time steps. Note that even though the audio example 302 is a real audio example, the conditioning text input 304 still corresponds to the audio example 302.
[0074] The discriminator neural network system 300 may include one or more unconditional discriminators, one or more conditional discriminators, or both. Figure 3 In the example depicted in , the discriminator neural network system 300 includes five unconditional discriminators (including unconditional discriminator 320) and five conditional discriminators (including conditional discriminator 340).
[0075] In some implementations, instead of processing the entire audio example 302, each discriminator in the discriminator neural network system 300 processes a different subset of the audio example. For example, each discriminator can randomly sample a subset of the audio example 302. That is, each discriminator processes only the amplitudes of a subsequence of consecutive output time steps of the audio sample 302, where the subsequence of consecutive output time steps is randomly sampled from the entire sequence of output time steps. In some implementations, the size of the random sample (i.e., the number of output time steps sampled) is the same for each discriminator. In some other implementations, the size of the random sample is different for each discriminator and is referred to as the "window size" of the discriminator.
[0076] In some implementations, one or more discriminators in the discriminator neural network system 300 may have the same network architecture and the same parameter values. For example, the training system may update the one or more discriminators in the same manner during training of the discriminator neural network system.
[0077] In some implementations, different discriminators have different window sizes. For example, for each window size in a set of multiple window sizes, the system can include one conditional discriminator with that window size and one unconditional discriminator with that window size. As a specific example, the discriminator neural network 300 includes one conditional discriminator and one unconditional discriminator for each window size in a set of five window sizes (240, 480, 960, 1920, and 3600 output time steps).
[0078] Each conditional discriminator also obtains samples of the conditioned text input 304 that correspond to the random samples of the audio samples 302 obtained by the conditional discriminator. Because the samples of the audio example 302 and the samples of the conditioned text input 304 must be aligned, the conditional discriminator can be constrained to sample a subsequence of the audio output 302 that begins at the same point as the point at which the input time step of the conditioned text input 304 begins. That is, because the number of input time steps in the conditioned text input 304 is less than the number of output time steps in the audio example 302, each conditional discriminator can be constrained to sample the audio example 302 at a point that aligns with the input time step of the conditioned text input 304. Because the unconditional discriminator does not process the conditioned text input 304, the unconditional discriminator does not have this constraint.
[0079] In some implementations, before processing a random sample of audio examples 302, each discriminator first downsamples the random sample using a "shaping" layer by a factor proportional to the discriminator's window size. Downsampling effectively allows discriminators with different window sizes to process audio examples 302 at different frequencies, where the frequency at which a particular discriminator operates is proportional to the window size of the particular discriminator. Downsampling by a factor proportional to the window size also allows each downsampled representation to have a common dimensionality, which allows each discriminator to have a similar architecture and similar computational complexity, despite having different window sizes. Figure 3 In the example depicted in , the common dimension is 240 time steps; each discriminator uses the factor Mark, the discriminator passes the factor (1, 2, 4, 8, and 15, respectively) to downsample their random samples.
[0080] In some implementations, the discriminator uses a strided convolutional neural network layer to downsample corresponding random samples of the audio examples 302.
[0081] The unconditional discriminator 320 randomly samples samples 312 comprising 480 output time steps from the audio example 302. The unconditional discriminator 320 includes a reshaping layer 322 that downsamples the random samples 312 by 2x to generate a network input comprising 240 time steps.
[0082] The unconditional discriminator 320 includes a sequence 324 of discriminator blocks ("DBlocks"). Although the unconditional discriminator 320 is depicted as including five discriminator blocks, in general, an unconditional discriminator can have any number of discriminator blocks. Each discriminator block in the unconditional discriminator 320 is an unconditional discriminator block. Figure 4 Describes an example unconditional discriminator block.
[0083] The first discriminator block in the discriminator block sequence 324 is configured to receive a network input from the shaping layer 322 and process the network input to generate a block output. Each subsequent discriminator block in the discriminator block sequence 324 is configured to receive as input the block output generated by the previous discriminator block in the sequence 324 and generate a subsequent block output.
[0084] In some implementations, one or more discriminator blocks in the unconditional discriminator 320 perform further downsampling. Figure 3 As shown, the second discriminator block in the sequence of discriminator blocks 324 performs 5x downsampling, while the third discriminator block in the sequence 324 performs 3x downsampling. In general, downsampling helps the unconditional discriminator 320 limit the dimensionality of the internal representation, thereby improving efficiency and allowing the unconditional discriminator 320 to learn relationships between more distant elements of the samples 312 of the audio example 302. In some implementations, the number of downsampling layers in a corresponding unconditional discriminator and the degree of downsampling performed by the corresponding unconditional discriminator depend on a factor k of the corresponding unconditional discriminator.
[0085] The unconditional discriminator 320 includes an output layer 326 that receives as input the block output of the last discriminator block in the discriminator block sequence 324 and generates a prediction 328. The prediction 328 predicts whether the audio example is a real audio example or a synthesized audio example. For example, the output layer 326 can be an average pooling layer that generates a scalar representing the likelihood that the audio example 302 is a real audio example, where a larger scalar value indicates a higher confidence that the audio example 302 is real.
[0086] The conditional discriminator 340 randomly samples samples 332 comprising 3600 output time steps from the audio example 302. The conditional discriminator 340 includes a reshaping layer 342 that downsamples the random samples 332 by 15x to generate a network input comprising 240 time steps.
[0087] The conditional discriminator 340 includes a sequence of discriminator blocks 344. Although the conditional discriminator 340 is depicted as including five unconditional discriminator blocks and one conditional discriminator block, in general the conditional discriminator may have any number of conditional and unconditional discriminator blocks.
[0088] The first discriminator block in the discriminator block sequence 344 is configured to receive a network input from the shaping layer 342 and process the network input to generate a block output. Each subsequent discriminator block in the discriminator block sequence 344 is configured to receive as input a block output generated by a previous discriminator block in the sequence 344 and generate a subsequent block output.
[0089] The conditional discriminator block is configured to receive as input i) the block output of the previous discriminator block in the sequence 344 and ii) a sample of the conditioned text input 304 corresponding to the sample 332 of the audio example. Figure 4 Describes an example conditional discriminator block.
[0090] Due to the frequency difference between the output time steps and the input time steps, one or more unconditional discriminator blocks preceding the conditional discriminator block in the discriminator block sequence 344 and / or the conditional discriminator block itself may perform downsampling. For example, the total degree of downsampling across all discriminator blocks in the conditional discriminator 340 may be proportional to the ratio of the frequency of the input time steps to the frequency of the output time steps. Figure 3 In the example depicted in , the conditional discriminator 340 downsamples the network input by 8x to achieve a frequency of 200 Hz, which is the same frequency as the conditioned text input.
[0091] The conditional discriminator 340 includes an output layer 346 that receives as input the block output of the last discriminator block in the sequence of discriminator blocks 344 and generates a prediction 348 of whether the audio example is a real audio example or a synthesized audio example. For example, the output layer 346 can be an average pooling layer that generates a scalar representing the likelihood that the audio example 302 is a real audio example.
[0092] After each discriminator in the discriminator neural network system 300 generates a prediction, the discriminator neural network system 300 can combine the corresponding predictions to generate a final prediction 306 of whether the audio example is a real audio example or a synthesized audio example. For example, the discriminator neural network system 300 can determine the sum or average of the corresponding predictions of the discriminators. As another example, the discriminator neural network system 300 can generate the final prediction 306 based on a voting algorithm, for example, predicting that the audio example 302 is a real audio example if and only if a majority of the discriminators predict that the audio example 302 is a real audio example.
[0093] Note that the discriminator neural network system 300 does not receive the class associated with the audio input as input. However, in some implementations, the discriminator neural network system 300 may do so, for example as a further conditioning input to the conditioning discriminator block of each conditional discriminator.
[0094] Figure 4 is a diagram of an example unconditional discriminator block 400 and an example conditional discriminator block 450. The discriminator blocks 400 and 450 may be discriminator neural network systems (e.g., Figure 1 Components of the discriminator neural network system 120 depicted in FIG. The discriminator neural network system may include one or more discriminators, and each discriminator may include one or more discriminator block sequences. In some implementations, each discriminator block in the sequence of discriminator blocks in the discriminator has the same architecture. In some other implementations, one or more discriminator blocks have a different architecture than other discriminator blocks in the sequence of discriminator blocks in the discriminator.
[0095] Unconditional discriminator block 400 can be a component of one or more unconditional discriminators, one or more conditional discriminators, or both. Conditional discriminator block 450 can simply be a component of one or more conditional discriminators. That is, a conditional discriminator can include an unconditional discriminator block, but an unconditional discriminator cannot include a conditional discriminator block.
[0096] The unconditional discriminator block 400 is configured to receive a block input 402 and generate a block output 404. In some implementations, if the unconditional discriminator block 400 is the first discriminator block in a sequence of discriminator blocks, the block input 402 is an audio example. In some other implementations, if the unconditional discriminator block 400 is the first discriminator block in a sequence of discriminator blocks, the block input 402 is an embedding of the audio example, for example, an embedding generated by an input convolutional neural network layer of the discriminator. If the unconditional discriminator block 400 is not the first discriminator block in the sequence, the block input 402 can be the block output of the previous discriminator block in the sequence.
[0097] The unconditional discriminator block 400 includes a first stack 412 of neural network layers, including a downsampling layer 412a, an activation layer 412b, and a convolutional neural network layer 412c.
[0098] For an unconditional discriminator block that downsamples the block input 402 (e.g., Figure 3 For example, the second and third discriminator blocks in the sequence of discriminator blocks 324 depicted in FIG, the downsampling layer 412a downsamples the block input 402 by the corresponding downsampling factor. Figure 4The unconditional discriminator block 400 depicted in FIG. 4 has a single downsampling layer 412 a , but in general the unconditional discriminator block may have any number of downsampling layers throughout the unconditional discriminator block.
[0099] In some implementations, the activation layer 412b may be a ReLU activation layer.
[0100] Convolutional neural network layer 412c processes the data including for each element channels of the layer input, and generates a The layer output has channels. corresponds to the number of channels in the block input 202, and Corresponds to the multiplier. In some cases, In general, any one or more convolutional layers in an unconditional discriminator block can change the number of channels of the layer input.
[0101] The unconditional discriminator block 400 includes a second stack of neural network layers 414 that processes the output of the first stack of neural network layers 412. The second stack of neural network layers 414 includes an activation layer 414a and a convolutional neural network layer 414b.
[0102] The unconditional discriminator block 400 includes a skip connection 426 that combines the input of the first stack of neural network layers 412 with the output of the second stack of neural network layers 414. For example, the skip connection 426 can add or concatenate the input of the first stack of neural network layers 412 with the output of the second stack of neural network layers 414.
[0103] In the case where the first stack 412 or the second stack 414 of neural network layers includes a convolutional neural network layer that changes the number of channels of the layer input (e.g., convolutional neural network layer 412c), the unconditional discriminator block may include another convolutional neural network layer 422 before the skip connection 426 such that the two inputs to the skip connection 426 have the same number of channels.
[0104] In cases where the first stack 412 or the second stack 414 of neural network layers includes a downsampling layer (e.g., downsampling layer 412a), the unconditional discriminator block may include another downsampling layer 424 before the skip connection 426 such that the two inputs to the skip connection 426 have the same dimension.
[0105] In some implementations, one or more convolutional neural network layers in the unconditional discriminator block 400 are dilated convolutional layers.
[0106] The conditional discriminator block 450 is configured to receive a block input 452 corresponding to an audio example and a conditioning text input 454 , and generate a block output 456 .
[0107] The conditional discriminator block 450 includes a first stack 462 of neural network layers, including a downsampling layer 462a, an activation layer 462b, and a convolutional neural network layer 462c.
[0108] For the conditional discriminator block that downsamples the block input 452, the downsampling layer 462a downsamples the block input 452 by the corresponding downsampling factor. Figure 4 The conditional discriminator block 450 shown in FIG has a single downsampling layer 462 a, but in general the conditional discriminator block can have any number of downsampling layers throughout the conditional discriminator block. In particular, the conditional discriminator block 450 downsamples the layer input 452 so that it has the same dimensions in the temporal dimension as the conditioned text input 454.
[0109] In some implementations, the activation layer 462b may be a ReLU activation layer.
[0110] Convolutional neural network layer 462c processes the data including for each element channels of the layer input, and generates a channels of the layer output; in this example, In general, any one or more convolutional layers in a conditional discriminator block can change the number of channels of the layer input.
[0111] The conditional discriminator block 450 includes a convolutional neural network layer 472 that processes the conditioned text input 454 and generates a conditional discriminator block 450 that includes a conditional discriminator block 450 for each element. That is, convolutional neural network layer 472 changes the number of channels (in this case, 567) of the conditioned text input to match the output of convolutional neural network layer 462c in the first stack of neural network layers 462.
[0112] The conditional discriminator block 450 includes a combining layer 474 that combines the outputs of the convolutional neural network layers 462c and 472, for example, by addition or by concatenation.
[0113] The conditional discriminator block 450 includes a second stack of neural network layers 482 that processes the output of the combination layer 474. The second stack of neural network layers 482 includes an activation layer 482a and a convolutional neural network layer 482b.
[0114] The conditional discriminator block 450 includes a skip connection 496 that combines the input of the first stack of neural network layers 462 with the output of the second stack of neural network layers 482. For example, the skip connection 496 can add or concatenate the input of the first stack of neural network layers 462 with the output of the second stack of neural network layers 482.
[0115] In the case where the first stack 462 or the second stack 482 of neural network layers includes a convolutional neural network layer that changes the number of channels of the layer input (e.g., convolutional neural network layer 462c), the conditional discriminator block 450 may include another convolutional neural network layer 492 before the skip connection 496 such that the two inputs to the skip connection 496 have the same number of channels.
[0116] In cases where the first stack 462 or the second stack 482 of neural network layers includes a downsampling layer (e.g., downsampling layer 462a), the conditional discriminator block 450 may include another downsampling layer 494 before the skip connection 496 such that the two inputs to the skip connection 496 have the same dimensions.
[0117] In some implementations, one or more convolutional neural network layers in the conditional discriminator block 450 are dilated convolutional layers.
[0118] Figure 5 is a flow chart of an example process 500 for training a generative neural network. For convenience, process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a training system (e.g., Figure 1 The depicted training system 100 ) may perform process 500 .
[0119] The generative neural network has a plurality of generative parameters. The generative neural network may include a sequence of groups of convolutional neural network layers, each group including one or more dilated convolutional neural network layers. One or more groups or blocks of the generative neural network may also include one or more upsampling layers to account for adjusting the ratio between input time steps of the text input and output time steps of the audio output.
[0120] The system obtains a training conditioning text input (step 502). The training conditioning text input may include a corresponding linguistic feature representation at each input time step in a plurality of input time steps. For example, the linguistic feature representation at each input time step may include a phoneme, duration, and logarithmic fundamental frequency at that time step.
[0121] The system processes the training generated input including the training conditioned text input using the generative neural network to generate a training audio output (step 504). The system generates the training audio output based on the current values of the generation parameters of the generative neural network. The training audio output can include a corresponding audio sample at each output time step in a plurality of output time steps.
[0122] The training generated input may also include a noise input. The generated neural network may include one or more conditional batch normalization neural network layers conditioned according to a linear embedding of the noise input. In some implementations, each conditional batch normalization neural network layer is conditioned according to a different linear embedding of the noise input.
[0123] In some implementations, generating the input includes an identification of a class to which the output wave should belong. In some such implementations, the conditional batch normalization neural network layer is further conditioned based on the identification of the class.
[0124] In some implementations, the system zero-pads the training conditioning text inputs (i.e., adds one or more zeros to the end of the training conditioning text inputs) before providing them to the generative neural network so that each training conditioning text input has the same dimensions. Batching fixed-size training examples can be more efficient for deep learning models than sequentially processing training examples of different sizes.
[0125] However, because the generative neural network has multiple convolutional neural network layers, convolution may propagate non-zero values into zero-padded elements, which causes interference with non-zero padded elements in later convolutions. To avoid this, the system can apply a convolution mask to the input of each convolutional neural network layer. The convolution mask can be a zero-one mask, where 0 is at the elements corresponding to the zero-padded elements of the input and 1 is at the elements corresponding to the non-zero padded elements of the input. Applying the convolution mask ensures that the zero-padded elements of the input do not interfere with the non-zero padded elements of the input. The zero-one mask can be upsampled at the same rate as the conditional input, so that the appropriate number of zero-padded elements are processed before each corresponding convolutional neural network layer. The system can remove the zero-padded elements before outputting the final audio example. By zero-padding and then applying the convolution mask, the system is able to generate audio examples of arbitrary length.
[0126] The system processes the training audio output using each of the plurality of discriminators to generate a corresponding prediction of whether the training audio output is real or synthesized (step 506).
[0127] In some implementations, the plurality of discriminators includes one or more conditional discriminators and one or more unconditional discriminators.
[0128] In some implementations, one or more of the plurality of discriminators processes different proper subsets of the training audio output. In some such implementations, the size of the proper subset for a particular discriminator is predetermined. In some such implementations, each discriminator downsamples the proper subset by a predetermined downsampling factor corresponding to the size of the proper subset. In some such implementations, each discriminator downsamples the proper subset using a strided convolution.
[0129] The system determines a combined prediction by combining the corresponding predictions of the multiple discriminators (step 508). For example, the system can determine the average of the predictions or use a voting algorithm to process the predictions.
[0130] The system determines updates to the current values of the generation parameters to increase the error in the combined prediction (step 510). In some implementations, the system may also determine updates to the current values of the discrimination parameters of the plurality of discriminators to reduce the error in the combined prediction.
[0131] This specification uses the term "configured" in connection with systems and computer program components. A system of one or more computers is configured to perform a particular operation or action if the system has software, firmware, hardware, or a combination thereof installed thereon that, when operated, causes the system to perform the operation or action. One or more computer programs are configured to perform a particular operation or action if the program or programs include instructions that, when executed by a data processing device, cause the device to perform the operation or action.
[0132] The embodiments of the subject matter and functional operations described in this specification may be implemented in digital electronic circuits, in tangibly implemented computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents) or in a combination of one or more thereof. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory or a combination of one or more thereof. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical or electromagnetic signal that is generated as encoded information to be transmitted to a suitable receiver device for execution by a data processing device).
[0133] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of apparatus, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. The apparatus may also be or further include special-purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the apparatus may optionally include code that creates an execution environment for a computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.
[0134] A computer program (which may also be referred to or described as a program, software, software application, app, module, software module, script, or code) may be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and may be deployed in any form, including as a standalone program or a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to it, or in multiple collaborating files (e.g., files that store portions of one or more modules, subroutines, or code). A computer program may be deployed to execute on one computer or on multiple computers that are located at one site or distributed among multiple sites and interconnected by a data communications network.
[0135] In this specification, the term "database" is used broadly to refer to any collection of data: the data need not be structured in any particular way, or at all, and may be stored on storage devices in one or more locations. Thus, for example, an index database may include multiple collections of data, each of which may be organized and accessed differently.
[0136] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a specific engine; in other cases, multiple engines can be installed and run on the same one or more computers.
[0137] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by a dedicated logic circuit (such as an FPGA or ASIC), or by a combination of a dedicated logic circuit and one or more programmed computers.
[0138] A computer suitable for executing a computer program can be based on a general or special microprocessor or both, or any other type of central processing unit. In general, the central processing unit will receive instructions and data from a read-only memory or random access memory or both. The basic elements of a computer are a central processing unit for executing or running instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by a dedicated logic circuit or incorporated into the dedicated logic circuit. In general, a computer can include one or more large-capacity storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data or can be operatively coupled to one or more large-capacity storage devices to receive data from them or transfer data to them or receive and transmit both. However, a computer does not necessarily have such a device. In addition, a computer can be embedded in another device (e.g., to give only a few examples, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive)).
[0139] Computer-readable media for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0140] To support interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices can also be used to support interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, voice, or tactile input. In addition, a computer can interact with a user by sending files to and receiving files from a device used by the user; for example, by sending a web page to a web browser in response to a request received from the web browser on the user's device. In addition, a computer can interact with a user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving a response message from the user in return.
[0141] The data processing apparatus for implementing the machine learning model may also include, for example, dedicated hardware accelerator units for processing the common and computationally intensive parts of machine learning training or generation (i.e., inference workloads).
[0142] The machine learning model can be implemented and deployed using a machine learning framework (e.g., the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework).
[0143] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component (e.g., as a data server) or includes a middleware component (e.g., an application server) or includes a front-end component (e.g., a client computer with a graphical user interface, a web browser, or an application through which a user can interact with an implementation of the subject matter described in this specification), or includes any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0144] A computing system may include a client and a server. Generally speaking, the client and the server are remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to a user device acting as a client, for example, for the purpose of displaying data to a user interacting with the device and receiving user input from the user. Data generated at the user device (e.g., the result of the user interaction) can be received from the device at the server.
[0145] Although this specification contains many specific implementation details, these details should not be interpreted as limiting the scope of any invention or the scope of possible protection requests, but should be interpreted as descriptions of features that may be specific to a particular embodiment of a particular invention. Certain features described in this specification in the context of separate embodiments may also be implemented in a single embodiment in combination. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may be described as working in some combinations and initially requested as such, in some cases, one or more features from the claimed combination may be excluded from the combination, and the claimed combination may involve a variant of a sub-combination or a sub-combination.
[0146] Similarly, although operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed to achieve a desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments. Instead, it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.
[0147] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As an example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desired results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method of training a generative neural network having a plurality of generative parameters and configured to generate output audio examples using conditioned text input, in, Each conditioned text input represents text at each input time step in a plurality of input time steps, wherein the generative neural network is configured to receive a generative input comprising a conditioned text input and process the generative input to generate an audio output comprising a corresponding audio sample at each of a plurality of output time steps, and The training includes: Get training adjustment text input; processing a training generation input comprising the training conditioned text input using the generation neural network according to current values of the generation parameters to generate a training audio output; Processing the training audio output using each of a plurality of discriminators includes: generating a respective discriminant input corresponding to each of the plurality of discriminators; processing each respective discriminant input using a corresponding discriminator to generate a respective prediction of whether the training audio output is a real audio example or a synthesized audio example; and An update to a current value of the generated parameter is determined based on the respective predictions of the plurality of discriminators.
2. The method according to claim 1, wherein Generating a respective discriminant input corresponding to each of the plurality of discriminators includes, for each discriminator, obtaining a respective sample of the training audio output, wherein the respective sample includes a plurality of consecutive audio samples.
3. The method according to claim 2, wherein: Each of the plurality of discriminators has a respective corresponding size, and wherein at least two of the plurality of discriminators have different respective corresponding sizes.
4. The method according to claim 3, wherein: Processing each corresponding discriminant input using the corresponding discriminator includes, for each discriminator, downsampling a corresponding sample of the training audio output to generate a downsampled representation, wherein each discriminator downsamples the corresponding sample by a corresponding predetermined downsampling factor.
5. The method according to claim 4, wherein: The corresponding predetermined downsampling factor for each discriminator corresponds to the corresponding size for the discriminator; as well as Each downsampled representation has a common dimension for all discriminators.
6. The method according to claim 4, wherein: Downsampling the corresponding samples of the training audio output includes processing the corresponding samples of the training audio output using a strided convolutional neural network layer.
7. The method according to claim 2, wherein: Obtaining corresponding samples of the training audio output includes obtaining random samples of the training audio output.
8. The method according to claim 1, wherein Each respective discriminator input comprises a respective proper subset of the training audio outputs.
9. The method according to claim 1, wherein At least two of the respective discriminator inputs comprise different proper subsets of the training audio outputs.
10. The method according to claim 1, wherein The plurality of discriminators includes one or more conditional discriminators, and wherein generating a respective discriminative input corresponding to each of the plurality of discriminators includes, for each of the one or more conditional discriminators: Obtaining a corresponding sample of the training audio output, wherein the corresponding sample comprises a plurality of consecutive audio samples; and Retrieve corresponding samples of the training adjusted text input corresponding to the corresponding samples of the training audio output.
11. The method according to claim 1, wherein The generation input also includes an identification of a category to which the audio output should belong.
12. The method according to claim 1, wherein The generative neural network is a feedforward generative neural network.
13. The method according to claim 1, wherein Determining an update to a current value of the generated parameter based on the respective predictions of the plurality of discriminators comprises: determining a combined prediction by combining the respective predictions; and The update to the current value of the generated parameter is determined to increase the error in the combined prediction.
14. The method of claim 1, wherein: Each discriminator has multiple corresponding discriminant parameters, Each discriminator processes the corresponding discriminant input according to the current value of the corresponding discriminant parameter, and The method also includes determining an update to the current value of the discrimination parameter based on the respective predictions of the plurality of discriminators.
15. The method according to claim 1, wherein The generative neural network comprises a sequence of groups of convolutional neural network layers, wherein each group comprises one or more dilated convolutional layers.
16. The method according to claim 1, wherein Each discriminator comprises a discriminator neural network comprising a sequence of groups of convolutional neural network layers, wherein each group comprises one or more dilated convolutional layers.
17. The method according to claim 1, wherein The generative neural network comprises a sequence of groups of convolutional neural network layers, wherein one or more groups comprise one or more respective upsampling layers to account for the first ratio between input time steps of the conditioned text input and output time steps of the audio output.
18. The method according to claim 1, wherein Each discriminator comprises a discriminator neural network comprising a sequence of groups of convolutional neural network layers, wherein one or more groups comprise one or more respective downsampling layers to account for a second ratio between output time steps of the audio output and input time steps of the conditioned text input.
19. A system comprising one or more computers and one or more storage devices storing instructions operable when executed by the one or more computers to cause the one or more computers to perform operations of training a generative neural network having a plurality of generative parameters and configured to generate output audio examples using conditioned text input, in, Each conditioned text input represents text at each input time step in a plurality of input time steps, wherein the generative neural network is configured to receive a generative input comprising a conditioned text input and process the generative input to generate an audio output comprising a corresponding audio sample at each of a plurality of output time steps, and The training includes: Get training adjustment text input; processing a training generation input comprising the training conditioned text input using the generation neural network according to current values of the generation parameters to generate a training audio output; Processing the training audio output using each of a plurality of discriminators includes: generating a respective discriminant input corresponding to each of the plurality of discriminators; processing each respective discriminant input using a corresponding discriminator to generate a respective prediction of whether the training audio output is a real audio example or a synthesized audio example; and An update to a current value of the generated parameter is determined based on the respective predictions of the plurality of discriminators.
20. One or more non-transitory computer storage media encoded with computer program instructions that, when executed by a plurality of computers, cause the plurality of computers to perform operations of training a generative neural network having a plurality of generative parameters and configured to generate output audio examples using conditioned text input, in, Each conditioned text input represents text at each input time step in a plurality of input time steps, wherein the generative neural network is configured to receive a generative input comprising a conditioned text input and process the generative input to generate an audio output comprising a corresponding audio sample at each of a plurality of output time steps, and The training includes: Get training adjustment text input; processing a training generation input comprising the training conditioned text input using the generation neural network according to current values of the generation parameters to generate a training audio output; Processing the training audio output using each of a plurality of discriminators includes: generating a respective discriminant input corresponding to each of the plurality of discriminators; processing each respective discriminant input using a corresponding discriminator to generate a respective prediction of whether the training audio output is a real audio example or a synthesized audio example; and An update to a current value of the generated parameter is determined based on the respective predictions of the plurality of discriminators.