Audio Generation Using an Auto-Regressive Generative Neural Network
The system addresses the limitations of conventional audio generation by using generative neural networks to condition on semantic representations and process acoustic signals, achieving high-quality, consistent audio generation across diverse data and speaker identities, reducing computational resources and training time.
Patent Information
- Application Number
- JP2024531121
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-01-26
- Filing Date
- 2023-09-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-09-07
AI Technical Summary
Conventional systems for generating audio struggle to produce high-quality, consistent speech and music without text annotation, and are limited by acoustic diversity and quality, often requiring clean speech training data and failing to maintain speaker identity and recording conditions.
A system utilizing generative neural networks to generate audio signals by conditioning on semantic representations, which include decoder neural networks to process acoustic representations, using a hierarchy of vector quantizers for residual vector quantization to enhance audio quality and consistency, allowing for robust training on diverse and noisy data.
The system generates high-quality, long-term consistent audio with maintained speaker identity and recording conditions, reducing computational resources and training time by using pre-trained components, and effectively producing text-conditioned music without large paired datasets.
Smart Images

Figure 0007701566000001 
Figure 0007701566000002 
Figure 0007701566000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Application No. 63 / 404,528, filed Sep. 7, 2022, and U.S. Provisional Application No. 63 / 441,412, filed Jan. 26, 2023. The disclosure of the prior applications is considered part of the disclosure of this application and is incorporated by reference into the disclosure of this application.
Background Art
[0002] This specification relates to using neural networks to generate audio.
[0003] A neural network is a machine learning model that uses one or more layers of non - linear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from the received input according to the current values of its respective set of parameters.
Summary of the Invention
[0004] This specification describes a system implemented as a computer program on one or more computers in one or more locations that uses one or more generative neural networks to generate an audio signal.
[0005] Generally, the output audio signal is an output audio example that includes samples of sound waves at each of a sequence of output time steps spanning a specified time window. For example, the output time steps can be arranged at regular intervals within the specified time window.
[0006] The audio samples at a given output time step can be the amplitude values of the sound wave, or amplitude values that have been compressed, companded, or both. For example, the audio samples can be the unprocessed amplitude values, or the mu-law companded representation of the amplitude values.
[0007] According to a first aspect, there is provided a method for generating a prediction of an audio signal, the method comprising receiving a request for generating an audio signal having respective audio samples at each of a plurality of output time steps spanning a time window, and obtaining a semantic representation of the audio signal specifying respective semantic tokens at each of a plurality of first time steps spanning the time window, each semantic token being selected from a vocabulary of semantic tokens and representing the semantic content of the audio signal at the corresponding first time step, using one or more generative neural networks to generate an acoustic representation of the audio signal conditioned at least on the semantic representation, the acoustic representation specifying a set of one or more respective acoustic tokens at each of a plurality of second time steps spanning the time window, each of the one or more respective acoustic tokens at each second time step representing the acoustic characteristics of the audio signal at the corresponding second time step, and using a decoder neural network to process at least the acoustic representation to generate a prediction of the audio signal.
[0008] In some embodiments, the decoder neural network is a decoder neural network of a neural audio codec that is co-trained with an encoder neural network for the purpose of measuring the reconstruction quality of the predicted audio signal generated by the decoder neural network from the acoustic representation generated using the output generated by the encoder neural network.
[0009] In some embodiments, the acoustic representation is a prediction of a ground truth acoustic representation that would be generated from the output of an encoder neural network by processing an audio signal.
[0010] In some embodiments, the encoder neural network outputs respective embeddings at each of a plurality of second time steps, and the ground truth acoustic representation is generated by applying quantization to each of the respective embeddings.
[0011] In some embodiments, the quantization is residual vector quantization that encodes each embedding using a hierarchy of a plurality of vector quantizers that each generate a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer, the hierarchy including one or more coarse vector quantizers at one or more first positions within the hierarchy and one or more fine quantizers at one or more last positions within the hierarchy, and the set of one or more respective acoustic tokens at each of the second time steps includes, for each vector quantizer, a respective acoustic token that is selected from the vocabulary of the vector quantizer and that is a prediction of a ground truth acoustic token that would be generated by the vector quantizer from the ground truth embedding generated by the encoder neural network at the second time step.
[0012] In some embodiments, the set of one or more respective acoustic tokens at each of a plurality of second time steps includes a plurality of acoustic tokens that collectively represent a prediction of the output of residual vector quantization applied to an embedding representing the acoustic characteristics of the audio signal at the second time step, where the residual vector quantization encodes the embedding using a hierarchy of a plurality of vector quantizers each generating a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer, the hierarchy including one or more coarse vector quantizers at one or more first positions within the hierarchy and one or more fine vector quantizers at one or more last positions within the hierarchy, and the set of acoustic tokens at each of the second time steps includes, for each vector quantizer, a respective acoustic token selected from the vocabulary of the vector quantizer.
[0013] In some embodiments, generating an acoustic representation of an audio signal using one or more generative neural networks, conditioned at least on a semantic representation, includes using a first generative neural network to generate, for each of one or more coarse vector quantizers within a hierarchy, a respective acoustic token at a second time step of the vector quantizer, conditioned at least on the semantic representation.
[0014] In some embodiments, the first generative neural network is an autoregressive neural network configured to autoregressively generate acoustic tokens according to a first generation order, and each particular acoustic token at each particular second time step for each particular coarse vector quantizer is conditioned on at least the semantic representation and any acoustic tokens preceding the particular acoustic token in the first generation order.
[0015] In some embodiments, for each particular acoustic token at each particular second time step for each particular coarse vector quantizer, prior to the particular acoustic token: (i) any acoustic token for any of the coarse vector quantizers at any second time step preceding the particular second time step, and (ii) any acoustic token at the particular second time step of any coarse vector quantizer preceding the particular vector quantizer within the hierarchy precede in a first generation order.
[0016] In some embodiments, the first generation neural network has a transformer architecture dedicated to the decoder, or an encoder-decoder transformer architecture.
[0017] In some embodiments, generating an acoustic representation of an audio signal using one or more generation neural networks, conditioned at least on a semantic representation, includes using a second generation neural network to generate, for each of one or more fine vector quantizers within the hierarchy, each acoustic token at the second time step of the vector quantizer, conditioned on each acoustic token at the second time step of one or more coarse vector quantizers within the hierarchy.
[0018] In some embodiments, the second generation neural network is not conditioned on a semantic representation.
[0019] In some embodiments, the second generation neural network is an autoregressive neural network configured to autoregressively generate acoustic tokens according to a second generation order, and for each particular acoustic token at each particular second time step for each particular fine vector quantizer, the particular acoustic token is conditioned on (i) each acoustic token for at least a subset of the second time steps of one or more coarse vector quantizers, and (ii) at least a subset of acoustic tokens preceding the particular acoustic token in the second generation order.
[0020] In some embodiments, prior to each particular acoustic token at each particular second time step for each particular fine vector quantizer, (i) any acoustic token for any of the fine vector quantizers at any second time step preceding the particular second time step, and (ii) any acoustic token at the particular second time step for any fine vector quantizer preceding the particular vector quantizer within the hierarchy precede in a second generation order.
[0021] In some embodiments, each particular acoustic token at each particular second time step for each particular fine vector quantizer is conditioned on (i) each respective acoustic token for one or more coarse vector quantizers that are at most a threshold number of second time steps before the second time step, and (ii) any acoustic token at a second time step that precedes the particular second time step in the second generation order and that is at most a threshold number of second time steps before the second time step.
[0022] In some embodiments, the second generation neural network has a transformer architecture dedicated to a decoder or an encoder-decoder transformer architecture.
[0023] In some embodiments, obtaining a semantic representation of an audio signal includes using a third generation neural network to autoregressively generate the semantic representation.
[0024] In some embodiments, the request specifies a context for the audio signal, and the audio signal is conditioned on the context.
[0025] In some embodiments, the context specifies semantic characteristics of the audio signal, and obtaining a semantic representation of the audio signal includes generating the semantic representation conditioned on the context.
[0026] In some embodiments, the context specifies acoustic characteristics of the audio signal, and using one or more generative neural networks to generate an acoustic representation of the audio signal, conditioned at least on the semantic representation, includes using one or more generative neural networks to generate an acoustic representation of the audio signal, conditioned on the semantic representation and the context.
[0027] In some embodiments, using a first generative neural network to generate, for each of one or more coarse vector quantizers within a hierarchy, each acoustic token of a second time step of the vector quantizer, conditioned at least on the semantic representation, includes using a first generative neural network to generate, for each of one or more coarse vector quantizers within a hierarchy, each acoustic token of a second time step of the vector quantizer, conditioned on the semantic representation and the context.
[0028] In some embodiments, using a decoder neural network to process at least the acoustic representation to generate a prediction of the audio signal includes using a decoder neural network to process the acoustic representation and the acoustic representation of the context to generate a prediction of the audio signal.
[0029] In some embodiments, the context includes the audio input.
[0030] In some embodiments, the context includes visual data.
[0031] In some embodiments, the context includes text data.
[0032] In some embodiments, the number of first time steps and the number of second time steps spanning a time window are less than the number of output time steps spanning the time window.
[0033] In some embodiments, the number of first time steps spanning the time window is less than the number of second time steps spanning the time window.
[0034] According to a second aspect, a method for generating a prediction of an audio signal is provided. The method includes receiving a request for generating an audio signal having respective audio samples at each of a plurality of output time steps spanning a time window, conditioned on an input; processing the input using an embedding neural network to map the input to one or more embedding tokens; generating a semantic representation of the audio signal that specifies respective semantic tokens at each of a plurality of first time steps spanning the time window, wherein each semantic token is selected from a vocabulary of semantic tokens conditioned on the embedding tokens and represents the semantic content of the audio signal at the corresponding first time step; using one or more generative neural networks to generate an acoustic representation of the audio signal, conditioned on at least the semantic representation and the embedding tokens, wherein the acoustic representation specifies a set of one or more respective acoustic tokens at each of a plurality of second time steps spanning the time window, and wherein the one or more respective acoustic tokens at each of the second time steps represent the acoustic characteristics of the audio signal at the corresponding second time step; and using a decoder neural network to process at least the acoustic representation to generate a prediction of the audio signal.
[0035] In some embodiments, processing the input using an embedding neural network to map the input to one or more embedding tokens includes using the embedding neural network to generate an embedding vector for the input in a joint embedding space and quantizing the embedding vector to generate the embedding tokens.
[0036] In some embodiments, the input includes a sequence of text, and the embedding neural network is trained to map text and audio into a joint embedding space.
[0037] In some embodiments, the input includes a sequence of text, and the prediction of the audio signal is a prediction of music described by the sequence of text.
[0038] In some embodiments, the input further includes an audio signal representing a melody, and the prediction of the audio signal is a prediction of music along the melody described by the sequence of text.
[0039] In some embodiments, the method further includes using a melody embedding neural network to process the audio signal to map the audio signal to one or more melody embedding tokens, and concatenating the melody embedding tokens with the embedding tokens.
[0040] In some embodiments, the input includes an audio signal representing a melody, and the prediction of the audio signal is a prediction of music along the melody.
[0041] In some embodiments, the method further includes using a melody embedding neural network to process the audio signal to map the audio signal to one or more melody embedding tokens, and each semantic token is selected from a vocabulary of semantic tokens conditioned on the melody embedding tokens.
[0042] In some embodiments, processing an audio signal using a melody-embedded neural network to map the audio signal to one or more melody-embedded tokens includes generating one or more melody-embedded vectors of the audio signal in a joint embedding space using the melody-embedded neural network and quantizing the one or more melody-embedded vectors to generate melody-embedded tokens.
[0043] In some embodiments, a sequence of text includes a plurality of subsequences of text, and predicting an audio signal is a prediction of music having a section of music corresponding to each of the subsequences and reflecting each of the subsequences.
[0044] In some embodiments, an embedded neural network is trained with training data including an audio signal.
[0045] In some embodiments, an embedded neural network is trained for the purpose such that text describing an audio signal and the corresponding audio signal have embeddings that are close to each other within a joint embedding space.
[0046] In some embodiments, a decoder neural network is a decoder neural network of a neural audio codec that is co-trained with an encoder neural network for the purpose of measuring the reconstruction quality of a predicted audio signal generated by the decoder neural network from an acoustic representation generated using the output generated by the encoder neural network.
[0047] In some embodiments, an acoustic representation is a prediction of a ground truth acoustic representation that would be generated from the output of an encoder neural network by processing an audio signal.
[0048] In some embodiments, the encoder neural network outputs respective embeddings at each of a plurality of second time steps, and the ground truth acoustic representation is generated by applying quantization to each of the respective embeddings.
[0049] In some embodiments, the quantization is residual vector quantization that encodes each embedding using a hierarchy of a plurality of vector quantizers each generating a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer, the hierarchy including one or more coarse vector quantizers at one or more first positions within the hierarchy and one or more fine quantizers at one or more last positions within the hierarchy, and the set of one or more respective acoustic tokens at each of the second time steps is selected, for each vector quantizer, from the vocabulary of the vector quantizer and includes each acoustic token that is a prediction of the ground truth acoustic token that would be generated by the vector quantizer from the ground truth embedding generated by the encoder neural network at the second time step.
[0050] In some embodiments, the set of one or more respective acoustic tokens at each of the plurality of second time steps includes a plurality of acoustic tokens that collectively represent a prediction of the output of residual vector quantization applied to an embedding representing acoustic characteristics of an audio signal at the second time step, the residual vector quantization encoding the embedding using a hierarchy of a plurality of vector quantizers each generating a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer, the hierarchy including one or more coarse vector quantizers at one or more first positions within the hierarchy and one or more fine vector quantizers at one or more last positions within the hierarchy, and the set of acoustic tokens at each of the second time steps includes each acoustic token selected from the vocabulary of the vector quantizer for each vector quantizer.
[0051] In some embodiments, generating an acoustic representation of an audio signal using one or more generative neural networks, conditioned on at least a semantic representation and an embedding token, includes using a first generative neural network to generate, for each of one or more coarse vector quantizers within a hierarchy, each acoustic token of a second time step of the vector quantizer, conditioned on at least the semantic representation and the embedding token.
[0052] In some embodiments, the first generative neural network is an autoregressive neural network configured to autoregressively generate acoustic tokens according to a first generation order, and each particular acoustic token at each particular second time step for each particular coarse vector quantizer is conditioned on at least the semantic representation, and the embedding token, and any acoustic tokens that precede the particular acoustic token in the first generation order.
[0053] In some embodiments, prior to each particular acoustic token at each particular second time step for each particular coarse vector quantizer, (i) any acoustic token for any of the coarse vector quantizers at any second time step that precedes the particular second time step, and (ii) any acoustic token at the particular second time step of any coarse vector quantizer that precedes the particular vector quantizer within the hierarchy precede in the generation order.
[0054] In some embodiments, the first generative neural network has a decoder-only transformer architecture, or an encoder-decoder transformer architecture.
[0055] In some embodiments, generating an acoustic representation of an audio signal using one or more generative neural networks, conditioned on at least a semantic representation and an embedding token, includes using a second generative neural network to generate each acoustic token of a second time step of a vector quantizer, conditioned on each acoustic token of a second time step of one or more coarse vector quantizers within the hierarchy, for each of one or more fine vector quantizers within the hierarchy.
[0056] In some embodiments, the second generative neural network is not conditioned on the semantic representation and the embedding token.
[0057] In some embodiments, the second generative neural network is an autoregressive neural network configured to autoregressively generate acoustic tokens according to a second generation order, where each particular acoustic token at each particular second time step for each particular fine vector quantizer is conditioned on (i) each acoustic token of at least a subset of the second time steps of one or more coarse vector quantizers, and (ii) at least a subset of acoustic tokens that precede the particular acoustic token in the second generation order.
[0058] In some embodiments, prior to each particular acoustic token at each particular second time step for each particular fine vector quantizer, (i) any acoustic token for any of the fine vector quantizers at any second time step preceding the particular second time step, and (ii) any acoustic token at the particular second time step for any fine vector quantizer preceding the particular vector quantizer within the hierarchy precede in the generation order.
[0059] In some embodiments, each particular acoustic token at each particular second time step for each particular fine vector quantizer is conditioned on (i) each acoustic token for one or more coarse vector quantizers that is at most the threshold number of the second time step before the second time step, and (ii) any acoustic token at a second time step that precedes the particular second time step in the second generation order and is at most the threshold number of the second time step before the second time step.
[0060] In some embodiments, the second generation neural network has a transformer architecture dedicated to the decoder or a transformer architecture of an encoder-decoder.
[0061] In some embodiments, generating a semantic representation of an audio signal includes autoregressively generating the semantic representation using a third generation neural network having an embedding token as a conditioning signal.
[0062] In some embodiments, the number of first time steps and the number of second time steps spanning the time window are less than the number of output time steps spanning the time window.
[0063] In some embodiments, the number of first time steps spanning the time window is less than the number of second time steps spanning the time window.
[0064] In some embodiments, the input includes a sequence of text, the embedding is an embedding in a joint embedding space, and one or more generation neural networks are trained at least in part with audio-specific training data, and during training, the semantic representation and the acoustic representation are conditioned on the embedding of the audio input in the joint embedding space.
[0065] According to a further aspect, a system is also described that includes one or more computers and one or more storage devices that store instructions that, when executed by the one or more computers, cause the one or more computers to perform the methods described herein.
[0066] According to a further aspect, one or more computer-readable storage media are also described that store instructions that, when executed by one or more computers, cause the one or more computers to perform the methods described herein.
[0067] Certain embodiments of the subject matter described herein can be implemented to realize one or more of the following advantages.
[0068] The systems described herein provide high-quality audio generation with a long-term consistent structure. For example, to generate predictions of audio signals, the system can obtain a semantic representation of the audio signal. The semantic representation can guarantee long-term consistency by representing features such as the linguistic content of speech, i.e., the melody and rhythm of music. The system can use one or more generative neural networks to generate an acoustic representation of the audio signal based on the semantic representation. The acoustic representation can guarantee high-quality audio synthesis by representing features such as acoustic details. The system can then use a decoder neural network to process the acoustic representation to generate predictions of the audio signal. In this way, the system can generate high-quality and consistent-structured audio using information from both the semantic and acoustic representations.
[0069] Conventional systems for generating audio cannot generate audio with consistent speech without conditioning or text annotation. The systems described herein can generate syntactically and semantically consistent speech even without text annotation. For example, the system can use semantic representations to capture local dependencies such as orthography, as well as long-term semantic information such as language content and temporal structure.
[0070] Furthermore, conventional systems for generating audio may generate audio with limited acoustic diversity or quality. For example, conventional systems may be trained only on clean speech and may generate audio with the voice of a single speaker. Other conventional systems may generate low-quality audio. The systems described herein can generate a sequence of audio that maintains the voice, intonation, and prosody of any unseen speaker. For example, the system can use acoustic representations to capture the identification of the speaker of an audio input for a given context and the recording conditions.
[0071] In some embodiments, the audio generated by the system is music. For example, the system can generate a sequence of audio inputs that includes music that is consistent with the input with respect to melody, harmony, tone, playing style, timbre, and rhythm. The system can use semantic representations to capture information such as harmony, rhythm, and melody. For example, the semantic representation can provide a melody and temporal structure consistent with the prompt. The system can then use the semantic representation to guide the generation of an acoustic representation.
[0072] The system can save computational resources during training. For example, a decoder neural network that generates predictions of audio signals from acoustic representations can be pre-trained and fixed earlier than training one or more generative neural networks. An audio representation neural network that can be used to obtain a target semantic representation, and a neural audio codec that can be used to obtain a target acoustic representation can also be pre-trained and fixed earlier than training one or more generative neural networks. Further, conventional systems are trained with clean audio speech samples, while the systems described herein exhibit strong performance when trained on more diverse and noisy samples. By increasing the robustness to the quality of the trained data, the computational resources and time required to create and clean up training data are reduced.
[0073] The system can also save computational resources during inference. For example, the system uses the semantic representation as a condition for generating the acoustic representation after obtaining the semantic representation. Thus, the number of tokens processed by the system is reduced compared to alternatives such as processing an interleaved sequence of semantic and acoustic tokens, enabling more efficient inference and training. Further, computational resources are saved compared to alternatives such as directly processing audio samples by autoregressively generating semantic and acoustic tokens and then mapping the acoustic tokens to the audio samples.
[0074] In some embodiments, the system can generate music given an input that includes text. The system can generate long, high-quality, and consistent music that conforms to a given high-quality text description with a higher level of complexity than conventional music generation systems. For example, because the system uses an embedding neural network having a joint embedding space for audio and text and the generative neural networks described herein, the system can more effectively generate text-conditioned music than conventional systems that generate music from text.
[0075] Furthermore, not all desired characteristics of music are easily described using text. For example, a melody can be an important characteristic of a musical work, but it is difficult to describe using text. Given an input audio clip that represents a melody, in some embodiments, the system can generate music along the melody. For example, the input audio clip can include a pitched melody. The system can use a melody embedding neural network to condition one or more generative neural networks to generate music along the melody. In some embodiments, the system can generate music that follows the input melody and reflects an input sequence of text. For example, the system can use an embedding neural network and a melody embedding neural network to generate embedding tokens, also referred to as audio embedding tokens.
[0076] The system can generate text-conditioned music without requiring a large training dataset of training data paired text-music. The system can be trained with a dataset of music only. Also, training the system with music only helps to make the system robust to noisy labels that can be a limitation when implementing conventional music generation systems.
[0077] In some embodiments where the system can generate music given an input that includes text, each component of the system can be trained separately, allowing for flexibility and efficiency during training. For example, the embedded neural network and the decoder neural network can be trained separately or simultaneously.
[0078] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description and drawings, and from the claims.
Brief Description of the Drawings
[0079]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Modes for Carrying Out the Invention
[0080] Like reference numerals and symbols in the various drawings refer to like elements.
[0081] FIG. 1 is a block diagram of an exemplary audio generation system 100. The audio generation system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations where the systems, components, and techniques described below are implemented.
[0082] The audio generation system 100 generates a prediction of an audio signal 104 given a request 102 for generating the audio signal. The audio signal 104 includes respective audio samples at each of a plurality of output time steps spanning a time window.
[0083] To generate the audio, the system 100 receives the request 102.
[0084] In some embodiments, the request 102 can specify the context of the audio signal 104. In these embodiments, the audio signal 104 is conditioned on the context.
[0085] For example, the context can include an audio input as the input audio signal. In some examples, the audio input can include words spoken by a particular speaker. In these examples, the audio signal 104 can be a continuation of the words spoken by the particular speaker. In some examples, the audio input can include music. In these examples, the audio signal 104 can be a continuation of the music within the audio input.
[0086] In some implementations, the input audio signal can include a melody, and the audio signal 104 can be music along the melody. The system 100 can generate the output audio signal 104 that is music along the melody, as will be described in more detail below with reference to FIGS. 5-8.
[0087] In some examples, the context can also include text data. The audio signal 104 can include speech that reflects the text data. In some embodiments, the audio signal 104 can include music that reflects the text data. In these embodiments, the system 100 generates an output audio signal 104 that reflects the text data, as will be described in more detail below with reference to FIGS. 5-8.
[0088] In some examples, the context can also include visual data. In these examples, the audio signal 104 can include speech that describes the visual data or music that reflects the visual data.
[0089] In some embodiments, the system 100 can also include an embedded neural network 120. In some examples where the context includes an input, the system 100 can process the input and map it to one or more embedding tokens 122, also referred to as audio embedding tokens.
[0090] For example, when the input includes a sequence of text, the embedded neural network 120 can be trained to map the text and audio to a joint embedding space of text and audio, also referred to as a joint audio embedding space. For example, the embedded neural network can include a neural network that maps text input to an embedding and a neural network that maps audio input to an embedding. In the joint embedding space of text and audio, both the text and the audio are mapped to embeddings within the same embedding space. That is, the embedding vectors of text and audio have the same dimensionality. Further, embeddings that are close to each other within the joint embedding space mean that the embeddings share semantics within and across modalities. For example, two embeddings that are close to each other can represent two text sequences that are semantically similar, two audio samples that are semantically similar, or an audio sample and a text sequence having semantically similar features. The embedded neural network 120 will be described in more detail below with reference to FIG. 8.
[0091] The system 100 can obtain a semantic representation 106 of the audio signal 104. The semantic representation 106 specifies each semantic token at each of a plurality of first time steps spanning a time window.
[0092] Each semantic token is selected from a vocabulary of semantic tokens and represents the semantic content of the audio signal 104 at the corresponding first time step. Examples of semantic content that a semantic token can represent include the linguistic content of speech, orthography, linguistic syntax, and prosodic features. Examples of semantic content can also include the genre, melody, harmony, and rhythm characteristics of music.
[0093] The system can generate a semantic representation 106 of the audio signal 104 in any of a variety of ways. Generating the semantic representation 106 will be described in more detail below with reference to FIGS. 3 and 7.
[0094] The number of first time steps, and the number of semantic tokens, may depend on the sampling rate of the audio representation neural network, i.e., the number of embeddings per unit time, which will be described in more detail below with reference to FIGS. 3 and 7.
[0095] The system 100 then uses one or more generative neural networks 108 to generate an acoustic representation 110 of the audio signal 104, conditioned at least on the semantic representation 106.
[0096] The acoustic representation 110 specifies a set of one or more respective acoustic tokens at each of a plurality of second time steps that span a time window. The one or more respective acoustic tokens at each of the second time steps represent the acoustic characteristics of the audio signal 104 at the corresponding second time step. The acoustic characteristics capture details of the audio waveform and enable high-quality synthesis. The acoustic characteristics can include, for example, speaker identification. The acoustic characteristics can also include recording conditions such as reverberation level, distortion, background noise, etc. Generating the acoustic representation 110 will be described in more detail below with reference to FIGS. 2 and 6.
[0097] The number of second time steps may depend on the sampling rate of the encoder neural network, which will be described in more detail below with reference to FIGS. 3 and 4. In some embodiments, the number of second time steps can be greater than the number of first time steps. For example, the number of second time steps can be twice the number of first time steps. Thus, for each semantic token at the first time step, the acoustic representation can include one or more acoustic tokens at each of two second time steps corresponding to the first time step.
[0098] System 100 then uses the decoder neural network 112 to process at least the acoustic representation 110 to generate a prediction of the audio signal 104. For example, each respective audio sample at each of a plurality of output time steps spanning a time window may be based on one or more acoustic tokens of the acoustic representation 110.
[0099] In some embodiments, the decoder neural network 112 can be the decoder neural network of a neural audio codec. The neural audio codec can be, for example, the SoundStream neural audio codec. The decoder neural network 112 is described in more detail below with reference to FIGS. 3 and 4.
[0100] Accordingly, system 100 can generate an audio signal 104 that meets the requirements from the acoustic tokens, thereby requiring fewer computational resources and less power than directly generating audio samples without tokens. For example, it is more computationally efficient to generate tokens that represent the features of audio samples autoregressively than to generate audio samples autoregressively. System 100 can generate tokens from the embedding at a sampling rate lower than the sampling rate of the audio signal. For example, since the audio signal can have a sampling rate of 16 kHz, the audio signal includes an audio sample every 0.0625 ms. The embedding for generating semantic tokens can be computed at a sampling rate of 25 Hz, so each embedding can represent the features of a 40 ms window of the audio signal. Further, the embedding for generating acoustic tokens can be computed at a sampling rate of 50 Hz, so each embedding can represent the features of a 20 ms window of the audio signal. After quantizing the embedding into acoustic tokens, system 100 can utilize the decoder neural network 112 to decode the acoustic tokens into an audio signal at a higher sampling rate.
[0101] System 100 can be configured to perform any of a variety of tasks that require generating an audio signal 104 as an output.
[0102] For example, the audio signal 104 may be a voice signal, and the system 100 can unconditionally generate a voice signal, for example, as a result, a voice signal extracted from a distribution represented by a training data set (s) on which a generative neural network (s) has been trained will be generated.
[0103] As another example, the audio signal 104 can be a different type of audio signal, such as music, animal voices, etc., and the system can unconditionally generate the audio signal 104, for example, as a result, an audio signal 104 extracted from a distribution represented by a training data set (s) on which a generative neural network (s) has been trained will be generated.
[0104] As another example, as described above, the system 100 can receive a context along with a request to generate the audio signal 104 and generate the audio signal 104 conditioned on the received context.
[0105] For example, the audio signal 104 can be a voice signal or another audio signal, that is, the system 100 generates an output audio signal 104 conditioned on a context that is an input audio signal.
[0106] As an example, the generated output audio signal 104 can be a prediction of an audio signal that follows the input audio signal. For example, the context can be an input voice signal that is a question asked by one speaker, and the output audio signal can be an output voice signal that is an answer to a question spoken by the same or another speaker. As another example, the context can be an input voice signal that is the first part of an utterance spoken by one speaker, and the output audio signal 104 can be an output voice signal that is the completion of the utterance spoken by the speaker or another speaker, or a response to the input utterance.
[0107] As an example, in some embodiments, the input audio signal can represent music, and the system 100 can generate an output audio signal 104 that is music that follows the input audio signal.
[0108] As another example, the system 100 can perform voice separation on the input audio signal to generate the output audio signal 104. For example, the input audio signal can include both voice and music, or other background noise, and the output audio signal 104 can represent only the voice. As another example, the input audio signal can include voices from multiple speakers (and optionally background noise), and the output audio signal 104 can include only the voice of one of the speakers. In some examples, the system 100 can perform audio-conditioned separation. That is, when the input audio signal can include additional audio input that is acoustically similar to one of the speakers, the output audio signal 104 can include only the voice of the speaker in the input audio signal that is acoustically similar to the additional audio input.
[0109] As another example, the system 100 can perform a conversion between voices where the input voice and the output voice represent the same semantic content but are spoken differently. For example, the input audio signal can include voice in one natural language, and the output audio signal can represent voice in another natural language that is a conversion of the input voice to the target language. As another example, the input audio signal can include voice spoken by a first speaker, and the output audio signal 104 can represent voice spoken by another speaker that represents the same semantic content as the input voice. As an example of this, the input audio signal can include voice spoken by a first speaker having a first accent of a natural language, and the output audio signal can represent voice spoken by another speaker having a different accent of the natural language that represents the same semantic content as the input voice. As another example of this, the input audio signal can include voice spoken by a first speaker having a speech disorder, and the output audio signal can be the same as the input voice but represent semantic content without a speech disorder. As another example, the input audio signal can include a first audio segment, and the output audio signal can include a second shorter audio segment that summarizes the semantic content of the first audio.
[0110] As another example, in some embodiments, the system 100 can perform melody-conditioned music generation where the input audio signal represents a melody, and the system 100 can generate an output audio signal 104 that is music along the melody.
[0111] As another example, the context can include both audio data and text data.
[0112] For example, the system 100 can perform transcript-conditioned voice enhancement where the context is a text transcript and noisy audio corresponding to the text transcript, and the output audio signal 104 is clean audio corresponding to the text transcript. For example, the clean audio can include less background noise than the noisy audio.
[0113] As another example, the system 100 can perform transcript-based audio infilling where the context is a text transcript and audio corresponding to a portion of the text transcript, and the output audio signal 104 corresponds to another portion of the text transcript. The system 100 can be trained, for example, to autoregressively predict semantic tokens and acoustic tokens using teacher forcing.
[0114] As another example, the system 100 can perform speaker-conditioned voice conversion where the context is a text transcript and the speaker's audio, and the output audio signal 104 is the verbalization of the text transcript spoken by the speaker.
[0115] As another example, in some embodiments, the system 100 can perform text and melody-conditioned music generation. The context can include a sequence of text and an input audio signal representing a melody. The system 100 can generate an output audio signal 104 that is music along the melody as described by the sequence of text.
[0116] As another example, the context can include both audio data and visual, such as image or video data.
[0117] For example, the system 100 can perform audio-video continuity, where the system receives a partial audio track along with the corresponding video, and the output audio signal 104 is the continuity of the partial audio track.
[0118] As another example, system 100 can perform cross-modal infill where the system receives a video and an audio track corresponding to a portion of the video, and output audio signal 104 is an audio track corresponding to another portion of the video.
[0119] As another example, the context input can include only visual data.
[0120] For example, system 100 can perform image-conditioned audio generation where the system receives an input image and generates an output audio signal 104 that describes the image.
[0121] As another example, the context input can include only text data.
[0122] For example, system 100 can perform text-description-based sound synthesis where, for example, the input is text that describes an audio signal and the output is an audio signal 104 characterized by the text.
[0123] As an example, system 100 can perform text-conditioned music generation where the input can include a sequence of text and system 100 can generate an output audio signal 104 that is music describable by the sequence of text. Generating music given text that describes the music will be described in more detail below with reference to FIGS. 5-8.
[0124] As another example, in some embodiments, system 100 can perform story mode generation. The input can include a plurality of subsequences of text. The system can generate an output audio signal 104 that includes a section of music corresponding to each of the subsequences. Each section of music can be described by the corresponding subsequence of text. Further, system 100 can generate a smooth transition between each section of music that has a consistent tempo and is semantically appropriate.
[0125] As another example, the system 100 can perform text-conditioned voice sampling where, for example, the input is a text transcript and the output audio signal 104 is a verbalization of the text transcript spoken by a random speaker with a random prosody.
[0126] In some embodiments where the context input includes non-audio data, the system 100 can be trained with training data that includes only audio, as will be described in more detail below with reference to FIG. 8.
[0127] FIG. 2 is a diagram of an exemplary process for generating acoustic tokens. For convenience, the process 200 is described as being performed by a system of one or more computers located in one or more locations. For example, an audio generation system, such as the audio generation system 100 of FIG. 1, appropriately programmed according to this specification, can execute the process 200.
[0128] The system can use one or more generative neural networks, such as a coarse generative neural network 210 and a fine generative neural network 220, to generate acoustic tokens of an acoustic representation. The coarse generative neural network 210 is also referred to as the first generative neural network. The fine generative neural network 220 is also referred to as the second generative neural network. The system can condition the generation of the acoustic representation based at least on the semantic representation 106.
[0129] An encoder neural network can be part of a neural audio codec that includes a decoder neural network and an encoder neural network. As will be described in more detail below with reference to FIG. 4, the decoder neural network and the encoder neural network can be jointly trained. The coarse generation neural network 210 and the fine generation neural network 220 can be trained to predict an acoustic representation generated based on the output of the vector quantizer of the neural audio codec.
[0130] In some embodiments, the set of one or more respective acoustic tokens at each of a plurality of second time steps includes a plurality of acoustic tokens that collectively represent a prediction of the output of residual vector quantization applied to an embedding that represents the acoustic characteristics of the audio signal at the second time step. Residual vector quantization encodes the embedding using a hierarchy of a plurality of vector quantizers, each of which generates a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer. The hierarchy includes one or more coarse vector quantizers at one or more first positions within the hierarchy and one or more fine vector quantizers at one or more last positions within the hierarchy. The set of acoustic tokens at each of the second time steps includes, for each vector quantizer, a respective acoustic token selected from the vocabulary of the vector quantizer.
[0131] For example, the hierarchy can include Q vector quantizers, where vector quantizers 1... Q' can be coarse vector quantizers and vector quantizers (Q'+1)... Q can be fine vector quantizers. The coarse vector quantizers generate coarse acoustic tokens or acoustic tokens of the coarse vector quantizer 212 that represent acoustic characteristics such as speaker identification and recording conditions. The fine vector quantizers generate fine acoustic tokens or acoustic tokens of the fine vector quantizer 222 that represent fine acoustic details. For example, the fine acoustic tokens can be used to remove irreversible compression artifacts within the coarse acoustic tokens.
[0132] In some embodiments, the number of first time steps spanning a time window and the number of second time steps are less than the number of output time steps spanning the time window. That is, the sampling rate of the encoder neural network that outputs respective embeddings at each of the second time steps can be lower than the sampling rate of the audio signal. Further, the sampling rate of the audio representation neural network described below with reference to FIG. 3 can be lower than the sampling rate of the audio signal. Thus, the number of embeddings capable of generating semantic tokens and the number of embeddings capable of generating acoustic tokens can be less than the number of output time steps.
[0133] In some embodiments, the number of first time steps spanning a time window is less than the number of second time steps spanning the time window. For example, for all semantic tokens, the system can generate two acoustic tokens for each coarse vector quantizer and two acoustic tokens for each fine vector quantizer. That is, the sampling rate of the encoder neural network that outputs respective embeddings at each of the second time steps can be higher than the sampling rate of the audio representation neural network described below with reference to FIG. 3.
[0134] To generate an acoustic representation, the coarse generation neural network 210 can generate acoustic tokens of the coarse vector quantizer 212 conditioned on at least the semantic representation 106. For example, the coarse generation neural network 210 can generate respective acoustic tokens of the second time step of the vector quantizer conditioned on at least the semantic representation for each of one or more coarse vector quantizers within the hierarchy. The acoustic tokens of the coarse vector quantizer 212 represent acoustic characteristics such as speaker identification and recording conditions.
[0135] In some embodiments where the request specifies a context, such as when the context includes an audio input, and the context specifies acoustic characteristics of an audio signal, the system can generate an acoustic representation of the audio signal conditioned on the semantic representation and the context. For example, the coarse generation neural network 210 can generate each acoustic token of a second time step of the vector quantizer conditioned on the semantic representation 106 and an acoustic token representing the context for each of one or more coarse vector quantizers within the hierarchy.
[0136] The coarse generation neural network 210 can be an autoregressive neural network configured to autoregressively generate acoustic tokens of the coarse vector quantizer 212 according to a first generation order.
[0137] Each particular acoustic token at each particular second time step for each particular coarse vector quantizer is conditioned on at least the semantic representation and any acoustic tokens that precede the particular acoustic token in the first generation order. Further, prior to each particular acoustic token at each particular second time step for each particular coarse vector quantizer, (i) any acoustic tokens for any of the coarse vector quantizers at any second time step that precedes the particular second time step, and (ii) any acoustic tokens at the particular second time step of any coarse vector quantizer that precedes the particular vector quantizer within the hierarchy precede in the generation order.
[0138] For example, the hierarchy may include a Q-vector quantizer, where vector quantizers 1...Q' can be coarse vector quantizers, while vector quantizers (Q'+1)...Q can be fine vector quantizers. The acoustic tokens at a particular second time step t and a particular coarse vector quantizer q ≤ Q' are conditioned on all of the semantic tokens of the semantic representation, any acoustic tokens at a second time step preceding the second time step t, the coarse vector quantizers preceding Q', and any acoustic tokens at the second time step t in the coarse vector quantizers preceding the particular coarse vector quantizer q.
[0139] In some embodiments, the coarse generation neural network 210 has a transformer architecture dedicated to the decoder. In some embodiments, the coarse generation neural network 210 has an encoder-decoder transformer architecture.
[0140] To generate an acoustic representation, the fine generation neural network 220 can generate the acoustic tokens of the fine vector quantizer 222 conditioned on at least the acoustic tokens of the coarse vector quantizer 212. For example, the fine generation neural network 220 can generate the acoustic tokens of each of one or more fine vector quantizers in the hierarchy conditioned on the respective acoustic tokens of each of the second time steps of one or more coarse vector quantizers in the hierarchy. Thus, the fine generation neural network 220 cannot be conditioned on the semantic representation 106. The acoustic tokens of the fine vector quantizer 222 can be used to further improve the audio quality, for example, by removing non-reversible compression artifacts.
[0141] The fine-grained generation neural network 220 can be an autoregressive neural network configured to autoregressively generate acoustic tokens according to a second generation order. Each specific acoustic token at each specific second time step for each specific fine-grained vector quantizer is conditioned on (i) each respective acoustic token for at least a subset of the second time steps of one or more coarse-grained vector quantizers, and (ii) at least a subset of the acoustic tokens preceding the specific acoustic token in the second generation order. Prior to each specific acoustic token at each specific second time step for each specific fine-grained vector quantizer, (i) any acoustic token for any of the fine-grained vector quantizers at any second time step preceding the specific second time step, and (ii) any acoustic token at the specific second time step for any fine-grained vector quantizer preceding the specific vector quantizer within the hierarchy precede in the second generation order.
[0142] For example, the hierarchy can include Q vector quantizers, and the vector quantizers (Q’ + 1)...Q can be fine-grained vector quantizers. The acoustic token at a specific second time step t and a specific fine-grained vector quantizer q > Q’ is conditioned on all of the acoustic tokens of the coarse-grained vector quantizer 312, any acoustic tokens at the second time steps preceding the second time step t and fine-grained vector quantizers following Q’, and any acoustic tokens at the second time step t in the coarse-grained vector quantizers and fine-grained vector quantizers preceding the specific vector quantizer q.
[0143] In some implementations, each particular acoustic token at each particular second time step for each particular fine vector quantizer is conditioned on (i) each respective acoustic token for one or more coarse vector quantizers that are at most the threshold number of the second time step before the second time step, and (ii) any acoustic token at a second time step that precedes the particular second time step in the second generation order and is at most the threshold number of the second time step before the second time step. That is, the second time steps can be divided into non-overlapping batches of consecutive second time steps.
[0144] For example, the acoustic token at a particular second time step t and a particular fine vector quantizer q > Q’ is conditioned on the acoustic token of the coarse vector quantizer 312 that precedes the particular second time step t within the corresponding batch of the same second time step, any acoustic token at a second time step that precedes the second time step t within the corresponding batch and for fine vector quantizers that follow Q’, and any acoustic token at the second time step t in the coarse vector quantizer and fine vector quantizers that precede the particular fine vector quantizer q.
[0145] In some embodiments, the fine generation neural network 220 has a transformer architecture dedicated to the decoder. In some embodiments, the fine generation neural network 220 has an encoder-decoder transformer architecture.
[0146] In some embodiments, the system can simultaneously generate acoustic tokens of the coarse vector quantizer 212 and acoustic tokens of the fine vector quantizer 222. For example, in the case of a hierarchical Q vector quantizer, the acoustic tokens at a particular second time step t of a particular coarse or fine vector quantizer q are conditioned on all of the acoustic tokens at a second time step preceding the second time step t and in a vector quantizer preceding q, as well as any acoustic tokens at the second time step t in a vector quantizer preceding the particular vector quantizer q.
[0147] The system can include acoustic tokens of the fine vector quantizer 222 within the acoustic representation. The system can use the acoustic representation to generate a prediction of the audio signal, as described below with reference to FIG. 3.
[0148] FIG. 3 is a flowchart of an exemplary process for generating a prediction of an audio signal. For convenience, process 300 is described as being performed by a system of one or more computers located in one or more locations. For example, an audio generation system, such as the audio generation system 100 of FIG. 1, appropriately programmed according to this specification, can execute process 300.
[0149] The system receives a request to generate an audio signal (step 310). The audio signal has respective audio samples at each of a plurality of output time steps spanning a time window. In some examples, the request specifies the context of the audio signal. The audio signal can be conditioned on the context. For example, the context can include audio input, visual data, and / or text data, as described above.
[0150] The system can obtain a semantic representation of the audio signal (step 320). The semantic representation specifies each semantic token at each of a plurality of first time steps spanning a time window. Each semantic token can be selected from a vocabulary of semantic tokens and can represent the semantic content of the audio signal at the corresponding first time step.
[0151] In some embodiments, the system can generate a semantic representation. For example, the system can autoregressively generate a semantic representation using a semantic representation generation neural network. The semantic representation generation neural network may also be referred to as a third generation neural network.
[0152] For example, the semantic representation generation neural network can be trained to sequentially and autoregressively generate semantic tokens. The semantic representation generation neural network can be, for example, a decoder only, or an encoder-decoder transformer-based neural network. For example, the semantic representation generation neural network may be trained to predict a semantic representation generated based on the output of one or more layers, such as one of the intermediate layers of an audio representation neural network.
[0153] The audio representation neural network can generate outputs or embeddings at regular time intervals. For example, the audio representation neural network can generate one embedding every 40 ms of the input audio signal.
[0154] The training of the semantic representation generation neural network and the audio representation neural network will be described in more detail below with reference to FIG. 4.
[0155] In some embodiments, the claim specifies a context that specifies the semantic characteristics of the audio signal. For example, the context can include an audio input. The system can generate a semantic representation conditioned on the context. For example, when the generation is conditioned on a context that includes an audio input, the system can use an audio representation neural network to generate a semantic representation of the audio input as described above.
[0156] In some examples, the audio signal spans the same time window as the audio input. The system can use the semantic representation of the audio input as the semantic representation of the audio signal. For example, the semantic representation of the audio signal can include the semantic tokens of the semantic representation of the audio input.
[0157] In some examples, the audio signal spans a window of time longer than the audio input. For example, the audio signal can be a continuation of the audio input. For example, the audio signal can be a continuation of speech or a continuation of music. The system can condition a semantic representation generation neural network based at least on the semantic representation of the audio input while autoregressively generating the semantic representation of the audio signal. For example, the semantic representation generation neural network can autoregressively generate the semantic tokens of the audio signal conditioned on the semantic tokens of the audio input. The semantic representation of the audio signal can include the semantic tokens of the semantic representation of the audio input followed by the autoregressively generated semantic tokens.
[0158] In some embodiments, the claim specifies a context that includes text or image data. When generation is conditioned on a context input that is at least partially non-audio, the system can map the context input from the vocabulary to semantic tokens using an appropriate encoder neural network for the non-audio part of the audio input, and use the semantic tokens as (at least part of) the semantic representation of the audio signal to generate semantic tokens, or condition a semantic representation generation neural network based at least on the semantic tokens while autoregressively generating the semantic representation of the audio signal.
[0159] For example, when the context includes image data, the system can map the image data from the vocabulary to semantic tokens using an encoder neural network configured to map the image data to audio. In some examples, the system can use the semantic tokens as the semantic representation of the audio signal. In some examples, the system can condition a semantic representation generation neural network based at least on the semantic tokens while autoregressively generating the semantic representation of the audio signal.
[0160] For example, when the context includes text data, the system can map the text data from vocabulary to semantic tokens using an encoder neural network configured to map the text data to audio. In some embodiments, the encoder neural network can be an embedding neural network that will be described in more detail with reference to FIGS. 1 and 8. For example, the text data can include a sequence of text describing music, and the embedding neural network can map the description of the music to audio. The system can use the semantic tokens corresponding to the audio as at least part of the semantic representation of the audio signal. In some examples, the system can condition a semantic representation generation neural network based at least on the semantic tokens while autoregressively generating the semantic representation of the audio signal.
[0161] The system generates an acoustic representation of the audio signal (step 330). The system can use one or more generation neural networks to generate the acoustic representation, conditioned at least on the semantic representation. The acoustic representation specifies a set of one or more respective acoustic tokens at each of a plurality of second time steps spanning a time window. The one or more respective acoustic tokens at each of the second time steps can represent the acoustic characteristics of the audio signal at the corresponding second time step. Generating the acoustic representation of the audio signal is described in more detail above with reference to FIG. 2.
[0162] In some embodiments, when the context specifies the target speaker of the output audio signal, when the output audio signal includes subsequent speech, or when the output audio signal includes audio that should be similar in another way, the context specifies the acoustic characteristics of the generated audio signal. In these embodiments, the system can map the context to an acoustic representation, for example using the neural audio codec described above, and can use at least a portion of the acoustic representation of the context when generating the acoustic representation of the audio signal. Generating an acoustic representation conditioned on the context is described in more detail above with reference to FIG. 2.
[0163] The system processes at least the acoustic representation to generate a prediction of the audio signal (step 340). The system can process at least the acoustic representation using a decoder neural network to generate a prediction of the audio signal. For example, each audio sample at each of a plurality of output time steps spanning a time window can be based on one or more acoustic tokens of all the vector quantizers within the hierarchy.
[0164] In some embodiments where the context specifies the acoustic characteristics of the audio signal, the system can process the acoustic representation and the acoustic representation of the context using a decoder neural network to generate a prediction of the audio signal. For example, the decoder neural network can process the acoustic tokens of the acoustic representation and the acoustic tokens representing the context.
[0165] In some embodiments, the number of first time steps and the number of second time steps spanning a time window are less than the number of output time steps spanning the time window. In some embodiments, as described in more detail above with reference to FIG. 2, the number of first time steps spanning a time window is less than the number of second time steps spanning the time window.
[0166] FIG. 4 is a diagram of an exemplary process 400 for training an exemplary audio generation system. For convenience, process 400 is described as being executed by a training system of one or more computers located in one or more locations.
[0167] The training system can train an audio generation system, such as audio generation system 100 of FIG. 1, with training data. In some embodiments, components of the audio generation system, such as semantic representation generation neural networks 130, one or more generation neural networks 108, neural audio codec 420, and audio representation neural network 410, can be trained separately.
[0168] The training system can train neural audio codec 420 and audio representation neural network 410 with training data that includes audio clips. An exemplary audio clip of the training data is shown as target audio 450 in FIG. 4.
[0169] Neural audio codec 420 can include a decoder neural network and an encoder neural network. For example, the encoder neural network can convert target audio 450 into a coded signal quantized into an acoustic representation. The decoder neural network can convert the acoustic representation into a predicted audio signal. In some embodiments, neural audio codec 420 can be pre-trained and fixed prior to training of the audio generation system.
[0170] The neural audio codec 420 can be trained to minimize adversarial loss and reconstruction loss. For example, the decoder neural network and the encoder neural network can be jointly trained for the purpose of measuring the reconstruction quality of the predicted audio signal generated by the decoder neural network from the acoustic representation generated using the output generated by the encoder neural network.
[0171] The audio representation neural network 410 may be trained to generate a representation of the input audio signal. In some embodiments, the audio representation neural network 410 can be pre-trained and fixed prior to training of the audio generation system. The audio representation neural network 410 can be trained to minimize mask language model (MLM) loss and contrastive loss. The audio representation neural network 410 can be, for example, a w2v-BERT model that maps the input audio signal to a set of linguistic features.
[0172] The training system can train the semantic representation generation neural networks 130 and one or more generation neural networks 108 with training data including audio clips. In some embodiments, the generation neural network can be trained using teacher forcing.
[0173] The training system can train the semantic representation generation neural network 130 to predict a semantic representation generated based on the output of one or more layers, such as one of the intermediate layers of the audio representation neural network 410.
[0174] For example, the audio representation neural network 410 can include a model having multiple layers. For example, the audio representation neural network 410 can include a self-attention-based model, such as a transformer-based model or a conformer-based model having multiple layers.
[0175] Self-attention based models may be trained on speech representation tasks, for example, through self-supervised learning. Self-attention based models may be trained on other tasks such as automatic speech recognition.
[0176] The output of one or more layers of the audio representation neural network 410 may include an embedding of the input audio. For example, a self-attention based model can generate high-density embeddings for some or all of the audio samples in the training data. The embeddings of the intermediate layers of the self-attention based model for some or all of the audio samples can be clustered into K clusters by the k-means method. The centroid of the cluster can be used as a semantic token. In some embodiments, the output can be normalized so that each dimension has zero mean and unit variance before clustering.
[0177] Exemplary target semantic representations for training the semantic representation generation neural network 130 may be generated by providing the target audio 450 to a self-attention based model. The training system can generate semantic tokens of the target semantic representation by assigning each output of the target audio 450 to the centroid of the nearest cluster.
[0178] In some embodiments, consecutive repetitions of semantic tokens within the target semantic representation can be removed. For example, the target semantic representation of the target audio 450 includes a sequence of semantic tokens. The training system can remove the consecutively repeated semantic tokens in the sequence and then use the target semantic representation without consecutively repeated semantic tokens for training.
[0179] The training system can train one or more generative neural networks 108 to generate an acoustic representation that is conditioned on at least a semantic representation. In some embodiments, the acoustic representation can be a prediction of a ground truth acoustic representation that would be generated from the output of an encoder neural network by processing an audio signal. For example, the encoder neural network can be a convolutional encoder that maps the audio signal to a sequence of embeddings. The encoder neural network can output a respective embedding at each of a plurality of second time steps. Each respective embedding at each of the plurality of second time steps can correspond to a feature of the audio signal at the second time step. The ground truth acoustic representation can be generated by applying quantization to each of the respective embeddings. As described above, the encoder neural network can be part of a neural audio codec, such as neural audio codec 420.
[0180] For example, the quantization can be residual vector quantization that encodes each embedding using a hierarchy of a plurality of vector quantizers, each of which generates a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer. The hierarchy can include Q-vector quantizers that each use a corresponding vocabulary. The hierarchy can include one or more coarse vector quantizers at one or more first positions within the hierarchy and one or more fine vector quantizers at one or more last positions within the hierarchy. The set of one or more respective acoustic tokens at each of the second time steps can include, for each vector quantizer, a respective acoustic token selected from the vocabulary of the vector quantizer, which is a prediction of a ground truth acoustic token that would be generated by the vector quantizer from the ground truth embedding generated by the encoder neural network at the second time step.
[0181] The training system can train one or more generative neural networks 108 to predict an acoustic representation generated based on the output of the residual vector quantizer of the neural audio codec 420.
[0182] The target acoustic representation for training one or more generative neural networks 108 may be generated by applying quantization to each of the respective embeddings output by the encoder neural network of the neural audio codec 420 at each of a plurality of second time steps. The quantization can be residual vector quantization (RVQ), as described above with reference to FIG. 2.
[0183] FIG. 5 is a diagram of an exemplary process for generating a prediction of an audio signal. For convenience, process 500 is described as being executed by a system of one or more computers located in one or more locations. For example, an audio generation system, such as the audio generation system 100 of FIG. 1, appropriately programmed according to this specification, can execute process 500.
[0184] The system receives a request to generate an audio signal conditioned on input 502. For example, input 502 can be included in the context of the audio signal described above with reference to FIGS. 1-3. In exemplary process 500, the input includes text data. The system can perform text-conditioned music generation such that the generated audio 504 is music described by the text data "hip-hop song including a violin solo". The text data can include descriptions of a comparable level of complexity, such as "impressive saxophone solo and enchanting jazz song by a solo singer" or "90s Berlin techno with bass and a powerful kick".
[0185] The system processes text data using the embedded neural network 120 described above with reference to FIG. 1. The embedded neural network 120 maps text to one or more embedded tokens 122. The embedded neural network 120 can be trained to map text and audio to a joint embedding space of text and audio, also referred to as a joint audio embedding space. The system can use the embedded neural network 120 to generate an embedding vector for an input of text data within the joint embedding space. The system can then quantize the embedding vector to generate the embedded tokens 122.
[0186] In some embodiments, the embedded neural network 120 can include a text network and a music network. For example, the embedded neural network 120 can include a neural network trained to map text to an embedding within the joint embedding space and a neural network trained to map audio to an embedding within the joint embedding space. For example, the embedded neural network 120 can be a joint embedding model having two embedding towers, one for text and one for music. The towers use contrastive learning to map text and music to a shared embedding space. The training of the embedded neural network 120 is described below with reference to FIG. 8.
[0187] The system generates a semantic representation 106 of the audio signal 504. Each semantic token of the semantic representation 106 is selected from a vocabulary of semantic tokens conditioned on the embedded tokens 122.
[0188] For example, the system can autoregressively generate the semantic representation 106 using a semantic representation generation neural network that has the embedded token 122 as a conditioning signal. For example, the semantic representation generation neural network can condition on the embedded token 122 while autoregressively generating the semantic representation 106. That is, each semantic token at the first time step can condition on the semantic tokens leading up to the first time step and the embedded token 122. The semantic representation generation neural network is S t |S <t , M T can be generated, where S t represents the semantic token at the first time step t, and M T represents the embedded token 122.
[0189] The system then generates an acoustic representation 110 of the audio signal 504 conditioned on at least the semantic representation 106 and the embedded token 122. The system can use one or more generation neural networks to generate an acoustic representation conditioned on at least the semantic representation 106 and the embedded token 122. For example, one or more generation neural networks can autoregressively generate the acoustic representation 110 conditioned on the semantic representation 106 and the embedded token 122. That is, each set of one or more acoustic tokens at the second time step can condition on the acoustic tokens at the preceding second time step, the semantic representation 106, and the embedded token 122. One or more generation neural networks are A t |A <t , S, M T can be generated, where A t represents the acoustic token at the first time step t, S represents the semantic representation, and M T represents the embedded token 122 generated from the text.
[0190] The system then uses the decoder neural network 112 to process at least the acoustic representation 110 to generate a prediction of the audio signal 504.
[0191] Accordingly, the generated audio 504 includes music that can be described by the text data of the input 502, "a hip-hop song including a violin solo". Although the input 502 in this example includes text data, as will be described below with reference to FIG. 8, the system can be trained with an audio signal.
[0192] FIG. 6 is a diagram of an exemplary process 600 for generating acoustic tokens. For convenience, process 600 is described as being performed by a system of one or more computers located in one or more locations. For example, an audio generation system appropriately programmed according to this specification, such as the audio generation system 100 of FIG. 1, can execute process 600.
[0193] Process 600 is similar to process 200 described above with reference to FIG. 2 in that the system can use one or more generative neural networks, such as a coarse generative neural network 610 and a fine generative neural network 620, to generate acoustic tokens of the acoustic representation. The coarse generative neural network 610, also referred to as the first generative neural network, is similar to the coarse generative neural network 210 described with reference to FIG. 2. The fine generative neural network 620, also referred to as the second generative neural network, is similar to the fine generative neural network 620 described with reference to FIG. 2. The system can condition the generation of the acoustic representation based on at least the semantic representation 106 and the embedding token 122, also referred to as the audio embedding token.
[0194] An encoder neural network can be part of a neural audio codec that includes a decoder neural network and an encoder neural network. As will be described in more detail below with reference to FIG. 8, the decoder neural network and the encoder neural network can be jointly trained. The coarse generation neural network 610 and the fine generation neural network 620 can be trained to predict an acoustic representation generated based on the output of the vector quantizer of the neural audio codec.
[0195] In some embodiments, the set of one or more respective acoustic tokens at each of a plurality of second time steps includes a plurality of acoustic tokens that collectively represent a prediction of the output of residual vector quantization applied to an embedding that represents the acoustic characteristics of the audio signal at the second time step. Residual vector quantization encodes the embedding using a hierarchy of a plurality of vector quantizers, each generating a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer. The hierarchy includes one or more coarse vector quantizers at one or more first positions within the hierarchy and one or more fine vector quantizers at one or more last positions within the hierarchy. The set of acoustic tokens at each of the second time steps includes a respective acoustic token selected from the vocabulary of the vector quantizer for each vector quantizer.
[0196] For example, the hierarchy can include Q vector quantizers, where vector quantizers 1... Q' can be coarse vector quantizers and vector quantizers (Q'+1)... Q can be fine vector quantizers. The coarse vector quantizers generate coarse acoustic tokens that represent acoustic characteristics, such as speaker identification and recording conditions, or acoustic tokens of the coarse vector quantizer 612. The fine vector quantizers generate fine acoustic tokens that represent fine acoustic details, or acoustic tokens of the fine vector quantizer 622. For example, the fine acoustic tokens can be used to remove irreversible compression artifacts within the coarse acoustic tokens.
[0197] In some embodiments, the number of first time steps and the number of second time steps spanning the time window are less than the number of output time steps spanning the time window. That is, the sampling rate of the encoder neural network that outputs respective embeddings at each of the second time steps can be lower than the sampling rate of the audio signal. Further, the sampling rate of the audio representation neural network described with reference to FIG. 7 can be lower than the sampling rate of the audio signal. Thus, the number of embeddings capable of generating semantic tokens and the number of embeddings capable of generating acoustic tokens can be less than the number of output time steps.
[0198] In some embodiments, the number of first time steps spanning the time window is less than the number of second time steps spanning the time window. For example, for all semantic tokens, the system can generate two acoustic tokens per coarse vector quantizer and two acoustic tokens per fine vector quantizer. That is, the sampling rate of the encoder neural network that outputs respective embeddings at each of the second time steps can be higher than the sampling rate of the audio representation neural network described with reference to FIG. 7.
[0199] To generate an acoustic representation, the coarse generation neural network 610 can generate acoustic tokens of the coarse vector quantizer 612 conditioned on at least the semantic representation 106 and the embedding token 122. In some examples, the embedding token 122 can include a melody embedding token, also referred to as a melody audio embedding token. In some examples, the embedding token 122 can be concatenated with the melody embedding token.
[0200] For example, the coarse generation neural network 610 can generate each acoustic token at the second time step of the vector quantizer for each of one or more coarse vector quantizers within the layer, conditioned on at least the semantic representation 106 and the embedding token 122. The acoustic tokens of the coarse vector quantizer 612 represent acoustic characteristics such as recording conditions.
[0201] The coarse generation neural network 610 can be an autoregressive neural network configured to autoregressively generate the acoustic tokens of the coarse vector quantizer 612 according to a first generation order.
[0202] Each specific acoustic token at each specific second time step for each specific coarse vector quantizer is conditioned on at least the semantic representation 106, and the embedding token 122, and any acoustic tokens preceding the specific acoustic token in the first generation order. Further, prior to each specific acoustic token at each specific second time step for each specific coarse vector quantizer, (i) any acoustic token for any of the coarse vector quantizers at any second time step preceding the specific second time step, and (ii) any acoustic token at the specific second time step of any coarse vector quantizer preceding the specific vector quantizer within the layer precede in the generation order.
[0203] For example, the layer can include Q vector quantizers, where vector quantizers 1...Q' can be coarse vector quantizers, while vector quantizers (Q'+1)...Q can be fine vector quantizers. The acoustic tokens at a specific second time step t and a specific coarse vector quantizer q ≤ Q' are conditioned on all of the semantic tokens of the semantic representation 106, all of the embedding tokens 122, any acoustic tokens at second time steps preceding the second time step t, and the coarse vector quantizers preceding Q', and any acoustic tokens at the second time step t in the coarse vector quantizers preceding the specific coarse vector quantizer q.
[0204] In some embodiments, the coarse generation neural network 610 has a transformer architecture dedicated to the decoder. In some embodiments, the coarse generation neural network 610 has an encoder-decoder transformer architecture.
[0205] To generate an acoustic representation, the fine generation neural network 620 can generate the acoustic tokens of the fine vector quantizer 622 conditioned on at least the acoustic tokens of the coarse vector quantizer 612. For example, the fine generation neural network 620 can generate the acoustic tokens of each of one or more fine vector quantizers in the hierarchy conditioned on the respective acoustic tokens of the second time step of one or more coarse vector quantizers in the hierarchy. Thus, the fine generation neural network 620 cannot be conditioned on the semantic representation 106 and the embedding tokens 122. The acoustic tokens of the fine vector quantizer 622 can be used to further improve the audio quality, for example, by removing non-reversible compression artifacts.
[0206] The fine-grained generation neural network 620 can be an autoregressive neural network configured to autoregressively generate acoustic tokens according to a second generation order. Each particular acoustic token at each particular second time step for each particular fine-grained vector quantizer is conditioned on (i) each respective acoustic token for at least a subset of the second time steps of one or more coarse vector quantizers, and (ii) at least a subset of the acoustic tokens preceding the particular acoustic token in the second generation order. Prior to each particular acoustic token at each particular second time step for each particular fine-grained vector quantizer, (i) any acoustic token for any of the fine-grained vector quantizers at any second time step preceding the particular second time step, and (ii) any acoustic token at the particular second time step for any fine-grained vector quantizer preceding the particular vector quantizer within the hierarchy precede in the second generation order.
[0207] For example, the hierarchy can include Q vector quantizers, and the vector quantizers (Q’ + 1)...Q can be fine-grained vector quantizers. The acoustic token at a particular second time step t and a particular fine-grained vector quantizer q > Q’ is conditioned on all of the acoustic tokens of the coarse vector quantizer 612, any acoustic token at the second time step t preceding the second time step and in the fine-grained vector quantizers following Q’, and any acoustic token at the second time step t in the coarse vector quantizer and fine-grained vector quantizers preceding the particular fine-grained vector quantizer q.
[0208] In some embodiments, each particular acoustic token at each particular second time step for each particular fine vector quantizer is conditioned on (i) each respective acoustic token for one or more coarse vector quantizers that is at most a threshold number of second time steps before the second time step, and (ii) any acoustic token at a second time step that precedes the particular second time step in the second generation order and that is at most a threshold number of second time steps before the second time step. That is, the second time steps can be divided into non-overlapping batches of consecutive second time steps.
[0209] For example, the acoustic token at a particular second time step t and a particular fine vector quantizer q>Q’ is conditioned on the acoustic token of the coarse vector quantizer 612 that precedes the particular second time step t within the corresponding batch of the same second time step, any acoustic token at a second time step that precedes the second time step t within the corresponding batch and that follows Q’ in the fine vector quantizers, and any acoustic token at the second time step t in the coarse vector quantizers and fine vector quantizers that precede the particular fine vector quantizer q.
[0210] In some embodiments, the fine generation neural network 620 has a decoder-only transformer architecture. In some embodiments, the fine generation neural network 620 has an encoder-decoder transformer architecture.
[0211] In some embodiments, the system can generate acoustic tokens of the coarse vector quantizer 612 and the acoustic tokens of the fine vector quantizer 622 simultaneously. For example, in the case of a hierarchy of Q vector quantizers, the acoustic token at a particular second time step t of a particular coarse or fine vector quantizer q is conditioned on all of the acoustic tokens at a second time step preceding the second time step t and in the vector quantizer preceding q, as well as any acoustic tokens at the second time step t in the vector quantizer preceding the particular vector quantizer q.
[0212] The system can include acoustic tokens of the fine vector quantizer 622 within the acoustic representation. The system can use the acoustic representation to generate a prediction of the audio signal, as described below with reference to FIG. 7.
[0213] FIG. 7 is a flowchart of an exemplary process 700 for generating a prediction of an audio signal. For convenience, process 700 is described as being performed by one or more computer systems located in one or more locations. For example, an audio generation system appropriately programmed in accordance with this specification, such as the audio generation system 100 of FIG. 1, can execute process 700.
[0214] The system receives a request to generate an audio signal conditioned on an input (step 710). The audio signal has respective audio samples at each of a plurality of output time steps spanning a time window. The audio signal can be conditioned on the input. The input can be included, for example, in the context specified by the request. The input can include an input audio signal, visual data, and / or text data.
[0215] In an example where the input includes text data, such as a sequence of text, the prediction of the audio signal can be a prediction of music described by the sequence of text.
[0216] In some examples, a sequence of text can include multiple subsequences of text. Predicting an audio signal can be a prediction of music having sections of music corresponding to and reflecting each of the subsequences.
[0217] In an example where the input includes an input audio signal representing text data and a melody, predicting the audio signal can be a prediction of music described by a sequence of text and along the melody.
[0218] In an example where the input includes an input audio signal representing a melody, predicting the audio signal can be a prediction of music along the melody. For example, the input audio signal can represent a whistle or humming.
[0219] In an example where the input includes an input audio signal representing music, predicting the audio signal can be a continuation of the input audio signal.
[0220] In some examples where the input includes visual data, the system can obtain a text description of the visual data. In some embodiments, the system can generate a text description of the visual data. Predicting the audio signal can be a prediction of music described by the text description.
[0221] The system processes the input to map the input to one or more embedding tokens (step 715). Embedding tokens are also referred to as audio embedding tokens. The system can use an embedding neural network to process the input. For example, the system can use an embedding neural network to generate an embedding vector for the input and quantize the embedding vector to generate an embedding token.
[0222] In some examples, the input includes a sequence of text. An embedded neural network may be trained to map text and audio to a joint embedding space, also referred to as a joint audio embedding space. In examples where the input includes a sequence of text, the embedded tokens are generated from the text.
[0223] The embedded neural network may be trained with training data that includes audio signals. The training of the embedded neural network is described in more detail below with reference to FIG. 8.
[0224] In some examples, the input can include multiple subsequences of text. In these examples, the system can generate embedded tokens for each subsequence, so that the conditional signal can be changed for each subsequence when generating semantic and acoustic representations.
[0225] In some examples, the input includes a sequence of text and an input audio signal representing a melody. The system can use a melody-embedded neural network to process the input audio signal and map the input audio signal to one or more melody-embedded tokens, also referred to as melody audio-embedded tokens. The system can then concatenate the melody-embedded tokens with the embedded tokens.
[0226] In some examples, the input includes an input audio signal representing a melody. The system can use a melody-embedded neural network to process the input audio signal and map the audio signal to one or more melody-embedded tokens. The system can use the melody-embedded tokens as embedded tokens used for conditioning to generate semantic and acoustic representations.
[0227] In these examples where the input audio signal represents a melody, to process the input audio signal using a melody-embedding neural network, the system can generate one or more melody-embedding vectors of the input audio signal in a joint embedding space, also referred to as a joint audio embedding space, using the melody-embedding neural network. In some embodiments, the melody-embedding neural network can be a vision transformer (ViT) that receives time frames of a mel spectrogram of the input audio signal and generates a melody-embedding vector of the input audio signal.
[0228] The system can quantize the melody-embedding vector to generate a melody-embedding token. For example, the system can use residual vector quantization to quantize the melody-embedding vector into a melody-embedding token. The training of the melody-embedding neural network is described in more detail below with reference to FIG. 8.
[0229] The system generates a semantic representation of the audio signal (step 720). The semantic representation specifies respective semantic tokens at each of a plurality of first time steps spanning a time window. Each semantic token is selected from a vocabulary of semantic tokens conditioned on the embedding token and represents the semantic content of the audio signal at the corresponding first time step.
[0230] In some embodiments, the system can autoregressively generate the semantic representation using a semantic representation generation neural network having the embedding token as a conditioning signal. The semantic representation generation neural network may also be referred to as a third generation neural network and is similar to the semantic representation generation neural network described above with reference to FIG. 3.
[0231] For example, a semantic representation generation neural network can be trained to autoregressively generate semantic tokens one after another. The semantic representation generation neural network can be, for example, a neural network based on a decoder only, or an encoder-decoder transformer. For example, the semantic representation generation neural network can be trained to predict a semantic representation generated based on the output of one or more layers, such as one of the intermediate layers of an audio representation neural network.
[0232] The audio representation neural network can generate outputs or embeddings at regular time intervals. For example, the audio representation neural network can generate one embedding every 40 ms of the input audio signal.
[0233] The training of the semantic representation generation neural network and the audio representation neural network will be described in more detail below with reference to FIG. 8.
[0234] In some examples, the input includes a sequence of text, and the embedding tokens are generated from the text. In these examples, the system can condition the semantic representation generation neural network based on at least the embedding tokens while autoregressively generating the semantic representation of the audio signal.
[0235] In some examples, the input can include multiple subsequences of text, and the embedding tokens include embedding tokens for each subsequence. For example, in the case of the first subsequence, the system can condition a semantic representation generation neural network based on the embedding tokens of the first subsequence while autoregressively generating a semantic representation of the audio signal. In the case of the second subsequence, the system can condition a semantic representation generation neural network based on the embedding tokens of the second subsequence and the semantic tokens generated for the first subsequence.
[0236] In some examples, the input includes a sequence of text and an input audio signal representing a melody, the embedding tokens are generated from the text, and the melody embedding tokens are generated from the input audio signal. The system can concatenate the melody embedding tokens with the embedding tokens. The system can select each semantic token from a vocabulary of semantic tokens conditioned on the melody embedding tokens and the embedding tokens. The system can condition a semantic representation generation neural network based on the embedding tokens and the melody embedding tokens while autoregressively generating a semantic representation of the audio signal.
[0237] In some examples, the input includes an input audio signal representing a melody, and the melody embedding tokens are generated from the input audio signal. In these examples, the system can use the melody embedding tokens as the embedding tokens. The system can select each semantic token from a vocabulary of semantic tokens conditioned on the embedding tokens. The system can condition a semantic representation generation neural network based on the embedding tokens while autoregressively generating a semantic representation of the audio signal.
[0238] In some examples, the system can generate predictions of audio signals that are longer than the audio signals on which the system was trained. For example, the system can be trained on 30 - second audio signals. To generate a longer audio signal, the semantic - representation - generating neural network can use 15 seconds as a prefix, advancing in 15 - second strides, conditioned on the same embedding tokens from the input text, to generate an additional 15 seconds.
[0239] The system generates an acoustic representation of the audio signal (step 730). The system can use one or more generative neural networks to generate the acoustic representation conditioned on at least the semantic representation and the embedding tokens. The acoustic representation specifies a set of one or more respective acoustic tokens at each of a plurality of second time steps spanning a time window. The one or more respective acoustic tokens at each of the second time steps represent the acoustic characteristics of the audio signal at the corresponding second time step. Generating the acoustic representation is described in more detail above with reference to FIG. 6.
[0240] In some embodiments, the input includes a sequence of text. The one or more generative neural networks may be trained at least in part on audio - specific training data. During training, the semantic and acoustic representations are conditioned on the embedding of the audio input within a joint embedding space of text and audio. Training is described in more detail below with reference to FIG. 8.
[0241] The system processes at least the acoustic representation to generate a prediction of the audio signal (step 740). The system can process the acoustic representation using a decoder neural network. For example, each respective audio sample at a plurality of output time steps spanning a time window can be based on one or more acoustic representations of all of the vector quantizers within the hierarchy.
[0242] In some embodiments, the number of first time steps and the number of second time steps spanning a time window are less than the number of output time steps spanning the time window. In some embodiments, as described in more detail above with reference to FIG. 6, the number of first time steps spanning a time window is less than the number of second time steps spanning the time window.
[0243] FIG. 8 is a diagram of an exemplary process 800 for training an exemplary audio generation system. For convenience, process 800 is described as being performed by a training system of one or more computers located in one or more locations.
[0244] Process 800 is similar to process 400 described with reference to FIG. 4 in that it can train an audio generation system, such as audio generation system 100 of FIG. 1, with training data. In some embodiments, components of the audio generation system, such as the embedded neural network 120, the semantic representation generation neural network 130 (shown in FIG. 4 as "generation neural network 130"), one or more generation neural networks 108, the neural audio codec 420, and the audio representation neural network 410, can be trained separately.
[0245] The embedded neural network 120 can be a joint embedding model having two embedding towers, namely one for text and one for music. The embedded neural network 120 can be, for example, the MuLan model. The towers use contrastive learning to map text and music into a shared embedding space. The text network can be BERT pre-trained on a large corpus of text-only data. The music network can be, for example, a residual convolutional network.
[0246] The embedded neural network 120 can be trained with training data including an audio signal. For example, the training data can include a pair of a music clip and a corresponding text annotation. The embedded neural network 120 can be trained to link music to an unconstrained natural language description. For example, the embedded neural network 120 can be trained for the purpose such that the text describing the audio signal and the corresponding audio signal have embeddings that are close to each other within the joint embedding space of audio and text. In some embodiments, the embedded neural network 120 can be pre-trained and fixed before training the audio generation system.
[0247] Since the audio generation system can be trained with training data that is only audio, the training data can be easily expanded. The training data is not limited to audio data with text captions. Further, by training the embedded neural network 120 with a contrastive loss, robustness against noisy text descriptions can be enhanced.
[0248] To generate an embedded token for training the audio generation system, the training system can provide the target audio signal 850 to the embedded neural network 120. The embedded neural network 120 can generate a representation of the target audio signal 850 within the joint embedding space. The training system can quantize the representation into individual embedded tokens. The embedded tokens are based on the embedding of the target audio signal 850 within the joint embedding space.
[0249] In some embodiments, the training data can include audio signals that are longer than the audio signals that the embedded neural network 120 was pre-trained to operate on. For example, the target audio signal 850 can have a length of 30 seconds, and the embedded neural network 120 may have been pre-trained to operate on 10-second sequences. The training system can use the embedded neural network 120 to calculate representations with a 1-second stride over 10-second windows of the target audio signal 850 and average the resulting representations. The training system can then quantize the averaged representations to the individual embedded tokens of the target audio signal 850.
[0250] The training system can train the neural audio codec 420 and the audio representation neural network 410 with training data that includes audio clips such as music clips. An exemplary audio clip of the training data is shown in FIG. 8 as the target audio 850.
[0251] As described above with reference to FIG. 4, the neural audio codec 420 can include a decoder neural network and an encoder neural network. For example, the encoder neural network can convert the target audio 850 into a coded signal quantized to an acoustic representation. The decoder neural network can convert the acoustic representation into a predicted audio signal. In some embodiments, the neural audio codec 420 can be pre-trained and fixed before training of the audio generation system.
[0252] The neural audio codec 420 can be trained to minimize adversarial loss and reconstruction loss. For example, the decoder neural network and the encoder neural network can be jointly trained for the purpose of measuring the reconstruction quality of the predicted audio signal generated by the decoder neural network from the acoustic representation generated using the output generated by the encoder neural network.
[0253] The audio representation neural network 410 may be trained to generate a representation of the input audio signal. In some embodiments, the audio representation neural network 410 can be pre-trained and fixed prior to training of the audio generation system. The audio representation neural network 410 can be trained to minimize masked language model (MLM) loss and contrastive loss. The audio representation neural network 410 can be, for example, a w2v-BERT model.
[0254] The training system can train the semantic representation generation neural networks 130 and one or more generation neural networks 108 with training data including audio clips. In some embodiments, the generation neural network can be trained using teacher forcing.
[0255] The training system can train the semantic representation generation neural network 130 to predict a semantic representation generated based on the output of one or more layers, such as one of the intermediate layers of the audio representation neural network 410. The training system can train the semantic representation generation neural network 130 conditioned on the embedded tokens. The semantic representation generation neural network 130 can model the distribution p(S t |S <t , M A ), where S t represents the semantic token at the first time step t and M Arepresents an embedded token generated from audio.
[0256] For example, the audio representation neural network 410 can include a network having multiple layers. For example, the audio representation neural network 410 can include an attention-based model, such as a transformer-based model or a conformer-based model having multiple layers.
[0257] The attention-based model may be trained on a music representation task, for example, through self-supervised learning.
[0258] The output of one or more layers of the audio representation neural network 410 can include an embedding of the input audio. For example, an attention-based model can generate high-density embeddings for some or all of the audio samples in the training data. The embeddings of the intermediate layers of the attention-based model for some or all of the audio samples can be clustered into K clusters by the k-means method. The centroid of the cluster can be used as a semantic token. In some embodiments, the output can be normalized so that each dimension has zero mean and unit variance before clustering.
[0259] An exemplary target semantic representation for training the semantic representation generation neural network 130 may be generated by providing the target audio 450 to an attention-based model. The training system can generate the semantic tokens of the target semantic representation by assigning each output of the target audio 450 to the centroid of the nearest cluster.
[0260] In some embodiments, consecutive repetitions of semantic tokens within a target semantic representation can be removed. For example, the target semantic representation of the target audio 850 includes a sequence of semantic tokens. The training system can remove the consecutively repeated semantic tokens within the sequence and then use the target semantic representation without consecutively repeated semantic tokens for training.
[0261] The training system can train one or more generative neural networks 108 to generate an acoustic representation that is conditioned on at least the semantic representation and the embedding tokens. In some embodiments, the acoustic representation is a prediction of a ground truth acoustic representation that would be generated from the output of an encoder neural network by processing an audio signal. For example, the encoder neural network can be a convolutional encoder that maps the audio signal to a sequence of embeddings. Each respective embedding at each of a plurality of second time steps can correspond to a feature of the audio signal at the second time step. The ground truth acoustic representation can be generated by applying quantization to each of the respective embeddings. As described above, the encoder neural network can be part of a neural audio codec such as a neural audio codec.
[0262] For example, quantization can be residual vector quantization that encodes each embedding using a hierarchy of multiple vector quantizers, each generating a respective acoustic token from a corresponding vocabulary of vector quantizers for the vector quantizer. The hierarchy can include one or more coarse vector quantizers at one or more first positions within the hierarchy and one or more fine vector quantizers at one or more last positions within the hierarchy. The set of one or more respective acoustic tokens at each of the second time steps can include, for each vector quantizer, a respective acoustic token selected from the vocabulary of the vector quantizer and predicted to be the ground truth acoustic token that would be generated by the vector quantizer from the ground truth embedding generated by the encoder neural network at the second time step.
[0263] The training system can train one or more generative neural networks 108 to predict an acoustic representation generated based on the output of the residual vector quantizer of the neural audio codec 420. The training system can train one or more generative neural networks 408 conditioned on the embedding token. The one or more generative neural networks 408 can model the distribution p(A t |A <t , S, M A ), where A t represents an acoustic token at the first time step t, S represents a semantic representation, and M A represents an embedding token generated from the audio.
[0264] The target acoustic representation for training the one or more generative neural networks 108 may be generated by applying quantization to each of the respective embeddings output by the encoder neural network of the neural audio codec 420 at each of the plurality of second time steps. The quantization can be residual vector quantization (RVQ), as described above with reference to FIG. 6.
[0265] In some embodiments, the audio generation system also includes a melody-embedded neural network. The melody-embedded neural network may be trained with audio pairs that have matching melodies but different acoustics. For example, the training data can include different versions of the same music clip, such as covers, instrumentals, vocals, humming, and singing.
[0266] The melody-embedded neural network can be trained so that when two audio clips contain the same melody, the corresponding embeddings in a joint embedding space, also referred to as a joint audio embedding space, are close to each other. For example, the melody-embedded neural network can be trained to generate an audio embedding that represents the melody in the input audio signal while being invariant to the acoustic characteristics associated with the instrument being played, using a semi-hard triplet loss.
[0267] This specification uses the term "configured" in relation to components of a system and computer program. In the case of one or more computer systems configured to perform a particular operation or action, it means that the system has installed thereon, during operation, software, firmware, hardware, or a combination thereof that causes the system to perform the operation or action. In the case of one or more computer programs configured to perform a particular operation or action, it means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operation or action.
[0268] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, or in combinations of one or more of them, including the structures disclosed in this specification and their structural equivalents. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random access memory device, or a serial access memory device, or one or more combinations thereof. Alternatively or in addition, the program instructions may be generated and encoded as an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, for transmission to a receiver device suitable for execution by a data processing apparatus.
[0269] The term “data processing apparatus” refers to data processing hardware and encompasses, by way of example, all kinds of devices, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus may also include, either as a further possibility or in addition to the hardware, code for creating an execution environment for computer programs, e.g., processor firmware, protocol stacks, database management systems, operating systems, or code constituting one or more combinations thereof.
[0270] A computer program, which may also be referred to as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages. A computer program can be deployed in any form, including as a stand-alone program or in the form of a module, component, subroutine, or other unit suitable for use in a computing environment. A program may or may not correspond to a file in a file system. A program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a document of a markup language), in a single file dedicated to the program of interest, or in multiple related files, such as files that store one or more modules, subprograms, or portions of code. A computer program can be deployed to be executed on one computer or placed in one location, or distributed across multiple computers interconnected by a data communication network and executed on those multiple computers.
[0271] As used herein, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine is implemented as one or more software modules or components installed on one or more computers located in one or more locations. In some cases, one or more computers are dedicated to a particular engine, and in other cases, multiple engines can be installed and executed on the same one or more computers.
[0272] The processes and logical flows described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be executed by special purpose logic circuits, such as FPGAs or ASICs, or by a combination of special purpose logic circuits and one or more programmed computers.
[0273] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors, or both, or other types of central processing units. In general, a central processing unit receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a central processing unit for performing or executing instructions, and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuits. In general, a computer also includes, or is operatively coupled to receive data from, or transfer data to, or both, one or more mass storage devices for storing data, such as, by way of example only, magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Further, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio player or a mobile video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive.
[0274] Computer-readable media suitable for storing computer program instructions and data include, by way of example, semiconductor memory devices such as, for example, EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and all forms of non-volatile memory, media devices and memory devices including CD-ROM and DVD-ROM disks.
[0275] To implement interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to implement interaction with a user. For example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user can be received in any form including acoustic, speech, or tactile input. Further, the computer can interact with the user by sending documents to and receiving documents from the device used by the user, such as, for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser. Also, the computer can interact with the user by sending a text message or other form of message to a personal device such as a smartphone running a messaging application and receiving a reply message from the user in response thereto.
[0276] A data processing apparatus for implementing a machine learning model can also include, for example, a dedicated hardware accelerator unit for processing common computationally intensive parts of machine learning training or generation, i.e., inference, workload.
[0277] The machine learning model can be implemented and deployed using a machine learning framework such as the TensorFlow framework or the Jax framework.
[0278] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes, for example, backend components such as a data server, or, for example, middleware components such as an application server, or frontend components such as a graphical user interface, a web browser, or a client computer having an app through which a user can interact with an implementation of the subject matter described in this specification, or in any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0279] The computing system can include clients and servers. Clients and servers are generally located far apart from each other and typically interact through a communication network. The relationship between a client and a server arises by virtue of computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server sends data, such as an HTML page, to a user device for the purpose of displaying data to a user who interacts with a device that functions as a client and receiving user input from the user. For example, data generated at a user device, such as as a result of a user interaction, can be received at the server from the device.
[0280] This specification includes details of many specific embodiments, which should not be construed as limitations on the scope of any invention or the scope of what can be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Specific features described in the context of individual embodiments herein may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Further, even if a feature is described as functioning in a particular combination and was initially claimed as such, in some cases, one or more features from the claimed combination may be deleted from that combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.
[0281] Similarly, operations are depicted in the drawings and recited in the claims in a particular order, but it should not be understood that such operations must be performed in the particular order or sequence shown, or that all of the operations shown must be performed, in order to obtain a desirable result. In certain circumstances, multitasking and parallel processing may be advantageous. Further, separating the various system modules and components in the embodiments described above should not be understood to be required in all embodiments, and it should be understood that the program components and systems described may generally be integrated into a single software product or packaged into multiple software products.
[0282] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims may be performed in a different order and still achieve desirable results. As one example, in the processes shown in the accompanying figures, it is not necessary to follow the particular order or sequence shown, or to perform all of the operations shown, in order to obtain a desirable result. In some cases, multitasking and parallel processing may be advantageous.
Claims
**Claim 1** A computer-implemented method for generating a prediction of an audio signal, comprising: receiving a request for generating an audio signal having respective audio samples at each of a plurality of output time steps spanning a time window; obtaining a semantic representation of the audio signal that specifies respective semantic tokens at each of a plurality of first time steps spanning the time window, each semantic token being selected from a vocabulary of semantic tokens and representing the semantic content of the audio signal at the corresponding first time step; generating an acoustic representation of the audio signal using one or more generative neural networks and conditioned at least on the semantic representation, the acoustic representation specifying a set of one or more respective acoustic tokens at each of a plurality of second time steps spanning the time window, each of the one or more respective acoustic tokens at each second time step representing the acoustic characteristics of the audio signal at the corresponding second time step; generating the prediction of the audio signal by processing at least the acoustic representation using a decoder neural network; and a method comprising the steps of: **Claim 2** The method according to claim 1, wherein the decoder neural network is jointly trained with an encoder neural network for the purpose of measuring the reconstruction quality of a predicted audio signal generated by the decoder neural network from an acoustic representation generated using an output generated by the encoder neural network. **Claim 3** The method according to claim 1, wherein the acoustic representation is a prediction of a ground-truth acoustic representation that would be generated from an output of an encoder neural network by processing the audio signal. **Claim 4** The method according to claim 3, wherein the encoder neural network outputs respective embeddings at each of the plurality of second time steps, and the ground-truth acoustic representation is generated by applying quantization to each of the respective embeddings. Claim 5 wherein the quantization is residual vector quantization that encodes each embedding using a hierarchy of a plurality of the vector quantizers each generating a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer, the hierarchy comprising one or more coarse vector quantizers at one or more first positions within the hierarchy and one or more fine quantizers at one or more last positions within the hierarchy, wherein each set of the one or more respective acoustic tokens at each second time step is selected, for each vector quantizer, from the vocabulary of the vector quantizer and includes a respective acoustic token that is a prediction of a ground truth acoustic token that would be generated by the vector quantizer from a ground truth embedding generated by the encoder neural network at the second time step, The method according to claim 4. Claim 6 wherein each set of the one or more respective acoustic tokens at each of the plurality of second time steps includes a plurality of acoustic tokens that collectively represent a prediction of an output of residual vector quantization applied to an embedding representing acoustic characteristics of the audio signal at the second time step, wherein the residual vector quantization encodes the embedding using a hierarchy of a plurality of the vector quantizers each generating a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer, the hierarchy comprising one or more coarse vector quantizers at one or more first positions within the hierarchy and one or more fine vector quantizers at one or more last positions within the hierarchy, wherein each set of the acoustic tokens at each second time step includes a respective acoustic token selected, for each vector quantizer, from the vocabulary of the vector quantizer, The method according to claim 1. Claim 7 generating an acoustic representation of the audio signal using one or more generative neural networks and conditioned at least on the semantic representation, Using the first generation neural network, for each of the one or more coarse vector quantizers within the layer, generating each respective acoustic token of the second time step of the coarse vector quantizer, conditioned at least on the semantic representation. The method according to claim 6. **Claim 8** The method according to claim 7, wherein the first generation neural network is an autoregressive neural network configured to autoregressively generate the acoustic tokens according to a first generation order, and each respective acoustic token at each respective second time step for each particular coarse vector quantizer is conditioned on at least the semantic representation and any acoustic tokens preceding the particular acoustic token in the first generation order. **Claim 9** Prior to each respective acoustic token at each respective second time step for each particular coarse vector quantizer, (i) any acoustic tokens for any of the coarse vector quantizers at any second time step preceding the particular second time step, and (ii) any acoustic tokens at the particular second time step of any coarse vector quantizer preceding the particular coarse vector quantizer within the layer precede in the first generation order. The method according to claim 8. **Claim 10** The method according to claim 7, wherein the first generation neural network has a decoder-only transformer architecture or an encoder-decoder transformer architecture. **Claim 11** Generating an acoustic representation of the audio signal using one or more generation neural networks and conditioned at least on the semantic representation. Using a second generation neural network, for each of the one or more fine vector quantizers within the layer, generating each respective acoustic token of the second time step of the fine vector quantizer, conditioned on each respective acoustic token of the second time step of the one or more coarse vector quantizers within the layer. The method according to claim 7. **Claim 12** The method according to claim 11, wherein the second generation neural network is not conditioned on the semantic representation. **Claim 13** The second generation neural network is an autoregressive neural network configured to autoregressively generate the acoustic tokens according to a second generation order, and each specific acoustic token at each specific second time step for each specific fine vector quantizer is based on (i) each respective acoustic token for at least a subset of the second time steps of the one or more coarse vector quantizers, and (ii) at least a subset of the acoustic tokens preceding the specific acoustic token in the second generation order. The method according to claim 11.
14. Before each specific acoustic token at each specific second time step for each specific fine vector quantizer, (i) any acoustic token for any of the fine vector quantizers at any second time step preceding the specific second time step, and (ii) any acoustic token at the specific second time step for any fine vector quantizer preceding the specific fine vector quantizer within the hierarchy precede in the second generation order. The method according to claim 13.
15. Each specific acoustic token at each specific second time step for each specific fine vector quantizer is based on (i) each respective acoustic token for the one or more coarse vector quantizers that are at most a threshold number of second time steps before the second time step, and (ii) any acoustic token at a second time step that precedes the specific second time step in the second generation order and is at most a threshold number of second time steps before the second time step. The method according to claim 13.
16. The second generation neural network has a decoder-only transformer architecture or an encoder-decoder transformer architecture. The method according to claim 11.
17. Obtaining the semantic representation of the audio signal includes using a third generation neural network to autoregressively generate the semantic representation. The method according to claim 1.
18. The method according to claim 1, wherein the requirement specifies the context of the audio signal, and the audio signal is conditional on the context.
19. wherein the context specifies semantic characteristics of the audio signal, and obtaining a semantic representation of the audio signal, generating the semantic representation conditional on the context, The method according to claim 18.
20. wherein the context specifies acoustic characteristics of the audio signal, generating an acoustic representation of the audio signal using one or more generative neural networks and conditional on at least the semantic representation, generating an acoustic representation of the audio signal using one or more generative neural networks and conditional on the semantic representation and the context, The method according to claim 18.
21. The set of respective acoustic tokens at each of the plurality of second time steps includes a plurality of acoustic tokens that collectively represent a prediction of an output of residual vector quantization applied to an embedding representing acoustic characteristics of the audio signal at the second time step, wherein the residual vector quantization encodes the embedding using a plurality of hierarchical levels of the vector quantizers, each of which generates a respective acoustic token from a corresponding vocabulary of acoustic tokens for a vector quantizer, the hierarchy comprising one or more coarse vector quantizers at one or more first positions within the hierarchy and one or more fine vector quantizers at one or more last positions within the hierarchy, the set of acoustic tokens at each of the second time steps includes respective acoustic tokens selected from the vocabulary of the vector quantizer for each vector quantizer, generating an acoustic representation of the audio signal using one or more generative neural networks and conditional on the semantic representation and the context, using a first generative neural network to generate, for each of the one or more coarse vector quantizers within the hierarchy, the respective acoustic tokens at the second time step of the coarse vector quantizer conditional on the semantic representation and the context, The method according to claim 20.
22. Processing at least the acoustic representation using a decoder neural network to generate the prediction of the audio signal, including processing the acoustic representation and the acoustic representation of the context using the decoder neural network to generate the prediction of the audio signal, The method according to claim 20.
23. The method according to claim 18, wherein the context includes an audio input.
24. The method according to claim 18, wherein the context includes visual data.
25. The method according to claim 18, wherein the context includes text data.
26. The method according to claim 1, wherein the number of first time steps and the number of second time steps spanning the time window are less than the number of output time steps spanning the time window.
27. The method according to claim 26, wherein the number of first time steps spanning the time window is less than the number of second time steps spanning the time window.
28. A system, one or more computers, one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform respective operations of the method according to claim 1 A system comprising.
29. One or more computer-readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform respective operations of the method according to claim 1.
Citation Information
Patent Citations
Speech translation device and speech translation method
JP2007148039A
Speech synthesis system, speech synthesis program and speech synthesis method
JP2018141915A
Expressive text-to-speech utilizing contextual word-level style tokens
US20220028367A1
Two-level speech prosody transfer
WO2022035586A1