Generating audio using generative neural networks
Patent Information
- Application Number
- PCT/EP2024/083041
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-20
- Filing Date
- 2024-11-20
- Publication Date
- 2025-07-10
AI Technical Summary
Existing audio generation technologies struggle to create high-quality, coherent audio that accurately reflects user inputs, especially when generating long audio pieces, as they often require significant computational resources and may result in poor coherence and quality.
The system employs a hierarchy of diffusion models, specifically using base and upsampling diffusion neural networks, to generate audio conditioned on user inputs. This approach allows for the iterative extension of audio chunks, ensuring temporal coherence and fine-grained details, even for longer audio pieces without degrading quality.
The system effectively generates high-quality audio that accurately reflects user inputs, maintaining coherence and quality across varying lengths, thereby overcoming the limitations of existing technologies in terms of computational efficiency and audio fidelity.
Smart Images

Figure EP2024083041_10072025_PF_FP_ABST
Abstract
Description
[0001] GENERATING AUDIO USING GENERATIVE NEURAL NETWORKS
[0002] CROSS-REFERENCE TO RELATED APPLICATION
[0003] This application claims priority to U.S. Provisional Application No. 63 / 601,183, filed on November 20, 2023. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.
[0004] BACKGROUND
[0005] This specification relates to generating audio conditioned on conditioning inputs using neural networks.
[0006] Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to or more other layers in the network, i.e., one or more other hidden layers, the output layer, or both. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.
[0007] SUMMARY
[0008] This specification describes a system implemented as computer programs on one or more computers in one or more locations that generates audio that represents an audio signal conditioned on a conditioning input using generative neural networks.
[0009] For example, the system can generate a spectrogram that represents the audio signal and then convert the spectrogram into an audio waveform.
[0010] As one particular example, the system can generate a musical composition, e.g., a song that includes lyrics or an instrumental track that does not have lyrics, conditioned on a user input that characterizes desired properties of the musical composition.
[0011] As another particular example, the system can generate a soundscape that represents a particular audio environment, conditioned on a user input that characterizes desired audio properties of the environment.
[0012] In some cases, the system can generate this musical composition using a hierarchy of diffusion models.
[0013] For example, the system can first generate a spectrogram representing a musical composition of a first length using a first set of diffusion neural networks. Optionally, the system can then generate an extended spectrogram representing a longer musical composition conditioned on the user input and on the spectrogram of the first length by using a second set of diffusion neural networks.
[0014] A “spectrogram” as used in this specification is a visual representation of the spectrum of frequencies of an audio signal over time. The frequencies in the spectrogram can be represented in any appropriate way. For example, the spectrogram can be a log-mel amplitude spectrogram that represents the frequencies in a log-mel scale or a different type of spectrogram that represents frequencies in a different scale.
[0015] In some cases, a “spectrogram” refers to a “stereo” spectrogram that is a combination of a respective spectrogram for two or more different, temporally-aligned audio signals, e.g., that can be played back simultaneously from two or more different sound sources to create a multi-dimensional perspective. In these cases, the stereo spectrogram can be, e.g., a concatenation of multiple different spectrograms along the depth dimension.
[0016] Additionally, in these cases, the system can generate a stereo audio signal from the stereo spectrogram, i.e., can generate two or more separate waveforms that are intended to be played back simultaneously from respective sound sources.
[0017] Optionally, the system can also generate an image (an “album cover”) that visually represents certain properties of the generated musical composition.
[0018] In some cases, the system can use a language model neural network to transform a user input, e.g., a natural language user input, into an input prompt for the generative neural network system that generates the musical composition and, optionally, another input prompt for the generative neural network system that generates the album cover image.
[0019] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.
[0020] By making use of the described techniques, the system can receive a user input and, in response, effectively generate high-quality music that accurately reflects the properties of the music that are characterized by the user input, e.g., by generating music that is described by a natural language text or audio input from a user.
[0021] In particular, the system can effectively leverage diffusion neural networks by first generating a lower-resolution spectrogram of the desired music using one diffusion neural network and then upsampling the lower-resolution spectrogram using another diffusion neural network. By making use of this hierarchy of diffusion neural networks, the system can ensure that the resulting audio is temporally coherent while still exhibiting fine-grained details that reflect the user input. Optionally, a longer piece of music may be required. In these cases, the system can extend an initially generated “chunk” of the music using another set of diffusion neural networks. By iteratively extending the initially generated chunk, the system can generate significantly longer audio without requiring an exceedingly large amount of memory and without degrading the quality or coherence of the audio as the length increases. That is, directly generating long, e.g., multiple minute long, audio may require a significant amount of computational resources, e.g., memory, and may require result in an output that has poor coherence. Iteratively generating new chunks that extend on already generated chunks addresses these issues, resulting in high-quality long audio.
[0022] As a result, the system can generate high-quality music at any of a variety lengths in a computationally efficient manner. That is, the system can generate high-quality music, even when the desired or target length for the music exceeds the modeling capacity of any given one of the diffusion neural networks employed by the system.
[0023] Additionally, by leveraging diffusion models, the system can effectively generate multiple high-quality, plausible musical compositions in response to a given user input, e.g., that represent multiple different plausible interpretations of the user input. That is, because diffusion models start from a noisy representation of the output, by sampling different noise to initialize the representation for any given one of the diffusion models used by the system, the system can generate a different musical composition that is high-quality, plausible interpretation of the user input.
[0024] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description and the drawings.
[0025] BRIEF DESCRIPTION OF THE DRAWINGS
[0026] FIG. l is a diagram of an example audio generation system.
[0027] FIG. 2 shows an example of the operation of the audio generative neural network system.
[0028] FIG. 3 shows another example of the operation of the audio generative neural network system.
[0029] FIG. 4 is a flow diagram of an example process for generating a high-resolution spectrogram that has a first length. FIG. 5 is a flow diagram of an example process for generating a high-resolution spectrogram that has an extended length.
[0030] FIG. 6 is a flow diagram of an example process for generating an audio input and an image input from a user input.
[0031] FIGS. 7A-7C show an example of the operation of the system.
[0032] Like reference numbers and designations in the various drawings indicate like elements.
[0033] DETAILED DESCRIPTION
[0034] This specification describes a system implemented as computer programs on one or more computers in one or more locations that generates audio conditioned on a conditioning input.
[0035] FIG. 1 is a diagram of an example audio generation system 100. The audio generation system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented.
[0036] The system 100 generates, conditioned on a conditioning input 110 and using an audio generative neural network system 130, audio 120 that represents an audio signal. For example, the audio 120 can be a waveform representing the audio signal, with the waveform including a respective amplitude value of the audio signal at each of multiple time points.
[0037] Generally, the system 100 can generate audio 120 that has a target length, i.e., that spans a time window of a target length. In some implementations, the target length is fixed, i.e., so that each audio signal generated by the system 100 has the same length. In some other implementations, the target length is variable, and different audio signals generated by the system 100 can have different lengths.
[0038] For example, the system 100 can generate a spectrogram that represents the audio signal and then convert the spectrogram into an audio waveform.
[0039] A “spectrogram” as used in this specification is a visual representation of the spectrum of frequencies of an audio signal over time. The frequencies in the spectrogram can be represented in any appropriate way. For example, the spectrogram can be a log-mel amplitude spectrogram that represents the frequencies in a log-mel scale or a different type of spectrogram that represents frequencies in a different scale.
[0040] In some cases, a “spectrogram” refers to a “stereo” spectrogram that is a combination of a respective spectrogram for two or more different, temporally-aligned audio signals, e.g., that can be played back simultaneously from two or more different sound sources to create a multi-dimensional perspective. In these cases, the stereo spectrogram can be, e.g., a concatenation of multiple different spectrograms along the depth dimension.
[0041] Additionally, in these cases, the system 100 can generate a stereo audio signal from the stereo spectrogram, i.e., can generate two or more separate waveforms that are intended to be played back simultaneously from respective sound sources.
[0042] The system 100 can then play back the audio to a user or save the audio waveform in memory for later playback.
[0043] As a particular example, the system 100 can generate a musical composition, e.g., a song that includes lyrics or an instrumental composition that does not have lyrics, conditioned on a user input that characterizes desired properties of the musical composition or a soundscape conditioned on a user input that characterizes desired properties of the soundscape. A “soundscape” is audio that characterizes a particular environment, such as a domestic home, office, hospital, shopping center or mall, beach, forest, and so forth.
[0044] That is, the audio 120 can be a musical composition or soundscape and the conditioning input 110 can be a user input that characterizes desired properties of the musical composition or soundscape.
[0045] For example, the user input can be natural language text that describes the desired content of the musical composition or soundscape.
[0046] As another example, the system 100 can present a user interface that allows the user to provide structured text that includes values for one or more properties of the musical composition or soundscape. Thus, in these cases, the user input is structured text that includes respective values for each of one or more properties of the musical composition or soundscape.
[0047] In some implementations, the system 100 can receive, from the user and as the user input, or generate, from the user input, an audio input 132 that includes (i) a positive input that specifies desired properties of the musical composition or soundscape and optionally (ii) lyrics to be sung in the musical composition. In some of these implementations, the input also includes (iii) a negative input that specifies properties that the musical composition or soundscape should not have.
[0048] The system 100 then processes the audio input 132 using the audio generative neural network system 130 to generate the audio 120.
[0049] Optionally, the system 100 can also generate an image 140 (an “album cover”) that visually represents certain properties of the generated musical composition. In particular, the system 100 can generate an image input 142 that characterizes desired properties of the image 140 and then use an image generative neural network system 150 to generate the image 140 conditioned on the image input 142.
[0050] For example, the system 100 can receive, from the user and as the user input, or generate, from the user input, an image input 142 that includes (i) a positive image input that specifies desired properties of the image 140. In some of these implementations, the image input 142 also includes (ii) a negative image input that specifies properties that the image 142 should not have.
[0051] The image generative neural network system 150 can use any appropriate generative neural network to generate the image 140.
[0052] For example, the system 150 can use a diffusion neural network to generate the image 140. One example of such a neural network is a latent diffusion model, e.g., Stable Diffusion. Another example of such a neural network is a diffusion model that uses a text-to- image diffusion model to generate a first image, and then applies one or more superresolution diffusion models to the first image to generate the final image 140. One example of such a model is Imagen.
[0053] As another example, the system 150 can use an auto-regressive generative model to generate the image 140. One example of such a model is Parti (Yu et al. “Scaling Autoregressive Models for Content-Rich Text-to-Image Generation”, arXiv:2206.10789vl, 2022).
[0054] As yet another example, the system 150 can use a masked token generative model that sequentially unmasks visual tokens during generation. One example of such a generative model is Muse (Chang et al. “Muse: Text-To-Image Generation via Masked Generative Transformers”, arXiv:2301.00704vl, 2023).
[0055] In some implementations, the system 100 can use a language model neural network to generate both the audio input 132 and the music input 142 from the user input.
[0056] For example, the system 100 can use the language model neural network to map a natural language user input describing a musical composition or soundscape to a structured text sequence that includes at least a positive audio input. Optionally, the sequence can also include a negative audio input. Further optionally, the structured text sequence can also include lyrics generated by the language model neural network. The system 100 can then generate at least a positive image prompt from the structured text sequence that is provided to the image generation model to generate the album cover. The language model neural network can have any appropriate neural network architecture that allows the neural network to map an input sequence of tokens from a vocabulary to an output sequence of tokens from the vocabulary.
[0057] The vocabulary of tokens can include any of a variety of tokens that represent text symbols or other symbols. For example, the vocabulary of text tokens can include one or more of characters, sub-words, words, punctuation marks, numbers, or other symbols that appear in a corpus of text in a natural language and / or a computer programming language.
[0058] For example, the language model neural network can be a Transformer-based language model neural network or a recurrent neural network-based language model. As a particular example, the language model neural network can be an auto-regressive Transformer-based neural network that has, e.g., an encoder-only Transformer architecture, an encoder-decoder Transformer architecture, or a decoder-only Transformer architecture.
[0059] Examples of such architectures include those described in Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv: 1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like opendomain chatbot. CoRR, abs / 2001.09977, 2020; Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, and Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020; Aakanksha Chowdhery, et al. PaLM: Scaling Language Modeling with Pathways, arXiv preprint arXiv: 2204.02311; and Rohan Anil, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
[0060] FIG. 2 shows an example audio generative neural network system 130 that generates a musical composition or soundscape 220 conditioned on the audio input 132.
[0061] The system 130 can generate this musical composition or soundscape 220 using a hierarchy of diffusion models that includes a base diffusion neural network 260 and an upsampling diffusion neural network 270.
[0062] In particular, the system 130 can first generate a low-resolution spectrogram 240 representing a musical composition or soundscape having a first resolution using the base diffusion neural network 260.
[0063] Generally, the low-resolution spectrogram 240 is referred to as a low-resolution spectrogram because the spectrogram spans the same time window as the musical composition or soundscape 220 but has a first resolution that is lower than the target resolution of the musical composition or soundscape 220.
[0064] For example, if the target resolution is x by y, the first resolution can be x / 4 by y / 4, x / 8 by y / 8, or x / 16 by y / 16.
[0065] The base diffusion neural network 260 is a diffusion neural network that is configured to receive a “base” diffusion input that includes a representation of an audio input 210 and a current representation of the low-resolution spectrogram 240 and to process the diffusion input to generate a “base” denoising output. For example, the representation of the audio input 210 can be embeddings of the audio input 210 that are generated by a text embedding neural network.
[0066] The system 130 uses the base diffusion neural network 260 to generate the low- resolution spectrogram 240 across multiple sampling steps.
[0067] Prior to performing the multiple sampling steps, the system 130 initializes a representation of the spectrogram 240 by sampling a respective noisy value for each value of the low-resolution spectrogram 240 from a noise distribution, e.g., a Gaussian distribution. Thus, the representation of the spectrogram 240 initially has only noisy values.
[0068] At each sampling step, the system 130 processes a base diffusion input that includes the representation of the audio input 210 and the current representation of the spectrogram 240 using the base diffusion neural network 260 to generate a base denoising output for the sampling step.
[0069] When the audio input 210 includes a positive audio input, a negative audio input, and lyrics, the system 130 can process two separate base diffusion inputs at each sampling iteration.
[0070] In particular, the system 130 can process a positive base diffusion input that includes a representation of the positive audio input and the lyrics and the current representation of the spectrogram 240 using the base diffusion neural network 260 to generate a positive base denoising output for the sampling step.
[0071] The system 130 can also process a negative base diffusion input that includes a representation of the negative audio input and the current representation of the spectrogram 240 using the base diffusion neural network 260 to generate a negative base denoising output for the sampling step.
[0072] The system 130 can then combine the negative and positive base denoising output in accordance with a guidance weight for the sampling step to generate a final base denoising output for the sampling step. The system 130 then uses the base denoising output to update the representation.
[0073] For example, the system 130 can compute, from the current representation and the base denoising output, an estimate of the spectrogram 240 and then use the estimate to update the current representation.
[0074] For each sampling step other than the last sampling step, the system 130 can apply a diffusion sampler to the estimate to generate the updated representation. The system can use any appropriate diffusion sampler, e.g., the DDPM (Denoising Diffusion Probabilistic Model) sampler, the DDIM (Denoising Diffusion Implicit Model) sampler or another appropriate sampler.
[0075] For the last sampling step, the system 130 can use the estimate as the updated representation.
[0076] The system 130 then uses the updated representation after the last sampling step as the final spectrogram 240. In other words, the system 130 gradually “de-noises” the initial noisy representation of the spectrogram 240 using the base diffusion neural network to generate the low-resolution spectrogram 240 from the audio input 210.
[0077] The system 130 can then use the low-resolution spectrogram 240 of the musical composition or soundscape of the first length to condition the upsampling diffusion neural network 270 to generate a spectrogram 280 of a second, higher resolution.
[0078] Optionally, the upsampling diffusion neural network 270 can also be conditioned on the audio input, e.g., as generated by the language model neural network from the user request.
[0079] The system 130 can condition the upsampling diffusion neural network 270 on the low-resolution spectrogram 240 in any of a variety of ways.
[0080] For example, the system 130 can up-sample the low-resolution spectrogram to have the higher resolution and then, at each sampling step during generation of the higher- resolution spectrogram 280, include, in the diffusion input for the upsampling diffusion neural network 270, the up-sampled low-resolution spectrogram. For example, the diffusion input can include a concatenation of the up-sampled low-resolution spectrogram and the current representation of the higher-resolution spectrogram 280.
[0081] The system 130 can use the upsampling diffusion neural network 270, conditioned on the low-resolution spectrogram and, optionally, the audio input, to generate the higher- resolution spectrogram 280 across multiple sampling steps, e.g., in the same manner described above for the generation of the low-resolution spectrogram. In some implementations, the system 130 then uses a vocoder 290 to generate the musical composition or soundscape 220 from the spectrogram 280.
[0082] The vocoder 290 can be any appropriate software that maps spectrograms to waveforms.
[0083] For example, the vocoder 290 can be one that applies a phase reconstruction method to the spectrogram to generate the waveform, e.g., a Griffin-Lim vocoder.
[0084] As another example, the vocoder 290 can be a trained neural network-based vocoder, e.g., a diffusion neural network or a WaveNet-based neural network.
[0085] In some other implementations, the spectrogram 280 can represent a musical composition or soundscape that has a shorter length (along the time dimension) than the target length (along the time dimension) of the final composition to be generated by the system 130.
[0086] That is, the spectrogram 280 can span a time window that is shorter than the target time window to be spanned by the final musical composition or soundscape. As a particular example, the spectrogram 280 can be 30 seconds long, while the target length for the final musical composition or soundscape can be 3 minutes.
[0087] For example, this discrepancy in length can arise because the target length may be longer than the length of audio that can be accurately modeled by the diffusion model hierarchy directly from the audio input.
[0088] To account for this, in some implementations, the system 130 can use the spectrogram 280 to iteratively generate the final musical composition or soundscape.
[0089] That is, at each of multiple iterations, the system 130 can use the already generated spectrogram(s) as of that iteration to generate a new spectrogram that spans a time window that immediately follows the most-recently generated spectrogram in the final musical composition or soundscape. In other words, the system 130 can iteratively extend the length of the generated musical composition or soundscape until the target length is reached.
[0090] FIG. 3 shows an example 300 of the operation of the audio generative neural network system 130 when iteratively extending generated audio.
[0091] In particular, as described above, the system 130 first generates a spectrogram 280 that represents the initial time window of a final, longer time window. More specifically, the initial time window has a first length while the final time window has a second, longer length. For example, the system can generate 15 second, 30 second, or 45 second “chunks” of a longer 3 minute, 4 minute, or 5 minute song. Prior to iteratively extending the spectrogram 280, the system 130 uses the spectrogram 280 to generate an initial extended spectrogram 310 that spans the entire final time window but that has a lower resolution than the target resolution for the musical composition or soundscape (and, therefore, than the resolution of the spectrogram 280).
[0092] The system 130 can generate the initial extended spectrogram 310 using an initial low-resolution extension diffusion neural network 320.
[0093] The diffusion neural network 320 is a diffusion neural network that is configured (trained) to process an initial low-resolution extension diffusion input that includes a representation of the audio input 210, a representation of the spectrogram 280, and a current representation of the spectrogram 310 and to process the diffusion input to generate a low- resolution extension denoising output.
[0094] The system 130 uses the initial low-resolution extension diffusion neural network 320 to generate the spectrogram 310 across multiple sampling steps.
[0095] Prior to performing the multiple sampling steps, the system 130 initializes the representation of the spectrogram 310.
[0096] In particular, the system 130 initializes the representation by sampling a respective noisy value for each of value of the spectrogram 310 from a noise distribution, e.g., a Gaussian distribution. Thus, the representation of the spectrogram 310 initially has only noisy values.
[0097] At each sampling step, the system 130 processes an initial low-resolution extension diffusion input that includes the current representation of the spectrogram 310 and the other data described above using the initial low-resolution extension diffusion neural network 320 to generate an initial low-resolution extension denoising output for the sampling step.
[0098] The system 130 then uses the initial low-resolution extension denoising output to update the representation.
[0099] For example, the system can compute, from the current representation and the low- resolution extension denoising output, an estimate of the spectrogram 310 and then use the estimate to update the current representation.
[0100] For each sampling step other than the last sampling step, the system 130 can apply a diffusion sampler to the estimate to generate the updated representation.
[0101] For the last sampling step, the system 130 can use the estimate as the updated representation.
[0102] After generating the spectrogram 310, the system 130 then iteratively extends the spectrogram 280 at each of multiple iterations. In particular, at each iteration, the system 130 generates a new high-resolution spectrogram 330 conditioned on (i) as will be described in more detail below, a low- resolution spectrogram that was used to generate the most-recently generated high-resolution spectrogram and (ii) the corresponding portion of the initial extended spectrogram 310.
[0103] At the first iteration, the most-recently generated high-resolution spectrogram is the spectrogram 280. At each subsequent iteration, the most-recently generated high-resolution spectrogram is the high-resolution spectrogram generated at the preceding iteration.
[0104] The corresponding portion of the initial extended spectrogram 310 is a portion that includes a portion that spans the same time window as the new high-resolution spectrogram 330. Optionally, the corresponding portion can also include a portion that precedes the time window spanned by the new high-resolution spectrogram 330 as additional context.
[0105] The system 130 can generate the new-high resolution spectrogram 330 using a high- resolution extension system 340.
[0106] The system 340 can in turn include (i) a low-resolution extension diffusion neural network and (ii) an upsampling diffusion neural network.
[0107] In some implementations, the low-resolution extension diffusion neural network is the same as the base diffusion neural network 260 and the upsampling diffusion neural network is the same as the upsampling diffusion neural network 270.
[0108] In some other implementations, the parameter values for one or both of the low- resolution extension diffusion neural network and the upsampling diffusion neural network have been learned separately from those of the neural networks 260 and 270.
[0109] The low-resolution extension diffusion neural network is a diffusion neural network that is configured to process a low-resolution extension diffusion input that includes a representation of the audio input 210, a representation of the corresponding portion of the initial extended spectrogram 310, a representation of the most-recently generated low- resolution spectrogram, and a current representation of the spectrogram 330 and to process the diffusion input to generate an high-resolution extension denoising output. When the audio input 210 includes lyrics, the representation of the audio input 210 can include only the portion of the lyrics that approximately corresponds to the time window spanned by the spectrogram 330.
[0110] The system 130 uses the low-resolution extension diffusion neural network 340 to generate a low-resolution spectrogram across multiple sampling steps, e.g., a spectrogram having the first resolution but that spans the same time window as the spectrogram 330. Prior to performing the multiple sampling steps, the system 130 initializes the representation of the low-resolution spectrogram.
[0111] In some implementations, the system 130 initializes the representation by sampling a respective noisy value for each of value of the spectrogram 240 from a noise distribution, e.g., a Gaussian distribution. Thus, the representation initially has only noisy values.
[0112] In some other implementations, the system 130 can upsample the corresponding portion of the initial extended spectrogram 310 to generate an upsampled spectrogram that has the first resolution. The system can then use the upsampled spectrogram as the initialized representation.
[0113] At each sampling step, the system 340 processes a low-resolution extension diffusion input that includes the current representation of the spectrogram and the other data described above using the low-resolution extension diffusion neural network to generate a low- resolution extension denoising output for the sampling step.
[0114] The system 340 then uses the low-resolution extension denoising output to update the representation.
[0115] For example, the system 340 can compute, from the current representation and the high-resolution extension denoising output, an estimate of the spectrogram and then use the estimate to update the current representation. For each sampling step other than the last sampling step, the system 340 can apply a diffusion sampler to the estimate to generate the updated representation. For the last sampling step, the system 340 can use the estimate as the updated representation.
[0116] The system 340 can then use the upsampling diffusion neural network to upsample the low-resolution spectrogram to have the target resolution, e.g., as described above with reference to FIG. 2.
[0117] The diffusion neural networks, i.e., the diffusion neural networks 260, 270, 320, and 340, described in this specification can have any appropriate architecture. For example, some or all of the diffusion neural networks can be convolutional neural network, e.g., a U-Net or other architecture that maps one input of a given dimensionality to an output of the same dimensionality, with either the same or different number of channels. More generally, because the diffusion neural networks generate spectrograms as output and spectrograms can be represented as images, the diffusion neural networks can have the architecture of any appropriate text-to-image diffusion neural network.
[0118] In any of the above examples, the data item generated using the diffusion neural network can either be an output in the output space, i.e., so that the values in the spectrogram are the values of a spectrogram of the appropriate resolution, or an output in a latent space, i.e., so that the values in the output data item are values in a latent representation of a spectrogram.
[0119] When the output data item is generated in a latent space, the system can generate a final spectrogram in output space by processing the output data item in the latent space using a decoder neural network, e.g., one that has been pre-trained in an auto-encoder framework. During training, the system can use an encoder neural network, e.g., one that has been pretrained jointly with the decoder in the auto-encoder framework, to encode target spectrograms in the output space to generate target outputs for the diffusion neural network in the latent space.
[0120] More specifically, as described above, for any given one of the diffusion neural networks, the system uses the diffusion neural network to perform a reverse diffusion process across multiple updating iterations to generate the output data item.
[0121] Any given one of the diffusion neural networks described above can be any appropriate diffusion neural network that has been trained, e.g., by the system or another training system, to, at any given updating iteration, process a diffusion input for the updating iteration that includes the current data item (as of the updating iteration) to generate a denoising output for the updating iteration.
[0122] In some implementations, the denoising output is an estimate of the noise component of the current data item, i.e., the noise that needs to be combined with, e.g., added to or subtracted to, a final data item, i.e., to the output data item being generated by the system, to generate the current data item.
[0123] In some other implementations, the denoising output is an estimate of the final data item given the current data item, i.e., an estimate of the data item that would result from removing the noise component of the current data item.
[0124] In some other implementations, the denoising output is a prediction of a v- parametrization of the noise component and the final data item (Salimans and Ho arXiv: 2202.00512, 2022, section 4; Appendix D).
[0125] For example, the system or another training system can have trained the diffusion neural network on a set of training data items using a denoising score-matching objective to generate the denoising output.
[0126] The denoising score-matching objective can measure an error, e.g., a mean-squared error, an LI error, an L2 error or a different type of error, between (i) a denoising output generated by processing an input that includes a noisy data item generated by adding sampled noise to a training data item and (ii) a target denoising output generated from the training data item, from the sampled noise, or both.
[0127] For example, when the denoising output is an estimate of the true noise component of the current data item, the target denoising output can be the sampled noise.
[0128] As another example, when the denoising output is an estimate of the target data item, the target denoising output can be the target data item.
[0129] As another example, when the denoising output is a prediction of a v-parametrization of the noise component and the final data item, the target denoising output can be the true v- parametrization of the sampled noise and the target data item.
[0130] That is, to train a given one of the diffusion neural networks, the system can obtain a training data set that includes spectrograms having the corresponding resolution and a respective conditioning input for each of the spectrograms. For example, the system obtain spectrograms representing synthetic or real-world musical compositions or soundscapes and can then downsample each spectrogram as required to generate a spectrogram having the corresponding resolution. The system can also obtain structured data, e.g. structured text as described elsewhere in this specification, describing properties of the musical compositions or soundscapes or natural language of the musical compositions or soundscapes for use as the conditioning inputs. The structured data, e.g. structured text can be obtained, e.g. by human labelling of the spectrograms in the training data set or of the audio represented by the spectrograms in the training data set. The system can then train the diffusion neural network on the objective described above using the spectrograms and the conditioning input.
[0131] As described above, the diffusion neural network can have any appropriate architecture that allows the neural network to map a diffusion input that includes an input data item that has the same dimensionality as the output data item to a denoising output that also has the same dimensionality as the output data item.
[0132] Moreover, as described above, the neural network can be conditioned on a conditioning input in any of a variety of ways.
[0133] As one example, the system can use an encoder neural network to generate one or more embeddings that represent the conditioning input and the diffusion neural network can include one or more cross-attention layers that each cross-attend into the one or more embeddings.
[0134] An embedding, as used in this specification, is an ordered collection of numerical values, e.g., a vector of floating point values or other types of values. For example, when the conditioning input is text, the system can use a text encoder neural network, e.g., a Transformer neural network, to generate a fixed or variable number of text embeddings that represent the conditioning input.
[0135] When the conditioning input is an image, e.g., a spectrogram, the system can use an image encoder neural network, e.g., a convolutional neural network or a vision Transformer neural network, to generate a set of embeddings that represent the image.
[0136] When the conditioning input is audio, the system can use, e.g., an audio encoder neural network, e.g., an audio encoder neural network that has been trained jointly with a decoder neural network as part of a neural audio codec (e.g. SoundStream, Zeghidour et al., arXiv: 2107.03312vl, 2021), to generate one or more embeddings that encode the audio.
[0137] When the conditioning input is a scalar value, the system can use, e.g., an embedding matrix to map the scalar value or a one-hot representation of the scalar value to an embedding.
[0138] In some cases, the conditioning input includes multiple different types of inputs, e.g., two or more of text, images, scalar values, or context embeddings.
[0139] In some of these cases, the system can generate one or more initial embeddings for each of the different types of inputs, i.e., using an appropriate encoder neural network as described above, and then process the initial embeddings for all of the different types of inputs using a Transformer encoder neural network to update each of the initial embeddings to generate a set of final embeddings. The one or more cross-attention layers within the diffusion neural network can then cross-attend into the set of final embeddings.
[0140] In others of these cases, different cross-attention layers within the diffusion neural network can cross-attend into embeddings of different types of conditioning inputs.
[0141] In yet others of these cases, the system can concatenate the initial embeddings of the different types of inputs along the sequence dimension and then the one or more crossattention layers can cross-attend into the concatenated set of final embeddings.
[0142] As another example, the diffusion neural network can include one or more other types of neural network layers that are conditioned on the one or more embeddings. Examples of such layers include Feature-wise Linear Modulation (FiLM) layers, layers with conditional gated activation functions, and so on.
[0143] The diffusion input at any given updating iteration can also include data defining a noise level for the iteration. Generally, each updating iteration has a corresponding time step t and the noise level for the iteration depends on the time step. For example, the noise level can be a decreasing function of the time step t. Examples of such functions include a linear function, a cosine function, and a sigmoid function. In these cases, data identifying the noise level, the time step, or both can be embedded using an appropriate neural network, e.g., a multi-layer perceptron (MLP) and used to condition the diffusion neural network 110 as described above for the conditioning input.
[0144] Moreover, as described above, at each updating iteration, the system uses the denoising output generated by the diffusion neural network to update the current data item as of the updating iteration.
[0145] For example, the system can determine an initial estimate of the final data item using the denoising output and then apply an appropriate diffusion sampler to the initial estimate to update the current data item.
[0146] As another example, the system can use classifier-free guidance or negative guidance to adjust the denoising output, determine an initial estimate of the final data item using the adjusted denoising output and then apply an appropriate diffusion sampler to the initial estimate to update the current data item. Classifier-free guidance is described in, for example, Ho and Salimans, arXiv:2207.12598.
[0147] The system can use any appropriate diffusion sampler to update the data item, e.g., the DDPM (Denoising Diffusion Probabilistic Model) sampler, the DDIM (Denoising Diffusion Implicit Model) sampler or another appropriate sampler, to the estimate to generate the updated current data item. DDPMs are, for example, discussed in Ho et al. arXiv:2006: 11239.
[0148] When the denoising output is a prediction of the data item, the system can directly use the denoising output (or the adjusted denoising output) as the initial estimate.
[0149] When the denoising output is a prediction of the noise component, the system can determine the initial estimate from the current data item, the denoising output, and the noise level for the current updating iteration, e.g., by combining the current data item and the denoising output in accordance with the noise level for the current updating iteration.
[0150] Optionally, after the last iteration, the system can refrain from using the diffusion sampler and can instead use the initial estimate as the updated current data item.
[0151] After the last updating iteration, the system outputs the current data item as the final output data item. As described above, when the data items are spectrograms in the output space, the system can use the final data item as the spectrogram generated using the diffusion neural network. When the data items are in the latent space, the system can map the final data item to a spectrogram by processing the final data item using the decoder neural network. FIG. 4 is a flow diagram of an example process 400 for generating an initial musical composition or soundscape. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, an audio generation system, e.g., the audio generation system 100 depicted in FIG. 1, appropriately programmed in accordance with this specification, can perform the process 400.
[0152] The system obtains an audio input (step 402).
[0153] Generally, the audio input is text that characterizes desired properties of the musical composition or soundscape. For example, the audio input can be unstructured natural language text or can be structured text that specifies respective values for one or more attributes of the composition or soundscape.
[0154] As described above, the audio input can include (i) a positive audio input, (ii) a negative audio input, and optionally (iii) lyrics for the musical composition.
[0155] In some implementations, the system receives the audio input directly from a user, while in other implementations, the system generates the audio input, e.g., using a language model neural network, from an original user input.
[0156] The system generates a low-resolution spectrogram that spans a first time window conditioned on the audio input (step 404). For example, the system can generate this low- resolution spectrogram using a base diffusion neural network as described above with reference to FIG. 2.
[0157] The system generates a high-resolution spectrogram that spans the first time window conditioned on the low-resolution spectrogram and the audio input (step 406). For example, the system can generate this high-resolution spectrogram using an upsampling diffusion neural network as described above with reference to FIG. 2.
[0158] Optionally, the system can then generate a waveform from the high-resolution spectrogram (step 408). That is, in some cases the first time window spanned by the high- resolution spectrogram matches the target length for the musical composition or soundscape. In these cases, the system can generate the waveform using a vocoder and provide the waveform as the final output of the system.
[0159] In some other cases, the first time window is shorter than the target length. In these cases, the system does not need to generate the waveform and can instead use the high- resolution spectrogram to iteratively extend the length of the musical composition or soundscape. FIG. 5 is a flow diagram of an example process 500 for generating a longer musical composition or soundscape. For convenience, the process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, an audio generation system, e.g., the audio generation system 100 depicted in FIG. 1, appropriately programmed in accordance with this specification, can perform the process 500.
[0160] The system obtains an audio input and an initial spectrogram that spans a first time window (step 502) that is shorter than the target time window spanned by the final musical composition or soundscape.
[0161] The system generates an initial extended spectrogram that spans the target time window from the audio input and the initial spectrogram (step 504). In some implementations, the initial extended spectrogram has a resolution that is lower than both the first resolution and the target resolution.
[0162] The system iteratively extends the initial spectrogram conditioned on the audio input and the initial extended spectrogram (step 506). That is, the system can generate new portions of a final spectrogram at each of a plurality of iterations until the final spectrogram reaches the target length, i.e., until the final spectrogram spans the target time window.
[0163] FIG. 6 is a flow diagram of an example process 600 for generating an audio input and an image input. For convenience, the process 600 will be described as being performed by a system of one or more computers located in one or more locations. For example, an audio generation system, e.g., the audio generation system 100 depicted in FIG. 1, appropriately programmed in accordance with this specification, can perform the process 600.
[0164] The system receives a user input (step 602). For example, the user input can be a natural language input that describes a musical composition or soundscape. As a particular example, the user input can be free-form natural language text that has been submitted through a user interface presented by the system on a user device.
[0165] The system generates an input sequence from the user input (step 604). In some implementations, the input sequence includes only the text tokens in the user input while, in some other implementations, the input sequence includes additional text in addition to the user input. For example, the input sequence can include a natural language instruction or a few-shot prompt.
[0166] The system processes the input sequence using a language model neural network to generate an output sequence (step 606).
[0167] The output sequence generally includes one or more subsequences. In particular, one or more of the subsequences define an audio input.
[0168] That is, the language model neural network has been configured, e.g., through finetuning, by virtue of the natural language instruction or few-shot prompt in the input sequence, or both to generate an output sequence that includes different portions that define a structured input to an audio generation neural network.
[0169] The input is referred to as “structured” because it adheres to a particular format or structure rather than being freeform natural language text. For example, the structured input can be structured to include a positive prompt and lyrics and, optionally, a negative prompt.
[0170] The positive prompt can be structured to include respective values for each of multiple properties that the to-be-generated audio should have. Examples of properties include genre, audio quality, style, tempo, rhythm, pitch, timbre, dynamics, theme, instruments used, year or other unit of time that the audio was generated, and so on. Examples of properties are described below with reference to FIGS. 7A-7C.
[0171] The negative prompt can be structured to include respective values for each of multiple properties that the audio should not have. The properties specified in the negative prompt can be the same as or different from the properties in the positive prompt.
[0172] The system processes the audio input using an audio generative neural network system to generate audio that is described by the user input (step 608).
[0173] Optionally, the system can also generate an image input from the output sequence, e.g., by using only the positive audio input in combination with a predetermined prompt or by using the positive audio and the negative audio input in combination with the predetermined prompt.
[0174] The system processes the image input using an image generative neural network system to generate an image that describes the generated audio (step 610). In particular, when the generated audio is a musical composition, as described above, the generated image can be an “album photo” that visually depicts certain properties of the generated musical composition.
[0175] FIGS. 7A-7C show an example 700 of the operation of the system. In particular, FIGS. 7A-7C show examples of user interfaces provided by the system that allow a user to submit inputs that result in music being generated.
[0176] As shown in FIG. 7A, a user submits an initial query in natural or free-form language (here: “Song about the weather in London”) into a query field (“Describe the music you want to hear”). As shown in FIG. 7B, the user selects “Generate prompt.” In response, the system uses the language model neural network to generate a positive prompt (“List elements to include”), a negative prompt (“list elements to exclude (optional)”), and lyrics (“Lyrics (optional)”). These fields are referred to collectively as “detailed prompts”.
[0177] In this case, the positive prompt is “London Weather, Blur, britpop, electric guitars, pop, 1998, HQ, pristine quality, Remastered 2023”, the negative prompt is “London Calling, The Clash, speech, audiobook, podcast, low quality”, and the lyrics are “There’s a cold wind blowing in the streets of London\nAnd I don’t know if I’ll ever get warm again\nThe rain is falling and the sky is grey\nAnd I’m wishing that I was back at home”.
[0178] Note that users are also able to directly input positive prompts, negative prompts, and lyrics or to edit the outputs of the language model through the user interface.
[0179] As shown in FIG. 7C, the user selects “Generate.” In response, the system generates a musical composition or soundscape and can generate a corresponding album photo using the image and audio generative models as described above and using the detailed prompts.
[0180] Note that if the advanced fields are hidden (e.g., by clicking on “Advanced”), the user can generate detailed prompts, images, and audio with a single input by selecting “Generate” directly. That is, the system can use the language model neural network to generate the detailed prompts and then generate the music or soundscape and images from the detailed prompts without surfacing the detailed prompts to the user.
[0181] Moreover, because of the use of diffusion models as part of both the image and audio generative neural networks, the system can generate multiple different plausible musical compositions or soundscapes and optionally images from the same detailed prompts by leveraging the stochasticity of the generation process, e.g., by sampling different noise. The system can display all of these different album covers and compositions or soundscapes to the user, allowing the user to view and listen to a set of multiple plausible, high-quality interpretations of their query.
[0182] Implementations of the described system have applications that extend beyond merely generating a musical composition. For example, as described above the system 130 can iteratively extend the length of the generated musical composition or soundscape until the target length is reached or in principle indefinitely. This facilitates using the generated musical composition or soundscape in therapeutic applications. As one particular example, the generated musical composition or soundscape can be used to mask tinnitus by providing the musical composition or soundscape continuously as background noise to a patient, e.g. through headphones or earbuds. As another example the generated musical composition or soundscape can be used to create privacy by sound masking, i.e. the generated musical composition or soundscape can be played in a public or other environment, such as an office or hospital, to mask private conversations taking place.
[0183] In some implementations of the described system is used to evaluate, calibrate, test, or modify an audio electronic device or audio communications system such as a mobile phone, smart speaker, video conferencing device or system, or an audio signal transmission system. As one example the described system can be used to generate music or a soundscape that is captured, transmitted, or played by the audio device or audio communications or signal transmission system, and an output of the audio device or audio communications or signal transmission system can then be compared with the generated music or soundscape to evaluate the fidelity of the output. Optionally the audio device or audio communications or signal transmission system can then be modified, e.g. calibrated or trained, to increase the fidelity. As another example, a generated soundscape can be combined with speech, and the combination captured, transmitted, or played by an audio device or audio communications or signal transmission system that is configured to filter out a background soundscape, e.g. noise from an office, shopping center or mall or other location. The output of the audio device or audio communications or signal transmission system can then be compared with the input speech and / or soundscape to evaluate the effectiveness of the filtering. Optionally the audio device or audio communications or signal transmission system can then be modified, e.g. calibrated or trained, to increase the effectiveness of the filtering (e.g. as measured by attenuation of an unwanted component of the signal, i.e. the soundscape).
[0184] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0185] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, e.g., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0186] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0187] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0188] In this specification, the term “database” is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.
[0189] Similarly, in this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0190] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0191] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0192] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.
[0193] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0194] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, e.g., inference, workloads.
[0195] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework or a Jax framework.
[0196] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0197] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0198] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0199] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0200] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
CLAIMS1. A method performed by one or more computers, the method comprising: obtaining an audio input characterizing a musical composition or soundscape having a target resolution; generating, from the audio input and using a base diffusion neural network, a first spectrogram having a first resolution that is lower than the target resolution; and generating, from the first spectrogram and using a first upsampling diffusion neural network, a second spectrogram having the target resolution.
2. The method of claim 1, further comprising: generating an audio waveform from the second spectrogram; and providing the audio waveform for playback.
3. The method of claim 1, wherein the second spectrogram spans a first time window, the method further comprising: iteratively extending the second spectrogram having the target resolution using one or more additional diffusion neural networks to generate a final spectrogram spanning a second, longer time window.
4. The method of claim 3, further comprising generating an audio waveform from the final spectrogram; and providing the audio waveform for playback.
5. The method of claim 3 or claim 4, wherein iteratively extending the second spectrogram having the target resolution using one or more additional diffusion neural networks to generate a final spectrogram spanning a second, longer time window comprises: generating, from the audio input and the second spectrogram, an initial extended spectrogram that spans the second, longer time window but has a lower resolution than the target resolution; and iteratively extending the second spectrogram using the initial extended spectrogram.
6. The method of claim 5, wherein generating, from the audio input and the second spectrogram, an initial extended spectrogram that spans the second, longer time window but has a lower resolution than the target resolution comprises:generating the initial extended spectrogram using an initial low-resolution extension diffusion neural network.
7. The method of claim 5 or claim 6, wherein iteratively extending the second spectrogram using the initial extended spectrogram comprises, at each of a plurality of iterations: generating a new high-resolution spectrogram conditioned on (i) a low-resolution spectrogram that was used to generate a most-recently generated high-resolution spectrogram and (ii) a corresponding portion of the initial extended spectrogram.
8. The method of claim 7, wherein, at the first iteration, the low-resolution spectrogram that was used to generate a most-recently generated high-resolution spectrogram is the second spectrogram.
9. The method of claim 7 or claim 8, wherein, at each subsequent iteration that is after the first iteration, the most-recently generated high-resolution spectrogram is the new high- resolution spectrogram generated at the preceding iteration.
10. The method of any one of claims 7-9, wherein the corresponding portion of the initial extended spectrogram comprises a portion of the initial extended spectrogram that spans a same time window as the new high-resolution spectrogram.
11. The method of any one of claims 7-10, wherein generating a new high-resolution spectrogram conditioned on (i) a low-resolution spectrogram that was used to generate a most-recently generated high-resolution spectrogram and (ii) a corresponding portion of the initial extended spectrogram comprises: generating the new high-resolution spectrogram using a low-resolution extension diffusion neural network and a second upsampling diffusion neural network.
12. The method of claim 11, wherein the low-resolution extension diffusion neural network is the base diffusion neural network and the second upsampling diffusion neural network is the first upsampling diffusion neural network.
13. The method of any preceding claim, wherein obtaining an audio input characterizing a musical composition or soundscape having a target resolution comprises: receiving a user input characterizing the musical composition or soundscape; andprocessing the user input using a language model neural network to generate the audio input.
14. The method of claim 13, wherein the audio input comprises one or more of a positive prompt, a negative prompt, or lyrics of the musical composition.
15. The method of any of claims 1-14, wherein generating, from the first spectrogram and using a first upsampling diffusion neural network, a second spectrogram having the target resolution comprises: up-sampling the first spectrogram to generate an upsampled spectrogram having the target resolution; and conditioning the first upsampling diffusion neural network on the upsampled spectrogram.
16. The method of any preceding claim, wherein generating, from the first spectrogram and using a first upsampling diffusion neural network, a second spectrogram having the target resolution comprises: conditioning the first upsampling diffusion neural network on the audio input.
17. The method of any preceding claim, further comprising: processing an image input characterizing the musical composition or soundscape using an image generation neural network to generate an image that characterizes the musical composition or soundscape.
18. A method performed by one or more computers, the method comprising: obtaining an initial spectrogram spanning a first time window; and iteratively extending the initial spectrogram using one or more diffusion neural networks to generate a final spectrogram that spans a longer time window.
19. The method of claim 18, further comprising: generating an audio waveform from the final spectrogram; and providing the audio waveform for playback.
20. The method of claim 18 or claim 19, wherein the initial spectrogram was generated from an audio input and has a target resolution, the final spectrogram has the target resolution, and iteratively extending the initial spectrogram having the target resolution usingone or more diffusion neural networks to generate a final spectrogram spanning a longer time window comprises: generating, from the audio input and the initial spectrogram, an initial extended spectrogram that spans the longer time window but has a lower resolution than the target resolution; and iteratively extending the initial spectrogram using the initial extended spectrogram.
21. The method of claim 20, wherein generating, from the audio input and the initial spectrogram, an initial extended spectrogram that spans the longer time window but has a lower resolution than the target resolution comprises: generating the initial extended spectrogram using an initial low-resolution extension diffusion neural network.
22. The method of claim 19 or claim 20, wherein iteratively extending the initial spectrogram using the initial extended spectrogram comprises, at each of a plurality of iterations: generating a new high-resolution spectrogram conditioned on (i) a low-resolution spectrogram that was used to generate a most-recently generated high-resolution spectrogram and (ii) a corresponding portion of the initial extended spectrogram.
23. The method of claim 22, wherein, at the first iteration, the low-resolution spectrogram that was used to generate a most-recently generated high-resolution spectrogram is the initial spectrogram.
24. The method of claim 22 or claim 23, wherein, at each subsequent iteration that is after the first iteration, the most-recently generated high-resolution spectrogram is the new high- resolution spectrogram generated at the preceding iteration.
25. The method of any one of claims 22-24, wherein the corresponding portion of the initial extended spectrogram comprises a portion of the initial extended spectrogram that spans a same time window as the new high-resolution spectrogram.
26. The method of any one of claims 22-25, wherein generating a new high-resolution spectrogram conditioned on (i) a low-resolution spectrogram that was used to generate a most-recently generated high-resolution spectrogram and (ii) a corresponding portion of the initial extended spectrogram comprises:generating the new high-resolution spectrogram using a low-resolution extension diffusion neural network and an upsampling diffusion neural network.
27. A method performed by one or more computers, the method comprising: processing a user input using a language model neural network to generate a structured audio input for an audio generation neural network; and processing the structured audio input using the audio generation neural network to generate an output that defines the musical composition or soundscape.
28. The method of any preceding claim, when dependent on claim 5, wherein the initial extended spectrogram has a resolution that is lower than both the first resolution and the target resolution.
29. A system comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the respective operations of any one of claims 1-28.
30. One or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the respective operations of the method of any one of claims 1-28.
Citation Information
Patent Citations
Autoencoder-based lyric generation
CA3132537A1