Generating music from images using generative neural networks
By using pre-trained generative neural networks and few-shot cues, the problem of image and audio embedding spatial alignment, which consumes a lot of resources in existing technologies, is solved, achieving high-quality image-matching music generation, improving user experience and film creation efficiency.
Patent Information
- Application Number
- CN202580003803.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-13
- Filing Date
- 2025-05-13
- Publication Date
- 2026-02-17
AI Technical Summary
Existing audio generation systems require significant resources to align the embedding space of images and audio, and struggle to quickly generate high-quality music that matches the visual content of a given image.
Using a pre-trained generative neural network and few-shot cues, images are processed by the generative neural network to generate musical explanations, and an audio generative neural network is used to generate audio signals suitable for the images, avoiding alignment training in the embedding space.
It enables high-quality, rapid generation of music that matches images, providing an enhanced user experience, especially suitable for visually impaired users and the filmmaking process, and simplifies the music generation process.
Smart Images

Figure CN121548854A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Utility Model Application No. 18 / 662,913, filed May 13, 2024. The disclosure of the prior application is considered part of the disclosure of this application and is incorporated herein by reference. Background Technology
[0003] This manual relates to the use of neural networks to generate audio.
[0004] A neural network is a machine learning model that uses one or more layers of non-linear units to predict the output from a given input. In addition to the output layer, some neural networks also include one or more hidden layers. The output of each hidden layer serves as the input to the next layer in the network (i.e., the next hidden layer or output layer). Each layer of the network generates its output from the received input based on the current values of its corresponding set of parameters. Summary of the Invention
[0005] This specification describes a system implemented as a computer program on one or more computers at one or more locations, which uses one or more generative neural networks to generate audio signals including music and / or soundscapes that sound appropriate to a given image. For example, as a particular example, the system can generate musical works conditioned on a given image, such as songs with lyrics or instrumental pieces without lyrics. As another particular example, the system can generate soundscapes conditioned on images that represent a specific audio environment (e.g., the soundscape could represent an audio environment reflecting a detailed description of the environment in the image).
[0006] Generally, the output audio signal is an output audio sample that includes samples of the audio wave at each output time step in a sequence of output time steps spanning a specified time window. For example, the output time steps can be arranged at regular intervals within the specified time window.
[0007] The audio sample at a given output time step can be the amplitude value of an audio wave, or the amplitude value that has been compressed, compressed, or both. For example, the audio sample can be the original amplitude value or a mu-law compressed representation of the amplitude value.
[0008] In general, an innovative aspect of the subject matter described in this specification can be embodied in a method comprising the actions of: receiving an input image; processing the input image using one or more generative neural networks to generate an audio description (e.g., a music and / or soundscape description) that describes one or more audio features corresponding to the input image; and processing the audio (e.g., music and / or soundscape) description using an audio generative neural network to generate an audio signal described by the audio (e.g., music and / or soundscape) description.
[0009] In some implementations, using one or more generative neural networks to process an input image to generate an audio (e.g., music and / or soundscape) description describing one or more audio features corresponding to the input image includes: using a first generative neural network to process the input image to generate an image description describing the input image; and using the first generative neural network to process a network input that includes at least the image description to generate an audio (e.g., music and / or soundscape) description describing one or more audio features.
[0010] In some implementations, using a first generative neural network to process the input image to generate an image description of the input image includes providing the input image and a request to describe the content of the input image as input to the first generative neural network.
[0011] In some implementations, using a first generative neural network to process network input that includes at least an image explanatory description to generate an audio (e.g., music and / or soundscape) explanatory description describing one or more audio features includes: providing network input and a request to rewrite the image explanatory description into an audio (e.g., music and / or soundscape) explanatory description as input to the first generative neural network.
[0012] In some implementations, the request further includes one or more examples, each example including an example image explanation and a corresponding example audio (e.g., music and / or soundscape) explanation.
[0013] In some implementations, the network input further includes an input image, and the use of a first generative neural network to process the network input, which includes at least an image explanatory description, to generate an audio (e.g., music and / or soundscape) explanation describing one or more audio features includes: providing network input and a request to rewrite the image explanatory description into an audio (e.g., music and / or soundscape) explanation for the input image as input to the first generative neural network.
[0014] In some implementations, the request further includes one or more examples, each example including an example image, a corresponding explanation of the example image, and a corresponding explanation of the example audio (e.g., music and / or soundscape).
[0015] In some implementations, using one or more generative neural networks to process an input image to generate an audio (e.g., music and / or soundscape) explanation describing one or more audio features corresponding to the input image includes: using a second generative neural network to process the input image to generate an image explanation describing the input image; and using a third generative neural network to process a network input that includes at least the image explanation to generate an audio (e.g., music and / or soundscape) explanation describing one or more audio features.
[0016] In some implementations, a second generative neural network is used to process the input image to generate an image description of the input image, including providing the input image and a request to describe the content of the input image as input to the second generative neural network.
[0017] In some implementations, using a third generative neural network to process network input that includes at least image explanatory descriptions to generate audio (e.g., music and / or soundscape) explanatory descriptions describing one or more audio features includes: providing network input and a request to rewrite the image explanatory descriptions into audio (e.g., music and / or soundscape) explanatory descriptions as input to the third generative neural network.
[0018] In some implementations, the request further includes one or more examples, each example including an example image explanation and a corresponding example audio (e.g., music and / or soundscape) explanation.
[0019] In some implementations, audio generative neural networks are configured to generate audio signals conditioned on at least the text.
[0020] In some implementations, receiving an input image includes receiving an input image from a user.
[0021] In some implementations, the method further includes providing an audio signal to be presented to the user.
[0022] In some implementations, one or more audio features describe any one or more of the following: style, rhythm, timing, pitch, mood, or instrument.
[0023] In some implementations, the audio explanation is a musical explanation. In other implementations, the audio explanation is a soundscape explanation.
[0024] Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each of which is configured to perform the actions of these methods.
[0025] Specific embodiments of the subject matter described in this specification may be implemented in order to achieve one or more of the following advantages.
[0026] The system described in this specification provides high-quality music generation that sounds appropriate for a given image. For example, for an image depicting swans on a calm lake, the system can generate classical-style music. For an image depicting a bustling street in the downtown area of a large city, the system can generate more intense, fast, and lively music.
[0027] To generate music from an input image, the system can use one or more generative neural networks to process the image to generate a musical description. The musical description describes one or more audio features corresponding to the image. The system can then use an audio generative neural network to process the musical description to generate an audio signal described by the musical description. Therefore, the system can use information from the musical description to generate music suitable for the image. By generating a musical description describing the audio features corresponding to the image, the system can provide the audio generative neural network with the musical description describing the audio features, thereby producing music that conforms to the audio features of the musical description. The musical description has audio-feature-specific details that the system can use to generate high-quality and diverse music. Therefore, the system can use the musical description to generate music of higher quality than music generated by, for example, using an aligned embedding space for the image and audio, or by providing an image description to the audio generative neural network.
[0028] Conventional systems for generating audio suitable for a given image may require alignment of the image and audio embedding spaces, which can demand significant resources for training and obtaining parallel training data for alignment. The system described in this specification can generate music suitable for a given image using a pre-trained generative neural network.
[0029] The system can also use few-shot cues to improve the performance of generative neural networks without having to further train the generative neural network.
[0030] The system described in this specification can be used to provide different user experiences for interacting with visual content such as visual art. For example, the system can generate music that evokes the same atmosphere, tone, and / or emotion as a given artwork, thereby allowing a user to experience the artwork not only visually or, alternatively, aurally. Therefore, the system can enable users such as visually impaired users to experience visual content.
[0031] Furthermore, the system described in this specification can be used to generate suitable music for videos such as those from movie scenes or for still images from movie scenes. For example, the system can generate music descriptions for frames of movie scenes. The system can process the music descriptions to generate audio signals described by the music descriptions. Therefore, the system allows users to add suitable music to videos without having to manually compose, search for, or create music—which may be difficult or impractical for some users. The system can also generate music suitable for videos faster than manually composing, searching for, or creating music, allowing users to easily experience different generated music paired with videos and use the generated music as inspiration during the creative process. Therefore, the system can provide a more efficient user experience for the film creation process.
[0032] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of this subject matter will become apparent from the description, drawings, and claims. Attached Figure Description
[0033] Figure 1 This is a block diagram of an example audio generation system.
[0034] Figure 2 This is a block diagram of another example audio generation system.
[0035] Figure 3 This is a block diagram of another example audio generation system.
[0036] Figure 4 An example of a few-sample hint is shown.
[0037] Figure 5 Example images and accompanying music are shown.
[0038] Figure 6 This is a flowchart of an example process for generating an audio signal given an image.
[0039] In the various figures, the same reference numerals and names indicate the same elements. Detailed Implementation
[0040] Figure 1This is a block diagram of an example audio generation system 100. Audio generation system 100 is an example of a system implemented as a computer program on one or more computers at one or more locations, wherein the systems, components, and techniques described below are implemented.
[0041] Given an input image 102, the audio generation system 100 generates a prediction of an audio signal 104. The audio signal 104 includes corresponding audio samples at each of a plurality of output time steps spanning a time window. The audio signal 104 may include music suitable for or reflecting the input image 102. Alternatively or additionally, the audio signal 104 may include a soundscape suitable for or reflecting the input image 102. A “soundscape” may include audio characterizing a specific environment such as a home, office, hospital, shopping mall or shopping center, beach, or forest (e.g., the environment shown or reflected in the image).
[0042] To generate audio, system 100 receives input image 102. Input image 102 comprises a plurality of pixels, each having one or more intensity values, such as one or more intensity values comprising the RGB color values of each pixel of an image or other color values in another colorization scheme. Input image 102 can be a real-world image or a synthetically generated image. In some examples, input image 102 can depict a physical artwork. Figure 1 In the example, input image 102 depicts a sea turtle in the ocean.
[0043] In some examples, system 100 receives input image 102 from a user. For instance, system 100 may receive input image 102 from a user through a user interface on a user device.
[0044] System 100 processes input image 102, for example, by processing the intensity values of pixels in input image 102, to generate an audio description (e.g., a music description and / or a soundscape description) that describes one or more audio features corresponding to input image 102. The audio description (e.g., a music description and / or a soundscape description) may include a sequence of natural language text describing one or more audio features. For example, audio features (e.g., music features) may include style, rhythm, melody, timing, pitch, mood, and / or instruments suitable for the content and / or atmosphere of input image 102. For example, audio features (e.g., soundscape features) may include environment, sound source, sound type, timing, rhythm, pitch, mood, and / or spatial characteristics suitable for the content and / or atmosphere of input image 102.
[0045] exist Figure 1In the example, the musical description used for input image 102 includes "Flowing instrumentalpiece featuring a combination of piano and strings, conveying a sense of grace and tranquility." See below for reference. Figure 4 Examples of music descriptions and input images.
[0046] exist Figure 1 In the example, the soundscape description for input image 102 could include “a gentle interplay of soft, low-frequency biological and hydroacoustic sounds, characterized by the rhythmic brushing of flippers against the backdrop of the murmuring ocean, creating an atmosphere of calm and deliberate movement within a spatially encompassing aquatic environment”.
[0047] Although some embodiments generate soundscapes based on soundscape interpretation, for the sake of brevity, embodiments related to music generation are described below.
[0048] In some implementations, system 100 may use a generative neural network to process the input image to generate musical narration, as shown in the following reference. Figure 2 As stated above.
[0049] In some other implementations, system 100 may use more than one generative neural network to process the input image to generate musical narration, as shown in the following references. Figure 3 As stated above.
[0050] System 100 processes at least the music interpretation description to generate an audio signal 104 described by the music interpretation description. Audio signal 104 includes music that can be described by the music interpretation description as “flowing instrumental piece featuring a combination of piano and strings, conveying a sense of grace and tranquility.” Because the music interpretation description describes audio features corresponding to the input image 102, audio signal 104 includes music suitable for the input image 102.
[0051] For example, system 100 can use an audio generative neural network to generate audio signal 104. The audio generative neural network is configured to generate audio signals conditioned on at least text. See below for reference. Figure 2 A further detailed description of the example audio generative neural network.
[0052] In some examples, system 100 provides an audio signal 104 to be presented to the user. For example, system 100 may provide data representing the audio signal 104 to the user device and cause the audio signal 104 to be played. Thus, system 100 may allow the user to experience the image 102 aurally, thereby providing different user experiences for interacting with the image 102.
[0053] In some examples, system 100 provides audio signal 104 to be presented to the user (e.g., through one or more speakers) while image 102 is presented to the user (e.g., through a display). Thus, system 100 can allow the user to experience image 102 both visually and aurally, thereby providing an enhanced user experience of interacting with image 102 compared to the user experience of simply viewing image 102.
[0054] Figure 2 It is for reference. Figure 1 A block diagram of an example audio generation system 100 is described. Specifically, the audio generation system 100 uses a generative neural network 210 and an audio generative neural network 220 to generate predictions of an audio signal 104 given an input image 102.
[0055] The system receives input image 102, as referenced above. Figure 1 As stated above.
[0056] The system uses a generative neural network 210 (also known as a first generative neural network) to process the input image 102 to generate an image description 212. The image description 212 is a sequence of natural language text describing the input image 102. For example, the image description 212 describes the content of the input image 102, such as descriptions of visual details, subject, background, context, mood, tone, and / or atmosphere.
[0057] To generate the image caption 212, system 100 may provide an input image 102 and a request to describe the content of the input image as input to the generative neural network 210. For example, the request may include a natural language request describing what can be seen in the input image. In some examples, the request may include a request to describe what can be seen in the input image in as much detail as possible. In some examples, the request may also specify the type of features that the image caption can describe. The generative neural network 210 is further described in detail below.
[0058] As an example, the image caption 212 could include: “A sea turtle floats calmly in clear turquoise ocean waters near the ocean floor. The ocean floor is teeming with marine life such as fish and plants. The sea turtle seems at peace with its surroundings.”
[0059] In some examples, the request to describe what can be seen in the input image includes one or more examples as few-sample cue examples. For instance, the request may include a request to describe what can be seen in the input image based on examples. In some examples, the request may include a natural language request to describe what can be seen in the input image based on examples. Each example may include an example image and a corresponding explanation of the example image. See below for reference. Figure 4 Describe the example image and provide an explanation of the image.
[0060] System 100 includes an image explanation 212 in the network input 214. System 100 uses a generative neural network 210 to process the network input 214 to generate a music explanation 216.
[0061] To generate the music description 216, system 100 can provide network input 214 and a request to rewrite the image description as a music description as input to the generative neural network 210. For example, the request may include a natural language request describing what a suitable musical accompaniment for the image description would sound like. In some examples, the natural language request may also include a definition of the music description, such as a definition of the types of audio features that the music description can describe. The generative neural network 210 is further described in detail below.
[0062] Music Explanation 216 is a reference to the above text. Figure 1 The description provides examples of musical interpretation and describes one or more audio features corresponding to the input image 102. For example, audio features may include style, rhythm, melody, timing, pitch, mood, and / or instrumentation.
[0063] exist Figure 2 In the example, musical description 216 includes "Flowing instrumental piece featuring a combination of piano and strings, conveying a sense of grace and tranquility." Musical description 216 describes characteristics such as instruments, style, and mood.
[0064] In some examples, the request to rewrite an image description as a musical description includes one or more examples as few-sample prompt examples. For instance, the request might include a request to rewrite an image description as a musical description based on an example. In some examples, the request might include a natural language request to rewrite an image description as a musical description based on an example. Each example in the examples might include an example image description and a corresponding example musical description. See below for reference. Figure 4 Describe the example images and music.
[0065] In some examples, system 100 also includes input image 102 in network input 214. In these examples, to generate music description 216, system 100 may provide network input 214 and a request to rewrite the image description into a music description for the input image as input to the generative neural network 210. For example, the request may include a natural language request describing what a suitable musical accompaniment for the image description would sound like. In some examples, the natural language request may also include a definition of the music description, such as a definition of the types of audio features that the music description can describe.
[0066] In some examples, the request may also include one or more examples as few-sample prompt examples. For instance, the request may include a request to rewrite an image description as a musical description based on an example. In some examples, the request may include a natural language request to rewrite an image description as a musical description based on an example. Each example may include an example image, a corresponding example image description, and a corresponding example musical description. See below for reference. Figure 4 Describe the example image, image description, and music description.
[0067] System 100 uses an audio generative neural network 220 to process the music interpretation description 216 to generate an audio signal 104. The audio generative neural network 220 is configured to perform text-conditional music generation, such that the generated audio is music described by the input text data. Therefore, the audio signal 104 includes music that can be described as “flowing instrumental piece featuring a combination of piano and strings, conveying asense of grace and tranquility.” An example audio generative neural network 220 is further described below.
[0068] The generative neural network 210 is configured to perform machine learning tasks involving manipulating inputs and / or generating outputs, such as image processing tasks and / or text processing tasks. For example, the image processing task to be performed can be specified by words in natural language or computer language.
[0069] For example, if the input includes an input image 102 and a request to describe the content of the input image 102, the output includes text describing the content of the input image. If the input includes an image description 212 and a request to rewrite the image description as a musical description, the output includes text describing the audio features corresponding to the image description. If the input includes an image description 212, an input image 102, and a request to rewrite the image description as a musical description of the input image, the output includes text describing the audio features corresponding to the image description and the input image.
[0070] Generative neural network 210 processes sequences of input data elements derived from input image data and / or input text, such as image embeddings and / or text terms, to generate an output that includes text. As used herein, an "embedding" is a sequence of one or more vectors of numerical values (e.g., floating-point values or other values), each vector having a predetermined dimension. For example, generative neural network 210 may use a language model neural network to process the sequence of input data elements to generate an output.
[0071] Language model neural networks can have any suitable architecture for processing text and / or images. Language model neural networks can have any suitable Transformer-based encoder-only Transformer architecture, encoder-decoder Transformer architecture, decoder-only Transformer architecture, or another attention-based architecture.
[0072] Generally, a Transformer-based architecture can be represented by a series of self-attention neural network layers. Each self-attention neural network layer has an attention layer input for each element of the input and is configured to apply an attention mechanism to the attention layer input to generate an attention layer output for each element of that input. Many different attention mechanisms are available.
[0073] As an example, the generative neural network 210 may include a visual language model neural network. A very large (but potentially noisy) dataset can be used to train the visual language model neural network, where text is paired with images and / or with one or more other types of data, such as audio data or data related to the actions of an agent acting in an environment to perform various tasks. Such a model can be trained, for example, using self-supervised learning. The pairings may often be imperfect, and the training dataset may include, but may not include, any actual examples of the specific task to be performed, but the ability to perform the specific task may emerge. Many examples of suitable publicly available training datasets exist.
[0074] Some example generative neural networks that can be used with the techniques described in this paper include: Flamingo (arXiv:2204.14198 by Alayrac et al.); ALIGN (arXiv:2102.05918 by Jia et al.); PaLI (arXiv:2209.06794 by Chen et al.) and PaLI-X (arXiv:2305.18565 by Chen et al.); and Gemini (“Gemini: A Family of Highly Capable Multimodal Models, Gemini Team, Google”). These citations also include indications of training datasets that can be used to train the respective models.
[0075] As a specific example, a visual language model neural network may include a lexical processing layer. The lexical processing layer can constitute a language model neural network trained on a large database of data (e.g., natural language data) such that when an input lexical string (e.g., an input lexical string representing input text and including a sequence of text lexical units from a vocabulary) is input to the first lexical processing layer, the output lexical string is an appropriate response. The visual language model neural network may also include gated cross-attention layers interleaved with the lexical processing layer. The gated cross-attention layers may apply an attention function to the image embeddings used for the input image and the input lexical string. In some examples, the input lexical string may also include "tags" (one or more lexical units, e.g., lexical units contained in the same vocabulary as the input lexical string) indicating the presence of the input image and optionally indicating location.
[0076] As used herein, an image can be any static or moving image, i.e., an image can be part of a 2D or 3D video, and can be a monochrome, color, or hyperspectral image (i.e., including monochrome or color pixels). As defined herein, an “image” includes, for example, a point cloud from a LiDAR system, and a “pixel” includes points in the point cloud. During or after (pre)training and / or fine-tuning, the image may have been captured, for example, by a real-world camera or other image sensor, and objects in the image or video may include physical objects represented by the image or video.
[0077] Generally, performing image processing tasks using a Visual Language Model (VLM) neural network, which includes a (trained) image encoder neural network, may involve feeding an image to the image encoder neural network to generate a representation of the image, specifically a representation of the image's features. The image representation is then processed to perform the image processing task. Techniques for processing image representations to perform various image processing tasks are generally well-known.
[0078] For example, system 100 can generate a representation of an image as a sequence of one or more embeddings. System 100 can divide an image into a sequence of patches (i.e., spatial regions) and use an image encoder neural network to generate a corresponding embedding for each patch. The system can then concatenate the corresponding embeddings of the image patches to generate a representation of the image as a sequence of embeddings.
[0079] In one example, the image encoder neural network can have a residual neural network architecture that includes a sequence of residual blocks, for example, where the input for each residual block is added to the output of the residual block. As another example, the image encoder neural network can have a convolutional neural network architecture that includes one or more convolutional layers. As yet another example, the image encoder neural network can have an attention-based neural network architecture. For example, the encoder neural network can include one or more self-attention neural network layers. The encoder neural network can repeatedly update the initial embedding of image patches, for example, using self-attention neural network layers, to generate a corresponding final embedding for each image patch. In a particular example, the encoder neural network can have a Vision Transformer architecture, for example, as described with reference to the following literature: A. Dosovitskiyyet al., “An image is worth 16x16 words: transformers for image recognition atscale,” arXiv:2010.11929v2, 2021.
[0080] Text can be received, for example, as a series of encoded characters (e.g., UTF-8 encoded characters); such "characters" can include Chinese characters and other similar characters, as well as logograms, syllabograms, and similar characters. The system can use a text encoder to represent the input text as a sequence of text words from a vocabulary of text words. The set of text words can include, for example, characters, n-grams, word fragments, words, or combinations thereof. In some examples, the system can map each text word to a corresponding numerical value according to a predefined mapping. In some examples, the system can represent each text word by a corresponding embedding, for example, according to a predefined mapping from numerical values to embeddings.
[0081] An example audio generative neural network 220 will now be described. For example, the audio generative neural network 220 can be configured to use one or more generative neural networks to generate audio signals.
[0082] For example, the system can provide a request to generate an audio signal with corresponding audio samples at each of multiple output time steps spanning a time window, conditioned on the input to the audio generative neural network 220.
[0083] The input may include, for example, text data containing musical explanations 216.
[0084] The audio generative neural network 220 uses an embedding neural network to process the input, mapping the input to one or more embedding terms. For example, the embedding neural network can be trained to map text and audio to a joint embedding space of text and audio, also known as a joint audio embedding space. The audio generative neural network 220 can use the embedding neural network to generate embedding vectors for the text data input in the joint embedding space. The audio generative neural network 220 can then quantize the embedding vectors to generate embedding terms.
[0085] For example, an embedding neural network can include a neural network that maps text input to embeddings, and a neural network that maps audio input to embeddings. In the joint embedding space of text and audio, both text and audio are mapped to embeddings in the same embedding space. That is, the embedding vectors for text and audio have the same dimension. Furthermore, embeddings that are close together in the joint embedding space imply that the embeddings share semantics intramodal and intermodal. For example, two embeddings that are close to each other can represent two semantically similar text sequences, two semantically similar audio samples, or audio samples and text sequences with semantically similar features. As an example, the embedding neural network could be a MuLAN model.
[0086] An audio generative neural network 220 generates a semantic representation of an audio signal 104. This semantic representation specifies a corresponding semantic unit at each of multiple first time steps spanning a time window. Each semantic unit is selected from a vocabulary of semantic units and represents the semantic content of the audio signal 104 at the corresponding first time step. Examples of semantic content that can be represented by semantic units include musical genres, melody, harmony, and rhythmic attributes.
[0087] The audio generative neural network 220 can use embedded words as conditional signals to autoregressively generate semantic representations using a semantic representation generative neural network. For example, the semantic representation generative neural network can conditionally generate semantic representations while simultaneously autoregressively, using embedded words as conditions. That is, the semantic words at each first time step can be conditional on the semantic words and embedded words used in the previous first time step. The semantic representation generative neural network can generate... ,in Semantic units representing the first time step ,and This represents embedded lexical units. For example, a semantic representation generative neural network can be trained to predict semantic representations generated from the output of one or more layers (e.g., one of the intermediate layers) of an audio representation neural network.
[0088] Audio representation neural networks can be trained to generate representations of input audio signals. For example, an audio representation neural network could be a w2v-BERT model that maps input audio signals to a set of language features.
[0089] The audio generative neural network 220 then uses one or more generative neural networks, conditioned at least on semantic representations, to generate an acoustic representation of the audio signal 104. The acoustic representation specifies a set of one or more corresponding acoustic terms at each of a plurality of second time steps spanning a time window. The one or more corresponding acoustic terms at each second time step represent the acoustic properties of the audio signal 104 at the corresponding second time step. The acoustic properties capture details of the audio waveform and allow for high-quality synthesis. Acoustic properties may include, for example, recording conditions such as reverberation level, distortion, and background noise.
[0090] For example, the audio generative neural network 220 can use both coarse-grained and fine-grained generative neural networks to generate acoustic lexical units for acoustic representations. The coarse-grained and fine-grained generative neural networks can be trained to predict acoustic representations generated from the output of the encoder neural network by processing audio signals.
[0091] For example, the encoder neural network could be a convolutional encoder that maps an audio signal to a sequence of embeddings. Each corresponding embedding at each of a plurality of second time steps could correspond to a feature of the audio signal at that second time step. A ground truth acoustic representation of the audio signal could be generated by applying quantization to each of the corresponding embeddings. For example, the encoder neural network could be part of a neural audio codec (such as the Soundstream Neural Audio Codec). For example, quantization could be residual vector quantization, which uses a hierarchy of multiple vector quantizers to encode each embedding, each of which generates a corresponding acoustic lexicon from a vocabulary corresponding to the acoustic lexicon used by that vector quantizer.
[0092] The set of one or more corresponding acoustic lexical units at each of the multiple second time steps comprises multiple acoustic lexical units that collectively represent a prediction of the output of the residual vector quantization applied to the embedding, the prediction representing the acoustic properties of the audio signal at the second time step. The residual vector quantization encodes the embedding using a hierarchy of multiple vector quantizers, each generating a corresponding acoustic lexical unit from a vocabulary of acoustic lexical units used by that vector quantizer. The hierarchy includes one or more coarse vector quantizers at one or more of the earliest positions in the hierarchy and one or more fine vector quantizers at one or more of the latest positions in the hierarchy. Therefore, for each vector quantizer, the set of acoustic lexical units at each second time step comprises a corresponding acoustic lexical unit selected from the vocabulary used by that vector quantizer.
[0093] For example, a hierarchy can include coarse-grained vector quantizers and fine-grained vector quantizers.
[0094] To generate acoustic representations, a coarse-grained generative neural network can generate acoustic lexics for a coarse-grained vector quantizer, conditioned at least on the semantic representation and embedded lexics. For example, for each of one or more coarse-grained vector quantizers in a hierarchy, the coarse-grained generative neural network can generate corresponding acoustic lexics for that vector quantizer at a second time step, conditioned at least on the semantic representation and embedded lexics.
[0095] A coarse-grained generative neural network can be an autoregressive neural network configured to autoregressively generate acoustic lexics for a coarse-grained vector quantizer according to a first generation order. In some implementations, the coarse-grained generative neural network has a decoder-only Transformer architecture. In some implementations, the coarse-grained generative neural network has an encoder-decoder Transformer architecture.
[0096] To generate acoustic representations, a fine-grained generative neural network can generate acoustic terms for a fine-grained vector quantizer, conditioned at least on acoustic terms and embedding terms used for a coarse-grained vector quantizer. For example, for each of one or more fine-grained vector quantizers in a hierarchy, the fine-grained generative neural network can generate corresponding acoustic terms for that vector quantizer for a second time step, conditioned at least on corresponding acoustic terms used for one or more coarse-grained vector quantizers in the hierarchy for a second time step. The fine-grained vector quantizer can employ a smaller quantization step size (e.g., higher resolution) than the coarse-grained vector quantizer.
[0097] Fine-grained generative neural networks can be autoregressive neural networks configured to autoregressively generate acoustic lexical units according to a second generation order. In some implementations, fine-grained generative neural networks have a decoder-only Transformer architecture. In other implementations, fine-grained generative neural networks have an encoder-decoder Transformer architecture.
[0098] The audio generative neural network 220 then uses a decoder neural network to process at least the acoustic representation to generate a prediction of the audio signal 104. For example, the corresponding audio sample at each of multiple output time steps across a time window can be based on one or more acoustic lexical units of the acoustic representation.
[0099] In some implementations, the decoder neural network can be a decoder neural network of a neural audio codec. For example, the neural audio codec can be the SoundStream neural audio codec.
[0100] Neural audio codecs can include decoder neural networks and encoder neural networks. For example, an encoder neural network can convert audio into a coded signal, which is then quantized into an acoustic representation. A decoder neural network can convert the acoustic representation into a predicted audio signal.
[0101] Figure 3 It is for reference. Figure 1 A block diagram of an example audio generation system 100 is described. Specifically, the audio generation system 100 uses generative neural networks 310, 315, and 220 to generate predictions of audio signals 104 given an input image 102.
[0102] The system receives input image 102, as referenced above. Figure 1 As stated above.
[0103] The system uses a generative neural network 310 (also known as a second generative neural network) to process the input image 102 to generate an image explanation 212. (See above for reference.) Figure 2 The image explanation 212 is a sequence of natural language text describing the input image 102.
[0104] To generate the image caption 212, system 100 can provide input image 102 and a request to describe the content of the input image as input to generative neural network 310. For example, the request may include a natural language request to describe what can be seen in the input image. In some examples, the request may include a request to describe what can be seen in the input image in as much detail as possible. In some examples, the request may also specify the type of features that the image caption can describe. Generative neural network 310 can be referenced above. Figure 2 The generative neural network 210 described herein is similar to or the same as that described above.
[0105] As an example, the image description 212 could include “A sea turtle floats calmly in clear turquoise ocean waters near the ocean floor. The ocean floor is teeming with marine life such as fish and plants. The sea turtle seems at peace with its surroundings.”
[0106] In some examples, the request to describe what can be seen in the input image includes one or more examples as few-sample cue examples. For instance, the request may include a request to describe what can be seen in the input image based on examples. In some examples, the request may include a natural language request to describe what can be seen in the input image based on examples. Each example may include an example image and a corresponding explanation of the example image. See below for reference. Figure 4 Describe the example image and provide an explanation of the image.
[0107] System 100 includes an image explanation 212 in network input 314. System 100 uses a generative neural network 315 (also known as a third generative neural network) to process network input 314 to generate a music explanation 216.
[0108] To generate the music description 216, system 100 can provide network input 314 and a request to rewrite the image description as a music description as input to generative neural network 315. For example, the request may include a natural language request describing what a suitable musical accompaniment for the image description would sound like. In some examples, the natural language request may also include a definition of the music description, such as a definition specifying the type of audio features that the music description can describe. The generative neural network 315 is described in further detail below. In some examples, generative neural network 310 may be configured to perform image and text processing tasks, but not plain text processing tasks. Therefore, the system can use generative neural network 315 to perform the plain text processing task of rewriting the image description as a music description. In some other examples, although generative neural network 310 may be configured to perform plain text processing tasks, the system can use generative neural network 315 to reduce the amount of computational resources and time required to generate the music description 216. For example, generative neural network 315 may be smaller than generative neural network 310, i.e., it may have fewer weights.
[0109] As referenced above Figure 2 The music interpretation description 216 describes one or more audio features that correspond to or are suitable for the input image 102. Figure 3 In the example, the musical explanation 216 includes "Flowing instrumental piece featuring a combination of piano and strings, conveying asense of grace and tranquility".
[0110] In some examples, the request to rewrite an image description as a musical description includes one or more examples as few-sample prompt examples. For instance, the request might include a request to rewrite an image description as a musical description based on an example. In some examples, the request might include a natural language request to rewrite an image description as a musical description based on an example. Each example in the examples might include an example image description and a corresponding example musical description. See below for reference. Figure 4 Describe the example images and music.
[0111] System 100 uses an audio generative neural network 220 to process the music interpretation 216 to generate an audio signal 104. (See above for reference.) Figure 2An audio generative neural network 220 is described. Therefore, the audio signal 104 includes music that can be described as “flowing instrumental piece featuring a combination of piano and strings, conveying a sense of grace and tranquility.”
[0112] The generative neural network 315 is configured to perform text processing tasks. The generative neural network 315 processes the input text to generate output. For example, if the input includes an image description 212 and a request to rewrite the image description as a music description, the output includes text describing the audio features corresponding to the image description.
[0113] The generative neural network 315 can have any suitable neural network architecture that allows the model to map an input sequence of text words from the vocabulary to an output sequence of text words from the vocabulary. The generative neural network 315 can have any suitable Transformer-based architecture, such as an encoder-decoder Transformer architecture, or a decoder-only Transformer architecture.
[0114] Specifically, the generative neural network 315 may be an autoregressive neural network that generates text words from an output sequence of regressively generated text words by generating each specific text word in the output sequence conditioned on the current input sequence, which includes (i) the input sequence, followed by (ii) any text words preceding the specific text word in the output sequence.
[0115] More specifically, to generate a specific text terminology, the generative neural network 315 can process the current input sequence to generate a score distribution, such as a probability distribution, which assigns a corresponding score, such as a corresponding probability, to each terminology in the vocabulary of text terms. The generative neural network 315 can then use the score distribution to select text terms from the vocabulary as specific text terms. For example, the generative neural network 315 can greedily select the terminology with the highest score, or it can sample terms from the distribution, for example, using top-k sampling, kernel sampling, or another sampling technique.
[0116] As a specific example, the generative neural network 315 can be an autoregressive Transformer-based neural network comprising multiple layers, each applying a self-attention operation. The generative neural network 315 can have any of a variety of Transformer-based neural network architectures. Examples of such architectures include those described in the following literature: J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, LA Hendricks, J. Welbl, A. Clark, et al., Training compute-optimal large language models, arXiv preprint arXiv:2203.15556, 2022; JW Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, HF Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, LA Hendricks, M. Rauh, P. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J.Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Elsen, SM Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M.Paganini, L. Sifre, L. Martens, XL Li, A. Kuncoro, A. Nematzadeh, E.Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J. Lespiau, M. Tsimpoukelli,N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D.Toyama, C. de Masson d’Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. A. Hechtman, L. Weidinger, I. Gabriel, W. S. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs / 2112.11446, 2021;Colin Raffel、Noam Shazeer、Adam Roberts、Katherine Lee、Sharan Narang、Michael Matena、Yanqi Zhou、Wei Li和Peter J Liu,Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan,Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, AmandaAskell, et al., Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. .
[0117] The lexical units in the vocabulary can be any suitable text lexical units, such as words, word fragments, punctuation marks, characters, bytes, etc., representing elements of text in one or more natural languages, as well as numbers and other text symbols optionally found in a corpus of text. For example, System 100 can lexicalize a given sequence of words by applying a lexer (such as the SentencePiece lexer (Kudo et al., arXiv:1808.06226) or another lexer) to divide the sequence into lexical units from the vocabulary.
[0118] Before using the generative neural network 315 to generate the music interpretation 216, the generative neural network 315 is pre-trained, for example by system 100 or by one or more other systems.
[0119] Specifically, System 100 or other systems pre-train a generative neural network 315 on a language modeling task (e.g., a task that requires predicting the next word in the training data given the current sequence of text words). Equivalently, for each given unlabeled text sequence in the training dataset, the language modeling task may require predicting the text sequence following the given unlabeled text sequence in the corresponding document. As a specific example, the generative neural network 315 can be pre-trained on a large dataset of text—e.g., text publicly available from the Internet or another text corpus—with regard to a maximum likelihood objective.
[0120] Figure 4 Examples of few-sample prompts, 400, 410, and 420, are shown. Each example includes an image, a corresponding image description, and a corresponding musical description. Each image description may include visual details, subject, background, setting, mood, hue, and / or atmosphere. Each musical description describes one or more audio features, such as style, rhythm, melody, timing, pitch, mood, and / or instrument.
[0121] Image 402 in Example 400 is a real-world image depicting a concert. The corresponding image description 404 includes: “Captured in this image is the dynamic energy of a live concert, with bandmembers engaging with the crowd, instruments at the ready. The stage glows under the spotlight, and the audience is a sea of faces, the air charged with anticipation and excitement.” Music description 406 includes: “Electrifying guitar riff, relentless drive of adrumbeat, and the raw power of a bass line, all combining into a high-energy rock anthem that commands head nods and foot taps in unison with the crowd's rhythmic pulse.”
[0122] Image caption 404 describes themes such as "band members engaging with the crowd, instruments at the ready" and "the audience is a sea of faces." Image caption 404 also describes scenes such as "live concert" and "stage." Image caption 404 further describes emotions and atmosphere, such as "dynamic energy" and "the air charged with anticipation and excitement." Image caption 404 also describes visual details, such as "[t]the stageglows under the spotlight."
[0123] In some examples, image descriptions can have varying amounts of detail for the same image. For instance, an alternative image description for image 402 could be “dynamic energy of a live concert, bandmembers engaging with the crowd, instruments at the ready.” Another alternative image description for image 402 could be “Live metal concert with electric guitars and a huge crowd.” The system can use image descriptions with varying amounts of detail in few-shot cue examples to influence the generation of potentially different image descriptions, and thus potentially different music descriptions and music. The system can also use image descriptions with less detail or shorter length in few-shot cue examples to enable the use of smaller generative neural networks, requiring fewer computational resources while maintaining the response time of larger generative neural networks.
[0124] Music Explanation 406 describes instruments such as “electrifying guitar riff,” “relentless drive of a drumbeat,” and “raw power of abass line.” Music Explanation 406 also describes the style as “rock anthem.” Music Explanation 406 further describes the mood and tone as “electrifying,” “relentless drive,” “high-energy,” and “commands headnods and foot taps in unison with the crowd’s rhythmic pulse.”
[0125] In some examples, for the same image and / or musical description, the musical description can have varying amounts of detail. For example, an alternative musical description for Example 400 could include “electrifying guitar riff, powerful drums and bass line”. The system can influence the generation of potentially different musical descriptions by using musical descriptions with varying amounts of detail in few-shot cue examples, and thus influence the generation of potentially different music. The system can also use musical descriptions with less detail or shorter length in few-shot cue examples to enable the use of smaller generative neural networks while maintaining the response time of larger generative neural networks.
[0126] Image 412 in Example 410 depicts a painting. The corresponding image description 414 includes: "A solitary figure stands enveloped by the quietude of a vibrant garden, basking in the gentle embrace of sunlight. She seems to be in a moment of tranquil reflection, as the world around her bursts with the life of untamed blooms and the soft whisper of leaves in the breeze. It is an image of peaceful solitude, where the clamor of the world falls away before the simple purity of nature's own artistry." Musical explanation 416 includes: "A soft piano melody, intertwined with the warm hum of a cello, creates an ambiance of gentle introspection, with an apace that breathes slowly like the hush of a summer's breeze through a sunlit garden."
[0127] The corresponding image caption 414 describes themes such as: “[a] a solitary figure stands enveloped by the quietude of a vibrant garden, basking in the gentle embrace of sunlight. She seems to be in a moment of tranquil reflection.” Image caption 414 also describes scenes such as: “a vibrant garden” and “the world around her bursts with the life of untamed blooms and the soft whisper of leaves in the breeze.” Image caption 414 further describes the atmosphere as “an image of peaceful solitude, where the clamor of the world falls away before the simple purity of nature's own artistry.”
[0128] Musical Explanation 416 describes instruments such as “[a] soft piano melody, intertwined with the warm hum of a cello.” It also describes a pace such as “a pace that breathes slowly like the hush of asummer's breeze through a sunlit garden.” Finally, it describes the mood as “an ambiance of gentle introspection.”
[0129] Image 422 in Example 412 is a real-world image depicting a parachute on the ground. The corresponding image caption 424 includes: "An open field scattered with wildflowers, where a collapsed yellow parachute lies on the ground, its mission complete. The surrounding landscape is tranquil, with a forest line in the distance under a bright, expansive sky, suggesting a narrative of adventure concluded in the calm of nature." Musical explanation 426 includes: "Acoustic guitar softly strumming, with a relaxed tempo and a melody that sings of freedom and the joy of a journey's end. The music carries alightness, akin to a gentle wind, that might have once carried the parachutealoft, now settling into a peaceful silence."
[0130] The corresponding image caption 424 describes themes such as: “[o]pen field scattered with wildflowers, where a collapsed yellow parachute lies on the ground, its mission complete.” Image caption 424 also describes scenes such as: “[t]the surrounding landscape is tranquil, with a forest line in the distance under a bright, expansive sky.” Image caption 424 further describes the atmosphere as “suggesting a narrative of adventure concluded in the calm of nature.”
[0131] Music Explanation 426 describes instruments such as “acoustic guitar softly strumming.” It also describes rhythms such as “a relaxed tempo.” Furthermore, it describes the melody as “a melody that sings of freedom and the joy of a journey’s end.” Finally, it describes the mood as “[t]the music carries a lightness, akin to a gentle wind, that might have once carried the parachute aloft, now settling into a peaceful silence.”
[0132] The system 100 described above can use few-sample cue examples, such as few-sample cue examples 400, 410 and 420, to generate image and music explanations that include similar types of detail, levels of detail and / or writing styles.
[0133] For example, as mentioned above (reference) Figure 2 The system 100 described herein may use images, corresponding image captions, and corresponding musical captions as few-shot cue examples for the generative neural network 210. The system 100 may also use image captions and corresponding musical captions as few-shot cue examples for the generative neural network 210.
[0134] Such as the above reference Figure 3 The system 100 described herein can use images and corresponding image explanations as few-shot cue examples for the generative neural network 310. The system 100 can also use image explanations and corresponding musical explanations as few-shot cue examples for the generative neural network 315.
[0135] Figure 5 Example images 500, 510, 520, 530, 540, 550, 560, and 570 are shown, along with musical explanations 512, 522, 532, 542, 552, 562, and 572. As described above, system 100 can generate music described by musical explanations 512, 522, 532, 542, 552, 562, and 572.
[0136] Image 500 depicts a tiger walking across a meadow. Example music description 502 may describe the rhythm, timing, and instruments, and includes "slow, suspenseful percussion." A piece of music with slow, suspenseful percussion is suitable for image 500, thus reflecting the tiger's gait depicted in image 500.
[0137] Image 510 depicts a bird flying across a lake. Example music description 512 may describe the instruments and timing, and includes "flute solo with slow, melodic wind instruments playing in the background." A piece of music featuring a flute solo and slow, melodic wind instruments is suitable for image 510, reflecting the bird's relaxed flight and the tranquil atmosphere of image 510.
[0138] Image 520 depicts a fireworks display atop city skyscrapers. Example music description 522 can describe the instruments, rhythm, and mood, and includes "pulsating electronic beats with a captivating drum and bass rhythm." A piece of music with a captivating drum and bass rhythm is suitable for image 520, reflecting its celebratory and modern atmosphere.
[0139] Image 530 depicts a cat reclining on a sofa. Example music description 532 can describe mood and melody, and includes "calming ambient sounds, a melody of birdsong, and a soft breeze." A piece of music with calming ambient sounds, a melody of birdsong, and a soft breeze is suitable for image 530, thus reflecting the relaxing atmosphere of image 530.
[0140] Image 540 depicts a sunset over a mountain. Example music description 542 can describe the instruments and mood, and includes "calm and peaceful guitar and flute song." A piece of music with a calm and peaceful guitar and flute melody is suitable for image 540, thus reflecting the tranquil atmosphere of image 540.
[0141] Image 550 depicts a snake. Example music description 552 can describe the instruments and mood, including "soft, mystical tune featuring flute and chimes." A piece of music with a soft, mystical tune featuring flute and chimes suits image 550, thus reflecting its tranquil atmosphere.
[0142] Image 560 depicts a busy intersection in a large city. Example music description 562 can describe the instruments, rhythm, and melody, and includes "up-tempo electronic beats, a hint of melody." A piece of music with up-tempo electronic beats and some melody is suitable for image 560, reflecting the busy streets and vibrant atmosphere of image 560.
[0143] Image 570 depicts two koalas sleeping in a tree. Example music description 572 can describe the instruments, melody, and mood, and includes "soothing and calming melody, featuring gentle piano and acoustic guitar." A piece of music with a soothing and calming melody and featuring gentle piano and acoustic guitar is suitable for image 570, thus reflecting the sleeping koalas and tranquil atmosphere of image 570.
[0144] Figure 6 This is a flowchart of an example process for generating an audio signal given an image. For convenience, process 600 will be described as being executed by a system of one or more computers located in one or more locations. For example, an audio generation system appropriately programmed according to this specification, such as... Figure 1 , Figure 2 and Figure 3 The audio generation system 100 can execute process 600.
[0145] The system receives an input image (step 610). In some examples, the system may receive an input image from a user.
[0146] The system uses one or more generative neural networks to process the input image to generate a musical description (step 620). The musical description describes one or more audio features corresponding to the input image. For example, the musical description may describe style, rhythm, melody, timing, pitch, mood, or instrument.
[0147] In some implementations, as referenced above... Figure 2 The system can use a generative neural network 210 to process the input image to generate an image description of the input image. The same generative neural network 210 can be used to process a network input that includes at least an image description to generate a music description. In some examples, the network input also includes an input image.
[0148] In some examples, the system can provide the generative neural network 210 with few-shot cue examples. For instance, to generate music explanations, the system can provide the generative neural network 210 with network input and a request to rewrite image explanations into music explanations based on the few-shot cue examples. (See above for reference.) Figure 4 This describes an example of a few samples prompting an example.
[0149] In some implementations, as referenced above... Figure 3The system can use a generative neural network 310 to process the input image to generate an image description of the input image. The system can also use different generative neural networks 315 to process network inputs that include at least the image description to generate music descriptions.
[0150] In some examples, the system can provide few-shot cue examples to the generative neural network 315. For instance, to generate music explanations, the system can provide the generative neural network 315 with network input and a request to rewrite image explanations into music explanations based on few-shot cue examples. (See above for reference.) Figure 4 This describes an example of a few samples prompting an example.
[0151] The system uses an audio generative neural network to process music descriptions to generate an audio signal described by the music descriptions (step 630). The audio generative neural network is configured to generate audio signals conditioned on at least the text. (See above for reference.) Figure 2 An example audio generative neural network is described.
[0152] The methods described in this paper can be used to generate audio representing an input image. The audio may include music and / or soundscape. Audio descriptions can be generated based on the image. The audio descriptions can describe audio features, such as music and / or soundscape features. Audio descriptions can be generated based on image descriptions of the image. Image descriptions can be generated based on the image. Image descriptions can be generated based on neural network input. The neural network input may include one or more examples, each example including an example image description and a corresponding example audio description. Each of the one or more examples may further include an example image and an example audio description corresponding to the example image description.
[0153] The described system has applications extending beyond simply generating musical works or soundscapes. For example, some implementations of the described system can be used to evaluate, calibrate, test, or modify audio electronics or audio communication systems (such as mobile phones, smart speakers, video conferencing devices or systems), or audio signal transmission systems. As an example, the described system can be used to generate music or soundscapes captured, transmitted, or played by an audio device or audio communication or signal transmission system, and the output of the audio device or audio communication or signal transmission system can then be compared with the generated music or soundscape to evaluate the fidelity of the output. Optionally, the audio device or audio communication or signal transmission system can then be modified, for example, calibrated or trained, to increase fidelity. As another example, the generated soundscape can be combined with speech, and this combination is captured, transmitted, or played by an audio device or audio communication or signal transmission system configured to filter out background soundscapes, such as noise from offices, shopping malls, or other locations. The output of the audio device or audio communication or signal transmission system can then be compared with the input speech and / or soundscape to evaluate the effectiveness of the filtering. Optionally, the audio device or audio communication or signal transmission system can then be modified, for example calibrated or trained, to increase the effectiveness of the filtering (e.g., as measured by the attenuation of unwanted components of the signal, i.e., soundscapes).
[0154] This specification uses the term "configured" in conjunction with system and computer program components. For configuring one or more computer systems to perform a specific operation or action, it means that software, firmware, hardware, or a combination thereof are installed on the system to cause the system to perform that operation or action during operation. For configuring one or more computer programs to perform a specific operation or action, it means that one or more programs include instructions that, when executed by a data processing device, cause that device to perform that operation or action.
[0155] Embodiments of the subject matter and functional operation described in this specification may be implemented in digital electronic circuit systems, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their equivalents), or in a combination of one or more of these. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these. Alternatively or additionally, program instructions may be encoded on artificially generated propagation signals (e.g., machine-generated electrical, optical, or electromagnetic signals) that are generated to encode information for transmission to a suitable receiver device for execution by the data processing device.
[0156] The term "data processing device" refers to data processing hardware and includes all kinds of devices, apparatuses, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. The device may also be or further include special-purpose logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the device may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.
[0157] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages); and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not necessarily, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on a single computer or on multiple computers located at a site or distributed across multiple sites and interconnected via a data communication network.
[0158] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines may be installed and run on the same one or more computers.
[0159] The processes and logic flows described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by a dedicated logic circuit system, such as an FPGA or ASIC, or by a combination of a dedicated logic circuit system and one or more programmed computers.
[0160] A computer suitable for executing computer programs can be based on a general-purpose or special-purpose microprocessor, or both, or any other type of central processing unit. Generally, the central processing unit receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are the central processing unit for executing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented by or incorporated into a special-purpose logic circuit system. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or will be operatively coupled to receive data from or transfer data to such one or more mass storage devices, or both. However, a computer need not have such devices. Furthermore, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few.
[0161] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0162] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser in response to a request received from a web browser on the user's device. Additionally, the computer can interact with the user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving responsive messages from the user in response.
[0163] Data processing devices used to implement machine learning models may also include, for example, dedicated hardware accelerator units for handling the general and computationally intensive portions of machine learning training or production (i.e., inference, workloads).
[0164] Machine learning models can be implemented and deployed using machine learning frameworks (such as the TensorFlow framework or the Jax framework).
[0165] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components (e.g., as a data server), or middleware components (e.g., an application server), or front-end components (e.g., a client computer having a graphical user interface, web browser, or app that a user can interact with through an implementation of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0166] A computing system may include clients and servers. Clients and servers are typically geographically separated and interact via a communication network. The client-server relationship is established by computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., HTML pages) to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from that user. Data generated at the user device, such as the result of user interaction, may be received at the server from the device.
[0167] In addition to the embodiments described above, the following embodiments are also innovative: Example 1 is a method comprising: Receive input image; One or more generative neural networks are used to process the input image to generate musical descriptions of one or more audio features corresponding to the input image; and An audio generative neural network is used to process the music interpretation to generate a music signal described by the audio interpretation.
[0168] Example 2 is the method as described in Example 1, wherein one or more generative neural networks are used to process the input image to generate a musical explanation describing one or more audio features corresponding to the input image, including: The input image is processed using a first generative neural network to generate an image explanation describing the input image; and The first generative neural network is used to process network inputs that include at least the image explanatory descriptions to generate music explanatory descriptions describing one or more audio features.
[0169] Example 3 is the method as described in Example 2, wherein using a first generative neural network to process the input image to generate an image explanation describing the input image includes providing the input image and a request to describe the content of the input image as input to the first generative neural network.
[0170] Example 4 is a method as described in any one of Examples 2 to 3, wherein using the first generative neural network to process network input including at least the image explanatory description to generate a music explanatory description describing one or more audio features includes providing the network input and a request to rewrite the image explanatory description into a music explanatory description as input to the first generative neural network.
[0171] Example 5 is the method as described in Example 4, wherein the request further includes one or more examples, each example including an example image explanation and a corresponding example music explanation.
[0172] Example 6 is the method as described in Example 2, wherein the network input further includes the input image, and wherein using the first generative neural network to process the network input, which includes at least the image caption, to generate a music caption describing one or more audio features includes providing the network input and a request to rewrite the image caption as a music caption for the input image as input to the first generative neural network.
[0173] Example 7 is the method as described in Example 6, wherein the request further includes one or more examples, each example including an example image, a corresponding explanation of the example image, and a corresponding explanation of the example music.
[0174] Example 8 is a method as described in any one of Examples 1 to 7, wherein using one or more generative neural networks to process the input image to generate a musical explanation describing one or more audio features corresponding to the input image includes: The input image is processed using a second generative neural network to generate an image explanation describing the input image; and A third generative neural network is used to process network inputs that include at least the image explanatory descriptions to generate music explanatory descriptions describing one or more audio features.
[0175] Example 9 is the method as described in Example 8, wherein using a second generative neural network to process the input image to generate an image explanation describing the input image includes providing the input image and a request to describe the content of the input image as input to the second generative neural network.
[0176] Example 10 is a method as described in any one of Examples 8 to 9, wherein using the third generative neural network to process network input including at least the image explanatory description to generate a music explanatory description describing one or more audio features includes providing the network input and a request to rewrite the image explanatory description as a music explanatory description as input to the third generative neural network.
[0177] Example 11 is the method as described in Example 10, wherein the request further includes one or more examples, each example including an example image explanation and a corresponding example music explanation.
[0178] Example 12 is a method as described in any one of Examples 1 to 11, wherein the audio generative neural network is configured to generate audio signals conditioned on at least text.
[0179] Example 13 is a method as described in any one of Examples 1 to 12, wherein receiving the input image includes receiving the input image from a user.
[0180] Example 14 is a method as described in any one of Examples 1 to 13, wherein the method further includes providing the audio signal to be presented to the user.
[0181] Example 15 is a method as described in any one of Examples 1 to 14, wherein the one or more audio features describe any one or more of the following: style, rhythm, timing, pitch, mood, or instrument.
[0182] Example 16 is a system comprising: One or more computers; and One or more storage devices communicatively coupled to one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the method as described in any one of Embodiments 1 to 15.
[0183] Example 17 is one or more non-transitory computer-readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the method as described in any one of Examples 1 to 15.
[0184] While this specification contains numerous details of specific implementations, these details should not be construed as limiting the scope of any invention or the scope that may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features described in this specification within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as operating in certain combinations and even initially claimed in this way, one or more features from a claimed combination may, in some cases, be removed from the combination, and the claimed combination may involve sub-combinations or variations thereof.
[0185] Similarly, although operations are depicted in the accompanying drawings and described in a specific order in the claims, this should not be construed as requiring such operations to be performed in the specific order shown or in an ordered sequence, or requiring all shown operations to achieve the desired result. In some contexts, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0186] Specific embodiments of this subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired result. As an example, the processes depicted in the drawings do not necessarily require a specific order or sequential sequence to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous.
Claims
1. A computer-implemented method, comprising: Receive input image; One or more generative neural networks are used to process the input image to generate an audio description of one or more audio features corresponding to the input image; as well as An audio generative neural network is used to process the audio description to generate an audio signal described by the audio description.
2. The method as described in claim 1, wherein, Using one or more generative neural networks to process the input image to generate an audio description of one or more audio features corresponding to the input image includes: The input image is processed using a first generative neural network to generate an image explanation describing the input image; and The first generative neural network is used to process network inputs that include at least the image explanatory descriptions to generate audio explanatory descriptions describing one or more audio features.
3. The method as described in claim 2, wherein, Using a first generative neural network to process the input image to generate an image description of the input image includes providing the input image and a request to describe the content of the input image as input to the first generative neural network.
4. The method as claimed in claim 2 or claim 3, wherein, Using the first generative neural network to process network input including at least the image captions to generate audio captions describing one or more audio features includes: providing the network inputs and a request to rewrite the image captions into audio captions as inputs to the first generative neural network.
5. The method of claim 4, wherein, The request further includes one or more examples, each example including an example image explanation and a corresponding example audio explanation.
6. The method as claimed in claim 2 or claim 3, wherein, The network input further includes the input image, and wherein using the first generative neural network to process the network input, which includes at least the image caption, to generate an audio caption describing one or more audio features includes: providing the network input and a request to rewrite the image caption into an audio caption for the input image as input to the first generative neural network.
7. The method of claim 6, wherein, The request further includes one or more examples, each example including an example image, a corresponding explanation of the example image, and a corresponding explanation of the example audio.
8. The method of claim 1, wherein, Using one or more generative neural networks to process the input image to generate an audio description of one or more audio features corresponding to the input image includes: The input image is processed using a second generative neural network to generate an image explanation describing the input image; and A third generative neural network is used to process network inputs that include at least the image explanatory descriptions to generate audio explanatory descriptions describing one or more audio features.
9. The method of claim 8, wherein, Using a second generative neural network to process the input image to generate an image explanation describing the input image includes: providing the input image and a request to describe the content of the input image as input to the second generative neural network.
10. The method of claim 8 or claim 9, wherein, Using the third generative neural network to process network input including at least the image explanatory description to generate an audio explanatory description describing one or more audio features includes: providing the network input and a request to rewrite the image explanatory description into an audio explanatory description as input to the third generative neural network.
11. The method of claim 10, wherein, The request further includes one or more examples, each example including an example image explanation and a corresponding example audio explanation.
12. The method as described in any of the preceding claims, wherein, The audio generative neural network is configured to generate audio signals conditioned on at least text.
13. The method as described in any of the preceding claims, wherein, Receiving an input image includes receiving the input image from the user.
14. The method of any of the preceding claims, further comprising providing the audio signal to a user.
15. The method as claimed in any of the preceding claims, wherein, The one or more audio features describe any one or more of the following: style, rhythm, timing, pitch, mood, or instrument.
16. The method as claimed in any of the preceding claims, wherein, The audio explanation is a music explanation.
17. A system comprising: One or more computers; as well as One or more storage devices communicatively coupled to one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operation of a corresponding method as described in any one of claims 1 to 16.
18. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operation of a corresponding method as claimed in any one of claims 1 to 16.