Method for synthesizing audio waveform from text description

By converting text descriptions into image embeddings and synthesizing spectrograms, the method solves the problem of scaling text-audio datasets, achieves efficient audio synthesis, and reduces computational complexity and data annotation costs.

CN121014077APending Publication Date: 2025-11-25DOLBY LABORATORIES LICENSING CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480027921.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2024-04-25
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing text-to-audio synthesis systems rely on training with large-scale labeled text-audio pairs, which makes it difficult to scale up the dataset, especially compared to text-image datasets.

Method used

By training a generative model using image and audio data, text descriptions are first converted into image embeddings, then spectrograms are synthesized and converted into audio waveforms. This avoids direct reliance on text-audio datasets and leverages the availability and ease of annotation of image-audio pairs. The CLIP model and diffusion model are used for training and conversion.

Benefits of technology

It reduces the complexity and computational requirements of generative models, decreases the reliance on large text-audio training datasets, and improves the efficiency and quality of audio synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121014077A_ABST
    Figure CN121014077A_ABST
Patent Text Reader

Abstract

One aspect of the present disclosure relates to a method for synthesizing an audio waveform from a textual description indicative of a desired sound, the method comprising: determining a text embedding from the textual description; determining image embedding according to the text embedding; synthesizing a spectrogram by inputting the image embedding into a generative model trained to synthesize the spectrogram given the input image embedding; and converting the synthesized spectrogram into an audio waveform.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to Spanish Application No. P202330333, filed April 27, 2023; U.S. Provisional Application No. 63 / 588,010, filed October 5, 2023; and U.S. Provisional Application No. 63 / 554,808, filed February 16, 2024, all of which are incorporated herein by reference in their entirety. Technical Field

[0003] This disclosure relates to methods, systems, and computer program products for synthesizing audio waveforms from text descriptions. Background Technology

[0004] Text-to-audio synthesis refers to the task of generating audio given a textual description of a desired sound. This task has wide applicability in fields such as content creation and sound engineering design.

[0005] With advancements in generative modeling and language-audio contrastive learning, various deep learning-based text-to-audio synthesis systems have emerged in recent years. Existing text-to-audio algorithms rely on large-scale labeled text-to-audio pairs to train such systems. For example, AudioLDM (Liu et al., “AudioLDM: Text-to-Audio Generation with LatentDiffusion Models”, https: / / arxiv.org / abs / 2301.12503v1) utilizes millions of text-to-audio pairs, while MusicLM (Agostinelli et al., “MusicLM: Generating Music From Text”, https: / / arxiv.org / abs / 2301.11325v1) collects tens of millions of text-music pairs to train its text-to-music model. However, labeling general audio (such as mixtures of sound effects, music, and speech) with accurate text subtitles is extremely laborious.

[0006] Despite significant human annotation efforts, the largest publicly available text-audio dataset known to the inventors currently contains only about 630,000 text-audio pairs. Given the relatively scarce availability of text-audio data compared to text-image data, it remains unclear whether it is feasible to scale text-audio datasets to a size comparable to large-scale text-image datasets, such as the LAION-5B dataset containing 5.85 billion text-image pairs. Summary of the Invention

[0007] In view of the above, it is desirable to provide a deep learning-based method that allows the synthesis of audio waveforms from text descriptions, which does not rely on text-audio data for training, but can be trained using image and audio data that are relatively easier to obtain than text-audio data.

[0008] According to a first aspect of this disclosure, a method is provided for synthesizing an audio waveform from a text description indicating a desired sound, the method comprising: determining a text embedding based on the text description; determining an image embedding based on the text embedding; synthesizing a spectrogram by inputting the image embedding into a generative model trained to synthesize a spectrogram given an input image embedding; and converting the synthesized spectrogram into an audio waveform.

[0009] According to the first aspect of the method, spectrograms can therefore be generated by a generative model trained to synthesize spectrograms given image embeddings (rather than text embeddings). Since the image embeddings input to the generative model are determined from text embeddings based on text descriptions, and the generative model is trained to synthesize spectrograms given image embeddings, this method does not rely on any text-audio pairs during inference and training. Obtaining video and text-image data is relatively easier than obtaining text-audio pairs, which facilitates the training of the generative model. Therefore, this method avoids the need for tedious labeling of large text-audio training datasets.

[0010] By synthesizing and then converting the spectrogram into an audio waveform, generative models can operate in a lower-dimensional space compared to directly synthesizing audio waveforms. This reduces the complexity and computational requirements of generative models.

[0011] The generative model trained to synthesize spectrograms may be referred to as the “second generative model” below, to distinguish it from the other generative models described in this paper.

[0012] In some embodiments, text embeddings can be generated using a text encoder trained to map input text descriptions to text embeddings. More specifically, the text encoder can be trained to map input text descriptions to text embeddings in a latent space shared with the image embeddings used to train the second generative model to synthesize spectrograms. Thus, the text embeddings can reflect a similar underlying data structure to the image embeddings used to train the second generative model. This indicates that the text encoder can provide text embeddings during inference that, in principle, are interchangeable with image embeddings due to their shared latent text-image embedding space, i.e., in zero-shot modality transfer. However, it has been observed that different modalities of text and images tend to create text-image modality gaps in the shared latent space. Since the second generative model is trained to generate spectrograms given image embeddings (i.e., without relying on any text embeddings), spectrogram generation may be particularly sensitive to this gap. By determining image embeddings based on text embeddings and using the determined image embeddings as input to the second generative model, spectrogram synthesis can be guided using embeddings that are closer to the image embeddings used during the training of the second generative model. This reduces performance degradation caused by potential text-image modal differences between the text embeddings and image embeddings in the training dataset.

[0013] The image embeddings used to train the second generative model can be generated using an image encoder trained to map input images into a latent space. The text encoder and image encoder can be jointly trained using a contrastive loss. Therefore, using a training dataset of text-image pairs, the text encoder and image encoder can be trained to predict matching text-image pairs.

[0014] The text encoder can be a contrastive language-image pre-trained (CLIP) text encoder. The image encoder can be a CLIP image encoder. That is, the text encoder and the image encoder can be the corresponding encoders of the CLIP model. Accordingly, the second generative model can be trained to synthesize spectrograms given CLIP image embeddings for the input.

[0015] In some embodiments, the text embedding may be a CLIP text embedding and the second generative model may be trained to synthesize a spectrogram given a CLIP image embedding.

[0016] In some embodiments, the image embedding is determined using a first generative model (i.e., different from a second generative model) trained to transform the input text embedding into an image embedding. Using a generative model allows for the imbuing of the image embedding with generative characteristics.

[0017] As mentioned above, even if the text embeddings and image embeddings used to train the second generative model reside in a shared latent space, a gap can still exist between the text and image modalities in practice. This can also be observed with contrastively trained text and image encoders (such as CLIP models). Therefore, the use of the first generative model allows for narrowing this gap by generating image embeddings that are closer to the image embeddings used during the training of the second generative model than the text embeddings.

[0018] Determining an image embedding based on a text embedding may involve sampling N image embeddings from a distribution of image embeddings generated by a first generative model given a text embedding as input, and selecting the image embedding with the highest similarity to the text embedding from these N image embeddings. Therefore, an image embedding can be determined by (e.g., randomly) sampling the distribution of image embeddings and selecting the image embedding closest to the text embedding. The selected image embedding may be the one among the N image embeddings that has the highest cosine similarity to the text embedding. Thus, the similarity to the text embedding can be determined without being affected by possible separations (e.g., Euclidean) within the latent space.

[0019] The first generative model can be a prior. Therefore, a model for image embeddings can be learned from a "prior" (interchangeably, a "prior model") and used to determine image embeddings based on text embeddings. For example, the prior could be a diffusion prior. Therefore, a diffusion-based process can be used to determine image embeddings.

[0020] The prior can be trained to model the distribution of image embeddings generated by the image encoder, conditioned on the text embeddings generated by the text encoder. Therefore, given the text embeddings of a matching text-image pair as input, the prior can be used to predict the image embeddings of the matching text-image pair.

[0021] In particular, the prior can be trained to generate CLIP image embeddings given CLIP text embeddings of the input.

[0022] In some embodiments, instead of using a generative model, determining image embeddings involves retrieving K ≥ 1 image embeddings from a set of image embeddings, wherein the retrieved K image embeddings are the K image embeddings in the set that have the maximum cosine similarity to the text embeddings, and the image embeddings are determined based on the retrieved K image embeddings. Therefore, a retrieval-based approach can be used to determine image embeddings in a computationally efficient manner. This set of image embeddings may include image embeddings known to be relatively close to the image embeddings used to train a second generative model (i.e., a generative model trained to synthesize spectrograms). In particular, this set of image embeddings may include the image embeddings used to train the second generative model.

[0023] The number of retrieved image embeddings can be K=1, where the image embedding (i.e., the image embedding to be input into the second generative model) can be determined as the image embedding with the maximum cosine similarity to the text embedding in the set of image embeddings.

[0024] The number of retrieved image embeddings can be K > 1, where the image embeddings (i.e., the image embeddings to be input into the second generative model) can be determined as the average of the K retrieved image embeddings, or can be randomly sampled from the K retrieved image embeddings. Therefore, a certain degree of generative characteristics can be assigned to the determination of image embeddings without relying on the generative model.

[0025] The second generative model for synthesizing spectrograms may include a diffusion decoder model, a generative adversarial network (GAN), or an autoregressive model, where the image embedding is input as a conditional vector into the second generative model.

[0026] The second generative model can be trained to synthesize Mel spectrograms. Mel spectrograms support compressed spectrogram representations, thereby further reducing the complexity of the synthesis task for the second generative model.

[0027] By feeding a synthesized spectrogram into a third generative model trained to convert the input spectrogram into an audio waveform, the synthesized spectrogram can be transformed into an audio waveform. Spectrogram-to-waveform conversion is particularly well-suited for generative methods, for example, because it can generate new waveforms not present in the original training dataset. This can be a desirable feature in areas such as content creation.

[0028] The third generative model can include diffusion models, generative adversarial networks, or autoregressive models.

[0029] In some embodiments, the second generative model (i.e., the generative model trained to synthesize spectrograms) is trained using spectrograms and image embeddings generated from video frames paired with audio. Therefore, the second generative model can be trained using audiovisual data obtained from a training dataset of videos.

[0030] Training the second generative model can include:

[0031] Extract sequences of video frames and audio frames from a training dataset of videos that include audio and video frames of objects;

[0032] Convert the sequence of audio frames into a spectrogram;

[0033] The training image embedding is determined based on the extracted video frames; and

[0034] Generative models are trained using training image embeddings as conditional vectors to synthesize spectrograms.

[0035] Therefore, training the second generative model does not require access to any image or audio captions. Videos depicting various types of sound-producing objects (i.e., objects that are the sound sources in the video) are readily available on the internet, for example. Accordingly, the second generative model can be trained on such video datasets without additional video labels to support text-based audio synthesis as illustrated in this paper. The training image embeddings can be determined using the image encoders mentioned above (e.g., CLIP image encoders). In particular, the training image embeddings can be CLIP image embeddings.

[0036] Training can include extracting a sequence of video frames from the video, determining a sequence of training image embeddings from the extracted sequence of video frames, and then determining the training image embeddings as the average image embeddings of the sequence of training image embeddings. This can make training more robust and reduce the impact of divergent and unrepresentative image frames when they appear in the video training dataset.

[0037] The training image embeddings can be determined using an image encoder jointly trained with the text encoder used during inference, as discussed above. The neural network model can be a contrastive model. In particular, the image encoder can be a CLIP model image encoder.

[0038] According to the second aspect, a system for synthesizing audio waveforms from text descriptions is provided, the system being configured to implement the method according to the first aspect.

[0039] According to a third aspect, a computer program product is provided that includes a computer program code portion configured to execute the method according to the first aspect when executed on a computer processor.

[0040] The features of the second and third aspects may contain the same effects and advantages as those discussed with reference to the method of the first aspect. Any functionality described in relation to the method may have a corresponding feature in a system or computer program product. Attached Figure Description

[0041] Various aspects of this disclosure will be described in more detail with reference to the accompanying drawings, which illustrate exemplary embodiments.

[0042] Figure 1 This is a flowchart of a method for synthesizing audio waveforms from text descriptions.

[0043] Figure 2 It is used to implement Figure 1 A schematic block diagram of the method system.

[0044] Figure 3This is a flowchart of a method for training a generative model to synthesize audio waveforms from image embeddings.

[0045] Figure 4 It is to achieve Figure 3 A schematic block diagram of the training method system.

[0046] Figure 5 A schematic block diagram of the example system during training is shown.

[0047] Figure 6 A schematic block diagram of the example system during testing is shown.

[0048] Figure 7 The illustration shows a schematic block diagram of an example device, system, or architecture that can be used to implement various aspects of this disclosure. Detailed Implementation

[0049] Now refer to Figure 1 Flowcharts and Figure 2 The present disclosure describes a method for synthesizing audio waveforms from text descriptions using a deep learning-based system 1. System 1 includes: a trained neural network model configured to determine a text embedding qtext based on the text description; a first generative model 12 for determining an image embedding q'img based on the text embedding qtext; a second generative model 13 for synthesizing a spectrogram X; and a third generative model 14 for converting the synthesized spectrogram X into an audio waveform A.

[0050] At S1, the text encoder 11 determines the text embedding qtext based on the text description. The text encoder 11 is trained to map the text description to the text embedding in the latent space. More specifically, the latent space can be shared with the image embedding used to train the second generative model 13.

[0051] As will be discussed in further detail below, the second generative model 13 can use an image encoder (e.g., Figure 4The image embeddings generated by the image encoder 21 shown are used for training. Hereinafter, the image embeddings generated by image encoder 21 are denoted as qimg. Since these image embeddings are used to train the second generative model 13, they can be interchangeably referred to as the training image embedding qimg. Image encoder 21 can be trained to map input images to image embedding qimg in the same latent space as the text embedding qtext generated by text encoder 11. This can be achieved, for example, by jointly training text encoder 11 and image encoder 21 using a contrastive loss. Text encoder 11 and image encoder 21 can be trained using a training dataset of matched text-image pairs, such that text encoder 11 and image encoder 21 can learn to predict matched text-image pairs, i.e., the correct pairing of text and image. More specifically, text encoder and image encoder can be trained to generate text and image embeddings qtext, qimg, such that the text and image embeddings of mismatched text-image pairs have lower cosine similarity than the text and image embeddings of matched text-image pairs.

[0052] The actual training of the text encoder 11 and image encoder 21 is not a prerequisite for the methods and systems disclosed herein. Instead, the text encoder 11 and image encoder 21 can be the text encoder and image encoder of any suitable existing pre-trained language-visual model. An example is the Contrastive Language-Image Pre-trained (CLIP) model (Radford et al., “Learning Transferable Visual Models From Natural Language Supervision”, arXiv:2103.00020, https: / / doi.org / 10.48550 / arXiv.2103.00020). The CLIP model is a pre-trained model trained on a very large dataset of image-text pairs using contrastive loss, such that text descriptions and images are mapped to a shared latent space, resulting in lower cosine similarity between the text and image embeddings of mismatched text-image pairs compared to those of matched text-image pairs. Accordingly, in some implementations, the text encoder 11 may be a CLIP text encoder trained to determine the CLIP text embedding qtext based on the text description, and the image encoder 21 that generates the image embedding qimg for training the second generative model 13 may be a CLIP image encoder.

[0053] The text description input to the text encoder 11 can be obtained, for example, from a text file that includes the text description and is stored in the memory of the computing system implementing system 1. The text description can also be obtained as interactive user input, for example, by entering it in an input box or by entering it at the prompts of the user interface of the computing system.

[0054] Although the goal is to synthesize audio, as discussed above, the text encoder 11 can be a text encoder for a language-visual model (such as CLIP). Therefore, to improve the relevance of the text embedding, it can be as follows: Figure 2 As shown, a text description is input to text encoder 11 using the syntax “[label] photo”, where [label] is a tag or descriptor indicating the type of audio to be synthesized (e.g., “train whistling” in the illustrated example). However, as can be appreciated, the specific syntax of the text description can depend on both the dataset used to train text encoder 11 and the labels of the dataset used to train the second generative model 13, as discussed further below. In any case, the text description can typically include natural language input describing the desired sound.

[0055] At S2, the first generative model 12 determines the image embedding q'img based on the text embedding qtext. The first generative model 12 is trained to generate the image embedding q'img given the input text embedding qtext. Therefore, the input text embedding qtext can be transformed into the image embedding q'img. It should be noted here that the image embedding generated by the first generative model 12 is represented as q'img, while the image embedding generated by the image encoder 21 during the training of the second generative model 13 (i.e., the training image embedding) is represented as qimg. Even though the text encoder 11 (e.g., the CLIP text encoder) is trained to map, for example, the text description "dog" to a text embedding close to the image embedding of a photograph of "dog" generated by the image encoder 21 in the embedding space, in practice, this text embedding is often not close enough to be directly used as input to the second generative model 13 to synthesize the spectrogram X of the audio waveform with the desired sound. Therefore, a first generative model 12 is introduced in system 1 as a way to generate an "improved" image embedding q'img that is closer to its corresponding training image embedding qimg than the text embedding qtext.

[0056] More specifically, the first generative model 12 can be a prior model, or a prior, trained to model the distribution of image embeddings q'img (e.g., CLIP image embeddings) generated by the image encoder 21 conditioned on text embeddings qtext (e.g., CLIP text embeddings) generated by the text encoder 11. Therefore, the first generative model 12 (i.e., the prior) can be used to predict the image embedding q'img of a matching text-image pair given its text embedding qtext as input. The prior can be a diffusion prior, making it a computationally efficient implementation. However, other generative priors are also possible, such as autoregressive priors.

[0057] While it is possible to train the prior from scratch, for example, using a training dataset of matched text-image pairs or text-image embeddings tailored for the audio synthesis task, it is also possible to use a pre-trained prior, such as a transformer-based diffusion prior model available at https: / / huggingface.co / laion / DALLE2-PyTorch that uses the CLIP model.

[0058] Using the first generative model 12, N≥1 image embeddings {q'img,1,…,q'img,N} generated by the first generative model 12 from a distribution {q'img}D of image embeddings generated by the first generative model 12 with a given text embedding qtext as input are sampled, and an image embedding with a high similarity to the text embedding qtext is selected from the N image embeddings {q'img,1,…,q'img,N≥1} as input to the second generative model 13, to determine the image embedding q'img to be used as input (e.g., a conditional vector) for the second generative model 13. Cosine similarity provides a convenient and relevant metric for evaluating similarity. As can be appreciated, in the implementation where N=1, the first generative model 12 can be run (i.e., executed) once to generate the image embedding q'img from the text embedding qtext, which can then be used as input to the second generative model 13. In an implementation where N>1, the first generative model 12 can be run N times to generate N image embeddings q'img,1,…,q'img,N (i.e., the distribution of the N image embeddings) from a (single) text embedding qtext. Then, an image embedding among the N image embeddings that has a large cosine similarity to the text embedding qtext can be selected as the image embedding q'img to be used as the input to the second generative model 13.

[0059] As an alternative to using the first generative model 12 to generate the image embedding q'img to be used as input to the second generative model 13, a retrieval-based method can be used. The reference numeral 12 is now used to indicate the image embedding-retrieval block. Figure 2The block diagram also represents this implementation. This retrieval block can be configured to retrieve K ≥ 1 image embeddings from a predefined set of image embeddings, where the retrieved K image embeddings are the K image embeddings in the set that have the highest cosine similarity to the text embedding qtext generated by the text encoder 11. Then, the image embedding q'img to be used as input to the second generative model 13 can be determined based on the retrieved K image embeddings. In the K > 1 implementation, the image embedding q'img can be determined as the average of the retrieved K image embeddings, or randomly sampled from the retrieved K image embeddings. In the K = 1 implementation, the image embedding q'img can simply be determined as the image embedding in the set that has the highest cosine similarity to the text embedding qtext. The set of image embeddings from which the retrieval module retrieves the K image embeddings can include the image embeddings used to train the second generative model 13. Therefore, it can be ensured that the image embedding q'img input to the second generative model 13 is or close to the image embedding used to train the second generative model 13.

[0060] At S3, the spectrogram X is synthesized by inputting the image embedding q'img into the second generative model 12. For example... Figure 2 As indicated herein, the second generative model 12 can be implemented as a diffusion decoder model, but other implementations are also possible, such as GANs or autoregressive models. In either case, the image embedding q'img can be input as a conditional vector to the second generative model 12. For example, the second generative model 12 can be trained to synthesize a Mel spectrogram, which can support a compressed representation of the audio waveform A. However, other spectrogram representations are also compatible with this disclosure.

[0061] At S4, the synthesized spectrogram X is transformed into an audio waveform A by the third generative model 14. The third generative model 14 is trained to transform the input spectrogram X into an audio waveform A, i.e., an audio signal in the waveform domain. Various implementations of the third generative model 14 are possible, such as GANs, autoregressive models, or diffusion models. An example of a GAN-based implementation is the BigVGAN neural vocoder model (S. Lee, W. Ping, B. Ginsburg, B. Catanzaro, and S. Yoon, “BigVGAN: A Universal Neural Vocoder with Large-Scale Training”, Proc. International Conference on Learning Representations, 2023). Examples of diffusion-based models include “DiffWave” (Kong et al., “DiffWave: A Versatile Diffusion Model for AudioSynthesis”, https: / / doi.org / 10.48550 / arXiv.2009.09761) and “SpecGrad” (Koizumi et al., “SpecGrad: Diffusion Probabilistic Model based Neural Vocoder with Adaptive Noise Spectral Shaping”, https: / / doi.org / 10.48550 / arXiv.2203.16749).

[0062] The audio synthesis method and system 1 are supported by both pre-trained priors and pre-trained text and image encoders 11, 21 (e.g., CLIP). However, to facilitate audio synthesis, a second generative model 13 needs to be trained. As indicated above, the second generative model 13 can be trained to synthesize a spectrogram X using training image embeddings and spectrograms generated from video frames paired with audio. Therefore, the second generative model 13 can be trained using audiovisual data obtained from a training dataset of videos.

[0063] Now refer to Figure 3 The flowchart and illustration of training system 2 are shown. Figure 4 The diagram illustrates an example implementation of training a second generative model 13 from a training dataset of videos.

[0064] In addition to the second generative model 13 to be trained, System 2 also includes an image encoder 21. The image encoder 21 is trained, as discussed above, to map the input image I to a (trained) image embedding qimg in a latent space shared with the text embedding qtext generated by the text encoder 11 of System 1 during testing / inference. The image encoder 21 may, for example, be a pre-trained CLIP image encoder that generates CLIP image embeddings. In any case, the image encoder 21 is trained before training the second generative model 13 and is frozen during the training of the second generative model 13.

[0065] At S5, video frames (i.e., images) or sequences of video frames I and sequences of audio frames are extracted from the corresponding video V in the training dataset {V}, which includes audio sound A' and video frame I. One or more video frames can be extracted from randomly selected segments of video V. The segments can have a predetermined or randomly selected length (e.g., within a suitable range).

[0066] At S6, the sequence of audio frames is converted into a spectrogram x0. Spectrogram x0 can be a Mel spectrogram as mentioned above. Spectrogram x0 can be calculated by the audio waveform to spectrogram conversion module 22.

[0067] At S7, the image embedding qimg to be input as a conditional vector to the second generative model 13 is determined based on the extracted video frame(s) I. In an example implementation, a single video frame is extracted from video V, wherein the image embedding qimg can be determined by mapping the extracted video frame to the image embedding qimg by the image encoder 21 and feeding the image embedding qimg to the second generative model 13.

[0068] In the example implementation, a sequence of video frames is extracted from video V. Image encoder 21 can map each video frame in the sequence to a corresponding image embedding. Then, the image embedding qimg input to the second generative model 13 can be determined as the average of the image embeddings of the sequence of video frames, or randomly selected from the image embeddings.

[0069] At S8, the second generative model 13 is trained to synthesize the ground-truth spectrogram xo determined in step S6 using the image embedding qimg as a conditional vector. The training process can be iterated on the training dataset of video V.

[0070] As mentioned earlier, the second generative model 12 can be implemented as a diffusion model (e.g. Figure 4(As shown in the figure). Denoising diffusion probability models can be employed (e.g., see “Improved Denoising Diffusion Probabilistic Models” by A. Nichol and P. Dhariwal, Proc. ICML, 2019) and classifier-free guidance (e.g., see “Classifier-Free Diffusion Guidance” by TS Jonathan Ho, NeurIPS Workshop on Deep Generative Models and Downstream Applications, 2021).

[0071] For example, the second generative model13 can be trained using the VGGSound dataset (Chen et al., “VGGSound: A Large-scale Audio-Visual Dataset”, arXiv:2004.14368, https: / / arxiv.org / pdf / 2004.14368.pdf). Each video in the VGGSound dataset is 10 seconds long and contains audio and visual frames of objects. There are approximately 300 audio classes in this dataset. However, the VGGSound dataset is just one example, and other datasets with similar characteristics can be used, such as MUSIC (available at https: / / github.com / roudimit / MUSIC_dataset), ACAV100M (available at https: / / acav100m.github.io / ), Ego4D (available at https: / / ego4d-data.org / ), or EPIC-SOUNDS (available at https: / / epic-kitchens.github.io / epic-sounds / ).

[0072] Although Figure 4 The second generative model 13 is shown as a diffusion model, but those skilled in the art will understand that other implementations of the second generative model 13 (e.g., GAN or autoregressive model) can be trained in a similar manner.

[0073] Now refer to Figures 5 to 6 Further illustrative implementations of the training and testing systems are described below. It should be noted that various numerical examples related to the model, training, and testing also apply to the above references. Figures 1 to 4 The methods and systems discussed.

[0074] Figure 5 There are three components to be trained. Each component and... Figure 6 The components of the testing system are described as follows:

[0075] CLIP Image / Text Encoder: The training system uses a pre-trained CLIP model to extract embeddings from input images (at training time) or text (at testing time). CLIP has been pre-trained on approximately 400 million image-text pairs. As mentioned earlier, it consists of an image encoder and a text encoder that map input images / text to (ideally) a shared latent space. Contrastive loss forces matching image-text pairs to have large cosine similarity embeddings, while non-matching image-text pairs have small cosine similarity embeddings. For example, cosine similarity can be a measure between 0 and 1, where values ​​closer to 1 have large cosine similarity. For instance, a matching text-image pair has a target cosine similarity of 1, while a non-matching text-image pair has a target cosine similarity of 0. As another example, a large cosine similarity can be a much larger value, in contrast to the cosine similarity of non-matching text-image pairs. In this way, the embeddings of the two modalities (image, text) can be used interchangeably. In other words, CLIP bridges the image and text modalities in audio synthesis.

[0076] Diffusion Decoder and BigVGAN: The diffusion model (corresponding to the second generative model 13) is trained on, for example, the VGGSound dataset mentioned above, to synthesize a Mel spectrogram given a CLIP image embedding. Since this diffusion model “unCLIPs” the CLIP embedding into a Mel spectrogram, it can be referred to as a diffusion decoder in this context. During training, segments of length X are randomly selected from the video (e.g., X could be 1-100 seconds, 1-1000 seconds, etc.). Image frames are extracted at Y frames per second (e.g., 1-10 frames per second, 1-100 frames per second, etc.), and audio frames are extracted every 8-32 ms in a window size of 40-128 ms. The audio is resampled to, for example, 16 kHz. Each audio frame is then converted to 40-128 Mel bands. A frame is randomly selected from the X*Y image frames, and a CLIP image embedding is extracted using a pre-trained CLIP image encoder. In some embodiments, the CLIP image embedding can be the average of the CLIP embeddings of all frames. The diffusion model follows the UNet DDPM architecture (Ho et al., “Denoising Diffusion Probabilistic Models”, arXiv:2006.11239v2, https: / / arxiv.org / pdf / 2006.11239.pdf) and is trained to synthesize the Mel spectrogram of the target audio given a CLIP image embedding as a conditional vector. Finally, BigVGAN is trained to transform the Mel spectrogram into an audio waveform. BigVGAN is trained on ground-based Mel spectrograms, also from VGGSound. Notably, this diffusion model generates a highly compressed Mel spectrogram compared to directly synthesizing a waveform. This option greatly simplifies the task of the diffusion model because it operates in a much lower dimensional space. Meanwhile, a powerful vocoder (BigVGAN) reconstructs spectral-temporal details from the lossy Mel spectrogram.

[0077] Diffusion Prior: At test time, in practice, a text description of the desired sound can be taken, converted into a text embedding using a pre-trained CLIP text encoder, and audio synthesized from that text. Ideally, the contrastive learning paradigm in CLIP projects the text embeddings and image embeddings of matched text-image pairs to close embedding vectors in the shared latent space of the CLIP text encoder and image encoder. However, as discussed above, in practice, there is often a gap between the two modalities in this space. Since the diffusion decoder only sees CLIP image embeddings during training, a large quality degradation can be observed when switching to CLIP text embeddings at test time. For example, the cosine similarity of matched text-image pairs can be 0.2–0.3 (showing the gap). Therefore, a further diffusion model, namely the diffusion prior, is used to learn a prior distribution of CLIP image embeddings generated by the CLIP image encoder given CLIP text embeddings (hence the name diffusion prior). By sampling from this prior distribution, CLIP embeddings that are closer to the CLIP image embeddings used during the training of the diffusion decoder in the previous step can be obtained. It can be noted that this approach is similar to the diffusion prior method proposed in "Hierarchical Text-Conditional Image Generation with CLIP Latents" (Ramesh et al., arXiv:2204.06125, https: / / arxiv.org / pdf / 2204.06125.pdf), where a text-to-image synthesis system is developed and CLIP image embeddings are generated from text descriptions using a diffusion prior, which are then used by a diffusion decoder to synthesize images. However, since the CLIP model itself is pre-trained on image-text pairs, the quality of the synthesized image may not be significantly degraded for text-to-image synthesis tasks, even without using a prior to bridge the text-image gap. Therefore, this is quite different from the audio synthesis task in this example, since the CLIP model has never been trained on audio.

[0078] Training the diffusion prior will require text-image pairs. However, to support unsupervised or self-supervised settings, a pre-trained diffusion prior model can be used, such as a transformer-based diffusion prior model, which is available at https: / / huggingface.co / laion / DALLE2-PyTorch. This diffusion prior model follows the transformer architecture proposed by Ramesh et al. and is trained on 2 billion text-image pairs. Note that all text is unrelated to audio.

[0079] Figure 6The processing during inference is illustrated. Given a text description, a pre-trained CLIP text encoder computes a text embedding, which is then transformed into a CLIP image embedding by sampling from a prior distribution of learned image embeddings. Given a CLIP text vector, N image embeddings (e.g., N=2) are randomly sampled from the prior distribution, and one image embedding with a large cosine similarity to the text embedding is selected. Finally, this image embedding is fed into a trained diffusion decoder to obtain a synthesized Mel spectrogram, which is then converted into an audio waveform by BigVGAN.

[0080] Instead of training a diffusion prior, a simple "zero-shot retrieval method" can be used to bridge the gap between CLIP text and image embeddings. During training, this method stores CLIP image embeddings for all images from the diffusion decoder training set or even a larger image dataset. During testing, given a CLIP text embedding, the K best-stored CLIP image embeddings with the highest cosine similarity to the text embedding can be retrieved. Audio can then be synthesized using either the top 1 image embedding or the average of the top K image embeddings.

[0081] The three specific generative models referenced in the example above (diffusion decoder, BigVGAN, and diffusion prior) are exemplary and can be replaced by any generative model, such as generative adversarial networks (GANs), autoregressive models, etc., all of which are within the scope of this disclosure.

[0082] The systems and methods disclosed in this disclosure can be implemented as software, firmware, hardware, or a combination thereof. In hardware implementations, the division of tasks does not necessarily correspond to the division of physical units; instead, a physical component can have multiple functionalities, and a task can be performed collaboratively by multiple physical components.

[0083] Computer hardware can be, for example, a server computer, client computer, personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, smartphone, web appliance, network router, switch, or bridge, or any machine capable of executing instructions (sequentially or otherwise) to perform actions specified for the computer hardware. Furthermore, this disclosure should relate to any collection of computer hardware that individually or in combination executes instructions to implement any one or more concepts discussed herein.

[0084] Figure 7 A schematic block diagram is shown of an example electronic device or architecture 200 (e.g., device 200) suitable for implementing example embodiments of this disclosure. Architecture 200 includes, but is not limited to, Figure 2 System 1 Figure 4 System 2 Figures 5 to 6The training and testing system, and the combination of Figures 1 to 6 The described method is implemented as follows. As shown in the figure, the architecture 200 includes a central processing unit (CPU) 201, which is capable of executing various processes according to a program stored, for example, in read-only memory (ROM) 202 or a program loaded from, for example, storage unit 208 into random access memory (RAM) 203. The CPU 201 may be, for example, an electronic processor 201, which may include one or more processor cores, and in some examples, the processor 201 may be multiple processors. In RAM 203, data required by the CPU 201 to execute various processes is also stored as needed. The CPU 201, ROM 202, and RAM 203 are connected via bus 204, and an input / output (I / O) interface 205 is also connected to bus 204.

[0085] The following components are connected to I / O interface 205: input unit 206, which may include a keyboard, mouse, etc.; output unit 207, which may include a display (such as a liquid crystal display (LCD)) and one or more speakers; storage unit 208, which includes a hard disk or another suitable storage device; and communication unit 209, which may include a network interface card, such as a network card (e.g., wired or wireless).

[0086] In some implementations, the input unit 206 includes one or more microphones (depending on the host device) located at different locations, enabling the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0087] In some implementations, output unit 207 includes a system with a different number of speakers. Output unit 207 (depending on the capabilities of the host device) can render audio signals in various formats, such as mono, stereo, immersive, binaural, and other suitable formats.

[0088] In some embodiments, communication unit 209 is configured to communicate with other devices (e.g., via a network). Drive 210 is also connected to I / O interface 205 as needed. Removable media 211 (such as a disk, optical disk, magneto-optical disk, flash drive, or other suitable removable media) is mounted on drive 210 to install computer programs read therefrom into storage unit 208 as needed. Those skilled in the art will understand that while apparatus 200 is described as including the components described above, in practical applications it is possible to add, remove, and / or replace some of these components, and all such modifications or changes fall within the scope of this disclosure.

[0089] According to exemplary embodiments of this disclosure, the above processes can be implemented as computer software programs or on computer-readable storage media. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly implemented on a machine-readable medium, the computer program including program code for performing methods. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 209, and / or installed from removable media 211, such as... Figure 7 As shown in the image.

[0090] Generally, the various example embodiments of this disclosure can be implemented in hardware or dedicated circuitry (e.g., control circuitry systems), software, logic, or any combination thereof. For example, those discussed above... Figure 2 and Figure 4 The various components can be controlled by a control circuit system (e.g., CPU 201 and...). Figure 7 The control circuit system can perform the actions described in this disclosure by combining other components in the system.

[0091] Some aspects may be implemented in hardware, while others may be implemented in firmware or software, which may be executed by a controller, processor, and / or (one or more) other computing devices (which may include control circuitry systems). While various aspects of exemplary embodiments of this disclosure have been illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it will be appreciated that, by way of non-limiting example, the blocks, apparatuses, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0092] Furthermore, the various blocks shown in the flowchart can be considered as method steps, and / or operations resulting from the operation of computer program code, and / or as circuit elements of multiple coupled logics performing one or more associated functions. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly implemented on a machine-readable medium, the computer program containing program code configured to perform the methods described above.

[0093] Computer program code used to perform the methods of this disclosure may be written in any combination of one or more programming languages. This computer program code may be provided to one or more processors of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus having a control circuitry system, such that when the program code is executed by one or more processors of the computer or other programmable data processing apparatus, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

[0094] One or more processors can operate as independent devices or can be connected (e.g., networked) to (one or more) other processors. Such a network can be built on a variety of different network protocols and can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0095] Software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is well known to those skilled in the art, the term "computer storage media" includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, various forms of physical (non-transitory) storage media, such as ROM, PROM, EPROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage devices, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. In addition, as is well known to those skilled in the art, communication media (transitory) typically implement computer-readable instructions, data structures, program modules or other data in modulated data signals (such as carrier waves or other transmission mechanisms), and includes any information delivery media.

[0096] The embodiments of the technology disclosed in the accompanying drawings are merely illustrative examples, and the invention is not limited thereto. For example, Figure 2 , Figure 4 , Figure 5 and Figure 6 The partitions shown (such as boxes) are merely illustrative logical partitions for ease of discussion. Without departing from the spirit of the invention, these partitions can be broken down into additional partitions, combined into fewer partitions, supplemented with additional partitions, or reduced by eliminating partitions. For Figure 1 and Figure 3 The flowchart shown may, without departing from the spirit of this disclosure, be divided into sections of operation steps (which may also be referred to as functions, steps, operations, processes or actions) into fewer steps or into additional steps, wherein steps may be reordered or eliminated in whole or in part.

[0097] Unless otherwise expressly stated, as is evident from the following discussion, it should be recognized that the use of terms such as “processing,” “computing,” “determining,” and “analyzing” throughout the public discussion can refer to the functions, actions, steps, and / or processes of computer hardware or computing systems or similar electronic computing devices that manipulate and / or transform data expressed in physical quantities (such as electronic quantities) into other data expressed in similar physical quantities.

[0098] It should be recognized that in the above description of exemplary embodiments of the invention, various features are sometimes grouped together in a single embodiment, drawing, or description in order to simplify the disclosure and aid in understanding one or more of the various inventive aspects. However, this mode of disclosure should not be construed as reflecting an intention to claim more features than expressly enumerated in each claim. Rather, as reflected in the following claims, inventive aspects are present in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, each claim being an independent embodiment of the invention in itself. Furthermore, while some embodiments described herein include some features included in other embodiments but not others, as those skilled in the art will understand, combinations of features from different embodiments are also implied within the scope of the invention and form different embodiments. For example, in the following claims, any claimed embodiment can be used in any combination.

[0099] Furthermore, some of the embodiments described herein are methods or combinations of methods that can be implemented by a processor of a computer system or by other means of performing functions. Thus, a processor with instructions for performing the elements of such a method forms a means for performing the elements of such a method. It should be noted that when a method includes several elements (e.g., several steps), the order of such elements is not implied unless otherwise stated. Furthermore, the elements of the apparatus embodiments described herein are examples of means for performing the functions performed by such elements in order to carry out embodiments of the invention. Numerous specific details are set forth in the description provided herein. However, it should be understood that embodiments of the invention can be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.

[0100] Various aspects of this disclosure can be understood from the following enumerated exemplary embodiments (“EEE”):

[0101] EEE 1. A method for synthesizing an audio waveform from a text description indicating a desired sound, the method comprising:

[0102] Determine the text embedding based on the text description;

[0103] Determine image embedding based on text embedding;

[0104] Spectrum maps are synthesized by feeding image embeddings into a generative model trained to synthesize spectromaps given input image embeddings; and

[0105] The synthesized spectrogram is converted into an audio waveform.

[0106] EEE 2. Based on the method of EEE 1, the text embedding is generated using a text encoder trained to map the input text description to a latent space shared with the image embedding used to train a second generative model to synthesize spectrograms.

[0107] EEE 3. Based on the method of EEE 2, the image embeddings used to train the generative model are generated using an image encoder trained to map the input image to the latent space, wherein the text encoder and the image encoder are jointly trained using a contrastive loss.

[0108] EEE 4. Based on the method of EEE 3, where the text encoder is a contrastive language-image pre-trained (CLIP) text encoder and the image encoder is a CLIP image encoder.

[0109] EEE 5. The method according to any of the preceding EEEs, wherein the text embedding is a CLIP text embedding, and wherein the generative model is trained to synthesize a spectrogram given a CLIP image embedding.

[0110] EEE 6. The method according to any of the preceding EEEs, wherein the generative model is a second generative model, and wherein the image embedding is determined using a first generative model trained to transform input text embeddings into image embeddings.

[0111] EEE 7. According to the method of EEE 6, wherein determining the image embedding based on the text embedding includes sampling N image embeddings from a distribution of image embeddings generated by a first generative model from a given text embedding as input, and selecting the image embedding with the maximum similarity to the text embedding from the N image embeddings.

[0112] EEE 8. Based on the method of EEE 7, wherein the selected image embedding is the image embedding with the largest cosine similarity to the text embedding among N image embeddings.

[0113] EEE 9. According to any of the methods in EEE 7-8, where the first generative model is a prior, such as a diffusion prior.

[0114] EEE 10. Based on the method of EEE 9, when subordinate to EEE 3, the prior is trained to model the distribution of image embeddings generated by the image encoder, conditioned on the text embeddings generated by the text encoder.

[0115] EEE 11. The method according to any of EEE 6-10, wherein the prior is trained to generate CLIP image embeddings given CLIP text embeddings of the input.

[0116] EEE 12. The method according to any one of EEE 1-5, wherein determining the image embedding comprises retrieving K ≥ 1 image embeddings from a set of image embeddings, wherein the retrieved K image embeddings are the K image embeddings in the set of image embeddings that have the maximum cosine similarity with respect to the text embeddings, and wherein the image embeddings are determined based on the retrieved K image embeddings.

[0117] EEE 13. According to the method of EEE 12, where K=1 and the image embedding is determined to be the retrieved image embedding, or where K>1 and the image embedding is determined to be the average of the retrieved K image embeddings, or randomly sampled from the retrieved K image embeddings.

[0118] EEE 14. According to the method of any one of EEE 12-13, when subordinate to EEE 2, the set of image embeddings includes image embeddings used to train a second generative model.

[0119] EEE 15. The method according to any of the preceding EEEs, wherein the generative model for synthesizing the spectrogram includes a diffusion decoder model, a generative adversarial network (GAN), or an autoregressive model, wherein the image embedding is input to the generative model as a conditional vector.

[0120] EEE 16. According to any of the methods in the preceding EEE, wherein the generative model is trained to synthesize a Mel spectrogram.

[0121] EEE 17. The method according to any of the preceding EEEs, wherein the synthesized spectrogram is converted into an audio waveform by inputting the synthesized spectrogram into a third generative model trained to convert the input spectrogram into an audio waveform.

[0122] EEE 18. Based on the approach of EEE 17, the third generative model includes diffusion models, generative adversarial networks, or autoregressive models.

[0123] EEE 19. The method according to any of the preceding EEEs, wherein the generative model is trained to synthesize a spectrogram using a spectrogram and training image embeddings generated from video frames paired with audio.

[0124] EEE 20. According to the method of EEE 19, training the generative model includes:

[0125] Extract sequences of video frames and audio frames from a training dataset of videos that include audio and video frames of objects;

[0126] Convert the sequence of audio frames into a spectrogram;

[0127] The training image embedding is determined based on the extracted video frames; and

[0128] Generative models are trained using training image embeddings as conditional vectors to synthesize spectrograms.

[0129] EEE 21. According to the method of EEE 20, training includes extracting a sequence of video frames from a video, determining a sequence of training image embeddings from the extracted sequence of video frames, and determining the training image embeddings as the average image embeddings of the sequence of training image embeddings.

[0130] EEE 22. According to any of the methods in EEE 19-21, when belonging to EEE 3, the training image embedding is determined using an image encoder.

[0131] EEE 23. A system for synthesizing audio waveforms from a text description, the system being configured to implement a method according to any of the preceding EEEs.

[0132] EEE 24. A computer program product including a computer program code portion configured to perform a method according to any one of EEE 1-22 when executed on a computer processor.

[0133] EEE 25. An unsupervised deep learning method for synthesizing audio waveforms from a text description of a desired sound, the method comprising:

[0134] Determine the text embedding based on the text description;

[0135] Text embeddings are transformed into image embeddings by sampling from the prior distribution of the learned image embeddings;

[0136] Select image embeddings that have a high cosine similarity to the text embeddings, sampled from the learned prior distribution;

[0137] The synthesized Mel spectrogram is determined based on the selected image embedding; and

[0138] The synthesized Mel spectrogram is converted into an audio waveform.

[0139] EEE 26. An unsupervised deep learning system for synthesizing audio waveforms from a text description describing a desired sound, the system comprising:

[0140] The neural network model is configured as follows:

[0141] Determine the text embedding based on the text description;

[0142] The first generative model is configured as follows:

[0143] Text embeddings are transformed into image embeddings by sampling from the prior distribution of the learned image embeddings;

[0144] Select image embeddings that have high cosine similarity to the text embeddings, randomly sampled from the learned prior distribution;

[0145] The second generative model is configured as follows:

[0146] The synthesized Mel spectrogram is determined based on the selected image embedding;

[0147] The third generative model is configured as follows:

[0148] The synthesized Mel spectrogram is converted into an audio waveform.

[0149] EEE 27. A system according to EEE 26, wherein at least one of a first generative model and / or a second generative model is configured to be trained via a dataset of text in which no audio-related information exists.

[0150] EEE 28. A system according to EEE 26 or EEE 27, wherein the neural network model includes a contrastive language-image pre-trained (CLIP) model, the first generative model includes a diffusion prior model, the second generative model includes a diffusion decoder model, and / or the third generative model includes a BigVGAN model.

[0151] EEE 29. A system according to any one of EEE 26-28, wherein the first generative model, the second generative model and the third generative model each include at least one of generative adversarial networks (GANs) and / or autoregressive models.

[0152] EEE 30. A system according to any one of EEE 26-29, wherein text embedding includes CLIP text embedding and image embedding includes CLIP image embedding.

[0153] EEE 31. An unsupervised deep learning system for synthesizing audio waveforms from a text description describing a desired sound, the system comprising:

[0154] The neural network model is configured as follows:

[0155] Determine the text embedding based on the text description;

[0156] The search module is configured as follows:

[0157] Retrieve image embeddings from a set of image embeddings, wherein the retrieved image embeddings have the maximum cosine similarity to text embeddings in the set of image embeddings;

[0158] The first generative model is configured as follows:

[0159] The synthesized Mel spectrogram is determined based on the selected image embedding;

[0160] The second generative model is configured as follows:

[0161] The synthesized Mel spectrogram is converted into an audio waveform.

[0162] EEE 32. A system according to EEE 31, wherein the neural network model includes a contrastive language-image pre-trained (CLIP) model, the first generative model includes a diffusion decoder model, and / or the second generative model includes a BigVGAN model.

[0163] EEE 33. A system according to EEE 31 or 32, wherein the first generative model and the second generative model each comprise at least one of a generative adversarial network (GAN) and / or an autoregressive model.

[0164] EEE 34. A system according to any one of EEE 31-33, wherein text embedding includes CLIP text embedding and image embedding includes CLIP image embedding.

[0165] EEE 35. An apparatus comprising a processor and a memory coupled to the processor, wherein the processor is adapted to perform a method according to EEE 25.

[0166] EEE 36. A program comprising instructions which, when executed by a processor, cause the processor to perform a method according to EEE1.

Claims

1. A method for synthesizing an audio waveform from a text description indicating a desired sound, the method comprising: The text embedding is determined based on the text description; The image embedding is determined based on the text embedding; The spectrogram is synthesized by feeding the image embedding into a generative model trained to synthesize spectrograms given the input image embedding; and The synthesized spectrogram is converted into an audio waveform.

2. The method of claim 1, wherein the text embedding is generated using a text encoder trained to map input text descriptions to text embeddings in a latent space shared with image embeddings used to train the generative model to synthesize spectrograms.

3. The method of claim 2, wherein the image embeddings used to train the generative model are generated using an image encoder trained to map input images to image embeddings in the aforementioned latent space, wherein the text encoder and the image encoder are jointly trained using a contrastive loss.

4. The method of claim 3, wherein the text encoder is a contrastive language-image pre-trained (CLIP) text encoder, and the image encoder is a CLIP image encoder.

5. The method according to any one of the preceding claims, wherein the text embedding is a CLIP text embedding, and wherein the generative model is trained to synthesize a spectrogram given a CLIP image embedding.

6. The method according to any one of the preceding claims, wherein the generative model is a second generative model, and wherein the image embedding is determined using a first generative model trained to transform input text embeddings into image embeddings.

7. The method of claim 6, wherein determining the image embedding based on the text embedding comprises sampling N image embeddings from a distribution of image embeddings generated by a first generative model, given the text embedding as input, and selecting from the N image embeddings the image embedding that has the greatest similarity to the text embedding.

8. The method of claim 7, wherein the selected image embedding is the image embedding among the N image embeddings that has the maximum cosine similarity to the text embedding.

9. The method according to any one of claims 7-8, wherein the first generative model is a prior, such as a diffusion prior.

10. The method of claim 9, wherein, when subordinate to claim 3, the prior is trained to model the distribution of image embeddings generated by the image encoder, conditioned on text embeddings generated by the text encoder.

11. The method of any one of claims 6-10, wherein the prior is trained to generate CLIP image embeddings given CLIP text embeddings of the input.

12. The method according to any one of claims 1-5, wherein determining the image embedding comprises retrieving K ≥ 1 image embeddings from a set of image embeddings, wherein the retrieved K image embeddings are the K image embeddings in the set of image embeddings that have the maximum cosine similarity to the text embedding, and wherein the image embeddings are determined based on the retrieved K image embeddings.

13. The method of claim 12, wherein K=1 and the image embedding is determined to be a retrieved image embedding, or wherein K>1 and the image embedding is determined to be the average of the K retrieved image embeddings, or randomly sampled from the K retrieved image embeddings.

14. The method according to any one of claims 12-13, wherein, when subordinate to claim 2, the set of image embeddings includes image embeddings used for training a second generative model.

15. The method according to any one of the preceding claims, wherein the generative model for synthesizing the spectrogram comprises a diffusion decoder model, a generative adversarial network (GAN), or an autoregressive model, wherein the image embedding is input to the generative model as a conditional vector.

16. The method according to any one of the preceding claims, wherein the generative model is trained to synthesize a Mel spectrogram.

17. The method according to any one of the preceding claims, wherein the synthesized spectrogram is converted into an audio waveform by inputting the synthesized spectrogram into a third generative model trained to convert the input spectrogram into an audio waveform.

18. The method of claim 17, wherein the third generative model comprises a diffusion model, a generative adversarial network, or an autoregressive model.

19. The method according to any one of the preceding claims, wherein the generative model is trained to synthesize a spectrogram using a spectrogram and training image embeddings generated from video frames paired with audio.

20. The method of claim 19, wherein training the generative model comprises: Extract sequences of video frames and audio frames from a training dataset of videos that include audio and video frames of objects; Convert the sequence of audio frames into a spectrogram; The training image embedding is determined based on the extracted video frames; as well as Generative models are trained using training image embeddings as conditional vectors to synthesize spectrograms.

21. The method of claim 20, wherein the training includes extracting a sequence of video frames from the video, determining a sequence of training image embeddings from the extracted sequence of video frames, and determining the training image embeddings as the average image embedding of the sequence of training image embeddings.

22. The method according to any one of claims 19-21, wherein, when subordinate to claim 3, the training image embedding is determined using an image encoder.

23. A system for synthesizing audio waveforms from a text description, the system being configured to implement the method according to any one of the preceding claims.

24. A computer program product including a computer program code portion, said computer program code portion being configured to perform the method according to any one of claims 1-22 when executed on a computer processor.