A diffusion model for generating audio data based on descriptive text prompts
JP2026503686AInactive Publication Date: 2026-01-29GOOGLE LLC
1 Cites -1 Cited by
Patent Information
- Application Number
- JP2025543330
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-26
- Filing Date
- 2024-01-26
- Publication Date
- 2026-01-29
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure 2026503686000001_ABST
Abstract
A corpus of text data is generated using a machine learning text generation model. The corpus of text data includes a plurality of sentences. Each sentence describes a type of audio. For each of a plurality of audio recordings, the audio recording is processed using a machine learning audio classification model to obtain training data including the audio recording and one or more sentences from the plurality of sentences that are closest to the audio recording in a joint audio-text embedding space of the machine learning audio classification model. The sentence(s) are processed using a machine learning generative model to obtain intermediate representations of the one or more sentences. The intermediate representations are processed using a machine learning cascade diffusion model to obtain audio data. The machine learning cascade diffusion model is trained based on differences between the audio data and the audio recording.
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
Diffusion models having improved accuracy and reduced consumption of computational resources
WO2022265992A1