A diffusion model for generating audio data based on descriptive text prompts

JP2026503686AInactive Publication Date: 2026-01-29GOOGLE LLC
1 Cites -1 Cited by

Patent Information

Application Number
JP2025543330
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-26
Filing Date
2024-01-26
Publication Date
2026-01-29
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure 2026503686000001_ABST
    Figure 2026503686000001_ABST
Patent Text Reader

Abstract

A corpus of text data is generated using a machine learning text generation model. The corpus of text data includes a plurality of sentences. Each sentence describes a type of audio. For each of a plurality of audio recordings, the audio recording is processed using a machine learning audio classification model to obtain training data including the audio recording and one or more sentences from the plurality of sentences that are closest to the audio recording in a joint audio-text embedding space of the machine learning audio classification model. The sentence(s) are processed using a machine learning generative model to obtain intermediate representations of the one or more sentences. The intermediate representations are processed using a machine learning cascade diffusion model to obtain audio data. The machine learning cascade diffusion model is trained based on differences between the audio data and the audio recording.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Diffusion models having improved accuracy and reduced consumption of computational resources

    WO2022265992A1