Audio Generation via Spectrogram Stacking and Generative Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating audio assets in video games are time-consuming, expensive, and computationally intensive, requiring manual crafting of variations and being difficult to implement on-the-fly, especially for creating multiple similar sounds with varying acoustic characteristics.

Innovation Solution

A method utilizing single-image generative models trained on graphical representations of audio, such as spectrograms, to efficiently generate multiple audio assets by converting input audio assets into graphical representations, stacking them, and using a generative adversarial network to produce output audio assets with controlled variations, reducing computational power and time requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual crafting methods are used to create audio asset variations, then each audio asset can be carefully designed with suitable variations, but the process becomes time-consuming and expensive

Engineering Contradiction:
Improveaudio asset variation qualityVSAvoidproduction time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of audio asset creation with an automated computational system. A generative model trained on spectral representations of audio data automatically generates varied audio assets, substituting human sound designers with an AI-based system that can produce multiple variations simultaneously without manual intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the approach from manual parameter adjustment to automated parameter generation. By training a generative model on spectral data, the system learns to generate audio assets with varied parameters (pitch, timbre, volume) automatically, transforming the creation process from manual parameter tweaking to automated parameter synthesis based on learned patterns.

Inventive Principle:
Principle #35Parameter changes

2Extent of automation

If computational audio generation is implemented, then audio assets can be generated automatically, but the process becomes complex and computationally intensive

Engineering Contradiction:
Improveaudio generation automationVSAvoidcomputational complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent extracts the essential features of audio data by converting audio assets into spectral representations (such as spectrograms or mel-spectrograms). This extraction process separates the core acoustic characteristics from the raw audio signal, allowing the generative model to work with compressed, feature-based data rather than full-resolution audio, thereby reducing computational complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system transforms the audio generation problem from the time domain to the frequency domain by using spectral representations. This dimensional change allows the generative model to operate on a different representation of the audio data that is more amenable to automated generation while reducing the computational burden compared to working directly with raw audio waveforms.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Manufacturing precision

If single output generative models are used, then each audio asset can be generated individually with high quality, but the process increases time and processing costs

Engineering Contradiction:
Improveaudio asset qualityVSAvoidgeneration speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent merges multiple audio asset generation tasks into a single batch processing operation. The generative model is trained to process multiple spectral representations simultaneously, allowing it to generate multiple varied audio assets in one computational pass. This combining of tasks maintains quality while significantly improving generation speed and reducing processing costs.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The generative model is designed with universal functionality to handle diverse audio generation tasks. By training on a dataset of spectral representations from various audio assets, the model learns general patterns that allow it to generate varied versions of different types of audio assets using the same system, making it multi-functional rather than requiring separate models for each asset type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12168176B2Audio generation methods and systems
Publication Date: 2024.12.17 SONY COMP ENTERTAINMENT EURO LTD
  • US12168176B2 patent drawing
  • US12168176B2 patent drawing
  • US12168176B2 patent drawing

AI summary

Generating audio assets by receiving a plurality of input audio assets, converting each input audio asset into an input graphical representation, generating an input multi-channel image by stacking each input graphical representation in a separate channel of the image, feeding the input multi-channel image into a generative model to train the generative model and generate one or more output multi-channel images, each output multi-channel image including an output graphical representation, extracting the output graphical representations from each output multi-channel image and converting each output graphical representation into an output audio asset.