Layered Audio Generation With Multi-Channel Spectrograms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio content generation in video games is time-consuming, expensive, and computationally intensive, particularly for creating varied audio assets, and existing computational solutions often require single-output generation, increasing time and costs.

Innovation Solution

A method using a generative model trained on multi-channel graphical representations of audio assets, specifically spectrograms, to efficiently generate layered audio assets with reduced computing power, allowing for quick and autonomous creation of varied audio assets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual crafting of audio assets is used to create variations, then audio quality and variation are improved, but time consumption and cost increase

Engineering Contradiction:
Improveaudio asset qualityVSAvoidtime consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent uses a trained generative model to automatically generate copies/variations of audio assets based on learned patterns from training data, replacing manual crafting while maintaining quality consistency

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system generates audio variations by manipulating parameters such as pitch, tempo, and timbre through the generative model, creating diverse audio assets from a single source without manual intervention

Inventive Principle:
Principle #35Parameter changes

2Loss of time

If computational audio generation is used to generate assets on the fly, then time consumption is reduced, but processing complexity and cost increase

Engineering Contradiction:
Improvegeneration timeVSAvoidprocessing complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The system performs preliminary training of the generative model during asset creation, storing the trained model for rapid generation of variations later, reducing real-time processing complexity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The trained model is saved and reused to generate multiple audio variations efficiently, avoiding repeated training processes and reducing computational overhead

Inventive Principle:
Principle #26Copying

3Device complexity

If single-output generative models are used, then model simplicity is maintained, but productivity decreases due to sequential generation

Engineering Contradiction:
Improvemodel simplicityVSAvoidgeneration throughput
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent segments the audio asset into multiple independent layers (e.g., different instrument tracks, voice channels), allowing parallel generation of each layer through the same generative model

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system combines multiple independently generated audio layers into a complete audio asset, achieving multi-output capability while reusing the same simple generative model for each layer

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12364924B2Audio generation methods and system
Publication Date: 2025.07.22 SONY COMP ENTERTAINMENT EURO LTD
  • US12364924B2 patent drawing
  • US12364924B2 patent drawing
  • US12364924B2 patent drawing

AI summary

A method of generating audio assets, comprising the steps of: receiving an input multi-layered audio asset comprising a plurality of audio layers, generating an input multi-channel image, wherein each channel of the input multi-channel image comprises an input image representative of one of the audio layers, training a generative model on the input multi-channel image and implementing the trained generative model to generate an output multi-channel image, wherein each channel of the output multi-channel image comprises an output image representative of an output audio layer, and generating an output multi-layered audio asset based on a combination of output audio layers derived from the output images.