Generative Audio Model Retargeting for Video Game Asset Creation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The generation of varied audio assets for video games is time-consuming and computationally expensive, as sound designers must manually create multiple variations of audio assets, and existing computational solutions are complex and inefficient, especially for generating assets of different durations.

Innovation Solution

A method using a single-image generative model trained on a graphical representation of audio, such as spectrograms, to retarget and generate new audio assets of different durations, allowing for efficient and autonomous production of similar audio assets with varying lengths while preserving key features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual crafting of audio assets is used to create variations, then audio quality and precision are improved, but time consumption and labor costs increase

Engineering Contradiction:
Improveaudio asset qualityVSAvoidtime consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system enables audio assets to generate their own variations automatically through the generative model. The model learns from the input audio asset and autonomously produces multiple variations with different durations, eliminating the need for manual crafting while maintaining quality consistency.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The generative model creates copies of the input audio asset with modified characteristics. By learning the underlying patterns from the original asset, the model generates multiple copies that preserve the essential features while introducing controlled variations in duration and other parameters.

Inventive Principle:
Principle #26Copying

2Productivity

If computational audio generation is used to reduce manual work, then productivity is improved, but complexity and processing time increase

Engineering Contradiction:
Improveaudio asset generation speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts the essential features and patterns from the input audio asset to create a simplified representation that the generative model can process. This extraction approach reduces the complexity of the generation task while maintaining the ability to produce high-quality variations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system controls the complexity by adjusting generation parameters such as duration scaling factors and variation intensity. These parameter changes allow the model to generate assets efficiently without requiring complex processing, balancing productivity with system simplicity.

Inventive Principle:
Principle #35Parameter changes

3Extent of automation

If existing computational solutions are used to generate audio assets, then automation is improved, but processing efficiency deteriorates

Engineering Contradiction:
Improveaudio generation automationVSAvoidprocessing time
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The generative model performs preliminary learning from the input audio asset during a training phase, storing the learned patterns for rapid generation. This preliminary action separates the complex learning process from the generation process, enabling fast automated production of audio variations without repeated heavy processing.

Inventive Principle:
Principle #10Preliminary action

4Manufacturing precision

If single output generation is used to ensure quality, then manufacturing precision is improved, but productivity deteriorates

Engineering Contradiction:
Improveaudio asset qualityVSAvoidgeneration throughput
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system merges multiple generation operations into a single batch processing workflow. The generative model can produce multiple audio assets with different durations and variations in one execution, combining what would otherwise require separate processing steps into a unified operation that maintains quality while increasing throughput.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12190851B2Audio generation methods and systems
Publication Date: 2025.01.07 SONY COMP ENTERTAINMENT EURO LTD
  • US12190851B2 patent drawing
  • US12190851B2 patent drawing
  • US12190851B2 patent drawing

AI summary

A method of generating audio assets, comprising the steps of: receiving an input audio asset having a first duration, generating an input image representative of the input audio asset, training a generative model on the input image and implementing the trained generative model to generate an output image representative of an output audio asset having a second duration different to the first duration, and generating the output audio asset based on the output image.