Audio Generation System Using Machine Learning Models for Sound Source Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio generation systems for virtual reality and other immersive applications face challenges in creating high-quality, varied audio that accurately localizes sound sources, leading to repetitive and less immersive experiences due to limitations in sound separation and masking techniques.

Innovation Solution

The system employs machine learning models to separate and generate new audio representations of sound sources, using techniques like discriminative algorithms and generative adversarial networks to isolate and recreate specific sound components, allowing for improved localization and variation in audio content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional sound separation and masking techniques are used, then audio content can be generated, but the audio becomes repetitive and lacks immersion due to limitations in sound separation quality

Engineering Contradiction:
Improveaudio immersion qualityVSAvoidsound separation technique complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical signal processing techniques (masking, subtraction-based separation) with machine learning models including autoencoders, generative adversarial networks, and discriminative algorithms. This substitution enables more accurate sound source separation and generation of varied audio content, directly resolving the contradiction between immersion quality and technique complexity by using intelligent systems instead of conventional mechanical processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the audio generation approach by changing parameters from fixed masking thresholds and subtraction operations to learnable parameters in neural networks. The system learns optimal separation parameters through training data, enabling dynamic adaptation to different audio scenarios and producing more immersive and varied audio content while managing complexity through parameter optimization.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If numerous audio variations are provided to increase variability, then user immersion improves, but the complexity of audio content management increases

Engineering Contradiction:
Improveaudio variabilityVSAvoidaudio content management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements self-service through generative models that automatically create audio variations without manual intervention. The system uses trained machine learning models to generate diverse audio content on-demand, eliminating the need for extensive manual audio asset creation and management. This resolves the contradiction by enabling high adaptability through automated generation while reducing management complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses copying principles through generative adversarial networks where a generator model creates copies/variations of audio content based on learned patterns from training data. This allows the system to produce numerous varied audio versions from a single source, increasing adaptability while keeping content management simple through automated synthesis rather than manual duplication.

Inventive Principle:
Principle #26Copying

3Measurement precision

If sound sources are accurately localized and separated, then audio quality improves, but the processing complexity and computational requirements increase

Engineering Contradiction:
Improvesound source localization accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical signal processing chains with integrated machine learning models that perform separation and localization jointly. The system uses end-to-end trained neural networks that achieve accurate sound source localization without requiring multiple separate processing stages, thereby improving measurement precision while managing processing complexity through unified model architecture.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Reliability

If traditional audio generation methods are used, then processing is simpler, but the audio quality and realism are insufficient for immersive applications

Engineering Contradiction:
Improveaudio qualityVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent systematically replaces traditional mechanical audio processing with intelligent machine learning systems. The use of autoencoders for separation, GANs for generation, and discriminative models for quality assessment creates a comprehensive intelligent processing chain that achieves superior audio quality and realism, resolving the contradiction by demonstrating that increased model complexity through AI substitution yields proportional gains in audio quality for immersive applications.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11653167B2Audio generation system and method
Publication Date: 2023.05.16 SONY INTERACTIVE ENTERTAINMENT LLC
  • US11653167B2 patent drawing
  • US11653167B2 patent drawing
  • US11653167B2 patent drawing

AI summary

A system for generating audio content in dependence upon an input audio track comprising audio corresponding to one or more sound sources, the system comprising an audio input unit operable to input the input audio track to one or more models, each representing one or more of the sound sources, and an audio generation unit operable to generate, using the one or more models, one or more audio tracks each comprising a representation of the audio contribution of the corresponding sound sources of the input audio track, wherein the generated audio tracks comprise one or more variations relative to the corresponding portion of the input audio track.