Spatial Audio Object Creation from Stereo Source Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio processing technologies struggle to effectively convert legacy audio content, mixed in stereo or mono formats, into spatially dynamic audio objects that preserve the original balance and intent of the mixing engineers, especially with the advent of spatial audio formats like Dolby Atmos and DTS-X.

Innovation Solution

An electronic device and method that analyze the results of stereo or multi-channel source separation to determine time-varying parameters, creating spatially dynamic audio objects by modifying audio object parameters such as side-mid ratio, and using monopole synthesis for rendering, which adapts positioning based on these parameters to create a more enveloping sound field.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing upmixing systems extract spectrally based features or add external effects to render legacy content spatially, then spatial audio rendering is achieved, but the original balance and intent of the mixing engineers is not preserved

Engineering Contradiction:
Improvespatial audio rendering capabilityVSAvoidoriginal mixing balance and intent
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system uses feedback from analyzing the original stereo mix to continuously adjust spatial positioning parameters. By monitoring the side-mid ratio and other time-varying parameters from the source separation results, the system feedback-adjusts the spatial rendering to maintain the original mixing balance while achieving spatial audio format compatibility.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system dynamically changes spatial rendering parameters based on time-varying parameters extracted from the audio signal. By modifying parameters such as spatial position, panning, and envelope characteristics in real-time according to the analyzed audio characteristics, the system preserves the original mixing intent while creating immersive spatial audio experiences.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If legacy audio content is converted to spatially dynamic audio objects, then immersive listening experience is improved, but the complexity of the processing system increases

Engineering Contradiction:
Improveimmersive listening experienceVSAvoidprocessing system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The processing system is segmented into distinct functional modules: source separation module, parameter analysis module, spatial rendering module, and feedback control module. Each module handles a specific aspect of the conversion process, making the overall complex task manageable and computationally efficient through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs dynamic spatial rendering that adapts to the content being processed. By making the spatial audio objects time-varying and responsive to the original audio characteristics, the system achieves high-quality immersive experiences while managing computational complexity through efficient dynamic adjustment rather than static over-engineering.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12170090B2Electronic device, method and computer program
Publication Date: 2024.12.17 SONY GROUP CORP
  • US12170090B2 patent drawing
  • US12170090B2 patent drawing
  • US12170090B2 patent drawing

AI summary

An electronic device comprising circuitry configured to analyze the results of a stereo or multi-channel source separation to determine one or more time-varying parameters, and to create spatially dynamic audio objects based on the one or more time-varying parameters.