Spatial Audio Object Creation from Stereo Source Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing technologies struggle to effectively convert legacy audio content, mixed in stereo or mono formats, into spatially dynamic audio objects that preserve the original balance and intent of the mixing engineers, especially with the advent of spatial audio formats like Dolby Atmos and DTS-X.
Innovation Solution
An electronic device and method that analyze the results of stereo or multi-channel source separation to determine time-varying parameters, creating spatially dynamic audio objects by modifying audio object parameters such as side-mid ratio, and using monopole synthesis for rendering, which adapts positioning based on these parameters to create a more enveloping sound field.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing upmixing systems extract spectrally based features or add external effects to render legacy content spatially, then spatial audio rendering is achieved, but the original balance and intent of the mixing engineers is not preserved
Solution Approach 1:
The system uses feedback from analyzing the original stereo mix to continuously adjust spatial positioning parameters. By monitoring the side-mid ratio and other time-varying parameters from the source separation results, the system feedback-adjusts the spatial rendering to maintain the original mixing balance while achieving spatial audio format compatibility.
Solution Approach 2:
The system dynamically changes spatial rendering parameters based on time-varying parameters extracted from the audio signal. By modifying parameters such as spatial position, panning, and envelope characteristics in real-time according to the analyzed audio characteristics, the system preserves the original mixing intent while creating immersive spatial audio experiences.
2Ease of operation
If legacy audio content is converted to spatially dynamic audio objects, then immersive listening experience is improved, but the complexity of the processing system increases
Solution Approach 1:
The processing system is segmented into distinct functional modules: source separation module, parameter analysis module, spatial rendering module, and feedback control module. Each module handles a specific aspect of the conversion process, making the overall complex task manageable and computationally efficient through modular architecture.
Solution Approach 2:
The system employs dynamic spatial rendering that adapts to the content being processed. By making the spatial audio objects time-varying and responsive to the original audio characteristics, the system achieves high-quality immersive experiences while managing computational complexity through efficient dynamic adjustment rather than static over-engineering.
Data Source
AI summary
An electronic device comprising circuitry configured to analyze the results of a stereo or multi-channel source separation to determine one or more time-varying parameters, and to create spatially dynamic audio objects based on the one or more time-varying parameters.


