Spatial Audio Augmentation Mixing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing spatial audio technologies face challenges in seamlessly integrating augmentation audio signals with spatial audio signals, particularly in immersive audio environments like virtual reality, where user movement affects audio rendering.
Innovation Solution
The proposed solution involves an apparatus and method that obtain spatial audio signals and augmentation audio signals, render them consistently with user movement, and mix them to generate an output audio signal. This is achieved by decoding spatial and augmentation audio signals from separate bit streams, applying user-locked or world-locked mixing modes, and adjusting gains based on user position and rotation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If spatial audio signals are rendered consistently with user movement in immersive environments, then audio fidelity and immersion quality are improved, but device complexity and processing requirements increase
Solution Approach 1:
The audio processing system is segmented into separate spatial audio signal processing path and augmentation audio signal processing path. The spatial audio decoder processes spatial audio signals with full 6DoF movement compensation, while the augmentation audio decoder handles augmentation signals with simplified processing. This segmentation allows high-fidelity spatial audio rendering without requiring the entire system to handle maximum processing complexity simultaneously.
Solution Approach 2:
The system dynamically adjusts the rendering approach based on the type of audio signal being processed. Spatial audio signals receive dynamic rendering that fully compensates for 6DoF user movement, while augmentation audio signals use a different rendering approach. This dynamic differentiation maintains audio fidelity for spatial content while managing overall processing complexity.
2Manufacturing precision
If separate bit streams are decoded for spatial and augmentation audio signals, then audio quality is improved, but processing time and computational resources increase
Solution Approach 1:
The spatial audio signal and augmentation audio signal are decoded from separate bit streams in parallel preliminary processing stages. By preparing both signals beforehand through separate decoding paths, the system avoids sequential processing delays and can mix the rendered outputs more efficiently, reducing overall processing time while maintaining high audio quality.
Solution Approach 2:
After separate decoding and rendering of spatial and augmentation audio signals, the processed signals are merged in the mixing stage. This combining of separately processed high-quality signals achieves both goals: maintains audio quality through dedicated processing paths while reducing processing time through parallel operation before the final mix.
3Adaptability or versatility
If user-locked or world-locked mixing modes are applied, then adaptability to different usage scenarios is improved, but device complexity increases
Solution Approach 1:
Different mixing modes (user-locked and world-locked) are applied locally to different audio signal components based on their specific requirements. Spatial audio signals can use one mixing mode while augmentation audio signals use another, allowing tailored processing for each signal type without requiring the entire system to support all modes simultaneously, thus managing complexity while maintaining versatility.
4Manufacturing precision
If gains are adjusted based on user position and rotation, then audio realism and immersion are improved, but computational load increases
Solution Approach 1:
The system applies gain adjustment based on user position and rotation selectively to spatial audio signals where it most impacts realism, rather than uniformly to all audio content. This partial application of the computationally intensive gain adjustment process maintains audio realism for spatial content while reducing overall computational energy consumption by not applying the same level of processing to all audio signals.
Data Source
AI summary
An apparatus configured to: obtain at least one spatial audio signal, the at least one spatial audio signal comprising at least one audio signal and at least one spatial parameter; determine at least one first audio signal based, at least partially, on the at least one spatial audio signal and the at least one of the user position or orientation; obtain at least one augmentation audio signal, wherein the at least one augmentation audio signal has a different audio format than an audio format of the at least one spatial audio signal, wherein the at least one augmentation audio signal provides a different type of media content; determine at least one second audio signal based, at least partially, on at least a part of the at least one augmentation audio signal; and mix the at least one first audio signal and the at least one second audio signal.


