Scalable Unified Audio Rendering for Low-Complexity Spatial Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio rendering technologies in computer-mediated reality systems face challenges in efficiently processing and rendering audio data to maintain coherence with video data, leading to potential audio artifacts and reduced immersion due to high processing demands and power consumption.
Innovation Solution
A scalable unified audio rendering technique that reduces processing complexity by transforming channel-based and object-based audio data into scene-based data, using spherical harmonic coefficients (SHC) and mixed order ambisonics (MOA) representations, allowing for efficient rendering across various speaker configurations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional channel-based and object-based audio rendering is used, then audio quality can be maintained, but processing complexity and power consumption increase significantly
Solution Approach 1:
The patent transforms audio data from traditional channel-based or object-based representations into scene-based data using spherical harmonic coefficients (SHC) and mixed order ambisonics (MOA). This parameter transformation enables the audio rendering system to process spatial audio information more efficiently while maintaining audio quality, thereby reducing processing complexity without compromising reliability.
2Reliability
If traditional audio rendering processing is used, then audio-visual coherence can be achieved, but power consumption increases reducing playback duration
Solution Approach 1:
By converting audio data into scene-based representations with SHC and MOA, the patent reduces the computational load required for audio rendering. This parameter transformation allows the system to maintain audio-visual coherence while consuming less power, thereby extending battery life and playback duration in mobile and wearable devices.
3Reliability
If high-fidelity audio rendering is performed, then user immersion is improved, but processing demands increase reducing system performance
Solution Approach 1:
The patent employs scene-based audio data representation using spherical harmonic coefficients and mixed order ambisonics, which transforms the audio processing workflow to reduce computational complexity. This enables high-fidelity audio rendering that enhances user immersion while maintaining system performance by reducing processing demands on the audio subsystem.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A device comprising an audio decoder, a memory, and a processor may be configured to perform various aspects of the techniques. The audio decoder may decode, from a bitstream, first audio data and second audio data. The memory may store the first audio data and the second audio data. The processor may render the first audio data into first spatial domain audio data for playback by virtual speakers at a set of virtual speaker locations, and render the second audio data into second spatial domain audio data for playback by the virtual speakers at the set of virtual speaker locations. The processor may also mix the first spatial domain audio data and the second spatial domain audio data to obtain mixed spatial domain audio data, and convert the mixed spatial domain audio data to scene-based audio data.