HOA Source Interpolation for Low-Bitrate 6DoF Audio Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current spatial audio capture techniques face challenges in meeting the requirements for both short and long wavelengths with a single microphone array, leading to limited spatial audio capture capabilities and high computational and bandwidth demands for 6 degrees of freedom (6DoF) rendering.
Innovation Solution
The generation of position-interpolated higher order ambisonics (HOA) sources during the rendering metadata creation phase, which reduces computational complexity and bandwidth requirements by generating additional HOA sources with spatial metadata, allowing for 6DoF rendering with a single HOA source.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a high-end microphone array is used for linear spatial audio capture, then spatial audio quality is improved, but device complexity and cost increase
Solution Approach 1:
The patent segments the spatial audio rendering task by separating capture (linear audio signals) from rendering (spatial processing). Instead of using a complex spherical microphone array for direct HOA capture, the system uses a simple linear array for capture and then applies computational rendering to generate the spatial audio experience, dividing the complexity between capture hardware and rendering software.
Solution Approach 2:
The patent replaces the mechanical complexity of a spherical microphone array with computational processing. Instead of physically capturing HOA signals with complex microphone geometry, the system captures linear audio signals and uses algorithms to synthesize the spatial audio experience, substituting mechanical complexity with computational processing.
2Measurement precision
If 6DoF rendering is implemented with multiple HOA sources, then rendering accuracy is improved, but computational complexity and bandwidth requirements increase
Solution Approach 1:
The patent extracts the spatial rendering computation from the audio capture process. By separating capture (linear audio) from rendering (spatial processing), the system reduces the computational burden on capture devices and enables flexible rendering at playback devices with varying computational capabilities.
Solution Approach 2:
The patent implements dynamic rendering where the number and positioning of HOA sources can be adjusted based on listener position and device capabilities. The system can dynamically select between different rendering modes (e.g., using 1-3 HOA sources instead of always using maximum sources) to adapt computational complexity to available resources while maintaining acceptable audio quality.
3Measurement precision
If multiple HOA sources are used for 6DoF rendering, then spatial audio quality is improved, but network bandwidth requirements increase
Solution Approach 1:
The patent applies partial action by using only the necessary number of HOA sources for the given listening position and spatial requirements. Instead of always transmitting and processing the maximum number of HOA sources, the system transmits and processes only 1-3 sources as needed, reducing bandwidth consumption while maintaining adequate spatial audio quality for the specific use case.
Data Source
AI summary
An apparatus for generating an immersive audio scene, the apparatus including circuitry configured to: obtain audio scene based sources, the audio scene based sources are associated with one or more positions in an audio scene, wherein each audio scene based source includes at least one spatial parameter and at least one audio signal; determine at least one position associated with at least one of the audio scene based sources; generate at least one audio source based on the determined at least one position, wherein the circuitry is configured to: generate at least one spatial audio parameter; and generate at least one audio source signal; and generate information about a relationship between the generated at least one spatial audio parameter and the at least one audio signals and the generated at least one audio source is selected based on a renderer preference.


