Unified Audio Coding Using Spherical Harmonic Coefficients
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current spatial audio coding technologies face challenges in efficiently encoding and transmitting complex sound fields, particularly with channel-based and object-based audio, due to bandwidth constraints and limitations in flexibility and audio quality, especially when dealing with multiple audio objects and varying acoustic conditions.
Innovation Solution
A unified audio coding approach that transforms channel-based and object-based audio into a hierarchical set of spherical harmonic basis function coefficients, allowing for scalable encoding and decoding independent of the number of audio objects, and adaptable to different loudspeaker geometries and acoustic conditions, using techniques such as time-frequency analysis and wavefront modeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If channel-based or object-based audio coding is used to represent sound fields, then audio quality and flexibility are improved, but bandwidth consumption and processing complexity increase
Solution Approach 1:
The patent transforms channel-based and object-based audio into spherical harmonic coefficients, changing the representation parameters from channel-specific or object-specific formats to a unified spatial basis function format. This parameter transformation enables flexible reconstruction for different loudspeaker configurations while maintaining manageable processing complexity through mathematical transformation.
Solution Approach 2:
Spherical harmonic coefficients serve as an intermediary representation between channel-based/object-based audio and final spatial sound field reconstruction. This intermediate form decouples the input format from the output configuration, allowing flexible adaptation to different loudspeaker geometries without reprocessing the original audio objects or channels.
2Adaptability or versatility
If multiple audio objects are encoded separately to maintain flexibility, then adaptability is improved, but bandwidth consumption increases
Solution Approach 1:
The patent merges multiple audio objects into a unified spherical harmonic coefficient representation. Instead of transmitting separate encoded streams for each audio object, the system combines them into a single set of coefficients that collectively describe the entire sound field, significantly reducing bandwidth consumption while preserving the ability to reconstruct individual objects or aggregate sound fields as needed.
3Reliability
If proprietary audio content is transmitted in original format, then audio quality is preserved, but content protection is compromised
Solution Approach 1:
The patent extracts the essential spatial and temporal characteristics of audio content by transforming them into spherical harmonic coefficients, separating the core auditory information from the proprietary source format. This extraction preserves audio quality for reconstruction purposes while preventing direct reversal to the original proprietary format, as the transformation to spherical harmonics is not easily invertible to the source representation.
Data Source
Figure 1A~1B
Figure 2A~2B
Figure 3A~3B
AI summary
Systems, methods, and apparatus for a unified approach to encoding different types of audio inputs are described.