Spatial Audio Encoding Using Transport Metadata for Flexible Down-Mix
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current spatial audio coding systems face challenges in efficiently synthesizing higher-order Ambisonics signals from DirAC parameters and reference signals, particularly in flexible down-mix signaling configurations, leading to suboptimal bitrate usage and audio quality due to inflexible and bitrate-consuming transport channel configurations.
Innovation Solution
Incorporating explicit transport metadata that indicates directional properties of down-mix signals, allowing for adaptive generation of transport representations and metadata that specify how down-mix signals are generated, enabling flexible encoding and decoding of spatial audio representations while maintaining high audio quality and reducing bitrate requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If explicit transport metadata is incorporated to indicate directional properties of down-mix signals, then adaptability and audio quality are improved, but device complexity and bitrate increase
Solution Approach 1:
The transport metadata is segmented into different types (e.g., down-mix type indication, channel configuration metadata, directional properties) that can be selectively included or excluded based on the specific application requirements. This allows the system to maintain high adaptability while controlling complexity by only processing the necessary metadata segments for each particular case.
Solution Approach 2:
The transport metadata structure is designed to be universal and multi-functional, serving multiple purposes simultaneously: indicating down-mix type, channel configuration, directional properties, and enabling various decoding strategies. This universal metadata framework reduces overall complexity by consolidating multiple information functions into a single integrated structure rather than requiring separate metadata for each function.
2Adaptability or versatility
If explicit transport metadata is incorporated to indicate directional properties of down-mix signals, then adaptability and audio quality are improved, but bitrate increases
Solution Approach 1:
The transport metadata implementation uses local quality by applying different levels of metadata detail to different frequency bands and time segments. Rather than uniformly applying high-detail metadata across the entire audio signal, the system adapts the metadata precision and granularity to match the local characteristics of the audio content, thereby maintaining adaptability while reducing overall bitrate requirements.
3Adaptability or versatility
If flexible down-mix signaling configurations are used, then adaptability is improved, but manufacturing precision and audio quality deteriorate
Solution Approach 1:
The system implements dynamic adaptation where the decoding precision and processing intensity are adjusted in real-time based on the specific down-mix configuration indicated by the transport metadata. Rather than using fixed rigid processing, the apparatus dynamically selects appropriate decoding strategies and precision levels matched to each specific down-mix type, thereby maintaining high audio quality across diverse flexible configurations.
Data Source
AI summary
An apparatus for encoding a spatial audio representation representing an audio scene to obtain an encoded audio signal includes: a transport representation generator for generating a transport representation from the spatial audio representation, and for generating transport metadata related to the generation of the transport representation or indicating one or more directional properties of the transport representation; and an output interface for generating the encoded audio signal, the encoded audio signal including information on the transport representation, and information on the transport metadata.


