Object Audio Encoding via Ambisonics Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital audio technologies face challenges in efficiently encoding and decoding object audio in the Ambisonics domain, particularly due to bandwidth limitations when aiming for high spatial resolution, which can result in large audio footprints and compatibility issues.
Innovation Solution
The method involves converting object audio into time-frequency domain Ambisonics audio and encoding it along with metadata as bit streams, allowing for efficient storage and transmission, and decoding it back into object audio using spatial information, with priority-based encoding for optimal resolution and bandwidth management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If object audio is encoded with high spatial resolution in Ambisonics domain, then spatial audio quality is improved, but bandwidth consumption increases
Solution Approach 1:
The patent segments the Ambisonics audio data into different components (first channel data and second channel data) with different spatial resolutions. The first channel data maintains high spatial resolution for critical spatial information, while the second channel data uses lower spatial resolution, thereby reducing overall bandwidth consumption while preserving essential spatial audio quality.
Solution Approach 2:
The patent applies different spatial resolution qualities to different parts of the audio data. Specifically, the first channel data is encoded with high spatial resolution to preserve critical spatial information, while the second channel data is encoded with lower spatial resolution. This local differentiation optimizes the balance between spatial audio quality and bandwidth efficiency.
2Measurement precision
If object audio is encoded with high spatial resolution, then spatial audio quality is improved, but encoded footprint increases
Solution Approach 1:
The patent divides the encoded audio data into two segments: first channel data with high spatial resolution and second channel data with lower spatial resolution. This segmentation allows the encoded footprint to be reduced by using lower resolution for less critical components while maintaining high resolution for essential spatial information.
Solution Approach 2:
Different quality levels are applied locally to different channel data. The first channel data retains high spatial resolution quality, while the second channel data uses reduced quality encoding. This approach minimizes the total encoded footprint while preserving the spatial audio quality where it matters most.
3Measurement precision
If object audio is converted to time-frequency domain Ambisonics audio, then spatial audio reproduction is enhanced, but processing complexity increases
Solution Approach 1:
The patent segments the processing task into converting object audio to time-frequency domain Ambisonics audio and then separately processing different channel data with different resolutions. This segmentation simplifies the overall processing complexity by breaking down the complex transformation into manageable stages with different processing requirements.
4Productivity
If priority-based encoding is applied to different object audio, then bandwidth efficiency is improved, but encoding complexity increases
Solution Approach 1:
The patent applies priority-based encoding by assigning different spatial resolution qualities to different channel data based on their importance. The first channel data, considered more important for spatial audio reproduction, receives high spatial resolution encoding, while the second channel data receives lower spatial resolution encoding. This local quality differentiation improves bandwidth efficiency while the systematic approach keeps encoding complexity manageable.
Data Source
AI summary
In one aspect, a computer-implemented method, includes obtaining object audio and metadata that spatially describes the object audio, converting the object audio to Ambisonics audio based on the metadata, encoding, in a first bit stream, the Ambisonics audio, and encoding, in a second bit stream, at least a subset of the metadata.


