Audio Object Conversion to HOA via Spatial Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding and decoding technologies face challenges in converting audio data into Higher-Order Ambisonics (HOA) coefficients, leading to issues with backward compatibility and the ability to play audio with arbitrary speaker configurations, as well as inefficiencies in storage and transmission due to the need for multiple audio formats.
Innovation Solution
The proposed solution involves encoding audio data in its original format alongside spatial positioning vectors that enable conversion into HOA coefficients, allowing for playback with arbitrary speaker configurations while maintaining backward compatibility, by determining and including spatial vectors in the bitstream that allow for the generation of HOA soundfields.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If audio data is converted into HOA coefficients for arbitrary speaker configuration playback, then adaptability to different speaker configurations is improved, but backward compatibility with traditional multi-channel formats deteriorates
Solution Approach 1:
The audio data is segmented into two separate representations: traditional multi-channel audio data and HOA coefficients. Each representation serves a specific purpose - multi-channel data ensures backward compatibility with legacy systems, while HOA coefficients enable flexible playback on arbitrary speaker configurations. This segmentation allows both requirements to coexist without conflict.
Solution Approach 2:
The system implements a universal audio representation that can serve multiple functions. The same audio content is encoded in both traditional multi-channel format (for backward compatibility) and HOA format (for adaptability), allowing the audio data to be universally playable across different system types and speaker configurations.
2Adaptability or versatility
If multiple audio formats are maintained for different playback scenarios, then adaptability is improved, but storage and transmission efficiency deteriorates
Solution Approach 1:
The patent merges multiple audio representations into a single integrated audio object structure. Instead of storing separate audio files for different formats, the system combines multi-channel audio data and HOA coefficients into one unified audio object that contains both representations, reducing overall storage requirements and simplifying transmission.
Solution Approach 2:
The unified audio object structure serves multiple functions simultaneously - it can be decoded to traditional multi-channel output for legacy systems or converted to HOA coefficients for flexible spatial audio playback. This multi-functional design eliminates the need for multiple separate format files, improving storage and transmission efficiency.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A device obtains an object-based representation of an audio signal of an audio object. The audio signal corresponds to a time interval. Additionally, the device obtains a representation of a spatial vector for the audio object, wherein the spatial vector is defined in a Higher-Order Ambisonics (HOA) domain and is based on a first plurality of loudspeaker locations. The device generates, based on the audio signal of the audio object and the spatial vector, a plurality of audio signals. Each respective audio signal of the plurality of audio signals corresponds to a respective loudspeaker in a plurality of local loudspeakers at the second plurality of loudspeaker locations different from the first plurality of loudspeaker locations.