Audio Bitstream Encoding with Spatial Positioning Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding and decoding technologies face challenges in converting multi-channel audio data into Higher-Order Ambisonics (HOA) coefficients, leading to issues with backward compatibility and efficient playback across various speaker configurations, as they often require multiple versions of audio formats, increasing bandwidth and storage needs, and may not support arbitrary speaker configurations.
Innovation Solution
The proposed solution involves encoding multi-channel audio data with spatial positioning vectors (SPVs) in its original format, allowing for conversion into HOA coefficients, enabling playback with arbitrary speaker configurations while maintaining backward compatibility by including SPVs in the bitstream, which are used to generate HOA soundfields and render audio signals for local loudspeakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple versions of audio formats are used to support different speaker configurations, then backward compatibility is improved, but bandwidth and storage requirements increase
Solution Approach 1:
The patent encodes spatial positioning vectors (SPVs) alongside multi-channel audio data in a universal bitstream format that can be decoded by both legacy systems and HOA-capable systems. This allows a single audio bitstream to serve multiple speaker configurations (5.1, 7.1, and arbitrary configurations) without requiring separate encoded versions, thereby reducing bandwidth and storage while maintaining backward compatibility.
Solution Approach 2:
The patent introduces spatial positioning vectors (SPVs) as intermediary data elements that bridge legacy multi-channel audio formats and HOA representations. These SPVs act as a mediator that enables HOA-capable systems to interpret and render audio for arbitrary speaker configurations while legacy systems can still decode the original multi-channel audio without the SPV information, thus resolving the contradiction between compatibility and efficiency.
2Ease of operation
If channel-based audio is used, then ease of operation is improved, but adaptability to arbitrary speaker configurations deteriorates
Solution Approach 1:
The patent transforms static channel-based audio into a dynamic representation by encoding spatial positioning vectors that describe the positions of virtual sound sources relative to the listener. This dynamic spatial information allows the same audio data to be adaptively rendered for any speaker configuration by calculating appropriate rendering matrices, thereby achieving both ease of operation (audio data remains in original format) and high adaptability (arbitrary configurations supported).
Solution Approach 2:
The patent adds a spatial dimension to traditional channel-based audio by incorporating directional information through spatial positioning vectors. This dimensional enhancement allows the audio system to transition from fixed channel assignments to a three-dimensional soundfield representation, enabling arbitrary speaker configurations while maintaining the simplicity of the original audio data structure.
3Adaptability or versatility
If HOA coefficients are generated from multi-channel audio, then adaptability to arbitrary speaker configurations is improved, but device complexity increases
Solution Approach 1:
The patent performs preliminary encoding of spatial positioning vectors alongside the multi-channel audio data during the encoding stage. This preliminary action stores the necessary spatial information in a compact form within the bitstream, eliminating the need for complex real-time HOA conversion and rendering operations at the playback device. The decoding process simply involves extracting the SPVs and applying pre-computed rendering matrices, significantly reducing device complexity while maintaining adaptability.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
In one example, a method includes obtaining a representation of a multi-channel audio signal for a source loudspeaker configuration; obtaining a representation of a plurality of spatial positioning vectors (SPVs), in a Higher-Order Ambisonics (HOA) domain, that are based on a source rendering matrix, which is based on the loudspeaker configuration; and generating a HOA soundfield based on the multi-channel audio signal and the plurality of spatial positioning vectors.