Rotated Spatial Audio Direction Quantization for Coherent Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing spatial audio encoding technologies are limited in their ability to effectively encode and decode sound-field related parameters from various input types, including microphone-array captured signals, loudspeaker signals, and audio objects, without accurately representing the spatial coherence of sounds.
Innovation Solution
The proposed apparatus and method for spatial audio signal encoding and decoding involve deriving and rotating indexed audio direction parameters, quantizing the rotation, and encoding spatial utilization and permutation indices to efficiently represent and reconstruct spatial audio signals from diverse input types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If spatial audio direction parameters are encoded using existing quantization methods, then the encoding process is simple, but the spatial coherence of sounds from diverse input types cannot be accurately represented
Solution Approach 1:
The patent transforms the encoding approach by changing the parameter representation method. Instead of directly quantizing azimuth and elevation angles, the invention derives new parameters (rotation angle and spatial extent) that better capture spatial coherence. The rotation angle represents the orientation of the spatial distribution, while spatial extent captures the spread, enabling accurate representation of sounds from diverse input types including microphone arrays, loudspeakers, and audio objects.
Solution Approach 2:
The invention introduces an additional dimensional transformation by rotating the coordinate system. The spatial distribution is reoriented so that the major axis aligns with the rotation angle, and the spatial extent is measured along this rotated axis. This dimensional transformation allows the encoding to capture spatial coherence properties that are not apparent in the standard azimuth-elevation representation.
2Adaptability or versatility
If the encoder processes multiple input types (microphone-array, loudspeaker, audio objects), then the encoder versatility is improved, but the complexity of accurately representing spatial coherence from these diverse inputs increases
Solution Approach 1:
The patent creates a universal encoding framework that handles multiple input types through a common parameter derivation process. The means for deriving rotation angle and spatial extent are designed to work with microphone-array captured signals, loudspeaker signals, and audio objects with directional metadata alike. This universal approach allows the encoder to process diverse inputs while maintaining consistent spatial coherence representation, avoiding the need for separate processing pipelines for each input type.
3Measurement precision
If directional parameters are quantized with high resolution, then the spatial accuracy is improved, but the bit rate increases
Solution Approach 1:
The invention changes the parameterization from direct azimuth-elevation quantization to a rotated coordinate system with spatial extent. This transformation allows for more efficient bit allocation because the rotation angle and spatial extent capture the essential spatial information with fewer bits. The spatial extent parameter, in particular, allows for coarser quantization while maintaining perceptual accuracy, reducing the overall bit rate requirement compared to high-resolution direct angle quantization.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method for spatial audio signal encoding comprising: obtaining a plurality of audio direction parameters, wherein each parameter comprises an elevation value and an azimuth value and wherein each parameter has an ordered position; deriving for each of the plurality of audio direction parameters a corresponding derived audio direction parameter (SP) comprising an elevation and an azimuth value, corresponding derived audio direction parameters (SP) being arranged in a manner determined by a spatial utilization defined by the elevation values and the azimuth values of the plurality of audio direction parameters; rotating each derived audio direction parameter (SP) by the azimuth value (φ0) of an audio direction parameter in the first position of the plurality of audio direction parameters and quantizing the rotation to determine for each a corresponding quantized rotated derived audio direction parameter; changing the ordered position of an audio direction parameter to a further position coinciding with a position of a rotated derived audio direction parameter when the azimuth value of the audio direction parameter is closest to the azimuth value of the further rotated derived audio direction parameter compared to the azimuth values of other rotated derived audio direction parameters, followed by determining for each of the plurality audio direction parameters a difference between each audio direction parameter and their corresponding quantized rotated derived audio direction parameter; and quantizing a difference for each of the plurality of audio direction parameters, wherein a difference quantization resolution for each of the plurality of audio direction parameters is defined based on a spatial extent of the audio direction parameters.