Spatial Audio Direction Quantization With Frame-Erasure Robustness
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio encoding systems fail to accurately represent spatially coherent features in synthetic sound scenes, such as 5.1 or 7.1 loudspeaker mixes, using direction and energy ratio parameters, and struggle with encoding audio objects with directional metadata like azimuth and elevation values, leading to non-uniform direction distributions and sensitivity to frame erasure errors.
Innovation Solution
The proposed system employs a directional index for audio objects, differential encoding with a prediction streak limiter, and audio object vector-based difference encoding, using a spherical quantizer and indexer to ensure accurate spatial audio parameterization and robust encoding of directional information, incorporating azimuth and elevation values, and adjusting quantization resolution based on spatial extent and available bitrate.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional direction and energy ratio parameters are used for spatial audio encoding, then the encoding process is simple, but the spatially coherent features in synthetic sound scenes cannot be accurately represented
Solution Approach 1:
The patent transforms the spatial representation from conventional direction vectors to a spherical coordinate system with azimuth and elevation parameters. This parameter transformation enables accurate representation of spatially coherent features while maintaining encoding efficiency through structured quantization of the spherical parameters.
Solution Approach 2:
The patent adds an elevation dimension to the traditional azimuth-only representation, creating a two-dimensional spherical coordinate system. This dimensional expansion allows accurate representation of sounds in three-dimensional space, including overhead and underhead positions, while the structured approach to this additional dimension manages the increased complexity.
2Manufacturing precision
If uniform quantization is used for azimuth and elevation values, then the encoding is straightforward, but non-uniform direction distributions occur
Solution Approach 1:
The patent applies different quantization resolutions to different regions of the spherical space. Higher resolution is allocated to regions where directional precision is more critical, while lower resolution is used in less critical regions. This local adaptation of quantization quality achieves uniform direction distribution while managing overall complexity.
Solution Approach 2:
The quantization resolution is made dynamic rather than static, allowing the encoder to adjust the number of bits allocated to azimuth and elevation based on the specific spatial configuration of audio objects. This dynamic allocation optimizes the distribution uniformity for each frame while adapting to changing scene complexity.
3Measurement precision
If high quantization resolution is used for directional parameters, then the spatial accuracy is improved, but the bitrate increases
Solution Approach 1:
The patent applies partial precision by using different quantization resolutions for different audio objects and different regions of the spherical space. Rather than uniformly applying high precision to all parameters, the system selectively applies higher resolution only where necessary, achieving adequate directional accuracy while controlling overall bitrate consumption.
Solution Approach 2:
The system dynamically changes the quantization parameters (number of bits for azimuth and elevation) based on the spatial extent and importance of audio objects. This parameter adaptation allows the encoder to maintain directional precision for critical objects while reducing precision for less important ones, optimizing the precision-bitrate tradeoff.
4Quantity of substance
If differential encoding is applied to reduce bitrate, then the bitrate efficiency is improved, but sensitivity to frame erasure errors increases
Solution Approach 1:
The patent prepares reference vectors in advance that can be used for error recovery. When frame erasure occurs, the decoder can utilize these pre-computed reference vectors and the structured spherical coordinate system to reconstruct missing directional information, cushioning against the impact of frame losses while maintaining bitrate efficiency through differential encoding.
Data Source
Figure 1
Figure 2a
Figure 2b
AI summary
A method for spatial audio signal encoding comprising: obtaining, for a first frame, a plurality of audio direction parameters, wherein each parameter comprises an elevation value and an azimuth value and wherein each parameter has an ordered position; determining whether, for a preceding frame, any of the plurality of audio direction parameters was differentially encoded based on a difference between the preceding frame parameter elevation value and a further preceding frame parameter elevation value and the preceding frame parameter azimuth value and a further preceding frame parameter azimuth value; generating, for any audio direction parameter which was not differentially encoded in the considered preceding frame, a differential parameter value based on a difference between the frame parameter elevation value and a preceding frame parameter elevation value and a difference between the frame parameter azimuth value and a preceding frame parameter azimuth value; generating for each of the plurality of audio direction parameters a difference parameter value based on a difference between the audio direction parameter and a rotated derived audio direction parameter; quantizing the difference between the audio direction parameter and a rotated derived audio direction parameter and the differential parameter value; and selecting for each of the plurality of audio direction parameters, either of the quantized difference or differential parameter value.