Spatial Audio Encoding Spherical Grid Quantization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current spatial audio encoding methods struggle to accurately represent directional components, such as elevation and azimuth, for multi-channel loudspeaker inputs, leading to inefficiencies in spatial metadata quantization and encoding, particularly near the 'poles' of the direction sphere.
Innovation Solution
The proposed solution involves defining a spherical grid with smaller spheres arranged in circles, where the elevation and azimuth components of direction parameters are converted to index values using a defined grid structure, allowing for uniform quantization and encoding with higher density near the 'poles', and utilizing these index values for spatial audio reproduction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If uniform quantization of elevation and azimuth is used, then the encoding is simple, but the representation accuracy near the poles is insufficient
Solution Approach 1:
The patent applies local quality by using different quantization step sizes for different regions of the spherical coordinate system. Specifically, smaller quantization steps are used near the poles (elevation angles close to 0° and 180°) where directional changes are more perceptible, while larger steps are used at the equator. This resolves the contradiction by maintaining encoding simplicity through a structured approach while improving accuracy where it matters most perceptually.
Solution Approach 2:
The patent changes the quantization parameter (step size) based on the elevation angle. The quantization step for azimuth is made dependent on the elevation angle, with the step size decreasing as the elevation approaches the poles. This dynamic parameter adjustment resolves the contradiction by adapting the encoding precision to the local perceptual requirements at different regions of the sphere.
2Measurement precision
If higher density quantization states are used near the poles, then the spatial audio reproduction quality is improved, but the encoded information volume increases
Solution Approach 1:
The patent dynamically changes the quantization parameter (step size) based on the elevation angle to achieve higher density near poles while controlling information volume. By making the azimuth quantization step dependent on elevation (smaller steps near poles, larger at equator), the system achieves perceptually uniform resolution without uniformly increasing the number of quantization states everywhere, thus managing the information volume efficiently.
Solution Approach 2:
The patent applies partial action by increasing quantization density only in the regions where it is most needed (near the poles) rather than uniformly across the entire sphere. This selective approach improves spatial reproduction quality in critical regions while avoiding the information overhead that would result from uniform high-density quantization everywhere.
3Measurement precision
If perceptual quantization with varying step sizes is used, then the spatial perceptual quality is improved, but the encoding complexity increases
Solution Approach 1:
The patent implements local quality by applying different quantization characteristics to different regions of the spherical coordinate system. The encoding process uses elevation-dependent quantization steps, where the azimuth quantization step is determined by the current elevation angle. This structured regional differentiation improves perceptual accuracy without requiring completely complex adaptive algorithms, as the variation follows a systematic pattern based on elevation.
Solution Approach 2:
The patent manages encoding complexity by using a systematic parameter change approach where the quantization step size is determined by a formula based on elevation angle rather than requiring complex adaptive algorithms. The relationship between elevation and quantization step follows a defined pattern, making the implementation more manageable while still achieving perceptual optimization.
Data Source
Figure 1
Figure 2
Figure 3a
AI summary
An apparatus for spatial audio signal encoding, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: determine, for two or more audio signals, at least one spatial audio parameter for providing spatial audio reproduction, the at least one spatial audio parameter comprising a direction parameter with an elevation and an azimuth component; define a spherical grid generated by covering a sphere with smaller spheres, the smaller spheres arranged in circles of spheres wherein a first circle of spheres comprises one of the smaller spheres located with a centre at an elevation of 90 degrees relative to a reference direction of the sphere; and convert the elevation and azimuth component of the direction parameter to an index value based on the defined spherical grid.