Spatial Audio Direction Quantization With Frame-Erasure Robustness

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio encoding systems fail to accurately represent spatially coherent features in synthetic sound scenes, such as 5.1 or 7.1 loudspeaker mixes, using direction and energy ratio parameters, and struggle with encoding audio objects with directional metadata like azimuth and elevation values, leading to non-uniform direction distributions and sensitivity to frame erasure errors.

Innovation Solution

The proposed system employs a directional index for audio objects, differential encoding with a prediction streak limiter, and audio object vector-based difference encoding, using a spherical quantizer and indexer to ensure accurate spatial audio parameterization and robust encoding of directional information, incorporating azimuth and elevation values, and adjusting quantization resolution based on spatial extent and available bitrate.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional direction and energy ratio parameters are used for spatial audio encoding, then the encoding process is simple, but the spatially coherent features in synthetic sound scenes cannot be accurately represented

Engineering Contradiction:
Improvespatial parameter accuracyVSAvoidencoding complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the spatial representation from conventional direction vectors to a spherical coordinate system with azimuth and elevation parameters. This parameter transformation enables accurate representation of spatially coherent features while maintaining encoding efficiency through structured quantization of the spherical parameters.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent adds an elevation dimension to the traditional azimuth-only representation, creating a two-dimensional spherical coordinate system. This dimensional expansion allows accurate representation of sounds in three-dimensional space, including overhead and underhead positions, while the structured approach to this additional dimension manages the increased complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Manufacturing precision

If uniform quantization is used for azimuth and elevation values, then the encoding is straightforward, but non-uniform direction distributions occur

Engineering Contradiction:
Improvedirection distribution uniformityVSAvoidquantization complexity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent applies different quantization resolutions to different regions of the spherical space. Higher resolution is allocated to regions where directional precision is more critical, while lower resolution is used in less critical regions. This local adaptation of quantization quality achieves uniform direction distribution while managing overall complexity.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The quantization resolution is made dynamic rather than static, allowing the encoder to adjust the number of bits allocated to azimuth and elevation based on the specific spatial configuration of audio objects. This dynamic allocation optimizes the distribution uniformity for each frame while adapting to changing scene complexity.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If high quantization resolution is used for directional parameters, then the spatial accuracy is improved, but the bitrate increases

Engineering Contradiction:
Improvedirectional parameter precisionVSAvoidbitrate
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies partial precision by using different quantization resolutions for different audio objects and different regions of the spherical space. Rather than uniformly applying high precision to all parameters, the system selectively applies higher resolution only where necessary, achieving adequate directional accuracy while controlling overall bitrate consumption.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system dynamically changes the quantization parameters (number of bits for azimuth and elevation) based on the spatial extent and importance of audio objects. This parameter adaptation allows the encoder to maintain directional precision for critical objects while reducing precision for less important ones, optimizing the precision-bitrate tradeoff.

Inventive Principle:
Principle #35Parameter changes

4Quantity of substance

If differential encoding is applied to reduce bitrate, then the bitrate efficiency is improved, but sensitivity to frame erasure errors increases

Engineering Contradiction:
Improvebitrate efficiencyVSAvoidrobustness to frame erasure
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent prepares reference vectors in advance that can be used for error recovery. When frame erasure occurs, the decoder can utilize these pre-computed reference vectors and the structured spherical coordinate system to reconstruct missing directional information, cushioning against the impact of frame losses while maintaining bitrate efficiency through differential encoding.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentEP4014234B1Quantization of spatial audio direction parameters
Publication Date: 2025.01.01 NOKIA TECHNOLOGIES OY
  • EP4014234B1 patent drawingFigure 1
  • EP4014234B1 patent drawingFigure 2a
  • EP4014234B1 patent drawingFigure 2b

AI summary

A method for spatial audio signal encoding comprising: obtaining, for a first frame, a plurality of audio direction parameters, wherein each parameter comprises an elevation value and an azimuth value and wherein each parameter has an ordered position; determining whether, for a preceding frame, any of the plurality of audio direction parameters was differentially encoded based on a difference between the preceding frame parameter elevation value and a further preceding frame parameter elevation value and the preceding frame parameter azimuth value and a further preceding frame parameter azimuth value; generating, for any audio direction parameter which was not differentially encoded in the considered preceding frame, a differential parameter value based on a difference between the frame parameter elevation value and a preceding frame parameter elevation value and a difference between the frame parameter azimuth value and a preceding frame parameter azimuth value; generating for each of the plurality of audio direction parameters a difference parameter value based on a difference between the audio direction parameter and a rotated derived audio direction parameter; quantizing the difference between the audio direction parameter and a rotated derived audio direction parameter and the differential parameter value; and selecting for each of the plurality of audio direction parameters, either of the quantized difference or differential parameter value.