Spatial Audio Encoder Bit-Rate Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing spatial audio codecs face challenges in efficiently encoding and decoding multi-channel audio signals that represent spatial audio images, particularly in scenarios with limited bandwidth, where accurate representation of spatial metadata requires excessive bit-rate, making it difficult to maintain high-quality reconstruction of directional sound components and ambience.

Innovation Solution

The spatial audio encoder processes multi-channel input audio signals into downmix signals and transforms them into encoded audio and spatial metadata, using techniques like short-time discrete Fourier transform and complex-modulated quadrature-mirror filters, while selectively encoding spatial parameters based on energy levels and criteria to optimize bit allocation for accurate reconstruction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If spatial metadata is encoded with high precision to accurately represent directional sound components and ambience, then reconstruction quality is improved, but bit-rate increases excessively

Engineering Contradiction:
Improvespatial metadata precisionVSAvoidbit-rate
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies parameter changes by dynamically adjusting the precision of spatial metadata encoding based on the characteristics of the audio signal. Different quantization parameters are used for different types of spatial information (e.g., direction of arrival, spatial dispersion) depending on their perceptual importance and the complexity of the audio scene, thereby achieving high reconstruction quality without excessive bit-rate

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements local quality by applying different encoding precision to different spatial parameters and different time-frequency regions. High-precision encoding is applied only where necessary (e.g., for prominent directional components), while lower precision is used for less critical information, optimizing the trade-off between quality and bit-rate

Inventive Principle:
Principle #3Local quality

2Measurement precision

If all spatial parameters are encoded to maintain accurate spatial audio image, then reconstruction quality is improved, but complexity of encoding and decoding increases

Engineering Contradiction:
Improvespatial audio reconstruction qualityVSAvoidencoding and decoding complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and encodes only the most essential spatial parameters needed for accurate reconstruction, such as direction of arrival and spatial dispersion characteristics. Less critical parameters are either omitted or represented more compactly, reducing encoding/decoding complexity while maintaining perceptual quality

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the spatial audio signal into distinct components (directional sound components and ambience) and processes each with appropriate spatial parameter encoding. This segmentation allows the decoder to reconstruct the spatial image more efficiently by handling different components with dedicated processing paths

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If variable bit-rate encoding is used to optimize quality for different audio content, then perceptual quality is improved, but difficulty of detecting and measuring signal characteristics increases

Engineering Contradiction:
Improveperceptual sound qualityVSAvoidsignal characteristic analysis
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent employs feedback mechanisms where the encoder analyzes the encoded bitstream and adjusts spatial parameter encoding precision based on the actual bit-rate consumption and the perceived quality impact. This feedback loop enables variable bit-rate operation while automatically adapting to maintain optimal perceptual quality without requiring complex manual analysis

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3762923B1Audio coding
Publication Date: 2024.07.10 NOKIA TECHNOLOGIES OY
  • EP3762923B1 patent drawingFigure 1~2
  • EP3762923B1 patent drawingFigure 3
  • EP3762923B1 patent drawingFigure 4

AI summary

According to an example embodiment, a method for encoding a multi-channel input audio signal that represents an audio scene as an encoded audio signal and spatial audio parameters, wherein the spatial audio parameters are descriptive of said audio scene is provided, the method comprising: encoding a frame of a downmix signal into a frame of the encoded audio signal, wherein the downmix signal is generated from the multi-channel input audio signal; deriving, from the frame of the multi-channel input audio signal, a plurality of spatial audio parameters that are descriptive of the audio scene, said spatial audio parameters comprising a plurality of direction of arrival (DOA) parameters, wherein a DOA parameter indicates a spatial position of a given directional sound component of the audio scene in a given frequency sub-band; and encoding said spatial audio parameters, comprising encoding a DOA parameter for a given directional sound component in a given frequency sub-band in dependence of an energy level of the given directional sound component in the given frequency sub-band meeting one or more criteria.