Spatial Audio Metadata Encoding Using DCT Codebooks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio encoding technologies face challenges in effectively encoding spatial audio metadata, particularly for inputs other than microphone-array captured signals, such as loudspeaker signals and Ambisonic signals, with existing methods focusing on direction and energy ratio parameters but lacking efficient compression of coherence data.

Innovation Solution

The proposed solution involves an apparatus and method that determine a codebook for encoding spread and surround coherence values based on energy ratio and azimuth values, using discrete cosine transformation and entropy encoding to efficiently compress and transmit spatial audio metadata, including coherence parameters, across various audio signal formats.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If coherence parameters are encoded using traditional methods, then encoding simplicity is maintained, but encoding efficiency and compression ratio deteriorate

Engineering Contradiction:
Improveencoding efficiencyVSAvoidencoding complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by transforming coherence parameters from the time domain to the frequency domain using Discrete Cosine Transform (DCT). This transformation changes the representation of the data, concentrating energy in fewer coefficients, which enables more efficient entropy coding and achieves better compression ratios without significantly increasing processing complexity.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the coherence parameters into different frequency sub-bands using DCT, allowing independent processing and coding of each sub-band. This segmentation enables selective coding strategies where important low-frequency components are preserved while less important high-frequency components are compressed more aggressively, improving overall encoding efficiency.

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If spatial audio metadata is compressed to reduce bitrate, then transmission efficiency is improved, but reconstruction quality deteriorates

Engineering Contradiction:
ImprovebitrateVSAvoidreconstruction quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent uses DCT to transform coherence parameters into a frequency domain representation where energy is concentrated in fewer coefficients. This parameter change enables selective retention of important coefficients during compression, allowing high-quality reconstruction even at reduced bitrates by preserving the most significant spectral components.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional time-domain compression methods with frequency-domain transformation and entropy coding. This substitution allows for more efficient information packing by exploiting the statistical properties of DCT coefficients, achieving better compression-to-quality ratio through mathematical transformation rather than simple quantization.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If coherence parameters are encoded with high precision, then spatial audio fidelity is improved, but data size increases

Engineering Contradiction:
Improvespatial audio fidelityVSAvoiddata size
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent transforms coherence parameters using DCT, which concentrates the information into fewer, more significant coefficients. This parameter transformation enables high-fidelity representation with reduced data size by retaining only the most important spectral components rather than preserving all time-domain samples at full precision.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent moves the representation of coherence parameters from the time domain to the frequency domain through DCT. This dimensional transformation reveals the underlying spectral structure of the data, allowing for more efficient compression while maintaining fidelity by preserving the essential spectral characteristics rather than all temporal details.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12009001B2Determination of spatial audio parameter encoding and associated decoding
Publication Date: 2024.06.11 NOKIA TECHNOLOGIES OY
  • US12009001B2 patent drawing
  • US12009001B2 patent drawing
  • US12009001B2 patent drawing

AI summary

An apparatus comprising means for: receiving values for sub-bands of a frame of an audio signal, the values comprising at least one azimuth value, at least one elevation value at least one energy ratio value and at least one spread and/or surround coherence value for each sub-band; determining a codebook for encoding at least one spread and/or surround coherence value for each sub-band based on the at least one energy ratio value and at least one azimuth value for each sub-band for a frame; discrete cosine transforming at least one vector, the at least one vector comprising the at least one spread and/or surround coherence value for a sub-band for the frame; and encoding a first number of components of the discrete cosine transformed vector based on the determined codebook.