Multichannel Audio Metadata Compression via KLT Eigenchannels

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Karhunen-Loève Transform (KLT)-based multichannel audio coding methods require high metadata bitrates for maintaining audio quality, limiting their applicability to specific numbers of audio channels and being inefficient for arbitrary channel configurations, especially with advanced recording devices like the Eigenmike.

Innovation Solution

An apparatus and method utilizing a KLT-based pre-processor to transform input audio channels into eigenchannels, with a metadata re-arrangement unit that re-arranges metadata elements into multi-dimensional blocks for efficient encoding, employing machine learning to determine correlation values and optimize the re-arrangement scheme for increased compression ratios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional KLT-based audio coding is used to process multichannel audio signals, then audio quality can be maintained, but metadata bitrate becomes excessively high

Engineering Contradiction:
Improveaudio qualityVSAvoidmetadata bitrate
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the metadata into multiple categories (spatial parameters, spectral parameters, temporal parameters, etc.) and applies different compression techniques to each category. This segmentation allows targeted optimization of each metadata type, reducing overall bitrate while preserving audio quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of organization by arranging metadata elements in a two-dimensional array structure (metadata matrix) with rows representing different metadata types and columns representing different time frames. This dimensional transformation enables more efficient compression through block-based processing and exploitation of temporal correlations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If conventional KLT-based audio coding is used, then audio quality is preserved, but the system is limited to specific numbers of audio channels

Engineering Contradiction:
Improveaudio qualityVSAvoidchannel configuration flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal metadata structure that can accommodate any number of input channels and arbitrary microphone configurations. The KLT-based processing and metadata organization are designed to be channel-agnostic, allowing the same framework to handle 5.1, 7.1, Ambisonics, and custom microphone array configurations uniformly.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs dynamic adaptation of the metadata structure based on the actual input signal characteristics. The number of eigenchannels, spatial parameters, and metadata elements are dynamically determined by the input configuration rather than being fixed, enabling flexible adaptation to different channel counts and microphone arrangements.

Inventive Principle:
Principle #15Dynamics

3Quantity of substance

If Vector Quantizer compression is applied to metadata, then bitrate is reduced, but implementation complexity increases and codebook training becomes difficult

Engineering Contradiction:
Improvemetadata bitrateVSAvoidimplementation complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent extracts and removes redundant information from the metadata before compression. By identifying and eliminating correlated parameters, temporal redundancies, and insignificant details, the metadata size is reduced inherently, requiring less aggressive compression and avoiding the need for complex codebook-based techniques.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses simple, computationally efficient compression techniques (such as quantization and entropy coding) that are easy to implement and require minimal training data. These lightweight methods achieve adequate compression ratios without the implementation complexity and training requirements of Vector Quantizers.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

4Device complexity

If small VQ codebook size is used for metadata compression, then implementation is simpler, but representation quality deteriorates

Engineering Contradiction:
Improveimplementation simplicityVSAvoidmetadata representation quality
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary processing of metadata including normalization, filtering, and redundancy removal before compression. This preliminary action reduces the effective information content that needs to be represented, allowing smaller codebooks or simpler compression schemes to achieve the same representation quality that would otherwise require large codebooks.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3469589B1Apparatuses and methods for encoding and decoding a multichannel audio signal
Publication Date: 2024.06.19 HUAWEI TECH DUESSELDORF
  • EP3469589B1 patent drawingFigure 1
  • EP3469589B1 patent drawingFigure 2
  • EP3469589B1 patent drawingFigure 3

AI summary

The invention relates to an apparatus (110) for encoding an input audio signal, wherein the input audio signal comprises a plurality of input audio channels. The apparatus (110) comprises a KLT-based pre-processor (111) configured to transform the plurality of input audio channels into a plurality of eigenchannels and to provide metadata in the form of a plurality of metadata elements, wherein the metadata allows reconstructing the plurality of input audio channels on the basis of the plurality of eigenchannels, a metadata rearrangement unit (114) configured to re-arrange the plurality of metadata elements on the basis of a re-arrangement scheme into one or more metadata blocks, wherein each of the one or more metadata blocks is a multi-dimensional array, and a metadata encoder (115) configured to encode each of the one or more metadata blocks.