KLT-Based Multichannel Audio Encoding for Arbitrary Channel Configurations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multichannel audio coding technologies are limited to specific numbers of audio channels and cannot efficiently process signals with an arbitrary number of channels or those acquired from arbitrarily arranged microphones, which is a constraint in emerging applications like immersive sound and virtual reality.

Innovation Solution

The use of a Karhunen-Loève Transform (KLT)-based apparatus for encoding and decoding multichannel audio signals, which transforms input audio channels into eigenchannels, selects a subset of eigenvectors based on eigenvalues, and encodes these eigenchannels along with metadata to reconstruct the original audio channels, allowing for flexible and efficient processing of signals with any number of channels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional multichannel audio coding schemes are used, then audio signals with specific channel configurations (5.1, 7.1, 22.2) can be encoded, but the system cannot process signals with an arbitrary number of channels or arbitrarily placed microphones

Engineering Contradiction:
Improveadaptability to arbitrary channel configurationsVSAvoidcoding scheme complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies universality by designing a coding scheme based on the Karhunen-Loève Transform that can handle any number of input channels and arbitrary microphone arrangements. The KLT-based approach transforms the audio signal into eigenchannels, allowing the system to universally process different channel configurations (5.1, 7.1, 22.2, or arbitrary numbers) through a single unified framework rather than requiring separate coding schemes for each configuration.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs parameter changes by utilizing the eigenvalues and eigenvectors obtained from the KLT transformation. By selecting a subset of eigenvectors based on eigenvalue thresholds and using these to reconstruct the audio signal, the system adapts to different channel configurations dynamically. The metadata encoding of eigenvalues and eigenvectors allows flexible parameter adjustment for arbitrary microphone arrangements.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If all eigenchannels are encoded, then complete audio information is preserved, but the bitrate increases significantly

Engineering Contradiction:
Improveaudio information completenessVSAvoidbitrate
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent applies the extraction principle by selecting only the most significant eigenchannels for encoding. The system extracts and identifies eigenchannels with eigenvalues above a certain threshold, discarding or reconstructing less significant eigenchannels from metadata. This selective extraction maintains essential audio information while significantly reducing the bitrate required for transmission and storage.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements partial action by encoding only a subset of eigenchannels rather than all of them. The system performs partial transformation and encoding of the audio signal, focusing computational resources and bitrate allocation on the most perceptually important eigenchannels, while using metadata to represent less critical components.

Inventive Principle:
Principle #16Partial or excessive action

3Manufacturing precision

If metadata for all eigenchannels is encoded, then perfect reconstruction is possible, but the encoding complexity and data volume increase

Engineering Contradiction:
Improvereconstruction accuracyVSAvoidencoding complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent extracts and encodes only the essential metadata parameters needed for reconstruction, such as eigenvalues and selected eigenvectors for significant eigenchannels. By identifying and encoding only the critical metadata components rather than all possible parameters, the system achieves good reconstruction accuracy with reduced encoding complexity and smaller data volume.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10916255B2Apparatuses and methods for encoding and decoding a multichannel audio signal
Publication Date: 2021.02.09 HUAWEI TECH DUESSELDORF
  • US10916255B2 patent drawing
  • US10916255B2 patent drawing
  • US10916255B2 patent drawing

AI summary

An input audio signal comprises a plurality of input audio channels. A KLT-based pre-processor transforms the plurality of input audio channels into a plurality of eigenchannels and provides metadata associated with the plurality of eigenchannels. Each eigenchannel is associated with an eigenvalue and an eigenvector. The metadata allows reconstructing the plurality of input audio channels on the basis of the plurality of eigenchannels. A selector selects a subset of the plurality of eigenvectors corresponding to a plurality of selected eigenchannels on the basis of a geometric mean of the eigenvalues. An eigenchannel encoder encodes the plurality of selected eigenchannels. A metadata encoder encodes the metadata.