KLT-Based Multichannel Audio Encoding for Arbitrary Channel Configurations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multichannel audio coding technologies are limited to specific numbers of audio channels and cannot efficiently process signals with an arbitrary number of channels or those acquired from arbitrarily arranged microphones, which is a constraint in emerging applications like immersive sound and virtual reality.
Innovation Solution
The use of a Karhunen-Loève Transform (KLT)-based apparatus for encoding and decoding multichannel audio signals, which transforms input audio channels into eigenchannels, selects a subset of eigenvectors based on eigenvalues, and encodes these eigenchannels along with metadata to reconstruct the original audio channels, allowing for flexible and efficient processing of signals with any number of channels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional multichannel audio coding schemes are used, then audio signals with specific channel configurations (5.1, 7.1, 22.2) can be encoded, but the system cannot process signals with an arbitrary number of channels or arbitrarily placed microphones
Solution Approach 1:
The patent applies universality by designing a coding scheme based on the Karhunen-Loève Transform that can handle any number of input channels and arbitrary microphone arrangements. The KLT-based approach transforms the audio signal into eigenchannels, allowing the system to universally process different channel configurations (5.1, 7.1, 22.2, or arbitrary numbers) through a single unified framework rather than requiring separate coding schemes for each configuration.
Solution Approach 2:
The patent employs parameter changes by utilizing the eigenvalues and eigenvectors obtained from the KLT transformation. By selecting a subset of eigenvectors based on eigenvalue thresholds and using these to reconstruct the audio signal, the system adapts to different channel configurations dynamically. The metadata encoding of eigenvalues and eigenvectors allows flexible parameter adjustment for arbitrary microphone arrangements.
2Loss of information
If all eigenchannels are encoded, then complete audio information is preserved, but the bitrate increases significantly
Solution Approach 1:
The patent applies the extraction principle by selecting only the most significant eigenchannels for encoding. The system extracts and identifies eigenchannels with eigenvalues above a certain threshold, discarding or reconstructing less significant eigenchannels from metadata. This selective extraction maintains essential audio information while significantly reducing the bitrate required for transmission and storage.
Solution Approach 2:
The patent implements partial action by encoding only a subset of eigenchannels rather than all of them. The system performs partial transformation and encoding of the audio signal, focusing computational resources and bitrate allocation on the most perceptually important eigenchannels, while using metadata to represent less critical components.
3Manufacturing precision
If metadata for all eigenchannels is encoded, then perfect reconstruction is possible, but the encoding complexity and data volume increase
Solution Approach 1:
The patent extracts and encodes only the essential metadata parameters needed for reconstruction, such as eigenvalues and selected eigenvectors for significant eigenchannels. By identifying and encoding only the critical metadata components rather than all possible parameters, the system achieves good reconstruction accuracy with reduced encoding complexity and smaller data volume.
Data Source
AI summary
An input audio signal comprises a plurality of input audio channels. A KLT-based pre-processor transforms the plurality of input audio channels into a plurality of eigenchannels and provides metadata associated with the plurality of eigenchannels. Each eigenchannel is associated with an eigenvalue and an eigenvector. The metadata allows reconstructing the plurality of input audio channels on the basis of the plurality of eigenchannels. A selector selects a subset of the plurality of eigenvectors corresponding to a plurality of selected eigenchannels on the basis of a geometric mean of the eigenvalues. An eigenchannel encoder encodes the plurality of selected eigenchannels. A metadata encoder encodes the metadata.


