Multi-channel Audio Encoding via Non-negative Tensor Factorization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional multi-channel audio encoding methods require significant data storage and transmission due to discrete bit stream encoding for each channel, leading to inefficiencies in data management and transmission.
Innovation Solution
The method involves transforming input signals into a frequency domain and performing non-negative tensor factorization to define object spectra, time-dependent gains, and channel-dependent gains, allowing for efficient parameterization and distribution of object spectra across channels, thereby reducing data redundancy and improving encoding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If discrete bit stream encoding is used for each channel, then high quality sound representation is achieved, but data storage and transmission requirements increase significantly
Solution Approach 1:
The patent merges multiple channel signals into a down-mixed signal for encoding, combining information from multiple channels into a single or reduced set of signals. This reduces the total data量 while maintaining the ability to reconstruct individual channels through spatial audio cues and object-based parameterization.
Solution Approach 2:
The patent extracts spatial audio cues and object spectra parameters from the multi-channel signal, separating the essential spatial information from the full channel data. This extracted parameter set is much smaller than the original multi-channel data but contains the key information needed for high-quality reconstruction.
2Quantity of substance
If down-mixing with single set of spatial cues per time-frequency block is used, then data reduction is achieved, but ability to represent complex auditory scenes with multiple overlapping sound sources deteriorates
Solution Approach 1:
The patent segments the audio scene into multiple independent sound objects, each with its own spectral characteristics and spatial parameters. This segmentation allows each object to be represented separately with its own set of parameters, enabling accurate representation of complex scenes with multiple overlapping sources while using fewer parameters than full multi-channel encoding.
Solution Approach 2:
The patent introduces object-based parameterization as an additional dimension beyond traditional time-frequency blocking. By organizing audio information into objects with spectral, temporal, and spatial dimensions, the system can represent complex auditory scenes more efficiently, capturing the essence of multiple sound sources without requiring discrete encoding of each channel.
Data Source
AI summary
A method comprising: receiving input signals for multiple channels; and parameterizing the received input signals into parameters defining multiple different object spectra and defining a distribution of the multiple different object spectra in the multiple channels.


