Multi-channel Audio Encoding via Non-negative Tensor Factorization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional multi-channel audio encoding methods require significant data storage and transmission due to discrete bit stream encoding for each channel, leading to inefficiencies in data management and transmission.

Innovation Solution

The method involves transforming input signals into a frequency domain and performing non-negative tensor factorization to define object spectra, time-dependent gains, and channel-dependent gains, allowing for efficient parameterization and distribution of object spectra across channels, thereby reducing data redundancy and improving encoding efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If discrete bit stream encoding is used for each channel, then high quality sound representation is achieved, but data storage and transmission requirements increase significantly

Engineering Contradiction:
Improvesound qualityVSAvoiddata amount
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent merges multiple channel signals into a down-mixed signal for encoding, combining information from multiple channels into a single or reduced set of signals. This reduces the total data量 while maintaining the ability to reconstruct individual channels through spatial audio cues and object-based parameterization.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent extracts spatial audio cues and object spectra parameters from the multi-channel signal, separating the essential spatial information from the full channel data. This extracted parameter set is much smaller than the original multi-channel data but contains the key information needed for high-quality reconstruction.

Inventive Principle:
Principle #2Taking out (Extraction)

2Quantity of substance

If down-mixing with single set of spatial cues per time-frequency block is used, then data reduction is achieved, but ability to represent complex auditory scenes with multiple overlapping sound sources deteriorates

Engineering Contradiction:
Improvedata amountVSAvoidrepresentation of complex auditory scenes
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent segments the audio scene into multiple independent sound objects, each with its own spectral characteristics and spatial parameters. This segmentation allows each object to be represented separately with its own set of parameters, enabling accurate representation of complex scenes with multiple overlapping sources while using fewer parameters than full multi-channel encoding.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces object-based parameterization as an additional dimension beyond traditional time-frequency blocking. By organizing audio information into objects with spectral, temporal, and spatial dimensions, the system can represent complex auditory scenes more efficiently, capturing the essence of multiple sound sources without requiring discrete encoding of each channel.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS9978379B2Multi-channel encoding and/or decoding using non-negative tensor factorization
Publication Date: 2018.05.22 PIECE FUTURE PTE LTD
  • US9978379B2 patent drawing
  • US9978379B2 patent drawing
  • US9978379B2 patent drawing

AI summary

A method comprising: receiving input signals for multiple channels; and parameterizing the received input signals into parameters defining multiple different object spectra and defining a distribution of the multiple different object spectra in the multiple channels.