Band-Wise Audio Downmixing to Avoid Phase Cancellation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for stereo-to-mono downmixing in audio signal processing, such as passive and active time-domain methods, suffer from phase cancellation effects and energy loss, and are suboptimal due to lack of frequency-dependent weighting, while frequency-domain methods like MPEG-H are not feasible for mobile communication codecs due to complexity and delay constraints.

Innovation Solution

A method involving a downmixer with a weighting value estimator, spectral weighter, converter, and mixer that performs spectral weighting followed by time-domain conversion, allowing individual channel processing and band-wise weighting based on target energy, enabling efficient and high-quality stereo-to-mono conversion without additional delay or complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If passive time-domain downmixing is used, then device complexity is low, but phase cancellation effects and energy loss occur degrading audio quality

Engineering Contradiction:
Improvedownmixing complexityVSAvoidaudio quality
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The downmixing process is segmented into multiple stages: spectral domain weighting applied independently to each channel's frequency spectrum, followed by time-domain conversion and mixing. This segmentation allows frequency-dependent processing without requiring full frequency-domain transformation of the entire signal, reducing complexity while improving audio quality through band-wise energy preservation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from purely time-domain processing to a hybrid approach by introducing spectral domain weighting. By applying weights in the frequency domain and then converting back to time domain, the system achieves frequency-dependent control without the complexity of full frequency-domain downmixing, effectively adding a dimensional aspect to the processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If frequency-domain methods like MPEG-H are used, then band-wise weighting and energy preservation are achieved, but complexity and processing delay increase making them unsuitable for mobile communication codecs

Engineering Contradiction:
Improveaudio qualityVSAvoiddownmixing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The complex frequency-domain transformation is segmented and replaced by applying spectral weights directly to the encoded spectral coefficients. This avoids the need for full inverse transform and re-encoding, achieving band-wise weighting with minimal complexity suitable for mobile communication constraints.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The spectral weighting is applied as a preliminary operation on the encoded spectral coefficients before time-domain conversion. By preparing the weighted spectral representation early in the decoding process, the system achieves frequency-dependent downmixing without requiring subsequent complex processing stages.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If frequency-domain methods like MPEG-H are used, then band-wise weighting and energy preservation are achieved, but processing delay increases making them unsuitable for mobile communication codecs

Engineering Contradiction:
Improveaudio qualityVSAvoidprocessing delay
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The downmixing operation is segmented into independent spectral weight applications on each channel's coefficients, followed by efficient time-domain conversion. This segmentation eliminates the need for lengthy frequency-domain processing chains, achieving fast processing suitable for real-time mobile communication applications.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses lightweight spectral weighting operations that can be applied directly to encoded coefficients without requiring expensive or time-consuming transform operations. This disposable approach to frequency-domain processing achieves the desired effect with minimal computational overhead and delay.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

4Device complexity

If stereo-to-mono downmix is performed without frequency-dependent weighting, then processing is simple, but phase cancellation effects severely degrade quality depending on the audio item

Engineering Contradiction:
Improvedownmixing complexityVSAvoidphase cancellation effects
Core Design Contradiction:
Device complexityVSObject-affected harmful factors

Solution Approach 1:

The patent applies different weighting factors to different frequency bands (spectral components) of each channel. This local quality approach allows frequency-dependent control of the downmixing process, preventing phase cancellation effects in critical frequency ranges while maintaining simplicity through direct coefficient weighting rather than complex signal processing.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3935630B1Audio downmixing
Publication Date: 2024.09.04 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP3935630B1 patent drawingFigure 1
  • EP3935630B1 patent drawingFigure 2
  • EP3935630B1 patent drawingFigure 3a

AI summary

A downmixer for downmixing a multi-channel signal having at least two channels, comprises: a weighting value estimator (100) for estimating band-wise weighting values for the at least two channels; a spectral weighter (200) for weighting spectral domain representations of the at least two channels using the band-wise weighting values; a converter (300) for converting weighted spectral domain representations of the at least two channels into time representations of the at least two channels; and a mixer (400) for mixing the time representations of the at least two chan- nels to obtain a down mix signal.