Multi-Channel Audio Time Alignment for Spatial Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-channel audio coding methods fail to accurately represent time and amplitude differences between channels, leading to suboptimal audio quality in binaural and stereo signals, especially when channels have different time alignments across frequency bands and time instants.

Innovation Solution

A method that involves dividing multi-channel audio signals into spectral bands, selecting a leading channel based on event detection, determining time shift values, and time aligning channels to remove and restore time differences, ensuring accurate reproduction of spatial audio properties.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If multi-channel audio signals are coded separately, then each channel can be independently processed, but the bit-rate increases and coding efficiency decreases

Engineering Contradiction:
Improvecoding efficiencyVSAvoidbit-rate
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent combines multiple audio channels into a reduced set of channels (e.g., mid and side channels) by merging redundant information. This is achieved through linear transformation where the mid channel is the sum of left and right channels, and the side channel is the difference between them. This merging reduces the number of independent channels that need to be coded separately, improving coding efficiency while maintaining audio quality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a compact representation of multi-channel audio by copying and transforming channel information into a reduced format. The down-mixed signals serve as compressed copies that contain essential spatial and temporal information, which can then be expanded back to full multi-channel output at the decoder side using stored transformation matrices.

Inventive Principle:
Principle #26Copying

2Productivity

If down-mixing is applied to reduce bit-rate, then coding efficiency improves, but spatial audio quality and time difference accuracy deteriorate

Engineering Contradiction:
Improvecoding efficiencyVSAvoidtime difference accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary time alignment to the input channels before down-mixing by computing and compensating for channel delays. This preliminary action ensures that time differences between channels are corrected prior to the down-mixing operation, preventing degradation of spatial accuracy. The delay compensation is calculated based on cross-correlation analysis and applied as a pre-processing step.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary time alignment compensation mechanism that operates between the input channels and the down-mixing process. This intermediary step computes optimal delay values for each channel based on their correlation with a reference channel, and applies these delays to align the channels before they are combined into down-mixed signals, thereby preserving spatial information.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If channels are time-aligned at the encoder, then spatial accuracy improves, but additional processing complexity is introduced

Engineering Contradiction:
Improvespatial accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the time alignment process into distinct, manageable steps: cross-correlation computation, delay estimation, and delay application. Each step operates independently on specific signal components, allowing for efficient implementation. The cross-correlation is computed separately for each channel pair, and delays are applied channel-by-channel, reducing overall computational complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the temporal parameters of the audio signal by applying computable delay values to align channels. Instead of complex time-warping or resampling operations, the invention uses simple sample-shifting with calculated delay values derived from cross-correlation peaks. This parameter-based approach maintains spatial accuracy while minimizing processing complexity.

Inventive Principle:
Principle #35Parameter changes

4Quantity of substance

If parametric coding is used to compress spatial information, then bit-rate decreases, but reconstruction accuracy of spatial cues deteriorates

Engineering Contradiction:
Improvebit-rateVSAvoidspatial cue accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent incorporates feedback by storing and transmitting the down-mixing transformation matrices and time alignment parameters from the encoder to the decoder. This feedback mechanism allows the decoder to reconstruct the spatial configuration accurately by applying the same transformations and delay compensations that were used during encoding, thereby maintaining spatial cue accuracy despite compression.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent compresses spatial information by changing the representation parameters from full multi-channel waveforms to compact transformation matrices and delay values. These parameters describe the spatial relationships between channels in a compressed form that requires fewer bits to transmit while preserving the essential spatial cues needed for accurate reconstruction.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP2291841B1Method, apparatus and computer program product for providing improved audio processing
Publication Date: 2014.08.20 NOKIA CORP
  • EP2291841B1 patent drawingFigure 1
  • EP2291841B1 patent drawingFigure 2
  • EP2291841B1 patent drawingFigure 3

AI summary

An apparatus for performing improved audio processing may include a processor. The processor may be configured to divide respective signals of each channel of a multi-channel audio input signal into one or more spectral bands corresponding to respective analysis frames, select a leading channel from among channels of the multi-channel audio input signalfor at least one spectral band, determine a time shift value for at least one spectral band of at least one channel, and time align the channels based at least in part on the time shift value.