Audio Signal Decoding Temporal Shift Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In audio signal decoding, the temporal offset between audio signals captured by multiple microphones leads to high entropy in side channel signals, requiring more bits for encoding and reducing coding efficiency due to misalignment and variations in shift estimates, causing sample repetition and artifact skipping.

Innovation Solution

A method that estimates and adjusts for temporal shifts between audio signals from multiple microphones using non-causal shifting, generating encoded signals with reduced bit usage by aligning the target audio signal with the reference signal, and dynamically adjusting shift values based on sound source location and microphone configurations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If temporal offset between audio signals is not corrected, then encoding is simpler, but side channel signal entropy increases requiring more bits for encoding

Engineering Contradiction:
Improveencoding complexityVSAvoidside channel signal entropy
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent applies preliminary action by estimating the temporal offset between audio signals from multiple microphones before encoding. The system calculates shift estimates for different frame types (voiced, unvoiced, transition frames) and applies non-causal shifting to align the signals in advance, thereby reducing side channel entropy before the encoding process begins.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If different frame types use different shift estimates, then frame-specific accuracy improves, but sample repetition and artifact skipping occur at frame boundaries

Engineering Contradiction:
Improveshift estimate accuracyVSAvoidframe boundary continuity
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies dynamics by making the shift estimation adaptive to different frame types. The system dynamically adjusts shift estimates based on whether frames are voiced, unvoiced, or transition frames, allowing optimal alignment for each frame type while maintaining overall signal continuity through careful boundary handling.

Inventive Principle:
Principle #15Dynamics

3Productivity

If temporal alignment is improved, then coding efficiency increases, but decoding complexity increases due to non-causal shifting

Engineering Contradiction:
Improvecoding efficiencyVSAvoiddecoding complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by performing the complex temporal alignment and non-causal shifting operations during the encoding phase rather than during decoding. The encoder pre-calculates and applies shift estimates to align audio signals, reducing side channel energies and improving coding efficiency, while the decoder receives pre-aligned signals requiring minimal processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3430622B1Two-channel audio signal decoding
Publication Date: 2021.07.14 QUALCOMM INC
  • EP3430622B1 patent drawingFigure 1
  • EP3430622B1 patent drawingFigure 2
  • EP3430622B1 patent drawingFigure 3

AI summary

An apparatus includes a receiver configured to receive at least one encoded signal that includes inter-channel bandwidth extension (BWE) parameters. The device also includes a decoder configured to generate a mid channel time-domain high-band signal by performing bandwidth extension based on the at least one encoded signal. The decoder is also configured to generate, based on the mid channel time-domain high-band signal and the inter-channel BWE parameters, a first channel time-domain high-band signal and a second channel time-domain high-band signal. The decoder is further configured to generate a target channel signal by combining the first channel time-domain high-band signal and a first channel low-band signal, and to generate a reference channel signal by combining the second channel time-domain high-band signal and a second channel low-band signal. The decoder is also configured to generate a modified target channel signal by modifying the target channel signal based on a temporal mismatch value.