Audio Signal Decoding Temporal Shift Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In audio signal decoding, the temporal offset between audio signals captured by multiple microphones leads to high entropy in side channel signals, requiring more bits for encoding and reducing coding efficiency due to misalignment and variations in shift estimates, causing sample repetition and artifact skipping.
Innovation Solution
A method that estimates and adjusts for temporal shifts between audio signals from multiple microphones using non-causal shifting, generating encoded signals with reduced bit usage by aligning the target audio signal with the reference signal, and dynamically adjusting shift values based on sound source location and microphone configurations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If temporal offset between audio signals is not corrected, then encoding is simpler, but side channel signal entropy increases requiring more bits for encoding
Solution Approach 1:
The patent applies preliminary action by estimating the temporal offset between audio signals from multiple microphones before encoding. The system calculates shift estimates for different frame types (voiced, unvoiced, transition frames) and applies non-causal shifting to align the signals in advance, thereby reducing side channel entropy before the encoding process begins.
2Measurement precision
If different frame types use different shift estimates, then frame-specific accuracy improves, but sample repetition and artifact skipping occur at frame boundaries
Solution Approach 1:
The patent applies dynamics by making the shift estimation adaptive to different frame types. The system dynamically adjusts shift estimates based on whether frames are voiced, unvoiced, or transition frames, allowing optimal alignment for each frame type while maintaining overall signal continuity through careful boundary handling.
3Productivity
If temporal alignment is improved, then coding efficiency increases, but decoding complexity increases due to non-causal shifting
Solution Approach 1:
The patent applies preliminary action by performing the complex temporal alignment and non-causal shifting operations during the encoding phase rather than during decoding. The encoder pre-calculates and applies shift estimates to align audio signals, reducing side channel energies and improving coding efficiency, while the decoder receives pre-aligned signals requiring minimal processing.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An apparatus includes a receiver configured to receive at least one encoded signal that includes inter-channel bandwidth extension (BWE) parameters. The device also includes a decoder configured to generate a mid channel time-domain high-band signal by performing bandwidth extension based on the at least one encoded signal. The decoder is also configured to generate, based on the mid channel time-domain high-band signal and the inter-channel BWE parameters, a first channel time-domain high-band signal and a second channel time-domain high-band signal. The decoder is further configured to generate a target channel signal by combining the first channel time-domain high-band signal and a first channel low-band signal, and to generate a reference channel signal by combining the second channel time-domain high-band signal and a second channel low-band signal. The decoder is also configured to generate a modified target channel signal by modifying the target channel signal based on a temporal mismatch value.