Inter-Channel Time Difference Estimation for Stereo Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing stereo coding techniques are not suitable for conversational speech, particularly failing to accurately determine inter-channel time differences in scenarios with non-coincident microphones and reverberation, leading to suboptimal ambience synthesis and increased bit-rates.
Innovation Solution
A method that estimates inter-channel time differences by smoothing the cross-correlation spectrum based on spectral characteristics, using adaptive weighting procedures to enhance robustness and accuracy, and combines broadband and narrowband alignment parameters for optimal channel alignment, incorporating inter-channel time differences and phase differences to improve coding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If conventional parametric stereo coding is used with decorrelators, then the stereo image width is artificially reproduced, but the natural ambience of speech is not appropriately recreated and quality is inconsistent for different conversational scenarios
Solution Approach 1:
The patent implements dynamic switching between different stereo coding modes (joint stereo, intensity stereo, and parametric stereo) based on the detected signal characteristics. The system adapts the coding approach according to whether the signal contains speech, music, or mixed content, and whether microphones are coincident or non-coincident, thereby maintaining consistent quality across different conversational scenarios while preserving natural speech ambience.
2Measurement precision
If joint stereo coding is used with high frequency resolution, then time-frequency transformation is achieved, but the bit-rate increases and compatibility with low delay time domain processing is lost
Solution Approach 1:
The patent segments the frequency spectrum into multiple bands and processes each band separately using appropriate stereo coding techniques. Instead of applying high-frequency resolution joint stereo coding to the entire spectrum, the system divides frequencies into low, mid, and high bands, applying intensity stereo to lower bands and parametric stereo to higher bands where needed, thereby achieving sufficient frequency resolution while maintaining low bit-rate and compatibility with time domain processing.
3Reliability
If non-coincident microphones are used for recording, then spatial capture is improved, but inter-channel time difference estimation becomes inaccurate and ambience synthesis fails
Solution Approach 1:
The patent performs preliminary time alignment of channels from non-coincident microphones before proceeding with stereo coding. The system estimates the inter-channel time difference using cross-correlation techniques, applies appropriate delay compensation to align the channels, and then proceeds with ambience synthesis. This preliminary alignment action ensures that subsequent processing steps can accurately estimate spatial parameters and synthesize natural ambience despite the non-coincident microphone configuration.
Data Source
AI summary
An apparatus for estimating an inter-channel time difference between a first channel signal and a second channel signal, includes a signal analyzer for estimating a signal characteristic of the first channel signal or the second channel signal or both signals or a signal derived from the first channel signal or the second channel signal; a calculator for calculating a cross-correlation spectrum for a time block from the first channel signal in the time block and the second channel signal in the time block; a weighter for weighting a smoothed or non-smoothed cross-correlation spectrum to obtain a weighted cross correlation spectrum using a first weighting procedure or using a second weighting procedure depending on a signal characteristic estimated by the signal analyzer, wherein the first weighting procedure is different from the second weighting procedure; and a processor for processing the weighted cross-correlation spectrum to obtain the inter-channel time difference.


