Audio Encoder Frequency Domain Spatial Parameter Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding methods require high processing power due to the need for inverse Fast Fourier Transform (IFFT) and extensive searches for maximum cross-correlation values in the time domain, which complicates the determination of spatial parameters like inter-channel phase difference and coherence.
Innovation Solution
Calculating a complex coherence value by summing cross-correlation function values in the frequency domain, eliminating the need for IFFT and allowing for simpler estimation of inter-channel phase difference and coherence, thereby reducing computational effort.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If inverse FFT and time domain maximum search are used to determine spatial parameters, then measurement precision is improved, but device complexity and processing power requirements increase
Solution Approach 1:
The patent applies dimensionality change by performing cross-correlation calculations in the frequency domain rather than transforming to the time domain via IFFT. The complex coherence value is calculated by summing cross-correlation function values across frequency bins, effectively solving the spatial parameter determination problem in a different dimensional space (frequency domain vs. time domain) while maintaining measurement precision and reducing processing complexity
Solution Approach 2:
The patent extracts only the essential information needed for spatial parameter determination by calculating a complex coherence value through summation of cross-correlation values in the frequency domain. This extraction approach avoids the computationally intensive IFFT operation and maximum search in time domain, removing unnecessary processing steps while preserving the critical spatial relationship information
2Measurement precision
If IFFT and extensive time domain processing are performed, then spatial parameter accuracy is improved, but processing time increases
Solution Approach 1:
The patent resolves the time-consuming nature of IFFT and time domain maximum search by performing the entire cross-correlation analysis in the frequency domain. The complex coherence value is obtained by summing cross-correlation function values across frequency bins, eliminating the need for time domain transformation and subsequent maximum search, thereby significantly reducing processing time while maintaining spatial parameter accuracy
Solution Approach 2:
The patent performs preliminary action by calculating the cross-correlation function in the frequency domain before any time domain transformation would be needed. By computing the complex coherence value through frequency domain summation upfront, the patent avoids subsequent IFFT operations and time domain processing, effectively performing the essential calculation in advance in a more efficient manner
Data Source
AI summary
The encoder transforms the audio signals (x(n),y(n)) from the time domain to audio signal (X(k),Y(k)) in the frequency domain, and determines the cross-correlation function (Ri, Pi) in the frequency domain. A complex coherence value (Qi) is calculated by summing the (complex) cross-correlation function values (Ri, Pi) in the frequency domain. The inter-channel phase difference (IPDi) is estimated by the argument of the complex coherence value (Qi), and the inter-channel coherence (ICi) is estimated by the absolute value of the complex coherence value (Qi).


