Multi-Channel Audio Encoder Parameter Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing parametric stereo audio coding technologies face challenges in achieving stable and fast parameter estimation, particularly for inter-aural time difference (ITD) and channel level difference (CLD), which leads to instability and poor tracking behavior, especially at low bitrates.
Innovation Solution
The approach involves using both strong and weak smoothing techniques on cross-correlation and energy functions for ITD and CLD estimation, respectively, and employing a quality criterion to select the most reliable encoding parameters, allowing for stable and reactive parameter estimation by switching between long-term and short-term evaluations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If only one full band ITD parameter is transmitted to reduce bitrate overhead, then bitrate efficiency is improved, but parameter estimation stability deteriorates
Solution Approach 1:
The patent segments the parameter estimation process into two independent smoothing paths: strong smoothing for stability and weak smoothing for reactivity. Each path processes the ITD parameter independently, allowing the system to select the most appropriate estimation based on current audio conditions, thereby maintaining stability even with limited bitrate overhead.
Solution Approach 2:
The patent changes the smoothing parameter dynamically by providing two distinct smoothing coefficients (strong and weak) and selecting between them based on a quality criterion. This allows the system to adapt the estimation stability and reactivity according to the actual audio scene changes, resolving the contradiction between stability and bitrate efficiency.
2Reliability
If strong smoothing is applied to improve parameter estimation stability, then stability is improved, but tracking behavior deteriorates
Solution Approach 1:
The patent makes the smoothing strength dynamic by providing two distinct smoothing paths (strong and weak) and selecting between them based on a quality criterion that detects actual source position changes. This dynamic adaptation allows the system to switch between stability-oriented and reactivity-oriented estimation, resolving the contradiction between stability and tracking behavior.
Solution Approach 2:
The system periodically evaluates the quality criterion to determine whether to switch between strong and weak smoothing modes. This periodic assessment allows the system to adapt to changing audio conditions, maintaining stability during static scenes and improving tracking during dynamic scenes.
3Speed
If weak smoothing is applied to improve tracking behavior, then reactivity is improved, but parameter estimation stability deteriorates
Solution Approach 1:
The patent makes the smoothing parameter dynamic by providing two distinct smoothing paths (strong and weak) and selecting between them based on a quality criterion. This allows the system to adapt the estimation stability and reactivity according to the actual audio scene changes, resolving the contradiction between stability and reactivity.
Solution Approach 2:
The patent changes the smoothing parameter dynamically by providing two distinct smoothing coefficients (strong and weak) and selecting between them based on a quality criterion. This allows the system to adapt the estimation stability and reactivity according to the actual audio scene changes, resolving the contradiction between stability and reactivity.
4Loss of time
If small frame size is used to improve time resolution, then time resolution is improved, but parameter estimation reliability deteriorates
Solution Approach 1:
The patent merges two smoothing paths (strong and weak) into a unified selection mechanism based on quality criterion evaluation. This combination allows the system to leverage both short-term (weak smoothing) and long-term (strong smoothing) information, achieving reliable parameter estimation even with small frame sizes that provide good time resolution.
Data Source
AI summary
The invention relates to a method for determining an encoding parameter for an audio channel signal of a multi-channel audio signal, the method comprising: determining for the audio channel signal a set of functions from the audio channel signal and a reference audio signal; determining a first set of encoding parameters based on a smoothing of the set of functions with respect to a frame sequence of the multi-channel audio signal, the smoothing being based on a first smoothing coefficient; determining a second set of encoding parameters based on a smoothing of the set of functions with respect to the frame sequence of the multi-channel audio signal, the smoothing being based on a second smoothing coefficient; and determining the encoding parameter based on a quality criterion with respect to the first set of encoding parameters and/or the second set of encoding parameters.


