Audio Object Separation Using Adaptive Time-Frequency Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio object coding techniques face challenges in achieving optimal time-frequency representation for different audio objects, leading to suboptimal separation performance and crosstalk issues, as they often use fixed time-frequency resolutions that do not match the characteristics of the audio objects, resulting in impaired sound quality.
Innovation Solution
The proposed solution involves an audio decoder and method that allow for the selection of the most suitable time-frequency representation for each audio object based on its specific characteristics, using an Enhanced Side Information Estimator and Enhanced Object Separator to compute and apply object-specific side information, enabling improved separation performance and subjective quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If fixed time-frequency resolution is used for audio object coding, then device complexity is reduced, but separation performance deteriorates due to mismatch with audio object characteristics
Solution Approach 1:
The patent applies dynamics by transitioning from fixed time-frequency resolution to adaptive time-frequency resolution that changes according to audio object characteristics. The system dynamically selects different time-frequency resolutions based on the temporal and spectral properties of each audio object, allowing the representation to adapt to the specific requirements of different audio content types.
Solution Approach 2:
The patent implements local quality by applying different time-frequency resolutions to different audio objects or different portions of audio objects based on their local characteristics. This allows high temporal resolution for transient components and high spectral resolution for tonal components, optimizing the representation quality for each specific region of the audio signal.
2Measurement precision
If adaptive time-frequency representation is used for each audio object, then separation performance is improved, but device complexity increases
Solution Approach 1:
The patent applies parameter changes by modifying the time-frequency resolution parameters according to the characteristics of each audio object. The system changes parameters such as window size, overlap factor, and filter bank configuration to match the temporal and spectral properties of different audio objects, thereby optimizing separation performance without requiring a complete redesign of the system architecture.
3Measurement precision
If high bitrate transmission is used for multi-channel audio content, then audio quality is improved, but resource load increases
Solution Approach 1:
The patent applies parameter changes by using parametric techniques to represent audio objects efficiently. Instead of transmitting full-resolution multi-channel audio data, the system transmits compressed parametric information including time-frequency resolution indicators, spatial parameters, and object characteristics, which can then be used to reconstruct the audio objects at the receiver side with high quality but low bitrate.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An audio decoder is proposed for decoding a multi-object audio signal consisting of a downmix signal X and side information PSI. The side information comprises object-specific side information PSIi, for an audio object Si in a time/frequency region R(tR,fR), and object-specific time/frequency resolution information TFRIi indicative of an object-specific time/frequency resolution TFRh of the object-specific side information for the audio object Si in the time/frequency region Κ(tR,fR). The audio decoder comprises an object-specific time/frequency resolution determiner 110 configured to determine the object-specific time/frequency resolution information TFRIi from the side information PSI for the audio object Si . The audio decoder further comprises an object separator 120 configured to separate the audio object si from the downmix signal X using the object-specific side information in accordance with the object-specific time/frequency resolution TFRIi. A corresponding encoder and corresponding methods for decoding or encoding are also described.