Transient Steering Decorrelator for Spatial Audio Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio processing technologies, such as MPEG Surround, struggle to maintain high perceptual quality for applause-like signals at low bitrates due to the inability of lattice allpass decorrelators to accurately represent the spatio-temporal structure of such signals, leading to temporal smearing and loss of immersive sound characteristics.
Innovation Solution
The implementation of a Transient Steering Decorrelator (TSD) within the USAC decoder, which separates transient and non-transient components in the QMF domain and applies dedicated decorrelators to each, allowing for improved spatial distribution and preservation of transient events.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If lattice allpass decorrelators are used in MPEG Surround to generate decorrelated signals, then spatial audio quality is improved, but temporal signal structure and transient events are smeared and degraded
Solution Approach 1:
The patent segments the audio signal into transient and non-transient components using a transient detector and separator. Transient components are identified and separated from the main signal, allowing different decorrelation processing to be applied to each segment. This segmentation resolves the contradiction by preserving transient temporal structure while maintaining spatial quality through component-specific processing.
Solution Approach 2:
The patent applies different decorrelation characteristics to different signal components. Transient components receive specialized decorrelation treatment that preserves their temporal structure, while non-transient components receive standard lattice allpass decorrelation for optimal spatial quality. This local quality approach allows simultaneous optimization for both spatial audio quality and temporal signal stability.
2Reliability
If decorrelators are applied to generate artificial reverberation or improve echo cancellation, then spatial impression is enhanced, but temporal signal structure and convergence behavior are affected
Solution Approach 1:
The patent segments signals into transient and non-transient portions, applying different decorrelation strategies to each. For echo cancellation applications, this segmentation allows transient events to be processed separately, preserving their temporal characteristics and improving convergence behavior while still enhancing spatial impression through appropriate decorrelation of non-transient components.
3Productivity
If bitrate is reduced for audio coding, then coding efficiency is improved, but perceptual quality of applause-like signals deteriorates
Solution Approach 1:
The patent segments applause-like signals into transient and non-transient components, allowing efficient coding of each type. Transient components are coded with parameters that preserve their temporal structure and immersive characteristics, while non-transient components are coded for spatial quality. This segmentation enables high perceptual quality at low bitrates by avoiding redundant coding of transient temporal information.
Solution Approach 2:
The patent changes the coding parameters based on signal characteristics. Different decorrelation parameters and coding strategies are applied to transient versus non-transient components, optimizing the bitrate-quality tradeoff. By adapting parameters to the specific characteristics of each signal segment, high perceptual quality is maintained at reduced bitrates.
Data Source
AI summary
An apparatus for decoding, an apparatus for encoding, a method for decoding and a method for encoding positions of slots having events in an audio signal frame and respective computer programs and encoded signals, wherein the apparatus for decoding has: an analyzing unit for analyzing a frame slots number indicating the total of slots of the audio signal frame, an event slots number indicating the number of slots having the events of the audio signal frame, and an event state number, and a generating unit for generating an indication of a plurality of positions of slots having the events in the audio signal frame using the frame slots number, the event slots number and the event state number.


