MDCT Stereo Audio Encoder Spectral Whitening
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding methods, particularly in MDCT-based systems, face challenges in efficiently processing panned signals and managing complex predictions across different frequencies, leading to high computational costs and inadequate handling of panned signals with varying panning in different frequencies.
Innovation Solution
A multi-channel audio encoder that applies spectral whitening to both separate-channel and mid-side representations, allowing for real or complex predictions to be made on these representations, and decides on the encoding based on the whitened signals to optimize bit allocation and reduce computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If complex prediction or coding of angle between mid and side channel is applied to handle panned signals, then the handling of panned signals is improved, but the computational cost increases
Solution Approach 1:
The patent applies spectral whitening to the mid and side channels before performing prediction or angle coding. This preliminary normalization of the spectral content simplifies subsequent processing operations, reducing the computational complexity while maintaining the ability to handle panned signals effectively.
Solution Approach 2:
The patent transforms the signal representation by applying spectral whitening, which changes the spectral parameters to have uniform power distribution. This parameter transformation simplifies the prediction process and reduces computational requirements for handling panned signals across different frequency ranges.
2Reliability
If band-wise M/S processing is applied, then stereo processing effectiveness is improved, but additional processing like complex prediction is required for panned signals
Solution Approach 1:
The patent applies spectral whitening as a preliminary step before band-wise M/S processing and prediction. This preprocessing normalizes the spectral content across bands, making the subsequent prediction operations more effective and reducing the complexity of handling panned signals in the frequency domain.
Solution Approach 2:
By transforming the mid and side channels through spectral whitening, the patent changes the spectral parameters to have flattened power spectra. This parameter transformation improves the effectiveness of band-wise processing while simplifying the prediction operations needed for panned signals.
3Productivity
If global ILD concept is used, then coding efficiency for panned channels is improved, but special whitening of M/S is required
Solution Approach 1:
The patent applies spectral whitening to the M/S channels as a preliminary step before implementing global ILD coding. This preprocessing ensures that the spectral content is normalized, which improves the efficiency of global ILD coding by providing a more uniform basis for inter-channel level difference calculation across all frequency bands.
Solution Approach 2:
The patent transforms the M/S channel parameters through spectral whitening, creating a normalized spectral representation. This parameter change improves global ILD coding efficiency by providing consistent spectral characteristics across bands, while the whitening operation itself becomes a standardized preprocessing step.
Data Source
AI summary
The invention refers to audio encoders, audio decoders, and audio encoding methods and audio decoding methods. In some examples, the invention refers to improved stereo coding. An encoder provides an encoded representation of an audio signal. The encoder applies a spectral whitening to a separate-channel representation of the input audio signal, to obtain a whitened separate-channel representation of the signal. The audio encoder applies a spectral whitening to a mid-side representation of the signal, to obtain a whitened mid-side representation of the signal. The audio encoder decides whether to encode the whitened separate-channel representation of the signal, to obtain the encoded representation of the signal, or to encode the whitened mid-side representation of the signal, to obtain the encoded representation of the signal.


