High-Frequency Envelope Processing for Transient Audio Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing technologies face challenges in encoding transient signals like applause and percussive sounds due to pre-echo artifacts, which are difficult to manage with current perceptual coding methods, especially at low bit-rates, leading to reduced audio quality.
Innovation Solution
The implementation of High Resolution Envelope Processing (HREP) which involves pre-processing the audio signal to attenuate high-frequency transient events and post-processing to restore the original dynamics, using time-variable high frequency gain information as side information to reduce the transient nature of the signal, thereby reducing the bit-rate demand and improving coding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If quantization is applied in the spectral domain using filterbank decomposition, then coding efficiency is improved, but quantization noise spreads out in time causing pre-echo artifacts
Solution Approach 1:
The patent applies temporal noise shaping in the spectral domain before quantization to pre-flatten the temporal envelope of transient signals. This preliminary action redistributes quantization noise in the spectral domain such that when transformed to time domain, the noise does not create pre-echo artifacts before transient onsets, while still maintaining coding efficiency.
Solution Approach 2:
The patent changes the parameter distribution of quantization noise from uniform to non-uniform in the spectral domain. By applying temporal noise shaping that modifies spectral coefficients based on their temporal characteristics, the noise is concentrated in frequency regions where it will be masked or less perceptible, resolving the pre-echo problem while maintaining spectral coding efficiency.
2Object-affected harmful factors
If coding precision is increased for transient signal portions, then pre-echo artifacts are reduced, but bit rate consumption increases
Solution Approach 1:
The patent applies local quality by differentiating treatment of spectral coefficients based on their temporal characteristics. Transient-associated spectral coefficients receive different quantization precision compared to stationary signal coefficients. Temporal noise shaping selectively shapes noise in spectral regions corresponding to transient events, reducing pre-echo artifacts only where needed without increasing overall bit rate.
Solution Approach 2:
The patent changes the quantization parameter distribution dynamically based on signal characteristics. By using temporal noise shaping to modify spectral coefficients before quantization, the system achieves reduced pre-echo artifacts through parameter optimization rather than uniform precision increase, thereby maintaining constant bit rate while improving transient handling.
3Loss of time
If short window sizes are used for transient coding, then temporal resolution is improved, but frequency resolution deteriorates
Solution Approach 1:
The patent applies dynamics by using adaptive window switching that changes window size based on signal characteristics. For transient signals, short windows are used to achieve good temporal resolution and minimize pre-echo. For stationary signals, long windows provide better frequency resolution. The system dynamically adapts window size to signal content, resolving the resolution trade-off.
Solution Approach 2:
The patent compensates for frequency resolution loss in short-window transient coding by applying temporal noise shaping in the spectral domain. This adds another dimension of processing that redistributes noise energy spectrally to compensate for the reduced frequency resolution, effectively recovering frequency information that would otherwise be lost with short windows.
Data Source
Figure 1
Figure 2
Figure 3a~3c
AI summary
An audio post-processor (100) for post-processing an audio signal (102) having a time-variable high frequency gain information (104) as side information comprises: a band extractor (110) for extracting a high frequency band (112) of the audio signal (102) and a low frequency band (114) of the audio signal (102); a high band processor (120) for performing a time-variable modification of the high frequency band (112) in accordance with the time-variable high frequency gain information (104) to obtain a processed high frequency band (122); and a combiner (130) for combining the processed high frequency band (122) and the low frequency band (114). Furthermore, a pre-processor for analyzing an audio signal to determine a time-variable high frequency gain information, perform modification of an high frequency band, and output a signal comprising the pre-processed audio signal and the high frequency gain information.