Audio Encoder Peak Spectral Region Attenuation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The MDCT-based TCX codec in the EVS codec faces challenges at low bitrates, where high spectral peaks in the upper frequency band dominate the quantization process, leading to poor perceptual quality due to the quantization of all spectral coefficients except those above the CELP frequency, resulting in missing psychoacoustically relevant low-frequency signal portions.
Innovation Solution
The proposed solution involves preprocessing the audio signal to detect peak spectral regions in the upper frequency band, attenuating these peaks, and using shaping information from the lower frequency band to shape both the lower and upper frequency bands, thereby reducing the dominance of high spectral peaks and ensuring bits are allocated to perceptually important lower frequency ranges.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high spectral peaks in the upper frequency band are encoded with high precision, then encoding precision for high frequencies is improved, but bit allocation for low frequencies is insufficient leading to poor perceptual quality
Solution Approach 1:
The patent applies different spectral resolutions to different frequency regions. Specifically, it identifies peak spectral regions in the upper frequency band and encodes only those peaks with high precision, while using lower precision for non-peak regions. This local differentiation ensures that perceptually important features (peaks) receive adequate encoding precision without sacrificing overall bit allocation efficiency, thereby maintaining perceptual quality while reducing the dominance of high-frequency peaks in bit consumption.
2Manufacturing precision
If spectral coefficients in the upper frequency band are quantized with high precision, then high frequency encoding accuracy is improved, but low frequency signal portions are lost due to insufficient bit allocation
Solution Approach 1:
The patent dynamically changes the quantization parameter (spectral resolution) based on the local spectral characteristics. By detecting peak spectral regions and adjusting the encoding precision accordingly, the system allocates bits adaptively: high precision is applied only to detected peaks in the upper frequency band, while lower precision is used for non-peak regions. This parameter adaptation ensures that low-frequency signal portions receive adequate bit allocation and are not lost, while still maintaining encoding accuracy for perceptually important high-frequency peaks.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
An audio encoder for encoding an audio signal having a lower frequency band and an upper frequency band, comprises: a detector (802) for detecting a peak spectral region in the upper frequency band of the audio signal; a shaper (804) for shaping the lower frequency band using shaping information for the lower band and for shaping the upper frequency band using at least a portion of the shaping information for the lower band, wherein the shaper (804) is configured to additionally attenuate spectral values in the detected peak spectral region in the upper frequency band; and a quantizer and coder stage (806) for quantizing a shaped lower frequency band and a shaped upper frequency band and for entropy coding quantized spectral values from the shaped lower frequency band and the shaped upper frequency band.