Sparse Spectral Peak Coding for Low Bit Rate Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional perceptual audio coding techniques at low bit rates result in unnatural distortion and muffled sound due to large missing spectral portions, particularly affecting high frequency components like those from string instruments, as they are often quantized to zero, leading to coarse sounding noise.
Innovation Solution
An efficient coding scheme for sparse spectral peak data is developed, where spectral peaks are identified, predictively coded as shifts over time, and quantized with higher precision, avoiding large zero runs and using a joint trio quantization of zero runs and non-zero coefficients, thereby reducing bit rate impact and improving audio quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional scalar quantization is used to compress audio at low bit rates, then compression ratio is improved, but spectral data is lost causing unnatural distortion and muffled sound
Solution Approach 1:
The patent segments the spectrum into multiple frequency bands and applies different quantization strategies to each band. Important frequency bands with tonal components use finer quantization to preserve spectral data, while less important bands use coarser quantization to maintain compression ratio. This segmentation resolves the contradiction by selectively preserving spectral information where it matters most.
Solution Approach 2:
The patent applies local quality by allocating different quantization resolutions to different frequency regions based on their perceptual importance. Critical frequency bands containing tonal components receive higher precision quantization to prevent spectral loss and distortion, while non-critical bands use lower precision to maintain overall compression efficiency. This localized approach preserves essential spectral data without sacrificing compression ratio.
2Quantity of substance
If vector quantization is used to represent high frequency components with few bits, then bit rate is reduced, but tonal components are distorted into coarse sounding noise
Solution Approach 1:
The patent applies different quantization precision to different spectral regions. Tonal components in critical frequency bands are identified and assigned finer quantization steps to preserve their fidelity, while non-tonal or less important components use coarser vector quantization to reduce bit rate. This local differentiation maintains tonal component quality while achieving overall bit rate reduction.
Solution Approach 2:
The patent dynamically adjusts quantization parameters based on the characteristics of each frequency band. When tonal components are detected, the quantization step size is reduced (finer quantization) to preserve precision. When no tonal components are present, the quantization step size is increased (coarser quantization) to reduce bit rate. This parameter adaptation resolves the contradiction between bit rate and tonal fidelity.
3Productivity
If run-length coding is used to encode zero-level coefficients, then compression efficiency is improved, but large zero runs require escape codes increasing complexity
Solution Approach 1:
The patent modifies the run-length coding parameters by setting a maximum run length threshold. When zero runs exceed this threshold, the coding scheme transitions to an escape code mechanism. This parameter-based approach optimizes compression efficiency for typical cases while managing complexity through predictable fallback behavior, resolving the contradiction between compression efficiency and coding complexity.
Data Source
AI summary
An audio encoder/decoder provides efficient compression of spectral transform coefficient data characterized by sparse spectral peaks. The audio encoder/decoder applies a temporal prediction of the frequency position of spectral peaks. The spectral peaks in the transform coefficients that are predicted from those in a preceding transform coding block are encoded as a shift in frequency position from the previous transform coding block and two non-zero coefficient levels. The prediction may avoid coding very large zero-level transform coefficient runs as compared to conventional run length coding. For spectral peaks not predicted from those in a preceding transform coding block, the spectral peaks are encoded as a value trio of a length of a run of zero-level spectral transform coefficients, and two non-zero coefficient levels.


