Flexible Frequency Time Partitioning Audio Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding techniques at low bit rates result in unnatural distortion and muffled sound due to large spectral holes and missing high frequencies, particularly when vector quantization is applied, which fails to effectively combine with window switching techniques for tonal and transient components.
Innovation Solution
The proposed techniques involve partitioning spectral holes and missing high frequencies into a band structure using threshold parameters for vector quantization, combining with adaptively varying transform block sizes to improve time and frequency resolution, and employing band partitioning procedures like hole-filling, frequency extension, and overlay to allocate bands efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If vector quantization is used to represent high frequency components with very few bits, then bitrate is reduced, but distortion occurs for tonal high frequency components
Solution Approach 1:
The patent changes the quantization parameter from scalar to vector-based, using code-vectors from a codebook to represent groups of spectral coefficients. This allows capturing the joint statistics and correlations of multiple coefficients simultaneously, improving representation accuracy for tonal components while maintaining low bitrate.
Solution Approach 2:
The patent uses a codebook containing pre-defined code-vectors that represent typical spectral patterns. Instead of encoding each coefficient individually, the system copies relevant code-vectors to represent spectral regions, efficiently capturing tonal characteristics with minimal bits.
2Productivity
If transform coding quantizes many coefficients to zero at low bit rates, then compression ratio is improved, but spectral holes produce unnatural distortion
Solution Approach 1:
The patent converts the harmful effect of quantization noise and spectral holes into a beneficial perceptual outcome by shaping the noise spectrum according to the masking threshold. Noise is deliberately placed in frequencies where it is masked by stronger components, transforming what would be audible distortion into imperceptible noise.
Solution Approach 2:
The patent applies different quantization strategies to different spectral regions based on their perceptual importance. Strong tonal components are preserved with higher precision, while weaker components in masked regions use coarser quantization or noise substitution, optimizing the balance between compression and quality.
3Ease of operation
If a fixed number of bands are allocated for spectral holes, then coding simplicity is maintained, but enough bands may not be left for missing high frequency data
Solution Approach 1:
The patent implements dynamic band allocation where the number of bands assigned to spectral holes versus high frequency extension varies based on the actual signal characteristics. The system analyzes the spectral content and adaptively adjusts the split, ensuring sufficient bands are allocated to high frequencies when needed while maintaining simple coding when the signal allows.
4Measurement precision
If larger transform window size is used for tonal components, then frequency resolution is improved, but time resolution of transients deteriorates
Solution Approach 1:
The patent segments the spectral data into multiple bands that can be processed with different window sizes. Tonal regions use larger windows for frequency resolution, while transient regions use smaller windows for time resolution. This segmentation allows each region to be optimized independently without compromising the other.
Solution Approach 2:
The patent applies different transform window sizes to different spectral bands based on the local signal characteristics. Bands containing transients use smaller windows to preserve time resolution, while bands with stationary tonal components use larger windows for better frequency resolution and compression.
Data Source
AI summary
An audio encoder/decoder performs band partitioning for vector quantization encoding of spectral holes and missing high frequencies that result from quantization when encoding at low bit rates. The encoder/decoder determines a band structure for spectral holes based on two threshold parameters: a minimum hole size threshold and a maximum band size threshold. Spectral holes wider than the minimum hole size threshold are partitioned evenly into bands not exceeding the maximum band size threshold in size. Such hole filling bands are configured up to a preset number of hole filling bands. The bands for missing high frequencies are then configured by dividing the high frequency region into bands having binary-increasing, linearly-increasing or arbitrarily-configured band sizes up to a maximum overall number of bands.


