Audio Encoding Subspace Projections Noise Shaping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal encoding methods at low bit-rates suffer from high computational complexity and perceptual degradation due to correlated quantization noise, leading to a muffled sound quality, especially in speech signals where higher frequencies are often quantized to zero.
Innovation Solution
The proposed solution employs subspace projections for noise shaping using uniform quantization combined with an iterative delta-coding scheme or arithmetic coding, which reduces the correlation between quantization noise and the original signal, allowing for perceptual noise shaping and minimizing energy loss in higher frequencies through dithering and Wiener filtering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If uniform quantization with arithmetic coding is used at low bit-rates, then coding efficiency is improved, but quantization noise becomes highly correlated to the original signal causing muffled sound quality
Solution Approach 1:
The patent applies preliminary dithering before quantization to randomize the quantization error and break the correlation between quantization noise and the original signal. This preliminary action prevents the muffled sound quality that would otherwise result from correlated quantization noise at low bit-rates.
Solution Approach 2:
The patent modifies the quantization process by introducing dither signals and using perceptual noise shaping that adjusts quantization parameters based on the spectral characteristics of the signal. This allows maintaining coding efficiency while reducing the harmful correlation between quantization noise and the original signal.
2Measurement precision
If arithmetic coding with rate-loop is used to scale accuracy to available bit-rate, then coding precision is improved, but computational complexity increases significantly
Solution Approach 1:
The patent extracts the essential function of rate-adaptive coding by using fixed-rate quantization with perceptual noise shaping, removing the need for complex arithmetic coding loops and rate-scaling mechanisms while maintaining coding precision through perceptual optimization.
Solution Approach 2:
The patent replaces complex arithmetic coding with simpler quantization schemes that use perceptual models to achieve similar precision. The complexity is reduced by using lightweight dithering and fixed-rate coding instead of computationally intensive arithmetic coding loops.
3Loss of information
If higher frequencies are quantized to zero to reduce bit-rate, then loss of information is reduced, but perceptual quality deteriorates due to loss of high-frequency content
Solution Approach 1:
The patent changes the quantization parameters dynamically based on perceptual importance, using perceptual noise shaping to allocate quantization precision according to the masking threshold. This ensures that quantization noise is shaped to be less perceptible, maintaining perceptual quality even when high frequencies are coarsely quantized.
Solution Approach 2:
The patent converts the harmful effect of quantization noise into a beneficial perceptual outcome by shaping the noise spectrum to follow the masking threshold. Quantization noise that would normally be harmful is transformed into perceptually acceptable noise that is masked by the signal itself, allowing aggressive bit-rate reduction without perceptual degradation.
Data Source
Figure 1~3
Figure 4
Figure 5
AI summary
An apparatus for encoding an audio input signal to obtain an encoded audio signal is provided. The apparatus comprises a transformation module (110) configured to transform the audio input signal from an original domain to a transform domain to obtain a transformed audio signal. Moreover, the apparatus comprises an encoding module (120), configured to quantize the transformed audio signal to obtain a quantized signal, and configured to encode the quantized signal to obtain the encoded audio signal. The transformation module (110) is configured to transform the audio input signal depending on a plurality of predefined power values of quantization noise in the original domain.