Audio Quantization Using Adaptive Masking Thresholds
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing methods fail to effectively allocate quantization noise due to insufficient bits in low bit rate codecs, leading to perceived distortion, especially in frequency bands with low energy, and do not adequately consider human auditory properties when processing mixed speech and music signals.
Innovation Solution
A method that adjusts the masking threshold based on the energy and sensitivity of quantization noise, using a psychoacoustic model to generate a modified masking threshold by applying weightings to audio signal bands, allowing for efficient quantization without additional bits, and improving sound quality by accounting for human auditory properties.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a low bit rate codec is used, then bit rate is reduced, but quantization noise becomes audible and perceived distortion increases
Solution Approach 1:
The patent applies different masking thresholds to different frequency bands based on local energy characteristics. By analyzing the energy distribution across frequency bands and applying band-specific masking thresholds, the system optimizes quantization noise allocation locally in each band rather than using a uniform threshold, thereby reducing audible quantization noise while maintaining low bit rate
Solution Approach 2:
The patent dynamically adjusts the masking threshold parameter based on the energy characteristics of the audio signal in different frequency bands. By changing the masking threshold parameter according to local energy levels and signal type (speech/music), the system adapts the quantization noise allocation to minimize perceived distortion at low bit rates
2Device complexity
If a psychoacoustic model is applied to mixed speech and music signals, then general audio processing is simplified, but speech signal quality deteriorates due to inadequate consideration of speech-specific auditory properties
Solution Approach 1:
The patent dynamically adapts the psychoacoustic model parameters based on the detected signal type (speech or music). By identifying whether the current frame contains speech or music signals and adjusting the masking threshold calculation accordingly, the system maintains simplicity for general audio while optimizing speech-specific characteristics when needed
Solution Approach 2:
The patent applies different processing characteristics to speech and music portions of the signal. By detecting speech-specific features and applying speech-optimized masking threshold calculations only to speech segments, the system preserves speech quality without significantly increasing overall processing complexity
Data Source
AI summary
A method for processing an audio signal is disclosed. The method for processing an audio signal includes frequency-transforming an audio signal to generate a frequency-spectrum, deciding a weighting per band corresponding to energy per band using the frequency spectrum, receiving a masking threshold based on a psychoacoustic model, applying the weighting to the masking threshold to generate a modified masking threshold, and quantizing the audio signal using the modified masking threshold.


