Masking Threshold Filterbank for Low-Bitrate Audio Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding methods fail to efficiently represent acoustic scenes while balancing perceived quality, accuracy, complexity, and computational complexity, as they do not adequately account for human auditory masking effects.
Innovation Solution
A masking threshold determinator using a filterbank with filters of varying bandwidths models human auditory frequency separation, determining masking thresholds based on bandpass signals to optimize quantization and encoding, considering both magnitude and phase information of complex bandpass signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the number of bits allocated for side information in the frequency domain is increased to improve masking threshold determination accuracy, then the coding precision is improved, but the bit rate increases
Solution Approach 1:
The frequency spectrum is divided into multiple frequency bands, and the masking threshold is determined separately for each band. This segmentation allows precise representation of masking characteristics in different frequency regions while limiting the total bit allocation per band, thus resolving the contradiction between accuracy and bit rate.
Solution Approach 2:
Different bit allocation strategies are applied to different frequency bands based on their specific characteristics. Frequency bands requiring higher precision for masking threshold determination are allocated more bits, while less critical bands receive fewer bits, optimizing the overall balance between accuracy and bit rate.
2Measurement precision
If the audio signal is transformed to the frequency domain for masking threshold determination, then the coding precision is improved, but the processing complexity increases
Solution Approach 1:
The frequency domain processing is segmented into discrete frequency bands with standardized transformation procedures. This segmentation simplifies the overall processing complexity by breaking down the complex frequency domain analysis into manageable, repeatable steps for each band.
Solution Approach 2:
The invention uses standardized parameter transformations between time and frequency domains, and within the frequency domain itself. These standardized parameter changes simplify the processing complexity while maintaining the precision benefits of frequency domain analysis.
3Quantity of substance
If a fixed number of bits is allocated for side information, then the bit rate is controlled, but the adaptability to different audio signals is reduced
Solution Approach 1:
The invention dynamically adjusts the number of bits allocated for side information in each frequency band based on the actual audio signal characteristics. This parameter adaptation allows the system to maintain controlled overall bit rate while achieving optimal representation accuracy for different signal types and complexity levels.
Solution Approach 2:
The bit allocation scheme transitions from a fixed approach to a dynamic approach where the number of bits per frequency band is adjusted according to signal characteristics. This dynamic adaptation maintains bit rate control while significantly improving versatility across different audio signals.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Embodiments are related to a masking threshold determinator (100, 200, 340, 1200, 1300, 3200), wherein the masking threshold determinator is configured to obtain a plurality of bandpass signals (111, 211, 311, 1211, 1311, 3211) using a plurality of filters (110, 210, 310, 1210, 1310, 3210) having different bandwidths; and wherein the masking threshold determinator is configured to obtain a masking threshold information associated with a given frequency region on the basis of bandpass signal values of at least two bandpass signals. Furthermore, audio encoders, methods and computer programs are disclosed.