Hybrid Audio Encoding for Low Bitrate Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional E-AC-3 encoding methods degrade in quality when bitrates drop below 192 kbps for multichannel audio signals, leading to significant spatial collapse and increased coding artifacts, making it difficult to maintain 'broadcast quality' at lower bitrates.
Innovation Solution
A hybrid encoding method that downmixes only the low-frequency components of a multichannel audio signal and applies waveform coding, while using parametric encoding for the remaining frequency components, thereby reducing coding noise and preserving spatial information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional E-AC-3 encoding is used at bitrates below 192 kbps, then bit rate reduction is achieved, but audio quality degrades with significant spatial collapse and increased coding artifacts
Solution Approach 1:
The audio signal is segmented into low-frequency components (below 3.5 kHz) and high-frequency components (3.5 kHz and above). The low-frequency components undergo waveform coding while the high-frequency components undergo parametric coding. This segmentation allows different coding strategies to be applied to different frequency ranges, maintaining audio quality at reduced bitrates by preventing spatial collapse in the low-frequency domain while efficiently coding the high-frequency domain.
2Device complexity
If waveform coding is applied to all frequency components, then coding simplicity is maintained, but coding efficiency decreases and spatial information is lost
Solution Approach 1:
Different coding methods are applied to different frequency components based on their local characteristics. The low-frequency components (below 3.5 kHz) use waveform coding to preserve spatial information and prevent collapse, while the high-frequency components (3.5 kHz and above) use parametric coding for efficiency. This local quality approach optimizes coding efficiency for each frequency range while maintaining overall coding simplicity through a clear division of responsibilities.
3Productivity
If parametric coding is used for all frequency components, then coding efficiency is improved, but spatial information is lost and quality degrades at lower bitrates
Solution Approach 1:
The audio spectrum is segmented into low-frequency and high-frequency bands. Waveform coding is applied to the low-frequency components to preserve spatial information and prevent collapse, while parametric coding is applied to the high-frequency components for efficiency. This segmentation resolves the contradiction by assigning the spatial information preservation function to the appropriate frequency range where it is most critical.
Solution Approach 2:
The coding method is adapted to the local requirements of different frequency components. Low-frequency components require waveform coding to maintain spatial information and prevent collapse, while high-frequency components can use parametric coding for efficiency. This local quality approach ensures that spatial information is preserved where it matters most while maximizing overall coding efficiency.
Data Source
Figure 1~4
Figure 2
Figure 3
AI summary
A method for encoding a multichannel audio input signal, including steps of generating a downmix of low frequency components of a subset of channels of the input signal, waveform coding each channel of the downmix, thereby generating waveform coded, downmixed data, performing parametric encoding on at least some higher frequency components of each channel of the input signal, thereby generating parametrically coded data, and generating an encoded audio signal (e.g., an E-AC-3 encoded signal) indicative of the waveform coded, downmixed data and the parametrically coded data. Other aspects are methods for decoding such an encoded signal, and systems configured to perform any embodiment of the inventive method.