Audio Coding Using Frequency and Time Domain Processors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding methods, particularly at low bitrates, suffer from reduced audio quality due to band limitation and inefficiencies in encoding high-frequency harmonics, leading to inaccuracies in reproducing signals with prominent high-frequency content.
Innovation Solution
A combined time and frequency domain encoding/decoding processor with intelligent gap filling (IGF) that operates in a single domain, allowing full-band encoding and decoding, where high-resolution encoding is applied to tonal portions and low-resolution parametric encoding to noisy components, with gap filling using frequency tiles and energy information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If bandwidth extension is used to reduce bitrate, then bitrate efficiency is improved, but audio quality deteriorates due to band limitation
Solution Approach 1:
The audio spectrum is segmented into multiple bands (first frequency band up to crossover frequency, second frequency band above crossover frequency). Different encoding strategies are applied to each segment: waveform coding for the first band and parametric coding for the second band, allowing optimized treatment of different frequency regions
Solution Approach 2:
Different encoding resolutions are applied to different frequency bands. The first band (up to crossover) uses high-resolution waveform coding while the second band (above crossover) uses low-resolution parametric coding. This local differentiation allows bitrate reduction in less critical high-frequency regions while preserving quality in the more important low-frequency range
2Loss of energy
If frequency domain encoding with bandwidth extension is used, then bitrate efficiency is improved, but encoding accuracy deteriorates for high-frequency harmonics
Solution Approach 1:
The encoding approach changes from waveform coding to parametric coding at the crossover frequency boundary. Parametric coding uses fewer parameters to represent the signal, reducing bitrate but also reducing accuracy for capturing fine spectral details like high-frequency harmonics
3Productivity
If time domain encoding is used for speech signals, then encoding efficiency is improved, but adaptability to non-speech signals deteriorates
Solution Approach 1:
The system dynamically switches between time domain encoding (for speech) and frequency domain encoding with parametric coding (for non-speech). The encoder adapts its operating mode based on the signal characteristics, providing both efficiency for speech and versatility for different signal types
4Loss of energy
If bandwidth extension operates above crossover frequency, then bitrate efficiency is improved, but frequency range coverage deteriorates
Solution Approach 1:
The frequency range is segmented into two parts: the first band (up to crossover) is fully covered by waveform coding, while the second band (above crossover) is parametrically coded. This segmentation allows the system to claim full-bandwidth operation while using bandwidth extension techniques in the upper frequencies
Data Source
Figure 1A~1B
Figure 2A
Figure 2B
AI summary
An audio encoder for encoding an audio signal, comprises: a first encoding processor (600) for encoding a first audio signal portion in a frequency domain, wherein the first encoding processor (600) comprises: a time frequency converter (602) for converting the first audio signal portion into a frequency domain representation having spectral lines up to a maximum frequency of the first audio signal portion; an analyzer (604) for analyzing the frequency domain representation up to the maximum frequency to determine first spectral portions to be encoded with a first spectral resolution and second spectral regions to be encoded with a second spectral resolution, the second spectral resolution being lower than the first spectral resolution; a spectral encoder (606) for encoding the first spectral portions with the first spectral resolution and for encoding the second spectral portions with the second spectral resolution; a second encoding processor (610) for encoding a second different audio signal portion in the time domain; a controller (620) configured for analyzing the audio signal and for determining, which portion of the audio signal is the first audio signal portion encoded in the frequency domain and which portion of the audio signal is the second audio signal portion encoded in the time domain; and an encoded signal former (630) for forming an encoded audio signal comprising a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion.