Audio Decoder Bandwidth Selection Spectral Leakage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio decoding technologies face challenges in maintaining audio quality due to spectral energy leakage, particularly when narrowband content is encoded and decoded using wideband coders, leading to degraded quality exacerbated by non-linear power amplification or dynamic range compression.
Innovation Solution
A decoder system that classifies audio frames as either wideband or band-limited, selectively removing spectral energy leakage from the high band to output band-limited content, using different thresholds to avoid frequent mode switching and ensure optimal quality by transitioning between wideband and band-limited modes based on consecutive frame classifications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If narrowband content is encoded and decoded using a wideband coder, then the decoder can handle both narrowband and wideband content, but spectral energy leakage occurs in frequency bands above the original bandwidth, degrading audio quality
Solution Approach 1:
The frequency spectrum is divided into multiple bands (e.g., low band 0-4kHz and high band 4-8kHz). The decoder processes each band separately, applying different output modes to different frequency segments. This allows the decoder to handle both narrowband and wideband content while preventing spectral energy leakage from affecting the overall audio quality.
Solution Approach 2:
Different parts of the frequency spectrum are treated with different quality requirements. The low band (0-4kHz) receives full processing while the high band (4-8kHz) has its spectral energy leakage removed or attenuated. This local differentiation ensures that each frequency band contributes optimally to the overall audio quality without introducing harmful artifacts.
2Object-affected harmful factors
If the decoder always outputs band-limited content, then audio quality for narrowband content is maintained, but wideband content quality is degraded due to removal of valid high-frequency information
Solution Approach 1:
The decoder dynamically adjusts its output mode based on the characteristics of each audio frame. By analyzing energy distribution across frequency bands and classifying frames as narrowband or wideband, the decoder adapts its processing in real-time. This allows optimal quality for narrowband content while preserving valid high-frequency information in wideband content.
Solution Approach 2:
The decoder uses feedback from spectral analysis to control its output mode. By continuously monitoring energy distribution in different frequency bands and comparing against thresholds, the decoder determines whether to apply band-limited output mode or wideband output mode, ensuring optimal quality for the current audio content type.
3Object-affected harmful factors
If the decoder frequently switches between wideband and band-limited modes, then optimal quality can be maintained for varying content types, but mode switching causes instability and potential quality degradation
Solution Approach 1:
The decoder performs preliminary classification of audio frames as narrowband or wideband before applying the corresponding output mode. By预先 determining the content type and setting appropriate mode transition thresholds, the decoder avoids frequent or unnecessary mode switching, ensuring stability while maintaining optimal quality for varying content types.
4Power
If non-linear power amplification or dynamic range compression is applied to narrowband content, then audio output level is adjusted, but degraded audio quality is magnified
Solution Approach 1:
The decoder removes spectral energy leakage from narrowband content before the signal passes through non-linear power amplification or dynamic range compression. By eliminating the source of degradation upfront, the subsequent power amplification and compression operations can adjust audio output levels without magnifying quality degradation, as the harmful spectral components have already been removed.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A device includes a receiver configured to receive an audio frame of an audio stream. The device also includes a decoder configured to generate first decoded speech associated with the audio frame and to determine a count of audio frames classified as being associated with band limited content. The decoder is further configured to output second decoded speech based on the first decoded speech. The second decoded speech may be generated according to an output mode of the decoder. The output mode may be selected based at least in part on the count of audio frames.