Audio Encoding via Subband Dimensionality Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio encoding technologies face challenges in improving encoding efficiency while maintaining audio quality, particularly in remote audio/video calls, where existing methods do not effectively balance data compression and signal fidelity.
Innovation Solution
The proposed solution involves decomposing audio signals into low-frequency and high-frequency subband signals, where the low-frequency subband retains higher feature dimensionality for quality and the high-frequency subband has lower dimensionality for efficient encoding, using quantization encoding on both to generate bitstreams, and employing neural networks for feature extraction and reconstruction to optimize data compression.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If audio signals are encoded using traditional uniform compression methods, then encoding simplicity is maintained, but audio quality deteriorates due to insufficient differentiation between frequency bands
Solution Approach 1:
The audio signal is segmented into multiple frequency subbands (low-frequency and high-frequency bands) using filter banks. Each subband is then processed independently with tailored encoding strategies, allowing differentiated compression approaches that preserve quality where needed while reducing complexity where tolerable.
Solution Approach 2:
Different encoding precision levels are applied to different frequency subbands based on their perceptual importance. Low-frequency subbands receive higher precision encoding to preserve fundamental audio quality, while high-frequency subbands use lower precision encoding where human perception is less sensitive, optimizing the quality-complexity tradeoff.
2Productivity
If feature dimensionality is reduced for high-frequency subbands, then encoding efficiency improves, but information loss increases
Solution Approach 1:
Instead of uniformly reducing feature dimensionality across all subbands, the method applies partial dimensionality reduction only to high-frequency subbands where it is less critical. Low-frequency subbands maintain full or higher dimensionality to preserve essential audio information, achieving efficiency gains without excessive information loss.
Solution Approach 2:
The feature dimensionality parameter is changed differently for different frequency subbands. High-frequency subbands use reduced dimensionality features to improve encoding efficiency, while low-frequency subbands maintain higher dimensionality to preserve audio quality and minimize information loss.
3Quantity of substance
If data compression is increased for remote communication, then bandwidth usage decreases, but audio fidelity deteriorates
Solution Approach 1:
The audio data is segmented into frequency subbands that are independently compressed. This allows different compression ratios to be applied to different subbands, reducing overall data amount while preserving fidelity in critical low-frequency regions where human hearing is most sensitive.
Solution Approach 2:
Different compression levels are applied locally to different frequency subbands. Low-frequency subbands maintain higher quality with less compression to preserve audio fidelity, while high-frequency subbands accept higher compression to reduce data amount, optimizing the overall quality-quantity tradeoff.
Data Source
AI summary
An audio processing method and apparatus, including decomposing an audio signal into a low-frequency subband signal and a high-frequency subband signal, obtaining a low-frequency feature of the low-frequency subband signal, obtaining a high-frequency feature of the high-frequency subband signal, feature dimensionality of the high-frequency feature being lower than feature dimensionality of the low-frequency feature, performing quantization encoding on the low-frequency feature to obtain a low-frequency bitstream of the audio signal, and performing quantization encoding on the high-frequency feature to obtain a high-frequency bitstream of the audio signal.


